mirror of
https://github.com/outbackdingo/optimclaw.git
synced 2026-08-27 08:00:17 +00:00
feat: structured fallback deliverables for failed/stuck jobs (#236)
* feat: structured fallback deliverables for failed/stuck jobs (#221) When a job fails or gets stuck, build a FallbackDeliverable that captures partial results, action statistics, cost, timing, and repair attempts. This replaces opaque error strings with structured data users can act on. - Add FallbackDeliverable, LastAction, ActionStats types in context/fallback.rs - Store fallback in JobContext.metadata["fallback_deliverable"] on failure - Surface fallback in job_status tool output and SSE job_result events - Update mark_failed() and mark_stuck() in worker to build fallback - 8 unit tests covering zero/mixed actions, truncation, timing, serialization Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address review comments on fallback deliverables - Fix doc comment: "200 chars" -> "200 bytes (UTF-8 safe)" since truncate_str operates on byte length, not character count. - Add code comment documenting that SSE fallback_deliverable is currently always None (forward-compatible infrastructure). Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor: take Option<&FallbackDeliverable> instead of &Option<…> Addresses Gemini review feedback: idiomatic Rust prefers Option<&T> over &Option<T> for borrowed optional values. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: guard against non-object metadata and add fallback test - store_fallback_in_metadata now resets metadata to {} when it's any non-object type (string, array, number), not just null. Prevents panic on index assignment. - Add test_job_status_includes_fallback_deliverable to verify the fallback field is surfaced in job_status tool output. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: use sanitized output in fallback preview + add integration tests Security fix: FallbackDeliverable::build() now uses output_sanitized instead of output_raw, preventing secrets/PII from leaking through the job_status tool and SSE job_result events. Also adds: - test_fallback_uses_sanitized_output: proves raw secrets don't leak - test_store_fallback_in_metadata_roundtrip: full serialize/deserialize - test_store_fallback_handles_non_object_metadata: edge case coverage - test_store_fallback_none_is_noop: None input is safe Addresses serrrfirat review feedback on PR #236. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: harden fallback deliverables against review findings - Truncate failure_reason to 1000 bytes to prevent metadata bloat - Add tracing::warn on fallback serialization failure (was silently discarded) - Fix module/struct docs to cover stuck jobs, remove stale SSE claim - Fix job.rs test to use real FallbackDeliverable field names - Add tests for failure_reason truncation and completed_at=None elapsed time - Fix pre-existing clippy warning in settings.rs (field_reassign_with_default) Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address Copilot review findings on fallback deliverables - Fix output_raw/output_sanitized field swap in ActionRecord::succeed() so sanitized data actually goes into the sanitized field (security) - Return None instead of empty Memory when get_memory fails in build_fallback, with tracing::warn for observability - Replace manual elapsed calculation with ctx.elapsed() which already clamps negative durations Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: resolve rebase conflicts and update tests for parameter swap - Add fallback field to SseEvent::JobResult in job_monitor - Fix type annotation in fallback deliverable test - Update test_action_record_succeed_sets_fields for new parameter order - Use create_job_for_user in test (API changed on main) Co-Authored-By: Claude Opus 4.6 <[email protected]> * chore: trigger CI re-check after rebase * fix: fall back to error message for failed action output_preview When the last action is a failed tool call, output_sanitized is None, leaving output_preview empty. Now falls back to the action's error message so users see what went wrong. [skip-regression-check] * ci: add safety comments to test code for no-panics check The CI no-panics grep check cannot distinguish test code inside src/ files from production code. Add // safety: test annotations to .unwrap(), .expect(), and assert!() calls in #[cfg(test)] modules. * fix: clarify succeed() doc and avoid clone in output_preview - Fix doc comment: output_raw is stored as pretty-printed JSON string, not a raw JSON value - Borrow string slice directly in fallback preview to avoid cloning potentially large sanitized outputs before truncation * refactor: reuse floor_char_boundary in truncate_str Replace hand-rolled UTF-8 boundary logic with existing crate::util::floor_char_boundary to reduce duplication. * fix: rename SSE fallback field to fallback_deliverable for consistency The SSE JobResult field was named `fallback` while everywhere else (metadata key, job_status tool) uses `fallback_deliverable`. Align the SSE wire format to avoid forcing clients to handle two names. --------- Co-authored-by: Claude Opus 4.6 <[email protected]>
This commit is contained in:
co-authored by
Claude Opus 4.6
parent
86ae12747b
commit
65062f3cc0
@@ -0,0 +1,319 @@
|
||||
//! Structured fallback deliverables for failed or stuck jobs.
|
||||
//!
|
||||
//! When a job fails or is detected as stuck, a [`FallbackDeliverable`] captures
|
||||
//! what was accomplished before the failure: partial results, action statistics,
|
||||
//! cost, and timing. This gives users visibility into terminal jobs instead of
|
||||
//! just an error string.
|
||||
//!
|
||||
//! Fallback deliverables are stored in `JobContext.metadata["fallback_deliverable"]`
|
||||
//! and surfaced through the `job_status` tool.
|
||||
|
||||
use serde::{Deserialize, Serialize};
|
||||
|
||||
use crate::context::memory::Memory;
|
||||
use crate::context::state::JobContext;
|
||||
|
||||
/// Structured summary of a failed or stuck job.
|
||||
///
|
||||
/// Stored in `JobContext.metadata["fallback_deliverable"]` when a job fails
|
||||
/// or is marked stuck. Surfaced through the `job_status` tool.
|
||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||
pub struct FallbackDeliverable {
|
||||
/// True if at least one action succeeded before failure.
|
||||
pub partial: bool,
|
||||
/// Why the job failed.
|
||||
pub failure_reason: String,
|
||||
/// Last action taken before failure.
|
||||
pub last_action: Option<LastAction>,
|
||||
/// Aggregate action statistics.
|
||||
pub action_stats: ActionStats,
|
||||
/// Total tokens consumed.
|
||||
pub tokens_used: u64,
|
||||
/// Total cost incurred (decimal as string for JSON safety).
|
||||
pub cost: String,
|
||||
/// Wall-clock elapsed time in seconds.
|
||||
pub elapsed_secs: f64,
|
||||
/// Number of self-repair attempts.
|
||||
pub repair_attempts: u32,
|
||||
}
|
||||
|
||||
/// Summary of the last action taken before failure.
|
||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||
pub struct LastAction {
|
||||
pub tool_name: String,
|
||||
/// Truncated to 200 bytes (UTF-8 safe).
|
||||
pub output_preview: String,
|
||||
pub success: bool,
|
||||
}
|
||||
|
||||
/// Aggregate action counts.
|
||||
#[derive(Debug, Clone, Serialize, Deserialize)]
|
||||
pub struct ActionStats {
|
||||
pub total: u32,
|
||||
pub successful: u32,
|
||||
pub failed: u32,
|
||||
}
|
||||
|
||||
impl FallbackDeliverable {
|
||||
/// Build a fallback deliverable from a job context and its memory.
|
||||
pub fn build(ctx: &JobContext, memory: &Memory, reason: &str) -> Self {
|
||||
let successful = memory.successful_actions() as u32;
|
||||
let failed = memory.failed_actions() as u32;
|
||||
let total = memory.actions.len() as u32;
|
||||
|
||||
let last_action = memory.last_action().map(|a| {
|
||||
// Use sanitized output to avoid leaking secrets through the fallback API surface.
|
||||
// For failed actions (no sanitized output), fall back to the error message.
|
||||
// Borrow the string slice directly when possible to avoid cloning
|
||||
// potentially large outputs just for truncation.
|
||||
let owned_fallback;
|
||||
let preview_str: &str = if let Some(v) = a.output_sanitized.as_ref() {
|
||||
match v {
|
||||
serde_json::Value::String(s) => s.as_str(),
|
||||
other => {
|
||||
owned_fallback = serde_json::to_string(other).unwrap_or_default();
|
||||
&owned_fallback
|
||||
}
|
||||
}
|
||||
} else if let Some(ref err) = a.error {
|
||||
err.as_str()
|
||||
} else {
|
||||
""
|
||||
};
|
||||
let preview = truncate_str(preview_str, 200);
|
||||
LastAction {
|
||||
tool_name: a.tool_name.clone(),
|
||||
output_preview: preview.to_string(),
|
||||
success: a.success,
|
||||
}
|
||||
});
|
||||
|
||||
let elapsed_secs = ctx.elapsed().map_or(0.0, |d| d.as_secs_f64());
|
||||
|
||||
Self {
|
||||
partial: successful > 0,
|
||||
failure_reason: truncate_str(reason, 1000).to_string(),
|
||||
last_action,
|
||||
action_stats: ActionStats {
|
||||
total,
|
||||
successful,
|
||||
failed,
|
||||
},
|
||||
tokens_used: ctx.total_tokens_used,
|
||||
cost: ctx.actual_cost.to_string(),
|
||||
elapsed_secs,
|
||||
repair_attempts: ctx.repair_attempts,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Truncate a string to at most `max_len` bytes on a char boundary.
|
||||
fn truncate_str(s: &str, max_len: usize) -> &str {
|
||||
&s[..crate::util::floor_char_boundary(s, max_len)]
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::context::memory::Memory;
|
||||
use crate::context::state::JobContext;
|
||||
use chrono::{Duration, Utc};
|
||||
use rust_decimal::Decimal;
|
||||
use std::time::Duration as StdDuration;
|
||||
|
||||
#[test]
|
||||
fn test_fallback_zero_actions() {
|
||||
let ctx = JobContext::new("Test", "Empty job");
|
||||
let memory = Memory::new(ctx.job_id);
|
||||
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, "timed out");
|
||||
|
||||
assert!(!fb.partial); // safety: test
|
||||
assert_eq!(fb.failure_reason, "timed out"); // safety: test
|
||||
assert!(fb.last_action.is_none()); // safety: test
|
||||
assert_eq!(fb.action_stats.total, 0); // safety: test
|
||||
assert_eq!(fb.action_stats.successful, 0); // safety: test
|
||||
assert_eq!(fb.action_stats.failed, 0); // safety: test
|
||||
assert_eq!(fb.tokens_used, 0); // safety: test
|
||||
assert_eq!(fb.cost, "0"); // safety: test
|
||||
assert_eq!(fb.repair_attempts, 0); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_fallback_mixed_actions() {
|
||||
let mut ctx = JobContext::new("Test", "Mixed job");
|
||||
ctx.total_tokens_used = 5000;
|
||||
ctx.actual_cost = Decimal::new(42, 2); // 0.42
|
||||
ctx.repair_attempts = 1;
|
||||
|
||||
let mut memory = Memory::new(ctx.job_id);
|
||||
|
||||
// 3 successes
|
||||
for _ in 0..3 {
|
||||
let action = memory
|
||||
.create_action("tool_a", serde_json::json!({}))
|
||||
.succeed(
|
||||
Some("output".to_string()),
|
||||
serde_json::json!({}),
|
||||
StdDuration::from_secs(1),
|
||||
);
|
||||
memory.record_action(action);
|
||||
}
|
||||
// 2 failures
|
||||
for _ in 0..2 {
|
||||
let action = memory
|
||||
.create_action("tool_b", serde_json::json!({}))
|
||||
.fail("broke", StdDuration::from_secs(1));
|
||||
memory.record_action(action);
|
||||
}
|
||||
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, "max iterations");
|
||||
|
||||
assert!(fb.partial); // safety: test
|
||||
assert_eq!(fb.action_stats.total, 5); // safety: test
|
||||
assert_eq!(fb.action_stats.successful, 3); // safety: test
|
||||
assert_eq!(fb.action_stats.failed, 2); // safety: test
|
||||
assert_eq!(fb.tokens_used, 5000); // safety: test
|
||||
assert_eq!(fb.cost, "0.42"); // safety: test
|
||||
assert_eq!(fb.repair_attempts, 1); // safety: test
|
||||
assert!(fb.last_action.is_some()); // safety: test
|
||||
let la = fb.last_action.unwrap(); // safety: test
|
||||
assert_eq!(la.tool_name, "tool_b"); // safety: test
|
||||
assert!(!la.success); // safety: test
|
||||
// Failed actions should surface the error message as the output preview
|
||||
assert_eq!(la.output_preview, "broke"); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_fallback_failed_action_shows_error() {
|
||||
let ctx = JobContext::new("Test", "Error preview");
|
||||
let mut memory = Memory::new(ctx.job_id);
|
||||
|
||||
let action = memory
|
||||
.create_action("broken_tool", serde_json::json!({}))
|
||||
.fail("connection timed out after 30s", StdDuration::from_secs(30));
|
||||
memory.record_action(action);
|
||||
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, "tool failure");
|
||||
let la = fb.last_action.unwrap(); // safety: test
|
||||
assert!(!la.success); // safety: test
|
||||
assert_eq!(la.output_preview, "connection timed out after 30s"); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_fallback_last_action_truncation() {
|
||||
let ctx = JobContext::new("Test", "Truncation");
|
||||
let mut memory = Memory::new(ctx.job_id);
|
||||
|
||||
let long_output = "x".repeat(500);
|
||||
let action = memory
|
||||
.create_action("tool_c", serde_json::json!({}))
|
||||
.succeed(
|
||||
Some(long_output.clone()),
|
||||
serde_json::Value::String(long_output),
|
||||
StdDuration::from_secs(1),
|
||||
);
|
||||
memory.record_action(action);
|
||||
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, "failed");
|
||||
let la = fb.last_action.unwrap(); // safety: test
|
||||
assert!(la.output_preview.len() <= 200); // safety: test
|
||||
assert!(!la.output_preview.is_empty()); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_fallback_uses_sanitized_output() {
|
||||
let ctx = JobContext::new("Test", "Sanitized");
|
||||
let mut memory = Memory::new(ctx.job_id);
|
||||
|
||||
let action = memory
|
||||
.create_action("tool_d", serde_json::json!({}))
|
||||
.succeed(
|
||||
Some("[REDACTED]".to_string()),
|
||||
serde_json::json!({"api_key": "sk-secret-key-12345"}),
|
||||
StdDuration::from_secs(1),
|
||||
);
|
||||
memory.record_action(action);
|
||||
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, "failed");
|
||||
let la = fb.last_action.unwrap(); // safety: test
|
||||
// Must use sanitized output, not raw
|
||||
assert!(!la.output_preview.contains("sk-secret")); // safety: test
|
||||
assert!(la.output_preview.contains("REDACTED")); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_fallback_elapsed_time() {
|
||||
let mut ctx = JobContext::new("Test", "Timing");
|
||||
let now = Utc::now();
|
||||
ctx.started_at = Some(now - Duration::seconds(10));
|
||||
ctx.completed_at = Some(now);
|
||||
|
||||
let memory = Memory::new(ctx.job_id);
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, "failed");
|
||||
|
||||
// Should be approximately 10 seconds
|
||||
assert!((fb.elapsed_secs - 10.0).abs() < 0.1); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_fallback_no_started_at() {
|
||||
let ctx = JobContext::new("Test", "Never started");
|
||||
let memory = Memory::new(ctx.job_id);
|
||||
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, "failed");
|
||||
assert!((fb.elapsed_secs - 0.0).abs() < 0.001); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_fallback_elapsed_time_no_completed_at() {
|
||||
let mut ctx = JobContext::new("Test", "Still running");
|
||||
ctx.started_at = Some(Utc::now() - Duration::seconds(5));
|
||||
// completed_at is None — should use Utc::now() as fallback
|
||||
|
||||
let memory = Memory::new(ctx.job_id);
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, "stuck");
|
||||
|
||||
// Should be approximately 5 seconds (using now as end time)
|
||||
assert!(fb.elapsed_secs >= 4.0 && fb.elapsed_secs <= 7.0); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_fallback_failure_reason_truncation() {
|
||||
let ctx = JobContext::new("Test", "Long reason");
|
||||
let memory = Memory::new(ctx.job_id);
|
||||
|
||||
let long_reason = "x".repeat(5000);
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, &long_reason);
|
||||
|
||||
assert!(fb.failure_reason.len() <= 1000); // safety: test
|
||||
assert!(!fb.failure_reason.is_empty()); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_truncate_str_ascii() {
|
||||
assert_eq!(truncate_str("hello", 10), "hello"); // safety: test
|
||||
assert_eq!(truncate_str("hello world", 5), "hello"); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_truncate_str_unicode() {
|
||||
// "é" is 2 bytes in UTF-8
|
||||
let s = "café";
|
||||
assert_eq!(truncate_str(s, 10), "café"); // safety: test
|
||||
// Truncating at 4 would split "é", should back up to 3
|
||||
assert_eq!(truncate_str(s, 4), "caf"); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_fallback_serialization() {
|
||||
let ctx = JobContext::new("Test", "Serialize");
|
||||
let memory = Memory::new(ctx.job_id);
|
||||
let fb = FallbackDeliverable::build(&ctx, &memory, "test error");
|
||||
|
||||
// Should serialize to JSON and back without error
|
||||
let json = serde_json::to_value(&fb).unwrap(); // safety: test
|
||||
let deserialized: FallbackDeliverable = serde_json::from_value(json).unwrap(); // safety: test
|
||||
assert_eq!(deserialized.failure_reason, "test error"); // safety: test
|
||||
}
|
||||
}
|
||||
+79
-70
@@ -58,15 +58,19 @@ impl ActionRecord {
|
||||
}
|
||||
|
||||
/// Mark the action as successful.
|
||||
///
|
||||
/// `output_sanitized` is the tool output after safety processing (string).
|
||||
/// `output_raw` is the original tool result (JSON value, stored as a
|
||||
/// pretty-printed JSON string in `ActionRecord.output_raw`).
|
||||
pub fn succeed(
|
||||
mut self,
|
||||
output_raw: Option<String>,
|
||||
output_sanitized: serde_json::Value,
|
||||
output_sanitized: Option<String>,
|
||||
output_raw: serde_json::Value,
|
||||
duration: Duration,
|
||||
) -> Self {
|
||||
self.success = true;
|
||||
self.output_raw = output_raw;
|
||||
self.output_sanitized = Some(output_sanitized);
|
||||
self.output_raw = Some(serde_json::to_string_pretty(&output_raw).unwrap_or_default());
|
||||
self.output_sanitized = output_sanitized.map(serde_json::Value::String);
|
||||
self.duration = duration;
|
||||
self
|
||||
}
|
||||
@@ -248,15 +252,15 @@ mod tests {
|
||||
#[test]
|
||||
fn test_action_record() {
|
||||
let action = ActionRecord::new(0, "test", serde_json::json!({"key": "value"}));
|
||||
assert_eq!(action.sequence, 0);
|
||||
assert!(!action.success);
|
||||
assert_eq!(action.sequence, 0); // safety: test
|
||||
assert!(!action.success); // safety: test
|
||||
|
||||
let action = action.succeed(
|
||||
Some("raw".to_string()),
|
||||
serde_json::json!({"result": "ok"}),
|
||||
Duration::from_millis(100),
|
||||
);
|
||||
assert!(action.success);
|
||||
assert!(action.success); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -267,7 +271,7 @@ mod tests {
|
||||
memory.add(ChatMessage::user("How are you?"));
|
||||
memory.add(ChatMessage::assistant("Good!"));
|
||||
|
||||
assert_eq!(memory.len(), 3); // Oldest removed
|
||||
assert_eq!(memory.len(), 3); // Oldest removed // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -286,9 +290,9 @@ mod tests {
|
||||
.with_cost(Decimal::new(20, 1));
|
||||
memory.record_action(action2);
|
||||
|
||||
assert_eq!(memory.total_cost(), Decimal::new(30, 1));
|
||||
assert_eq!(memory.total_duration(), Duration::from_secs(3));
|
||||
assert_eq!(memory.successful_actions(), 2);
|
||||
assert_eq!(memory.total_cost(), Decimal::new(30, 1)); // safety: test
|
||||
assert_eq!(memory.total_duration(), Duration::from_secs(3)); // safety: test
|
||||
assert_eq!(memory.successful_actions(), 2); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -296,11 +300,11 @@ mod tests {
|
||||
let action = ActionRecord::new(1, "broken_tool", serde_json::json!({"x": 1}));
|
||||
let action = action.fail("something went wrong", Duration::from_millis(50));
|
||||
|
||||
assert!(!action.success);
|
||||
assert_eq!(action.error.as_deref(), Some("something went wrong"));
|
||||
assert_eq!(action.duration, Duration::from_millis(50));
|
||||
assert!(action.output_raw.is_none());
|
||||
assert!(action.output_sanitized.is_none());
|
||||
assert!(!action.success); // safety: test
|
||||
assert_eq!(action.error.as_deref(), Some("something went wrong")); // safety: test
|
||||
assert_eq!(action.duration, Duration::from_millis(50)); // safety: test
|
||||
assert!(action.output_raw.is_none()); // safety: test
|
||||
assert!(action.output_sanitized.is_none()); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -308,9 +312,9 @@ mod tests {
|
||||
let action = ActionRecord::new(0, "risky_tool", serde_json::json!({}));
|
||||
let action = action.with_warnings(vec!["suspicious pattern".into(), "possible xss".into()]);
|
||||
|
||||
assert_eq!(action.sanitization_warnings.len(), 2);
|
||||
assert_eq!(action.sanitization_warnings[0], "suspicious pattern");
|
||||
assert_eq!(action.sanitization_warnings[1], "possible xss");
|
||||
assert_eq!(action.sanitization_warnings.len(), 2); // safety: test
|
||||
assert_eq!(action.sanitization_warnings[0], "suspicious pattern"); // safety: test
|
||||
assert_eq!(action.sanitization_warnings[1], "possible xss"); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -319,41 +323,46 @@ mod tests {
|
||||
let cost = Decimal::new(42, 2); // 0.42
|
||||
let action = action.with_cost(cost);
|
||||
|
||||
assert_eq!(action.cost, Some(Decimal::new(42, 2)));
|
||||
assert_eq!(action.cost, Some(Decimal::new(42, 2))); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_action_record_new_defaults() {
|
||||
let action = ActionRecord::new(5, "my_tool", serde_json::json!({"key": "val"}));
|
||||
|
||||
assert_eq!(action.sequence, 5);
|
||||
assert_eq!(action.tool_name, "my_tool");
|
||||
assert_eq!(action.input, serde_json::json!({"key": "val"}));
|
||||
assert!(!action.success);
|
||||
assert!(action.output_raw.is_none());
|
||||
assert!(action.output_sanitized.is_none());
|
||||
assert!(action.sanitization_warnings.is_empty());
|
||||
assert!(action.cost.is_none());
|
||||
assert_eq!(action.duration, Duration::ZERO);
|
||||
assert!(action.error.is_none());
|
||||
assert_eq!(action.sequence, 5); // safety: test
|
||||
assert_eq!(action.tool_name, "my_tool"); // safety: test
|
||||
assert_eq!(action.input, serde_json::json!({"key": "val"})); // safety: test
|
||||
assert!(!action.success); // safety: test
|
||||
assert!(action.output_raw.is_none()); // safety: test
|
||||
assert!(action.output_sanitized.is_none()); // safety: test
|
||||
assert!(action.sanitization_warnings.is_empty()); // safety: test
|
||||
assert!(action.cost.is_none()); // safety: test
|
||||
assert_eq!(action.duration, Duration::ZERO); // safety: test
|
||||
assert!(action.error.is_none()); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_action_record_succeed_sets_fields() {
|
||||
let action = ActionRecord::new(0, "tool", serde_json::json!({}));
|
||||
let action = action.succeed(
|
||||
Some("raw output here".into()),
|
||||
Some("sanitized output".into()),
|
||||
serde_json::json!({"clean": true}),
|
||||
Duration::from_secs(7),
|
||||
);
|
||||
|
||||
assert!(action.success);
|
||||
assert_eq!(action.output_raw.as_deref(), Some("raw output here"));
|
||||
assert!(action.success); // safety: test
|
||||
// output_raw is the JSON value pretty-printed
|
||||
let expected_raw =
|
||||
serde_json::to_string_pretty(&serde_json::json!({"clean": true})).unwrap(); // safety: test
|
||||
assert_eq!(action.output_raw.as_deref(), Some(expected_raw.as_str())); // safety: test
|
||||
// output_sanitized wraps the string in a JSON string value
|
||||
assert_eq!(
|
||||
/* safety: test */
|
||||
action.output_sanitized,
|
||||
Some(serde_json::json!({"clean": true}))
|
||||
Some(serde_json::json!("sanitized output"))
|
||||
);
|
||||
assert_eq!(action.duration, Duration::from_secs(7));
|
||||
assert_eq!(action.duration, Duration::from_secs(7)); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -361,13 +370,13 @@ mod tests {
|
||||
let mut mem = ConversationMemory::new(10);
|
||||
mem.add(ChatMessage::user("hello"));
|
||||
mem.add(ChatMessage::assistant("hi"));
|
||||
assert_eq!(mem.len(), 2);
|
||||
assert!(!mem.is_empty());
|
||||
assert_eq!(mem.len(), 2); // safety: test
|
||||
assert!(!mem.is_empty()); // safety: test
|
||||
|
||||
mem.clear();
|
||||
assert_eq!(mem.len(), 0);
|
||||
assert!(mem.is_empty());
|
||||
assert!(mem.messages().is_empty());
|
||||
assert_eq!(mem.len(), 0); // safety: test
|
||||
assert!(mem.is_empty()); // safety: test
|
||||
assert!(mem.messages().is_empty()); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -379,20 +388,20 @@ mod tests {
|
||||
mem.add(ChatMessage::assistant("four"));
|
||||
|
||||
let last_2 = mem.last_n(2);
|
||||
assert_eq!(last_2.len(), 2);
|
||||
assert_eq!(last_2[0].content, "three");
|
||||
assert_eq!(last_2[1].content, "four");
|
||||
assert_eq!(last_2.len(), 2); // safety: test
|
||||
assert_eq!(last_2[0].content, "three"); // safety: test
|
||||
assert_eq!(last_2[1].content, "four"); // safety: test
|
||||
|
||||
// Requesting more than available returns all
|
||||
let last_100 = mem.last_n(100);
|
||||
assert_eq!(last_100.len(), 4);
|
||||
assert_eq!(last_100.len(), 4); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_conversation_memory_last_n_empty() {
|
||||
let mem = ConversationMemory::new(10);
|
||||
let result = mem.last_n(5);
|
||||
assert!(result.is_empty());
|
||||
assert!(result.is_empty()); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -405,13 +414,13 @@ mod tests {
|
||||
// At capacity (3). Adding one more should trim, but keep system.
|
||||
mem.add(ChatMessage::user("msg3"));
|
||||
|
||||
assert_eq!(mem.len(), 3);
|
||||
assert_eq!(mem.len(), 3); // safety: test
|
||||
// System message must survive
|
||||
assert_eq!(mem.messages()[0].role, crate::llm::Role::System);
|
||||
assert_eq!(mem.messages()[0].content, "You are helpful");
|
||||
assert_eq!(mem.messages()[0].role, crate::llm::Role::System); // safety: test
|
||||
assert_eq!(mem.messages()[0].content, "You are helpful"); // safety: test
|
||||
// Oldest non-system message (msg1) should be gone
|
||||
assert_eq!(mem.messages()[1].content, "msg2");
|
||||
assert_eq!(mem.messages()[2].content, "msg3");
|
||||
assert_eq!(mem.messages()[1].content, "msg2"); // safety: test
|
||||
assert_eq!(mem.messages()[2].content, "msg3"); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -422,9 +431,9 @@ mod tests {
|
||||
// Now at capacity. Add another.
|
||||
mem.add(ChatMessage::user("b"));
|
||||
|
||||
assert_eq!(mem.len(), 2);
|
||||
assert_eq!(mem.messages()[0].role, crate::llm::Role::System);
|
||||
assert_eq!(mem.messages()[1].content, "b");
|
||||
assert_eq!(mem.len(), 2); // safety: test
|
||||
assert_eq!(mem.messages()[0].role, crate::llm::Role::System); // safety: test
|
||||
assert_eq!(mem.messages()[1].content, "b"); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -440,7 +449,7 @@ mod tests {
|
||||
mem.add(ChatMessage::user("hello"));
|
||||
// Should have broken out rather than looping forever.
|
||||
// The system message is protected, so len may exceed max.
|
||||
assert!(mem.len() <= 2);
|
||||
assert!(mem.len() <= 2); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -459,14 +468,14 @@ mod tests {
|
||||
.fail("oops", Duration::from_millis(2));
|
||||
memory.record_action(err);
|
||||
|
||||
assert_eq!(memory.successful_actions(), 1);
|
||||
assert_eq!(memory.failed_actions(), 1);
|
||||
assert_eq!(memory.successful_actions(), 1); // safety: test
|
||||
assert_eq!(memory.failed_actions(), 1); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_memory_last_action() {
|
||||
let mut memory = Memory::new(Uuid::new_v4());
|
||||
assert!(memory.last_action().is_none());
|
||||
assert!(memory.last_action().is_none()); // safety: test
|
||||
|
||||
let a1 = memory
|
||||
.create_action("first", serde_json::json!({}))
|
||||
@@ -478,8 +487,8 @@ mod tests {
|
||||
.fail("nope", Duration::ZERO);
|
||||
memory.record_action(a2);
|
||||
|
||||
let last = memory.last_action().unwrap();
|
||||
assert_eq!(last.tool_name, "second");
|
||||
let last = memory.last_action().unwrap(); // safety: test
|
||||
assert_eq!(last.tool_name, "second"); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -499,9 +508,9 @@ mod tests {
|
||||
);
|
||||
memory.record_action(a);
|
||||
|
||||
assert_eq!(memory.actions_by_tool("shell").len(), 3);
|
||||
assert_eq!(memory.actions_by_tool("http").len(), 1);
|
||||
assert_eq!(memory.actions_by_tool("nonexistent").len(), 0);
|
||||
assert_eq!(memory.actions_by_tool("shell").len(), 3); // safety: test
|
||||
assert_eq!(memory.actions_by_tool("http").len(), 1); // safety: test
|
||||
assert_eq!(memory.actions_by_tool("nonexistent").len(), 0); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -509,25 +518,25 @@ mod tests {
|
||||
let mut memory = Memory::new(Uuid::new_v4());
|
||||
|
||||
let a0 = memory.create_action("t", serde_json::json!({}));
|
||||
assert_eq!(a0.sequence, 0);
|
||||
assert_eq!(a0.sequence, 0); // safety: test
|
||||
|
||||
let a1 = memory.create_action("t", serde_json::json!({}));
|
||||
assert_eq!(a1.sequence, 1);
|
||||
assert_eq!(a1.sequence, 1); // safety: test
|
||||
|
||||
let a2 = memory.create_action("t", serde_json::json!({}));
|
||||
assert_eq!(a2.sequence, 2);
|
||||
assert_eq!(a2.sequence, 2); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_memory_add_message_delegates_to_conversation() {
|
||||
let mut memory = Memory::new(Uuid::new_v4());
|
||||
assert!(memory.conversation.is_empty());
|
||||
assert!(memory.conversation.is_empty()); // safety: test
|
||||
|
||||
memory.add_message(ChatMessage::user("hello"));
|
||||
memory.add_message(ChatMessage::assistant("hi"));
|
||||
|
||||
assert_eq!(memory.conversation.len(), 2);
|
||||
assert_eq!(memory.conversation.messages()[0].content, "hello");
|
||||
assert_eq!(memory.conversation.len(), 2); // safety: test
|
||||
assert_eq!(memory.conversation.messages()[0].content, "hello"); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -540,7 +549,7 @@ mod tests {
|
||||
.succeed(None, serde_json::json!({}), Duration::ZERO);
|
||||
memory.record_action(a);
|
||||
|
||||
assert_eq!(memory.total_cost(), Decimal::ZERO);
|
||||
assert_eq!(memory.total_cost(), Decimal::ZERO); // safety: test
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -560,6 +569,6 @@ mod tests {
|
||||
memory.record_action(a2);
|
||||
|
||||
// Both successful and failed actions contribute to total duration
|
||||
assert_eq!(memory.total_duration(), Duration::from_millis(300));
|
||||
assert_eq!(memory.total_duration(), Duration::from_millis(300)); // safety: test
|
||||
}
|
||||
}
|
||||
|
||||
@@ -6,10 +6,12 @@
|
||||
//! - State machine
|
||||
//! - Resource tracking
|
||||
|
||||
pub mod fallback;
|
||||
mod manager;
|
||||
mod memory;
|
||||
mod state;
|
||||
|
||||
pub use fallback::FallbackDeliverable;
|
||||
pub use manager::ContextManager;
|
||||
pub use memory::{ActionRecord, ConversationMemory, Memory};
|
||||
pub use state::{JobContext, JobState, StateTransition, TokenBudgetExceeded};
|
||||
|
||||
Reference in New Issue
Block a user