mirror of
https://github.com/outbackdingo/optimclaw.git
synced 2026-08-31 08:39:24 +00:00
* refactor: extract shared assertion helpers to support/assertions.rs Move 5 assertion helpers from e2e_spot_checks.rs to a shared module. Add assert_all_tools_succeeded and assert_tool_succeeded for eliminating false positives in E2E tests. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add tool output capture via tool_results() accessor Extract (name, preview) from ToolResult status events in TestChannel and TestRig, enabling content assertions on tool outputs. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: correct tool parameters in 3 broken trace fixtures - tool_time.json: add missing "operation": "now" for time tool - robust_correct_tool.json: same fix - memory_full_cycle.json: change "path" to "target" for memory_write Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: add tool success and output assertions to eliminate false positives Every E2E test that exercises tools now calls assert_all_tools_succeeded. Added tool output content assertions where tool results are predictable (time year, read_file content, memory_read content). Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: capture per-tool timing from ToolStarted/ToolCompleted events Record Instant on ToolStarted and compute elapsed duration on ToolCompleted, wiring real timing data into collect_metrics() instead of hardcoded zeros. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor: add RAII CleanupGuard for temp file/dir cleanup in tests Replace manual cleanup_test_dir() calls and inline remove_file() with Drop-based CleanupGuard that ensures cleanup even if a test panics. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: add Drop impl and graceful shutdown for TestRig Wrap agent_handle in Option so Drop can abort leaked tasks. Signal the channel shutdown before aborting for future cooperative shutdown. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: replace agent startup sleep with oneshot ready signal Use a oneshot channel fired in Channel::start() instead of a fixed 100ms sleep, eliminating the race condition on slow systems. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: replace fragile string-matching iteration limit with count-based detection Use tool completion count vs max_tool_iterations instead of scanning status messages for "iteration"/"limit" substrings. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: use assert_all_tools_succeeded for memory_full_cycle test Remove incorrect comment about memory_tree failing with empty path (it actually succeeds). Omit empty path from fixture and use the standard assert_all_tools_succeeded instead of per-tool assertions. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor: promote benchmark metrics types to library code Move TraceMetrics, ScenarioResult, RunResult, MetricDelta, and compare_runs() from tests/support/metrics.rs to src/benchmark/metrics.rs. Existing tests use re-export for backward compatibility. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add Scenario and Criterion types for agent benchmarking Scenario defines a task with input, success criteria, and resource limits. Criterion is an enum of programmatic checks (tool_used, response_contains, etc.) evaluated without LLM judgment. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add initial benchmark scenario suite (12 scenarios across 5 categories) Scenarios cover tool_selection, tool_chaining, error_recovery, efficiency, and memory_operations. All loaded from JSON with deserialization validation test. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add benchmark runner with BenchChannel and InstrumentedLlm BenchChannel is a minimal Channel implementation for benchmarks. InstrumentedLlm wraps any LlmProvider to capture per-call metrics. Runner creates a fresh agent per scenario, evaluates success criteria, and produces RunResult with timing, token, and cost metrics. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add baseline management, reports, and benchmark entry point - baseline.rs: load/save/promote benchmark results - report.rs: format comparison reports with regression detection - benchmark_runner.rs: integration test with real LLM (feature-gated) - Add benchmark feature flag to Cargo.toml Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: apply cargo fmt to benchmark module Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add multi-turn scenario types with setup, judge, ResponseNotContains Add BenchScenario, Turn, TurnAssertions, JudgeConfig, ScenarioSetup, WorkspaceSetup, SeedDocument types for multi-turn benchmark scenarios. Add ResponseNotContains criterion variant. Add TurnAssertions::to_criteria() converter for backward compat with existing evaluation engine. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add JSON scenario loader with recursive discovery and tag filter Add load_bench_scenarios() for the new BenchScenario format with recursive directory traversal and tag-based filtering. Create 4 initial trajectory scenarios across tool-selection, multi-turn, and efficiency categories. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): multi-turn runner with workspace seeding and per-turn metrics Add run_bench_scenario() that loops over BenchScenario turns, seeds workspace documents, collects per-turn metrics (tokens, tool calls, wall time), and evaluates per-turn assertions. Add TurnMetrics to metrics.rs and clear_for_next_turn() to BenchChannel. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add LLM-as-judge scoring with prompt formatting and score parsing Create judge.rs with format_judge_prompt, parse_judge_score, and judge_turn. Wire into run_bench_scenario for turns with judge config -- scores below min_score fail the turn. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add CLI subcommand (ironclaw benchmark) Add BenchmarkCommand with --tags, --scenario, --no-judge, --timeout, --update-baseline flags. Wire into Command enum and main.rs dispatch. Feature-gated behind benchmark flag. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): per-scenario JSON output with full trajectory Add save_scenario_results() that writes per-scenario JSON files alongside the run summary. Each scenario gets its own file with turn_metrics trajectory. Update CLI to use new output format. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add ToolRegistry::retain_only and wire tool filtering in scenarios Add a retain_only() method to ToolRegistry that filters tools down to a given allowlist. Wire this into run_bench_scenario() so that when a scenario specifies a tools list in its setup, only those tools are available during the benchmark run. Includes two tests for the new method: one verifying filtering works and one verifying empty input is a no-op. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): wire identity overrides into workspace before agent start Add seed_identity() helper that writes identity files (IDENTITY.md, USER.md, etc.) into the workspace before the agent starts, so that workspace.system_prompt() picks them up. Wire it into run_bench_scenario() after workspace seeding. Include a test that verifies identity files are written and readable. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add --parallel and --max-cost CLI flags Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(benchmark): use feature-conditional snapshot names for CLI help tests Prevents snapshot conflicts between default (no benchmark) and all-features (with benchmark) builds by using separate snapshot names per feature set. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): parallel execution with JoinSet and budget cap enforcement Replace sequential loop in run_all_bench() with parallel execution using JoinSet + semaphore when config.parallel > 1. Add budget cap enforcement that skips remaining scenarios when max_total_cost_usd is exceeded. Track skipped count in RunResult.skipped_scenarios and display it in format_report(). Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add tool restriction and identity override test scenarios Co-Authored-By: Claude Opus 4.6 <[email protected]> * chore: fix formatting for Phase 3 Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add SkillRegistry::retain_only and wire skill filtering in scenarios Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add --json flag for machine-readable output Co-Authored-By: Claude Opus 4.6 <[email protected]> * ci: add GitHub Actions benchmark workflow (manual trigger) Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(benchmark): remove in-tree benchmark harness, keep retain_only utilities Move benchmark-specific code out of ironclaw in preparation for the nearai/benchmarks trajectory adapter. This removes: - src/benchmark/ (runner, scenarios, metrics, judge, report, etc.) - src/cli/benchmark.rs and the Benchmark CLI subcommand - benchmarks/ data directory (scenarios + trajectories) - .github/workflows/benchmark.yml - The "benchmark" Cargo feature flag What remains: - ToolRegistry::retain_only() and SkillRegistry::retain_only() - Test support types (TraceMetrics, InstrumentedLlm) inlined into tests/support/ instead of re-exporting from the deleted module Co-Authored-By: Claude Opus 4.6 <[email protected]> * docs: add README for LLM trace fixture format Documents the trajectory JSON format, response types, request hints, directory structure, and how to write new traces. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(test): unify trace format around turns, add multi-turn support Introduce TraceTurn type that groups user_input with LLM response steps, making traces self-contained conversation trajectories. Add run_trace() to TestRig for automatic multi-turn replay. Backward-compatible: flat "steps" JSON is deserialized as a single turn transparently. Includes all trace fixtures (spot, coverage, advanced), plan docs, and new e2e tests for steering, error recovery, long chains, memory, and prompt injection resilience. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): fix CI failures after merging main - Fix tool_json fixture: use "data" parameter (not "input") to match JsonTool schema - Fix status_events test: remove assertion for "time" tool that isn't in the fixture (only "echo" calls are used) - Allow dead_code in test support metrics/instrumented_llm modules (utilities for future benchmark tests) [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * Working on recording traces and testing them * feat(test): add declarative expects to trace fixtures, split infra tests Add TraceExpects struct with 9 optional assertion fields (response_contains, tools_used, all_tools_succeeded, etc.) that can be declared in fixture JSON instead of hand-written Rust. Add verify_expects() and run_recorded_trace() so recorded trace tests become one-liners. Split trace infra tests (deserialization, backward compat) into tests/trace_format.rs which doesn't require the libsql feature gate. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(test): add expects to all trace fixtures, simplify e2e tests Add declarative expects blocks to all 19 trace fixture JSONs across spot/, coverage/, advanced/, and root directories. Update all 8 e2e test files to use verify_trace_expects() / run_and_verify_trace(), replacing ~270 lines of hand-written assertions with fixture-driven verification. Tests that check things beyond expects (file content on disk, metrics, event ordering) keep those extra assertions alongside the declarative ones. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): adapt tests to AppBuilder refactor, fix formatting Update test files to work with refactored TestRigBuilder that uses AppBuilder::build_all() (removing with_tools/with_workspace methods). Update telegram_check fixture to use tool_list instead of echo. Fix cargo fmt issues in src/llm/mod.rs and src/llm/recording.rs. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(test): deduplicate support unit tests into single binary Support modules (assertions, cleanup, test_channel, test_rig, trace_llm) had #[cfg(test)] mod tests blocks that were compiled and run 12 times — once per e2e test binary that declares `mod support;`. Extracted all 29 support unit tests into a dedicated `tests/support_unit_tests.rs` so they run exactly once. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: fix trailing newlines in support files Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(test): unify trace types and fix recorded multi-turn replay Import shared types (TraceStep, TraceResponse, TraceToolCall, RequestHint, ExpectedToolResult, MemorySnapshotEntry, HttpExchange*) from ironclaw::llm::recording instead of redefining them in trace_llm.rs. Fix the flat-steps deserializer to split at UserInput boundaries into multiple turns, instead of filtering them out and wrapping everything into a single turn. This enables recorded multi-turn traces to be replayed as proper multi-turn conversations via run_trace(). [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): fix CI failures - unused imports and missing struct fields - Add #[allow(unused_imports)] on pub use re-exports in trace_llm.rs (types are re-exported for downstream test files, not used locally) - Add `..` to ToolCompleted pattern in test_channel.rs to match new `error` and `parameters` fields Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): fix CI failures after merging main - Add missing `error` and `parameters` fields to ToolCompleted constructors in support_unit_tests.rs - Add `..` to ToolCompleted pattern match in support_unit_tests.rs - Add #[allow(dead_code)] to CleanupGuard, LlmTrace impl, and TraceLlm impl (only used behind #[cfg(feature = "libsql")]) Co-Authored-By: Claude Opus 4.6 <[email protected]> * Adding coverage running script * fix(test): address review feedback on E2E test infrastructure - Increase wait_for_responses polling to exponential backoff (50ms-500ms) and raise default timeout from 15s to 30s to reduce CI flakiness (#1) - Strengthen prompt_injection_resilience test with positive safety layer assertion via has_safety_warnings(), enable injection_check (#2) - Add assert_tool_order() helper and tools_order field in TraceExpects for verifying tool execution ordering in multi-step traces (#3) - Document TraceLlm sequential-call assumption for concurrency (#6) - Clean up CleanupGuard with PathKind enum instead of shotgun remove_file + remove_dir_all on every path (#8) - Fix coverage.sh: default to --lib only, fix multi-filter syntax, add COV_ALL_TARGETS option - Add coverage/ to .gitignore - Remove planning docs from PR [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address PR review - use HashSet in retain_only, improve skill test - Use HashSet for O(N+M) lookup in SkillRegistry::retain_only and ToolRegistry::retain_only instead of linear scan - Strengthen test_retain_only_empty_is_noop in SkillRegistry to pre-populate with a skill before asserting the no-op behavior [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): revert incorrect safety layer assertion in injection test The safety layer sanitizes tool output, not user input. The injection test sends a malicious user message with no tools called, so the safety layer never fires. Reverted to the original test which correctly validates the LLM refuses via trace expects. Also fixed case-sensitive request hint ("ignore" -> "Ignore") to suppress noisy warning. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: clean stale profdata before coverage run Adds `cargo llvm-cov clean` before each run to prevent "mismatched data" warnings from stale instrumentation profiles. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: fix formatting in retain_only test [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Co-authored-by: Claude Opus 4.6 <[email protected]> Co-authored-by: Illia Polosukhin <[email protected]>
400 lines
13 KiB
Rust
400 lines
13 KiB
Rust
//! Job state machine.
|
|
|
|
use std::collections::HashMap;
|
|
use std::sync::Arc;
|
|
use std::time::Duration;
|
|
|
|
use chrono::{DateTime, Utc};
|
|
use rust_decimal::Decimal;
|
|
use serde::{Deserialize, Serialize};
|
|
use uuid::Uuid;
|
|
|
|
use crate::llm::recording::HttpInterceptor;
|
|
|
|
/// State of a job.
|
|
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Serialize, Deserialize)]
|
|
#[serde(rename_all = "snake_case")]
|
|
pub enum JobState {
|
|
/// Job is waiting to be started.
|
|
Pending,
|
|
/// Job is currently being worked on.
|
|
InProgress,
|
|
/// Job work is complete, awaiting submission.
|
|
Completed,
|
|
/// Job has been submitted for review.
|
|
Submitted,
|
|
/// Job was accepted/paid.
|
|
Accepted,
|
|
/// Job failed and cannot be completed.
|
|
Failed,
|
|
/// Job is stuck and needs repair.
|
|
Stuck,
|
|
/// Job was cancelled.
|
|
Cancelled,
|
|
}
|
|
|
|
impl JobState {
|
|
/// Check if this state allows transitioning to another state.
|
|
pub fn can_transition_to(&self, target: JobState) -> bool {
|
|
use JobState::*;
|
|
|
|
matches!(
|
|
(self, target),
|
|
// From Pending
|
|
(Pending, InProgress) | (Pending, Cancelled) |
|
|
// From InProgress
|
|
(InProgress, Completed) | (InProgress, Failed) |
|
|
(InProgress, Stuck) | (InProgress, Cancelled) |
|
|
// From Completed
|
|
(Completed, Submitted) | (Completed, Failed) |
|
|
// From Submitted
|
|
(Submitted, Accepted) | (Submitted, Failed) |
|
|
// From Stuck (can recover or fail)
|
|
(Stuck, InProgress) | (Stuck, Failed) | (Stuck, Cancelled)
|
|
)
|
|
}
|
|
|
|
/// Check if this is a terminal state.
|
|
pub fn is_terminal(&self) -> bool {
|
|
matches!(self, Self::Accepted | Self::Failed | Self::Cancelled)
|
|
}
|
|
|
|
/// Check if the job is active (not terminal).
|
|
pub fn is_active(&self) -> bool {
|
|
!self.is_terminal()
|
|
}
|
|
}
|
|
|
|
impl std::fmt::Display for JobState {
|
|
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
|
let s = match self {
|
|
Self::Pending => "pending",
|
|
Self::InProgress => "in_progress",
|
|
Self::Completed => "completed",
|
|
Self::Submitted => "submitted",
|
|
Self::Accepted => "accepted",
|
|
Self::Failed => "failed",
|
|
Self::Stuck => "stuck",
|
|
Self::Cancelled => "cancelled",
|
|
};
|
|
write!(f, "{}", s)
|
|
}
|
|
}
|
|
|
|
/// A state transition event.
|
|
#[derive(Debug, Clone, Serialize, Deserialize)]
|
|
pub struct StateTransition {
|
|
/// Previous state.
|
|
pub from: JobState,
|
|
/// New state.
|
|
pub to: JobState,
|
|
/// When the transition occurred.
|
|
pub timestamp: DateTime<Utc>,
|
|
/// Reason for the transition.
|
|
pub reason: Option<String>,
|
|
}
|
|
|
|
/// Context for a running job.
|
|
#[derive(Debug, Clone, Serialize)]
|
|
pub struct JobContext {
|
|
/// Unique job ID.
|
|
pub job_id: Uuid,
|
|
/// Current state.
|
|
pub state: JobState,
|
|
/// User ID that owns this job (for workspace scoping).
|
|
pub user_id: String,
|
|
/// Conversation ID if linked to a conversation.
|
|
pub conversation_id: Option<Uuid>,
|
|
/// Job title.
|
|
pub title: String,
|
|
/// Job description.
|
|
pub description: String,
|
|
/// Job category.
|
|
pub category: Option<String>,
|
|
/// Budget amount (if from marketplace).
|
|
pub budget: Option<Decimal>,
|
|
/// Budget token (e.g., "NEAR", "USD").
|
|
pub budget_token: Option<String>,
|
|
/// Our bid amount.
|
|
pub bid_amount: Option<Decimal>,
|
|
/// Estimated cost to complete.
|
|
pub estimated_cost: Option<Decimal>,
|
|
/// Estimated time to complete.
|
|
pub estimated_duration: Option<Duration>,
|
|
/// Actual cost so far.
|
|
pub actual_cost: Decimal,
|
|
/// Total tokens consumed by LLM calls in this job.
|
|
pub total_tokens_used: u64,
|
|
/// Maximum tokens allowed per job (0 = unlimited).
|
|
pub max_tokens: u64,
|
|
/// When the job was created.
|
|
pub created_at: DateTime<Utc>,
|
|
/// When the job was started.
|
|
pub started_at: Option<DateTime<Utc>>,
|
|
/// When the job was completed.
|
|
pub completed_at: Option<DateTime<Utc>>,
|
|
/// Number of repair attempts.
|
|
pub repair_attempts: u32,
|
|
/// State transition history.
|
|
pub transitions: Vec<StateTransition>,
|
|
/// Metadata.
|
|
pub metadata: serde_json::Value,
|
|
/// Extra environment variables to inject into spawned child processes.
|
|
///
|
|
/// Used by the worker runtime to pass fetched credentials to tools
|
|
/// (e.g., shell commands) without mutating the global process environment
|
|
/// via `std::env::set_var`, which is unsafe in multi-threaded programs.
|
|
///
|
|
/// Wrapped in `Arc` for cheap cloning on every tool invocation.
|
|
#[serde(skip)]
|
|
pub extra_env: Arc<HashMap<String, String>>,
|
|
/// Optional HTTP interceptor for trace recording/replay.
|
|
///
|
|
/// When set, tools that make outgoing HTTP requests should check this
|
|
/// interceptor before sending real requests. During recording, the
|
|
/// interceptor captures request/response pairs. During replay, it
|
|
/// returns pre-recorded responses.
|
|
#[serde(skip)]
|
|
pub http_interceptor: Option<Arc<dyn HttpInterceptor>>,
|
|
}
|
|
|
|
impl JobContext {
|
|
/// Create a new job context.
|
|
pub fn new(title: impl Into<String>, description: impl Into<String>) -> Self {
|
|
Self::with_user("default", title, description)
|
|
}
|
|
|
|
/// Create a new job context with a specific user ID.
|
|
pub fn with_user(
|
|
user_id: impl Into<String>,
|
|
title: impl Into<String>,
|
|
description: impl Into<String>,
|
|
) -> Self {
|
|
Self {
|
|
job_id: Uuid::new_v4(),
|
|
state: JobState::Pending,
|
|
user_id: user_id.into(),
|
|
conversation_id: None,
|
|
title: title.into(),
|
|
description: description.into(),
|
|
category: None,
|
|
budget: None,
|
|
budget_token: None,
|
|
bid_amount: None,
|
|
estimated_cost: None,
|
|
estimated_duration: None,
|
|
actual_cost: Decimal::ZERO,
|
|
total_tokens_used: 0,
|
|
max_tokens: 0,
|
|
created_at: Utc::now(),
|
|
started_at: None,
|
|
completed_at: None,
|
|
repair_attempts: 0,
|
|
transitions: Vec::new(),
|
|
extra_env: Arc::new(HashMap::new()),
|
|
http_interceptor: None,
|
|
metadata: serde_json::Value::Null,
|
|
}
|
|
}
|
|
|
|
/// Transition to a new state.
|
|
pub fn transition_to(
|
|
&mut self,
|
|
new_state: JobState,
|
|
reason: Option<String>,
|
|
) -> Result<(), String> {
|
|
if !self.state.can_transition_to(new_state) {
|
|
return Err(format!(
|
|
"Cannot transition from {} to {}",
|
|
self.state, new_state
|
|
));
|
|
}
|
|
|
|
let transition = StateTransition {
|
|
from: self.state,
|
|
to: new_state,
|
|
timestamp: Utc::now(),
|
|
reason,
|
|
};
|
|
|
|
self.transitions.push(transition);
|
|
|
|
// Cap transition history to prevent unbounded memory growth
|
|
const MAX_TRANSITIONS: usize = 200;
|
|
if self.transitions.len() > MAX_TRANSITIONS {
|
|
let drain_count = self.transitions.len() - MAX_TRANSITIONS;
|
|
self.transitions.drain(..drain_count);
|
|
}
|
|
|
|
self.state = new_state;
|
|
|
|
// Update timestamps
|
|
match new_state {
|
|
JobState::InProgress if self.started_at.is_none() => {
|
|
self.started_at = Some(Utc::now());
|
|
}
|
|
JobState::Completed | JobState::Accepted | JobState::Failed | JobState::Cancelled => {
|
|
self.completed_at = Some(Utc::now());
|
|
}
|
|
_ => {}
|
|
}
|
|
|
|
Ok(())
|
|
}
|
|
|
|
/// Add to the actual cost.
|
|
pub fn add_cost(&mut self, cost: Decimal) {
|
|
self.actual_cost += cost;
|
|
}
|
|
|
|
/// Record token usage from an LLM call. Returns an error string if the
|
|
/// token budget has been exceeded after this addition.
|
|
pub fn add_tokens(&mut self, tokens: u64) -> Result<(), String> {
|
|
self.total_tokens_used += tokens;
|
|
if self.max_tokens > 0 && self.total_tokens_used > self.max_tokens {
|
|
Err(format!(
|
|
"Token budget exceeded: used {} of {} allowed tokens",
|
|
self.total_tokens_used, self.max_tokens
|
|
))
|
|
} else {
|
|
Ok(())
|
|
}
|
|
}
|
|
|
|
/// Check whether the monetary budget has been exceeded.
|
|
pub fn budget_exceeded(&self) -> bool {
|
|
if let Some(ref budget) = self.budget {
|
|
self.actual_cost > *budget
|
|
} else {
|
|
false
|
|
}
|
|
}
|
|
|
|
/// Get the duration since the job started.
|
|
pub fn elapsed(&self) -> Option<Duration> {
|
|
self.started_at.map(|start| {
|
|
let end = self.completed_at.unwrap_or_else(Utc::now);
|
|
let duration = end.signed_duration_since(start);
|
|
Duration::from_secs(duration.num_seconds().max(0) as u64)
|
|
})
|
|
}
|
|
|
|
/// Mark the job as stuck.
|
|
pub fn mark_stuck(&mut self, reason: impl Into<String>) -> Result<(), String> {
|
|
self.transition_to(JobState::Stuck, Some(reason.into()))
|
|
}
|
|
|
|
/// Attempt to recover from stuck state.
|
|
pub fn attempt_recovery(&mut self) -> Result<(), String> {
|
|
if self.state != JobState::Stuck {
|
|
return Err("Job is not stuck".to_string());
|
|
}
|
|
self.repair_attempts += 1;
|
|
self.transition_to(JobState::InProgress, Some("Recovery attempt".to_string()))
|
|
}
|
|
}
|
|
|
|
impl Default for JobContext {
|
|
fn default() -> Self {
|
|
Self::with_user("default", "Untitled", "No description")
|
|
}
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
|
|
#[test]
|
|
fn test_state_transitions() {
|
|
assert!(JobState::Pending.can_transition_to(JobState::InProgress));
|
|
assert!(JobState::InProgress.can_transition_to(JobState::Completed));
|
|
assert!(!JobState::Completed.can_transition_to(JobState::Pending));
|
|
assert!(!JobState::Accepted.can_transition_to(JobState::InProgress));
|
|
}
|
|
|
|
#[test]
|
|
fn test_terminal_states() {
|
|
assert!(JobState::Accepted.is_terminal());
|
|
assert!(JobState::Failed.is_terminal());
|
|
assert!(JobState::Cancelled.is_terminal());
|
|
assert!(!JobState::InProgress.is_terminal());
|
|
}
|
|
|
|
#[test]
|
|
fn test_job_context_transitions() {
|
|
let mut ctx = JobContext::new("Test", "Test job");
|
|
assert_eq!(ctx.state, JobState::Pending);
|
|
|
|
ctx.transition_to(JobState::InProgress, None).unwrap();
|
|
assert_eq!(ctx.state, JobState::InProgress);
|
|
assert!(ctx.started_at.is_some());
|
|
|
|
ctx.transition_to(JobState::Completed, Some("Done".to_string()))
|
|
.unwrap();
|
|
assert_eq!(ctx.state, JobState::Completed);
|
|
}
|
|
|
|
#[test]
|
|
fn test_transition_history_capped() {
|
|
let mut ctx = JobContext::new("Test", "Transition cap test");
|
|
// Cycle through Pending -> InProgress -> Stuck -> InProgress -> Stuck ...
|
|
ctx.transition_to(JobState::InProgress, None).unwrap();
|
|
for i in 0..250 {
|
|
ctx.mark_stuck(format!("stuck {}", i)).unwrap();
|
|
ctx.attempt_recovery().unwrap();
|
|
}
|
|
// 1 initial + 250*2 = 501 transitions, should be capped at 200
|
|
assert!(
|
|
ctx.transitions.len() <= 200,
|
|
"transitions should be capped at 200, got {}",
|
|
ctx.transitions.len()
|
|
);
|
|
}
|
|
|
|
#[test]
|
|
fn test_add_tokens_enforces_budget() {
|
|
let mut ctx = JobContext::new("Test", "Budget test");
|
|
ctx.max_tokens = 1000;
|
|
assert!(ctx.add_tokens(500).is_ok());
|
|
assert_eq!(ctx.total_tokens_used, 500);
|
|
assert!(ctx.add_tokens(600).is_err());
|
|
assert_eq!(ctx.total_tokens_used, 1100); // tokens still recorded
|
|
}
|
|
|
|
#[test]
|
|
fn test_add_tokens_unlimited() {
|
|
let mut ctx = JobContext::new("Test", "No budget");
|
|
// max_tokens = 0 means unlimited
|
|
assert!(ctx.add_tokens(1_000_000).is_ok());
|
|
}
|
|
|
|
#[test]
|
|
fn test_budget_exceeded() {
|
|
let mut ctx = JobContext::new("Test", "Money test");
|
|
ctx.budget = Some(Decimal::new(100, 0)); // $100
|
|
assert!(!ctx.budget_exceeded());
|
|
ctx.add_cost(Decimal::new(50, 0));
|
|
assert!(!ctx.budget_exceeded());
|
|
ctx.add_cost(Decimal::new(60, 0));
|
|
assert!(ctx.budget_exceeded());
|
|
}
|
|
|
|
#[test]
|
|
fn test_budget_exceeded_none() {
|
|
let ctx = JobContext::new("Test", "No budget");
|
|
assert!(!ctx.budget_exceeded()); // No budget = never exceeded
|
|
}
|
|
|
|
#[test]
|
|
fn test_stuck_recovery() {
|
|
let mut ctx = JobContext::new("Test", "Test job");
|
|
ctx.transition_to(JobState::InProgress, None).unwrap();
|
|
ctx.mark_stuck("Timed out").unwrap();
|
|
assert_eq!(ctx.state, JobState::Stuck);
|
|
|
|
ctx.attempt_recovery().unwrap();
|
|
assert_eq!(ctx.state, JobState::InProgress);
|
|
assert_eq!(ctx.repair_attempts, 1);
|
|
}
|
|
}
|