mirror of
https://github.com/outbackdingo/optimclaw.git
synced 2026-08-26 15:40:18 +00:00
* refactor: extract shared assertion helpers to support/assertions.rs Move 5 assertion helpers from e2e_spot_checks.rs to a shared module. Add assert_all_tools_succeeded and assert_tool_succeeded for eliminating false positives in E2E tests. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add tool output capture via tool_results() accessor Extract (name, preview) from ToolResult status events in TestChannel and TestRig, enabling content assertions on tool outputs. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: correct tool parameters in 3 broken trace fixtures - tool_time.json: add missing "operation": "now" for time tool - robust_correct_tool.json: same fix - memory_full_cycle.json: change "path" to "target" for memory_write Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: add tool success and output assertions to eliminate false positives Every E2E test that exercises tools now calls assert_all_tools_succeeded. Added tool output content assertions where tool results are predictable (time year, read_file content, memory_read content). Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: capture per-tool timing from ToolStarted/ToolCompleted events Record Instant on ToolStarted and compute elapsed duration on ToolCompleted, wiring real timing data into collect_metrics() instead of hardcoded zeros. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor: add RAII CleanupGuard for temp file/dir cleanup in tests Replace manual cleanup_test_dir() calls and inline remove_file() with Drop-based CleanupGuard that ensures cleanup even if a test panics. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: add Drop impl and graceful shutdown for TestRig Wrap agent_handle in Option so Drop can abort leaked tasks. Signal the channel shutdown before aborting for future cooperative shutdown. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: replace agent startup sleep with oneshot ready signal Use a oneshot channel fired in Channel::start() instead of a fixed 100ms sleep, eliminating the race condition on slow systems. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: replace fragile string-matching iteration limit with count-based detection Use tool completion count vs max_tool_iterations instead of scanning status messages for "iteration"/"limit" substrings. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: use assert_all_tools_succeeded for memory_full_cycle test Remove incorrect comment about memory_tree failing with empty path (it actually succeeds). Omit empty path from fixture and use the standard assert_all_tools_succeeded instead of per-tool assertions. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor: promote benchmark metrics types to library code Move TraceMetrics, ScenarioResult, RunResult, MetricDelta, and compare_runs() from tests/support/metrics.rs to src/benchmark/metrics.rs. Existing tests use re-export for backward compatibility. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add Scenario and Criterion types for agent benchmarking Scenario defines a task with input, success criteria, and resource limits. Criterion is an enum of programmatic checks (tool_used, response_contains, etc.) evaluated without LLM judgment. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add initial benchmark scenario suite (12 scenarios across 5 categories) Scenarios cover tool_selection, tool_chaining, error_recovery, efficiency, and memory_operations. All loaded from JSON with deserialization validation test. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add benchmark runner with BenchChannel and InstrumentedLlm BenchChannel is a minimal Channel implementation for benchmarks. InstrumentedLlm wraps any LlmProvider to capture per-call metrics. Runner creates a fresh agent per scenario, evaluates success criteria, and produces RunResult with timing, token, and cost metrics. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add baseline management, reports, and benchmark entry point - baseline.rs: load/save/promote benchmark results - report.rs: format comparison reports with regression detection - benchmark_runner.rs: integration test with real LLM (feature-gated) - Add benchmark feature flag to Cargo.toml Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: apply cargo fmt to benchmark module Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add multi-turn scenario types with setup, judge, ResponseNotContains Add BenchScenario, Turn, TurnAssertions, JudgeConfig, ScenarioSetup, WorkspaceSetup, SeedDocument types for multi-turn benchmark scenarios. Add ResponseNotContains criterion variant. Add TurnAssertions::to_criteria() converter for backward compat with existing evaluation engine. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add JSON scenario loader with recursive discovery and tag filter Add load_bench_scenarios() for the new BenchScenario format with recursive directory traversal and tag-based filtering. Create 4 initial trajectory scenarios across tool-selection, multi-turn, and efficiency categories. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): multi-turn runner with workspace seeding and per-turn metrics Add run_bench_scenario() that loops over BenchScenario turns, seeds workspace documents, collects per-turn metrics (tokens, tool calls, wall time), and evaluates per-turn assertions. Add TurnMetrics to metrics.rs and clear_for_next_turn() to BenchChannel. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add LLM-as-judge scoring with prompt formatting and score parsing Create judge.rs with format_judge_prompt, parse_judge_score, and judge_turn. Wire into run_bench_scenario for turns with judge config -- scores below min_score fail the turn. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add CLI subcommand (ironclaw benchmark) Add BenchmarkCommand with --tags, --scenario, --no-judge, --timeout, --update-baseline flags. Wire into Command enum and main.rs dispatch. Feature-gated behind benchmark flag. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): per-scenario JSON output with full trajectory Add save_scenario_results() that writes per-scenario JSON files alongside the run summary. Each scenario gets its own file with turn_metrics trajectory. Update CLI to use new output format. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add ToolRegistry::retain_only and wire tool filtering in scenarios Add a retain_only() method to ToolRegistry that filters tools down to a given allowlist. Wire this into run_bench_scenario() so that when a scenario specifies a tools list in its setup, only those tools are available during the benchmark run. Includes two tests for the new method: one verifying filtering works and one verifying empty input is a no-op. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): wire identity overrides into workspace before agent start Add seed_identity() helper that writes identity files (IDENTITY.md, USER.md, etc.) into the workspace before the agent starts, so that workspace.system_prompt() picks them up. Wire it into run_bench_scenario() after workspace seeding. Include a test that verifies identity files are written and readable. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add --parallel and --max-cost CLI flags Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(benchmark): use feature-conditional snapshot names for CLI help tests Prevents snapshot conflicts between default (no benchmark) and all-features (with benchmark) builds by using separate snapshot names per feature set. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): parallel execution with JoinSet and budget cap enforcement Replace sequential loop in run_all_bench() with parallel execution using JoinSet + semaphore when config.parallel > 1. Add budget cap enforcement that skips remaining scenarios when max_total_cost_usd is exceeded. Track skipped count in RunResult.skipped_scenarios and display it in format_report(). Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add tool restriction and identity override test scenarios Co-Authored-By: Claude Opus 4.6 <[email protected]> * chore: fix formatting for Phase 3 Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add SkillRegistry::retain_only and wire skill filtering in scenarios Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add --json flag for machine-readable output Co-Authored-By: Claude Opus 4.6 <[email protected]> * ci: add GitHub Actions benchmark workflow (manual trigger) Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(benchmark): remove in-tree benchmark harness, keep retain_only utilities Move benchmark-specific code out of ironclaw in preparation for the nearai/benchmarks trajectory adapter. This removes: - src/benchmark/ (runner, scenarios, metrics, judge, report, etc.) - src/cli/benchmark.rs and the Benchmark CLI subcommand - benchmarks/ data directory (scenarios + trajectories) - .github/workflows/benchmark.yml - The "benchmark" Cargo feature flag What remains: - ToolRegistry::retain_only() and SkillRegistry::retain_only() - Test support types (TraceMetrics, InstrumentedLlm) inlined into tests/support/ instead of re-exporting from the deleted module Co-Authored-By: Claude Opus 4.6 <[email protected]> * docs: add README for LLM trace fixture format Documents the trajectory JSON format, response types, request hints, directory structure, and how to write new traces. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(test): unify trace format around turns, add multi-turn support Introduce TraceTurn type that groups user_input with LLM response steps, making traces self-contained conversation trajectories. Add run_trace() to TestRig for automatic multi-turn replay. Backward-compatible: flat "steps" JSON is deserialized as a single turn transparently. Includes all trace fixtures (spot, coverage, advanced), plan docs, and new e2e tests for steering, error recovery, long chains, memory, and prompt injection resilience. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): fix CI failures after merging main - Fix tool_json fixture: use "data" parameter (not "input") to match JsonTool schema - Fix status_events test: remove assertion for "time" tool that isn't in the fixture (only "echo" calls are used) - Allow dead_code in test support metrics/instrumented_llm modules (utilities for future benchmark tests) [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * Working on recording traces and testing them * feat(test): add declarative expects to trace fixtures, split infra tests Add TraceExpects struct with 9 optional assertion fields (response_contains, tools_used, all_tools_succeeded, etc.) that can be declared in fixture JSON instead of hand-written Rust. Add verify_expects() and run_recorded_trace() so recorded trace tests become one-liners. Split trace infra tests (deserialization, backward compat) into tests/trace_format.rs which doesn't require the libsql feature gate. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(test): add expects to all trace fixtures, simplify e2e tests Add declarative expects blocks to all 19 trace fixture JSONs across spot/, coverage/, advanced/, and root directories. Update all 8 e2e test files to use verify_trace_expects() / run_and_verify_trace(), replacing ~270 lines of hand-written assertions with fixture-driven verification. Tests that check things beyond expects (file content on disk, metrics, event ordering) keep those extra assertions alongside the declarative ones. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): adapt tests to AppBuilder refactor, fix formatting Update test files to work with refactored TestRigBuilder that uses AppBuilder::build_all() (removing with_tools/with_workspace methods). Update telegram_check fixture to use tool_list instead of echo. Fix cargo fmt issues in src/llm/mod.rs and src/llm/recording.rs. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(test): deduplicate support unit tests into single binary Support modules (assertions, cleanup, test_channel, test_rig, trace_llm) had #[cfg(test)] mod tests blocks that were compiled and run 12 times — once per e2e test binary that declares `mod support;`. Extracted all 29 support unit tests into a dedicated `tests/support_unit_tests.rs` so they run exactly once. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: fix trailing newlines in support files Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(test): unify trace types and fix recorded multi-turn replay Import shared types (TraceStep, TraceResponse, TraceToolCall, RequestHint, ExpectedToolResult, MemorySnapshotEntry, HttpExchange*) from ironclaw::llm::recording instead of redefining them in trace_llm.rs. Fix the flat-steps deserializer to split at UserInput boundaries into multiple turns, instead of filtering them out and wrapping everything into a single turn. This enables recorded multi-turn traces to be replayed as proper multi-turn conversations via run_trace(). [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): fix CI failures - unused imports and missing struct fields - Add #[allow(unused_imports)] on pub use re-exports in trace_llm.rs (types are re-exported for downstream test files, not used locally) - Add `..` to ToolCompleted pattern in test_channel.rs to match new `error` and `parameters` fields Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): fix CI failures after merging main - Add missing `error` and `parameters` fields to ToolCompleted constructors in support_unit_tests.rs - Add `..` to ToolCompleted pattern match in support_unit_tests.rs - Add #[allow(dead_code)] to CleanupGuard, LlmTrace impl, and TraceLlm impl (only used behind #[cfg(feature = "libsql")]) Co-Authored-By: Claude Opus 4.6 <[email protected]> * Adding coverage running script * fix(test): address review feedback on E2E test infrastructure - Increase wait_for_responses polling to exponential backoff (50ms-500ms) and raise default timeout from 15s to 30s to reduce CI flakiness (#1) - Strengthen prompt_injection_resilience test with positive safety layer assertion via has_safety_warnings(), enable injection_check (#2) - Add assert_tool_order() helper and tools_order field in TraceExpects for verifying tool execution ordering in multi-step traces (#3) - Document TraceLlm sequential-call assumption for concurrency (#6) - Clean up CleanupGuard with PathKind enum instead of shotgun remove_file + remove_dir_all on every path (#8) - Fix coverage.sh: default to --lib only, fix multi-filter syntax, add COV_ALL_TARGETS option - Add coverage/ to .gitignore - Remove planning docs from PR [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address PR review - use HashSet in retain_only, improve skill test - Use HashSet for O(N+M) lookup in SkillRegistry::retain_only and ToolRegistry::retain_only instead of linear scan - Strengthen test_retain_only_empty_is_noop in SkillRegistry to pre-populate with a skill before asserting the no-op behavior [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): revert incorrect safety layer assertion in injection test The safety layer sanitizes tool output, not user input. The injection test sends a malicious user message with no tools called, so the safety layer never fires. Reverted to the original test which correctly validates the LLM refuses via trace expects. Also fixed case-sensitive request hint ("ignore" -> "Ignore") to suppress noisy warning. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: clean stale profdata before coverage run Adds `cargo llvm-cov clean` before each run to prevent "mismatched data" warnings from stale instrumentation profiles. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: fix formatting in retain_only test [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Co-authored-by: Claude Opus 4.6 <[email protected]> Co-authored-by: Illia Polosukhin <[email protected]>
455 lines
16 KiB
Rust
455 lines
16 KiB
Rust
//! TraceLlm -- a replay-based LLM provider for E2E testing.
|
|
//!
|
|
//! Replays canned responses from a JSON trace, advancing through steps
|
|
//! sequentially. Supports both text and tool-call responses with optional
|
|
//! request-hint validation.
|
|
|
|
use std::path::Path;
|
|
use std::sync::Mutex;
|
|
use std::sync::atomic::{AtomicUsize, Ordering};
|
|
|
|
use async_trait::async_trait;
|
|
use rust_decimal::Decimal;
|
|
use serde::{Deserialize, Serialize};
|
|
|
|
use ironclaw::error::LlmError;
|
|
use ironclaw::llm::{
|
|
ChatMessage, CompletionRequest, CompletionResponse, FinishReason, LlmProvider, Role, ToolCall,
|
|
ToolCompletionRequest, ToolCompletionResponse,
|
|
};
|
|
|
|
// Re-export shared types from recording module so existing test code can
|
|
// still import them from here.
|
|
// Re-export all shared types so downstream test files can import from here.
|
|
#[allow(unused_imports)]
|
|
pub use ironclaw::llm::recording::{
|
|
ExpectedToolResult, HttpExchange, HttpExchangeRequest, HttpExchangeResponse,
|
|
MemorySnapshotEntry, RequestHint, TraceResponse, TraceStep, TraceToolCall,
|
|
};
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Trace types (test-only wrappers around shared recording types)
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/// A single turn in a trace: one user message and the LLM response steps that follow.
|
|
#[derive(Debug, Clone, Serialize, Deserialize)]
|
|
pub struct TraceTurn {
|
|
pub user_input: String,
|
|
pub steps: Vec<TraceStep>,
|
|
/// Declarative expectations for this turn (optional).
|
|
#[serde(default, skip_serializing_if = "TraceExpects::is_empty")]
|
|
pub expects: TraceExpects,
|
|
}
|
|
|
|
/// A complete LLM trace: a model name and an ordered list of turns.
|
|
///
|
|
/// Each turn pairs a user message with the LLM response steps that follow it.
|
|
/// For JSON backward compatibility, traces with a flat top-level `"steps"` array
|
|
/// (no `"turns"`) are deserialized into turns by splitting at `UserInput` boundaries.
|
|
///
|
|
/// Recorded traces (from `RecordingLlm`) may also include `memory_snapshot`,
|
|
/// `http_exchanges`, and `user_input` response steps.
|
|
#[derive(Debug, Clone, Serialize)]
|
|
pub struct LlmTrace {
|
|
pub model_name: String,
|
|
pub turns: Vec<TraceTurn>,
|
|
/// Workspace memory documents captured before the recording session.
|
|
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
|
pub memory_snapshot: Vec<MemorySnapshotEntry>,
|
|
/// HTTP exchanges recorded during the session, in order.
|
|
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
|
pub http_exchanges: Vec<HttpExchange>,
|
|
/// Declarative expectations for the whole trace (optional).
|
|
#[serde(default, skip_serializing_if = "TraceExpects::is_empty")]
|
|
pub expects: TraceExpects,
|
|
/// Raw steps before turn conversion (populated only for recorded traces).
|
|
/// Used by `playable_steps()` for recorded-format inspection.
|
|
#[serde(skip)]
|
|
#[allow(dead_code)]
|
|
pub steps: Vec<TraceStep>,
|
|
}
|
|
|
|
/// Declarative expectations for a trace or turn.
|
|
///
|
|
/// All fields are optional and default to empty/None, so traces without
|
|
/// `expects` work unchanged (backward compatible).
|
|
#[derive(Debug, Clone, Default, Serialize, Deserialize)]
|
|
pub struct TraceExpects {
|
|
/// Each string must appear in the response (case-insensitive).
|
|
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
|
pub response_contains: Vec<String>,
|
|
/// None of these may appear in the response (case-insensitive).
|
|
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
|
pub response_not_contains: Vec<String>,
|
|
/// Regex that must match the response.
|
|
#[serde(default, skip_serializing_if = "Option::is_none")]
|
|
pub response_matches: Option<String>,
|
|
/// Each tool name must appear in started calls.
|
|
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
|
pub tools_used: Vec<String>,
|
|
/// None of these tool names may appear.
|
|
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
|
pub tools_not_used: Vec<String>,
|
|
/// If true, all tools must succeed.
|
|
#[serde(default, skip_serializing_if = "Option::is_none")]
|
|
pub all_tools_succeeded: Option<bool>,
|
|
/// Upper bound on tool call count.
|
|
#[serde(default, skip_serializing_if = "Option::is_none")]
|
|
pub max_tool_calls: Option<usize>,
|
|
/// Minimum response count.
|
|
#[serde(default, skip_serializing_if = "Option::is_none")]
|
|
pub min_responses: Option<usize>,
|
|
/// Tool result preview must contain substring (tool_name -> substring).
|
|
#[serde(default, skip_serializing_if = "std::collections::HashMap::is_empty")]
|
|
pub tool_results_contain: std::collections::HashMap<String, String>,
|
|
/// Tools must have been called in this relative order.
|
|
#[serde(default, skip_serializing_if = "Vec::is_empty")]
|
|
pub tools_order: Vec<String>,
|
|
}
|
|
|
|
impl TraceExpects {
|
|
/// Returns true if no expectations are set.
|
|
pub fn is_empty(&self) -> bool {
|
|
self.response_contains.is_empty()
|
|
&& self.response_not_contains.is_empty()
|
|
&& self.response_matches.is_none()
|
|
&& self.tools_used.is_empty()
|
|
&& self.tools_not_used.is_empty()
|
|
&& self.all_tools_succeeded.is_none()
|
|
&& self.max_tool_calls.is_none()
|
|
&& self.min_responses.is_none()
|
|
&& self.tool_results_contain.is_empty()
|
|
&& self.tools_order.is_empty()
|
|
}
|
|
}
|
|
|
|
/// Raw deserialization helper -- accepts either `turns` or flat `steps`.
|
|
#[derive(Deserialize)]
|
|
struct RawLlmTrace {
|
|
model_name: String,
|
|
#[serde(default)]
|
|
steps: Vec<TraceStep>,
|
|
#[serde(default)]
|
|
turns: Vec<TraceTurn>,
|
|
#[serde(default)]
|
|
memory_snapshot: Vec<MemorySnapshotEntry>,
|
|
#[serde(default)]
|
|
http_exchanges: Vec<HttpExchange>,
|
|
#[serde(default)]
|
|
expects: TraceExpects,
|
|
}
|
|
|
|
impl<'de> Deserialize<'de> for LlmTrace {
|
|
fn deserialize<D>(deserializer: D) -> Result<Self, D::Error>
|
|
where
|
|
D: serde::Deserializer<'de>,
|
|
{
|
|
let raw = RawLlmTrace::deserialize(deserializer)?;
|
|
// Keep the raw steps for `playable_steps()` inspection.
|
|
let raw_steps = raw.steps.clone();
|
|
let turns = if !raw.turns.is_empty() {
|
|
raw.turns
|
|
} else if !raw.steps.is_empty() {
|
|
// Split flat steps at UserInput boundaries into turns.
|
|
let mut turns = Vec::new();
|
|
let mut current_input = "(test input)".to_string();
|
|
let mut current_steps: Vec<TraceStep> = Vec::new();
|
|
|
|
for step in raw.steps {
|
|
if let TraceResponse::UserInput { ref content } = step.response {
|
|
// Flush accumulated steps as a turn (if any).
|
|
if !current_steps.is_empty() {
|
|
turns.push(TraceTurn {
|
|
user_input: current_input.clone(),
|
|
steps: std::mem::take(&mut current_steps),
|
|
expects: TraceExpects::default(),
|
|
});
|
|
}
|
|
current_input = content.clone();
|
|
} else {
|
|
current_steps.push(step);
|
|
}
|
|
}
|
|
|
|
// Flush remaining steps.
|
|
if !current_steps.is_empty() {
|
|
turns.push(TraceTurn {
|
|
user_input: current_input,
|
|
steps: current_steps,
|
|
expects: TraceExpects::default(),
|
|
});
|
|
}
|
|
|
|
turns
|
|
} else {
|
|
vec![]
|
|
};
|
|
Ok(LlmTrace {
|
|
model_name: raw.model_name,
|
|
turns,
|
|
memory_snapshot: raw.memory_snapshot,
|
|
http_exchanges: raw.http_exchanges,
|
|
expects: raw.expects,
|
|
steps: raw_steps,
|
|
})
|
|
}
|
|
}
|
|
|
|
#[allow(dead_code)]
|
|
impl LlmTrace {
|
|
/// Create a trace from turns.
|
|
pub fn new(model_name: impl Into<String>, turns: Vec<TraceTurn>) -> Self {
|
|
Self {
|
|
model_name: model_name.into(),
|
|
turns,
|
|
memory_snapshot: Vec::new(),
|
|
http_exchanges: Vec::new(),
|
|
expects: TraceExpects::default(),
|
|
steps: Vec::new(),
|
|
}
|
|
}
|
|
|
|
/// Convenience: create a single-turn trace (for simple tests).
|
|
pub fn single_turn(
|
|
model_name: impl Into<String>,
|
|
user_input: impl Into<String>,
|
|
steps: Vec<TraceStep>,
|
|
) -> Self {
|
|
Self {
|
|
model_name: model_name.into(),
|
|
turns: vec![TraceTurn {
|
|
user_input: user_input.into(),
|
|
steps,
|
|
expects: TraceExpects::default(),
|
|
}],
|
|
memory_snapshot: Vec::new(),
|
|
http_exchanges: Vec::new(),
|
|
expects: TraceExpects::default(),
|
|
steps: Vec::new(),
|
|
}
|
|
}
|
|
|
|
/// Load a trace from a JSON file.
|
|
pub fn from_file(path: impl AsRef<Path>) -> Result<Self, Box<dyn std::error::Error>> {
|
|
let contents = std::fs::read_to_string(path)?;
|
|
let trace: Self = serde_json::from_str(&contents)?;
|
|
Ok(trace)
|
|
}
|
|
|
|
/// Return only the playable steps from the raw steps (text + tool_calls),
|
|
/// skipping `user_input` markers. Only meaningful for recorded traces that
|
|
/// were deserialized from a flat `steps` array.
|
|
#[allow(dead_code)]
|
|
pub fn playable_steps(&self) -> Vec<&TraceStep> {
|
|
self.steps
|
|
.iter()
|
|
.filter(|s| !matches!(s.response, TraceResponse::UserInput { .. }))
|
|
.collect()
|
|
}
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// TraceLlm provider
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/// An `LlmProvider` that replays canned responses from a trace.
|
|
///
|
|
/// Steps from all turns are flattened into a single sequence at construction
|
|
/// time. The provider advances through them linearly regardless of turn
|
|
/// boundaries.
|
|
///
|
|
/// **Concurrency assumption:** Uses `AtomicUsize` for step indexing, so
|
|
/// concurrent calls to `complete`/`complete_with_tools` may consume steps
|
|
/// in non-deterministic order. Current tests are single-threaded per rig;
|
|
/// if parallel tool execution is ever enabled, steps may interleave.
|
|
pub struct TraceLlm {
|
|
model_name: String,
|
|
steps: Vec<TraceStep>,
|
|
index: AtomicUsize,
|
|
hint_mismatches: AtomicUsize,
|
|
captured_requests: Mutex<Vec<Vec<ChatMessage>>>,
|
|
}
|
|
|
|
#[allow(dead_code)]
|
|
impl TraceLlm {
|
|
/// Create from an in-memory trace.
|
|
pub fn from_trace(trace: LlmTrace) -> Self {
|
|
let steps: Vec<TraceStep> = trace.turns.into_iter().flat_map(|t| t.steps).collect();
|
|
Self {
|
|
model_name: trace.model_name,
|
|
steps,
|
|
index: AtomicUsize::new(0),
|
|
hint_mismatches: AtomicUsize::new(0),
|
|
captured_requests: Mutex::new(Vec::new()),
|
|
}
|
|
}
|
|
|
|
/// Load from a JSON file and create the provider.
|
|
pub fn from_file(path: impl AsRef<Path>) -> Result<Self, Box<dyn std::error::Error>> {
|
|
let trace = LlmTrace::from_file(path)?;
|
|
Ok(Self::from_trace(trace))
|
|
}
|
|
|
|
/// Number of calls made so far.
|
|
pub fn calls(&self) -> usize {
|
|
self.index.load(Ordering::Relaxed)
|
|
}
|
|
|
|
/// Number of request-hint mismatches observed (warnings only).
|
|
pub fn hint_mismatches(&self) -> usize {
|
|
self.hint_mismatches.load(Ordering::Relaxed)
|
|
}
|
|
|
|
/// Clone of all captured request message lists.
|
|
pub fn captured_requests(&self) -> Vec<Vec<ChatMessage>> {
|
|
self.captured_requests.lock().unwrap().clone()
|
|
}
|
|
|
|
// -- internal helpers ---------------------------------------------------
|
|
|
|
/// Advance the step index and return the current step, or an error if exhausted.
|
|
fn next_step(&self, messages: &[ChatMessage]) -> Result<TraceStep, LlmError> {
|
|
// Capture the request messages.
|
|
self.captured_requests
|
|
.lock()
|
|
.unwrap()
|
|
.push(messages.to_vec());
|
|
|
|
let idx = self.index.fetch_add(1, Ordering::Relaxed);
|
|
let step = self
|
|
.steps
|
|
.get(idx)
|
|
.ok_or_else(|| LlmError::RequestFailed {
|
|
provider: self.model_name.clone(),
|
|
reason: format!(
|
|
"TraceLlm exhausted: called {} times but only {} steps",
|
|
idx + 1,
|
|
self.steps.len()
|
|
),
|
|
})?
|
|
.clone();
|
|
|
|
// Soft-validate request hints.
|
|
if let Some(ref hint) = step.request_hint {
|
|
self.validate_hint(hint, messages);
|
|
}
|
|
|
|
Ok(step)
|
|
}
|
|
|
|
fn validate_hint(&self, hint: &RequestHint, messages: &[ChatMessage]) {
|
|
if let Some(ref expected_substr) = hint.last_user_message_contains {
|
|
let last_user = messages.iter().rev().find(|m| matches!(m.role, Role::User));
|
|
let matched = last_user
|
|
.map(|m| m.content.contains(expected_substr.as_str()))
|
|
.unwrap_or(false);
|
|
if !matched {
|
|
self.hint_mismatches.fetch_add(1, Ordering::Relaxed);
|
|
eprintln!(
|
|
"[TraceLlm WARN] Request hint mismatch: expected last user message to contain {:?}, \
|
|
got {:?}",
|
|
expected_substr,
|
|
last_user.map(|m| &m.content),
|
|
);
|
|
}
|
|
}
|
|
|
|
if let Some(min_count) = hint.min_message_count
|
|
&& messages.len() < min_count
|
|
{
|
|
self.hint_mismatches.fetch_add(1, Ordering::Relaxed);
|
|
eprintln!(
|
|
"[TraceLlm WARN] Request hint mismatch: expected >= {} messages, got {}",
|
|
min_count,
|
|
messages.len(),
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
#[async_trait]
|
|
impl LlmProvider for TraceLlm {
|
|
fn model_name(&self) -> &str {
|
|
&self.model_name
|
|
}
|
|
|
|
fn cost_per_token(&self) -> (Decimal, Decimal) {
|
|
(Decimal::ZERO, Decimal::ZERO)
|
|
}
|
|
|
|
async fn complete(&self, request: CompletionRequest) -> Result<CompletionResponse, LlmError> {
|
|
let step = self.next_step(&request.messages)?;
|
|
match step.response {
|
|
TraceResponse::Text {
|
|
content,
|
|
input_tokens,
|
|
output_tokens,
|
|
} => Ok(CompletionResponse {
|
|
content,
|
|
input_tokens,
|
|
output_tokens,
|
|
finish_reason: FinishReason::Stop,
|
|
}),
|
|
TraceResponse::ToolCalls { .. } => Err(LlmError::RequestFailed {
|
|
provider: self.model_name.clone(),
|
|
reason: "TraceLlm::complete() called but current step is a tool_calls response; \
|
|
use complete_with_tools() instead"
|
|
.to_string(),
|
|
}),
|
|
TraceResponse::UserInput { .. } => Err(LlmError::RequestFailed {
|
|
provider: self.model_name.clone(),
|
|
reason: "TraceLlm::complete() encountered a user_input step; \
|
|
these should have been filtered out during construction"
|
|
.to_string(),
|
|
}),
|
|
}
|
|
}
|
|
|
|
async fn complete_with_tools(
|
|
&self,
|
|
request: ToolCompletionRequest,
|
|
) -> Result<ToolCompletionResponse, LlmError> {
|
|
let step = self.next_step(&request.messages)?;
|
|
match step.response {
|
|
TraceResponse::Text {
|
|
content,
|
|
input_tokens,
|
|
output_tokens,
|
|
} => Ok(ToolCompletionResponse {
|
|
content: Some(content),
|
|
tool_calls: Vec::new(),
|
|
input_tokens,
|
|
output_tokens,
|
|
finish_reason: FinishReason::Stop,
|
|
}),
|
|
TraceResponse::ToolCalls {
|
|
tool_calls,
|
|
input_tokens,
|
|
output_tokens,
|
|
} => {
|
|
let calls: Vec<ToolCall> = tool_calls
|
|
.into_iter()
|
|
.map(|tc| ToolCall {
|
|
id: tc.id,
|
|
name: tc.name,
|
|
arguments: tc.arguments,
|
|
})
|
|
.collect();
|
|
Ok(ToolCompletionResponse {
|
|
content: None,
|
|
tool_calls: calls,
|
|
input_tokens,
|
|
output_tokens,
|
|
finish_reason: FinishReason::ToolUse,
|
|
})
|
|
}
|
|
TraceResponse::UserInput { .. } => Err(LlmError::RequestFailed {
|
|
provider: self.model_name.clone(),
|
|
reason: "TraceLlm::complete_with_tools() encountered a user_input step; \
|
|
these should have been filtered out during construction"
|
|
.to_string(),
|
|
}),
|
|
}
|
|
}
|
|
}
|