* feat(db): add list_dispatched_routine_runs to RoutineStore trait
Add method to query routine runs with status='running' AND job_id IS NOT NULL,
enabling the routine engine to sync completion status from background jobs.
Implements for both PostgreSQL and libSQL backends.
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(routines): sync dispatched full-job runs with background job status (#697)
Full-job routines were immediately marked Ok on dispatch, so
failures/completions were never reflected in the routine run record.
Now dispatch returns Running status, and a periodic sync checks linked
jobs to update the run when the job completes, fails, or is cancelled.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(routines): fail fast when sandbox unavailable at dispatch time (#697)
Thread sandbox_available bool from Docker detection through AgentDeps
to RoutineEngine. Full-job routines now fail immediately with a clear
error message when sandbox is enabled but Docker is not available,
instead of dispatching a job that silently fails.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* feat(startup): notify user when sandbox unavailable (#697)
When sandbox is enabled but Docker is not installed or not running,
send a user-visible warning through all channels at startup (with a
2s delay to let channels connect). Previously this was only logged
via tracing::warn, invisible to TUI/web users.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* style: fix formatting in routine_engine.rs
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(tests): set sandbox_available=true in test rig for full_job traces
Test rig doesn't use real Docker — full_job routines execute via trace
replay. Setting sandbox_available=true allows the routine_news_digest
trace test to dispatch full_job routines as before.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(routines): address review feedback on sync_dispatched_runs (#697)
- Sanitize last_reason from job transitions before using in
notifications (truncate to 500 chars, strip control characters)
- Treat Submitted as in-progress (can still transition to Failed),
only Completed and Accepted are terminal success states
- Add test for sanitize_summary
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(tests): add missing sandbox_available field to test constructors
Staging added sandbox_available to AgentDeps and RoutineEngine::new.
Add the missing field/argument in test files to fix CI compilation.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix: sanitize job reason in notifications, fix state handling for Submitted/Accepted
- Enhance sanitize_summary to strip HTML tags and collapse whitespace,
preventing injection via untrusted container job reasons
- Use char-boundary-safe truncation to avoid panics on multi-byte strings
- Treat Submitted and Accepted as in-progress states (continue polling)
rather than terminal success, since they can still transition to Failed
- Increase channel-connect delay from 2s to 5s and add debug log for
sandbox-unavailable warning delivery
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* Replace sandbox_available bool with SandboxReadiness enum
Distinguishes DisabledByConfig from DockerUnavailable so full-job
routine errors give actionable guidance instead of a generic message.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* ci: re-trigger CI with latest changes
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix: add missing owner_id arg to send_notification call
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix: update e2e tests to use SandboxReadiness enum
Co-Authored-By: Claude Opus 4.6 <[email protected]>
---------
Co-authored-by: Claude Opus 4.6 <[email protected]>
Co-authored-by: [email protected] <[email protected]>
* feat(self-repair): wire stuck_threshold, store, and builder (#647)
Wire the previously dead-code fields in DefaultSelfRepair:
- stuck_threshold: detect_stuck_jobs() now filters by duration, only
reporting jobs stuck longer than the configured threshold
- with_store(): wired in agent_loop.rs from AgentDeps.store for
tool failure tracking via Database trait
- with_builder(): wired from register_builder_tool() return value
through AppComponents and AgentDeps for automatic tool rebuilding
- tools: passed alongside builder for hot-reload logging
Remove all #[allow(dead_code)] annotations. Add regression tests for
threshold-based filtering (both above and below threshold).
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix: add missing `builder` field to AgentDeps in gateway workflow harness
After rebase onto staging, AgentDeps gained a `builder` field for
self-repair tool rebuilding. The gateway workflow test harness was
missing this field, causing CI compilation failure.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* ci: retrigger CI
* fix: force CI refresh after path_routing_tests dedup
* test: add E2E test for stuck job repair and tool rebuild cycle
Tests the full self-repair flow requested in review:
1. Job transitions Pending -> InProgress -> Stuck
2. detect_stuck_jobs() finds it (zero threshold)
3. repair_stuck_job() recovers it back to InProgress
4. A broken tool is repaired via MockBuilder
5. Verify builder was invoked and repair succeeded
Uses a MockBuilder (impl SoftwareBuilder) that returns successful
BuildResult without requiring an LLM or filesystem. Uses libsql
test database for the store (increment_repair_attempts, mark_tool_repaired).
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix(self-repair): measure stuck_duration from Stuck transition, not started_at
- Use ctx.transitions to find the most recent Stuck transition timestamp
instead of ctx.started_at (which reflects job start, not stuck time)
- Fix StuckJob.last_activity to use stuck transition timestamp
- Remove misleading "hot-reloaded into registry" log
- Remove stray "// ci fix" comment in memory.rs
- Add regression test: backdated started_at must not inflate stuck_duration
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* ci: re-trigger CI with latest changes
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix: add type annotation to Ok(()) in test to resolve E0282
Co-Authored-By: Claude Opus 4.6 <[email protected]>
---------
Co-authored-by: Claude Opus 4.6 <[email protected]>
* feat(testing): add FaultInjector framework for StubLlm (#1220)
Adds a configurable fault injection framework for testing retry, failover,
and circuit breaker behavior. The FaultInjector attaches to StubLlm and
provides per-call control over failure type, timing, and sequencing.
Components:
- FaultType: maps to LlmError variants (RequestFailed, RateLimited,
AuthFailed, InvalidResponse, IoError, ContextLengthExceeded, SessionExpired)
- FaultAction: Succeed, Fail(FaultType), Delay(Duration)
- FaultMode: SequenceOnce (play then succeed), SequenceLoop (repeat forever),
Random (seeded xorshift64 PRNG for reproducibility)
- FaultInjector: thread-safe (AtomicU32 counter + Mutex RNG)
Integration:
- StubLlm gains optional fault_injector field via with_fault_injector()
- When set, takes precedence over should_fail/error_kind
- Backward compatible: existing StubLlm usage unchanged
Closes#1220
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* refactor(testing): address review feedback on FaultInjector
- Remove redundant .abs() in random fault comparison
- Extract check_faults() helper to DRY up StubLlm methods
- Guard xorshift seed=0 (fixed point) by mapping to 1
- Add StubLlm integration test (stub_llm_fault_injector_sequence)
- Remove dead seed field from FaultMode::Random
- Move pub mod fault_injection to top of mod.rs
- Add Debug impl for FaultInjector
- Add empty_sequence_always_succeeds test
- Add random_seed_zero_does_not_always_fail test
* fix(testing): address #1233 review -- seed-0 bug, reset(), Debug derive
- Store seed in FaultMode::Random so reset() can re-init the RNG
- Add reset() method for test reproducibility (re-seeds RNG, zeros counter)
- Strengthen seed=0 regression test to 100 iterations with stricter assertion
- Add reset_restores_random_rng_from_stored_seed test
- Debug impl and empty_sequence test were already present from prior commit
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* ci: re-trigger CI with latest changes
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* ci: trigger new run with skip-regression-check label
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(testing): address PR #1233 review -- error_rate validation and edge cases
- Validate error_rate is in 0.0..=1.0 and not NaN (panics on invalid input)
- Fix error_rate==1.0 edge case: use <= instead of < so 1.0 always fails
- Add regression tests for error_rate validation (NaN, negative, >1.0)
- Add tests for error_rate boundary values (0.0 never fails, 1.0 always fails)
- Add delay action test using tokio::time::pause() for deterministic timing
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <[email protected]>
* Rebase onto staging
* fix(routines): prevent autonomy-escalation in lightweight routines
- Add ROUTINE_TOOL_DENYLIST to block routine_create/update/delete/fire
and restart from being callable by lightweight routines
- Deduplicate sentinel logic by reusing handle_text_response() in the
no-tools path
- Filter tool definitions sent to LLM to only include callable tools,
avoiding wasted tokens on tools that would be rejected
* refactor: centralize test credential constants into testing::credentials
Scattered test credential strings (API keys, OAuth tokens, crypto keys,
Telegram tokens, session tokens) across ~25 files made security auditing
harder and created unnecessary duplication. Centralize all test-only fake
credentials into a new `src/testing/credentials.rs` module with named
constants and a shared `test_secrets_store()` helper.
- Convert `src/testing.rs` to directory module (`src/testing/mod.rs`)
- Add `src/testing/credentials.rs` with ~30 named constants
- Replace hardcoded literals in 24 source files
- Deduplicate `test_store()` helper (was copy-pasted in 3 files)
- Leave leak_detector/shell/signature tests as-is (inline values
aid readability for pattern detection tests)
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* refactor: replace real Telegram bot token with obviously fake test stub
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
* Update src/testing/credentials.rs
Co-authored-by: Copilot <[email protected]>
* Update src/testing/credentials.rs
Co-authored-by: Copilot <[email protected]>
* refactor: address PR review feedback on test credentials
- Fix TEST_CRYPTO_KEY doc comment ("32-byte hex" → "32-character key string")
- Rename confusing "real"/"fake" Anthropic constant names and values
- Change TEST_STRIPE_KEY from "sk-live" to "sk_test_fake123" to avoid scanners
- Use test_secrets_store() helper in orchestrator and http tool tests
- Clarify config_round_trip.rs doc comment about integration test visibility
Co-Authored-By: Claude Opus 4.6 <[email protected]>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <[email protected]>
Co-authored-by: Copilot <[email protected]>