mirror of
https://github.com/outbackdingo/optimclaw.git
synced 2026-08-25 14:53:34 +00:00
* refactor: extract shared assertion helpers to support/assertions.rs Move 5 assertion helpers from e2e_spot_checks.rs to a shared module. Add assert_all_tools_succeeded and assert_tool_succeeded for eliminating false positives in E2E tests. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add tool output capture via tool_results() accessor Extract (name, preview) from ToolResult status events in TestChannel and TestRig, enabling content assertions on tool outputs. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: correct tool parameters in 3 broken trace fixtures - tool_time.json: add missing "operation": "now" for time tool - robust_correct_tool.json: same fix - memory_full_cycle.json: change "path" to "target" for memory_write Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: add tool success and output assertions to eliminate false positives Every E2E test that exercises tools now calls assert_all_tools_succeeded. Added tool output content assertions where tool results are predictable (time year, read_file content, memory_read content). Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: capture per-tool timing from ToolStarted/ToolCompleted events Record Instant on ToolStarted and compute elapsed duration on ToolCompleted, wiring real timing data into collect_metrics() instead of hardcoded zeros. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor: add RAII CleanupGuard for temp file/dir cleanup in tests Replace manual cleanup_test_dir() calls and inline remove_file() with Drop-based CleanupGuard that ensures cleanup even if a test panics. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: add Drop impl and graceful shutdown for TestRig Wrap agent_handle in Option so Drop can abort leaked tasks. Signal the channel shutdown before aborting for future cooperative shutdown. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: replace agent startup sleep with oneshot ready signal Use a oneshot channel fired in Channel::start() instead of a fixed 100ms sleep, eliminating the race condition on slow systems. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: replace fragile string-matching iteration limit with count-based detection Use tool completion count vs max_tool_iterations instead of scanning status messages for "iteration"/"limit" substrings. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: use assert_all_tools_succeeded for memory_full_cycle test Remove incorrect comment about memory_tree failing with empty path (it actually succeeds). Omit empty path from fixture and use the standard assert_all_tools_succeeded instead of per-tool assertions. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor: promote benchmark metrics types to library code Move TraceMetrics, ScenarioResult, RunResult, MetricDelta, and compare_runs() from tests/support/metrics.rs to src/benchmark/metrics.rs. Existing tests use re-export for backward compatibility. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add Scenario and Criterion types for agent benchmarking Scenario defines a task with input, success criteria, and resource limits. Criterion is an enum of programmatic checks (tool_used, response_contains, etc.) evaluated without LLM judgment. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add initial benchmark scenario suite (12 scenarios across 5 categories) Scenarios cover tool_selection, tool_chaining, error_recovery, efficiency, and memory_operations. All loaded from JSON with deserialization validation test. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add benchmark runner with BenchChannel and InstrumentedLlm BenchChannel is a minimal Channel implementation for benchmarks. InstrumentedLlm wraps any LlmProvider to capture per-call metrics. Runner creates a fresh agent per scenario, evaluates success criteria, and produces RunResult with timing, token, and cost metrics. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: add baseline management, reports, and benchmark entry point - baseline.rs: load/save/promote benchmark results - report.rs: format comparison reports with regression detection - benchmark_runner.rs: integration test with real LLM (feature-gated) - Add benchmark feature flag to Cargo.toml Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: apply cargo fmt to benchmark module Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add multi-turn scenario types with setup, judge, ResponseNotContains Add BenchScenario, Turn, TurnAssertions, JudgeConfig, ScenarioSetup, WorkspaceSetup, SeedDocument types for multi-turn benchmark scenarios. Add ResponseNotContains criterion variant. Add TurnAssertions::to_criteria() converter for backward compat with existing evaluation engine. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add JSON scenario loader with recursive discovery and tag filter Add load_bench_scenarios() for the new BenchScenario format with recursive directory traversal and tag-based filtering. Create 4 initial trajectory scenarios across tool-selection, multi-turn, and efficiency categories. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): multi-turn runner with workspace seeding and per-turn metrics Add run_bench_scenario() that loops over BenchScenario turns, seeds workspace documents, collects per-turn metrics (tokens, tool calls, wall time), and evaluates per-turn assertions. Add TurnMetrics to metrics.rs and clear_for_next_turn() to BenchChannel. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add LLM-as-judge scoring with prompt formatting and score parsing Create judge.rs with format_judge_prompt, parse_judge_score, and judge_turn. Wire into run_bench_scenario for turns with judge config -- scores below min_score fail the turn. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add CLI subcommand (ironclaw benchmark) Add BenchmarkCommand with --tags, --scenario, --no-judge, --timeout, --update-baseline flags. Wire into Command enum and main.rs dispatch. Feature-gated behind benchmark flag. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): per-scenario JSON output with full trajectory Add save_scenario_results() that writes per-scenario JSON files alongside the run summary. Each scenario gets its own file with turn_metrics trajectory. Update CLI to use new output format. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add ToolRegistry::retain_only and wire tool filtering in scenarios Add a retain_only() method to ToolRegistry that filters tools down to a given allowlist. Wire this into run_bench_scenario() so that when a scenario specifies a tools list in its setup, only those tools are available during the benchmark run. Includes two tests for the new method: one verifying filtering works and one verifying empty input is a no-op. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): wire identity overrides into workspace before agent start Add seed_identity() helper that writes identity files (IDENTITY.md, USER.md, etc.) into the workspace before the agent starts, so that workspace.system_prompt() picks them up. Wire it into run_bench_scenario() after workspace seeding. Include a test that verifies identity files are written and readable. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add --parallel and --max-cost CLI flags Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(benchmark): use feature-conditional snapshot names for CLI help tests Prevents snapshot conflicts between default (no benchmark) and all-features (with benchmark) builds by using separate snapshot names per feature set. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): parallel execution with JoinSet and budget cap enforcement Replace sequential loop in run_all_bench() with parallel execution using JoinSet + semaphore when config.parallel > 1. Add budget cap enforcement that skips remaining scenarios when max_total_cost_usd is exceeded. Track skipped count in RunResult.skipped_scenarios and display it in format_report(). Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add tool restriction and identity override test scenarios Co-Authored-By: Claude Opus 4.6 <[email protected]> * chore: fix formatting for Phase 3 Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add SkillRegistry::retain_only and wire skill filtering in scenarios Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(benchmark): add --json flag for machine-readable output Co-Authored-By: Claude Opus 4.6 <[email protected]> * ci: add GitHub Actions benchmark workflow (manual trigger) Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(benchmark): remove in-tree benchmark harness, keep retain_only utilities Move benchmark-specific code out of ironclaw in preparation for the nearai/benchmarks trajectory adapter. This removes: - src/benchmark/ (runner, scenarios, metrics, judge, report, etc.) - src/cli/benchmark.rs and the Benchmark CLI subcommand - benchmarks/ data directory (scenarios + trajectories) - .github/workflows/benchmark.yml - The "benchmark" Cargo feature flag What remains: - ToolRegistry::retain_only() and SkillRegistry::retain_only() - Test support types (TraceMetrics, InstrumentedLlm) inlined into tests/support/ instead of re-exporting from the deleted module Co-Authored-By: Claude Opus 4.6 <[email protected]> * docs: add README for LLM trace fixture format Documents the trajectory JSON format, response types, request hints, directory structure, and how to write new traces. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(test): unify trace format around turns, add multi-turn support Introduce TraceTurn type that groups user_input with LLM response steps, making traces self-contained conversation trajectories. Add run_trace() to TestRig for automatic multi-turn replay. Backward-compatible: flat "steps" JSON is deserialized as a single turn transparently. Includes all trace fixtures (spot, coverage, advanced), plan docs, and new e2e tests for steering, error recovery, long chains, memory, and prompt injection resilience. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): fix CI failures after merging main - Fix tool_json fixture: use "data" parameter (not "input") to match JsonTool schema - Fix status_events test: remove assertion for "time" tool that isn't in the fixture (only "echo" calls are used) - Allow dead_code in test support metrics/instrumented_llm modules (utilities for future benchmark tests) [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * Working on recording traces and testing them * feat(test): add declarative expects to trace fixtures, split infra tests Add TraceExpects struct with 9 optional assertion fields (response_contains, tools_used, all_tools_succeeded, etc.) that can be declared in fixture JSON instead of hand-written Rust. Add verify_expects() and run_recorded_trace() so recorded trace tests become one-liners. Split trace infra tests (deserialization, backward compat) into tests/trace_format.rs which doesn't require the libsql feature gate. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(test): add expects to all trace fixtures, simplify e2e tests Add declarative expects blocks to all 19 trace fixture JSONs across spot/, coverage/, advanced/, and root directories. Update all 8 e2e test files to use verify_trace_expects() / run_and_verify_trace(), replacing ~270 lines of hand-written assertions with fixture-driven verification. Tests that check things beyond expects (file content on disk, metrics, event ordering) keep those extra assertions alongside the declarative ones. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): adapt tests to AppBuilder refactor, fix formatting Update test files to work with refactored TestRigBuilder that uses AppBuilder::build_all() (removing with_tools/with_workspace methods). Update telegram_check fixture to use tool_list instead of echo. Fix cargo fmt issues in src/llm/mod.rs and src/llm/recording.rs. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(test): deduplicate support unit tests into single binary Support modules (assertions, cleanup, test_channel, test_rig, trace_llm) had #[cfg(test)] mod tests blocks that were compiled and run 12 times — once per e2e test binary that declares `mod support;`. Extracted all 29 support unit tests into a dedicated `tests/support_unit_tests.rs` so they run exactly once. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: fix trailing newlines in support files Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(test): unify trace types and fix recorded multi-turn replay Import shared types (TraceStep, TraceResponse, TraceToolCall, RequestHint, ExpectedToolResult, MemorySnapshotEntry, HttpExchange*) from ironclaw::llm::recording instead of redefining them in trace_llm.rs. Fix the flat-steps deserializer to split at UserInput boundaries into multiple turns, instead of filtering them out and wrapping everything into a single turn. This enables recorded multi-turn traces to be replayed as proper multi-turn conversations via run_trace(). [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): fix CI failures - unused imports and missing struct fields - Add #[allow(unused_imports)] on pub use re-exports in trace_llm.rs (types are re-exported for downstream test files, not used locally) - Add `..` to ToolCompleted pattern in test_channel.rs to match new `error` and `parameters` fields Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): fix CI failures after merging main - Add missing `error` and `parameters` fields to ToolCompleted constructors in support_unit_tests.rs - Add `..` to ToolCompleted pattern match in support_unit_tests.rs - Add #[allow(dead_code)] to CleanupGuard, LlmTrace impl, and TraceLlm impl (only used behind #[cfg(feature = "libsql")]) Co-Authored-By: Claude Opus 4.6 <[email protected]> * Adding coverage running script * fix(test): address review feedback on E2E test infrastructure - Increase wait_for_responses polling to exponential backoff (50ms-500ms) and raise default timeout from 15s to 30s to reduce CI flakiness (#1) - Strengthen prompt_injection_resilience test with positive safety layer assertion via has_safety_warnings(), enable injection_check (#2) - Add assert_tool_order() helper and tools_order field in TraceExpects for verifying tool execution ordering in multi-step traces (#3) - Document TraceLlm sequential-call assumption for concurrency (#6) - Clean up CleanupGuard with PathKind enum instead of shotgun remove_file + remove_dir_all on every path (#8) - Fix coverage.sh: default to --lib only, fix multi-filter syntax, add COV_ALL_TARGETS option - Add coverage/ to .gitignore - Remove planning docs from PR [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address PR review - use HashSet in retain_only, improve skill test - Use HashSet for O(N+M) lookup in SkillRegistry::retain_only and ToolRegistry::retain_only instead of linear scan - Strengthen test_retain_only_empty_is_noop in SkillRegistry to pre-populate with a skill before asserting the no-op behavior [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix(test): revert incorrect safety layer assertion in injection test The safety layer sanitizes tool output, not user input. The injection test sends a malicious user message with no tools called, so the safety layer never fires. Reverted to the original test which correctly validates the LLM refuses via trace expects. Also fixed case-sensitive request hint ("ignore" -> "Ignore") to suppress noisy warning. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: clean stale profdata before coverage run Adds `cargo llvm-cov clean` before each run to prevent "mismatched data" warnings from stale instrumentation profiles. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: fix formatting in retain_only test [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Co-authored-by: Claude Opus 4.6 <[email protected]> Co-authored-by: Illia Polosukhin <[email protected]>
316 lines
11 KiB
Rust
316 lines
11 KiB
Rust
//! Configuration for IronClaw.
|
|
//!
|
|
//! Settings are loaded with priority: env var > database > default.
|
|
//! `DATABASE_URL` lives in `~/.ironclaw/.env` (loaded via dotenvy early
|
|
//! in startup). Everything else comes from env vars, the DB settings
|
|
//! table, or auto-detection.
|
|
|
|
mod agent;
|
|
mod builder;
|
|
mod channels;
|
|
mod database;
|
|
mod embeddings;
|
|
mod heartbeat;
|
|
pub(crate) mod helpers;
|
|
mod hygiene;
|
|
mod llm;
|
|
mod routines;
|
|
mod safety;
|
|
mod sandbox;
|
|
mod secrets;
|
|
mod skills;
|
|
mod tunnel;
|
|
mod wasm;
|
|
|
|
use std::collections::HashMap;
|
|
use std::sync::OnceLock;
|
|
|
|
use crate::error::ConfigError;
|
|
use crate::settings::Settings;
|
|
|
|
// Re-export all public types so `crate::config::FooConfig` continues to work.
|
|
pub use self::agent::AgentConfig;
|
|
pub use self::builder::BuilderModeConfig;
|
|
pub use self::channels::{ChannelsConfig, CliConfig, GatewayConfig, HttpConfig, SignalConfig};
|
|
pub use self::database::{DatabaseBackend, DatabaseConfig, SslMode, default_libsql_path};
|
|
pub use self::embeddings::EmbeddingsConfig;
|
|
pub use self::heartbeat::HeartbeatConfig;
|
|
pub use self::hygiene::HygieneConfig;
|
|
pub use self::llm::{
|
|
AnthropicDirectConfig, LlmBackend, LlmConfig, NearAiConfig, OllamaConfig,
|
|
OpenAiCompatibleConfig, OpenAiDirectConfig, TinfoilConfig,
|
|
};
|
|
pub use self::routines::RoutineConfig;
|
|
pub use self::safety::SafetyConfig;
|
|
pub use self::sandbox::{ClaudeCodeConfig, SandboxModeConfig};
|
|
pub use self::secrets::SecretsConfig;
|
|
pub use self::skills::SkillsConfig;
|
|
pub use self::tunnel::TunnelConfig;
|
|
pub use self::wasm::WasmConfig;
|
|
|
|
/// Thread-safe overlay for injected env vars (secrets loaded from DB).
|
|
///
|
|
/// Used by `inject_llm_keys_from_secrets()` to make API keys available to
|
|
/// `optional_env()` without unsafe `set_var` calls. `optional_env()` checks
|
|
/// real env vars first, then falls back to this overlay.
|
|
static INJECTED_VARS: OnceLock<HashMap<String, String>> = OnceLock::new();
|
|
|
|
/// Main configuration for the agent.
|
|
#[derive(Debug, Clone)]
|
|
pub struct Config {
|
|
pub database: DatabaseConfig,
|
|
pub llm: LlmConfig,
|
|
pub embeddings: EmbeddingsConfig,
|
|
pub tunnel: TunnelConfig,
|
|
pub channels: ChannelsConfig,
|
|
pub agent: AgentConfig,
|
|
pub safety: SafetyConfig,
|
|
pub wasm: WasmConfig,
|
|
pub secrets: SecretsConfig,
|
|
pub builder: BuilderModeConfig,
|
|
pub heartbeat: HeartbeatConfig,
|
|
pub hygiene: HygieneConfig,
|
|
pub routines: RoutineConfig,
|
|
pub sandbox: SandboxModeConfig,
|
|
pub claude_code: ClaudeCodeConfig,
|
|
pub skills: SkillsConfig,
|
|
pub observability: crate::observability::ObservabilityConfig,
|
|
}
|
|
|
|
impl Config {
|
|
/// Create a full Config for integration tests without reading env vars.
|
|
///
|
|
/// Requires the `libsql` feature. Sets up:
|
|
/// - libSQL database at the given path
|
|
/// - WASM and embeddings disabled
|
|
/// - Skills enabled with the given directories
|
|
/// - Heartbeat, routines, sandbox, builder all disabled
|
|
/// - Safety with injection check off, 100k output limit
|
|
#[cfg(feature = "libsql")]
|
|
pub fn for_testing(
|
|
libsql_path: std::path::PathBuf,
|
|
skills_dir: std::path::PathBuf,
|
|
installed_skills_dir: std::path::PathBuf,
|
|
) -> Self {
|
|
Self {
|
|
database: DatabaseConfig {
|
|
backend: DatabaseBackend::LibSql,
|
|
url: secrecy::SecretString::from("unused://test".to_string()),
|
|
pool_size: 1,
|
|
ssl_mode: SslMode::Disable,
|
|
libsql_path: Some(libsql_path),
|
|
libsql_url: None,
|
|
libsql_auth_token: None,
|
|
},
|
|
llm: LlmConfig::for_testing(),
|
|
embeddings: EmbeddingsConfig::default(),
|
|
tunnel: TunnelConfig::default(),
|
|
channels: ChannelsConfig {
|
|
cli: CliConfig { enabled: false },
|
|
http: None,
|
|
gateway: None,
|
|
signal: None,
|
|
wasm_channels_dir: std::path::PathBuf::from("/tmp/ironclaw-test-channels"),
|
|
wasm_channels_enabled: false,
|
|
wasm_channel_owner_ids: HashMap::new(),
|
|
},
|
|
agent: AgentConfig::for_testing(),
|
|
safety: SafetyConfig {
|
|
max_output_length: 100_000,
|
|
injection_check_enabled: false,
|
|
},
|
|
wasm: WasmConfig {
|
|
enabled: false,
|
|
..WasmConfig::default()
|
|
},
|
|
secrets: SecretsConfig::default(),
|
|
builder: BuilderModeConfig {
|
|
enabled: false,
|
|
..BuilderModeConfig::default()
|
|
},
|
|
heartbeat: HeartbeatConfig::default(),
|
|
hygiene: HygieneConfig::default(),
|
|
routines: RoutineConfig {
|
|
enabled: false,
|
|
..RoutineConfig::default()
|
|
},
|
|
sandbox: SandboxModeConfig {
|
|
enabled: false,
|
|
..SandboxModeConfig::default()
|
|
},
|
|
claude_code: ClaudeCodeConfig::default(),
|
|
skills: SkillsConfig {
|
|
enabled: true,
|
|
local_dir: skills_dir,
|
|
installed_dir: installed_skills_dir,
|
|
..SkillsConfig::default()
|
|
},
|
|
observability: crate::observability::ObservabilityConfig::default(),
|
|
}
|
|
}
|
|
|
|
/// Load configuration from environment variables and the database.
|
|
///
|
|
/// Priority: env var > TOML config file > DB settings > default.
|
|
/// This is the primary way to load config after DB is connected.
|
|
pub async fn from_db(
|
|
store: &(dyn crate::db::SettingsStore + Sync),
|
|
user_id: &str,
|
|
) -> Result<Self, ConfigError> {
|
|
Self::from_db_with_toml(store, user_id, None).await
|
|
}
|
|
|
|
/// Load from DB with an optional TOML config file overlay.
|
|
pub async fn from_db_with_toml(
|
|
store: &(dyn crate::db::SettingsStore + Sync),
|
|
user_id: &str,
|
|
toml_path: Option<&std::path::Path>,
|
|
) -> Result<Self, ConfigError> {
|
|
let _ = dotenvy::dotenv();
|
|
crate::bootstrap::load_ironclaw_env();
|
|
|
|
// Load all settings from DB into a Settings struct
|
|
let mut db_settings = match store.get_all_settings(user_id).await {
|
|
Ok(map) => Settings::from_db_map(&map),
|
|
Err(e) => {
|
|
tracing::warn!("Failed to load settings from DB, using defaults: {}", e);
|
|
Settings::default()
|
|
}
|
|
};
|
|
|
|
// Overlay TOML config file (values win over DB settings)
|
|
Self::apply_toml_overlay(&mut db_settings, toml_path)?;
|
|
|
|
Self::build(&db_settings).await
|
|
}
|
|
|
|
/// Load configuration from environment variables only (no database).
|
|
///
|
|
/// Used during early startup before the database is connected,
|
|
/// and by CLI commands that don't have DB access.
|
|
/// Falls back to legacy `settings.json` on disk if present.
|
|
///
|
|
/// Loads both `./.env` (standard, higher priority) and `~/.ironclaw/.env`
|
|
/// (lower priority) via dotenvy, which never overwrites existing vars.
|
|
pub async fn from_env() -> Result<Self, ConfigError> {
|
|
Self::from_env_with_toml(None).await
|
|
}
|
|
|
|
/// Load from env with an optional TOML config file overlay.
|
|
pub async fn from_env_with_toml(
|
|
toml_path: Option<&std::path::Path>,
|
|
) -> Result<Self, ConfigError> {
|
|
let _ = dotenvy::dotenv();
|
|
crate::bootstrap::load_ironclaw_env();
|
|
let mut settings = Settings::load();
|
|
|
|
// Overlay TOML config file (values win over JSON settings)
|
|
Self::apply_toml_overlay(&mut settings, toml_path)?;
|
|
|
|
Self::build(&settings).await
|
|
}
|
|
|
|
/// Load and merge a TOML config file into settings.
|
|
///
|
|
/// If `explicit_path` is `Some`, loads from that path (errors are fatal).
|
|
/// If `None`, tries the default path `~/.ironclaw/config.toml` (missing
|
|
/// file is silently ignored).
|
|
fn apply_toml_overlay(
|
|
settings: &mut Settings,
|
|
explicit_path: Option<&std::path::Path>,
|
|
) -> Result<(), ConfigError> {
|
|
let path = explicit_path
|
|
.map(std::path::PathBuf::from)
|
|
.unwrap_or_else(Settings::default_toml_path);
|
|
|
|
match Settings::load_toml(&path) {
|
|
Ok(Some(toml_settings)) => {
|
|
settings.merge_from(&toml_settings);
|
|
tracing::debug!("Loaded TOML config from {}", path.display());
|
|
}
|
|
Ok(None) => {
|
|
if explicit_path.is_some() {
|
|
return Err(ConfigError::ParseError(format!(
|
|
"Config file not found: {}",
|
|
path.display()
|
|
)));
|
|
}
|
|
}
|
|
Err(e) => {
|
|
if explicit_path.is_some() {
|
|
return Err(ConfigError::ParseError(format!(
|
|
"Failed to load config file {}: {}",
|
|
path.display(),
|
|
e
|
|
)));
|
|
}
|
|
tracing::warn!("Failed to load default config file: {}", e);
|
|
}
|
|
}
|
|
Ok(())
|
|
}
|
|
|
|
/// Build config from settings (shared by from_env and from_db).
|
|
async fn build(settings: &Settings) -> Result<Self, ConfigError> {
|
|
Ok(Self {
|
|
database: DatabaseConfig::resolve()?,
|
|
llm: LlmConfig::resolve(settings)?,
|
|
embeddings: EmbeddingsConfig::resolve(settings)?,
|
|
tunnel: TunnelConfig::resolve(settings)?,
|
|
channels: ChannelsConfig::resolve(settings)?,
|
|
agent: AgentConfig::resolve(settings)?,
|
|
safety: SafetyConfig::resolve()?,
|
|
wasm: WasmConfig::resolve()?,
|
|
secrets: SecretsConfig::resolve().await?,
|
|
builder: BuilderModeConfig::resolve()?,
|
|
heartbeat: HeartbeatConfig::resolve(settings)?,
|
|
hygiene: HygieneConfig::resolve()?,
|
|
routines: RoutineConfig::resolve()?,
|
|
sandbox: SandboxModeConfig::resolve()?,
|
|
claude_code: ClaudeCodeConfig::resolve()?,
|
|
skills: SkillsConfig::resolve()?,
|
|
observability: crate::observability::ObservabilityConfig {
|
|
backend: std::env::var("OBSERVABILITY_BACKEND").unwrap_or_else(|_| "none".into()),
|
|
},
|
|
})
|
|
}
|
|
}
|
|
|
|
/// Load API keys from the encrypted secrets store into a thread-safe overlay.
|
|
///
|
|
/// This bridges the gap between secrets stored during onboarding and the
|
|
/// env-var-first resolution in `LlmConfig::resolve()`. Keys in the overlay
|
|
/// are read by `optional_env()` before falling back to `std::env::var()`,
|
|
/// so explicit env vars always win.
|
|
pub async fn inject_llm_keys_from_secrets(
|
|
secrets: &dyn crate::secrets::SecretsStore,
|
|
user_id: &str,
|
|
) {
|
|
let mappings = [
|
|
("llm_openai_api_key", "OPENAI_API_KEY"),
|
|
("llm_anthropic_api_key", "ANTHROPIC_API_KEY"),
|
|
("llm_compatible_api_key", "LLM_API_KEY"),
|
|
("llm_nearai_api_key", "NEARAI_API_KEY"),
|
|
];
|
|
|
|
let mut injected = HashMap::new();
|
|
|
|
for (secret_name, env_var) in mappings {
|
|
match std::env::var(env_var) {
|
|
Ok(val) if !val.is_empty() => continue,
|
|
_ => {}
|
|
}
|
|
match secrets.get_decrypted(user_id, secret_name).await {
|
|
Ok(decrypted) => {
|
|
injected.insert(env_var.to_string(), decrypted.expose().to_string());
|
|
tracing::debug!("Loaded secret '{}' for env var '{}'", secret_name, env_var);
|
|
}
|
|
Err(_) => {
|
|
// Secret doesn't exist, that's fine
|
|
}
|
|
}
|
|
}
|
|
|
|
let _ = INJECTED_VARS.set(injected);
|
|
}
|