mirror of
https://github.com/outbackdingo/optimclaw.git
synced 2026-08-26 15:40:18 +00:00
* feat(setup): add Anthropic OAuth and Codex OAuth onboarding flows Add OAuth token authentication as an alternative to API keys during onboarding for both Anthropic (via `claude login`) and OpenAI/Codex (via `~/.codex/auth.json`). Key changes: - New `AnthropicOAuthProvider` using `Authorization: Bearer` header (rig-core hardcodes `x-api-key` which rejects OAuth tokens) - Wizard auth method selector: "Direct API Key" vs "OAuth Token" for both Anthropic and OpenAI providers - Codex token extraction from `$CODEX_HOME/auth.json` / `~/.codex/auth.json` - Claude Code sandbox sub-step in Docker setup (checks for credentials) - Secret injection mappings for `ANTHROPIC_OAUTH_TOKEN` and `CODEX_OAUTH_TOKEN` - `CODEX_OAUTH_TOKEN` falls back to `OPENAI_API_KEY` (same Bearer auth) Supersedes #143 which had a broken auth flow (OAuth token sent as x-api-key → 401). Credit to @bigguybobby for the original approach. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: persist OAuth tokens in bootstrap .env and re-extract at startup OAuth tokens stored only in the secrets DB were invisible to Config::from_env() which runs before the DB connects (chicken-and-egg). Two fixes: 1. write_bootstrap_env() now persists ANTHROPIC_OAUTH_TOKEN and CODEX_OAUTH_TOKEN to ~/.ironclaw/.env (same pattern as NEARAI_API_KEY) 2. main.rs re-extracts a fresh token from the OS credential store (macOS Keychain / ~/.claude/.credentials.json) before config resolution, handling token expiry (8-12h) gracefully Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: persist all LLM credentials in bootstrap .env, not just NEAR AI All providers had the same chicken-and-egg issue: API keys stored in the secrets DB were invisible to Config::from_env() which runs before DB connects. Only NEARAI_API_KEY was written to bootstrap .env. Now write_bootstrap_env() persists all credential env vars: NEARAI_API_KEY, ANTHROPIC_API_KEY, ANTHROPIC_OAUTH_TOKEN, OPENAI_API_KEY, CODEX_OAUTH_TOKEN, LLM_API_KEY, TINFOIL_API_KEY. Also: setup_api_key_provider() now sets the env var during the wizard session so write_bootstrap_env() can pick it up. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address security review findings for OAuth onboarding - Extract "oauth-placeholder" to named OAUTH_PLACEHOLDER constant shared across config and wizard to prevent silent drift - Document plaintext credential tradeoff in write_bootstrap_env (API keys stored with 0o600 permissions, recommend full-disk encryption) - Add blocking "Press Enter" wait in Anthropic OAuth retry flow so user has time to run `claude login` in another terminal - Add escape hatch from manual OAuth paste back to API key flow (empty input switches to setup_api_key_provider) - Fix Retry-After header: parse u64 seconds into Duration before passing to LlmError::RateLimited - Make config::llm module pub(crate) for constant visibility - Use .bearer_auth() instead of manual format!("Bearer {}") - Remove response body from debug log (may contain PII) - Update Anthropic API version to 2024-10-22 Co-Authored-By: Claude Opus 4.6 <[email protected]> * security: remove plaintext credentials from bootstrap .env Credentials (API keys, OAuth tokens) were being written in plaintext to ~/.ironclaw/.env to work around a chicken-and-egg problem: Config::from_env() runs before the encrypted secrets DB is connected. Instead of storing secrets on disk, LlmConfig::resolve() now defers gracefully when credentials are missing — it returns None for the provider config instead of hard-erroring with MissingRequired. After the DB connects, AppBuilder::build_all() loads secrets from encrypted storage via inject_llm_keys_from_secrets() and re-resolves the config. For Anthropic OAuth tokens (which expire in 8-12h), the secret injection step also tries the OS credential store (macOS Keychain / Linux credentials.json) for a fresh token, overriding the potentially stale copy in the DB. Changes: - LlmConfig::resolve(): OpenAI, Anthropic, OpenAI-compatible, and Tinfoil all return None instead of MissingRequired when credentials are absent - write_bootstrap_env(): no longer writes any credential env vars - inject_llm_keys_from_secrets(): refreshes Anthropic OAuth from OS credential store before overlay is finalized - main.rs: removed OAuth re-extraction hack (no longer needed) Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: load OS credential store tokens even without secrets DB The OAuth token extraction from macOS Keychain / Linux credentials files was only running inside inject_llm_keys_from_secrets(), which requires the encrypted secrets DB. When no master key is configured, init_secrets() returned early — skipping both DB secret loading AND OS credential store extraction, leaving the Anthropic OAuth token unavailable. Split into two paths: - inject_llm_keys_from_secrets(): loads from encrypted DB + OS stores - inject_os_credentials(): loads from OS stores only (no DB needed) init_secrets() now calls inject_os_credentials() and re-resolves config even in the no-master-key early-return path, so `claude login` tokens are always available regardless of secrets DB state. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: add anthropic-beta header required for OAuth authentication Anthropic's api.anthropic.com requires the `anthropic-beta: oauth-2025-04-20` header to accept OAuth Bearer tokens. Without it, the API returns 401 "OAuth authentication is currently not supported." Also reverts API version to 2023-06-01 since the OAuth beta flag does not support the 2024-10-22 version (returns 400 "not a valid version"). This was the same bug that caused PR #143's 401 errors — the beta header was missing entirely. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: Anthropic and OpenAI model resolution respects selected_model The Anthropic and OpenAI config resolution ignored settings.selected_model entirely, only checking the provider-specific env var (ANTHROPIC_MODEL, OPENAI_MODEL) and falling back to a hardcoded default. This meant the model chosen during onboarding wizard was silently overridden. Now follows the same pattern as NearAI and OpenAI-compatible: env var > settings.selected_model > hardcoded default. Also deduplicated the Anthropic config construction (two identical branches for API key vs OAuth now share model/base_url resolution). Co-Authored-By: Claude Opus 4.6 <[email protected]> * test: add provider resolution tests for all LLM backends Covers deferred resolution (no credentials → None instead of error), credential presence, model selection fallback chain, and OAuth token routing for Anthropic, OpenAI, Tinfoil, Ollama, and NearAI. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: handle nested tokens.access_token format in Codex auth.json Codex CLI stores OAuth tokens in a nested format under tokens.access_token (ChatGPT OAuth flow), not at the top level. Also adds ENV_MUTEX to Codex token tests for thread safety. Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor: remove Codex OAuth onboarding (incompatible with OpenAI API) Codex CLI OAuth tokens use a different endpoint (chatgpt.com/backend-api/codex) and the Responses API wire format, not api.openai.com with Chat Completions. The tokens lack the model.request scope needed for the platform API, so they can't be used as drop-in OPENAI_API_KEY replacements. Removes: extract_codex_oauth_token(), wizard Codex OAuth flow, CODEX_OAUTH_TOKEN env var support, and related tests. OpenAI onboarding now uses direct API key only. Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: fix formatting for CI (cargo fmt) Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address Gemini review feedback - Use ? operator for ANTHROPIC_MODEL/BASE_URL env resolution instead of .ok().flatten() to propagate ConfigErrors consistently - Skip Tool messages without tool_call_id with a warning instead of using unwrap_or_default() which would send empty string to Anthropic - Extract credential check into closure to reduce duplication in Claude Code sandbox setup Co-Authored-By: Claude Opus 4.6 <[email protected]> * refactor(review): address PR review feedback for OAuth onboarding - Gate ANTHROPIC_OAUTH_TOKEN resolution to Anthropic provider only (was needlessly checked for all registry providers) - Add 3 regression tests for OAuth config resolution: - oauth_token sets placeholder api_key - real api_key takes priority over oauth - non-Anthropic providers don't pick up oauth_token - Validate OAuth token prefix (sk-ant-oat) in wizard to catch accidentally pasted API keys - Improve error body read handling in AnthropicOAuthProvider (was silently swallowing read errors with unwrap_or_default) - Remove extra blank line in write_bootstrap_env - Remove stale blank line in RegistryProviderConfig doc comment [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address PR #384 review comments Blocker: - Replace OnceLock<HashMap> with LazyLock<Mutex<HashMap>> for INJECTED_VARS so both inject_os_credentials() and inject_llm_keys_from_secrets() merge data instead of the second caller silently dropping its entries. High: - Add 401 retry with OS credential store re-extraction in AnthropicOAuthProvider, recovering from expired OAuth tokens (~8-12h) without manual intervention. - Fix comment in app.rs: ~/.codex/auth.json → ~/.claude/.credentials.json. Medium: - Remove unsafe { std::env::set_var } from wizard; use thread-safe inject_single_var() overlay instead (safe on multi-threaded Tokio). - Add post-init validation in AppBuilder: fail early with clear error when LLM_BACKEND is set but no credentials were resolved after secret injection. - Add sk-ant-oat prefix validation in parse_oauth_access_token(). - Only route to AnthropicOAuthProvider when api_key is missing or equals OAUTH_PLACEHOLDER (API key takes priority over OAuth token). - Teach fetch_anthropic_models() to use Bearer auth when only OAuth token is available (model listing no longer fails for OAuth-only users). Low: - Use optional_env() in wizard credential checks to read from injected overlay, not just raw env vars. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: cargo fmt Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Co-authored-by: Claude Opus 4.6 <[email protected]> Co-authored-by: [email protected] <[email protected]>
546 lines
18 KiB
Rust
546 lines
18 KiB
Rust
//! LLM integration for the agent.
|
|
//!
|
|
//! Supports multiple backends:
|
|
//! - **NEAR AI** (default): Session token or API key auth via Chat Completions API
|
|
//! - **OpenAI**: Direct API access with your own key
|
|
//! - **Anthropic**: Direct API access with your own key
|
|
//! - **Ollama**: Local model inference
|
|
//! - **OpenAI-compatible**: Any endpoint that speaks the OpenAI API
|
|
|
|
mod anthropic_oauth;
|
|
pub mod circuit_breaker;
|
|
pub mod costs;
|
|
pub mod failover;
|
|
mod nearai_chat;
|
|
mod provider;
|
|
mod reasoning;
|
|
pub mod recording;
|
|
pub mod registry;
|
|
pub mod response_cache;
|
|
pub mod retry;
|
|
mod rig_adapter;
|
|
pub mod session;
|
|
pub mod smart_routing;
|
|
|
|
pub use circuit_breaker::{CircuitBreakerConfig, CircuitBreakerProvider};
|
|
pub use failover::{CooldownConfig, FailoverProvider};
|
|
pub use nearai_chat::{ModelInfo, NearAiChatProvider};
|
|
pub use provider::{
|
|
ChatMessage, CompletionRequest, CompletionResponse, ContentPart, FinishReason, ImageUrl,
|
|
LlmProvider, ModelMetadata, Role, ToolCall, ToolCompletionRequest, ToolCompletionResponse,
|
|
ToolDefinition, ToolResult,
|
|
};
|
|
pub use reasoning::{
|
|
ActionPlan, Reasoning, ReasoningContext, RespondOutput, RespondResult, SILENT_REPLY_TOKEN,
|
|
TOOL_INTENT_NUDGE, TokenUsage, ToolSelection, is_silent_reply, llm_signals_tool_intent,
|
|
};
|
|
pub use recording::RecordingLlm;
|
|
pub use registry::{ProviderDefinition, ProviderProtocol, ProviderRegistry};
|
|
pub use response_cache::{CachedProvider, ResponseCacheConfig};
|
|
pub use retry::{RetryConfig, RetryProvider};
|
|
pub use rig_adapter::RigAdapter;
|
|
pub use session::{SessionConfig, SessionManager, create_session_manager};
|
|
pub use smart_routing::{SmartRoutingConfig, SmartRoutingProvider, TaskComplexity};
|
|
|
|
use std::sync::Arc;
|
|
|
|
use rig::client::CompletionClient;
|
|
use secrecy::ExposeSecret;
|
|
|
|
use crate::config::{LlmConfig, NearAiConfig, RegistryProviderConfig};
|
|
use crate::error::LlmError;
|
|
|
|
/// Create an LLM provider based on configuration.
|
|
///
|
|
/// - NearAI backend: Uses session manager for authentication
|
|
/// - Registry providers: Looked up by protocol and constructed generically
|
|
pub fn create_llm_provider(
|
|
config: &LlmConfig,
|
|
session: Arc<SessionManager>,
|
|
) -> Result<Arc<dyn LlmProvider>, LlmError> {
|
|
if config.backend == "nearai" || config.backend == "near_ai" || config.backend == "near" {
|
|
return create_llm_provider_with_config(&config.nearai, session);
|
|
}
|
|
|
|
let reg_config = config
|
|
.provider
|
|
.as_ref()
|
|
.ok_or_else(|| LlmError::AuthFailed {
|
|
provider: config.backend.clone(),
|
|
})?;
|
|
|
|
create_registry_provider(reg_config)
|
|
}
|
|
|
|
/// Create an LLM provider from a `NearAiConfig` directly.
|
|
///
|
|
/// This is useful when constructing additional providers for failover,
|
|
/// where only the model name differs from the primary config.
|
|
pub fn create_llm_provider_with_config(
|
|
config: &NearAiConfig,
|
|
session: Arc<SessionManager>,
|
|
) -> Result<Arc<dyn LlmProvider>, LlmError> {
|
|
let auth_mode = if config.api_key.is_some() {
|
|
"API key"
|
|
} else {
|
|
"session token"
|
|
};
|
|
tracing::info!(
|
|
model = %config.model,
|
|
base_url = %config.base_url,
|
|
auth = auth_mode,
|
|
"Using NEAR AI (Chat Completions API)"
|
|
);
|
|
Ok(Arc::new(NearAiChatProvider::new(config.clone(), session)?))
|
|
}
|
|
|
|
/// Create a provider from a registry-resolved config.
|
|
///
|
|
/// Dispatches on `RegistryProviderConfig::protocol` to build the appropriate
|
|
/// rig-core client. This single function replaces what used to be 5 separate
|
|
/// `create_*_provider` functions.
|
|
fn create_registry_provider(
|
|
config: &RegistryProviderConfig,
|
|
) -> Result<Arc<dyn LlmProvider>, LlmError> {
|
|
match config.protocol {
|
|
ProviderProtocol::OpenAiCompletions => create_openai_compat_from_registry(config),
|
|
ProviderProtocol::Anthropic => create_anthropic_from_registry(config),
|
|
ProviderProtocol::Ollama => create_ollama_from_registry(config),
|
|
}
|
|
}
|
|
|
|
fn create_openai_compat_from_registry(
|
|
config: &RegistryProviderConfig,
|
|
) -> Result<Arc<dyn LlmProvider>, LlmError> {
|
|
use rig::providers::openai;
|
|
|
|
let mut extra_headers = reqwest::header::HeaderMap::new();
|
|
for (key, value) in &config.extra_headers {
|
|
let name = match reqwest::header::HeaderName::from_bytes(key.as_bytes()) {
|
|
Ok(n) => n,
|
|
Err(e) => {
|
|
tracing::warn!(header = %key, error = %e, "Skipping extra header: invalid name");
|
|
continue;
|
|
}
|
|
};
|
|
let val = match reqwest::header::HeaderValue::from_str(value) {
|
|
Ok(v) => v,
|
|
Err(e) => {
|
|
tracing::warn!(header = %key, error = %e, "Skipping extra header: invalid value");
|
|
continue;
|
|
}
|
|
};
|
|
extra_headers.insert(name, val);
|
|
}
|
|
|
|
let api_key = config
|
|
.api_key
|
|
.as_ref()
|
|
.map(|k| k.expose_secret().to_string())
|
|
.unwrap_or_else(|| {
|
|
tracing::warn!(
|
|
provider = %config.provider_id,
|
|
"No API key configured for {}. Requests will likely fail with 401. \
|
|
Check your .env or secrets store.",
|
|
config.provider_id,
|
|
);
|
|
"no-key".to_string()
|
|
});
|
|
|
|
let mut builder = openai::Client::builder().api_key(&api_key);
|
|
if !config.base_url.is_empty() {
|
|
builder = builder.base_url(&config.base_url);
|
|
}
|
|
if !extra_headers.is_empty() {
|
|
builder = builder.http_headers(extra_headers);
|
|
}
|
|
|
|
let client: openai::Client = builder.build().map_err(|e| LlmError::RequestFailed {
|
|
provider: config.provider_id.clone(),
|
|
reason: format!("Failed to create OpenAI-compatible client: {e}"),
|
|
})?;
|
|
|
|
// Use CompletionsClient (Chat Completions API) instead of the default
|
|
// Client (Responses API). The Responses API path in rig-core handles
|
|
// tool results differently, which breaks IronClaw's tool call flow.
|
|
let client = client.completions_api();
|
|
let model = client.completion_model(&config.model);
|
|
|
|
tracing::info!(
|
|
provider = %config.provider_id,
|
|
model = %config.model,
|
|
base_url = %config.base_url,
|
|
"Using OpenAI-compatible provider"
|
|
);
|
|
|
|
Ok(Arc::new(RigAdapter::new(model, &config.model)))
|
|
}
|
|
|
|
fn create_anthropic_from_registry(
|
|
config: &RegistryProviderConfig,
|
|
) -> Result<Arc<dyn LlmProvider>, LlmError> {
|
|
// Route to OAuth provider when an OAuth token is present and no real API
|
|
// key was provided. When both are set, the API key takes priority (standard
|
|
// x-api-key auth via rig-core).
|
|
let api_key_is_placeholder = config
|
|
.api_key
|
|
.as_ref()
|
|
.is_some_and(|k| k.expose_secret() == crate::config::llm::OAUTH_PLACEHOLDER);
|
|
if config.oauth_token.is_some() && (config.api_key.is_none() || api_key_is_placeholder) {
|
|
tracing::info!(
|
|
provider = %config.provider_id,
|
|
model = %config.model,
|
|
base_url = if config.base_url.is_empty() { "default" } else { &config.base_url },
|
|
"Using Anthropic OAuth API"
|
|
);
|
|
let provider = anthropic_oauth::AnthropicOAuthProvider::new(config)?;
|
|
return Ok(Arc::new(provider));
|
|
}
|
|
|
|
use crate::config::CacheRetention;
|
|
use crate::config::helpers::optional_env;
|
|
use rig::providers::anthropic;
|
|
|
|
let api_key = config
|
|
.api_key
|
|
.as_ref()
|
|
.map(|k| k.expose_secret().to_string())
|
|
.ok_or_else(|| LlmError::AuthFailed {
|
|
provider: config.provider_id.clone(),
|
|
})?;
|
|
|
|
let client: anthropic::Client = if config.base_url.is_empty() {
|
|
anthropic::Client::new(&api_key)
|
|
} else {
|
|
anthropic::Client::builder()
|
|
.api_key(&api_key)
|
|
.base_url(&config.base_url)
|
|
.build()
|
|
}
|
|
.map_err(|e| LlmError::RequestFailed {
|
|
provider: config.provider_id.clone(),
|
|
reason: format!("Failed to create Anthropic client: {e}"),
|
|
})?;
|
|
|
|
// Resolve prompt cache retention from env (default: Short).
|
|
// Injects top-level cache_control via additional_params for Anthropic
|
|
// automatic caching (the API auto-places the breakpoint at the last
|
|
// cacheable block).
|
|
let cache_retention: CacheRetention = optional_env("ANTHROPIC_CACHE_RETENTION")
|
|
.ok()
|
|
.flatten()
|
|
.and_then(|val| match val.parse::<CacheRetention>() {
|
|
Ok(r) => Some(r),
|
|
Err(e) => {
|
|
tracing::warn!("Invalid ANTHROPIC_CACHE_RETENTION: {e}; defaulting to short");
|
|
None
|
|
}
|
|
})
|
|
.unwrap_or_default();
|
|
|
|
let model = client.completion_model(&config.model);
|
|
|
|
if cache_retention != CacheRetention::None {
|
|
tracing::info!(
|
|
model = %config.model,
|
|
retention = %cache_retention,
|
|
"Anthropic automatic prompt caching enabled"
|
|
);
|
|
}
|
|
|
|
tracing::info!(
|
|
provider = %config.provider_id,
|
|
model = %config.model,
|
|
base_url = if config.base_url.is_empty() { "default" } else { &config.base_url },
|
|
"Using Anthropic provider"
|
|
);
|
|
|
|
Ok(Arc::new(
|
|
RigAdapter::new(model, &config.model).with_cache_retention(cache_retention),
|
|
))
|
|
}
|
|
|
|
fn create_ollama_from_registry(
|
|
config: &RegistryProviderConfig,
|
|
) -> Result<Arc<dyn LlmProvider>, LlmError> {
|
|
use rig::client::Nothing;
|
|
use rig::providers::ollama;
|
|
|
|
let client: ollama::Client = ollama::Client::builder()
|
|
.base_url(&config.base_url)
|
|
.api_key(Nothing)
|
|
.build()
|
|
.map_err(|e| LlmError::RequestFailed {
|
|
provider: config.provider_id.clone(),
|
|
reason: format!("Failed to create Ollama client: {e}"),
|
|
})?;
|
|
|
|
let model = client.completion_model(&config.model);
|
|
|
|
tracing::info!(
|
|
provider = %config.provider_id,
|
|
model = %config.model,
|
|
base_url = %config.base_url,
|
|
"Using Ollama provider"
|
|
);
|
|
|
|
Ok(Arc::new(RigAdapter::new(model, &config.model)))
|
|
}
|
|
|
|
/// Create a cheap/fast LLM provider for lightweight tasks (heartbeat, routing, evaluation).
|
|
///
|
|
/// Uses `NEARAI_CHEAP_MODEL` if set, otherwise falls back to the main provider.
|
|
/// Currently only supports NEAR AI backend.
|
|
pub fn create_cheap_llm_provider(
|
|
config: &LlmConfig,
|
|
session: Arc<SessionManager>,
|
|
) -> Result<Option<Arc<dyn LlmProvider>>, LlmError> {
|
|
let Some(ref cheap_model) = config.nearai.cheap_model else {
|
|
return Ok(None);
|
|
};
|
|
|
|
if config.backend != "nearai" {
|
|
tracing::warn!(
|
|
"NEARAI_CHEAP_MODEL is set but LLM_BACKEND is '{}', not nearai. \
|
|
Cheap model setting will be ignored.",
|
|
config.backend
|
|
);
|
|
return Ok(None);
|
|
}
|
|
|
|
let mut cheap_config = config.nearai.clone();
|
|
cheap_config.model = cheap_model.clone();
|
|
|
|
Ok(Some(Arc::new(NearAiChatProvider::new(
|
|
cheap_config,
|
|
session,
|
|
)?)))
|
|
}
|
|
|
|
/// Build the full LLM provider chain with all configured wrappers.
|
|
///
|
|
/// Applies decorators in this order:
|
|
/// 1. Raw provider (from config)
|
|
/// 2. RetryProvider (per-provider retry with exponential backoff)
|
|
/// 3. SmartRoutingProvider (cheap/primary split when cheap model is configured)
|
|
/// 4. FailoverProvider (fallback model when primary fails)
|
|
/// 5. CircuitBreakerProvider (fast-fail when backend is degraded)
|
|
/// 6. CachedProvider (in-memory response cache)
|
|
///
|
|
/// Also returns a separate cheap LLM provider for heartbeat/evaluation (not
|
|
/// part of the chain — it's a standalone provider for explicitly cheap tasks).
|
|
///
|
|
/// This is the single source of truth for provider chain construction,
|
|
/// called by both `main.rs` and `app.rs`.
|
|
#[allow(clippy::type_complexity)]
|
|
pub fn build_provider_chain(
|
|
config: &LlmConfig,
|
|
session: Arc<SessionManager>,
|
|
) -> Result<
|
|
(
|
|
Arc<dyn LlmProvider>,
|
|
Option<Arc<dyn LlmProvider>>,
|
|
Option<Arc<RecordingLlm>>,
|
|
),
|
|
LlmError,
|
|
> {
|
|
let llm = create_llm_provider(config, session.clone())?;
|
|
tracing::info!("LLM provider initialized: {}", llm.model_name());
|
|
|
|
// 1. Retry
|
|
let retry_config = RetryConfig {
|
|
max_retries: config.nearai.max_retries,
|
|
};
|
|
let llm: Arc<dyn LlmProvider> = if retry_config.max_retries > 0 {
|
|
tracing::info!(
|
|
max_retries = retry_config.max_retries,
|
|
"LLM retry wrapper enabled"
|
|
);
|
|
Arc::new(RetryProvider::new(llm, retry_config.clone()))
|
|
} else {
|
|
llm
|
|
};
|
|
|
|
// 2. Smart routing (cheap/primary split)
|
|
let llm: Arc<dyn LlmProvider> = if let Some(ref cheap_model) = config.nearai.cheap_model {
|
|
let mut cheap_config = config.nearai.clone();
|
|
cheap_config.model = cheap_model.clone();
|
|
let cheap = create_llm_provider_with_config(&cheap_config, session.clone())?;
|
|
let cheap: Arc<dyn LlmProvider> = if retry_config.max_retries > 0 {
|
|
Arc::new(RetryProvider::new(cheap, retry_config.clone()))
|
|
} else {
|
|
cheap
|
|
};
|
|
tracing::info!(
|
|
primary = %llm.model_name(),
|
|
cheap = %cheap.model_name(),
|
|
"Smart routing enabled"
|
|
);
|
|
Arc::new(SmartRoutingProvider::new(
|
|
llm,
|
|
cheap,
|
|
SmartRoutingConfig {
|
|
cascade_enabled: config.nearai.smart_routing_cascade,
|
|
..SmartRoutingConfig::default()
|
|
},
|
|
))
|
|
} else {
|
|
llm
|
|
};
|
|
|
|
// 3. Failover
|
|
let llm: Arc<dyn LlmProvider> = if let Some(ref fallback_model) = config.nearai.fallback_model {
|
|
if fallback_model == &config.nearai.model {
|
|
tracing::warn!(
|
|
"fallback_model is the same as primary model, failover may not be effective"
|
|
);
|
|
}
|
|
let mut fallback_config = config.nearai.clone();
|
|
fallback_config.model = fallback_model.clone();
|
|
let fallback = create_llm_provider_with_config(&fallback_config, session.clone())?;
|
|
tracing::info!(
|
|
primary = %llm.model_name(),
|
|
fallback = %fallback.model_name(),
|
|
"LLM failover enabled"
|
|
);
|
|
let fallback: Arc<dyn LlmProvider> = if retry_config.max_retries > 0 {
|
|
Arc::new(RetryProvider::new(fallback, retry_config.clone()))
|
|
} else {
|
|
fallback
|
|
};
|
|
let cooldown_config = CooldownConfig {
|
|
cooldown_duration: std::time::Duration::from_secs(config.nearai.failover_cooldown_secs),
|
|
failure_threshold: config.nearai.failover_cooldown_threshold,
|
|
};
|
|
Arc::new(FailoverProvider::with_cooldown(
|
|
vec![llm, fallback],
|
|
cooldown_config,
|
|
)?)
|
|
} else {
|
|
llm
|
|
};
|
|
|
|
// 4. Circuit breaker
|
|
let llm: Arc<dyn LlmProvider> = if let Some(threshold) = config.nearai.circuit_breaker_threshold
|
|
{
|
|
let cb_config = CircuitBreakerConfig {
|
|
failure_threshold: threshold,
|
|
recovery_timeout: std::time::Duration::from_secs(
|
|
config.nearai.circuit_breaker_recovery_secs,
|
|
),
|
|
..CircuitBreakerConfig::default()
|
|
};
|
|
tracing::info!(
|
|
threshold,
|
|
recovery_secs = config.nearai.circuit_breaker_recovery_secs,
|
|
"LLM circuit breaker enabled"
|
|
);
|
|
Arc::new(CircuitBreakerProvider::new(llm, cb_config))
|
|
} else {
|
|
llm
|
|
};
|
|
|
|
// 5. Response cache
|
|
let llm: Arc<dyn LlmProvider> = if config.nearai.response_cache_enabled {
|
|
let rc_config = ResponseCacheConfig {
|
|
ttl: std::time::Duration::from_secs(config.nearai.response_cache_ttl_secs),
|
|
max_entries: config.nearai.response_cache_max_entries,
|
|
};
|
|
tracing::info!(
|
|
ttl_secs = config.nearai.response_cache_ttl_secs,
|
|
max_entries = config.nearai.response_cache_max_entries,
|
|
"LLM response cache enabled"
|
|
);
|
|
Arc::new(CachedProvider::new(llm, rc_config))
|
|
} else {
|
|
llm
|
|
};
|
|
|
|
// 6. Recording (trace capture for replay testing)
|
|
let recording_handle = RecordingLlm::from_env(llm.clone());
|
|
let llm: Arc<dyn LlmProvider> = if let Some(ref recorder) = recording_handle {
|
|
Arc::clone(recorder) as Arc<dyn LlmProvider>
|
|
} else {
|
|
llm
|
|
};
|
|
|
|
// Standalone cheap LLM for heartbeat/evaluation (not part of the chain)
|
|
let cheap_llm = create_cheap_llm_provider(config, session)?;
|
|
if let Some(ref cheap) = cheap_llm {
|
|
tracing::info!("Cheap LLM provider initialized: {}", cheap.model_name());
|
|
}
|
|
|
|
Ok((llm, cheap_llm, recording_handle))
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
use crate::config::NearAiConfig;
|
|
|
|
fn test_nearai_config() -> NearAiConfig {
|
|
NearAiConfig {
|
|
model: "test-model".to_string(),
|
|
cheap_model: None,
|
|
base_url: "https://api.near.ai".to_string(),
|
|
api_key: None,
|
|
fallback_model: None,
|
|
max_retries: 3,
|
|
circuit_breaker_threshold: None,
|
|
circuit_breaker_recovery_secs: 30,
|
|
response_cache_enabled: false,
|
|
response_cache_ttl_secs: 3600,
|
|
response_cache_max_entries: 1000,
|
|
failover_cooldown_secs: 300,
|
|
failover_cooldown_threshold: 3,
|
|
smart_routing_cascade: true,
|
|
}
|
|
}
|
|
|
|
fn test_llm_config() -> LlmConfig {
|
|
LlmConfig {
|
|
backend: "nearai".to_string(),
|
|
session: SessionConfig::default(),
|
|
nearai: test_nearai_config(),
|
|
provider: None,
|
|
}
|
|
}
|
|
|
|
#[test]
|
|
fn test_create_cheap_llm_provider_returns_none_when_not_configured() {
|
|
let config = test_llm_config();
|
|
let session = Arc::new(SessionManager::new(SessionConfig::default()));
|
|
|
|
let result = create_cheap_llm_provider(&config, session);
|
|
assert!(result.is_ok());
|
|
assert!(result.unwrap().is_none());
|
|
}
|
|
|
|
#[test]
|
|
fn test_create_cheap_llm_provider_creates_provider_when_configured() {
|
|
let mut config = test_llm_config();
|
|
config.nearai.cheap_model = Some("cheap-test-model".to_string());
|
|
|
|
let session = Arc::new(SessionManager::new(SessionConfig::default()));
|
|
let result = create_cheap_llm_provider(&config, session);
|
|
|
|
assert!(result.is_ok());
|
|
let provider = result.unwrap();
|
|
assert!(provider.is_some());
|
|
assert_eq!(provider.unwrap().model_name(), "cheap-test-model");
|
|
}
|
|
|
|
#[test]
|
|
fn test_create_cheap_llm_provider_ignored_for_non_nearai_backend() {
|
|
let mut config = test_llm_config();
|
|
config.backend = "openai".to_string();
|
|
config.nearai.cheap_model = Some("cheap-test-model".to_string());
|
|
|
|
let session = Arc::new(SessionManager::new(SessionConfig::default()));
|
|
let result = create_cheap_llm_provider(&config, session);
|
|
|
|
assert!(result.is_ok());
|
|
assert!(result.unwrap().is_none());
|
|
}
|
|
}
|