mirror of
https://github.com/outbackdingo/optimclaw.git
synced 2026-08-25 14:53:34 +00:00
* feat(llm): add smart model routing based on request complexity Automatically selects optimal model tier (flash/standard/pro/frontier) for each request based on 13-dimension complexity scoring: - Reasoning words, multi-step signals, code indicators - Domain-specific terms, creativity, precision - Safety sensitivity, tool likelihood, question complexity - Token estimate, context dependency, sentence complexity Features: - Pattern overrides for fast-path routing (greetings → flash, security audits → frontier) - Configurable tier-to-model mappings (defaults to -latest aliases) - Thinking mode per tier (pro: low, frontier: medium) - User-configurable pattern overrides - Zero-config for default benefits, full control for power users Expected cost savings: 50-70% vs always-using-frontier baseline. Refs: smart-routing-spec.md * fix(routing): address Gemini Code Assist review feedback - Add tracing warnings for invalid tier/regex in user overrides (router.rs) - Use unreachable!() for tier hint match since regex enforces valid tiers (scorer.rs) - Refactor weighted total to array iteration for maintainability (scorer.rs) - Add TODO for making domain keywords configurable (scorer.rs) Refs: PR #208 * feat(routing): make domain keywords configurable - Add ScorerConfig with optional domain_keywords field - Add DEFAULT_DOMAIN_KEYWORDS constant (exported for reference) - Add domain_keywords to RouterConfig for top-level configuration - Build domain regex at runtime from config, fallback to defaults - Add score_complexity_with_config() function - Add test for custom domain keywords Users can now provide project-specific keywords: RouterConfig { domain_keywords: Some(vec!["mycompany".into(), "myproduct".into()]), ..Default::default() } Addresses Gemini Code Assist review feedback on PR #208. Tests: 20/20 passing * docs: add domain_keywords to routing config example * feat: integrate 13-dimension complexity scorer into smart routing (takeover #208) Folds the 13-dimension complexity scorer and pattern overrides from PR #208 into the existing SmartRoutingProvider, replacing the simpler keyword-based classifier. Adds 4-tier system (Flash/Standard/Pro/Frontier), configurable scorer weights, domain keywords, regex pattern overrides, tier hints, and multi-dimensional boost. Removes separate routing/ directory and lazy_static dependency in favor of std::sync::LazyLock. Includes 44 tests covering all scoring dimensions, tier boundaries, pattern overrides, and provider routing. Co-Authored-By: onlyamicrowave <[email protected]> Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address review feedback on smart routing PR (#529) - Cache compiled domain regex in SmartRoutingProvider (built once at construction, not per-request) and add score_complexity_with_regex() API - Check explicit tier hints before pattern overrides so user intent wins (e.g. "[tier:flash] security audit" routes as Flash, not Frontier) - Trim input before matching/scoring so trailing whitespace doesn't break anchored override regexes or skew token-length scoring - Fix token estimate comment (>=520 chars = 100, not >500) - Update spec: check implementation plan boxes, fix file paths, add note that llm.routing YAML schema is target design (current config uses env vars) - Add regression tests for tier hint precedence and trimmed greeting matching Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: restore Cargo.lock from main to fix html_to_markdown test The lockfile was fully regenerated during the PR #208 merge conflict resolution, which bumped html-to-markdown-rs from 2.25.1 to 2.27.2. The new version produces different output that breaks the golden-file snapshot test. Restore the original lockfile from main — lazy_static was never in main's lockfile, so no further changes needed. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address second round of review feedback (#529) - Tighten quick-lookup override regex with end anchor to prevent matching complex questions like "What time complexity is merge sort?" - Handle empty domain keywords list by falling back to defaults instead of producing a broken regex that matches empty strings everywhere - Clarify spec architecture diagram: current impl uses 2-provider split (cheap/primary), per-tier model mapping is target design - Add regression tests for both fixes Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Co-authored-by: Microwave <[email protected]> Co-authored-by: Joe <[email protected]> Co-authored-by: onlyamicrowave <[email protected]> Co-authored-by: Claude Opus 4.6 <[email protected]>
196 lines
6.2 KiB
Markdown
196 lines
6.2 KiB
Markdown
# Smart Model Routing for IronClaw
|
|
|
|
**Status:** Implemented
|
|
**Author:** Microwave
|
|
**Date:** 2026-02-19
|
|
|
|
## What
|
|
|
|
Automatic model selection based on request complexity. The router analyzes each user message and selects an appropriate model tier (flash/standard/pro/frontier), then maps that tier to a configured model.
|
|
|
|
## Why
|
|
|
|
1. **Cost optimization** — Simple requests ("hi", "what time is it") don't need expensive models
|
|
2. **User experience** — Simple requests return faster with lightweight models
|
|
3. **NEAR AI native** — Default backend uses NEAR AI inference where costs vary by model
|
|
4. **Zero-config value** — Users benefit immediately without configuration
|
|
5. **Not just power users** — Everyone gets smart defaults, power users can override
|
|
|
|
## How
|
|
|
|
### Architecture
|
|
|
|
```
|
|
User Message
|
|
│
|
|
▼
|
|
┌──────────────────┐
|
|
│ Pattern Overrides │ ← Fast-path for obvious cases (greetings, security audits)
|
|
└────────┬─────────┘
|
|
│ no match
|
|
▼
|
|
┌──────────────────┐
|
|
│ Complexity Scorer │ ← 13-dimension analysis
|
|
└────────┬─────────┘
|
|
│ score 0-100
|
|
▼
|
|
┌──────────────────┐
|
|
│ Tier Mapping │ ← 0-15: flash, 16-40: standard, 41-65: pro, 66+: frontier
|
|
└────────┬─────────┘
|
|
│ tier
|
|
▼
|
|
┌──────────────────┐
|
|
│ Model Selection │ ← Currently: cheap provider (Flash/Standard/Pro) vs primary (Frontier)
|
|
└────────┬─────────┘ Target: per-tier model mapping via config
|
|
│
|
|
▼
|
|
LLM Provider
|
|
```
|
|
|
|
### Complexity Scorer (13 Dimensions)
|
|
|
|
Each dimension produces a 0-100 score. Weighted sum determines total.
|
|
|
|
| Dimension | Weight | Signals |
|
|
|-----------|--------|---------|
|
|
| Reasoning Words | 14% | "why", "explain", "compare", "trade-offs" |
|
|
| Token Estimate | 12% | Prompt length |
|
|
| Code Indicators | 10% | Backticks, syntax, "implement", "PR" |
|
|
| Multi-Step | 10% | "first", "then", "after", "steps" |
|
|
| Domain Specific | 10% | Technical terms (configurable) |
|
|
| Creativity | 7% | "write", "summarize", "tweet", "blog" |
|
|
| Question Complexity | 7% | Multiple questions, open-ended starters |
|
|
| Precision | 6% | Numbers, "exactly", "calculate" |
|
|
| Ambiguity | 5% | Vague references |
|
|
| Context Dependency | 5% | "previous", "you said" |
|
|
| Sentence Complexity | 5% | Commas, conjunctions, clause depth |
|
|
| Tool Likelihood | 5% | "read", "deploy", "install" |
|
|
| Safety Sensitivity | 4% | "password", "auth", "vulnerability" |
|
|
|
|
**Multi-dimensional boost:** +30% when 3+ dimensions score above threshold.
|
|
|
|
### Tier Boundaries
|
|
|
|
| Score | Tier | Typical Use Case |
|
|
|-------|------|------------------|
|
|
| 0-15 | flash | Greetings, acknowledgments, quick lookups |
|
|
| 16-40 | standard | Writing, comparisons, defined tasks |
|
|
| 41-65 | pro | Multi-step analysis, code review |
|
|
| 66+ | frontier | Critical decisions, security audits |
|
|
|
|
### Pattern Overrides
|
|
|
|
Fast-path rules that bypass scoring for obvious cases:
|
|
|
|
```yaml
|
|
# Force flash tier
|
|
- "^(hi|hello|hey|thanks|ok|sure|yes|no)$"
|
|
- "^what.*(time|date|day)"
|
|
|
|
# Force frontier tier
|
|
- "security.*(audit|review|scan)"
|
|
- "vulnerabilit(y|ies).*(review|scan|check|audit)"
|
|
|
|
# Force pro tier
|
|
- "deploy.*(mainnet|production)"
|
|
```
|
|
|
|
### Configuration
|
|
|
|
> **Note:** The current implementation supports smart routing via
|
|
> `NEARAI_CHEAP_MODEL` and `SMART_ROUTING_CASCADE` env vars, plus
|
|
> `domain_keywords` on `SmartRoutingConfig`. The full `llm.routing` YAML
|
|
> schema below is the target design — not all knobs are wired yet.
|
|
|
|
**Default (zero-config):**
|
|
```yaml
|
|
llm:
|
|
routing:
|
|
enabled: true # default
|
|
```
|
|
|
|
**Power user overrides (target schema):**
|
|
```yaml
|
|
llm:
|
|
routing:
|
|
enabled: true
|
|
tiers:
|
|
flash: "claude-3-5-haiku-latest"
|
|
standard: "claude-sonnet-4-5-latest"
|
|
pro: "claude-sonnet-4-5-latest"
|
|
frontier: "claude-opus-4-5-latest"
|
|
thinking:
|
|
pro: "low"
|
|
frontier: "medium"
|
|
overrides:
|
|
- pattern: "my-custom-pattern"
|
|
tier: "pro"
|
|
domain_keywords: # Custom keywords for your domain
|
|
- "mycompany"
|
|
- "myproduct"
|
|
- "internal-tool"
|
|
```
|
|
|
|
If `domain_keywords` is not set, uses `DEFAULT_DOMAIN_KEYWORDS` which covers common web3/infra terms.
|
|
|
|
**Disable routing (pin model):**
|
|
```yaml
|
|
llm:
|
|
routing:
|
|
enabled: false
|
|
model: "claude-opus-4-5"
|
|
```
|
|
|
|
**Bring your own keys:**
|
|
```yaml
|
|
llm:
|
|
backend: anthropic
|
|
api_key: "sk-..."
|
|
routing:
|
|
enabled: true # still works with external providers
|
|
```
|
|
|
|
### Integration Points
|
|
|
|
1. **RoutingProvider** — New wrapper implementing `LlmProvider` trait (like `FailoverProvider`)
|
|
2. **Scorer** — Pure function, no I/O, fast (~1ms)
|
|
3. **Config schema** — Extend `LlmConfig` with `routing` section
|
|
4. **Telemetry** — Log routing decisions for observability
|
|
|
|
### Model Agnosticism
|
|
|
|
**Critical:** No hardcoded model names in the router logic itself.
|
|
|
|
- Tier→model mappings come from config
|
|
- Default mappings use `-latest` patterns where supported
|
|
- NEAR AI backend handles actual model resolution
|
|
- Router only knows about tiers
|
|
|
|
### Layers of Control
|
|
|
|
| Layer | User Type | Config |
|
|
|-------|-----------|--------|
|
|
| 1. Zero-config | Everyone | `routing.enabled: true` (default) |
|
|
| 2. Tier tuning | Power users | Custom `routing.tiers` mapping |
|
|
| 3. Pattern overrides | Power users | Custom `routing.overrides` |
|
|
| 4. Model pinning | Power users | `routing.enabled: false` + `model: X` |
|
|
| 5. Own API keys | Power users | `backend: anthropic` + `api_key` |
|
|
|
|
## Implementation Plan
|
|
|
|
1. [x] Port scorer to Rust (`src/llm/smart_routing.rs`)
|
|
2. [x] Implement router wrapper (`src/llm/smart_routing.rs`)
|
|
3. [x] Extend config schema (`src/config.rs`)
|
|
4. [x] Wire into provider creation (`src/llm/mod.rs`)
|
|
5. [x] Add telemetry/logging
|
|
6. [x] Tests with real conversation samples
|
|
7. [x] Codex + Gemini security review
|
|
8. [x] Documentation updated (this spec)
|
|
|
|
## Expected Outcomes
|
|
|
|
- **50-70% cost reduction** for typical usage patterns
|
|
- **Faster responses** for simple requests
|
|
- **Zero config required** for default benefits
|
|
- **Full control** for power users who want it
|