Files
optimclaw/docs/smart-routing-spec.md
T
69cddb10fd feat: integrate 13-dimension complexity scorer into smart routing (#529)
* feat(llm): add smart model routing based on request complexity

Automatically selects optimal model tier (flash/standard/pro/frontier) for each
request based on 13-dimension complexity scoring:

- Reasoning words, multi-step signals, code indicators
- Domain-specific terms, creativity, precision
- Safety sensitivity, tool likelihood, question complexity
- Token estimate, context dependency, sentence complexity

Features:
- Pattern overrides for fast-path routing (greetings → flash, security audits → frontier)
- Configurable tier-to-model mappings (defaults to -latest aliases)
- Thinking mode per tier (pro: low, frontier: medium)
- User-configurable pattern overrides
- Zero-config for default benefits, full control for power users

Expected cost savings: 50-70% vs always-using-frontier baseline.

Refs: smart-routing-spec.md

* fix(routing): address Gemini Code Assist review feedback

- Add tracing warnings for invalid tier/regex in user overrides (router.rs)
- Use unreachable!() for tier hint match since regex enforces valid tiers (scorer.rs)
- Refactor weighted total to array iteration for maintainability (scorer.rs)
- Add TODO for making domain keywords configurable (scorer.rs)

Refs: PR #208

* feat(routing): make domain keywords configurable

- Add ScorerConfig with optional domain_keywords field
- Add DEFAULT_DOMAIN_KEYWORDS constant (exported for reference)
- Add domain_keywords to RouterConfig for top-level configuration
- Build domain regex at runtime from config, fallback to defaults
- Add score_complexity_with_config() function
- Add test for custom domain keywords

Users can now provide project-specific keywords:

  RouterConfig {
      domain_keywords: Some(vec!["mycompany".into(), "myproduct".into()]),
      ..Default::default()
  }

Addresses Gemini Code Assist review feedback on PR #208.

Tests: 20/20 passing

* docs: add domain_keywords to routing config example

* feat: integrate 13-dimension complexity scorer into smart routing (takeover #208)

Folds the 13-dimension complexity scorer and pattern overrides from PR #208
into the existing SmartRoutingProvider, replacing the simpler keyword-based
classifier. Adds 4-tier system (Flash/Standard/Pro/Frontier), configurable
scorer weights, domain keywords, regex pattern overrides, tier hints, and
multi-dimensional boost. Removes separate routing/ directory and lazy_static
dependency in favor of std::sync::LazyLock. Includes 44 tests covering all
scoring dimensions, tier boundaries, pattern overrides, and provider routing.

Co-Authored-By: onlyamicrowave <[email protected]>
Co-Authored-By: Claude Opus 4.6 <[email protected]>

* fix: address review feedback on smart routing PR (#529)

- Cache compiled domain regex in SmartRoutingProvider (built once at
  construction, not per-request) and add score_complexity_with_regex() API
- Check explicit tier hints before pattern overrides so user intent wins
  (e.g. "[tier:flash] security audit" routes as Flash, not Frontier)
- Trim input before matching/scoring so trailing whitespace doesn't break
  anchored override regexes or skew token-length scoring
- Fix token estimate comment (>=520 chars = 100, not >500)
- Update spec: check implementation plan boxes, fix file paths, add note
  that llm.routing YAML schema is target design (current config uses env vars)
- Add regression tests for tier hint precedence and trimmed greeting matching

Co-Authored-By: Claude Opus 4.6 <[email protected]>

* fix: restore Cargo.lock from main to fix html_to_markdown test

The lockfile was fully regenerated during the PR #208 merge conflict
resolution, which bumped html-to-markdown-rs from 2.25.1 to 2.27.2.
The new version produces different output that breaks the golden-file
snapshot test. Restore the original lockfile from main — lazy_static
was never in main's lockfile, so no further changes needed.

Co-Authored-By: Claude Opus 4.6 <[email protected]>

* fix: address second round of review feedback (#529)

- Tighten quick-lookup override regex with end anchor to prevent matching
  complex questions like "What time complexity is merge sort?"
- Handle empty domain keywords list by falling back to defaults instead of
  producing a broken regex that matches empty strings everywhere
- Clarify spec architecture diagram: current impl uses 2-provider split
  (cheap/primary), per-tier model mapping is target design
- Add regression tests for both fixes

Co-Authored-By: Claude Opus 4.6 <[email protected]>

---------

Co-authored-by: Microwave <[email protected]>
Co-authored-by: Joe <[email protected]>
Co-authored-by: onlyamicrowave <[email protected]>
Co-authored-by: Claude Opus 4.6 <[email protected]>
2026-03-05 09:14:07 +00:00

6.2 KiB

Smart Model Routing for IronClaw

Status: Implemented Author: Microwave Date: 2026-02-19

What

Automatic model selection based on request complexity. The router analyzes each user message and selects an appropriate model tier (flash/standard/pro/frontier), then maps that tier to a configured model.

Why

  1. Cost optimization — Simple requests ("hi", "what time is it") don't need expensive models
  2. User experience — Simple requests return faster with lightweight models
  3. NEAR AI native — Default backend uses NEAR AI inference where costs vary by model
  4. Zero-config value — Users benefit immediately without configuration
  5. Not just power users — Everyone gets smart defaults, power users can override

How

Architecture

User Message
     │
     ▼
┌──────────────────┐
│ Pattern Overrides │  ← Fast-path for obvious cases (greetings, security audits)
└────────┬─────────┘
         │ no match
         ▼
┌──────────────────┐
│ Complexity Scorer │  ← 13-dimension analysis
└────────┬─────────┘
         │ score 0-100
         ▼
┌──────────────────┐
│   Tier Mapping   │  ← 0-15: flash, 16-40: standard, 41-65: pro, 66+: frontier
└────────┬─────────┘
         │ tier
         ▼
┌──────────────────┐
│  Model Selection │  ← Currently: cheap provider (Flash/Standard/Pro) vs primary (Frontier)
└────────┬─────────┘    Target: per-tier model mapping via config
         │
         ▼
    LLM Provider

Complexity Scorer (13 Dimensions)

Each dimension produces a 0-100 score. Weighted sum determines total.

Dimension Weight Signals
Reasoning Words 14% "why", "explain", "compare", "trade-offs"
Token Estimate 12% Prompt length
Code Indicators 10% Backticks, syntax, "implement", "PR"
Multi-Step 10% "first", "then", "after", "steps"
Domain Specific 10% Technical terms (configurable)
Creativity 7% "write", "summarize", "tweet", "blog"
Question Complexity 7% Multiple questions, open-ended starters
Precision 6% Numbers, "exactly", "calculate"
Ambiguity 5% Vague references
Context Dependency 5% "previous", "you said"
Sentence Complexity 5% Commas, conjunctions, clause depth
Tool Likelihood 5% "read", "deploy", "install"
Safety Sensitivity 4% "password", "auth", "vulnerability"

Multi-dimensional boost: +30% when 3+ dimensions score above threshold.

Tier Boundaries

Score Tier Typical Use Case
0-15 flash Greetings, acknowledgments, quick lookups
16-40 standard Writing, comparisons, defined tasks
41-65 pro Multi-step analysis, code review
66+ frontier Critical decisions, security audits

Pattern Overrides

Fast-path rules that bypass scoring for obvious cases:

# Force flash tier
- "^(hi|hello|hey|thanks|ok|sure|yes|no)$"
- "^what.*(time|date|day)"

# Force frontier tier
- "security.*(audit|review|scan)"
- "vulnerabilit(y|ies).*(review|scan|check|audit)"

# Force pro tier
- "deploy.*(mainnet|production)"

Configuration

Note: The current implementation supports smart routing via NEARAI_CHEAP_MODEL and SMART_ROUTING_CASCADE env vars, plus domain_keywords on SmartRoutingConfig. The full llm.routing YAML schema below is the target design — not all knobs are wired yet.

Default (zero-config):

llm:
  routing:
    enabled: true  # default

Power user overrides (target schema):

llm:
  routing:
    enabled: true
    tiers:
      flash: "claude-3-5-haiku-latest"
      standard: "claude-sonnet-4-5-latest"
      pro: "claude-sonnet-4-5-latest"
      frontier: "claude-opus-4-5-latest"
    thinking:
      pro: "low"
      frontier: "medium"
    overrides:
      - pattern: "my-custom-pattern"
        tier: "pro"
    domain_keywords:  # Custom keywords for your domain
      - "mycompany"
      - "myproduct"
      - "internal-tool"

If domain_keywords is not set, uses DEFAULT_DOMAIN_KEYWORDS which covers common web3/infra terms.

Disable routing (pin model):

llm:
  routing:
    enabled: false
  model: "claude-opus-4-5"

Bring your own keys:

llm:
  backend: anthropic
  api_key: "sk-..."
  routing:
    enabled: true  # still works with external providers

Integration Points

  1. RoutingProvider — New wrapper implementing LlmProvider trait (like FailoverProvider)
  2. Scorer — Pure function, no I/O, fast (~1ms)
  3. Config schema — Extend LlmConfig with routing section
  4. Telemetry — Log routing decisions for observability

Model Agnosticism

Critical: No hardcoded model names in the router logic itself.

  • Tier→model mappings come from config
  • Default mappings use -latest patterns where supported
  • NEAR AI backend handles actual model resolution
  • Router only knows about tiers

Layers of Control

Layer User Type Config
1. Zero-config Everyone routing.enabled: true (default)
2. Tier tuning Power users Custom routing.tiers mapping
3. Pattern overrides Power users Custom routing.overrides
4. Model pinning Power users routing.enabled: false + model: X
5. Own API keys Power users backend: anthropic + api_key

Implementation Plan

  1. Port scorer to Rust (src/llm/smart_routing.rs)
  2. Implement router wrapper (src/llm/smart_routing.rs)
  3. Extend config schema (src/config.rs)
  4. Wire into provider creation (src/llm/mod.rs)
  5. Add telemetry/logging
  6. Tests with real conversation samples
  7. Codex + Gemini security review
  8. Documentation updated (this spec)

Expected Outcomes

  • 50-70% cost reduction for typical usage patterns
  • Faster responses for simple requests
  • Zero config required for default benefits
  • Full control for power users who want it