mirror of
https://github.com/outbackdingo/optimclaw.git
synced 2026-08-25 14:53:34 +00:00
* fix(ci): secrets can't be used in step if conditions [skip-regression-check] (#787) GitHub Actions step-level `if:` doesn't have access to `secrets` context. Replace `if: secrets.X != ''` with `continue-on-error: true` and let the Set token step handle the fallback. Co-authored-by: Claude Sonnet 4.6 <[email protected]> * fix(ci): clean up staging pipeline — remove hacks, skip redundant checks [skip-regression-check] (#794) - Remove continue-on-error from staging-ci.yml app token steps (secrets are configured) - Skip test.yml and code_style.yml on PRs targeting staging (staging-ci.yml already runs tests before promoting, promotion PR gets full CI on main) - Allow ironclaw-ci[bot] in Claude Code review for bot-created promotion PRs Co-authored-by: Claude Opus 4.6 <[email protected]> * fix(ci): run fmt + clippy on staging PRs, skip Windows clippy [skip-regression-check] (#802) - Remove branches:[main] filter from code_style.yml so it runs on all PRs - Gate clippy-windows with `if: github.base_ref == 'main'` (skip on staging PRs) - Update rollup job to allow skipped clippy-windows - Simplify claude-review.yml to only trigger on labeled event (avoids duplicate runs) Co-authored-by: Claude Opus 4.6 <[email protected]> * feat: persist user_id in save_job and expose job_id on routine runs (#709) * feat: persist worker events to DB and fix activity tab rendering In-process Worker (used by Scheduler::dispatch_job) now persists events via save_job_event at key execution points: plan creation, LLM responses, tool_use, tool_result, and job completion/failure/stuck. Event data shapes match the container worker format so the gateway activity tab renders them correctly. Frontend: tool_result errors now show a red X icon with danger styling instead of a silent empty output. The result event falls back to the error field when message is absent. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat: wire RoutineEngine into gateway for direct manual trigger firing Replace the message-channel hack in routines_trigger_handler with a direct call to RoutineEngine::fire_manual(), ensuring FullJob routines dispatch correctly when triggered from the web UI. Inject the engine into GatewayState from Agent::run after construction. Also persists user_id in save_job for both PG and libSQL backends, removes the source='sandbox' filter so all jobs are visible, and exposes job_id on RoutineRunInfo for the frontend job link. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: remove stale gateway_state argument from Agent::new test call sites The gateway_state parameter was removed from Agent::new during rebase (replaced by post-construction set_routine_engine_slot), but three test call sites still passed the extra None argument. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address PR review — restore sandbox source filter, remove blank lines - Revert removal of `source = 'sandbox'` filter in all SandboxStore queries (8 sites across PG and libSQL). Sandbox-specific APIs should stay scoped to sandbox jobs; unified job listing for the Jobs tab should use a separate query path. - Remove extra blank lines in agent_loop.rs and worker.rs that caused formatting CI failure. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address review — regenerate Cargo.lock, add user_id regression test - Regenerate Cargo.lock from main's lockfile to eliminate dependency version downgrades (anyhow, syn, etc.) that were churn from rebase. - Add regression test verifying user_id round-trips through save_job and get_job in the libSQL backend. Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: remove trailing blank line in libsql jobs.rs [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <[email protected]> * test: add Postgres-side regression test for user_id persistence in save_job Mirrors the existing libSQL test (test_save_job_persists_user_id) for the Postgres backend. Gated behind #[cfg(feature = "postgres")] + #[ignore] since it requires a running PostgreSQL instance (integration tier). Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Co-authored-by: Claude Opus 4.6 <[email protected]> * refactor: unify three agentic loops into single AgenticLoop engine (#654) Replace three independent copy-pasted agentic loops (dispatcher, worker, container runtime) with a single shared engine in `agentic_loop.rs` that all consumers customize via the `LoopDelegate` trait. Phase 1 — Shared engine (`src/agent/agentic_loop.rs`, 205 lines): - `run_agentic_loop()` owns the core LLM → tool exec → repeat cycle - `LoopDelegate` trait (Send + Sync, &dyn dispatch) with 6 hook points - Tool intent nudge logic consolidated (was duplicated in 3 files) - Iteration limit + force-text behavior preserved Phase 2 — Three delegate implementations: - `ChatDelegate` (dispatcher.rs): 3-phase approval flow, hooks, cost guard, context compaction, skill attenuation, interruption - `JobDelegate` (worker/job.rs): planning pre-loop phase, parallel JoinSet exec, mark_completed/stuck/failed, SSE streaming, self-repair - `ContainerDelegate` (worker/container.rs): sequential tool exec, HTTP-proxied LLM, container-safe tools, credential injection Phase 3 — File moves and cleanup: - Delete `src/agent/worker.rs` — job logic moved to `src/worker/job.rs` - Rename `src/worker/runtime.rs` → `src/worker/container.rs` - Re-export `Worker`/`WorkerDeps` from `crate::worker` in `agent/mod.rs` - Update `scheduler.rs` imports to new worker location Shared helpers (`src/tools/execute.rs`): - `execute_tool_with_safety()` replaces 4 copies of validate → timeout → execute → serialize - `process_tool_result()` replaces 3 copies of sanitize → wrap → ChatMessage (also used by thread_ops.rs approval resume paths) Net result: -2,408 lines, zero duplicated loop logic, single code path for tool intent nudge and completion detection. Closes #654 Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address review feedback from Copilot 1. scheduler.rs: Replace `unwrap_or` fallback with proper error propagation when parsing tool output JSON — surfaces bugs instead of silently changing the output type. 2. worker/job.rs: Drop MutexGuard before the cancellation `.await` in `check_signals()` to avoid holding a lock across an async I/O call (prevents `await_holding_lock` lint). 3. worker/job.rs: Restore consecutive rate-limit counter (MAX_CONSECUTIVE_RATE_LIMITS = 10) so sustained rate limiting marks the job stuck with "Persistent rate limiting" instead of silently burning through max_iterations. Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: incorporate staging changes — token budget tracking + mark_failed Merge staging's changes into the refactored JobDelegate: - Add token budget tracking in call_llm (update_context/add_tokens) - mark_stuck → mark_failed for iteration cap and rate-limit exhaustion (aligns with staging's #788 fix) Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address zmanian's PR review — eliminate type erasure, clean up Address all 6 review points from zmanian on PR #800: 1. Replace LoopOutcome::Custom(Box<dyn Any>) with typed LoopOutcome::NeedApproval(Box<PendingApproval>) — eliminates type erasure and downcast, resolves clippy large_enum_variant. 2. Remove dead max_tool_iterations field from ChatDelegate struct. 3. Add on_tool_intent_nudge() hook to LoopDelegate trait with implementations in Job and Container delegates for observability. 4. Fix SSE events in job worker to emit raw sanitized content instead of XML-wrapped <tool_output> tags. 5. Remove 4 duplicate completion tests from job.rs that were already covered by the shared util module. 6. Avoid logging full tool results — use result_size_bytes in debug logs (execute.rs, job.rs). Also updates path references in CLAUDE.md, COVERAGE_PLAN.md, and add-sse-event.md command. Co-Authored-By: Claude Opus 4.6 <[email protected]> * feat(doctor): expand diagnostics from 7 to 16 health checks * test: add unit tests for agentic_loop and execute shared modules Add 16 tests covering the two new critical shared modules: agentic_loop.rs (10 tests): - Text response exits loop immediately - Tool call → text response continuation - LoopSignal::Stop exits before LLM call - LoopSignal::InjectMessage adds user message to context - Max iterations terminates with LoopOutcome::MaxIterations - Tool intent nudge fires twice then caps - before_llm_call early exit bypasses LLM - truncate_for_preview: short string, long string, multibyte safety execute.rs (6 tests): - execute_tool_with_safety success path - Missing tool returns ToolError::NotFound - Tool execution failure propagates - Per-tool timeout enforcement (50ms) - process_tool_result XML wrapping on success - process_tool_result error formatting All 2,777 unit tests pass, 0 clippy warnings. Co-Authored-By: Claude Opus 4.6 <[email protected]> * style: cargo fmt Co-Authored-By: Claude Opus 4.6 <[email protected]> * fix: address code review — 9 issues across agentic loop, job worker, container CRITICAL fixes: - Rate-limit exhaustion now returns Err(LlmError::RateLimited) instead of Ok(Text("")), stopping the loop immediately with no ghost iteration. Below-threshold retries still use Text("") with an explicit empty-string guard in handle_text_response to skip injection. - check_signals drains the entire message channel before returning, prioritizing Stop over UserMessage. Previously returned early on first UserMessage, silently dropping any queued Stop or additional messages. - check_signals now detects all non-progressing job states (Cancelled, Failed, Stuck, Completed, Submitted, Accepted) instead of only Cancelled and Failed. HIGH fixes: - Error path in process_tool_result_job applies truncate_for_preview to bound error strings in SSE/DB events (was unbounded). - Document Send+Sync lifetime constraint on LoopDelegate trait. - Test mock before_llm_call refactored from double-lock to single lock acquisition, eliminating deadlock risk on refactor. MEDIUM fixes: - CompletionReport includes actual iteration count via shared Arc<Mutex<u32>> tracker (was hardcoded 0). - process_tool_result_job return type changed from Result<bool> to Result<()> — the bool was always false (dead API). - Deduplicate truncate in container.rs; now uses truncate_for_preview from agentic_loop. Verified: 0 clippy warnings, 2781 tests pass, cargo fmt clean. Co-Authored-By: Claude Opus 4.6 <[email protected]> --------- Co-authored-by: Henry Park <[email protected]> Co-authored-by: Claude Sonnet 4.6 <[email protected]> Co-authored-by: Illia Polosukhin <[email protected]> Co-authored-by: Umesh Kumar Singh <[email protected]> Co-authored-by: reidliu41 <[email protected]>
241 lines
13 KiB
Markdown
241 lines
13 KiB
Markdown
# IronClaw Development Guide
|
|
|
|
**IronClaw** is a secure personal AI assistant — user-first security, self-expanding tools, defense in depth, multi-channel access with proactive background execution.
|
|
|
|
## Build & Test
|
|
|
|
```bash
|
|
cargo fmt # format
|
|
cargo clippy --all --benches --tests --examples --all-features # lint (zero warnings)
|
|
cargo test # unit tests
|
|
cargo test --features integration # + PostgreSQL tests
|
|
RUST_LOG=ironclaw=debug cargo run # run with logging
|
|
```
|
|
|
|
E2E tests: see `tests/e2e/CLAUDE.md`.
|
|
|
|
## Code Style
|
|
|
|
- Prefer `crate::` for cross-module imports; `super::` is fine in tests and intra-module refs
|
|
- No `pub use` re-exports unless exposing to downstream consumers
|
|
- No `.unwrap()` or `.expect()` in production code (tests are fine)
|
|
- Use `thiserror` for error types in `error.rs`
|
|
- Map errors with context: `.map_err(|e| SomeError::Variant { reason: e.to_string() })?`
|
|
- Prefer strong types over strings (enums, newtypes)
|
|
- Keep functions focused, extract helpers when logic is reused
|
|
- Comments for non-obvious logic only
|
|
|
|
## Architecture
|
|
|
|
Prefer generic/extensible architectures over hardcoding specific integrations. Ask clarifying questions about the desired abstraction level before implementing.
|
|
|
|
Key traits for extensibility: `Database`, `Channel`, `Tool`, `LlmProvider`, `SuccessEvaluator`, `EmbeddingProvider`, `NetworkPolicyDecider`, `Hook`, `Observer`, `Tunnel`.
|
|
|
|
All I/O is async with tokio. Use `Arc<T>` for shared state, `RwLock` for concurrent access.
|
|
|
|
## Project Structure
|
|
|
|
```
|
|
src/
|
|
├── lib.rs # Library root, module declarations
|
|
├── main.rs # Entry point, CLI args, startup
|
|
├── app.rs # App startup orchestration (channel wiring, DB init)
|
|
├── bootstrap.rs # Base directory resolution (~/.ironclaw), early .env loading
|
|
├── settings.rs # User settings persistence (~/.ironclaw/settings.json)
|
|
├── service.rs # OS service management (launchd/systemd daemon install)
|
|
├── tracing_fmt.rs # Custom tracing formatter
|
|
├── util.rs # Shared utilities
|
|
├── config/ # Configuration from env vars (split by subsystem)
|
|
│ ├── mod.rs # Re-exports all config types; top-level Config struct
|
|
│ ├── agent.rs, llm.rs, channels.rs, database.rs, sandbox.rs, skills.rs
|
|
│ ├── heartbeat.rs, routines.rs, safety.rs, embeddings.rs, wasm.rs
|
|
│ ├── tunnel.rs # Tunnel provider config (TUNNEL_PROVIDER, TUNNEL_URL, etc.)
|
|
│ └── secrets.rs, hygiene.rs, builder.rs, helpers.rs
|
|
├── error.rs # Error types (thiserror)
|
|
│
|
|
├── agent/ # Core agent loop, dispatcher, scheduler, sessions — see src/agent/CLAUDE.md
|
|
│
|
|
├── channels/ # Multi-channel input
|
|
│ ├── channel.rs # Channel trait, IncomingMessage, OutgoingResponse
|
|
│ ├── manager.rs # ChannelManager merges streams
|
|
│ ├── cli/ # Full TUI with Ratatui
|
|
│ ├── http.rs # HTTP webhook (axum) with secret validation
|
|
│ ├── webhook_server.rs # Unified HTTP server composing all webhook routes
|
|
│ ├── repl.rs # Simple REPL (for testing)
|
|
│ ├── web/ # Web gateway (browser UI) — see src/channels/web/CLAUDE.md
|
|
│ └── wasm/ # WASM channel runtime
|
|
│ ├── mod.rs
|
|
│ ├── bundled.rs # Bundled channel discovery
|
|
│ ├── capabilities.rs # Channel-specific capabilities (HTTP endpoint, emit rate)
|
|
│ ├── error.rs # WASM channel error types
|
|
│ ├── runtime.rs # WASM channel execution runtime
|
|
│ ├── setup.rs # WasmChannelSetup, setup_wasm_channels(), inject_channel_credentials()
|
|
│ └── wrapper.rs # Channel trait wrapper for WASM modules
|
|
│
|
|
├── cli/ # CLI subcommands (clap)
|
|
│ ├── mod.rs # Cli struct, Command enum (run/onboard/config/tool/registry/mcp/memory/pairing/service/doctor/status/completion)
|
|
│ └── config.rs, tool.rs, registry.rs, mcp.rs, memory.rs, pairing.rs, service.rs, doctor.rs, status.rs, completion.rs
|
|
│
|
|
├── registry/ # Extension registry catalog
|
|
│ ├── manifest.rs # ExtensionManifest, ArtifactSpec, BundleDefinition types
|
|
│ ├── catalog.rs # RegistryCatalog: load from filesystem and embedded JSON
|
|
│ └── installer.rs # RegistryInstaller: download, verify, install WASM artifacts
|
|
│
|
|
├── hooks/ # Lifecycle hooks (6 points: BeforeInbound, BeforeToolCall, BeforeOutbound, OnSessionStart, OnSessionEnd, TransformResponse)
|
|
│
|
|
├── tunnel/ # Tunnel abstraction for public internet exposure
|
|
│ ├── mod.rs # Tunnel trait, TunnelProviderConfig, create_tunnel(), start_managed_tunnel()
|
|
│ ├── cloudflare.rs # CloudflareTunnel (cloudflared binary)
|
|
│ ├── ngrok.rs # NgrokTunnel
|
|
│ ├── tailscale.rs # TailscaleTunnel (serve/funnel modes)
|
|
│ ├── custom.rs # CustomTunnel (arbitrary command with {host}/{port})
|
|
│ └── none.rs # NoneTunnel (local-only, no exposure)
|
|
│
|
|
├── observability/ # Pluggable event/metric recording (noop, log, multi)
|
|
│
|
|
├── orchestrator/ # Internal HTTP API for sandbox containers
|
|
│ ├── api.rs # Axum endpoints (LLM proxy, events, prompts)
|
|
│ ├── auth.rs # Per-job bearer token store
|
|
│ └── job_manager.rs # Container lifecycle (create, stop, cleanup)
|
|
│
|
|
├── worker/ # Runs inside Docker containers
|
|
│ ├── container.rs # Container worker runtime (ContainerDelegate + shared agentic loop)
|
|
│ ├── job.rs # Background job worker (JobDelegate + shared agentic loop)
|
|
│ ├── claude_bridge.rs # Claude Code bridge (spawns claude CLI)
|
|
│ └── proxy_llm.rs # LlmProvider that proxies through orchestrator
|
|
│
|
|
├── safety/ # Prompt injection defense
|
|
│ ├── sanitizer.rs # Pattern detection, content escaping
|
|
│ ├── validator.rs # Input validation (length, encoding, patterns)
|
|
│ ├── policy.rs # PolicyRule system with severity/actions
|
|
│ ├── leak_detector.rs # Secret detection (API keys, tokens, etc.)
|
|
│ └── credential_detect.rs # HTTP request credential detection
|
|
│
|
|
├── llm/ # Multi-provider LLM integration — see src/llm/CLAUDE.md
|
|
│
|
|
├── tools/ # Extensible tool system
|
|
│ ├── tool.rs # Tool trait, ToolOutput, ToolError
|
|
│ ├── registry.rs # ToolRegistry for discovery
|
|
│ ├── rate_limiter.rs # Shared sliding-window rate limiter
|
|
│ ├── builtin/ # Built-in tools (echo, time, json, http, web_fetch, file, shell, memory, message, job, routine, extension_tools, skill_tools, secrets_tools)
|
|
│ ├── builder/ # Dynamic tool building
|
|
│ │ ├── core.rs # BuildRequirement, SoftwareType, Language
|
|
│ │ ├── templates.rs # Project scaffolding
|
|
│ │ ├── testing.rs # Test harness integration
|
|
│ │ └── validation.rs # WASM validation
|
|
│ ├── mcp/ # Model Context Protocol
|
|
│ │ ├── client.rs # MCP client over HTTP
|
|
│ │ ├── factory.rs # create_client_from_config() — transport dispatch factory
|
|
│ │ ├── protocol.rs # JSON-RPC types
|
|
│ │ └── session.rs # MCP session management (Mcp-Session-Id header, per-server state)
|
|
│ └── wasm/ # Full WASM sandbox (wasmtime)
|
|
│ ├── runtime.rs # Module compilation and caching
|
|
│ ├── wrapper.rs # Tool trait wrapper for WASM modules
|
|
│ ├── host.rs # Host functions (logging, time, workspace)
|
|
│ ├── limits.rs # Fuel metering and memory limiting
|
|
│ ├── allowlist.rs # Network endpoint allowlisting
|
|
│ ├── credential_injector.rs # Safe credential injection
|
|
│ ├── loader.rs # WASM tool discovery from filesystem
|
|
│ ├── rate_limiter.rs # Per-tool rate limiting
|
|
│ ├── error.rs # WASM-specific error types
|
|
│ └── storage.rs # Linear memory persistence
|
|
│
|
|
├── db/ # Dual-backend persistence (PostgreSQL + libSQL) — see src/db/CLAUDE.md
|
|
│
|
|
├── workspace/ # Persistent memory system — see src/workspace/README.md
|
|
│
|
|
├── context/ # Job context isolation (JobState, JobContext, ContextManager)
|
|
├── estimation/ # Cost/time/value estimation with EMA learning
|
|
├── evaluation/ # Success evaluation (rule-based, LLM-based)
|
|
│
|
|
├── sandbox/ # Docker execution sandbox
|
|
│ ├── config.rs # SandboxConfig, SandboxPolicy enum (ReadOnly/WorkspaceWrite/FullAccess)
|
|
│ ├── manager.rs # SandboxManager orchestration
|
|
│ ├── container.rs # ContainerRunner, Docker lifecycle
|
|
│ └── proxy/ # Network proxy: domain allowlist, credential injection, CONNECT tunnel
|
|
│
|
|
├── secrets/ # Secrets management (AES-256-GCM, OS keychain for master key)
|
|
│
|
|
├── setup/ # 7-step onboarding wizard — see src/setup/README.md
|
|
│
|
|
├── skills/ # SKILL.md prompt extension system — see .claude/rules/skills.md
|
|
│
|
|
└── history/ # Persistence (PostgreSQL repositories, analytics)
|
|
|
|
tests/
|
|
├── *.rs # Integration tests (workspace, heartbeat, WS gateway, pairing, etc.)
|
|
├── test-pages/ # HTML→Markdown conversion fixtures
|
|
└── e2e/ # Python/Playwright E2E scenarios (see tests/e2e/CLAUDE.md)
|
|
```
|
|
|
|
## Database
|
|
|
|
Dual-backend: PostgreSQL + libSQL/Turso. **All new persistence features must support both backends.** See `src/db/CLAUDE.md` and `.claude/rules/database.md`.
|
|
|
|
## Module Specs
|
|
|
|
When modifying a module with a spec, read the spec first. Code follows spec; spec is the tiebreaker.
|
|
|
|
**Module-owned initialization:** Module-specific initialization logic (database connection, transport creation, channel setup) must live in the owning module as a public factory function — not in `main.rs` or `app.rs`. These entry-point files orchestrate calls to module factories. Feature-flag branching (`#[cfg(feature = ...)]`) must be confined to the module that owns the abstraction.
|
|
|
|
| Module | Spec |
|
|
|--------|------|
|
|
| `src/agent/` | `src/agent/CLAUDE.md` |
|
|
| `src/channels/web/` | `src/channels/web/CLAUDE.md` |
|
|
| `src/db/` | `src/db/CLAUDE.md` |
|
|
| `src/llm/` | `src/llm/CLAUDE.md` |
|
|
| `src/setup/` | `src/setup/README.md` |
|
|
| `src/tools/` | `src/tools/README.md` |
|
|
| `src/workspace/` | `src/workspace/README.md` |
|
|
| `tests/e2e/` | `tests/e2e/CLAUDE.md` |
|
|
|
|
## Job State Machine
|
|
|
|
```
|
|
Pending -> InProgress -> Completed -> Submitted -> Accepted
|
|
\-> Failed
|
|
\-> Stuck -> InProgress (recovery)
|
|
\-> Failed
|
|
```
|
|
|
|
## Skills System
|
|
|
|
SKILL.md files extend the agent's prompt with domain-specific instructions. See `.claude/rules/skills.md` for full details.
|
|
|
|
- **Trust model**: Trusted (user-placed in `~/.ironclaw/skills/` or workspace `skills/`, full tool access) vs Installed (registry, read-only tools)
|
|
- **Selection pipeline**: gating (check bin/env/config requirements) -> scoring (keywords/patterns/tags) -> budget (fit within `SKILLS_MAX_TOKENS`) -> attenuation (trust-based tool ceiling)
|
|
- **Skill tools**: `skill_list`, `skill_search`, `skill_install`, `skill_remove`
|
|
|
|
## Configuration
|
|
|
|
See `.env.example` for all environment variables. LLM backends (`nearai`, `openai`, `anthropic`, `ollama`, `openai_compatible`, `tinfoil`, `bedrock`) documented in `src/llm/CLAUDE.md`.
|
|
|
|
## Adding a New Channel
|
|
|
|
1. Create `src/channels/my_channel.rs`
|
|
2. Implement the `Channel` trait
|
|
3. Add config in `src/config/channels.rs`
|
|
4. Wire up in `src/app.rs` channel setup section
|
|
|
|
## Workspace & Memory
|
|
|
|
Persistent memory with hybrid search (FTS + vector via RRF). Four tools: `memory_search`, `memory_write`, `memory_read`, `memory_tree`. Identity files (AGENTS.md, SOUL.md, USER.md, IDENTITY.md) injected into system prompt. Heartbeat system runs proactive periodic execution (default: 30 minutes), reading `HEARTBEAT.md` and notifying via channel if findings. See `src/workspace/README.md`.
|
|
|
|
## Debugging
|
|
|
|
```bash
|
|
RUST_LOG=ironclaw=trace cargo run # verbose
|
|
RUST_LOG=ironclaw::agent=debug cargo run # agent module only
|
|
RUST_LOG=ironclaw=debug,tower_http=debug cargo run # + HTTP request logging
|
|
```
|
|
|
|
## Current Limitations
|
|
|
|
1. Domain-specific tools (`marketplace.rs`, `restaurant.rs`, etc.) are stubs
|
|
2. Integration tests need testcontainers for PostgreSQL
|
|
3. MCP: no streaming support; stdio/HTTP/Unix transports all use request-response
|
|
4. WIT bindgen: auto-extract tool schema from WASM is stubbed
|
|
5. Built tools get empty capabilities; need UX for granting access
|
|
6. No tool versioning or rollback
|
|
7. Observability: only `log` and `noop` backends (no OpenTelemetry)
|