mirror of
https://github.com/outbackdingo/optimclaw.git
synced 2026-08-25 14:53:34 +00:00
feat: chat onboarding and routine advisor (#927)
* feat: port NPA psychographic profiling system into IronClaw
Port the complete psychographic profiling system from NPA into IronClaw,
including enriched profile schema, conversational onboarding, profile
evolution, and three-tier prompt augmentation.
Personal onboarding moved from wizard Step 9 to first assistant
interaction per maintainer feedback — the First Contact system prompt
block now instructs the LLM to conduct a natural onboarding conversation
that builds the psychographic profile via memory_write.
Changes:
- Enrich profile.rs with 5 new structs, 9-dimension analysis framework,
custom deserializers for backward compatibility, and rendering methods
- Add conversational onboarding engine with one-step-removed questioning
technique, personality framework, and confidence-scored profile generation
- Add profile evolution with confidence gating, analysis metadata tracking,
and weekly update routine
- Replace thin interaction style injection with three-tier system gated on
confidence > 0.6 and profile recency
- Replace wizard Step 9 with First Contact system prompt block that drives
conversational onboarding during the user's first interaction
- Add autonomy progression to SOUL.md seed and personality framework to
AGENTS.md seed
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* feat: replace chat-based onboarding with bootstrap greeting and workspace seeds
Remove the interactive onboarding_chat.rs engine in favor of a simpler
bootstrap flow: fresh workspaces get a proactive LLM greeting that
naturally profiles the user. Identity files are now seeded from
src/workspace/seeds/ instead of being hardcoded. Also removes the
identity-file write protection (seeds are now managed), adds routine
advisor integration, and includes an e2e trace for bootstrap greeting.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* feat(safety): sanitize identity file writes via Sanitizer to prevent prompt injection
Identity files (SOUL.md, AGENTS.md, USER.md, IDENTITY.md) are injected into
every system prompt. Rather than hard-blocking writes (which broke onboarding),
scan content through the existing Sanitizer and reject writes with High/Critical
severity injection patterns. Medium/Low warnings are logged but allowed.
Also clarifies AGENTS.md identity file roles (USER.md = user info, IDENTITY.md =
agent identity) and adds IDENTITY.md setup as an explicit bootstrap step.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* docs: update profile_onboarding_completed comment to reflect current wiring
The field is now actively used by the agent loop to suppress BOOTSTRAP.md
injection — remove the stale "not yet wired" TODO.
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(setup): use env_or_override for NEARAI_API_KEY in model fetch config
When the user authenticates via NEAR AI Cloud API key (option 4),
api_key_login() stores the key via set_runtime_env(). But
build_nearai_model_fetch_config() was using std::env::var() which
doesn't check the runtime overlay — so model listing fell back to
session-token auth and re-triggered the interactive NEAR AI
authentication menu.
Switch to env_or_override() which checks both real env vars and the
runtime overlay.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(agent): correct channel/user_id in bootstrap greeting persist call
persist_assistant_response was called with channel="default",
user_id="system" but the assistant thread was created via
get_or_create_assistant_conversation("default", "gateway") which owns
the conversation as user_id="default", channel="gateway". The mismatch
caused ensure_writable_conversation to reject the write with:
WARN Rejected write for unavailable thread id user=system channel=default
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(web): remove all inline event handlers for CSP compliance
The Content-Security-Policy header (added in f48fe95) blocks inline JS
via script-src 'self'. All onclick/onchange attributes in index.html
are replaced with getElementById().addEventListener() calls. Dynamic
inline handlers in app.js (jobs, routines, memory breadcrumb, code
blocks, TEE report) are replaced with data-action attributes and a
single delegated click handler on document.
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(agent): align bootstrap message user/channel and update fixture schema field
- Bootstrap IncomingMessage now uses ("default", "gateway") consistently
with persist and session registration calls
- Update bootstrap_greeting.json fixture: schema_version → version to
match current PROFILE_JSON_SCHEMA
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* style: cargo fmt
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(safety): address PR review — expand injection scanning and harden profile sync
- BOOTSTRAP.md: fix target "profile" → "context/profile.json" so the
write hits the correct path and triggers profile sync
- IDENTITY_FILES: add context/assistant-directives.md to the scanned
set since it is also injected into the system prompt
- sync_profile_documents(): scan derived USER.md and assistant-directives
content through Sanitizer before writing, rejecting High/Critical
injection patterns
- profile_evolution_prompt(): wrap recent_messages_summary in <user_data>
delimiters with untrusted-data instruction to mitigate indirect
prompt injection
- routine-advisor skill: update cron examples from 6-field to standard
5-field format for consistency with routine_create tool docs
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* style: cargo fmt
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(setup): detect env-provided LLM keys during quick-mode onboarding
Quick-mode wizard now checks LLM_BACKEND, NEARAI_API_KEY,
ANTHROPIC_API_KEY, and OPENAI_API_KEY env vars to pre-populate
the provider setting, so users aren't re-prompted for credentials
they already supplied. Also teaches setup_nearai() to recognize
NEARAI_API_KEY from env (previously only checked session tokens).
Includes web UI cleanup (remove duplicate event listeners) and
e2e test response count adjustment.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix(test): update routine_create_list to expect 7-field normalized cron
The cron normalizer now always expands to 7-field format, so the
stored schedule is "0 0 9 * * * *" not "0 0 9 * * *".
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* feat(setup): skip LLM provider prompts when NEARAI_API_KEY is present
In quick mode, if NEARAI_API_KEY is set in the environment and the
backend was auto-detected as nearai, skip the interactive inference
provider and model selection steps. The API key is persisted to the
secrets store and a default model is set automatically.
Also simplify the static fallback model list for nearai to a single
default entry.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: unify default model, static bootstrap greeting, and web UI cleanup
- Add DEFAULT_MODEL const and default_models() fallback list in
llm/nearai_chat.rs; use from config, wizard, and .env.example so the
default model is defined in one place
- Restore multi-model fallback list in setup wizard (was reduced to 1)
- Move BOOTSTRAP_GREETING to module-level const (out of run() body)
- Replace LLM-based bootstrap with static greeting (persist to DB before
channels start, then broadcast — eliminates startup LLM call and race)
- Fix double env::var read for NEARAI_API_KEY in quick setup path
- Move thread sidebar buttons into threads-section-header (web UI)
- Remove orphaned .thread-sidebar-header CSS and fix double blank line
- Update bootstrap e2e test for static greeting (no LLM trace needed)
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix(safety): move prompt injection scanning into Workspace write/append
Addresses PR #927 review comments (#1, #3) — identity file write
protection and unsanitized profile fields in system prompt.
Instead of scanning at the tool layer (memory.rs) or the sync layer
(sync_profile_documents), injection scanning now lives in
Workspace::write() and Workspace::append() for all files that are
injected into the system prompt. This ensures every code path that
writes to these files is protected, including future ones.
- Add SYSTEM_PROMPT_FILES const and reject_if_injected() in workspace
- Add WorkspaceError::InjectionRejected variant
- Add map_write_err() in memory.rs to convert InjectionRejected to
ToolError::NotAuthorized
- Remove redundant IDENTITY_FILES/Sanitizer from memory.rs
- Remove redundant sanitizer calls from sync_profile_documents()
- Move sanitization tests to workspace::tests
- Existing integration test (test_memory_write_rejects_injection)
continues to pass through the new path
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* style: cargo fmt
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: address Copilot review — merge marker order, orphan thread, stale fixture
- merge_profile_section: search for END marker after BEGIN position to
avoid matching a stray END earlier in the file
- Bootstrap phase 2: use get_or_create_session + Thread::with_id instead
of resolve_thread(None) to avoid creating an orphan thread
- setup_nearai: use env_or_override for NEARAI_API_KEY consistency with
runtime overlay
- Delete orphaned bootstrap_greeting.json fixture (no test references it)
- Add test_merge_end_marker_must_follow_begin regression test
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* style: cargo fmt
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* style: fmt agent_loop.rs (CI stable rustfmt)
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: lazy-init sanitizer, check profile non-empty before skipping bootstrap
Address Copilot review:
- Use LazyLock<Sanitizer> to avoid rebuilding Aho-Corasick + regexes
on every workspace write
- has_profile check now requires non-empty content, not just file
existence, to prevent empty profile.json from suppressing onboarding
- Add seed_tests integration tests (libsql-backed) verifying:
- Empty profile.json does not suppress BOOTSTRAP.md seeding
- Non-empty profile.json correctly suppresses bootstrap for upgrades
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* style: cargo fmt
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: duplicate language handler, empty LLM_BACKEND, test_rig style
Address Copilot review on PR #927:
- Remove duplicate language-option click listeners (delegated
data-action handler already covers them)
- Guard LLM_BACKEND env prefill against empty string to prevent
suppressing API-key-based auto-detection
- Use destructured local `keep_bootstrap` instead of `self.keep_bootstrap`
in test_rig for consistency after destructure
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: update stale BOOTSTRAP.md write-protection comment [skip-regression-check]
BOOTSTRAP.md is now in SYSTEM_PROMPT_FILES and gets injection scanning
on write. The old comment incorrectly stated it was not write-protected.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: replace debug_assert panics with graceful error returns [skip-regression-check]
debug_assert! in execute_tool_with_safety and JobContext::transition_to
panicked in test builds before the graceful error path could run.
Existing tests (test_cancel_job_completed, test_execute_empty_tool_name_returns_not_found)
already cover these paths — they were the ones failing.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: address Copilot review — schema label, env var check, path normalization, profile validation
1. Label ANALYSIS_FRAMEWORK and PROFILE_JSON_SCHEMA sections separately
in bootstrap prompt so the LLM knows which blob is the target structure.
2. Wizard quick-mode backend auto-detection now rejects empty env vars
(std::env::var().is_ok_and(|v| !v.is_empty())) to avoid selecting the
wrong backend when e.g. NEARAI_API_KEY="" is set.
3. Normalize the target path before comparing with paths::PROFILE in
memory_write so non-canonical variants like "context//profile.json"
still trigger profile sync.
4. seed_if_empty now requires valid JSON parse of context/profile.json
before treating it as a populated profile. Corrupted content no longer
permanently suppresses bootstrap seeding.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* style: cargo fmt
* fix: address Copilot review — append scan, profile validation, env_or_override
1. Workspace::append() now scans the combined content (existing + new)
for prompt injection, not just the appended chunk. Prevents split-
injection evasion across multiple appends.
2. seed_if_empty() now deserializes into PsychographicProfile instead of
serde_json::Value for profile validation. Stray/legacy JSON that
doesn't match the expected schema no longer suppresses bootstrap.
3. Wizard quick-mode backend auto-detection now uses env_or_override()
to honor runtime overlays and injected secrets. LLM_BACKEND value
is trimmed before storage.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* test: add bootstrap_onboarding_clears_bootstrap E2E trace test
Exercises the full onboarding flow end-to-end:
1. Bootstrap greeting fires automatically on fresh workspace
2. User converses for 3 turns (name, tools, work style)
3. Agent writes psychographic profile to context/profile.json
4. Profile sync generates USER.md and assistant-directives.md
5. Agent writes IDENTITY.md (chosen persona)
6. Agent clears BOOTSTRAP.md via memory_write(target: "bootstrap")
Verifies:
- BOOTSTRAP.md is non-empty before onboarding, empty after
- bootstrap_completed flag is set
- Profile contains expected user data (name, profession, interests)
- USER.md contains profile-derived content (name, tone, profession)
- Assistant-directives.md references user and communication style
- IDENTITY.md contains agent's chosen persona name
- All memory_write calls succeed
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: address Copilot review — slash collapse, env_or_override, cron trim [skip-regression-check]
1. memory.rs path normalization now uses the same char-by-char loop as
Workspace::normalize_path() to fully collapse consecutive slashes
(e.g. "context///profile.json" → "context/profile.json").
2. Quick-mode NEARAI_API_KEY check (line 239) now uses env_or_override()
consistently with the backend auto-detection block above it.
3. normalize_cron_expression() trims input before field counting so the
passthrough branch (7+ fields) also strips whitespace.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
---------
Co-authored-by: Jay Zalowitz <[email protected]>
Co-authored-by: Claude Opus 4.6 <[email protected]>
This commit is contained in:
co-authored by
Jay Zalowitz
Claude Opus 4.6
parent
3a523347b0
commit
806d402876
@@ -705,4 +705,210 @@ mod advanced {
|
||||
mock_server.shutdown().await;
|
||||
rig.shutdown();
|
||||
}
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// 9. Bootstrap greeting fires on fresh workspace
|
||||
// -----------------------------------------------------------------------
|
||||
|
||||
/// Verifies that a fresh workspace triggers a static bootstrap greeting
|
||||
/// before the user sends any message (no LLM call needed).
|
||||
#[tokio::test]
|
||||
async fn bootstrap_greeting_fires() {
|
||||
let rig = TestRigBuilder::new().with_bootstrap().build().await;
|
||||
|
||||
// The static bootstrap greeting should arrive without us sending any
|
||||
// message and without an LLM call.
|
||||
let responses = rig.wait_for_responses(1, TIMEOUT).await;
|
||||
assert!(
|
||||
!responses.is_empty(),
|
||||
"bootstrap greeting should produce a response"
|
||||
);
|
||||
let greeting = &responses[0].content;
|
||||
assert!(
|
||||
greeting.contains("chief of staff"),
|
||||
"bootstrap greeting should contain the static text, got: {greeting}"
|
||||
);
|
||||
|
||||
// The bootstrap greeting must carry a thread_id so the gateway can
|
||||
// route it to the correct assistant conversation.
|
||||
assert!(
|
||||
responses[0].thread_id.is_some(),
|
||||
"bootstrap greeting response should have a thread_id set"
|
||||
);
|
||||
|
||||
rig.shutdown();
|
||||
}
|
||||
|
||||
// -----------------------------------------------------------------------
|
||||
// 10. Bootstrap onboarding completes and clears BOOTSTRAP.md
|
||||
// -----------------------------------------------------------------------
|
||||
|
||||
/// Exercises the full onboarding flow: bootstrap greeting fires, user
|
||||
/// converses for 3 turns, agent writes profile + memory + identity,
|
||||
/// clears BOOTSTRAP.md, and the workspace reflects all writes.
|
||||
#[tokio::test]
|
||||
async fn bootstrap_onboarding_clears_bootstrap() {
|
||||
use ironclaw::workspace::paths;
|
||||
|
||||
let trace = LlmTrace::from_file(format!("{FIXTURES}/bootstrap_onboarding.json")).unwrap();
|
||||
let rig = TestRigBuilder::new()
|
||||
.with_trace(trace.clone())
|
||||
.with_bootstrap()
|
||||
.build()
|
||||
.await;
|
||||
|
||||
// 1. Wait for the static bootstrap greeting (no user message needed).
|
||||
let greeting_responses = rig.wait_for_responses(1, TIMEOUT).await;
|
||||
assert!(
|
||||
!greeting_responses.is_empty(),
|
||||
"bootstrap greeting should arrive"
|
||||
);
|
||||
assert!(
|
||||
greeting_responses[0].content.contains("chief of staff"),
|
||||
"expected bootstrap greeting, got: {}",
|
||||
greeting_responses[0].content
|
||||
);
|
||||
|
||||
// 2. BOOTSTRAP.md should exist (non-empty) before onboarding completes.
|
||||
let ws = rig.workspace().expect("workspace should exist");
|
||||
let bootstrap_before = ws.read(paths::BOOTSTRAP).await;
|
||||
assert!(
|
||||
bootstrap_before.is_ok_and(|d| !d.content.is_empty()),
|
||||
"BOOTSTRAP.md should be non-empty before onboarding"
|
||||
);
|
||||
|
||||
// 3. Run the 3-turn conversation. The trace has the agent write
|
||||
// profile, memory, identity, and then clear bootstrap.
|
||||
let mut total = 1; // already have the greeting
|
||||
for turn in &trace.turns {
|
||||
rig.send_message(&turn.user_input).await;
|
||||
total += 1;
|
||||
let _ = rig.wait_for_responses(total, TIMEOUT).await;
|
||||
}
|
||||
|
||||
// 4. Verify all memory_write calls succeeded.
|
||||
let completed = rig.tool_calls_completed();
|
||||
let memory_writes: Vec<_> = completed
|
||||
.iter()
|
||||
.filter(|(name, _)| name == "memory_write")
|
||||
.collect();
|
||||
assert!(
|
||||
memory_writes.len() >= 4,
|
||||
"expected at least 4 memory_write calls (profile, memory, identity, bootstrap), got: {memory_writes:?}"
|
||||
);
|
||||
assert!(
|
||||
memory_writes.iter().all(|(_, ok)| *ok),
|
||||
"all memory_write calls should succeed: {memory_writes:?}"
|
||||
);
|
||||
|
||||
// 5. BOOTSTRAP.md should now be empty (cleared by memory_write target=bootstrap).
|
||||
let bootstrap_after = ws.read(paths::BOOTSTRAP).await.expect("read BOOTSTRAP");
|
||||
assert!(
|
||||
bootstrap_after.content.is_empty(),
|
||||
"BOOTSTRAP.md should be empty after onboarding, got: {:?}",
|
||||
bootstrap_after.content
|
||||
);
|
||||
|
||||
// 6. The bootstrap-completed flag should be set (prevents re-injection).
|
||||
assert!(
|
||||
ws.is_bootstrap_completed(),
|
||||
"bootstrap_completed flag should be set after profile write"
|
||||
);
|
||||
|
||||
// 7. Profile should exist in workspace with expected fields.
|
||||
let profile = ws.read(paths::PROFILE).await.expect("read profile");
|
||||
assert!(
|
||||
!profile.content.is_empty(),
|
||||
"profile.json should not be empty"
|
||||
);
|
||||
assert!(
|
||||
profile.content.contains("Alex"),
|
||||
"profile should contain preferred_name, got: {:?}",
|
||||
&profile.content[..profile.content.len().min(200)]
|
||||
);
|
||||
|
||||
// Try parsing the stored profile to catch deserialization issues early.
|
||||
let stored = ws
|
||||
.read(paths::PROFILE)
|
||||
.await
|
||||
.expect("read profile for deser test");
|
||||
let deser_result =
|
||||
serde_json::from_str::<ironclaw::profile::PsychographicProfile>(&stored.content);
|
||||
assert!(
|
||||
deser_result.is_ok(),
|
||||
"profile should deserialize: {:?}\ncontent: {:?}",
|
||||
deser_result.err(),
|
||||
&stored.content[..stored.content.len().min(300)]
|
||||
);
|
||||
let parsed = deser_result.unwrap();
|
||||
assert!(
|
||||
parsed.is_populated(),
|
||||
"profile should be populated: name={:?}, profession={:?}, goals={:?}",
|
||||
parsed.preferred_name,
|
||||
parsed.context.profession,
|
||||
parsed.assistance.goals
|
||||
);
|
||||
|
||||
// Manually trigger sync.
|
||||
let synced = ws
|
||||
.sync_profile_documents()
|
||||
.await
|
||||
.expect("sync_profile_documents");
|
||||
assert!(
|
||||
synced,
|
||||
"sync_profile_documents should return true for a populated profile"
|
||||
);
|
||||
assert!(
|
||||
profile.content.contains("backend engineer"),
|
||||
"profile should contain profession"
|
||||
);
|
||||
assert!(
|
||||
profile.content.contains("distributed systems"),
|
||||
"profile should contain interests"
|
||||
);
|
||||
|
||||
// 8. USER.md should have been synced from the profile via sync_profile_documents().
|
||||
let user_doc = ws.read(paths::USER).await.expect("read USER.md");
|
||||
assert!(
|
||||
user_doc.content.contains("Alex"),
|
||||
"USER.md should contain user name from profile, got: {:?}",
|
||||
&user_doc.content[..user_doc.content.len().min(300)]
|
||||
);
|
||||
assert!(
|
||||
user_doc.content.contains("direct"),
|
||||
"USER.md should contain communication tone from profile, got: {:?}",
|
||||
&user_doc.content[..user_doc.content.len().min(300)]
|
||||
);
|
||||
assert!(
|
||||
user_doc.content.contains("backend engineer"),
|
||||
"USER.md should contain profession from profile, got: {:?}",
|
||||
&user_doc.content[..user_doc.content.len().min(300)]
|
||||
);
|
||||
|
||||
// 9. Assistant directives should have been synced from the profile.
|
||||
let directives = ws
|
||||
.read(paths::ASSISTANT_DIRECTIVES)
|
||||
.await
|
||||
.expect("read assistant-directives.md");
|
||||
assert!(
|
||||
directives.content.contains("Alex"),
|
||||
"assistant-directives should reference user name, got: {:?}",
|
||||
&directives.content[..directives.content.len().min(300)]
|
||||
);
|
||||
assert!(
|
||||
directives.content.contains("direct"),
|
||||
"assistant-directives should reflect communication style, got: {:?}",
|
||||
&directives.content[..directives.content.len().min(300)]
|
||||
);
|
||||
|
||||
// 10. IDENTITY.md should have been written by the agent.
|
||||
let identity = ws.read(paths::IDENTITY).await.expect("read IDENTITY.md");
|
||||
assert!(
|
||||
identity.content.contains("Claw"),
|
||||
"IDENTITY.md should contain the chosen agent name, got: {:?}",
|
||||
identity.content
|
||||
);
|
||||
|
||||
rig.shutdown();
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,122 @@
|
||||
{
|
||||
"model_name": "bootstrap-onboarding-flow",
|
||||
"expects": {
|
||||
"tools_used": ["memory_write"],
|
||||
"all_tools_succeeded": true,
|
||||
"min_responses": 2
|
||||
},
|
||||
"turns": [
|
||||
{
|
||||
"user_input": "Hey! I'm Alex, a backend engineer working on distributed systems. I like direct, concise communication.",
|
||||
"steps": [
|
||||
{
|
||||
"request_hint": { "last_user_message_contains": "Alex" },
|
||||
"response": {
|
||||
"type": "text",
|
||||
"content": "Nice to meet you, Alex. Distributed systems — solid domain. I'll keep things direct.\n\nWhat's on your plate right now that I can help with? And any tools or services you use daily that I should know about?",
|
||||
"input_tokens": 800,
|
||||
"output_tokens": 40
|
||||
}
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"user_input": "I mainly use GitHub and Slack. I need help staying on top of PRs and incident response.",
|
||||
"steps": [
|
||||
{
|
||||
"request_hint": { "last_user_message_contains": "GitHub" },
|
||||
"response": {
|
||||
"type": "text",
|
||||
"content": "Got it — GitHub PRs and incident response are your priorities. I can set up routines for both.\n\nOne more thing: when something goes wrong at work, what's your instinct — jump in and debug, or step back and coordinate?",
|
||||
"input_tokens": 900,
|
||||
"output_tokens": 45
|
||||
}
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"user_input": "I usually jump in and start debugging. I'm pretty hands-on.",
|
||||
"steps": [
|
||||
{
|
||||
"request_hint": { "last_user_message_contains": "debugging" },
|
||||
"response": {
|
||||
"type": "tool_calls",
|
||||
"tool_calls": [
|
||||
{
|
||||
"id": "call_profile_write",
|
||||
"name": "memory_write",
|
||||
"arguments": {
|
||||
"content": "{\"version\":2,\"preferred_name\":\"Alex\",\"personality\":{\"empathy\":50,\"problem_solving\":50,\"emotional_intelligence\":50,\"adaptability\":50,\"communication\":50},\"communication\":{\"detail_level\":\"concise\",\"formality\":\"casual\",\"tone\":\"direct\",\"learning_style\":\"unknown\",\"social_energy\":\"unknown\",\"decision_making\":\"unknown\",\"pace\":\"fast\",\"response_speed\":\"unknown\"},\"cohort\":{\"cohort\":\"other\",\"confidence\":0,\"indicators\":[]},\"behavior\":{\"frictions\":[],\"desired_outcomes\":[],\"time_wasters\":[],\"pain_points\":[\"staying on top of PRs\",\"incident response\"],\"strengths\":[],\"suggested_support\":[]},\"friendship\":{\"style\":\"unknown\",\"values\":[],\"support_style\":\"unknown\",\"qualities\":{\"user_values\":[],\"friends_appreciate\":[],\"consistency_pattern\":null,\"primary_role\":null,\"secondary_roles\":[],\"challenging_aspects\":[]}},\"assistance\":{\"proactivity\":\"moderate\",\"formality\":\"unknown\",\"focus_areas\":[],\"routines\":[],\"goals\":[\"PR management\",\"incident response\"],\"interaction_style\":\"unknown\",\"notification_preferences\":\"moderate\"},\"context\":{\"profession\":\"backend engineer\",\"interests\":[\"distributed systems\"],\"life_stage\":null,\"challenges\":[]},\"relationship_values\":{\"primary\":[],\"secondary\":[],\"deal_breakers\":[]},\"interaction_preferences\":{\"proactivity_style\":\"reactive\",\"feedback_style\":\"direct\",\"decision_making\":\"guided\"},\"analysis_metadata\":{\"message_count\":0,\"analysis_date\":null,\"time_range\":null,\"model_used\":null,\"confidence_score\":0.0,\"analysis_method\":null,\"update_type\":null},\"confidence\":0.5,\"created_at\":\"2026-03-19T00:00:00Z\",\"updated_at\":\"2026-03-19T00:00:00Z\"}",
|
||||
"target": "context/profile.json",
|
||||
"append": false
|
||||
}
|
||||
}
|
||||
],
|
||||
"input_tokens": 1000,
|
||||
"output_tokens": 300
|
||||
}
|
||||
},
|
||||
{
|
||||
"response": {
|
||||
"type": "tool_calls",
|
||||
"tool_calls": [
|
||||
{
|
||||
"id": "call_memory_save",
|
||||
"name": "memory_write",
|
||||
"arguments": {
|
||||
"content": "User: Alex, backend engineer. Uses GitHub and Slack daily. Priorities: PR management, incident response. Communication: direct, concise, casual. Hands-on debugger.",
|
||||
"target": "daily_log"
|
||||
}
|
||||
}
|
||||
],
|
||||
"input_tokens": 1100,
|
||||
"output_tokens": 50
|
||||
}
|
||||
},
|
||||
{
|
||||
"response": {
|
||||
"type": "tool_calls",
|
||||
"tool_calls": [
|
||||
{
|
||||
"id": "call_identity_write",
|
||||
"name": "memory_write",
|
||||
"arguments": {
|
||||
"content": "# Identity\n\nName: Claw\nVibe: Sharp, efficient, no-nonsense\nEmoji: ⚡",
|
||||
"target": "IDENTITY.md",
|
||||
"append": false
|
||||
}
|
||||
}
|
||||
],
|
||||
"input_tokens": 1200,
|
||||
"output_tokens": 30
|
||||
}
|
||||
},
|
||||
{
|
||||
"response": {
|
||||
"type": "tool_calls",
|
||||
"tool_calls": [
|
||||
{
|
||||
"id": "call_bootstrap_clear",
|
||||
"name": "memory_write",
|
||||
"arguments": {
|
||||
"content": "",
|
||||
"target": "bootstrap"
|
||||
}
|
||||
}
|
||||
],
|
||||
"input_tokens": 1300,
|
||||
"output_tokens": 20
|
||||
}
|
||||
},
|
||||
{
|
||||
"response": {
|
||||
"type": "text",
|
||||
"content": "All set, Alex. I'm Claw — I'll keep things fast and direct.\n\nI've saved your profile and preferences. Want me to set up a morning PR review routine for your GitHub repos?",
|
||||
"input_tokens": 1400,
|
||||
"output_tokens": 35
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -25,6 +25,8 @@ use ironclaw::error::ChannelError;
|
||||
/// A `Channel` implementation for injecting messages and capturing responses
|
||||
/// in integration tests.
|
||||
pub struct TestChannel {
|
||||
/// Channel name returned by `Channel::name()`.
|
||||
channel_name: String,
|
||||
/// Sender half for injecting `IncomingMessage`s into the stream.
|
||||
tx: mpsc::Sender<IncomingMessage>,
|
||||
/// Receiver half, wrapped in Option so `start()` can take it exactly once.
|
||||
@@ -59,6 +61,7 @@ impl TestChannel {
|
||||
let (tx, rx) = mpsc::channel(256);
|
||||
let (ready_tx, ready_rx) = oneshot::channel();
|
||||
Self {
|
||||
channel_name: "test".to_string(),
|
||||
tx,
|
||||
rx: Mutex::new(Some(rx)),
|
||||
responses: Arc::new(Mutex::new(Vec::new())),
|
||||
@@ -72,6 +75,12 @@ impl TestChannel {
|
||||
}
|
||||
}
|
||||
|
||||
/// Override the channel name (default: "test").
|
||||
pub fn with_name(mut self, name: impl Into<String>) -> Self {
|
||||
self.channel_name = name.into();
|
||||
self
|
||||
}
|
||||
|
||||
/// Signal the channel (and any listening agent) to shut down.
|
||||
pub fn signal_shutdown(&self) {
|
||||
self.shutdown.store(true, Ordering::SeqCst);
|
||||
@@ -87,7 +96,7 @@ impl TestChannel {
|
||||
|
||||
/// Inject a user message into the channel stream.
|
||||
pub async fn send_message(&self, content: &str) {
|
||||
let msg = IncomingMessage::new("test", &self.user_id, content);
|
||||
let msg = IncomingMessage::new(&self.channel_name, &self.user_id, content);
|
||||
self.tx.send(msg).await.expect("TestChannel tx closed");
|
||||
}
|
||||
|
||||
@@ -98,7 +107,8 @@ impl TestChannel {
|
||||
|
||||
/// Inject a user message with a specific thread ID.
|
||||
pub async fn send_message_in_thread(&self, content: &str, thread_id: &str) {
|
||||
let msg = IncomingMessage::new("test", &self.user_id, content).with_thread(thread_id);
|
||||
let msg =
|
||||
IncomingMessage::new(&self.channel_name, &self.user_id, content).with_thread(thread_id);
|
||||
self.tx.send(msg).await.expect("TestChannel tx closed");
|
||||
}
|
||||
|
||||
@@ -281,7 +291,7 @@ impl Channel for TestChannelHandle {
|
||||
#[async_trait]
|
||||
impl Channel for TestChannel {
|
||||
fn name(&self) -> &str {
|
||||
"test"
|
||||
&self.channel_name
|
||||
}
|
||||
|
||||
async fn start(&self) -> Result<MessageStream, ChannelError> {
|
||||
@@ -291,7 +301,7 @@ impl Channel for TestChannel {
|
||||
.await
|
||||
.take()
|
||||
.ok_or_else(|| ChannelError::StartupFailed {
|
||||
name: "test".to_string(),
|
||||
name: self.channel_name.clone(),
|
||||
reason: "start() already called".to_string(),
|
||||
})?;
|
||||
|
||||
|
||||
@@ -354,6 +354,7 @@ pub struct TestRigBuilder {
|
||||
enable_routines: bool,
|
||||
http_exchanges: Vec<HttpExchange>,
|
||||
extra_tools: Vec<Arc<dyn Tool>>,
|
||||
keep_bootstrap: bool,
|
||||
}
|
||||
|
||||
impl TestRigBuilder {
|
||||
@@ -369,6 +370,7 @@ impl TestRigBuilder {
|
||||
enable_routines: false,
|
||||
http_exchanges: Vec::new(),
|
||||
extra_tools: Vec::new(),
|
||||
keep_bootstrap: false,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -426,6 +428,12 @@ impl TestRigBuilder {
|
||||
self
|
||||
}
|
||||
|
||||
/// Keep `bootstrap_pending` so the proactive greeting fires on startup.
|
||||
pub fn with_bootstrap(mut self) -> Self {
|
||||
self.keep_bootstrap = true;
|
||||
self
|
||||
}
|
||||
|
||||
/// Add pre-recorded HTTP exchanges for the `ReplayingHttpInterceptor`.
|
||||
///
|
||||
/// When set, all `http` tool calls will return these responses in order
|
||||
@@ -457,6 +465,7 @@ impl TestRigBuilder {
|
||||
enable_routines,
|
||||
http_exchanges: explicit_http_exchanges,
|
||||
extra_tools,
|
||||
keep_bootstrap,
|
||||
} = self;
|
||||
|
||||
// 1. Create temp dir + libSQL database + run migrations.
|
||||
@@ -537,6 +546,12 @@ impl TestRigBuilder {
|
||||
.await
|
||||
.expect("AppBuilder::build_all() failed in test rig");
|
||||
|
||||
// Clear bootstrap flag so tests don't get an unexpected proactive greeting
|
||||
// (unless the test explicitly wants to test the bootstrap flow).
|
||||
if !keep_bootstrap && let Some(ref ws) = components.workspace {
|
||||
ws.take_bootstrap_pending();
|
||||
}
|
||||
|
||||
// AppBuilder may re-resolve config from env/TOML and override test defaults.
|
||||
// Force test-rig agent flags to the requested deterministic values.
|
||||
components.config.agent.auto_approve_tools = auto_approve_tools.unwrap_or(true);
|
||||
@@ -648,7 +663,13 @@ impl TestRigBuilder {
|
||||
};
|
||||
|
||||
// 7. Create TestChannel and ChannelManager.
|
||||
let test_channel = Arc::new(TestChannel::new());
|
||||
// When testing bootstrap, the channel must be named "gateway" because
|
||||
// the bootstrap greeting targets only the gateway channel.
|
||||
let test_channel = if keep_bootstrap {
|
||||
Arc::new(TestChannel::new().with_name("gateway"))
|
||||
} else {
|
||||
Arc::new(TestChannel::new())
|
||||
};
|
||||
let handle = TestChannelHandle::new(Arc::clone(&test_channel));
|
||||
let channel_manager = ChannelManager::new();
|
||||
channel_manager.add(Box::new(handle)).await;
|
||||
|
||||
Reference in New Issue
Block a user