mirror of
https://github.com/outbackdingo/optimclaw.git
synced 2026-09-02 17:49:20 +00:00
feat(agent): thread per-tool reasoning through provider, session, and all surfaces (#1513)
* feat(agent): thread per-tool reasoning from LLM through to REPL, HTTP, SSE, and DB Add end-to-end agent reasoning summaries so users can see *why* the agent chose specific tools, not just what it did. - Add `reasoning: Option<String>` to `ToolCall` (all providers) - Populate from LLM response content in `Reasoning::respond_with_tools` and `select_tools`, with per-tool override when providers supply it - Extend `Turn` with `narrative` and `TurnToolCall` with `rationale` + `tool_call_id` for identity-based result matching - Persist reasoning in DB via existing tool_calls JSON (no migration) - Add `StatusUpdate::ReasoningUpdate` and `SseEvent::ReasoningUpdate` + `SseEvent::JobReasoning` for real-time streaming - Emit reasoning events in both chat dispatcher and worker job path - Add `/reasoning [N|all]` command for inspecting turn reasoning - Surface `narrative` and `rationale` in HTTP `/api/chat/history` Based on the design from #361 and #456, reconstructed cleanly with Option<String> to minimize blast radius (vs mandatory String that broke compilation in #456). Closes #456 Co-Authored-By: panosAthDBX <[email protected]> Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> * fix: address PR review feedback from Gemini and Copilot - Fix `_ => Ok(None)` in agent_loop.rs to avoid accidental shutdown - Fix fallback in record_tool_result_for/record_tool_error_for to use first pending call instead of last_mut (parallel execution safety) - Include per-tool decisions in WASM channel reasoning messages - Apply truncate_at_tool_tags + clean_response to shared_reasoning in select_tools (parity with respond_with_tools) - Persist turn-level narrative to DB in tool_calls JSON wrapper - Parse both old (array) and new (object) tool_calls formats in build_turns_from_db_messages for backward compatibility - Populate reasoning from action.reasoning in execute_plan ToolCalls [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> * fix: address second round of review comments + merge fixes - Add reasoning: None to new github_copilot.rs ToolCall sites (from staging merge) - Run cargo fmt on 4 files with formatting diffs - Truncate narrative to 1000 chars before DB persistence - Clone turn data and drop session lock in /reasoning command - Extract ToolDecisionDto::from_json_array shared helper (deduplicate worker/job.rs and orchestrator/api.rs) - Add unit tests for wrapped tool_calls JSON format with narrative [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> * fix: address third round of review comments (Copilot + serrrfirat) - Reword ToolCall.reasoning docstring to reflect provider-supplied or fallback contract - Sanitize narrative through SafetyLayer before storage/emission - Clean per-tool reasoning via truncate_at_tool_tags + clean_response in select_tools (parity with shared reasoning) - Convert 4 approval-path recording sites in thread_ops.rs to identity-based record_tool_result_for/record_tool_error_for - Preserve tool_call_id and reasoning through restore_from_messages - Fix has_result/has_error to reject JSON null values - Truncate tool_call_id to 128 chars before DB persistence - Add 4 unit tests for record_tool_result_for/error_for edge cases Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> * fix: address zmanian review — sanitize JobDelegate reasoning + warn on dropped results - Sanitize narrative and per-tool rationale through SafetyLayer in JobDelegate reasoning events (parity with ChatDelegate) - Add tracing::warn when record_tool_result_for/error_for drops a result because no matching or pending tool call exists - Add 3 unit tests for reasoning normalization (thinking tags, tool tags, empty-after-cleaning) Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> * fix: address 4 remaining unreplied review comments - Clean per-tool reasoning in respond_with_tools via truncate_at_tool_tags + clean_response (parity with select_tools) - Handle wrapped JSON format in rebuild_chat_messages_from_db so cold hydration works after persist_tool_calls format change - Update persist_tool_calls doc comment to describe new JSON shape - Sanitize per-tool rationale through SafetyLayer in ChatDelegate before emission and storage (parity with JobDelegate) Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> * fix: address zmanian review round 2 - Add tracing::debug on fallback-to-pending path in record_tool_result_for and record_tool_error_for (item 1) - Add comment explaining why /reasoning is special-cased in agent_loop.rs (item 4) - Items 2 (narrative persistence), 3 (rationale sanitization), and 5 (catch-all fix) were already addressed in prior commits Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]> --------- Co-authored-by: panosAthDBX <[email protected]> Co-authored-by: Claude Opus 4.6 (1M context) <[email protected]>
This commit is contained in:
co-authored by
panosAthDBX
Claude Opus 4.6
parent
6daa2f155f
commit
41ed0a0f98
@@ -265,6 +265,15 @@ impl OutgoingResponse {
|
||||
}
|
||||
}
|
||||
|
||||
/// A single tool decision within a reasoning update.
|
||||
#[derive(Debug, Clone)]
|
||||
pub struct ToolDecision {
|
||||
/// Tool name.
|
||||
pub tool_name: String,
|
||||
/// Agent's reasoning for choosing this tool.
|
||||
pub rationale: String,
|
||||
}
|
||||
|
||||
/// Status update types for showing agent activity.
|
||||
#[derive(Debug, Clone)]
|
||||
pub enum StatusUpdate {
|
||||
@@ -333,6 +342,13 @@ pub enum StatusUpdate {
|
||||
},
|
||||
/// Suggested follow-up messages for the user.
|
||||
Suggestions { suggestions: Vec<String> },
|
||||
/// Agent reasoning update (why it chose specific tools).
|
||||
ReasoningUpdate {
|
||||
/// Human-readable summary of the agent's decision.
|
||||
narrative: String,
|
||||
/// Per-tool decisions.
|
||||
decisions: Vec<ToolDecision>,
|
||||
},
|
||||
/// Per-turn token usage and cost summary (shown as subtle metadata).
|
||||
TurnCost {
|
||||
input_tokens: u64,
|
||||
|
||||
+1
-1
@@ -39,7 +39,7 @@ mod webhook_server;
|
||||
|
||||
pub use channel::{
|
||||
AttachmentKind, Channel, ChannelSecretUpdater, IncomingAttachment, IncomingMessage,
|
||||
MessageStream, OutgoingResponse, StatusUpdate, routing_target_from_metadata,
|
||||
MessageStream, OutgoingResponse, StatusUpdate, ToolDecision, routing_target_from_metadata,
|
||||
};
|
||||
pub use http::{HttpChannel, HttpChannelState};
|
||||
pub use manager::ChannelManager;
|
||||
|
||||
@@ -75,6 +75,7 @@ const SLASH_COMMANDS: &[&str] = &[
|
||||
"/suggest",
|
||||
"/thread",
|
||||
"/resume",
|
||||
"/reasoning",
|
||||
];
|
||||
|
||||
/// Rustyline helper for slash-command tab completion.
|
||||
@@ -841,6 +842,19 @@ impl Channel for ReplChannel {
|
||||
StatusUpdate::Suggestions { .. } => {
|
||||
// Suggestions are only rendered by the web gateway
|
||||
}
|
||||
StatusUpdate::ReasoningUpdate {
|
||||
narrative,
|
||||
decisions,
|
||||
} => {
|
||||
if !narrative.is_empty() {
|
||||
let display = truncate_for_preview(&narrative, CLI_STATUS_MAX);
|
||||
eprintln!(" \x1b[94m\u{25B6} {display}\x1b[0m");
|
||||
}
|
||||
for d in &decisions {
|
||||
let display = truncate_for_preview(&d.rationale, CLI_STATUS_MAX);
|
||||
eprintln!(" \x1b[90m\u{2192} {}: {display}\x1b[0m", d.tool_name);
|
||||
}
|
||||
}
|
||||
StatusUpdate::TurnCost { .. } => {
|
||||
// Cost display is handled by the TUI channel
|
||||
}
|
||||
|
||||
@@ -3061,6 +3061,20 @@ fn status_to_wit(
|
||||
},
|
||||
// Suggestions and turn cost are web-gateway-only; skip for WASM channels
|
||||
StatusUpdate::Suggestions { .. } | StatusUpdate::TurnCost { .. } => return None,
|
||||
StatusUpdate::ReasoningUpdate {
|
||||
narrative,
|
||||
decisions,
|
||||
} => {
|
||||
let mut msg = narrative.clone();
|
||||
for d in decisions {
|
||||
msg.push_str(&format!("\n → {}: {}", d.tool_name, d.rationale));
|
||||
}
|
||||
wit_channel::StatusUpdate {
|
||||
status: wit_channel::StatusType::Status,
|
||||
message: msg,
|
||||
metadata_json,
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
|
||||
@@ -398,8 +398,10 @@ pub async fn chat_history_handler(
|
||||
truncate_preview(&s, 500)
|
||||
}),
|
||||
error: tc.error.clone(),
|
||||
rationale: tc.rationale.clone(),
|
||||
})
|
||||
.collect(),
|
||||
narrative: t.narrative.clone(),
|
||||
})
|
||||
.collect();
|
||||
|
||||
|
||||
@@ -489,6 +489,20 @@ impl Channel for GatewayChannel {
|
||||
},
|
||||
StatusUpdate::Suggestions { suggestions } => AppEvent::Suggestions {
|
||||
suggestions,
|
||||
thread_id: thread_id.clone(),
|
||||
},
|
||||
StatusUpdate::ReasoningUpdate {
|
||||
narrative,
|
||||
decisions,
|
||||
} => AppEvent::ReasoningUpdate {
|
||||
narrative,
|
||||
decisions: decisions
|
||||
.into_iter()
|
||||
.map(|d| crate::channels::web::types::ToolDecisionDto {
|
||||
tool_name: d.tool_name,
|
||||
rationale: d.rationale,
|
||||
})
|
||||
.collect(),
|
||||
thread_id,
|
||||
},
|
||||
StatusUpdate::TurnCost {
|
||||
|
||||
@@ -231,6 +231,7 @@ pub fn convert_messages(messages: &[OpenAiMessage]) -> Result<Vec<ChatMessage>,
|
||||
name: tc.function.name.clone(),
|
||||
arguments: serde_json::from_str(&tc.function.arguments)
|
||||
.unwrap_or(serde_json::Value::Object(Default::default())),
|
||||
reasoning: None,
|
||||
})
|
||||
.collect();
|
||||
Ok(ChatMessage::assistant_with_tool_calls(
|
||||
@@ -954,6 +955,7 @@ mod tests {
|
||||
id: "call_abc".to_string(),
|
||||
name: "search".to_string(),
|
||||
arguments: serde_json::json!({"query": "rust"}),
|
||||
reasoning: None,
|
||||
}];
|
||||
|
||||
let converted = convert_tool_calls_to_openai(&calls);
|
||||
|
||||
@@ -1725,8 +1725,10 @@ async fn chat_history_handler(
|
||||
truncate_preview(&s, 500)
|
||||
}),
|
||||
error: tc.error.clone(),
|
||||
rationale: tc.rationale.clone(),
|
||||
})
|
||||
.collect(),
|
||||
narrative: t.narrative.clone(),
|
||||
})
|
||||
.collect();
|
||||
|
||||
|
||||
@@ -63,6 +63,9 @@ pub struct TurnInfo {
|
||||
pub started_at: String,
|
||||
pub completed_at: Option<String>,
|
||||
pub tool_calls: Vec<ToolCallInfo>,
|
||||
/// Agent's reasoning narrative for this turn.
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub narrative: Option<String>,
|
||||
}
|
||||
|
||||
#[derive(Debug, Serialize)]
|
||||
@@ -74,6 +77,9 @@ pub struct ToolCallInfo {
|
||||
pub result_preview: Option<String>,
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub error: Option<String>,
|
||||
/// Agent's reasoning for choosing this tool.
|
||||
#[serde(skip_serializing_if = "Option::is_none")]
|
||||
pub rationale: Option<String>,
|
||||
}
|
||||
|
||||
#[derive(Debug, Serialize)]
|
||||
@@ -116,7 +122,7 @@ pub struct ApprovalRequest {
|
||||
|
||||
// --- App Event (re-exported from ironclaw_common) ---
|
||||
|
||||
pub use ironclaw_common::AppEvent;
|
||||
pub use ironclaw_common::{AppEvent, ToolDecisionDto};
|
||||
|
||||
// --- Memory ---
|
||||
|
||||
|
||||
+87
-12
@@ -4,6 +4,21 @@ use crate::channels::web::types::{ToolCallInfo, TurnInfo};
|
||||
|
||||
pub use ironclaw_common::truncate_preview;
|
||||
|
||||
/// Parse tool call summary JSON objects into `ToolCallInfo` structs.
|
||||
fn parse_tool_call_infos(calls: &[serde_json::Value]) -> Vec<ToolCallInfo> {
|
||||
calls
|
||||
.iter()
|
||||
.map(|c| ToolCallInfo {
|
||||
name: c["name"].as_str().unwrap_or("unknown").to_string(),
|
||||
has_result: c.get("result_preview").is_some_and(|v| !v.is_null()),
|
||||
has_error: c.get("error").is_some_and(|v| !v.is_null()),
|
||||
result_preview: c["result_preview"].as_str().map(String::from),
|
||||
error: c["error"].as_str().map(String::from),
|
||||
rationale: c["rationale"].as_str().map(String::from),
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// Build TurnInfo pairs from flat DB messages (user/tool_calls/assistant triples).
|
||||
///
|
||||
/// Handles three message patterns:
|
||||
@@ -27,6 +42,7 @@ pub fn build_turns_from_db_messages(
|
||||
started_at: msg.created_at.to_rfc3339(),
|
||||
completed_at: None,
|
||||
tool_calls: Vec::new(),
|
||||
narrative: None,
|
||||
};
|
||||
|
||||
// Check if next message is a tool_calls record
|
||||
@@ -34,18 +50,28 @@ pub fn build_turns_from_db_messages(
|
||||
&& next.role == "tool_calls"
|
||||
{
|
||||
let tc_msg = iter.next().expect("peeked");
|
||||
match serde_json::from_str::<Vec<serde_json::Value>>(&tc_msg.content) {
|
||||
Ok(calls) => {
|
||||
turn.tool_calls = calls
|
||||
.iter()
|
||||
.map(|c| ToolCallInfo {
|
||||
name: c["name"].as_str().unwrap_or("unknown").to_string(),
|
||||
has_result: c.get("result_preview").is_some(),
|
||||
has_error: c.get("error").is_some(),
|
||||
result_preview: c["result_preview"].as_str().map(String::from),
|
||||
error: c["error"].as_str().map(String::from),
|
||||
})
|
||||
.collect();
|
||||
// Parse tool_calls JSON — supports two formats:
|
||||
// safety: no byte-index slicing; comment describes JSON shape
|
||||
match serde_json::from_str::<serde_json::Value>(&tc_msg.content) {
|
||||
Ok(serde_json::Value::Array(calls)) => {
|
||||
// Old format: plain array
|
||||
turn.tool_calls = parse_tool_call_infos(&calls);
|
||||
}
|
||||
Ok(serde_json::Value::Object(obj)) => {
|
||||
// New wrapped format with narrative
|
||||
turn.narrative = obj
|
||||
.get("narrative")
|
||||
.and_then(|v| v.as_str())
|
||||
.map(String::from);
|
||||
if let Some(serde_json::Value::Array(calls)) = obj.get("calls") {
|
||||
turn.tool_calls = parse_tool_call_infos(calls);
|
||||
}
|
||||
}
|
||||
Ok(_) => {
|
||||
tracing::warn!(
|
||||
message_id = %tc_msg.id,
|
||||
"Unexpected tool_calls JSON shape in DB, skipping"
|
||||
);
|
||||
}
|
||||
Err(e) => {
|
||||
tracing::warn!(
|
||||
@@ -83,6 +109,7 @@ pub fn build_turns_from_db_messages(
|
||||
started_at: msg.created_at.to_rfc3339(),
|
||||
completed_at: Some(msg.created_at.to_rfc3339()),
|
||||
tool_calls: Vec::new(),
|
||||
narrative: None,
|
||||
});
|
||||
turn_number += 1;
|
||||
}
|
||||
@@ -201,4 +228,52 @@ mod tests {
|
||||
assert!(turns[0].tool_calls.is_empty());
|
||||
assert_eq!(turns[0].state, "Completed");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_build_turns_with_wrapped_tool_calls_format() {
|
||||
let tc_json = serde_json::json!({
|
||||
"narrative": "Searching memory for context before proceeding.",
|
||||
"calls": [
|
||||
{"name": "memory_search", "result_preview": "found 3 items", "rationale": "consult prior context"},
|
||||
{"name": "shell", "error": "permission denied"}
|
||||
]
|
||||
});
|
||||
let messages = vec![
|
||||
make_msg("user", "Find info", 0),
|
||||
make_msg("tool_calls", &tc_json.to_string(), 500),
|
||||
make_msg("assistant", "Here's what I found", 1000),
|
||||
];
|
||||
let turns = build_turns_from_db_messages(&messages);
|
||||
assert_eq!(turns.len(), 1);
|
||||
assert_eq!(
|
||||
turns[0].narrative.as_deref(),
|
||||
Some("Searching memory for context before proceeding.")
|
||||
);
|
||||
assert_eq!(turns[0].tool_calls.len(), 2);
|
||||
assert_eq!(turns[0].tool_calls[0].name, "memory_search");
|
||||
assert_eq!(
|
||||
turns[0].tool_calls[0].rationale.as_deref(),
|
||||
Some("consult prior context")
|
||||
);
|
||||
assert!(turns[0].tool_calls[0].has_result);
|
||||
assert_eq!(turns[0].tool_calls[1].name, "shell");
|
||||
assert!(turns[0].tool_calls[1].has_error);
|
||||
assert_eq!(turns[0].response.as_deref(), Some("Here's what I found"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_build_turns_wrapped_format_without_narrative() {
|
||||
let tc_json = serde_json::json!({
|
||||
"calls": [{"name": "echo", "result_preview": "hello"}]
|
||||
});
|
||||
let messages = vec![
|
||||
make_msg("user", "Say hi", 0),
|
||||
make_msg("tool_calls", &tc_json.to_string(), 500),
|
||||
make_msg("assistant", "Done", 1000),
|
||||
];
|
||||
let turns = build_turns_from_db_messages(&messages);
|
||||
assert_eq!(turns.len(), 1);
|
||||
assert!(turns[0].narrative.is_none());
|
||||
assert_eq!(turns[0].tool_calls.len(), 1);
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user