* fix(safety): escape tool output XML content and remove misleading sanitized attr
The `sanitized="true/false"` attribute on `<tool_output>` misled LLMs into
treating unfiltered content as pre-sanitized. Remove it and add
`escape_xml_content()` to escape `<`, `>`, `&` in tool output body text,
preventing injected XML from breaking the structural boundary.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix(safety): replace contains assertions with exact assert_eq checks
Address Gemini review feedback on PR #1067: replace weak `contains`
assertions with precise `assert_eq!` comparisons in three safety tests
(wrap_for_llm escaping, XML boundary escape, escape_xml_content).
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: replace full XML escaping with targeted </tool_output escape to preserve JSON content
The previous approach escaped all XML metacharacters (<, >, &) in tool
output, which corrupted JSON content visible to the LLM. This was the
same issue that caused PR #598 to be reverted.
Now only the closing </tool_output sequence is neutralized (via a
zero-width space insertion), matching the pattern already used by
escape_skill_content(). All other content including JSON with angle
brackets and ampersands passes through unchanged.
Also:
- Remove unused _sanitized parameter from wrap_for_llm()
- Add unwrap_tool_output() with reverse escaping for round-trip fidelity
- Add round-trip tests verifying JSON content survives wrap/unwrap
- Update trace_llm test helper to use the new unwrap_tool_output()
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix: remove unwrap/expect from escape_tool_output_close to pass CI
Replace regex-based escaping with simple string search to avoid
.unwrap()/.expect() in production code (enforced by CI).
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* ci: re-trigger CI with latest changes
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix: remove stale 3rd arg from wrap_for_llm bench call
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix: address PR review - remove stale 3-arg call, add JSON round-trip test
Fix the test_wrap_for_llm_escapes_attr_chars test that still passed a
third `_sanitized` argument to wrap_for_llm (removed in earlier commit).
Add explicit JSON round-trip test with XML metacharacters
({"query": "a < b & c > d"}) confirming they survive wrap/unwrap intact,
as requested in PR #1067 review.
https://claude.ai/code/session_017ckCCurNiBL8uzE4dJg59K
* fix: remove stale sanitized= references from test fixtures, fix clippy warning
Update web/util.rs test fixtures to use the new tool_output format
without the removed sanitized="..." attribute. Remove redundant
#![cfg(test)] in codex_test_helpers.rs (already gated in mod.rs).
https://claude.ai/code/session_01Q4bRgRy96cqfmVPao4XiX8
* test: add round-trip JSON parsing regression gate for PR #598
Adds a test that verifies JSON content with XML metacharacters (<, >, &)
survives the full wrap_for_llm -> unwrap_tool_output -> serde_json::from_str
pipeline intact. This guards against the exact corruption scenario that
motivated reverting full XML escaping in PR #598.
https://claude.ai/code/session_01R2Zt832cV1xxDf7NXNq5GV
* fix(safety): harden wrap_external_content against boundary injection
Address reviewer feedback: apply the same targeted escaping strategy
to wrap_external_content() that was applied to wrap_for_llm(). The
closing delimiter "--- END EXTERNAL CONTENT ---" is now neutralized
in content bodies using a zero-width space, preventing an attacker
from injecting a fake closing delimiter to break out of the wrapper.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
---------
Co-authored-by: Claude Opus 4.6 <[email protected]>
* perf(safety): make XML attribute escaping single-pass
* test(safety): annotate assertion for no-panics CI
* test(safety): inline no-panics suppression comment
These tests guard against catastrophic regex backtracking (seconds/minutes),
not 12ms differences. CI runners with coverage instrumentation (cargo-llvm-cov)
consistently exceed the 100ms threshold due to overhead, causing flaky failures.
500ms still catches real regressions while tolerating CI variability.
[skip-regression-check]
Co-authored-by: Claude Opus 4.6 (1M context) <[email protected]>
* fix: eliminate panic paths in production code and document infallible operations
PolicyRule::new() now returns Result instead of panicking on invalid
caller-supplied regex. CreateJobTool returns ToolError when job_manager
is unconfigured instead of panicking. Remaining infallible unwrap/expect
calls (hardcoded regexes, compile-time constants, guarded accesses)
are annotated with SAFETY comments. Where possible, unwraps are replaced
with safer patterns: split_last(), if-let, match-destructure, and
reusing peek() values.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: use inline lowercase safety comments to match CI pattern
The no-panics CI check greps for '// safety:' (lowercase, inline)
to suppress false positives. Switch from block SAFETY comments to
inline safety comments on the .unwrap() lines.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* test: add regression tests for panic-path fixes
- PolicyRule::new returns Err on invalid regex (not panic)
- CreateJobTool::execute_sandbox returns ToolError when job_manager is None
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* fix: add inline // safety: comments on all infallible unwrap/expect lines
The CI no-panics check requires '// safety:' on the same line as
unwrap()/expect() to suppress false positives. Move safety annotations
from block comments to inline comments on every infallible production
unwrap/expect across all touched files.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* chore: trigger CI with skip-regression-check label
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
* refactor: remove redundant block-level SAFETY comments
Each unwrap/expect now carries its own inline // safety: annotation,
making the standalone block comments above them redundant.
Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <[email protected]>
* refactor: extract safety module into ironclaw_safety crate
Move prompt injection defense, input validation, secret leak detection,
and safety policy enforcement into a standalone crate under crates/.
The safety module was a leaf dependency with no async, no database, and
no other ironclaw traits — only pure computation with pattern matching.
SafetyConfig (2 fields) moves into the crate; env-var resolution stays
in ironclaw's config module as a free function. src/safety/mod.rs becomes
a thin re-export so all existing `crate::safety::*` imports keep working.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* docs: update CLAUDE.md for ironclaw_safety crate extraction
Add guidance to migrate imports from crate::safety to ironclaw_safety
when touching files. Update project structure to reflect crates/ dir.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* refactor: move safety fuzz targets into ironclaw_safety crate
Split fuzz infrastructure:
- crates/ironclaw_safety/fuzz/ — 5 safety-only targets (sanitizer,
validator, leak_detector, credential_detect, config_env) depending
only on ironclaw_safety for faster builds
- fuzz/ — keeps fuzz_tool_params which needs ironclaw::tools
Add seed corpus files (51 total) covering each pattern family:
sanitizer injection patterns, validator edge cases, leak detector
secret formats, credential detect HTTP param shapes.
Add new fuzz_credential_detect target exercising
params_contain_manual_credentials with arbitrary JSON.
Co-Authored-By: Claude Opus 4.6 <[email protected]>
* fix: address PR review — single-pass XML escaping and versioned path dep
Rewrite escape_xml_attr from chained .replace() to single-pass char
iteration (O(n) instead of O(4n) with intermediate allocations). Add
version = "0.1.0" to ironclaw_safety path dep to satisfy cargo-deny
wildcards = "deny".
Co-Authored-By: Claude Opus 4.6 <[email protected]>
---------
Co-authored-by: Claude Opus 4.6 <[email protected]>