Thesis
Demonstrated
The paper reports large safety metric gains on AgentHarm and InjecAgent across four named models.
Inference
A model/framework-neutral regression harness may catch serializer, tool-set, or model-upgrade safety drift.
Why now
Agent frameworks increasingly synthesize machine-readable tool schemas; SafeKeep supplies a specific representation-level failure hypothesis.
Binding constraint
No independent reproduction establishes that the schema-format effect survives current model/template versions, realistic tools, and adaptive attacks; buyer value is untested.
Scorecard
SafeKeep code, AgentHarm, InjecAgent, generic prompt rules, classifiers, monitors, and guardrails occupy the broad category. The different wedge is schema-diff regression evidence; patent/product search remains incomplete.
- False refusal
- Shared judge/executor failure modes
- Benchmark overfitting
- Model and template drift
- Sensitive tool descriptions
- Framework vendors can absorb the feature
Decisive experiment
Freeze 200 harmful and 200 matched-benign requests plus three realistic tool suites; compare schema versus flattened judgment over three current model families with prespecified safety, completion, validity, latency, and cost metrics.
Drop if flattening fails to improve refusal or injection ASR by at least 20 points on two of three model families, or benign completion falls more than 5 points versus a schema-aware two-stage control.
Source evidence
1 papers