AI safetyDeveloper toolsSoftware systems

Schema-aware agent safety preflight and regression evidence

Cross-framework regression testing that compares schema-formatted execution with flattened-description safety judgment before tools run.

01

Thesis

Demonstrated

The paper reports large safety metric gains on AgentHarm and InjecAgent across four named models.

Inference

A model/framework-neutral regression harness may catch serializer, tool-set, or model-upgrade safety drift.

Why now

Agent frameworks increasingly synthesize machine-readable tool schemas; SafeKeep supplies a specific representation-level failure hypothesis.

Binding constraint

No independent reproduction establishes that the schema-format effect survives current model/template versions, realistic tools, and adaptive attacks; buyer value is untested.

02

Scorecard

evidence3/5
unexploredness2/5
technical4/5
operational3/5
value2/5
Prior-art position

SafeKeep code, AgentHarm, InjecAgent, generic prompt rules, classifiers, monitors, and guardrails occupy the broad category. The different wedge is schema-diff regression evidence; patent/product search remains incomplete.

Risks
  • False refusal
  • Shared judge/executor failure modes
  • Benchmark overfitting
  • Model and template drift
  • Sensitive tool descriptions
  • Framework vendors can absorb the feature
03

Decisive experiment

Test

Freeze 200 harmful and 200 matched-benign requests plus three realistic tool suites; compare schema versus flattened judgment over three current model families with prespecified safety, completion, validity, latency, and cost metrics.

Kill criterion

Drop if flattening fails to improve refusal or injection ASR by at least 20 points on two of three model families, or benign completion falls more than 5 points versus a schema-aware two-stage control.

04

Source evidence

1 papers