Evidence confidence
low. This source suggests a useful experiment but does not establish a reliable benefit in a real agent workflow.
What was observed
Flag hazards in proposed agent tool calls before execution.
600 authored records, 320 held out; full hazard battery compared with a broad dangerousness question.
Baseline
Single generic hazard question.
Finding
Decomposed questions sharply reduced false blocks on hard negatives in this corpus. Calibration and latency targets still failed.
- Hard-negative block rate: full battery 0%; generic question 39.2%.
- Reported ECE 0.156; p95 added latency 595ms.
What the result does not establish
Author-designed synthetic attacks are not evidence of adversarial security in production. Approval friction must be measured separately.
What we would test in System One
Use independent, explicit checks as advisory signals. Keep authorization, execution safeguards and human approval outside Jev.
This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.
Try a related workflow
Primary sources
Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.
Metric definitions and review method · Submit a correction or new result