All evidence records

reported · public-dataset study

Find the likely failure in an agent trace

Trace classification is promising, but most full diagnoses were wrong.

TokenTrim; external author. We reviewed the report and did not rerun it.

Evidence confidence

moderate. A described comparison supports a bounded conclusion. Workload transfer and independent reproduction remain unresolved.

What was observed

Identify the responsible agent, step and error category in a failed run.

6,257 text traces from Who&When Pro. Jev answers three choice questions. The author uses the official scorer and compares with published paper baselines.

Baseline

GPT-5.4 results from the original paper, not a fresh matched API run.

Finding

Reported error-category macro-F1 was 23.7 for Jev and 15.3 for GPT-5.4 on the 100-point scale. Joint accuracy was 31.3% and 21.3%.

  • Error-category macro-F1: Jev 23.7; paper GPT-5.4 baseline 15.3.
  • All three labels correct: Jev 31.3%; paper GPT-5.4 baseline 21.3%.
  • Reported Jev input-price estimate: $1.28 for 6,257 traces.

What the result does not establish

Failures are injected. Jev selects enumerated agents and steps while paper baselines generate them. Mode-confidence ECE is 0.287. No latency comparison is reported.

What we would test in System One

Try advisory trace triage with a fixed error taxonomy. A low joint success rate does not justify automatic repair or blame assignment.

This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.

Primary sources

Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.

Metric definitions and review method · Submit a correction or new result