Evidence confidence
moderate. The comparison documents its method and limits. It does not establish performance across unfamiliar sites.
What was observed
Complete one Google Flights search.
Three alternating pairs compare two runtimes using Jev and the same text helper. A separate checker verifies results.
Baseline
The earlier runtime, also using Jev. This is not a frontier-agent comparison.
Finding
The optimized reader reduced median task time and browser protocol calls on this one task.
- Reported medians: 9.450s versus 7.092s; 3/3 verified in each arm.
- Median browser protocol calls: 1,092 versus 101.
What the result does not establish
Only three pairs. Initial navigation and independent verification are outside the task clock. Full task dollar cost is not reported.
What we would test in System One
Our inference: batch operation and target selection over current control state, then verify the actual result.
This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.
Try a related workflow
Primary sources
Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.
Metric definitions and review method · Submit a correction or new result