All evidence records

reported · independent experiment

Browser actions from observed control IDs

Less browser overhead helped a small matched comparison.

Browser Use repository authors; external implementation, not rerun here.

Evidence confidence

moderate. The comparison documents its method and limits. It does not establish performance across unfamiliar sites.

What was observed

Complete one Google Flights search.

Three alternating pairs compare two runtimes using Jev and the same text helper. A separate checker verifies results.

Baseline

The earlier runtime, also using Jev. This is not a frontier-agent comparison.

Finding

The optimized reader reduced median task time and browser protocol calls on this one task.

  • Reported medians: 9.450s versus 7.092s; 3/3 verified in each arm.
  • Median browser protocol calls: 1,092 versus 101.

What the result does not establish

Only three pairs. Initial navigation and independent verification are outside the task clock. Full task dollar cost is not reported.

What we would test in System One

Our inference: batch operation and target selection over current control state, then verify the actual result.

This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.

Try a related workflow

Primary sources

Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.

Metric definitions and review method · Submit a correction or new result