All capability theses

A use case to evaluate

Delegate routine browser decisions

With Jev, an agent can choose the next action from a fresh list of visible browser controls and review the resulting screen.

A human example

You ask your coding agent to save a light-mode preview named Alpha in a web app.

What the caller supplies

The agent supplies the goal and exact values. The companion supplies Chrome-computed names, roles, group context, field states and a screenshot for the agent.

What happens next

Jev recommends type, select or click with an observed target. The companion can perform a short authorized sequence and return the final screenshot.

Illustrative workflow. The linked original fixture records successful attempts and uncertainty stops.

Potential value: high

Repeated interface choices may fit a small decision model and avoid a full main-agent turn for each click. The end-to-end benefit still needs measurement.

Evidence confidence: low

External implementations and our six-task development diagnostic establish feasibility. They do not establish reliable completion or savings across unfamiliar websites.

The rating describes support for this claim. It is separate from Jev's returned probability. How we assign ratings.

Evidence, including disagreement

The next test

This protocol is planned. Its outcome is not yet known.

At least 40 unfamiliar tasks across original forms, filters, item lists and changing layouts. Freeze tasks and verifiers before testing. Include missing controls, loading, duplicate labels and misleading page instructions.

Compare against

  • Calling agent with its normal browser tools
  • Calling agent with System One single-step advice
  • Calling agent with the bounded companion loop
  • Known deterministic scripts where applicable
  • Paired identical browser tasks with and without AX enrichment

Measure

  • Independently verified task completion
  • Unintended writes and duplicate submissions
  • Agent escalations and manual interventions
  • Complete task median and p95 including startup and host review
  • All model usage, provider errors and stopped attempts
  • Screenshot and text volume sent to each model

Decision after the test

Keep the companion opt-in. Require no unintended writes in the reviewed suite and an equal-quality complete-workflow benefit before a savings claim. Publish all attempts and confidence intervals.

The report will retain inputs, question versions, every attempt and failure examples. We will update the confidence rating after reviewing the result.

Use a related Engine recipe

Recipes are implementation starting points. Their presence does not mean the protocol above has passed.

Read or improve this thesis on GitHub.