All capabilities

Verification · Decision

Score a small rubric

Evaluate a bounded quality scale with explicit anchors.

What this is worth testing

Your agent keeps the larger task and any action that follows. System One asks Jev for a typed answer on this check.

Research value
Not ratedNo research record is pinned. Compare this check with your own task before treating the answer as settled.
Sample size
About 101 input tokensThe recipe sample length divided by four. This is a planning estimate, not a measured tokenizer count.

$0.04 at Jev's $0.042 per million list rate

$3.03 at a $3 per million example rate

$2.99 input-cost difference for 10,000 uses of this sample

Excludes output charges, host tool calls, retries, hosting, and paid plans. Equal input volume does not establish equal answer quality or savings on a fixed-price subscription. $3 per million is an example rate, not a quoted model price.

Open the shared cost comparison

When an agent uses this

System One asks Jev for a typed answer. Your agent keeps the larger task, the original evidence, and any action that follows.

Example evidence

This is the recipe sample. Replace it with the evidence from your task. It is not a measured result.

Acceptance: error message should explain the failed action and recovery. Message: Could not save because the connection is offline. Reconnect and try Save again.

What your agent keeps

A rubric score is not a confidence probability. Validate it against human labels before automating decisions.

Finding

This capability is a contract and an example. A related research record has not been pinned for it. Compare it with your own task before treating the answer as settled.

Try it

Connect System One, then ask your agent for recipe rubric. Inspect the input on the recipe page before you send your own evidence.

More in Verification