The dollar figure uses 10,000 uses of the recipe sample. Jev is the published $0.042 per million input list rate. The other figure uses a $3 per million example rate. Token count is the sample length divided by four. It is a planning estimate, not a tokenizer measurement. Output charges, retries, hosting, and task quality are not included.
Agent flow
8 capabilities
Decision
Ask or proceed
Identify missing information that blocks a concrete next step.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.06Jev list rate. About $3.99 at the $3 per million example rate. About 133 input tokens in the sample.
Decision
Choose the next tool
Choose among a small set of tools with explicit capabilities.
- Research value
- MediumEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.39 at the $3 per million example rate. About 113 input tokens in the sample.
Decision
Choose who should handle a request
Send a request to deterministic code, a specialist model, or a person.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.10Jev list rate. About $6.99 at the $3 per million example rate. About 233 input tokens in the sample.
Decision
Interrupt or stay quiet
Decide whether a background result deserves attention now.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.03Jev list rate. About $2.16 at the $3 per million example rate. About 72 input tokens in the sample.
Decision
Match a command to a request
Map an informal request to an available command ID.
- Research value
- MediumEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $2.76 at the $3 per million example rate. About 92 input tokens in the sample.
Decision
Recognize the current intent
Separate a question, a change request and a status request.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $2.82 at the $3 per million example rate. About 94 input tokens in the sample.
Decision
Review a proposed tool call
Classify a proposed tool call as clear or caution before anything runs.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.10Jev list rate. About $6.99 at the $3 per million example rate. About 233 input tokens in the sample.
Decision
Understand a tool result
Recognize success, failure or missing evidence without another long response.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.30 at the $3 per million example rate. About 110 input tokens in the sample.
Classification
2 capabilities
Catalog selection
Navigate a decision tree
Use jev-tree to choose a known category with a bounded call budget.
- Research value
- MediumEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.03Jev list rate. About $2.49 at the $3 per million example rate. About 83 input tokens in the sample.
Decision
Route a support request
Choose a queue and flag an explicit time-sensitive problem.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.08Jev list rate. About $5.85 at the $3 per million example rate. About 195 input tokens in the sample.
Coding
2 capabilities
Decision
Compare issue reports
Flag whether two reports describe the same observed failure.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $2.97 at the $3 per million example rate. About 99 input tokens in the sample.
Decision
Route a code change
Choose the right review path from a short change description.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $2.88 at the $3 per million example rate. About 96 input tokens in the sample.
Computer use
6 capabilities
Decision
Check a browser outcome
Ask whether the final screen supports the result the user requested.
- Research value
- HighEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $2.94 at the $3 per million example rate. About 98 input tokens in the sample.
Decision
Check a form against the task
Compare requested values with a visible form before the agent submits it.
- Research value
- HighEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $2.76 at the $3 per million example rate. About 92 input tokens in the sample.
Decision
Choose the next browser operation
Choose a small next step from the current page state and task.
- Research value
- HighEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.07Jev list rate. About $5.10 at the $3 per million example rate. About 170 input tokens in the sample.
Decision
Match a visible control
Choose a control ID from descriptions captured by your browser tool.
- Research value
- HighEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $3.06 at the $3 per million example rate. About 102 input tokens in the sample.
Decision
Read what changed on screen
Classify observed progress after an action without inventing a result.
- Research value
- HighEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.06Jev list rate. About $4.29 at the $3 per million example rate. About 143 input tokens in the sample.
Decision
Return a stuck browser task
Choose whether to wait, observe again or return control to the main agent.
- Research value
- HighEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.78 at the $3 per million example rate. About 126 input tokens in the sample.
Context
5 capabilities
Decision
Find a useful preference
Distinguish a reusable explicit preference from incidental conversation.
- Research value
- HighEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $3.12 at the $3 per million example rate. About 104 input tokens in the sample.
Decision
Flag how a passage relates to a task
Separate useful evidence, contradiction and embedded instructions.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.06Jev list rate. About $4.02 at the $3 per million example rate. About 134 input tokens in the sample.
Decision
Keep useful context
Evaluate candidate context before adding it to a longer prompt.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $3.00 at the $3 per million example rate. About 100 input tokens in the sample.
Decision
Select a useful source snippet
Select a source ID using the meaning of the request.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $3.03 at the $3 per million example rate. About 101 input tokens in the sample.
Decision
Spot stale evidence
Identify whether a supplied fact needs a fresh lookup.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.03Jev list rate. About $1.95 at the $3 per million example rate. About 65 input tokens in the sample.
Corpus search
5 capabilities
Decision
Choose the next corpus search step
Choose among retrieval expansion, keyword filtering, excerpt reading, reformulation or conclusion.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.09Jev list rate. About $6.69 at the $3 per million example rate. About 223 input tokens in the sample.
Decision
Reformulate agent search query
Choose how to refine a search query when BM25 or grep yields zero or noisy matches.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.07Jev list rate. About $5.10 at the $3 per million example rate. About 170 input tokens in the sample.
Decision
Select top candidate document
Pick the most promising candidate document from BM25 retrieval to stage or inspect first.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.06Jev list rate. About $4.14 at the $3 per million example rate. About 138 input tokens in the sample.
Decision
Triage grep match lines
Separate high-signal matches containing relevant evidence from routine boilerplate.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.60 at the $3 per million example rate. About 120 input tokens in the sample.
Decision
Verify claim against candidate passage
Compare a factual claim with a concrete line range or passage retrieved from the corpus.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $2.85 at the $3 per million example rate. About 95 input tokens in the sample.
Design
1 capability
Decision
Match a color palette
Choose an approved palette that fits a written brief.
- Research value
- MediumEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.36 at the $3 per million example rate. About 112 input tokens in the sample.
Experiences
1 capability
Dialogue
Choose a natural moment
Let an optional adviser choose an eligible recorded reaction or silence.
- Research value
- MediumEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.03Jev list rate. About $1.95 at the $3 per million example rate. About 65 input tokens in the sample.
Operations
2 capabilities
Decision
Choose a diagnostic test
Rank proposed checks using current failure evidence.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.45 at the $3 per million example rate. About 115 input tokens in the sample.
Log triage
Prioritize log inspection
Use jevlogs to separate routine events from records that deserve investigation.
- Research value
- LowEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.02Jev list rate. About $1.29 at the $3 per million example rate. About 43 input tokens in the sample.
Research
1 capability
Decision
Screen a document against criteria
Check an abstract against explicit inclusion criteria.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.57 at the $3 per million example rate. About 119 input tokens in the sample.
Selection
3 capabilities
Decision
Compare two item records
Flag whether two descriptions likely refer to the same item.
- Research value
- MediumEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $3.06 at the $3 per million example rate. About 102 input tokens in the sample.
Decision
Match an item to a description
Choose a supplied asset, template or catalog item by its description.
- Research value
- HighEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.33 at the $3 per million example rate. About 111 input tokens in the sample.
Decision
Select values for known fields
Map a request to allowed field values without generating arguments.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.24 at the $3 per million example rate. About 108 input tokens in the sample.
Verification
7 capabilities
Decision
Check a citation in context
Compare a claim with the supplied source passage.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.48 at the $3 per million example rate. About 116 input tokens in the sample.
Decision
Check a claim
Compare a short claim with the evidence already collected.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $2.58 at the $3 per million example rate. About 86 input tokens in the sample.
Decision
Check a repair against its invariant
Distinguish temporary recovery from a repair that meets stated requirements.
- Research value
- HighEvidence confidence is moderate. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.66 at the $3 per million example rate. About 122 input tokens in the sample.
Decision
Check launcher acceptance
Batch independent checks over one observed result.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.54 at the $3 per million example rate. About 118 input tokens in the sample.
Decision
Review a draft before it is sent
Check a draft reply against the evidence, then recommend send, hold, or review.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.10Jev list rate. About $7.44 at the $3 per million example rate. About 248 input tokens in the sample.
Decision
Review a proposed refund
Check whether a refund request names an amount and looks like a duplicate charge.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.11Jev list rate. About $7.62 at the $3 per million example rate. About 254 input tokens in the sample.
Decision
Score a small rubric
Evaluate a bounded quality scale with explicit anchors.
- Research value
- Not ratedNo research record is pinned. Compare it with your own task.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $3.03 at the $3 per million example rate. About 101 input tokens in the sample.
Writing
2 capabilities
Decision
Check named writing rules
Flag vague wording and missing specifics before an agent edits a draft.
- Research value
- MediumEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.05Jev list rate. About $3.87 at the $3 per million example rate. About 129 input tokens in the sample.
Decision
Label a text block
Assign a known structural label to supplied text.
- Research value
- MediumEvidence confidence is low. This is a research judgment, not a measured return.
- Input estimate for 10,000 uses
- $0.04Jev list rate. About $2.85 at the $3 per million example rate. About 95 input tokens in the sample.
No capability matches. Try a shorter task name, or search guides and findings.