capability · Agent flow · Decision
Ask or proceed
Identify missing information that blocks a concrete next step.
Research value not rated. About 133 input tokens in the sample. At 10,000 uses, about $0.06 at the Jev list rate.
Open Ask or proceedcapability · Agent flow · Decision
Choose the next tool
Choose among a small set of tools with explicit capabilities.
Research value medium. About 113 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Choose the next toolcapability · Agent flow · Decision
Choose who should handle a request
Send a request to deterministic code, a specialist model, or a person.
Research value not rated. About 233 input tokens in the sample. At 10,000 uses, about $0.10 at the Jev list rate.
Open Choose who should handle a requestcapability · Agent flow · Decision
Interrupt or stay quiet
Decide whether a background result deserves attention now.
Research value not rated. About 72 input tokens in the sample. At 10,000 uses, about $0.03 at the Jev list rate.
Open Interrupt or stay quietcapability · Agent flow · Decision
Match a command to a request
Map an informal request to an available command ID.
Research value medium. About 92 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Match a command to a requestcapability · Agent flow · Decision
Recognize the current intent
Separate a question, a change request and a status request.
Research value not rated. About 94 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Recognize the current intentcapability · Agent flow · Decision
Review a proposed tool call
Classify a proposed tool call as clear or caution before anything runs.
Research value not rated. About 233 input tokens in the sample. At 10,000 uses, about $0.10 at the Jev list rate.
Open Review a proposed tool callcapability · Agent flow · Decision
Understand a tool result
Recognize success, failure or missing evidence without another long response.
Research value high. About 110 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Understand a tool resultcapability · Classification · Catalog selection
Navigate a decision tree
Use jev-tree to choose a known category with a bounded call budget.
Research value medium. About 83 input tokens in the sample. At 10,000 uses, about $0.03 at the Jev list rate.
Open Navigate a decision treecapability · Classification · Decision
Route a support request
Choose a queue and flag an explicit time-sensitive problem.
Research value not rated. About 195 input tokens in the sample. At 10,000 uses, about $0.08 at the Jev list rate.
Open Route a support requestcapability · Coding · Decision
Compare issue reports
Flag whether two reports describe the same observed failure.
Research value not rated. About 99 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Compare issue reportscapability · Coding · Decision
Route a code change
Choose the right review path from a short change description.
Research value not rated. About 96 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Route a code changecapability · Computer use · Decision
Check a browser outcome
Ask whether the final screen supports the result the user requested.
Research value high. About 98 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Check a browser outcomecapability · Computer use · Decision
Check a form against the task
Compare requested values with a visible form before the agent submits it.
Research value high. About 92 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Check a form against the taskcapability · Computer use · Decision
Choose the next browser operation
Choose a small next step from the current page state and task.
Research value high. About 170 input tokens in the sample. At 10,000 uses, about $0.07 at the Jev list rate.
Open Choose the next browser operationcapability · Computer use · Decision
Match a visible control
Choose a control ID from descriptions captured by your browser tool.
Research value high. About 102 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Match a visible controlcapability · Computer use · Decision
Read what changed on screen
Classify observed progress after an action without inventing a result.
Research value high. About 143 input tokens in the sample. At 10,000 uses, about $0.06 at the Jev list rate.
Open Read what changed on screencapability · Computer use · Decision
Return a stuck browser task
Choose whether to wait, observe again or return control to the main agent.
Research value high. About 126 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Return a stuck browser taskcapability · Context · Decision
Find a useful preference
Distinguish a reusable explicit preference from incidental conversation.
Research value high. About 104 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Find a useful preferencecapability · Context · Decision
Flag how a passage relates to a task
Separate useful evidence, contradiction and embedded instructions.
Research value not rated. About 134 input tokens in the sample. At 10,000 uses, about $0.06 at the Jev list rate.
Open Flag how a passage relates to a taskcapability · Context · Decision
Keep useful context
Evaluate candidate context before adding it to a longer prompt.
Research value high. About 100 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Keep useful contextcapability · Context · Decision
Select a useful source snippet
Select a source ID using the meaning of the request.
Research value high. About 101 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Select a useful source snippetcapability · Context · Decision
Spot stale evidence
Identify whether a supplied fact needs a fresh lookup.
Research value not rated. About 65 input tokens in the sample. At 10,000 uses, about $0.03 at the Jev list rate.
Open Spot stale evidencecapability · Corpus search · Decision
Choose the next corpus search step
Choose among retrieval expansion, keyword filtering, excerpt reading, reformulation or conclusion.
Research value high. About 223 input tokens in the sample. At 10,000 uses, about $0.09 at the Jev list rate.
Open Choose the next corpus search stepcapability · Corpus search · Decision
Reformulate agent search query
Choose how to refine a search query when BM25 or grep yields zero or noisy matches.
Research value high. About 170 input tokens in the sample. At 10,000 uses, about $0.07 at the Jev list rate.
Open Reformulate agent search querycapability · Corpus search · Decision
Select top candidate document
Pick the most promising candidate document from BM25 retrieval to stage or inspect first.
Research value high. About 138 input tokens in the sample. At 10,000 uses, about $0.06 at the Jev list rate.
Open Select top candidate documentcapability · Corpus search · Decision
Triage grep match lines
Separate high-signal matches containing relevant evidence from routine boilerplate.
Research value high. About 120 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Triage grep match linescapability · Corpus search · Decision
Verify claim against candidate passage
Compare a factual claim with a concrete line range or passage retrieved from the corpus.
Research value high. About 95 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Verify claim against candidate passagecapability · Design · Decision
Match a color palette
Choose an approved palette that fits a written brief.
Research value medium. About 112 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Match a color palettecapability · Experiences · Dialogue
Choose a natural moment
Let an optional adviser choose an eligible recorded reaction or silence.
Research value medium. About 65 input tokens in the sample. At 10,000 uses, about $0.03 at the Jev list rate.
Open Choose a natural momentcapability · Operations · Decision
Choose a diagnostic test
Rank proposed checks using current failure evidence.
Research value high. About 115 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Choose a diagnostic testcapability · Operations · Log triage
Prioritize log inspection
Use jevlogs to separate routine events from records that deserve investigation.
Research value low. About 43 input tokens in the sample. At 10,000 uses, about $0.02 at the Jev list rate.
Open Prioritize log inspectioncapability · Research · Decision
Screen a document against criteria
Check an abstract against explicit inclusion criteria.
Research value not rated. About 119 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Screen a document against criteriacapability · Selection · Decision
Compare two item records
Flag whether two descriptions likely refer to the same item.
Research value medium. About 102 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Compare two item recordscapability · Selection · Decision
Match an item to a description
Choose a supplied asset, template or catalog item by its description.
Research value high. About 111 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Match an item to a descriptioncapability · Selection · Decision
Select values for known fields
Map a request to allowed field values without generating arguments.
Research value not rated. About 108 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Select values for known fieldscapability · Verification · Decision
Check a citation in context
Compare a claim with the supplied source passage.
Research value high. About 116 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Check a citation in contextcapability · Verification · Decision
Check a claim
Compare a short claim with the evidence already collected.
Research value high. About 86 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Check a claimcapability · Verification · Decision
Check a repair against its invariant
Distinguish temporary recovery from a repair that meets stated requirements.
Research value high. About 122 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Check a repair against its invariantcapability · Verification · Decision
Check launcher acceptance
Batch independent checks over one observed result.
Research value not rated. About 118 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Check launcher acceptancecapability · Verification · Decision
Review a draft before it is sent
Check a draft reply against the evidence, then recommend send, hold, or review.
Research value not rated. About 248 input tokens in the sample. At 10,000 uses, about $0.10 at the Jev list rate.
Open Review a draft before it is sentcapability · Verification · Decision
Review a proposed refund
Check whether a refund request names an amount and looks like a duplicate charge.
Research value not rated. About 254 input tokens in the sample. At 10,000 uses, about $0.11 at the Jev list rate.
Open Review a proposed refundcapability · Verification · Decision
Score a small rubric
Evaluate a bounded quality scale with explicit anchors.
Research value not rated. About 101 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Score a small rubriccapability · Writing · Decision
Check named writing rules
Flag vague wording and missing specifics before an agent edits a draft.
Research value medium. About 129 input tokens in the sample. At 10,000 uses, about $0.05 at the Jev list rate.
Open Check named writing rulescapability · Writing · Decision
Label a text block
Assign a known structural label to supplied text.
Research value medium. About 95 input tokens in the sample. At 10,000 uses, about $0.04 at the Jev list rate.
Open Label a text blockguide · Coding and chat agents
Check a claim against tool output
Ask whether the evidence supports a statement before an agent repeats it as fact.
Open Check a claim against tool outputguide · Coding and research agents
Shortlist context for an agent
Evaluate a small set of retrieved passages without losing the source material.
Open Shortlist context for an agentguide · Agents with several connected tools
Choose the next available tool
Give an agent a tool recommendation from a list it can actually use.
Open Choose the next available toolguide · Anyone creating a System One pattern
Write a question Jev can evaluate
Define the evidence, possible answers and review policy before adjusting a threshold.
Open Write a question Jev can evaluateguide · Individual coding and chat agent users
Measure whether offloading helps
Compare completed tasks with and without System One, including review and fallback costs.
Open Measure whether offloading helpsguide · Individual coding-agent users with a local stdio MCP client
Give an agent a short browser subtask
Use current control IDs, caller-supplied text and final-screen review with the System One browser companion.
Open Give an agent a short browser subtaskguide · Agent and computer-use adapter developers
Give Jev useful accessibility evidence
Use labels, group context and exact field facts to turn a screen into a small decision.
Open Give Jev useful accessibility evidencefinding · reported · independent experiment
Model routing: test whether Jev adds value
No-Jev ablation matched the hybrid result.
Open Model routing: test whether Jev adds valuefinding · reported · independent experiment
Choosing tools from real MCP inventories
Better tool prediction can still take longer.
Open Choosing tools from real MCP inventoriesfinding · reported · synthetic test
Separate questions for separate hazards
Explicit checks helped; calibration and latency limits remain.
Open Separate questions for separate hazardsfinding · first-party · public-dataset study
Log triage: filtering is not automatically saving
Conservative triage can add cost; reuse may help more.
Open Log triage: filtering is not automatically savingfinding · first-party · synthetic test
Navigating beyond a flat choice limit
Trees avoid truncation; structured data may need no model.
Open Navigating beyond a flat choice limitfinding · reported · vendor benchmark
Typed evaluation versus chat-model wrappers
A reason to test typed decisions, not a savings guarantee.
Open Typed evaluation versus chat-model wrappersfinding · reference · official guidance
Define ambiguity before evaluating a workflow
Specify what uncertain and mixed cases should do.
Open Define ambiguity before evaluating a workflowfinding · first-party · synthetic test
Batching two evidence checks
Fewer calls and tokens, with one extra error in this small test.
Open Batching two evidence checksfinding · reported · public-dataset study
Rank retrieved passages before reading them
Ranking gains depend on the dataset and how scores are averaged.
Open Rank retrieved passages before reading themfinding · reported · public-dataset study
Find the likely failure in an agent trace
Trace classification is promising, but most full diagnoses were wrong.
Open Find the likely failure in an agent tracefinding · reported · synthetic-data study
Separate signals from a final verdict
Question design helped, but a simple rule was a strong baseline.
Open Separate signals from a final verdictfinding · reported · public-dataset and synthetic study
Compare decision types before choosing a model
A broad comparison supports task-specific testing, not one universal winner.
Open Compare decision types before choosing a modelfinding · anecdotal · firsthand Reddit report
A personal app routes two recipe requests
Descriptions can distinguish two plausible routes; reliability remains untested.
Open A personal app routes two recipe requestsfinding · anecdotal · firsthand Reddit report
A user tries Jev to reduce a long agent history
A user found history selection useful; quality and cache costs need a test.
Open A user tries Jev to reduce a long agent historyfinding · anecdotal · X anecdotes and criticism
X posts disagree about compaction by filtering
Conflicting firsthand views identify a question to test, not a settled result.
Open X posts disagree about compaction by filteringfinding · anecdotal · builder report
Choose a card design from a page description
A concrete design-selection integration exists; preference quality is unmeasured.
Open Choose a card design from a page descriptionfinding · reported · builder experiment
Select an editor command from an informal request
An authored demo maps descriptions to commands better than its name matcher.
Open Select an editor command from an informal requestfinding · reported · application experiment
Choose diagnostic tests and review repair evidence
More attempts passed in one small study; some incidents regressed.
Open Choose diagnostic tests and review repair evidencefinding · reference · official examples
TypeSafe examples for decisions over supplied text
Official examples show how to frame the question; each adaptation needs testing.
Open TypeSafe examples for decisions over supplied textfinding · first-party · synthetic test
Parable dialogue advice chose silence in every first-run case
Conservative silence is not evidence of better dialogue.
Open Parable dialogue advice chose silence in every first-run casefinding · reported · independent experiment
Browser actions from observed control IDs
Less browser overhead helped a small matched comparison.
Open Browser actions from observed control IDsfinding · reported · independent experiment
Mac actions from OCR and accessibility
Structured screen extraction is part of the workload.
Open Mac actions from OCR and accessibilityfinding · first-party · synthetic test
Shared page state improved our first browser diagnostic
Four of six saved states after a shared-input fix.
Open Shared page state improved our first browser diagnosticfinding · reference · implementation
What Jev needs from a computer-use adapter
Supply explicit accessible controls and shared state.
Open What Jev needs from a computer-use adapterfinding · first-party · synthetic test
Six accessibility-informed next-action choices
Six correct choices; complete-task benefit still untested.
Open Six accessibility-informed next-action choicesNothing matches. Try the task in fewer words, or browse the capability directory.