All evidence records

reported · public-dataset study

Rank retrieved passages before reading them

Ranking gains depend on the dataset and how scores are averaged.

Aness Belbati; external author. We reviewed the report and did not rerun it.

Evidence confidence

moderate. A described comparison supports a bounded conclusion. Workload transfer and independent reproduction remain unresolved.

What was observed

Reorder retrieved passages by relevance to a question.

Eight English datasets, 1,617 scored queries. Models see the same 30 BM25 candidates, each truncated to 2,000 characters. Jev 1.13.0 uses a four-level rubric.

Baseline

BM25, Cohere Rerank 4 Pro and other published rerankers.

Finding

Dataset-weighted nDCG@10 was 0.692 for Jev and 0.691 for Cohere Pro. The difference interval crosses zero. Weighting each query equally favors Cohere.

  • Jev minus Cohere nDCG@10: 0.001; 95% interval -0.009 to 0.012.
  • Query-weighted nDCG@10: Jev 0.738; Cohere 0.756.
  • Mean of dataset median request times: Jev 422 ms; Cohere 844 ms.

What the result does not establish

Scoring excludes queries without a relevant candidate. Providers use different routes. Times are means of dataset medians, not pooled medians. These rankings do not measure final answer quality.

What we would test in System One

Test a context shortlist with original snippets still available. System One allows eight questions per request, so this 30-question setup needs adaptation and a new evaluation.

This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.

Try a related workflow

Primary sources

Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.

Metric definitions and review method · Submit a correction or new result