All capability theses

A use case to evaluate

Check a draft against named writing rules

With Jev, an agent can flag specific writing issues before revising the text.

A human example

A product description uses vague claims and never says what the app does.

What the caller supplies

The caller supplies the draft and separate rules for clarity, specificity and actionability.

What happens next

Jev returns advisory checks. The writing agent revises the draft and a reader reviews the result.

Illustrative example, not a recorded result.

Potential value: medium

Repeated editorial checks may be useful when a team has explicit rules, including the Unslop rules used in this project.

Evidence confidence: low

Official guidance proposes semantic linting. We have no blinded study of these writing checks.

The rating describes support for this claim. It is separate from Jev's returned probability. How we assign ratings.

Evidence, including disagreement

The next test

This protocol is planned. Its outcome is not yet known.

40 original paragraphs with blind human labels per rule, including intentional quotations and technical prose.

Compare against

  • Deterministic style checks
  • Direct agent review

Measure

  • Agreement per rule
  • Unnecessary revisions
  • Blind preference after editing
  • Time

Decision after the test

Keep only rules with useful agreement and low false-positive rates. Do not combine them into an unexplained writing score.

The report will retain inputs, question versions, every attempt and failure examples. We will update the confidence rating after reviewing the result.

Use a related Engine recipe

Recipes are implementation starting points. Their presence does not mean the protocol above has passed.

Read or improve this thesis on GitHub.