🔍 Read the full analysis: What Are 24 Ways To Use Jev In AI Decision-Making? on ThorstenMeyerAI.com
Get privacy and security gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer says he has mapped 24 uses for Jev, a tool that returns typed, confidence-scored answers for software to act on. Three uses are running in his publishing operation, 12 are labeled strong fits, seven need measurement, and two are poor fits; the source lists only some of the 24 use cases.
Thorsten Meyer published a map of 24 potential uses for Jev, a tool that returns typed, confidence-scored answers to questions about text or JSON so software can make routine decisions. Meyer says three applications are already running in his publishing operation, while 12 other uses meet his stated fit criteria; the assessment matters to teams considering whether to automate frequent, low-cost decisions with an AI component.
Meyer divides the 24 uses into three live applications, 12 strong fits, seven that need measurement and two poor fits. The source describes the assessment as a map across publishing, commerce, software, business operations and the home, but the supplied material details only the three live publishing uses and part of the publishing list. It does not provide the full 24-item inventory.
In Meyer’s account, Jev takes a state, such as text or JSON, and typed questions, then returns answers that code can use directly. A yes-or-no question returns a probability; a choice question returns options with probabilities and confidence; a score question places a case on ordered levels. He says one call can carry several questions, takes about 0.3 to 0.9 seconds, and costs about $0.04 per million input tokens. These are claims in Meyer’s source, not independently verified figures.
The three reported live uses are a relevance gate, a language check and a topic-classifier fallback. Meyer says a scan of 78,889 articles cost $2.01, found 1,576 non-English items and fixed 1,553. He also reports about 10,000 site-and-story relevance pairings judged in three days, and 89% agreement between Jev’s fallback classifier and a frontier model across 31 topics. At confidence of 0.8 or higher, he says agreement was 97% to 99%; below 0.5, it was 42%. These results are from his own operation and measurement.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks May Help
The proposal centers on decisions that occur thousands of times but do not require a person or a large model to reason through every case. If the stated cost and speed hold in a particular workflow, using a smaller decision step for clear cases could make checks affordable at a scale where teams might otherwise sample or skip them. That could matter for publishing pipelines handling language, relevance, disclosure and moderation decisions.
Meyer’s approach also makes confidence part of the workflow. The tool supplies an answer, but the surrounding code sets the action: proceed on a clear result and send uncertain cases to a person or a more capable model. That division keeps policy and consequences with the system’s designers. For example, Meyer proposes that disclosure misses go to human review rather than being automatically published.
The categories are a useful constraint on adoption. Seven ideas remain unproven against existing heuristics, and two are labeled poor fits. Meyer’s example of event deduplication is instructive: a canary test found no duplicates, so he says the use case lacks evidence of a problem to solve. An inexpensive model call does not establish that a workflow needs one.
Meyer’s Four Tests for Jev
Meyer says a candidate should pass four conditions before a team wires Jev into a workflow: high volume, a narrow question, low-cost errors or a route for uncertain cases, and a heuristic that has been shown to fail. He recommends measuring that last condition rather than assuming it. If a simple keyword rule works, his guidance is to keep it.
For a promising use, Meyer proposes a shadow evaluation using 300 to 500 past decisions, comparisons overall and by confidence band, and review of 20 disagreements. He says to wire the tool in only where the high-confidence band reaches 95%. His rollout suggestion is a dedicated feature flag that starts off, followed by a canary covering 5% to 10% of units.
The detailed publishing examples illustrate the distinction between proposed and live uses. A thin-source detector and product matching for roundups are marked measure first; headline quality also needs measurement. Disclosure checks and comment moderation are strong fits in the source. For moderation, Meyer suggests auto-approving acceptable comments or hiding spam only at confidence of 0.9 or higher, with other cases queued for review.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer
Evidence Still Varies by Use
The supplied source does not show the full list of 24 use cases, so readers cannot assess every category or the evidence behind all 12 strong-fit labels from this material alone. It also does not provide the underlying evaluation data, the identity of the frontier model used for comparison, or independent replication of the reported accuracy, speed and cost figures.
Several publishing proposals are explicitly unvalidated. Meyer says the thin-source detector, product-fit matcher and headline-quality check need measurement; for the latter, he describes it as a pre-publication nudge rather than a sole gate. The source gives no measured error rate for the current heuristics in those cases. Results from Meyer’s publishing operation may also depend on its content, setup and decision rules, so they do not establish how Jev would perform in other organizations.
Measure Before Wider Rollout
Meyer’s proposed next step for unproven applications is to replay historical decisions in shadow mode, compare outcomes across confidence bands and examine disagreements before enabling automation. Where the high-confidence band meets his stated 95% threshold, he recommends a feature flag and a 5% to 10% canary. The supplied material does not announce a release date or a broader deployment schedule.
For readers evaluating the wider map, the missing details are the remaining use cases and evidence for each. The central question is whether a measured weakness in an existing rule justifies an automated decision step, and whether uncertain cases can be handled safely by the surrounding workflow.
Key Questions
What is Jev, according to Meyer?
Meyer describes Jev as a tool that takes text or JSON plus typed questions and returns probabilities, choices or scores that application code can use to make decisions.
How many Jev applications does Meyer say are live?
He reports three live uses in his publishing operation: relevance gating, checking whether article text is English and fallback topic classification.
What does Meyer mean by a strong fit?
He says a use should have high volume, a narrow question, low-cost errors or a safe path for uncertain cases, and an existing heuristic that demonstrably fails.
Are the reported performance figures independently verified?
The source presents them as Meyer’s own measurements. It does not include underlying data or independent verification.
What should happen when Jev is uncertain?
Meyer’s proposed pattern is to route unclear cases to a person or a more capable model, while code handles cases that meet a chosen confidence threshold.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
