AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: 24 Jev Techniques For Exploring AI Decisions on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Thorsten Meyer published a map of 24 ways to use Jev, a tool that returns typed judgments for software workflows. He says three uses are already live in his publishing operation, 12 meet his four-part fit test, seven need measurement and two are poor fits.

Thorsten Meyer published a guide on September 29 mapping 24 proposed uses for Jev, a tool that returns typed answers to software questions, and classifying them by fit. Meyer says three uses are live in his publishing operation, 12 meet his criteria, seven need measurement first and two are poor fits; the examples matter to teams weighing automated decisions against human review.

Meyer describes Jev as a system that accepts text or JSON state and typed questions, then returns answers that code can use for branching. The output types include a yes-or-no probability, a choice among options with probabilities and confidence, or a score on ordered levels. He says Jev does not write or summarize content. A call carrying the state and questions takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens, according to his account.

The three live examples concern publishing: checking whether a story fits a site, checking whether an article is in English, and classifying a headline into one of 31 topics when a primary language model fails. Meyer reports that a scan of 78,889 articles cost $2.01; it found 1,576 non-English articles, of which 1,553 were fixed. For the classifier fallback, he reports 89% agreement with a frontier model overall and 97% to 99% agreement when Jev’s confidence was at least 0.8. These are figures from Meyer’s own measurements, not independently verified results.

His other examples span publishing, commerce, software, business operations and home use. Publishing candidates include checking disclosures, moderating comments, detecting thin sourcing and judging whether products fit a roundup. Meyer labels some of these strong fits, but says others need baseline measurements to show that existing rules fail in practice.

At a glance
reportWhen: Published September 29, 2026
The developmentThorsten Meyer published a guide assessing 24 proposed Jev applications, including three he says are running in his publishing operation.

24 use cases for Jev at a glance

Publishing, commerce, software, business operations and the home, sorted by fit.

Every use case, coloured by how well it fits

Start in the green. Amber needs a measurement first. Red fails at least one of the four conditions.
livestrong fitmeasure firstpoor fit

Proven in production

1Relevance gate: story and site2Language check3Classifier fallback

Publishing and content

4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderation

Commerce and support

10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triage

Software and AI systems

15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triage

Business ops and home

21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent

15 of 24 are ready to build or already running

3
12
7
2
Live
Strong fit
Measure first
Poor fit
Live: in my fleet today. Strong fit: meets high volume, narrow question, cheap errors and a visibly failing heuristic. Measure first: the failing heuristic is unproven.
From “24 Ways to Use Jev” on thorstenmeyerai.com. Figures are my own production measurements, September 2026, rounded, unless marked illustrative.

Where Confidence Can Guide Review

The proposed workflow is to let software handle clear, low-cost decisions and send uncertain cases to a stronger model or a person. That could make frequent checks affordable enough to apply broadly, while limiting automation’s reach where confidence is low. Meyer says the application code, rather than Jev, sets the action for each answer.

The distinction between a promising idea and a measured problem is central to the guide. A duplicate-story detector, for example, is labeled a poor fit because Meyer’s canary found no duplicates. If a tool finds nothing to correct, it may add cost and complexity without addressing a demonstrated failure. The examples therefore speak to evaluation and workflow design as much as to Jev itself.

Amazon

Top picks for "techniqu explor decision"

As an affiliate, we earn on qualifying purchases.

Meyer’s Four Tests for Fit

Meyer says a use case should meet four conditions before teams wire Jev into a workflow: high volume, a narrow question, errors that are cheap or can be routed for review, and a heuristic that has been shown to fail. He recommends replaying 300 to 500 past decisions in shadow mode, comparing results overall and by confidence band, and reviewing 20 disagreements to judge which answer was right.

For deployment, he proposes enabling automation only where the high-confidence group reaches 95% accuracy, using a separate feature flag that starts off, testing it on 5% to 10% of units, then expanding. These are Meyer’s recommendations. His article also reports measurements from his own operation, including about 90,000 decisions across the three live uses; the supplied material does not include independent audits or detailed methods for every result.

“Use Jev only when all four conditions hold: High volume. Narrow question. Cheap errors. A heuristic fails visibly.”

— Thorsten Meyer

Evidence Still Needed by Use Case

The article’s supplied material does not establish whether the reported cost, speed or agreement rates will hold in other organizations, content types or systems. It also does not provide independent validation of the measurements or full evaluation details, such as the composition of the 31-topic sample. Meyer’s results should be read as his own operational account.

Seven ideas are marked for measurement because a failing existing heuristic has not been demonstrated, and two are described as poor fits. The available source text ends partway through the section on commerce and customer operations, so it does not show all 24 use cases or their individual rules and evidence. The total category counts are reported, but the missing material limits assessment of the full list.

Measure Before Wider Deployment

Meyer recommends that teams test candidate workflows against historical decisions in shadow mode, inspect disagreements, and check performance within confidence bands before switching on automation. For approved uses, his proposed next steps are a feature flag, a 5% to 10% canary and a staged rollout. The article does not announce a new product launch or a broader deployment schedule; it presents a set of use cases and a method for evaluating them.

Key Questions

What is Jev, according to Meyer?

Meyer describes Jev as a tool that takes text or JSON state plus typed questions and returns answers, probabilities and confidence that software can use to make routing or classification decisions.

How many of the 24 uses does Meyer say are ready?

He says three are live in his publishing operation and 12 are strong fits. Seven need measurement first, while two are poor fits. He describes 15 as ready to build or already running.

What results does Meyer report from the live uses?

He reports scanning 78,889 articles for $2.01, finding 1,576 non-English articles and fixing 1,553. He also reports 97% to 99% agreement with a frontier model for topic classifications where Jev confidence was at least 0.8. These are his measurements.

How does Meyer recommend testing a use case?

He recommends replaying 300 to 500 past decisions in shadow mode, comparing results by confidence band and reviewing 20 disagreements. He says teams should enable a use only where the high-confidence band reaches 95% accuracy, then start with a small canary.

Source: ThorstenMeyerAI.com

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

So Reddit Has Decided That Plain HTML Is Unsafe

Reddit has announced it will disable the use of plain HTML in posts to enhance security, citing potential vulnerabilities and abuse.

Is Ticketmaster down? Ticketmaster outage for some

Ticketmaster reports a service outage affecting select users, causing difficulties in purchasing tickets. The issue is ongoing and being addressed.

The Question No To-Do App Can Answer

A new productivity tool, Threlmark, aims to prioritize work across multiple projects but cannot answer the fundamental question of what to do next.

Fox to buy Roku

Fox plans to buy Roku, marking a significant move in streaming and media consolidation, confirmed by sources. Details on the deal are still emerging.