🔍 Read the full analysis: Jev On The Potential Benefits Of 'System One' AI That Doesn’t Write Sentences on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
TypeSafe unveiled Jev, an AI model that replaces traditional text generation with structured, typed decisions. It aims to enhance automation speed and accuracy by focusing on decisions rather than language output. This shift challenges the assumption that all AI tasks require large language models, especially in enterprise automation.
TypeSafe has introduced Jev, a new AI model that does not generate text but instead provides structured, typed decisions with calibrated probabilities, aiming to transform enterprise automation. This development challenges the common assumption that large language models are necessary for all AI tasks, especially decision-making processes.
Jev is part of TypeSafe’s broader category of ‘System One’ models, inspired by Daniel Kahneman’s concept of fast, intuitive thinking. Unlike traditional large language models (LLMs) that produce prose, Jev processes structured questions and returns typed answers, such as ‘team: billing, confidence: 0.94,’ enabling direct software action without parsing text.
The model is built for decision automation, handling choices, scores, and yes/no probabilities, and is claimed to operate at speeds between 70 and 500 milliseconds while costing approximately $0.042 per million tokens. It is marketed as having ‘zero hallucinations’ in output, meaning it cannot produce off-schema or malformed responses, reducing errors in automated pipelines.
Jev’s launch was supported by $40 million in funding led by DCVC, with its creator Diogo Almeida, known for co-inventing RLHF at OpenAI. Almeida’s team argues that the reinforcement learning techniques used in chatbots are less suitable for automation, proposing instead a method called Reinforcement Learning for Calibrated Decisions (RLCD), which emphasizes decision accuracy over language fluency.
Jev vs. LLMs: who should make the call?
Jev, from TypeSafe AI, is a “System One” model. It doesn’t write text. It returns a typed decision with a confidence score that your software can act on directly.
Same support ticket, two kinds of answer
“This ticket appears most likely related to billing, although it could also concern account settings or a recent plan change. I would suggest reviewing the invoice history before…”
A person reads it, or code has to parse the prose.
team: "billing"Software reads it and acts. Nothing to parse.
How they differ
| LLM | Jev | |
|---|---|---|
| Output | Text written for people | A choice, a score or a yes/no probability |
| Speed | Seconds per call | 70–500 ms* |
| Price | Input and (pricier) output tokens | $0.042 per million input tokens, output free* |
| Knows when it’s unsure | Often sounds confident when wrong | Confidence score on every answer |
| Explains its answer | Yes | No, which matters for audits |
| Best at | Reasoning, writing, open questions | Routing, tagging, scoring, duplicate checks |
* Vendor-reported. TypeSafe also claims up to 194× faster and 445× cheaper on its own selected workflows.
Accuracy is something you build
Jev is far cheaper and faster, but not more accurate than frontier models. How you phrase the question matters a lot.
TypeSafe’s benchmark scores agreement with two frontier models rather than verified ground truth. The five-question result used weights fitted on 1,000 labelled examples.
The real idea: a confidence dial you control
“duplicate listing”, confidence 0.62
Raise the threshold for fewer mistakes and more manual review. Lower it for more automation and more risk.
Only use Jev when all four hold
Good fits
- Routing tens of thousands of support tickets a day
- Flagging duplicate listings in a product catalogue
- Replacing a keyword filter that mis-tags half its matches
Poor fits
- Drafting customer emails or release notes
- Reviewing a few high-stakes contracts a month
- Anything that needs a written explanation
Implications for Enterprise Automation and AI Decision-Making
The introduction of Jev signifies a major shift in enterprise AI, moving away from text-based outputs toward structured decision-making. This approach promises faster, more reliable automation, reducing the need for human oversight and parsing errors. The model’s speed and low cost could expand automation use cases, especially in high-volume, decision-critical workflows.
By focusing on typed decisions with probabilistic confidence, Jev aims to improve the reliability of automated systems, potentially lowering operational risks associated with hallucinations and formatting errors common in traditional LLMs. However, its accuracy depends heavily on how well it is trained and integrated into specific workflows, and it does not eliminate judgment errors entirely.
enterprise decision automation AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on AI Models and Enterprise Automation Trends
Over the past three years, the AI industry has focused heavily on developing larger, more capable language models like GPT-4 and Claude, promising better reasoning, longer context, and improved code generation. These models have become central to enterprise AI strategies, often used for support, content creation, and decision support.
However, critics argue that LLMs suffer from issues such as mode dropping, overconfidence, and hallucinations, which can compromise reliability in critical applications. In response, some experts, including Diogo Almeida, have questioned whether language models are always the best fit for automation tasks that require precise decision-making rather than natural language output.
TypeSafe’s Jev is a direct response to these concerns, emphasizing decision precision, schema conformance, and speed, and challenging the industry assumption that text generation is necessary for all AI-driven automation.
“Our goal with Jev is to produce decisions that software can act on directly, eliminating the ambiguity and errors associated with text-based outputs.”
— Diogo Almeida, CEO of TypeSafe
As an affiliate, we earn on qualifying purchases.
Unconfirmed Aspects of Jev’s Performance and Reliability
While TypeSafe claims Jev has ‘zero hallucinations’ and demonstrates high speed and low cost, independent benchmarks show mixed results. In a phishing email detection test, Jev scored 62.6%, significantly lower than Claude Haiku 4.5’s 81.3%. Its accuracy depends on how well it is trained and the specific task context. The model’s overconfidence on some questions and underconfidence on others indicate that its calibration may require further refinement. It is also not yet clear how Jev performs across diverse real-world tasks or how it handles complex, ambiguous decisions.
decision-making AI models for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Adoption and Validation
TypeSafe is expected to expand testing of Jev in various enterprise workflows to validate its accuracy and reliability outside initial benchmarks. Industry observers will look for independent evaluations and real-world case studies demonstrating its effectiveness in automating decision-critical tasks. The company may also refine its training methods and calibration techniques to improve performance further.
Meanwhile, competitors and AI researchers will monitor whether structured decision models like Jev can scale across different domains and whether their advantages outweigh the limitations compared to traditional language models. The broader industry will assess whether this approach can become a new standard for enterprise automation.
automated decision processing software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
How does Jev differ from traditional language models?
Jev produces structured, typed decisions with probabilities instead of natural language text, enabling direct software actions and reducing parsing errors.
Is Jev more reliable than existing AI models?
Jev aims to be more reliable by eliminating hallucinations and formatting errors, but its accuracy still depends on proper training and task-specific calibration.
Can Jev handle complex or ambiguous decisions?
Jev is designed for decision tasks that can be broken into clear, small questions. Its effectiveness on complex or ambiguous issues remains to be seen and is subject to further testing.
What are the cost and speed advantages of Jev?
Jev answers in 70-500 milliseconds at about $0.042 per million tokens, making it significantly faster and cheaper than traditional LLMs for decision tasks.
Will Jev replace large language models in enterprise AI?
Jev is positioned as a complementary approach for decision automation, especially where speed and reliability are critical, but it may not replace language models in all applications.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
