Jev is designed as a 'System 1' reflex classifier returning JSON decisions in milliseconds, not a text-generating LLM, making it viable for routing, tagging, and verification tasks. The model costs roughly $0.04 per million input tokens with near-zero output costs, enabling developers to embed structured decisions throughout software pipelines at negligible expense.
Jev Reframes AI Economics: The Rise of the $0.04-Per-Million Classifier
A smart if statement, not a reasoning engine
So Jev is just a classifier? Why does that matter?
Because it's the first time someone has built a model specifically for that job instead of forcing a general-purpose model to do it. It's 444 times cheaper and 193 times faster. That changes what becomes economically viable.
But wait—the benchmarks are LLM-generated ground truth. How do we know it's actually right?
That's the real question. The zero hallucination claim is about format, not correctness. It will never return malformed JSON, but it can be confidently wrong.
So when would you actually use this instead of just writing an if statement?
When the decision is too complex for a simple rule but doesn't need reasoning. Spam detection, email routing, content tagging. Things that happen thousands of times a day.
And the cost structure—is that sustainable? Or does TypeSafe just have a loss leader?
That's the bet. They're betting on Jevons paradox. Make it cheap enough and developers embed it everywhere, so volume makes up for margin.
What's the biggest misuse you've seen?
People trying to use it as a judge for comparing outputs from reasoning models. It just doesn't have the context to do that well.
Is there a risk that this becomes a crutch? That developers stop thinking about whether they actually need a model at all?
Absolutely. Browne's point is that the real value is in the orchestration layer—knowing when to use Jev and when to use something smarter.
What happens next?
Image support. Once Jev can see, you can scan video frames for sensitive data before it gets published. That's where it gets interesting.
Der Puls
- Jev costs $0.04 per million input tokens; frontier models cost $0.20 to $10 per million
- Jev processes classification tasks in 70-500 milliseconds; traditional LLMs take 3-329 seconds
- Browne classified 32,311 messages across 1,118 chat threads for $37 total cost
- TypeSafe AI founded by Diogo Almeida, ChatGPT co-inventor
Jev is designed as a 'System 1' reflex classifier returning JSON decisions in milliseconds, not a text-generating LLM, making it viable for routing, tagging, and verification tasks. The model costs roughly $0.04 per million input tokens with near-zero output costs, enabling developers to embed structured decisions throughout software pipelines at negligible expense.
TypeSafe AI's new Jev model prioritizes speed and cost efficiency for classification tasks over general reasoning, operating 193.6x faster and 444.6x cheaper than frontier models for structured decision-making.
Theo Browne has spent the last few years watching artificial intelligence labs pitch the same story: one model, smart enough to do everything. That assumption carries a hidden tax. Every time a piece of software needs to make a simple choice—is this email spam? Which tool runs next?—it summons a model with billions of parameters, waits several seconds for it to think, then discards the reasoning. The economics of that approach, Browne argues, are about to shift.
The catalyst is Jev, a new model from TypeSafe AI, founded by Diogo Almeida, who helped create ChatGPT. Jev does not write essays or generate code. It does not stream words across a screen. Instead, it takes structured input and returns a decision in JSON format—a typed, probabilistic answer to a yes-or-no question. Browne describes it with deliberate bluntness: roughly as intelligent as a switch statement. He means it as no insult. It is the entire business model.
The numbers tell the story. In TypeSafe's testing, Jev runs 193.6 times faster than frontier models for classification work and costs 444.6 times less. Input tokens cost about four cents per million; output tokens are, for practical purposes, free. The machine is built for the unglamorous middle of every software pipeline: routing, ranking, tagging, verification. Browne frames this using Daniel Kahneman's distinction between System 2 thinking—deliberation, planning, writing—and System 1 thinking—the instant recognition that a shirt is blue. Jev is a System 1 machine. If you can answer a question in under ten seconds by looking at something, Jev can probably handle it. If the answer takes longer, it probably cannot.
The speed is not marketing rhetoric. In one real test, a developer classified 100 emails using eight parallel workers. The average response time was 200 milliseconds per email; the 95th percentile was 240 milliseconds. The system processed 38 emails per second. That is fast enough for a spam filter or routing layer to run invisibly, before a human notices any delay. Browne ran his own experiment, feeding 32,311 messages across 1,118 Claude Code chat threads into Jev for classification. The total cost was thirty-seven dollars. The output revealed that nearly half of his coding work was bug fixing and PR management—a discovery that would have been prohibitively expensive with a traditional language model. Jev also booked a flight by reading raw HTML in under 7.1 seconds, navigating links and selecting actions without generating a single explanatory sentence.
The model's marketing emphasizes "zero hallucinations," and Browne agrees with the framing, though with precision. Jev cannot return malformed JSON or invent an option outside a predefined schema. If you ask it to choose between cat, dog, or bird, it will never say elephant. That is a guarantee about format, not correctness. What matters in production is that a hallucinated tool call in an agent modifying a database is catastrophic. A model that returns confidently wrong answers is equally dangerous, but the failure mode differs. Jev returns a calibrated probability with every output. If it says 90 percent confidence, a developer can build a threshold around that number. Traditional language models, by contrast, tend to be overconfident and inconsistent.
Browne has grown frustrated watching the community misuse Jev in two viral demonstrations. The first is using it as an LLM judge—replacing expensive human evaluation with Jev for scoring outputs from other models. Browne's response is sharp: those evaluation prompts take thirty seconds to write, and judging between multiple complex implementations requires navigating a codebase, growing context, and synthesizing nuance. Handing three intricate implementations to a classifier that emits JSON in 200 milliseconds is not evaluation; it is a guess. The second misuse is context compaction—using Jev to decide which lines of conversation history to delete to save tokens. Browne's objection is technical and layered. Compaction is not filtering; it is synthesizing a summary. Jev lacks the reasoning traces behind decisions because labs no longer share them. And the economics break: editing early history invalidates downstream cache writes, which are far more expensive than reads.
The right mental model, Browne argues, is a function in a codebase—a smart if statement, not a tool you invoke in a chat window. It is a library you install. In a complex agent task with 100 model calls, 90 of them are routing, status checks, and verification. Every one is a candidate for Jev. The remaining 10—planning, code generation, deep reasoning—go to a frontier model. The orchestration layer that routes between them becomes the valuable asset. Browne sees the most exciting frontier as image support. Once Jev can see, he predicts it will enable scanning video frames for personally identifiable information before publication. Computer use will become far more compelling when a fast model can read a screen and choose the next click without streaming text.
Browne's deepest argument is not about Jev itself but what it signals. He rejects the claim that AI is getting more expensive. Cost per token for top models has risen five to ten times. But cost per unit of intelligence has collapsed. DeepSeek V4 Flash scores roughly equal to GPT-5 Codex High on overlapping benchmarks, at 21 times cheaper on input and 55 times cheaper on output. His own spending tells the story. In 2025, he spent about $1,000 on AI tokens. In 2026, he has days where he pulls $1,000 in a single day. The bills are not rising because models are more expensive; they are rising because developers are using expensive models more. That is Jevons paradox applied to cognition. When reasoning gets cheap, you do not spend less. You consume more. The companies that win are not necessarily the ones with the smartest single model. They are the ones that consume intelligence in bulk and route it intelligently. A classifier like Jev does not replace a reasoning model. It replaces the waste of using a reasoning model for a reflex.
Bemerkenswerte Zitate
If I show you a picture of a person wearing a shirt, how many seconds does it take for you to say what color it is? If the answer is under 10 seconds, this model's probably good for it.— Theo Browne, describing Jev's appropriate use cases
Our bills aren't going up because the models are more expensive. They're going up because we're using the expensive models more.— Theo Browne, on the economics of cheap AI