Rendered at 08:05:14 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
AnthusAI 4 days ago [-]
That was a pretty simple task they gave it, and sure you can use BERT with sequence classification for simple classification tasks.
In our benchmarks, Jev did a LOT better at multi-step reasoning tasks than any open decision model we have tested so far, and it was also better than GLiDE which was specifically designed for that kind of task. And also better than Luna. On accuracy and also confidence calibration but also time and cost.
Just looking through your results, seems like gpt-6 luna was run with reasoning:off for a lot (all?) results. Seems like an unfair comparison.
zurfer 21 hours ago [-]
Not if you care about latency and cost. Reasoning is slow and expensive.
Sha1rholder 20 hours ago [-]
That looks super fair to me
NeumannGod 22 hours ago [-]
This is a false comparison. What Jev does is fundamentally different from what LLMs are doing as a judge.
For a quick primer, Jev is able to provide confidence scores on its classification, i.e. it is able to calibrate how well it is able to predict. Being able to predict in a distribution is different from being able to calibrate confidence of the predictions which should happen from the question or domain distribution from which the decisions are predicted - being able to do that is tough and is not same as using LLMs logit probabilities which are predictions in the vocab space. Though both are loosely correlated and might converge as LLMs keep getting better, the former is a much stronger decision-making signal than the latter. Jev not beating LLM-as-a-judge might be due to various other reasons such as world knowledge etc, but Jev as a concept will always provide more reliable decisions / outputs than LLM-as-a-judge giving a scalar score.
serbuvlad 19 hours ago [-]
This feels more like mythology though.
Claude and GPT tell me all the time what they are confident and what they're not confident in. This is partially useful, but it's not a guarantee of anything, since they can be confidently wrong or unsure and correct.
Same with Jav.
Arguing that Jav in some way more useful than current LLMs because it includes a confidence score smells of someone who hasn't updated their priors since GPT-4/4o and it's particular issues.
Jav is cool because it's very fast, very cheap, and "good enough". That's it.
esafak 19 hours ago [-]
> Same with Jav.
There are benchmarks for these things, which Jev and others have been put through; it's not a marketing boast!
tedivm 16 hours ago [-]
The benchmarks show it is on par or worse than even older LLM models.
* With jev-sec-bench the confidence numbers reached roughly the same as Qwen2.5-7b-instruct (confidence was off by about 6% on average in both cases).
* On the OpenRouter Banking 77 benchmark Jev did the worst out of all the models, being about 25 points off from reality.
I also think the fact that Jev hasn't published any real papers or actual technical details should make people a bit more skeptical than they have been.
If you don't care about that why wouldn't you use the flagships??
serbuvlad 14 hours ago [-]
I'm not arguing against Jev. Jev is very cool.
I'm arguing against the idea that Jev fixes a fundamental deficiency in LLMs by showing it's confidence.
It doesn't.
Jev can do a strict subset of what LLMs can do, but incredibly quickly and incredibly cheaply, which is certainly valuable.
esafak 14 hours ago [-]
Yes, it does. A low-latency calibrated classifier is valuable.
exe34 15 hours ago [-]
I think the point is if you can't get the right answer, it might not matter how fast you can make something up.
esafak 15 hours ago [-]
It had an accuracy of 96.5% with context! Come on, guys.
19 hours ago [-]
jvanderbot 20 hours ago [-]
Partially so. But the fact that we're comparing Jev to SOA LLM and pre-trained classification means it's essentially satisfied its reason for existence.
The benefits of Jev, and the reason I'm excited about them:
* Pre-trained/tuned to provide only structured output
* Small, fast, essentially trivial to locally run - this alone makes them an actual candidate for real-world planning systems. Laya is a local-first 400 million parameter model, for example.
* Their confidence scores are useful, even if those scores are not correct/true probabilities.
For me and my work, they look like nearly turn-key, tiny, fast, locally-hostable models that can manage state transitions in a deeply autonomous system.
LLM-as judge requires, well, what we know as a full LLM. A fully trained classifier model requires fully training - something you cannot / won't do for a embedded autonomous planning system (by contradiction - if this worked, we'd have used it everywhere already!). Totally unfair comparison in my mind.
monocasa 18 hours ago [-]
With llama.cpp, you can give it an ebnf grammar and constrain the output to any formal grammar you wish.
And how is the confidence score any different than the softmaxed token probabilities you get out of running every llm?
ekidd 17 hours ago [-]
> With llama.cpp, you can give it an ebnf grammar and constrain the output to any formal grammar you wish.
With llama-server, you can use basically any GGUF model as a Jev-like classifier. See Pi.dev's codemode and the accompanying llama backend to the "classify" function. The advantages of Jev-like inference are that:
- You generate 1 token per question in parallel, instead of sequentially generating a bunch of JSON punctuation. So you get much better latency and utilization.
- You can use a custom "sampler" that sees the raw logprobs for these tokens, including all the tokens that might have been generated. So you can generate some "probability" or "confidence" numbers based on how likely the model was to have generated a different answer.
So really, no difference? It's just a nice API with some computational efficiencies.
jvanderbot 18 hours ago [-]
> And how is the confidence score any different than the softmaxed token probabilities you get out of running every llm?
You're asking me how a structured output which assigns probabilities to classifications of the input text is different than the next-token weights which are calculated while generating that output?
One is internal (token prob), one is output generated by that internal mechanism e.g.,
You have a ebnf grammar constrained to evade, land, search etc. Quite literally what's the difference between the log probs between those options and whatever number a very tiny model is coming up for the confidence.
jvanderbot 15 hours ago [-]
Copy. The mechanism is probably similar once we abandon the "LLM that is asked to generate json" path. But the grammar alone doesn't ensure the output is a probability vector does it?
Both use the softmax operation, but "probability of emitting a given label" is not the same as "probability the decided action is correct". OTOH, taking an action to sample something to come up with "probability of emitting a label" is in fact the whole shebang if the label is the next action. If it works it works.
Who cares, this distinction is small.
For me, a planning guy, either your suggestion or Jev-like would work just fine, and you better believe I'll be testing both as I'm quite excited about being able to use less-structured, untrained systems in decision pipelines. Thanks for the heads up.
S0y 16 hours ago [-]
ebnf grammar is a truly underrated feature of llamacpp.
mrkn1 19 hours ago [-]
[flagged]
kwinkunks 21 hours ago [-]
I agree that there are differences between Jev and LLM-as-a-judge (e.g. Almeida's assertion that LLM probabilities have been irrevocably biased by RLHF), but I am not sure about your description of 'confidence'. Perhaps I misinterpret you or the docs, but I understand it as simply being computed from the probability distribution: https://docs.typesafe.ai/confidence#how-confidence-is-calcul...
fifilura 20 hours ago [-]
I don't know about Jev, but in my experience, 90% of the time, the confidence score such machine outputs (e.g. simplest case a kalman filter) is bogus because the model is wrong. Or not even wrong, but just not perfect.
Mathematicians build this, and they love the beauty of it so much that they loose touch of reality.
bunderbunder 18 hours ago [-]
Usably good calibration is hard. It's hard with logistic regression, it's even harder with linear support vector machines, and it makes me question my life choices with non-linear models.
It's not necessarily because the model is wrong. I've had trouble getting useful calibration out of models with an F1 of 0.9. The fundamental problem is twofold. First, it turns out that [0.0, 1.0] is a much larger set than {0, 1}. Second, the kinds of use cases where you care about calibration tend to be fussy and demanding.
As a fun anecdote, I got a chance to ask the person who invented the method scikit-learn uses for calibrated SVMs if he had any advice, and his answer was basically, "good luck."
All that said, sometimes you don't actually need calibration; you just need decent ranking. "Items with a score of 0.9 should be more likely to be positive than ones with a score of 0.5," is an easier requirement than "90% of items with a score of 0.9 should be positive." But I've had trouble getting that out of highly non-linear neural models, too, because they oftentimes produce results where the relationship between score and probability of being in the positive class is not even remotely monotonic.
But I haven't poked at Jev like this either, so I do have to allow that maybe they've found the secret sauce.
ViscountPenguin 20 hours ago [-]
Fine-tuning language models for calibrated probability predictions has been a thing since the Bert era though, that's really nothing special.
torginus 18 hours ago [-]
Yes, but they are not able to act on their decisions, or ask for elaboration.
Maybe someone did order those enlargement pills, and an agentic model would be able to scan email history and determine that, when classifying an email as spam.
qurren 16 hours ago [-]
The things that I'm actually most excited about Jev-like models are:
* Latency: when you want to make a decision within 100ms
* Robotics: I want to try make a VLA version of it and see what happens
* Costs: For something that needs a yes/no answer every 10 seconds, it seems like it would be cheaper than an LLM API.
21 hours ago [-]
saberience 18 hours ago [-]
This is hocus-pocus, Jev is basically a dumber LLM constrained to certain problems, it's not magic.
Jev's "confidence" scores are just as valid as an LLMs confidence score, i.e. just as likely to be wrong or hallucinated as any LLM.
Ultimately, people are somehow thinking Jev is mysteriously more likely to be right when ultimately it's another non-deterministic magic box except this time with a stricter interface slapped on top of it.
anthonypasq 17 hours ago [-]
you know absolutely nothing about the internal architecture of Jev. why are you being so confident?
ramoz 17 hours ago [-]
a less intelligent decision is not "more reliable"
wat10000 17 hours ago [-]
LLMs can provide confidence scores. Just ask it and it will provide.
Are those scores accurate? Probably not. But are Jev's scores accurate? They certainly aren't with a torture test like "I've rolled a fair die. What is the result?" (It comes back with ~80% confidence on 1 in my testing.)
It seems like people are very enthusiastic about Jev giving scores for its results, but I'm not seeing a whole lot of attention paid to how good those scores are.
wredcoll 16 hours ago [-]
That's an interesting test, but is it really what jev is supposed to be good at? I mean, this is a genuine question because I don't understand jev, but wouldn't it need a more constrained question?
wat10000 15 hours ago [-]
Probably not. But it does seem to have some idea of probabilities if I get a little more specific. "I've rolled a fair die. I see several dots on the top." Gives me 46% for 6, and about 16% for 4 and 5. If I say "The top is mostly white." then it says 1 with 50% confidence... and 6 with 31%.
Interestingly, if I add a "unknown" option then my original question says unknown with 100% confidence. "I see several dots on top" gives 89% unknown. "The top is mostly white" gives 30% unknown and 37% 1.
So it definitely has some ability in this area, but it's iffy. This isn't exactly what it's made for but it doesn't seem too far off from the examples like taking in a customer email and deciding which team it belongs with.
actualwitch 18 hours ago [-]
LLMs had ability to do this from the start, but at some point providers stopped exposing this api to prevent distillation. Just write the prompt, pass the data and sample 1 token returning logits, OSS inference engines can do it easily.
yieldcrv 21 hours ago [-]
for the uninitiated:
LLM’s are not able to give confidence scores, they make them up.
Your AI driven app is making that up. Your product manager and executive team’s demand for confidence in the UI is a totally fictional cosmetic telling them nothing. Your company sold bullshit confidence to your clients.
I’ve done this for many organizations that “formed a new team to work with the CTO on their AI strategy”, and the trappings are the same
You can have an LLM tell you how much of a schema it was able to get information about. And derive a “confidence” or level of compliance from the completeness of the schema
But this is layers upon layers of cruft that a classification model wouldn’t need
andy99 20 hours ago [-]
Softmax over logits doesn’t give calibrated probabilities either as a rule. I don’t want to comment specifically on Jev but as a rule it’s very hard to get good calibration because it’s somewhat in tension with minimizing training loss for neural networks, e.g. Guo et al (2017) https://arxiv.org/abs/1706.04599
While I know there are ways to improve calibration, I’d personally want to see a lot of evidence the probabilities were actually more meaningful before trusting them. I agree of course that asking an LLM to provide a confidence estimate is meaningless.
wood_spirit 17 hours ago [-]
My experience is that sota LLMs are pretty bad at confidence and priorities and things.
I write harnesses that do decision steps and regularly use competing models - from the major vendors and open source models.
And I have learned that if I can isolate a problem and put it into a straightforward prompt then Gemini 2.5 flash - the most basic and cheap commercial model - is actually very reliable. The newer models take longer, cost more and are really bad at saying they don’t know or nothing. You can ask them for a confidence score but they just hallucinate it!
I would like to try out decision models like Jev. The way the decision models approach the problems I work with seems a much more promising fit.
gopalv 16 hours ago [-]
> I write harnesses that do decision steps and regularly use competing models
Competing models is table stakes for decision making.
While doing the ablations for the "Team of Rivals" paper, I found that the models were very happy to ignore errors[1] when the data it is classifying looks like its own output.
Jev slots into the middle critique in the paper. The only reason we haven't flipped to Jev last week is because we're a HIPAA shop and Typesafe hasn't responded to our emails.
* They are worst at critiquing what has happened in the same model and same conversation
* they are slightly better at critiquing the output of another model it not given all the conversation
* they are better at critiquing a weaker or earloer model by the same vendor
* they are best at critiquing the output of another vendors model
I imagine it is about correlated bias in the training data and regimes etc.
I often run adversarial models in dyadic reasoning. But then have a feedback loop to get that same outcome from a cheap model. A kind of poor man’s distillation :)
pugio 12 hours ago [-]
I use Gemini-2.5-flash extensively, mostly for its low latency, and haven't found anything which works better at that speed/cost. Which is why I'm really bummed that Google is deprecaty/removing it next week.
Have you found anything that works as a suitable replacement?
wood_spirit 11 hours ago [-]
Hmm maybe I misread the mail - I thought they were removing it for new users next week, but turning it off for existing users in January or something?
Now I worry.
And of course I have no easy replacement for when it does go away :(
psadri 16 hours ago [-]
They are better at comparisons between two items vs scoring it in isolation. You can then turn a set of pairwise comparisons into numeric scores.
ActivePattern 16 hours ago [-]
Post-training of LLMs often destroys calibration. The objective is no longer to predict the next token, so it doesn't matter as much whether the probability you assign to next-token "A" is correct, as long as the overall answer is good.
So it's not too surprising that the most heavily post-trained frontier models are actually pretty bad at outputting calibrated decisions.
segmondy 4 days ago [-]
duh, this is not news. (general, fast and cheap) before decision models, you could pick only 2.
LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.
traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.
decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3
xfalcox 22 hours ago [-]
Doesn't the article covers the speed part by showing that Qwen 3.6 35A3B has lower latency and same accuracy?
jvanderbot 20 hours ago [-]
Laya, a local Jev alternative, is a 400Million parameter model. It can play doom.
Find me a another class of 0.4 B model that can handle structured output decision problems with the same latency and accuracy as Qwen 3.6.
lostmsu 3 days ago [-]
Is Jev any faster, cheaper, or more precise than medium LLMs like Luna 6?
ShinTakuya 24 hours ago [-]
Yes. Not as fast/cheap as Typesafe claims, but lots of benchmarks suggest around 3-4 times cheaper, and 7-14 times faster.
If you can do it offline, batching solves this. We used a test dataset from CFPB and at n=20, it was 1.6x faster and 1.2x more expensive with no statistically meaningful accuracy dropoff. Did not tune for batch size, but it's possible that we could get to 30 and see better perf.
ShinTakuya 13 hours ago [-]
That's not one if, it's two big ifs. If you can to it offline. If you can batch. I wouldn't say that's all that common in the use cases Jev sells itself for. You often want the answer now, not "averaged across 20 items in a batch that took 20 times longer".
It's fair to note it, but I wouldn't say "batching solves this" so much as "batching can solve time insensitive use cases".
saberience 18 hours ago [-]
No shit its cheaper and faster, thats the case with every small model.
It's a trade off between wanting something dumber but fast and cheap, or something slower and more expensive that is vastly more capable and smarter.
andy12_ 18 hours ago [-]
But the thing is that across a lot of classification benchmarks, Jev is actually faster, cheaper, and smarter than Luna (at least without reasoning). For example, look at this one from another user in this thread [1]
The reason the comparison is being made is because Jev and Luna 6 have similar levels of accuracy...
22 hours ago [-]
IanCal 24 hours ago [-]
Luna 6 is 10c per million input tokens and charges 5x that for output. Not sure if it’s still true but it used to be the case that structured outputs took time to process and cache which is relevant if the structure changes. It doesn’t give a percentage you can use for thresholds, and I’d want to know if Jen treats the questions as independent (they aren’t with luna, order of questions will change the result).
Jev is 4.2c/m tokens in and free out.
pokeapallascat 3 days ago [-]
not really
nicce 23 hours ago [-]
I don't think Luna is fast enough by any means
22 hours ago [-]
reexpressionist 4 days ago [-]
The key properties for using such models for conditional-branching decisions in agentic stacks (and related) is that they should be well-calibrated (under the definition chosen for the task) and informative (e.g., always predicting the mean might be "well-calibrated" in a theoretical sense for some chosen quantities of interest, but isn't particularly useful in practice).
The tricky thing with the neural networks is that the output logits are in effect a highly lossy compression of the epistemic (reducible) uncertainty, so even if the target calibration quantity is well-specified, it can be difficult to obtain in practice. A side-effect of this is that estimates in the high probability regions are not particularly stable under even modest co-variate shifts, which is a real problem if the estimates are being used for decision-making in a multi-step search graph that can lead to branches that are unlike what the model/estimator saw at training/calibration (if not altogether out-of-distribution). Here are a couple papers that describe how to approach those challenges:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics.
[2] Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following. In Proceedings of the ACM Conference on AI and Agentic Systems (CAIS '26). Association for Computing Machinery, New York, NY, USA, 1259--1269.
Garlef 23 hours ago [-]
I think it's a bit early to call the race.
I think the abstraction is a useful one - a general purpose classifier that does not need to be specifically trained: Unstructured signal in, structured judgement out - with a focus on speed and cost efficiency.
And since there is not yet a large body of benchmarks, I don't think we have sufficiently explored how to measure these things.
But since there's a market and some hype this will soon happen.
(And it's not like TypesafeAI has a real moat or invented something entirely new here ~ they just managed to put things into one coherent perspecive)
bjord 22 hours ago [-]
am I missing something or is there an incredible amount of title editorialization here? on the page itself (and within the slug), the title is:
"Benchmarking AI decision models against traditional guardrails"
kwinkunks 21 hours ago [-]
Agree. In fact Jev does reasonably well in both tests. In the article, the assertion is hedged, and followed up with a sentiment I liked:
> decision models like Jev do not reliably outperform [other] models in speed or accuracy. However, they rightly refocus industry attention on lightweight, task-specific inference paradigm that more closely resembles predictive machine learning
(I also noticed some issues with bolding in Table 4 that downplay Jev's performance a little.)
mertcikla 21 hours ago [-]
Jev may or may not have truly innovated on AI architecture but it still kick-started a new paradigm.
its a breather after waves of llm wrappers.
poincareball 20 hours ago [-]
[dead]
Havoc 23 hours ago [-]
Traditional classifier isn’t a direct equivalent though. Jev has some light abstraction/reasoning ability.
eg feed it a weather forecast and ask it whether I need an umbrella. It’s smart enough to make the connection between rain and umbrella.
So somewhere between classifier and fat LLM.
Ultimately boils down to right tool for the job
ibgeek 14 hours ago [-]
This blog post reads like the authors have a chip on their shoulder. They do present evidence that Jev may not be the best choice for these two particular scenarios, but then they draw broad conclusions from their tests that are not justified.
deepsquirrelnet 3 days ago [-]
BART is quite an old model for this kind of test, and probably not a good very good choice for much these days. I'm working on replicating their benchmark on my own NLI model that targets zero-shot guardrail applications. I don't think it'll beat much larger models, but should give a better baseline for what a crossencoder can do.
Jev (or equivalent) seem to me ideally placed for prototyping - prove the classification (say) works or is useful, then decide on how that should be made robust and economic with other options (eg local) brought in to the evaluation. Much quicker to get up and running, especially in greenfield scenarios.
fg137 20 hours ago [-]
Who removed "guardrail" from HN submission title?
6thbit 4 days ago [-]
Shouldn't LLMs intuitively be better with a high number of available options?
This article only does simple prompts with only options to block or not block.
What is openai doing for their decisions API, a finetuned luna?
dominotw 4 days ago [-]
prompts that these evaluations were done are too trivial
deadbabe 22 hours ago [-]
If you build a traditional classifier, and you have the data set for training curated or created hy an LLM, then you would essentially be building a classifier that judges the same way the LLM would?
petesergeant 23 hours ago [-]
There are plenty of benchmarks that show they do, too, though, so this is a single data point.
pfeifemc 15 hours ago [-]
[flagged]
chatformmycusto 18 hours ago [-]
[flagged]
chelseahermes 4 days ago [-]
[flagged]
22 hours ago [-]
AlphanimbleAI 18 hours ago [-]
[flagged]
agenttavern 21 hours ago [-]
[flagged]
aidiveyt 2 days ago [-]
The block/allow framing is the part I'd push on. I had a batch where automated checks passed all 99 outputs and reading each one by hand found 8 broken. The scorer only catches the failure modes its rubric already names, and a two-option guardrail bench inherits that ceiling whichever model sits behind it.
In our benchmarks, Jev did a LOT better at multi-step reasoning tasks than any open decision model we have tested so far, and it was also better than GLiDE which was specifically designed for that kind of task. And also better than Luna. On accuracy and also confidence calibration but also time and cost.
https://hard-decisions.anth.us/models/
For a quick primer, Jev is able to provide confidence scores on its classification, i.e. it is able to calibrate how well it is able to predict. Being able to predict in a distribution is different from being able to calibrate confidence of the predictions which should happen from the question or domain distribution from which the decisions are predicted - being able to do that is tough and is not same as using LLMs logit probabilities which are predictions in the vocab space. Though both are loosely correlated and might converge as LLMs keep getting better, the former is a much stronger decision-making signal than the latter. Jev not beating LLM-as-a-judge might be due to various other reasons such as world knowledge etc, but Jev as a concept will always provide more reliable decisions / outputs than LLM-as-a-judge giving a scalar score.
Claude and GPT tell me all the time what they are confident and what they're not confident in. This is partially useful, but it's not a guarantee of anything, since they can be confidently wrong or unsure and correct.
Same with Jav.
Arguing that Jav in some way more useful than current LLMs because it includes a confidence score smells of someone who hasn't updated their priors since GPT-4/4o and it's particular issues.
Jav is cool because it's very fast, very cheap, and "good enough". That's it.
There are benchmarks for these things, which Jev and others have been put through; it's not a marketing boast!
* With jev-sec-bench the confidence numbers reached roughly the same as Qwen2.5-7b-instruct (confidence was off by about 6% on average in both cases).
* On the OpenRouter Banking 77 benchmark Jev did the worst out of all the models, being about 25 points off from reality.
I also think the fact that Jev hasn't published any real papers or actual technical details should make people a bit more skeptical than they have been.
If you don't care about that why wouldn't you use the flagships??
I'm arguing against the idea that Jev fixes a fundamental deficiency in LLMs by showing it's confidence.
It doesn't.
Jev can do a strict subset of what LLMs can do, but incredibly quickly and incredibly cheaply, which is certainly valuable.
The benefits of Jev, and the reason I'm excited about them:
* Pre-trained/tuned to provide only structured output
* Small, fast, essentially trivial to locally run - this alone makes them an actual candidate for real-world planning systems. Laya is a local-first 400 million parameter model, for example.
* Their confidence scores are useful, even if those scores are not correct/true probabilities.
For me and my work, they look like nearly turn-key, tiny, fast, locally-hostable models that can manage state transitions in a deeply autonomous system.
LLM-as judge requires, well, what we know as a full LLM. A fully trained classifier model requires fully training - something you cannot / won't do for a embedded autonomous planning system (by contradiction - if this worked, we'd have used it everywhere already!). Totally unfair comparison in my mind.
And how is the confidence score any different than the softmaxed token probabilities you get out of running every llm?
With llama-server, you can use basically any GGUF model as a Jev-like classifier. See Pi.dev's codemode and the accompanying llama backend to the "classify" function. The advantages of Jev-like inference are that:
- You generate 1 token per question in parallel, instead of sequentially generating a bunch of JSON punctuation. So you get much better latency and utilization.
- You can use a custom "sampler" that sees the raw logprobs for these tokens, including all the tokens that might have been generated. So you can generate some "probability" or "confidence" numbers based on how likely the model was to have generated a different answer.
So really, no difference? It's just a nice API with some computational efficiencies.
You're asking me how a structured output which assigns probabilities to classifications of the input text is different than the next-token weights which are calculated while generating that output?
One is internal (token prob), one is output generated by that internal mechanism e.g.,
You have a ebnf grammar constrained to evade, land, search etc. Quite literally what's the difference between the log probs between those options and whatever number a very tiny model is coming up for the confidence.
Both use the softmax operation, but "probability of emitting a given label" is not the same as "probability the decided action is correct". OTOH, taking an action to sample something to come up with "probability of emitting a label" is in fact the whole shebang if the label is the next action. If it works it works.
Who cares, this distinction is small.
For me, a planning guy, either your suggestion or Jev-like would work just fine, and you better believe I'll be testing both as I'm quite excited about being able to use less-structured, untrained systems in decision pipelines. Thanks for the heads up.
Mathematicians build this, and they love the beauty of it so much that they loose touch of reality.
It's not necessarily because the model is wrong. I've had trouble getting useful calibration out of models with an F1 of 0.9. The fundamental problem is twofold. First, it turns out that [0.0, 1.0] is a much larger set than {0, 1}. Second, the kinds of use cases where you care about calibration tend to be fussy and demanding.
As a fun anecdote, I got a chance to ask the person who invented the method scikit-learn uses for calibrated SVMs if he had any advice, and his answer was basically, "good luck."
All that said, sometimes you don't actually need calibration; you just need decent ranking. "Items with a score of 0.9 should be more likely to be positive than ones with a score of 0.5," is an easier requirement than "90% of items with a score of 0.9 should be positive." But I've had trouble getting that out of highly non-linear neural models, too, because they oftentimes produce results where the relationship between score and probability of being in the positive class is not even remotely monotonic.
But I haven't poked at Jev like this either, so I do have to allow that maybe they've found the secret sauce.
Maybe someone did order those enlargement pills, and an agentic model would be able to scan email history and determine that, when classifying an email as spam.
* Latency: when you want to make a decision within 100ms
* Robotics: I want to try make a VLA version of it and see what happens
* Costs: For something that needs a yes/no answer every 10 seconds, it seems like it would be cheaper than an LLM API.
Jev's "confidence" scores are just as valid as an LLMs confidence score, i.e. just as likely to be wrong or hallucinated as any LLM.
Ultimately, people are somehow thinking Jev is mysteriously more likely to be right when ultimately it's another non-deterministic magic box except this time with a stricter interface slapped on top of it.
Are those scores accurate? Probably not. But are Jev's scores accurate? They certainly aren't with a torture test like "I've rolled a fair die. What is the result?" (It comes back with ~80% confidence on 1 in my testing.)
It seems like people are very enthusiastic about Jev giving scores for its results, but I'm not seeing a whole lot of attention paid to how good those scores are.
Interestingly, if I add a "unknown" option then my original question says unknown with 100% confidence. "I see several dots on top" gives 89% unknown. "The top is mostly white" gives 30% unknown and 37% 1.
So it definitely has some ability in this area, but it's iffy. This isn't exactly what it's made for but it doesn't seem too far off from the examples like taking in a customer email and deciding which team it belongs with.
LLM’s are not able to give confidence scores, they make them up.
Your AI driven app is making that up. Your product manager and executive team’s demand for confidence in the UI is a totally fictional cosmetic telling them nothing. Your company sold bullshit confidence to your clients.
I’ve done this for many organizations that “formed a new team to work with the CTO on their AI strategy”, and the trappings are the same
You can have an LLM tell you how much of a schema it was able to get information about. And derive a “confidence” or level of compliance from the completeness of the schema
But this is layers upon layers of cruft that a classification model wouldn’t need
While I know there are ways to improve calibration, I’d personally want to see a lot of evidence the probabilities were actually more meaningful before trusting them. I agree of course that asking an LLM to provide a confidence estimate is meaningless.
I write harnesses that do decision steps and regularly use competing models - from the major vendors and open source models.
And I have learned that if I can isolate a problem and put it into a straightforward prompt then Gemini 2.5 flash - the most basic and cheap commercial model - is actually very reliable. The newer models take longer, cost more and are really bad at saying they don’t know or nothing. You can ask them for a confidence score but they just hallucinate it!
I would like to try out decision models like Jev. The way the decision models approach the problems I work with seems a much more promising fit.
Competing models is table stakes for decision making.
While doing the ablations for the "Team of Rivals" paper, I found that the models were very happy to ignore errors[1] when the data it is classifying looks like its own output.
Jev slots into the middle critique in the paper. The only reason we haven't flipped to Jev last week is because we're a HIPAA shop and Typesafe hasn't responded to our emails.
[1] - https://github.com/t3rmin4t0r/critique-evals
* They are worst at critiquing what has happened in the same model and same conversation
* they are slightly better at critiquing the output of another model it not given all the conversation
* they are better at critiquing a weaker or earloer model by the same vendor
* they are best at critiquing the output of another vendors model
I imagine it is about correlated bias in the training data and regimes etc.
I often run adversarial models in dyadic reasoning. But then have a feedback loop to get that same outcome from a cheap model. A kind of poor man’s distillation :)
Have you found anything that works as a suitable replacement?
Now I worry.
And of course I have no easy replacement for when it does go away :(
So it's not too surprising that the most heavily post-trained frontier models are actually pretty bad at outputting calibrated decisions.
LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.
traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.
decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3
Find me a another class of 0.4 B model that can handle structured output decision problems with the same latency and accuracy as Qwen 3.6.
- https://www.ml6.eu/en/blog/jev-vs-gpt-6-luna-vs-bert-text-cl... - https://tessl.io/blog/jev-is-136x-faster-and-27x-cheaper-tha... - https://x.com/fazxes/status/2100300097695232164 (this last one is Luna 5.6 but that isn't too different from 6 besides accuracy and cost)
It's fair to note it, but I wouldn't say "batching solves this" so much as "batching can solve time insensitive use cases".
It's a trade off between wanting something dumber but fast and cheap, or something slower and more expensive that is vastly more capable and smarter.
[1] https://hard-decisions.anth.us/models/
Jev is 4.2c/m tokens in and free out.
The tricky thing with the neural networks is that the output logits are in effect a highly lossy compression of the epistemic (reducible) uncertainty, so even if the target calibration quantity is well-specified, it can be difficult to obtain in practice. A side-effect of this is that estimates in the high probability regions are not particularly stable under even modest co-variate shifts, which is a real problem if the estimates are being used for decision-making in a multi-step search graph that can lead to branches that are unlike what the model/estimator saw at training/calibration (if not altogether out-of-distribution). Here are a couple papers that describe how to approach those challenges:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics.
[2] Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following. In Proceedings of the ACM Conference on AI and Agentic Systems (CAIS '26). Association for Computing Machinery, New York, NY, USA, 1259--1269.
I think the abstraction is a useful one - a general purpose classifier that does not need to be specifically trained: Unstructured signal in, structured judgement out - with a focus on speed and cost efficiency.
And since there is not yet a large body of benchmarks, I don't think we have sufficiently explored how to measure these things.
But since there's a market and some hype this will soon happen.
(And it's not like TypesafeAI has a real moat or invented something entirely new here ~ they just managed to put things into one coherent perspecive)
"Benchmarking AI decision models against traditional guardrails"
> decision models like Jev do not reliably outperform [other] models in speed or accuracy. However, they rightly refocus industry attention on lightweight, task-specific inference paradigm that more closely resembles predictive machine learning
(I also noticed some issues with bolding in Table 4 that downplay Jev's performance a little.)
its a breather after waves of llm wrappers.
eg feed it a weather forecast and ask it whether I need an umbrella. It’s smart enough to make the connection between rain and umbrella.
So somewhere between classifier and fat LLM.
Ultimately boils down to right tool for the job
https://huggingface.co/dleemiller/crossingguard-nli-l
What is openai doing for their decisions API, a finetuned luna?