Every Choice and Score answer from Jev, TypeSafe's decision model, includes a confidence value between 0 and 1 that indicates how concentrated its returned probabilities are. The guide explains how to split confidence into bands in code — act automatically above a high threshold, ask for human confirmation in a middle band, and hand off to a person below a low threshold — with the exact cutoffs tuned against labeled data rather than guessed. It covers why binary Noul (yes/no) answers skip confidence entirely since the single probability already captures certainty, why reading the full probabilities distribution matters beyond the single confidence number (e.g. detecting a torn Score answer), how to compute thresholds from labeled examples using a cost-based target, using a label hierarchy to fall back to a broader category when confidence is low, pinning a specific model version (e.g. jev-1.13.0) once thresholds are tuned since jev-latest can shift, and logging predictions in shadow mode before automating any decisions.

•16m read time•From flaviocopes.com
Post cover image
Table of contents
The quick answerWhat does confidence mean in Jev?Why doesn’t a Noul answer have a confidence score?Why read probabilities and not only confidence?How do I turn confidence into act, confirm, or hand off?How do I pick the thresholds from labeled data?What should I do when confidence is low?Why should I pin the model version once thresholds are tuned?How do I log answers in shadow mode before automating?

Questions this post answers

Why doesn't a Noul yes/no answer from Jev include a confidence score?

A Noul answer returns a single number, the probability that the answer is yes, and with only two possible outcomes that number already describes the whole distribution, so a separate confidence value would add nothing. The distance from 0.5 represents certainty: 0.95 is a confident yes, 0.04 a confident no, and 0.5 means the model cannot distinguish between them. In TypeScript, trying to read .confidence on a Noul response does not even compile. Readers building yes/no automation gates with daily.dev can dig deeper into Jev's Noul response design.

How do I pick a confidence threshold for automatically assigning support tickets to a team?

Start from the cost of a mistake: if a misrouted ticket costs 5 minutes to fix versus 1 minute for manual triage, automation pays off once more than 80% of automatic assignments are correct. Collect a few hundred labeled past messages, run them through the same questions, bucket the results by confidence, and pick the lowest threshold where accuracy still clears that 80% bar — refining as more data comes in. Developers tuning automation thresholds for support routing can use daily.dev to track this kind of practical AI-decision workflow.

Why should I pin a specific Jev model version like jev-1.13.0 instead of using jev-latest?

Because jev-latest is an alias that moves automatically when a new release ships, and thresholds tuned against one model's probability distribution can become wrong once that distribution shifts under a new version. Pinning the exact version, such as jev-1.13.0, keeps thresholds valid until you deliberately re-tune them by rerunning collection and threshold scripts against the new model and comparing results before switching. daily.dev helps developers stay on top of model versioning practices before an upgrade breaks tuned thresholds.

Share this post