A practical guide to classifying text (support tickets, reviews, comments) using TypeSafe's Jev decision model, which returns typed decisions instead of generated text. Covers choosing between Noul (yes/no probability), Choice (one label from up to 255 options), and Score (position on a described scale) questions, writing effective label descriptions with examples, batching classification for hundreds of items via concurrency and speculative fan-out, handling large label sets with hierarchical tree walks or beam search, and building a hand-labeled test set to measure accuracy at different confidence thresholds. Also lists common mistakes like asking Jev to do math, vague score levels, and overstuffed state.

•16m read time•From flaviocopes.com
Post cover image
Table of contents
The quick answerHow do I set up the SDK?How do I classify one piece of text?Which question type should I use?How do I write labels that work?When should I add examples to a label?How do I classify hundreds of items?What if I have hundreds of labels?How do I know the labels are accurate?What are the common mistakes?Where do I go from here?

Questions this post answers

When should I use a Noul, Choice, or Score question with Jev for text classification?

Use a Noul for yes/no probability questions phrased so a high value means yes, a Choice for picking exactly one label from an unordered set of up to 255 options, and a Score for placing text on a described scale with 2 to 10 levels. For multi-label tagging where a comment can match none or several tags, use one Noul per label instead of a single Choice, since Choice always picks exactly one winner. Choosing the right question type matters before shipping a classifier; daily.dev surfaces more Jev implementation patterns as they appear.

How many labels can a single Choice question handle in Jev before I need a hierarchical approach?

A single Choice question supports up to 255 options and works reliably up to roughly 240 according to TypeSafe's classification cookbook, which sorted SEC annual reports into 75 industry groups with one question. Beyond that, or for tree-shaped categories like product taxonomies, walk the tree with one Choice per level, optionally using beam search to keep multiple candidate paths instead of a single greedy path. Developers designing large-scale classification taxonomies can track best practices like this on daily.dev.

What are common mistakes that reduce accuracy when classifying text with an LLM decision model like Jev?

Asking it to do math (counting links, date comparisons) fails since it isn't a calculator; instead count matches with one Noul per item in code. Vague score levels like 'somewhat' or 'very' underperform descriptive situational levels. Questions measuring two things at once (polite AND on-topic) need two separate Nouls combined in code, and stuffing unrelated content into the state lowers accuracy. Anyone debugging flaky AI classification results can compare these pitfalls against their own setup via daily.dev.

Share this post