# Jev for prompt classification and routing

**Evidence checked 29 September 2026.** Jev looks useful as a fast decision step when an application already knows its possible routes and can supply concise context. The strongest directly relevant public head-to-head test found it **3.3 percentage points less accurate** than Claude Opus 5 on 77-way intent classification, but about **13 times faster** at median client-observed latency and **22 times cheaper** per 1,000 requests in that setup. A separate comparison against a *cheap* LLM found a smaller or absent cost advantage. These results do not establish a universal accuracy ranking: label descriptions, retrieved examples, model version, and the fallback policy materially change the outcome. [OpenRouter Banking77 test](https://openrouter.ai/blog/insights/jev-vs-claude-opus-5-classification/); [independent Jev benchmark](https://github.com/PavelRavvich/jev-bench).

## What Jev actually does

Jev 1.13 accepts text as a string, object, or array of text fields in a `state`, plus one or more independently evaluated questions. **Choice** selects one supplied label and returns a distribution and confidence; **Noul** returns a probability for a yes/no question; **Score** returns a position on a supplied ordered scale. It does not write a reply, code, or an explanation. It cannot choose a label outside the supplied Choice set, which prevents a type or enum error but does not prevent a wrong route. Add an explicit `other` or review option when none of the business routes fits. [TypeSafe primitives](https://docs.typesafe.ai/primitives); [System One documentation](https://docs.typesafe.ai/concepts/system-one).

TypeSafe lists `jev-1.13.0` at **$0.042 per million input tokens**, with no billed output tokens, a **64k-token request budget** and a **32k budget for state plus the longest question**. Its published limits are 1,200 requests/minute and 250,000 tokens/second, subject to change. `jev-latest` is a moving alias; pin a version when comparing results or setting confidence thresholds. Jev is a hosted API model; I found no publicly released weights or documented local deployment. Qwen below supplies a genuine local alternative. [TypeSafe model reference](https://docs.typesafe.ai/models).

Jev's probability is a signal to validate on *your* labeled traffic, not a guarantee that an individual answer is correct. TypeSafe itself documents weaker performance on literal or contradictory instructions, counting and arithmetic, date comparisons, several reasoning hops, irrelevant long context, adversarial text, and non-English input relative to English. Keep exact calculations and policy rules in code, and do not treat untrusted prompt text as routing instructions. [TypeSafe confidence guide](https://docs.typesafe.ai/confidence); [Jev 1.13 known limitations](https://docs.typesafe.ai/model-jaggedness/jev-1.13); [model language support](https://docs.typesafe.ai/models).

## Comparable options

| Route decision method | Accuracy evidence | Latency and cost evidence | Deployment and characteristic failure |
| --- | --- | --- | --- |
| **Jev 1.13 Choice** | **81.0%** on all 3,080 Banking77 test messages with 77 one-line label descriptions; zero invalid labels. Another Jev experiment reached **92.40%** on the same test split after developing definitions and retrieving 24 relevant labeled examples per message; that is a different, more informed setup. | In OpenRouter's paired Banking77 run: **175 ms p50**, **270 ms p95**, **$0.11/1,000** calls. Per-token price above is only one cost component; every option description consumes input tokens. | Hosted TypeSafe API or OpenRouter Decisions API. Narrow, typed judgments; can confidently choose the wrong label, especially when taxonomy is incomplete or instructions are misleading. |
| **GPT-6 Luna, structured JSON** | In one independent 500-item Banking77 sample, **3.4 points more accurate** than Jev (paired interval for Jev minus Luna: −5.8 to −1.2 points). Both produced valid labels. No comparable full-test result was found. | Independent study: Jev was about **5–6× faster** at median, while cost on 77-way Banking77 was roughly **at parity**. Current first-party rate: **$0.10/M input**, **$0.50/M output**; caching and short outputs matter. | Hosted OpenAI API; generated JSON needs schema validation and retry/fallback handling. Its reported confidence is verbalized rather than Jev's native distribution. |
| **Claude Opus 5, constrained enum** | **84.4%** on the same 3,080 Banking77 messages, 3.3 points above Jev; zero invalid labels in that run. Tested with reasoning off and prompt caching. | Same run: **2,266 ms p50**, **3,004 ms p95**, **$2.42/1,000** calls with the shared label list cached. These figures are historical measurements for Opus 5; Opus 5.5 is now available and should be benchmarked separately. | Hosted Claude API and listed cloud platforms. More capacity for ambiguous or multi-step cases, with higher latency and cost; a prompt or taxonomy error can still cause misrouting. |
| **Qwen3.5-4B, constrained output** | No trustworthy Jev head-to-head routing result found for this exact model and task. | No comparable latency or per-call cost: hardware, quantization, batching, and hosting dominate. | Apache-2.0 weights can be served locally with vLLM, SGLang or other compatible runtimes. Offers data locality and custom tuning, but requires model operations and a task-specific benchmark. |
| **Deterministic rules** | 13/16 on the small disclosed control below. This is an illustration, not a population estimate. | Negligible marginal token cost; local timing would not be comparable with hosted API round trips. | Fully local and auditable. Exact patterns miss paraphrases and need a review route for overlap. |

The Banking77 Jev–Opus numbers and methods come from [OpenRouter's full-test report](https://openrouter.ai/blog/insights/jev-vs-claude-opus-5-classification/). The higher Jev score with retrieved examples comes from a [separate public protocol and run](https://github.com/simonmesmith/jev-banking77-experiment). GPT-6 Luna's relative results come from an [independent two-task benchmark with code and predictions](https://github.com/PavelRavvich/jev-bench); its current list price and API features are in [OpenAI's model reference](https://developers.openai.com/api/docs/models/gpt-6-luna). [Anthropic documents the newer Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5), and [Qwen's model card](https://huggingface.co/Qwen/Qwen3.5-4B) documents its license and serving options.

### What the evaluations establish

OpenRouter used the **full Banking77 test split**, 40 items per intent, with the same one-line criterion for each of 77 labels. Jev used Choice; Opus used a strict JSON enum, reasoning disabled, and a cached system prompt. They ran from the same client with eight concurrent requests. OpenRouter reports 175 cases only Opus got right and 72 only Jev got right; the paired difference is 103/3,080 = **3.34 points**. Confidence-based fallback to Opus at a Jev threshold of 0.90 yielded **84.0% accuracy at $0.69/1,000** in that same dataset, but the threshold was assessed on the data used to choose it, so treat that as exploratory. [OpenRouter methods and cascade](https://openrouter.ai/blog/insights/jev-vs-claude-opus-5-classification/).

The independent 500-item comparison found Jev roughly on par with GPT-6 Luna for binary spam, while Luna led on 77-way intents. Jev's cost advantage was large against an expensive frontier model but roughly disappeared against Luna on the long 77-label schema. The study also reported that Jev's aggregate calibration error was *not* consistently better than the LLM baselines; nevertheless, its confidence ranked easier and harder cases usefully. These are one-run, public-dataset results from one prompt design and one client location. [Independent study and limitations](https://github.com/PavelRavvich/jev-bench/blob/main/article/jev-article.md).

TypeSafe's headline **193.6× faster and 444.6× cheaper** figure comes from its own multi-question workflow comparison against frontier models, using frontier-model consensus as the reference rather than independently adjudicated ground truth. It is evidence for that workload, not a general routing accuracy claim. LangChain's reported 500/500 Jev pass/fail agreements were **100 repetitions of only five fixed weather-agent cases**, so they measure stability on five examples more than broad generalization. A separate OpenRouter closed-rubric judge test was near tied on 88 items (Jev 84 correct, GPT-5.6 Luna 83), while the LLM followed expert ratings better on 50 longer-summary consistency judgments (Spearman 0.72 versus Jev 0.47). None of those judge tasks substitutes for a prompt-routing test. [TypeSafe workflow evaluation](https://evals.typesafe.ai/); [LangChain methods](https://www.langchain.com/blog/jev-agent-evals-langsmith); [OpenRouter judge study](https://openrouter.ai/blog/tutorials/jev-vs-llm-as-a-judge/).

## Small reproducible check

I made **16 original English support prompts** with a fixed rubric: route a single clear request to `billing`, `technical`, or `account`; route vague or cross-team requests to `review`. A disclosed keyword rule got **13/16** right. Its three misses were clear paraphrases without its trigger words: an unexpectedly doubled total, a blank screen after login, and a request to change a contact address. This demonstrates a rule-maintenance failure case; the sample is deliberately tiny and hand constructed and gives **no Jev or LLM accuracy estimate**. The attached script defines every prompt and label and writes per-case predictions.

As a separate reanalysis of *published* model results, I used OpenRouter's paired Banking77 discordance counts (Opus-only 175; Jev-only 72) in the exact two-sided McNemar/binomial test. The result is **p = 4.51 × 10⁻¹¹** for equal error rates on that set. It independently checks the arithmetic and the strength of the *reported* paired gap; it does not validate the underlying predictions or fix benchmark design choices. The attached script recomputes this value. No live Jev or LLM call was made: this session had no provider API credentials, and notebook access to GitHub's raw prediction files was rejected by the network policy. Latency and cost in the table are therefore cited measurements, not new measurements from this session.

## Deployment decision

For a **closed, stable taxonomy at high volume**, trial Jev with an explicit `other/review` route, concise state, version pinning, and confidence thresholds calibrated on held-out production labels. Budget the *whole* request, including every option description and any fallback call. For **77-way or otherwise fine-grained routing**, test a cheap structured LLM such as GPT-6 Luna first as well: its quality lead and near-cost parity in one independent study make Jev's advantage task dependent. Use a stronger LLM when errors are expensive or the decision needs multi-step reasoning or a written explanation. Use local Qwen if data locality or control over weights is decisive, after measuring its actual accuracy and throughput on your traffic. Keep exact conditions in code and send uncertain or consequential cases to review. These are recommendations inferred from the cited evidence, not a vendor benchmark result.
