Routing evaluation with Jev
Not every request deserves the same handler. Jev classifies cheaply; code sends work to a lookup, a specialist LLM, a costlier model, or a human.
This unofficial page is the evaluation slice of the intent and model routing pack. Intent: apply the Jev (TypeSafe System One) decision model to intent and model routing evaluation. Primary search language: Routing Jev evaluation. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Independent angle (cover ≠ clone): Routing glossary + compare vs rules/queues; ranked-route floors. Not a clone of a rival intent-routing glossary page.
Routing use-case context
Evaluation for intent and model routing is a frozen harness, not a vibe check and not an opinion-blog “Jev review.” Labels: handler gold and whether a human would have taken the complaint. We publish no unofficial accuracy.
Hub: Intent routing. Compare, when the other tool is the real job: rules and queues.
Evaluation inputs
Replay the same contract you ship:
{
"user": { "message": "Where is order 8831?", "locale": "en" },
"session": { "authenticated": true, "plan": "free" },
"catalog": { "handlers": ["order_status", "product_question", "return_exchange", "complaint", "other"] }
}
Freeze questions, criteria, and jev-1.13.0. Record the response model.
Decision signals and actions
Score these, not a blog-grade star rating:
- Wrong-handler rate at FLOOR
- Share of traffic that still hits the expensive model
- Human take-rate vs target — over-gating hides calibration
Pair auto-act errors with handoff rate. If the front-door router never acts, you have not evaluated intent and model routing — you have evaluated a human queue.
Do not treat a Noul of 0.5 as a “medium” intent and model routing score — it means yes and no are equally likely. Conjunctions stay in your code.
Guardrails and escalation
Promote a threshold only when the harness says auto-act error ≤ SLA and reviewers still catch the residual. TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For intent and model routing, treat complaint_auto_llm as the high bar (opening a refund/complaint LLM path or jumping to a costly model). Tune on labels — see offline evaluation.
Evaluation and rollout notes
After any criteria edit, rerun before production. Cookbook lifts you see on TypeSafe pages are vendor claims — re-measure on your latest user requests. Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.
Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Pack map
| Slice | Page |
|---|---|
| Graph and primitives | decision workflow |
What may enter state |
input contracts |
| What to gather first | evidence collection |
| Atomic rules | policy checks |
| Act / review / abstain | confidence thresholds |
| Reviewer payload | human handoff |
| What to persist | audit trail |
| How it breaks | failure modes |
| Labeled replay | you are here |
| Shadow → canary | production rollout |
FAQ
Will jev.pro publish a leaderboard for this use case? No. Measure on your labels. Vendor cookbook figures stay labeled as vendor claims.
What must stay frozen? Questions, criteria, and the pinned model id. Aliases can move.
Where is the rest of the Routing pack? Start with Routing failure modes and Routing production rollout. Cluster hub: Use cases.
Is intent a reserved API field?
No. You name the key. See glossary: intent.
When do rules beat Jev? When the route is already in structured data. Compare traditional routing.
What this page does not claim
- No claimed win-rate versus regex routers.
- LangChain
ModelRouterMiddlewareAPIs are LangChain’s — believe their docs if they move. - Not official TypeSafe.
- Official TypeSafe status, or that jev.pro issues API keys.
- That a schema-constrained answer is automatically factually correct.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.
Primary documentation: https://docs.typesafe.ai. Hub: Use cases.
Sources
Public TypeSafe or adjacent documentation only. No private claims.