Invoice confidence thresholds with Jev
Pixels never go to System One. An OCR or vendor extractor turns the PDF into strings; Jev then asks whether that text packet looks like a duplicate, a PO mismatch, or a complete AP file. Code posts to the ERP. Jev does not key invoices.
This unofficial page is the confidence thresholds slice of the invoice AP decisions pack. Intent: apply the Jev (TypeSafe System One) decision model to invoice AP decisions confidence thresholds. Primary search language: Invoice Jev confidence thresholds. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Independent angle (cover ≠ clone): OCR extracts; Jev judges already-extracted text (duplicate/vendor/PO language). Amounts and dates stay in code — not a clone of invoice-OCR IA or a rival AP-recipe page. Treat floors as production gates and price the false-reject cost — official 0.5/0.9 sketches are illustrations.
Invoice use-case context
Thresholds turn invoice AP decisions answers into act / review / abstain. They are product policy, not a hyperparameter TypeSafe ships. Official 0.5 / 0.9 sketches are illustrations. This slice also carries the false-reject discussion: over-gating invoice AP decisions hides calibration.
Hub: Use cases. Compare, when the other tool is the real job: invoice OCR.
Confidence Thresholds inputs
You need (1) pinned answers on a frozen contract and (2) labels for route_match / hold / duplicate gold from AP clerks, plus same-vendor gold. State shape:
{
"invoice": { "id": "INV-1044", "vendor_name": "Northwind Paper LLC", "memo": "Q3 copier paper — matches PO 88." },
"po": { "id": "PO-88", "vendor_name": "Northwind Paper" },
"extract": { "amount_cents": 128500, "invoice_date": "2026-03-12" },
"policy": { "duplicate": "Same vendor + same amount_cents within 14 days is a duplicate suspect." }
}
Decision signals and actions
| Axis | Where it lives | Invoice use |
|---|---|---|
choice / score / noul |
answer payload | What to do with the extracted invoice packet + PO text |
confidence |
Choice & Score only | Whether to trust the argmax |
| Distance from 0.5 | Noul | Whether same_vendor is decided |
FLOORS = {
"park_in_inbox": 0.52, # illustrations — replace
"post_to_erp": 0.90,
}
NOUL_TAU = 0.75 # for same_vendor
def allow(ans, action):
return ans.confidence >= FLOORS[action]
Do not treat a Noul of 0.5 as a “medium” invoice AP decisions score — it means yes and no are equally likely. Conjunctions stay in your code.
Guardrails and escalation
TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For invoice AP decisions, treat post_to_erp as the high bar (posting an invoice to the ERP). Tune on labels — see offline evaluation.
Band around 0.5 on same_vendor always reviews. Do not copy 0.75 onto Choice confidence.
Evaluation and rollout notes
- False-post rate on planted vendor mismatches (must be ~0 at your floor)
- False-hold throughput
- Handoff rate — if everything parks, floors are theater
Fit loop: pin jev-1.13.0 → replay → plot error vs confidence → pick floors where auto-act error ≤ your SLA. Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.
Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Pack map
| Slice | Page |
|---|---|
| Graph and primitives | decision workflow |
What may enter state |
input contracts |
| What to gather first | evidence collection |
| Atomic rules | policy checks |
| Act / review / abstain | you are here |
| Reviewer payload | human handoff |
| What to persist | audit trail |
| How it breaks | failure modes |
| Labeled replay | evaluation |
| Shadow → canary | production rollout |
FAQ
Should post_to_erp use 0.9 everywhere? No. Over-gating hides calibration and dumps the queue on humans. Fit per action.
Can I reuse a Noul τ as Choice confidence? No. Jaggedness: they are not interchangeable. See confidence.
Where is the rest of the Invoice pack? Start with Invoice decision workflow and Invoice human handoff. Cluster hub: Use cases.
Can we skip OCR and send the PDF? No. Official models page: no image input. Extract first. See Jev vs invoice OCR.
Should Jev decide if $1,285.00 matches the PO? No. Compare cents in code. Jev judges leftover language (vendor strings, memos).
What this page does not claim
- Not an OCR, ERP, or AP product.
- No invented extraction F1 or cycle-time lifts.
- Not official TypeSafe.
- Official TypeSafe status, or that jev.pro issues API keys.
- That a schema-constrained answer is automatically factually correct.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.
Primary documentation: https://docs.typesafe.ai. Hub: Use cases.
Sources
Public TypeSafe or adjacent documentation only. No private claims.