Citation failure modes with Jev
A generator produced a claim and a source pointer. Jev answers whether the source supports the claim. Code decides publish, hedge, or strip the citation.
This unofficial page is the failure modes slice of the citation checking pack. Intent: apply the Jev (TypeSafe System One) decision model to citation checking failure modes. Primary search language: Citation Jev failure modes. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Independent angle (cover ≠ clone): Cite vs generate boundary + failure modes. Cover the citation-check intent; do not mirror a citation-check-noul recipe slug.
Citation use-case context
Citation checking breaks in product-specific ways. This page lists those modes so you can write tests — not a generic “AI can be wrong” essay, and not a rival limitations-page clone.
Hub: RAG passage classification. Compare, when the other tool is the real job: citation services.
Failure Modes inputs
Many failures start as contract violations (distractors, missing claim + source excerpt text). Canonical shape:
{
"claim": "Pro plans include a 14-day refund window.",
"source": { "id": "kb-refunds", "quote": "Pro subscribers may request a refund within 14 days of purchase." },
"answer_draft": "Yes — you have two weeks on pro."
}
Decision signals and actions
- Citation services that only format APA/MLA do a different job — compare citation services.
- Support ≠ truth. The source can be wrong; Jev will not web-verify.
- Paraphrase drift: the draft says “two weeks” while the quote says “14 days” — ask an atomic Noul if you care.
- Generator-written citations without a quote in state.
- Unofficial accuracy % copied from blogs — do not.
HTTP vs application:
| You see | Class | Citation move |
|---|---|---|
| 401 / 422 / 429 / 529 | Documented HTTP | Fix key/body or back off — errors |
| 200 + flat confidence or Noul ≈ 0.5 | Low confidence | Hold; do not show a public citation next to the customer answer |
| Empty gather | Missing evidence | Skip Jev or ask “is enough information present?” |
Do not treat a Noul of 0.5 as a “medium” citation checking score — it means yes and no are equally likely. Conjunctions stay in your code.
Guardrails and escalation
Fail closed: do not show a public citation next to the customer answer. Schema-safe answers are not factual correctness. TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For citation checking, treat public_citation as the high bar (showing a public citation next to a customer answer). Tune on labels — see offline evaluation.
Evaluation and rollout notes
Your canary set should include each bullet above.
- False publish rate (unsupported citation shown)
- False strip rate (good citations removed)
- Reviewer disagreement — tighten criteria, not τ theater
Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.
Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Pack map
| Slice | Page |
|---|---|
| Graph and primitives | decision workflow |
What may enter state |
input contracts |
| What to gather first | evidence collection |
| Atomic rules | policy checks |
| Act / review / abstain | confidence thresholds |
| Reviewer payload | human handoff |
| What to persist | audit trail |
| How it breaks | you are here |
| Labeled replay | evaluation |
| Shadow → canary | production rollout |
FAQ
If the API returns 200, is the decision good? 200 only means the call parsed. Low confidence, Noul ≈ 0.5, or a policy miss are application failures.
Where do official weaknesses live? TypeSafe’s jev-1.13 jaggedness note — distractors, arithmetic, adversarial content. We do not invent more.
Where is the rest of the Citation pack? Start with Citation evaluation and Citation decision workflow. Cluster hub: Use cases.
Can Jev write the bibliography? No. It judges support. Formatting and URL fetching stay in code or another tool.
Is this the same as RAG passage relevance? Related but not the same intent. Relevance is “can this passage help?” Citation is “does this quote support this claim?”
What this page does not claim
- Not a plagiarism checker or fact API.
- No invented support accuracy.
- Not official TypeSafe.
- Official TypeSafe status, or that jev.pro issues API keys.
- That a schema-constrained answer is automatically factually correct.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.
Primary documentation: https://docs.typesafe.ai. Hub: Use cases.
Sources
Public TypeSafe or adjacent documentation only. No private claims.