Use cases· Last updated

Moderation failure modes with Jev

UGC needs a category, a severity, and an allow/review/remove decision. Jev scores the text you provide against your policy excerpt. Code enforces.

This unofficial page is the failure modes slice of the content moderation pack. Intent: apply the Jev (TypeSafe System One) decision model to content moderation failure modes. Primary search language: Moderation Jev failure modes. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.

Independent angle (cover ≠ clone): Policy-as-criteria + confidence abort + human pack — not a clone of a moderation-API landing page or rival recipe IA.

Moderation use-case context

Content moderation breaks in product-specific ways. This page lists those modes so you can write tests — not a generic “AI can be wrong” essay, and not a rival limitations-page clone.

Hub: Use cases. Compare, when the other tool is the real job: moderation APIs.

Failure Modes inputs

Many failures start as contract violations (distractors, missing user-generated post or message text). Canonical shape:

{
  "post": { "id": "p-209", "text": "…", "locale": "en" },
  "policy": { "hate": "…", "spam": "…", "illegal": "…" },
  "author": { "strikes": 1, "age_gate": "18+" }
}

Decision signals and actions

HTTP vs application:

You see Class Moderation move
401 / 422 / 429 / 529 Documented HTTP Fix key/body or back off — errors
200 + flat confidence or Noul ≈ 0.5 Low confidence Hold; do not remove content or issue a ban
Empty gather Missing evidence Skip Jev or ask “is enough information present?”

Do not treat a Noul of 0.5 as a “medium” content moderation score — it means yes and no are equally likely. Conjunctions stay in your code.

Guardrails and escalation

Fail closed: do not remove content or issue a ban. Schema-safe answers are not factual correctness. TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For content moderation, treat remove_or_ban as the high bar (removing content or issuing a ban). Tune on labels — see offline evaluation.

Evaluation and rollout notes

Your canary set should include each bullet above.

Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.

Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.

Pack map

Slice Page
Graph and primitives decision workflow
What may enter state input contracts
What to gather first evidence collection
Atomic rules policy checks
Act / review / abstain confidence thresholds
Reviewer payload human handoff
What to persist audit trail
How it breaks you are here
Labeled replay evaluation
Shadow → canary production rollout

FAQ

If the API returns 200, is the decision good? 200 only means the call parsed. Low confidence, Noul ≈ 0.5, or a policy miss are application failures.

Where do official weaknesses live? TypeSafe’s jev-1.13 jaggedness note — distractors, arithmetic, adversarial content. We do not invent more.

Where is the rest of the Moderation pack? Start with Moderation evaluation and Moderation decision workflow. Cluster hub: Use cases.

Should we replace our moderation vendor with Jev? Only after a labeled bake-off you run. This page does not publish one. See Jev vs moderation APIs.

Can Jev moderate images? Not directly. State is text. Run a vision system, put labels/transcripts in state, then ask typed questions.

What this page does not claim

Disclaimer

This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.

Primary documentation: https://docs.typesafe.ai. Hub: Use cases.

Sources

Public TypeSafe or adjacent documentation only. No private claims.