Moderation evidence collection with Jev
UGC needs a category, a severity, and an allow/review/remove decision. Jev scores the text you provide against your policy excerpt. Code enforces.
This unofficial page is the evidence collection slice of the content moderation pack. Intent: apply the Jev (TypeSafe System One) decision model to content moderation evidence collection. Primary search language: Moderation Jev evidence collection. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Independent angle (cover ≠ clone): Policy-as-criteria + confidence abort + human pack — not a clone of a moderation-API landing page or rival recipe IA.
Moderation use-case context
Evidence collection for content moderation happens before POST /v1/systemone. Jev does not browse your warehouse, retriever, or ESP. You gather the user-generated post or message facts, filter them, then ask snap questions. This slice is where fan-out cost math belongs: batch questions, do not re-send state.
Hub: Use cases. Compare, when the other tool is the real job: moderation APIs.
Evidence Collection inputs
Collect:
- The post text (or transcript you already extracted)
- The policy paragraphs the questions name
- Age-gate / locale if policy branches on them
Never send:
- Assuming Jev knows your community guidelines without pasting them
- Passing hashes of images and expecting a visual judgment
- Ten posts in one Choice
Shape the payload like this once the gather step finishes:
{
"post": { "id": "p-209", "text": "…", "locale": "en" },
"policy": { "hate": "…", "spam": "…", "illegal": "…" },
"author": { "strikes": 1, "age_gate": "18+" }
}
Decision signals and actions
Each evidence field should change a named answer:
| Id | Type | Job |
|---|---|---|
category |
Choice | ok / spam / hate / harassment / illegal / other |
severity |
Score | nuisance → severe harm |
allow |
Noul | Would a trained moderator leave this up given policy.*? |
Category + severity + allow on one post, one call. Extra Nouls for each policy atom beat one vague “is this bad?” Score.
Do not treat a Noul of 0.5 as a “medium” content moderation score — it means yes and no are equally likely. Conjunctions stay in your code.
Guardrails and escalation
If the gather step fails (empty user-generated post or message, redaction stripped everything, retriever empty), fail closed on removing content or issuing a ban. Do not invent evidence so Jev has something to say. TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For content moderation, treat remove_or_ban as the high bar (removing content or issuing a ban). Tune on labels — see offline evaluation.
Evaluation and rollout notes
Your eval set should include thin-evidence cases, not only happy user-generated post or messages. Label keep / review / remove gold from trained mods, plus category gold. Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.
Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Pack map
| Slice | Page |
|---|---|
| Graph and primitives | decision workflow |
What may enter state |
input contracts |
| What to gather first | you are here |
| Atomic rules | policy checks |
| Act / review / abstain | confidence thresholds |
| Reviewer payload | human handoff |
| What to persist | audit trail |
| How it breaks | failure modes |
| Labeled replay | evaluation |
| Shadow → canary | production rollout |
FAQ
Should evidence live in the question text?
Put facts in state and point instructions at post.text, policy.hate, policy.spam. Criteria stay stable so you can replay.
When do I split calls? Category + severity + allow on one post, one call. Extra Nouls for each policy atom beat one vague “is this bad?” Score.
Where is the rest of the Moderation pack? Start with Moderation input contracts and Moderation decision workflow. Cluster hub: Use cases.
Should we replace our moderation vendor with Jev? Only after a labeled bake-off you run. This page does not publish one. See Jev vs moderation APIs.
Can Jev moderate images? Not directly. State is text. Run a vision system, put labels/transcripts in state, then ask typed questions.
What this page does not claim
- Not a trust-and-safety certification.
- No published precision/recall.
- Not official TypeSafe.
- Official TypeSafe status, or that jev.pro issues API keys.
- That a schema-constrained answer is automatically factually correct.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.
Primary documentation: https://docs.typesafe.ai. Hub: Use cases.
Sources
Public TypeSafe or adjacent documentation only. No private claims.