Use cases· Last updated

Moderation policy checks with Jev

UGC needs a category, a severity, and an allow/review/remove decision. Jev scores the text you provide against your policy excerpt. Code enforces.

This unofficial page is the policy checks slice of the content moderation pack. Intent: apply the Jev (TypeSafe System One) decision model to content moderation policy checks. Primary search language: Moderation Jev policy checks. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.

Independent angle (cover ≠ clone): Policy-as-criteria + confidence abort + human pack — not a clone of a moderation-API landing page or rival recipe IA.

Moderation use-case context

A policy check is a typed question whose instructions + criteria are your rules about the user-generated post or message. Jev scores compliance; the moderation worker enforces. This is not a certification, and it is not a photocopy of a rival “policy engine” page — we keep rules atomic and ANDed in code.

Hub: Use cases. Compare, when the other tool is the real job: moderation APIs.

Policy Checks inputs

Put policy text and the artifact in structured state (never hope the model memorized last quarter’s PDF):

{
  "post": { "id": "p-209", "text": "…", "locale": "en" },
  "policy": { "hate": "…", "spam": "…", "illegal": "…" },
  "author": { "strikes": 1, "age_gate": "18+" }
}

Name post.text, policy.hate, policy.spam.

Decision signals and actions

Id Rule Enforce
illegal Illegal category → remove + escalate; do not auto-keep Choice + code
minors Age-gated surfaces use a stricter τ you own code floors
appeals Removes must be replayable from the audit pack logging

Typical primitives on the same request:

Id Type Job
category Choice ok / spam / hate / harassment / illegal / other
severity Score nuisance → severe harm
allow Noul Would a trained moderator leave this up given policy.*?
violations = [name for name, ans in policy_nouls.items() if ans.noul >= T_VIOLATION]
if violations:
    return review(violations)

Do not treat a Noul of 0.5 as a “medium” content moderation score — it means yes and no are equally likely. Conjunctions stay in your code.

Guardrails and escalation

Policy-in-state can be attacked (“ignore the policy”). High-risk removing content or issuing a ban still needs deterministic checks. TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For content moderation, treat remove_or_ban as the high bar (removing content or issuing a ban). Tune on labels — see offline evaluation.

Evaluation and rollout notes

Gold labels are policy-versioned. A criteria edit without replay is how silent false-allows ship. Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.

Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.

Pack map

Slice Page
Graph and primitives decision workflow
What may enter state input contracts
What to gather first evidence collection
Atomic rules you are here
Act / review / abstain confidence thresholds
Reviewer payload human handoff
What to persist audit trail
How it breaks failure modes
Labeled replay evaluation
Shadow → canary production rollout

FAQ

One Score for “compliant”? No. Atomic Nouls per rule, AND/OR in code. Money and dates: extract in code first (jaggedness).

If a regex can enforce it, should I still call Jev? Skip Jev. Official “how to build” guidance: keep deterministic rules in code when you can.

Where is the rest of the Moderation pack? Start with Moderation decision workflow and Moderation human handoff. Cluster hub: Use cases.

Should we replace our moderation vendor with Jev? Only after a labeled bake-off you run. This page does not publish one. See Jev vs moderation APIs.

Can Jev moderate images? Not directly. State is text. Run a vision system, put labels/transcripts in state, then ask typed questions.

What this page does not claim

Disclaimer

This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.

Primary documentation: https://docs.typesafe.ai. Hub: Use cases.

Sources

Public TypeSafe or adjacent documentation only. No private claims.