Incident response evaluation with Jev
Prometheus rules and PagerDuty already page on numeric thresholds. Jev is optional on messy customer-reported incidents or multi-alert narratives: which SEV band, which runbook class? It does not roll back deploys or page people.
This unofficial page is the evaluation slice of the incident response classification pack. Intent: apply the Jev (TypeSafe System One) decision model to incident response classification evaluation. Primary search language: Incident response Jev evaluation. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Independent angle (cover ≠ clone): Runbooks and pagers stay; Jev classifies messy alert/customer text into SEV / runbook class, then code pages. Not a runbook-clone or status-page IA photocopy.
Incident response use-case context
Evaluation for incident response classification is a frozen harness, not a vibe check and not an opinion-blog “Jev review.” Labels: SEV gold from ICs, runbook-class gold, and whether a page was warranted. We publish no unofficial accuracy.
Hub: Use cases. Compare, when the other tool is the real job: incident runbooks.
Evaluation inputs
Replay the same contract you ship:
{
"report": { "id": "INC-44", "text": "Checkout 500s since 14:02 UTC after the payments deploy. Status page still green." },
"signals": { "error_rate_bucket": "high", "payments_deploy_recent": true },
"policy": { "sev1": "SEV1 = complete checkout loss or safety." }
}
Freeze questions, criteria, and jev-1.13.0. Record the response model.
Decision signals and actions
Score these, not a blog-grade star rating:
- Missed SEV1 language on planted checkout-down reports
- False SEV1 pages (pager fatigue)
- IC override rate
Pair auto-act errors with handoff rate. If the incident classifier never acts, you have not evaluated incident response classification — you have evaluated a human queue.
Do not treat a Noul of 0.5 as a “medium” incident response classification score — it means yes and no are equally likely. Conjunctions stay in your code.
Guardrails and escalation
Promote a threshold only when the harness says auto-act error ≤ SLA and reviewers still catch the residual. TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For incident response classification, treat page_sev1 as the high bar (paging SEV1 / rolling back via automation). Tune on labels — see offline evaluation.
Evaluation and rollout notes
After any criteria edit, rerun before production. Cookbook lifts you see on TypeSafe pages are vendor claims — re-measure on your alert/customer incident narratives. Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.
Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Pack map
| Slice | Page |
|---|---|
| Graph and primitives | decision workflow |
What may enter state |
input contracts |
| What to gather first | evidence collection |
| Atomic rules | policy checks |
| Act / review / abstain | confidence thresholds |
| Reviewer payload | human handoff |
| What to persist | audit trail |
| How it breaks | failure modes |
| Labeled replay | you are here |
| Shadow → canary | production rollout |
FAQ
Will jev.pro publish a leaderboard for this use case? No. Measure on your labels. Vendor cookbook figures stay labeled as vendor claims.
What must stay frozen? Questions, criteria, and the pinned model id. Aliases can move.
Where is the rest of the Incident response pack? Start with Incident response failure modes and Incident response production rollout. Cluster hub: Use cases.
If metrics already say critical, should we wait for Jev? No. Metrics page now. Jev is for leftover messy text. See safe defaults.
Can Jev write the status-page update? Not in this workflow. Classification only. Generation is a different job.
What this page does not claim
- Not a pager, status page, or IR retainer.
- No MTTR or uptime claims.
- Not official TypeSafe.
- Official TypeSafe status, or that jev.pro issues API keys.
- That a schema-constrained answer is automatically factually correct.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.
Primary documentation: https://docs.typesafe.ai. Hub: Use cases.
Sources
Public TypeSafe or adjacent documentation only. No private claims.