Incident response decision workflow with Jev
Prometheus rules and PagerDuty already page on numeric thresholds. Jev is optional on messy customer-reported incidents or multi-alert narratives: which SEV band, which runbook class? It does not roll back deploys or page people.
This unofficial page is the decision workflow slice of the incident response classification pack. Intent: apply the Jev (TypeSafe System One) decision model to incident response classification decision workflow. Primary search language: Incident response Jev decision workflow. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Independent angle (cover ≠ clone): Runbooks and pagers stay; Jev classifies messy alert/customer text into SEV / runbook class, then code pages. Not a runbook-clone or status-page IA photocopy. Fan-out extra atoms on one request; open a second HTTP call only for a new artifact, not the same state.
Incident response use-case context
Incident response classification is a workflow, not a chat. Assemble a narrow state, ask the primitives below, and let the incident classifier branch. TypeSafe’s docs say a good question is a snap decision a knowledgeable person could make in a few seconds — not an open-ended analysis of the alert/customer incident narrative.
Hub: Use cases. Compare, when the other tool is the real job: incident runbooks.
Decision Workflow inputs
Keep only fields the questions name:
{
"report": { "id": "INC-44", "text": "Checkout 500s since 14:02 UTC after the payments deploy. Status page still green." },
"signals": { "error_rate_bucket": "high", "payments_deploy_recent": true },
"policy": { "sev1": "SEV1 = complete checkout loss or safety." }
}
Point instructions at report.text, policy.sev1, signals.error_rate_bucket. Drop raw Prometheus matrices and flame graphs; entire Slack incident channel history.
Decision signals and actions
| Id | Type | Job |
|---|---|---|
sev |
Choice | sev1 / sev2 / sev3 / other |
runbook |
Choice | payments_deploy / dependency / capacity / unknown / other |
customer_impact |
Noul | Does the narrative describe user-visible loss vs policy.sev1? |
All of these share state and run in parallel. Code owns the graph:
def incident_class(ans, error_rate_bucket):
if error_rate_bucket == "critical":
return page("sev1_payments") # metrics win
if ans["sev"].confidence < FLOOR or ans["sev"].choice == "other":
return page("ic_unknown") # fail toward a human IC
if ans["customer_impact"].noul >= T_IMPACT and ans["sev"].choice == "sev1":
return page("sev1_" + ans["runbook"].choice)
return "ticket_only"
Do not treat a Noul of 0.5 as a “medium” incident response classification score — it means yes and no are equally likely. Conjunctions stay in your code.
Guardrails and escalation
TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For incident response classification, treat page_sev1 as the high bar (paging SEV1 / rolling back via automation). Tune on labels — see offline evaluation.
Low confidence, sev == other, or a policy miss → human or safe default; do not page SEV1 from Jev alone; prefer a human IC when unsure.
Evaluation and rollout notes
Shadow: Pages follow today’s rules; log Jev SEV.
Canary: Suggest runbook class in the ticket for one service; paging stays rules-based.
Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.
Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Pack map
| Slice | Page |
|---|---|
| Graph and primitives | you are here |
What may enter state |
input contracts |
| What to gather first | evidence collection |
| Atomic rules | policy checks |
| Act / review / abstain | confidence thresholds |
| Reviewer payload | human handoff |
| What to persist | audit trail |
| How it breaks | failure modes |
| Labeled replay | evaluation |
| Shadow → canary | production rollout |
FAQ
Does Jev execute the incident classifier action? No. It returns typed answers. Your incident classifier code calls queues, models, or humans.
Why several questions in one request? TypeSafe’s fan-out pattern: extra questions are cheap versus another HTTP call. SEV + runbook + customer_impact in one call. Hierarchical product trees (if you have them) are a second Choice cascade with confidence abort — do not stuff 200 services into one Choice.
Where is the rest of the Incident response pack? Start with Incident response input contracts and Incident response confidence thresholds. Cluster hub: Use cases.
If metrics already say critical, should we wait for Jev? No. Metrics page now. Jev is for leftover messy text. See safe defaults.
Can Jev write the status-page update? Not in this workflow. Classification only. Generation is a different job.
What this page does not claim
- Not a pager, status page, or IR retainer.
- No MTTR or uptime claims.
- Not official TypeSafe.
- Official TypeSafe status, or that jev.pro issues API keys.
- That a schema-constrained answer is automatically factually correct.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.
Primary documentation: https://docs.typesafe.ai. Hub: Use cases.
Sources
Public TypeSafe or adjacent documentation only. No private claims.