Jev success metrics
Success for a Jev workflow is did the product action match gold at a cost you accept? It is not a star rating, not a jev.pro leaderboard, and not TypeSafe’s vendor speed charts copied into your QBR. Independent. docs.typesafe.ai.
Independent angle (cover, do not clone)
Win delta: policy-action metrics (auto precision, review load, 422s, pin drift) with vendor-claim labels. We do not clone a benchmark explainer or invent 193×.
Question and context
TypeSafe publishes workflow evals and jaggedness notes. Those are vendor stories (including high-end speed multiples). Your success metrics live on your traces and gold file. Evaluation how-to: evaluation guide and offline evaluation.
Core concepts
Score these
| Metric | Definition | Why |
|---|---|---|
| Auto-act precision | Among auto actions, share that match gold | Irreversible mistakes |
| Missed-irreversible rate | Gold said hold/escalate; you auto-acted | Safety |
| Review rate | Share sent to humans | Floor theater vs capacity |
| Override rate | Reviewer ≠ Jev argmax | Criteria or floor wrong |
| 422 / 429 / 529 rate | Documented HTTP | Contract and door health |
| Pin drift | response.model ≠ expected |
Silent alias |
| Input tokens p95 | usage.input_tokens |
Cost + distractors |
Do not score as “Jev success”
- Copied vendor 40×–200× without a footnote.
- Argmax match that ignores a bad action.
- Revenue lift you cannot attribute (confounded by the CRM change you shipped the same week).
List price $0.042 / M input (models page) is a vendor claim for budgeting, not a quality metric.
Practical approach
Date every report: door, pin, question SHA, language mix, who labeled gold. TypeSafe even warns some of their evals were written by the capabilities team — your labels are biased too. Write it down.
Pair auto-act errors with review rate. If the worker never acts, you measured a human queue, not Jev.
Risks and trade-offs
A dashboard that only shows “Noul mean” will be gamed. A dashboard that only shows savings will hide false-allows. Success metrics are a product policy (operating model).
What this page does not claim
- No unofficial accuracy for any vertical.
- No invented search rankings.
- Schema-safe ≠ correct.
FAQ
Will jev.pro publish a public leaderboard? Public evals only if we list a harness — still not “Jev-as-a-score.”
Online metrics first? After offline gates. See vs online experimentation.
Who owns the numbers? The eval hat on team roles.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai. Never treat jev.pro as TypeSafe official documentation. We do not sell, issue, or proxy API keys.
Open-cluster pages are independent field-guide notes. Replicas and third-party interfaces mentioned anywhere on jev.pro are not Jev and not endorsed. Hub: Open. Siblings: evaluation guide, operating model, team roles. Canonical: https://docs.typesafe.ai.
Sources
Public TypeSafe or adjacent documentation only. No private claims.