Questions for analysts
Analysts should refuse dashboards that plot a TypeSafe blog multiple as our KPI. Ask for action-level metrics on pinned models.
Unofficial. docs.typesafe.ai.
Ask
- What is the unit — label or action?
- Which model id?
- How was gold labeled?
- Review rate vs auto precision tradeoff?
- Token p95 and 429s?
- Split by language?
- Did we reuse a Noul threshold on a Choice?
- Are Score expectations abused as magnitudes?
Plot these first
Auto precision, review rate, 422 rate, token p95, response.model mix. Then, if you must, calibration curves per question id. Never a single site-wide “AI accuracy.”
What this page does not claim
- No canned ECE.
FAQ
Can I average Nouls across questions? Only as a hack you validate. Official: questions are different.
Vendor cookbook lifts? Cite them as vendor, on their pages.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai. Never treat jev.pro as TypeSafe official documentation. We do not sell, issue, or proxy API keys.
Open-cluster pages are independent field-guide notes. Replicas and third-party interfaces mentioned anywhere on jev.pro are not Jev and not endorsed. Hub: Open. Siblings: evaluation guide, for data teams, glossary calibration. Canonical: https://docs.typesafe.ai.
Sources
Public TypeSafe or adjacent documentation only. No private claims.