Open· Last updated

Questions for analysts

Analysts should refuse dashboards that plot a TypeSafe blog multiple as our KPI. Ask for action-level metrics on pinned models.

Unofficial. docs.typesafe.ai.

Ask

  1. What is the unit — label or action?
  2. Which model id?
  3. How was gold labeled?
  4. Review rate vs auto precision tradeoff?
  5. Token p95 and 429s?
  6. Split by language?
  7. Did we reuse a Noul threshold on a Choice?
  8. Are Score expectations abused as magnitudes?

Plot these first

Auto precision, review rate, 422 rate, token p95, response.model mix. Then, if you must, calibration curves per question id. Never a single site-wide “AI accuracy.”

What this page does not claim

FAQ

Can I average Nouls across questions? Only as a hack you validate. Official: questions are different.

Vendor cookbook lifts? Cite them as vendor, on their pages.

Disclaimer

This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai. Never treat jev.pro as TypeSafe official documentation. We do not sell, issue, or proxy API keys.

Open-cluster pages are independent field-guide notes. Replicas and third-party interfaces mentioned anywhere on jev.pro are not Jev and not endorsed. Hub: Open. Siblings: evaluation guide, for data teams, glossary calibration. Canonical: https://docs.typesafe.ai.

Sources

Public TypeSafe or adjacent documentation only. No private claims.