Jev evaluation guide
Evaluation means: did the product action match gold, at a cost you accept? TypeSafe’s workflow charts are a vendor story (including high-end speed multiples). Your harness is the review.
Unofficial. How-to depth: offline evaluation. Canonical APIs: docs.typesafe.ai.
Independent angle (cover, do not clone)
Win delta (review-hands-on): reproducible eval notes with dates; methodology > opinion blog.
Harness recap
Gold JSONL → pin jev-1.13.0 → freeze questions → score label and act/review/abstain → CI on diffs. Cookbook lifts stay on cookbook pages.
Dated methodology (2026-09-21)
Record door, language mix, whether state was filtered, and who labeled gold. Author-bias exists (TypeSafe even warns their evals were written by the capabilities team). Yours will too — write it down.
What “good” means
Good is policy: fewer expensive false autos, acceptable review load, no silent 422s. A higher argmax match rate with more irreversible mistakes is a regression.
Date every report. Include door + pin + question SHA. If you change two of those at once, you cannot attribute.
What this page does not claim
- No invented 193× on jev.pro dashboards.
FAQ
Online eval? After offline gates. See vs online experimentation.
Public leaderboards? public evals if we list any — still not Jev-as-a-score.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai. Never treat jev.pro as TypeSafe official documentation. We do not sell, issue, or proxy API keys.
Open-cluster pages are independent field-guide notes. Replicas and third-party interfaces mentioned anywhere on jev.pro are not Jev and not endorsed. Hub: Open. Siblings: testing guide, offline evaluation, vs offline evaluation. Canonical: https://docs.typesafe.ai.
Sources
Public TypeSafe or adjacent documentation only. No private claims.