Open· Last updated

Jev evaluation guide

Evaluation means: did the product action match gold, at a cost you accept? TypeSafe’s workflow charts are a vendor story (including high-end speed multiples). Your harness is the review.

Unofficial. How-to depth: offline evaluation. Canonical APIs: docs.typesafe.ai.

Independent angle (cover, do not clone)

Win delta (review-hands-on): reproducible eval notes with dates; methodology > opinion blog.

Harness recap

Gold JSONL → pin jev-1.13.0 → freeze questions → score label and act/review/abstain → CI on diffs. Cookbook lifts stay on cookbook pages.

Dated methodology (2026-09-21)

Record door, language mix, whether state was filtered, and who labeled gold. Author-bias exists (TypeSafe even warns their evals were written by the capabilities team). Yours will too — write it down.

What “good” means

Good is policy: fewer expensive false autos, acceptable review load, no silent 422s. A higher argmax match rate with more irreversible mistakes is a regression.

Date every report. Include door + pin + question SHA. If you change two of those at once, you cannot attribute.

What this page does not claim

FAQ

Online eval? After offline gates. See vs online experimentation.

Public leaderboards? public evals if we list any — still not Jev-as-a-score.

Disclaimer

This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai. Never treat jev.pro as TypeSafe official documentation. We do not sell, issue, or proxy API keys.

Open-cluster pages are independent field-guide notes. Replicas and third-party interfaces mentioned anywhere on jev.pro are not Jev and not endorsed. Hub: Open. Siblings: testing guide, offline evaluation, vs offline evaluation. Canonical: https://docs.typesafe.ai.

Sources

Public TypeSafe or adjacent documentation only. No private claims.