Code review confidence thresholds with Jev
Linters, typecheckers, and SAST prove or pattern-match. Jev can decide whether leftover PR language looks like “needs security eyes” or “reviewer load is high.” It will not compile the repo or write the fix.
This unofficial page is the confidence thresholds slice of the code review routing pack. Intent: apply the Jev (TypeSafe System One) decision model to code review routing confidence thresholds. Primary search language: Code review Jev confidence thresholds. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Independent angle (cover ≠ clone): Analyzers own AST truth; Jev routes review judgment on PR prose + small hunks. Not a SAST clone and not a rival “agent skill” IA photocopy. Treat floors as production gates and price the false-reject cost — official 0.5/0.9 sketches are illustrations.
Code review use-case context
Thresholds turn code review routing answers into act / review / abstain. They are product policy, not a hyperparameter TypeSafe ships. Official 0.5 / 0.9 sketches are illustrations. This slice also carries the false-reject discussion: over-gating code review routing hides calibration.
Hub: Use cases. Compare, when the other tool is the real job: code analyzers.
Confidence Thresholds inputs
You need (1) pinned answers on a frozen contract and (2) labels for security_review / request_changes / merge_ok gold from staff engineers. State shape:
{
"pr": { "id": "1234", "title": "Relax auth on internal debug route", "body": "Temp bypass for the oncall drill." },
"hunks": ["- if (!user) return 401;", "+ // skip auth in staging"],
"ci": { "sast_blockers": 0, "lint_errors": 0 },
"policy": { "secrets": "Flag prose that describes committing keys or disabling auth in prod-shaped paths." }
}
Decision signals and actions
| Axis | Where it lives | Code review use |
|---|---|---|
choice / score / noul |
answer payload | What to do with the PR description + selected hunks |
confidence |
Choice & Score only | Whether to trust the argmax |
| Distance from 0.5 | Noul | Whether needs_security is decided |
FLOORS = {
"comment_only": 0.50, # illustrations — replace
"auto_approve_merge": 0.92,
}
NOUL_TAU = 0.70 # for needs_security
def allow(ans, action):
return ans.confidence >= FLOORS[action]
Do not treat a Noul of 0.5 as a “medium” code review routing score — it means yes and no are equally likely. Conjunctions stay in your code.
Guardrails and escalation
TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For code review routing, treat auto_approve_merge as the high bar (approving a merge or skipping required review). Tune on labels — see offline evaluation.
Band around 0.5 on needs_security always reviews. Do not copy 0.70 onto Choice confidence.
Evaluation and rollout notes
- Missed auth-bypass language on a planted set
- False security pages (reviewer fatigue)
- Override rate after humans read the same hunks
Fit loop: pin jev-1.13.0 → replay → plot error vs confidence → pick floors where auto-act error ≤ your SLA. Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.
Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.
Pack map
| Slice | Page |
|---|---|
| Graph and primitives | decision workflow |
What may enter state |
input contracts |
| What to gather first | evidence collection |
| Atomic rules | policy checks |
| Act / review / abstain | you are here |
| Reviewer payload | human handoff |
| What to persist | audit trail |
| How it breaks | failure modes |
| Labeled replay | evaluation |
| Shadow → canary | production rollout |
FAQ
Should auto_approve_merge use 0.9 everywhere? No. Over-gating hides calibration and dumps the queue on humans. Fit per action.
Can I reuse a Noul τ as Choice confidence? No. Jaggedness: they are not interchangeable. See confidence.
Where is the rest of the Code review pack? Start with Code review decision workflow and Code review human handoff. Cluster hub: Use cases.
Does the TypeSafe agent skill replace our linter? No. It teaches agents to batch questions. Analyzers still own AST truth. See agent skill.
Can Jev score cyclomatic complexity? Do not use a 2–10 Score as complexity math. Count in code or skip.
What this page does not claim
- Not a SAST, linter, or merge bot product.
- No published precision on vulnerability finding.
- Not official TypeSafe.
- Official TypeSafe status, or that jev.pro issues API keys.
- That a schema-constrained answer is automatically factually correct.
Disclaimer
This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.
Primary documentation: https://docs.typesafe.ai. Hub: Use cases.
Sources
Public TypeSafe or adjacent documentation only. No private claims.