Use cases· Last updated

RAG evidence collection with Jev

Retrievers hope. After retrieval, Jev marks relevance, contradiction, or injection; code keeps, flags, or drops passages before a generator sees them.

This unofficial page is the evidence collection slice of the RAG passage decisions pack. Intent: apply the Jev (TypeSafe System One) decision model to RAG passage decisions evidence collection. Primary search language: RAG Jev evidence collection. Confirm patterns on docs.typesafe.ai. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.

Independent angle (cover ≠ clone): Filter/rerank recipes with failure modes; cite vs generate boundary. We cover the intent, not a rival rerank-passages-score URL tree.

RAG use-case context

Evidence collection for RAG passage decisions happens before POST /v1/systemone. Jev does not browse your warehouse, retriever, or ESP. You gather the query + retrieved passage facts, filter them, then ask snap questions. This slice is where fan-out cost math belongs: batch questions, do not re-send state.

Hub: Classifying RAG passages. Compare, when the other tool is the real job: RAG pipelines.

Evidence Collection inputs

Collect:

Never send:

Shape the payload like this once the gather step finishes:

{
  "query": "What is the refund window for pro plans?",
  "passage": { "id": "doc-88#p3", "text": "Pro subscribers may request a refund within 14 days." },
  "corpus": { "trust": "internal_kb" }
}

Decision signals and actions

Each evidence field should change a named answer:

Id Type Job
relevant Noul Does passage.text answer query?
contradiction Noul Does it contradict other kept passages you include?
injection Noul Hidden instructions / prompt injection in the passage?
support Score How completely does it support an extractive answer?

TypeSafe’s parallel-questions cookbook: batch questions on one state. For many candidates, loop pairs (query, passage) or shortlist with BM25 first — official rerank cookbook — instead of one giant Choice over 200 ids unless you followed their line-search pattern.

Do not treat a Noul of 0.5 as a “medium” RAG passage decisions score — it means yes and no are equally likely. Conjunctions stay in your code.

Guardrails and escalation

If the gather step fails (empty query + retrieved passage, redaction stripped everything, retriever empty), fail closed on showing a passage to a customer-facing answerer. Do not invent evidence so Jev has something to say. TypeSafe’s confidence-gated examples use a lower bar for recoverable reads than for irreversible actions. Those numbers are illustrations. For RAG passage decisions, treat feed_to_answerer as the high bar (showing a passage to a customer-facing answerer). Tune on labels — see offline evaluation.

Evaluation and rollout notes

Your eval set should include thin-evidence cases, not only happy query + retrieved passages. Label passage relevant / not, plus injection gold on a hostile slice. Pin jev-1.13.0 (the versioned id) after you fit thresholds. jev-latest and the marketing line jev-1.13 can move. Log the response model. TypeSafe’s published list price for jev-1.13 is $0.042 per million input tokens (vendor claim — confirm on the models page); output tokens are free on that same page. Unused distractors still bill as input.

Official Python and JavaScript SDKs read TYPESAFE_API_KEY and retry documented 429/529. This site does not sell, issue, or proxy TypeSafe keys. Use a credential you already have from the console or a documented gateway.

Pack map

Slice Page
Graph and primitives decision workflow
What may enter state input contracts
What to gather first you are here
Atomic rules policy checks
Act / review / abstain confidence thresholds
Reviewer payload human handoff
What to persist audit trail
How it breaks failure modes
Labeled replay evaluation
Shadow → canary production rollout

FAQ

Should evidence live in the question text? Put facts in state and point instructions at query, passage.text. Criteria stay stable so you can replay.

When do I split calls? TypeSafe’s parallel-questions cookbook: batch questions on one state. For many candidates, loop pairs (query, passage) or shortlist with BM25 first — official rerank cookbook — instead of one giant Choice over 200 ids unless you followed their line-search pattern.

Where is the rest of the RAG pack? Start with RAG input contracts and RAG decision workflow. Cluster hub: Use cases.

Should Jev generate the RAG answer? No. Classify or score passages; another model (or extractive code) writes. That is the cite-vs-generate boundary.

Do we publish rerank lifts? No. TypeSafe’s cookbooks may show measurements — treat those as vendor figures and re-run on your corpus.

What this page does not claim

Disclaimer

This is an independent unofficial site and is not affiliated with TypeSafe AI; official documentation is available at https://docs.typesafe.ai.

Primary documentation: https://docs.typesafe.ai. Hub: Use cases.

Sources

Public TypeSafe or adjacent documentation only. No private claims.