RAG Evaluation Harness
Splits retrieval and generation metrics against a frozen 420-question golden set so a regression is attributed to the right stage.
Rendering…
Make it your own.
flowchart LR
A[Golden set - 420 questions with human-cited answers] --> B[Freeze the corpus snapshot and the chunker version]
B --> C[Retrieve top-k chunks]
C --> D["Retrieval metrics: recall@10, nDCG@10, chunk overlap"]
C --> E[Generate with the candidate model and the pinned prompt]
E --> F[Answer metrics: faithfulness, answer relevance, citation precision]
D --> G{"Recall@10 above 0.85?"}
F --> H{Faithfulness above 0.90 and no uncited claims?}
G -->|No| I[Tune chunk size, hybrid search weights, reranker depth]
H -->|No| J[Tune the prompt and tighten the refuse-if-unsupported rule]
I --> C
J --> E
G -->|Yes| K[Promote to shadow traffic]
H -->|Yes| K
K --> L[Compare against production on 7 days of live queries]
L --> M{Win rate above 55 percent with no faithfulness regression?}
M -->|Yes| N[Roll out behind a flag at 10 percent]
M -->|No| I