Skip to content
Mermaid templates

RAG Evaluation Harness

Splits retrieval and generation metrics against a frozen 420-question golden set so a regression is attributed to the right stage.

Template previewMermaid
Rendering…

Make it your own.

flowchart LR
  A[Golden set - 420 questions with human-cited answers] --> B[Freeze the corpus snapshot and the chunker version]
  B --> C[Retrieve top-k chunks]
  C --> D["Retrieval metrics: recall@10, nDCG@10, chunk overlap"]
  C --> E[Generate with the candidate model and the pinned prompt]
  E --> F[Answer metrics: faithfulness, answer relevance, citation precision]
  D --> G{"Recall@10 above 0.85?"}
  F --> H{Faithfulness above 0.90 and no uncited claims?}
  G -->|No| I[Tune chunk size, hybrid search weights, reranker depth]
  H -->|No| J[Tune the prompt and tighten the refuse-if-unsupported rule]
  I --> C
  J --> E
  G -->|Yes| K[Promote to shadow traffic]
  H -->|Yes| K
  K --> L[Compare against production on 7 days of live queries]
  L --> M{Win rate above 55 percent with no faithfulness regression?}
  M -->|Yes| N[Roll out behind a flag at 10 percent]
  M -->|No| I