ML — retrieval-augmented LLM ablation study
Component ablation for a retrieval-augmented LLM: each row removes or swaps one part of the pipeline and reports the change in exact-match accuracy, ranked by size, with protocol, search budget and compute stated. Typed FlowScript for machine-learning evaluation. Keywords: ablation, benchmark, evaluation, model comparison.
Make it your own.
// Component ablation for a retrieval-augmented LLM: each row removes or swaps one part of the pipeline and reports the change in exact-match accuracy, ranked by size, with protocol, search budget and compute stated.
//
// Every string below is EXAMPLE text from a fictional evaluation: replace it.
//
// How to read it. The baseline is the full pipeline scored on a held-out
// split that was deduplicated against the retrieval corpus and the
// fine-tuning data. Each delta is the ablated pipeline's score minus the
// baseline's, as a fraction (−0.047 is 4.7 points of exact match), and
// is the mean of three seeds; the resulting score is computed as baseline
// plus delta, so it cannot drift from the delta beside it. Across seeds
// the baseline's standard deviation was 0.004, so a delta smaller than
// about 0.01 is within run-to-run noise and should be reported as such,
// not as a finding.
//
// Change ONE thing per ablation. A row that removes the reranker and also
// shortens the context measures neither.
model rag_full {
title: "Retrieval-augmented reader, 8B decoder"
params: 8000000000
family: "decoder-only transformer, instruction-tuned"
}
dataset heldout_qa {
title: "Internal QA benchmark, held-out split"
n: 2400
split: "test"
description: "2400 questions written by domain experts after the corpus snapshot; near-duplicates of corpus passages removed by MinHash (Jaccard above 0.8)."
}
eval full_pipeline {
model: rag_full
dataset: heldout_qa
metric: "exact-match accuracy"
score: 0.684
}
ablation closed_book {
base: full_pipeline
delta: -0.211
change: "No retrieval: answer from model weights alone (closed book)"
}
ablation bm25_only {
base: full_pipeline
delta: -0.047
change: "Sparse retrieval only: BM25 instead of hybrid BM25 + dense"
}
ablation no_reranker {
base: full_pipeline
delta: -0.032
change: "Remove the cross-encoder reranker; use first-stage order"
}
ablation top_k_5 {
base: full_pipeline
delta: -0.018
change: "Pass 5 passages to the reader instead of 20"
}
ablation no_query_rewrite {
base: full_pipeline
delta: -0.012
change: "Remove LLM query rewriting before retrieval"
}
ablation no_citation_prompt {
base: full_pipeline
delta: -0.006
change: "Drop the cite-your-passage instruction from the prompt"
}
ablation context_32k {
base: full_pipeline
delta: 0.004
change: "Extend the reader context window from 8k to 32k tokens"
}
ml_protocol protocol {
task: "Open-domain question answering over an internal document corpus; one short answer per question."
seeds: [11, 23, 47]
metric_definition: "Exact match after lower-casing and stripping punctuation and articles; the reported value is the mean over three seeds, and the interval for a paired difference is a 10,000-sample bootstrap over questions within each seed."
preprocessing: "Corpus chunked into 256-token passages with 32-token overlap, so 20 passages and the prompt fit the reader's 8k context; chunking and the dense index were built on the corpus only, never on test questions."
deduplication: "Test questions checked against the fine-tuning set and the corpus by MinHash; 37 near-duplicates removed before any run."
code: "Evaluation harness at tag eval-v2.3"
}
ml_search tuning {
space: "Reranker depth 20 to 100; passages to reader 5 to 25; temperature 0 to 0.3"
trials: 24
selection_metric: "Exact match"
selection_split: dev_qa
chosen: "Rerank top 50, 20 passages to the reader, temperature 0"
note: "Selection used the dev split only; the test split was first scored after the configuration was frozen."
}
dataset dev_qa {
title: "Internal QA benchmark, development split"
n: 600
split: "dev"
}
ml_compute compute {
hardware: "8 × 80 GB GPUs, one node"
accelerator_hours: 412
total_runs: 48
note: "24 tuning trials on the dev split, plus 7 ablations and the baseline at 3 seeds each on test."
}
view ablations: ablation_table(full_pipeline)