Skip to content
FlowScript templates

ML — retrieval-augmented LLM ablation study

Component ablation for a retrieval-augmented LLM: each row removes or swaps one part of the pipeline and reports the change in exact-match accuracy, ranked by size, with protocol, search budget and compute stated. Typed FlowScript for machine-learning evaluation. Keywords: ablation, benchmark, evaluation, model comparison.

Template previewFlowScript
ABLATION STUDY · EXACT-MATCH ACCURACYfull_pipeline — baseline 0.684#CHANGEΔ vs BASESCORE1No retrieval: answer from model weights alone (closed book)closed_book−0.2110.4732Sparse retrieval only: BM25 instead of hybrid BM25 + densebm25_only−0.0470.6373Remove the cross-encoder reranker; use first-stage orderno_reranker−0.0320.6524Pass 5 passages to the reader instead of 20top_k_5−0.0180.6665Remove LLM query rewriting before retrievalno_query_rewrite−0.0120.6726Drop the cite-your-passage instruction from the promptno_citation_prompt−0.0060.6787Extend the reader context window from 8k to 32k tokenscontext_32k+0.0040.688

Make it your own.

// Component ablation for a retrieval-augmented LLM: each row removes or swaps one part of the pipeline and reports the change in exact-match accuracy, ranked by size, with protocol, search budget and compute stated.
//
// Every string below is EXAMPLE text from a fictional evaluation: replace it.
//
// How to read it. The baseline is the full pipeline scored on a held-out
// split that was deduplicated against the retrieval corpus and the
// fine-tuning data. Each delta is the ablated pipeline's score minus the
// baseline's, as a fraction (−0.047 is 4.7 points of exact match), and
// is the mean of three seeds; the resulting score is computed as baseline
// plus delta, so it cannot drift from the delta beside it. Across seeds
// the baseline's standard deviation was 0.004, so a delta smaller than
// about 0.01 is within run-to-run noise and should be reported as such,
// not as a finding.
//
// Change ONE thing per ablation. A row that removes the reranker and also
// shortens the context measures neither.

model rag_full {
  title: "Retrieval-augmented reader, 8B decoder"
  params: 8000000000
  family: "decoder-only transformer, instruction-tuned"
}

dataset heldout_qa {
  title: "Internal QA benchmark, held-out split"
  n: 2400
  split: "test"
  description: "2400 questions written by domain experts after the corpus snapshot; near-duplicates of corpus passages removed by MinHash (Jaccard above 0.8)."
}

eval full_pipeline {
  model: rag_full
  dataset: heldout_qa
  metric: "exact-match accuracy"
  score: 0.684
}

ablation closed_book {
  base: full_pipeline
  delta: -0.211
  change: "No retrieval: answer from model weights alone (closed book)"
}
ablation bm25_only {
  base: full_pipeline
  delta: -0.047
  change: "Sparse retrieval only: BM25 instead of hybrid BM25 + dense"
}
ablation no_reranker {
  base: full_pipeline
  delta: -0.032
  change: "Remove the cross-encoder reranker; use first-stage order"
}
ablation top_k_5 {
  base: full_pipeline
  delta: -0.018
  change: "Pass 5 passages to the reader instead of 20"
}
ablation no_query_rewrite {
  base: full_pipeline
  delta: -0.012
  change: "Remove LLM query rewriting before retrieval"
}
ablation no_citation_prompt {
  base: full_pipeline
  delta: -0.006
  change: "Drop the cite-your-passage instruction from the prompt"
}
ablation context_32k {
  base: full_pipeline
  delta: 0.004
  change: "Extend the reader context window from 8k to 32k tokens"
}

ml_protocol protocol {
  task: "Open-domain question answering over an internal document corpus; one short answer per question."
  seeds: [11, 23, 47]
  metric_definition: "Exact match after lower-casing and stripping punctuation and articles; the reported value is the mean over three seeds, and the interval for a paired difference is a 10,000-sample bootstrap over questions within each seed."
  preprocessing: "Corpus chunked into 256-token passages with 32-token overlap, so 20 passages and the prompt fit the reader's 8k context; chunking and the dense index were built on the corpus only, never on test questions."
  deduplication: "Test questions checked against the fine-tuning set and the corpus by MinHash; 37 near-duplicates removed before any run."
  code: "Evaluation harness at tag eval-v2.3"
}

ml_search tuning {
  space: "Reranker depth 20 to 100; passages to reader 5 to 25; temperature 0 to 0.3"
  trials: 24
  selection_metric: "Exact match"
  selection_split: dev_qa
  chosen: "Rerank top 50, 20 passages to the reader, temperature 0"
  note: "Selection used the dev split only; the test split was first scored after the configuration was frozen."
}

dataset dev_qa {
  title: "Internal QA benchmark, development split"
  n: 600
  split: "dev"
}

ml_compute compute {
  hardware: "8 × 80 GB GPUs, one node"
  accelerator_hours: 412
  total_runs: 48
  note: "24 tuning trials on the dev split, plus 7 ablations and the baseline at 3 seeds each on test."
}

view ablations: ablation_table(full_pipeline)