When the corpus is too large to read in full, retrieve evidence for the question. Inspect source → parsing/chunking → retrieval → selection → context → answer. Finding a configuration does not establish its exceptions; retrieval scores and support for the answer need separate checks.
Introducing External Evidence into Generation
RAG (Retrieval-Augmented Generation) obtains external materials before or during generation, allowing the model to use these materials to answer. Classic research combines parametric models with retrievable non-parametric knowledge; engineering RAG can also use keyword search, SQL, document reading, or multi-turn tool queries, and is not limited to a single vector database. Retrieval-Augmented Generation
It decouples material updates from model weight updates, but it does not make the generation model unimportant. Materials may not exist at all, retrieval may miss critical conditions, assembly may truncate, and the model may misinterpret. These stages need to be verified separately; one cannot pre-assume that "the answer problem is definitely a retrieval problem."
For a small number of complete documents, providing the full text directly may be simpler, especially when cross-chapter holistic judgment is needed; for large-scale, frequently updated, or finely permissioned knowledge bases, obtaining materials on demand is easier to control. Fine-tuning can improve behavior, style, or domain handling, and can still be combined with retrieval. The choice depends on the task and evidence requirements, not a binary choice of "full text or fine-tuning won't work."
From Documents to Retrievable Evidence
Parsing Quality Determines What's in the Index
First, identify the body text, heading hierarchy, tables, code, and citation locations. PDF two-column order confusion, missing table headers, and OCR errors can all destroy meaning before vectorization. The retriever cannot recover non-existent original text from incorrectly parsed results.
Each evidence unit saves the document ID, version or content hash, chapter path, original text location, access scope, and update time. Original text location cannot rely solely on the chunk number, which changes with chunking adjustments; it needs to be able to return to the specific content in that document at that time.
Chunking is a Trade-off Between Integrity and Selection Precision
Short chunks are easy to match precisely but may lose references, definitions, and limitations; long chunks preserve relationships but may mix multiple topics and consume more window space. Chunking by chapter, paragraph, or code structure is a common starting point, but fixed length and overlap ratios should be validated against materials and tasks; one cannot assume all knowledge bases are suitable for 200–500 tokens.
A feasible design is to use smaller chunks for retrieval and parent chapters or adjacent paragraphs to provide answer evidence. For example, after hitting a configuration item, supplement its title, table header, and applicable limitations. Overlap can mitigate boundary breaks but causes duplicate hits; deduplicate before entering the context and check if different versions are incorrectly merged into one segment.
For snippets like "it increased by 10%" that are hard to understand without the original text, attach document and chapter context. Engineering cases of Contextual Retrieval use methods to supplement context for snippets to improve retrieval; the generated background also needs verification, and one cannot treat model-generated explanations as original text facts. Contextual Retrieval
Embeddings Provide Representation, Not Guaranteed Correct Understanding
Dense retrieval encodes queries and documents into vectors, sorting them based on dot product, cosine similarity, etc. The training objective makes relevant queries and evidence closer in the corresponding space, but this is not a guarantee of "identical meaning"; negations, exact numbers, and version differences may still be difficult to distinguish.
Queries and documents must use mutually compatible encoding schemes; they don't need to literally be the same encoder: dual-encoder designs can encode questions and paragraphs separately, provided the training and usage methods match. Dense Passage Retrieval model versions, vector dimensions, normalization, and query prompt formats should be recorded; switching encoding schemes usually requires regenerating the corresponding index; same dimension does not mean same space.
The choice between exact traversal and approximate nearest neighbor (ANN) indexing depends on data volume, dimensions, filtering, hardware, and latency goals; there is no unified threshold like "tens of thousands must be brute force, millions must use ANN." Approximate indexes also introduce candidate omissions, requiring separate checks for index recall and model representation capabilities.
Recall, Fusion, and Evidence Selection
Keywords and Vectors Can Complement Each Other
Error codes, product numbers, and function names are suitable for preserving literal matches; when users rephrase, semantic representation may fill in keyword omissions. Scores from methods like BM25 and vector retrieval are not on the same scale and cannot be added directly without calibration. Rank-based fusion can be used, or weighted methods calibrated on a validation set.
RRF accumulates 1/(c+r) for the rank r of candidate d in each result list; lists where it does not appear do not contribute. c is a smoothing constant, which is not the same parameter as how many results are returned at the end. RRF Original Paper
Assume keyword results are A, B, C, and vector results are B, D, A, with c=60. B's score is 1/62+1/61≈0.03252, A is 1/61+1/63≈0.03227, so B ranks before A. Fusion leverages signals that both paths support an item, but does not prove B's content is definitely correct.
Reranking Cannot Retrieve Materials Never Recalled
Cross-encoder reranking processes the query and candidate content jointly, usually costing more than a single vector similarity comparison, so it is often used for smaller candidate sets. It can bring correctly ranked but low-ranked recalled snippets into the final context, but it cannot add documents outside the candidate set out of thin air.
Set candidate count, post-reranking count, and context token budget separately. Expanding the candidate set may improve coverage but increases latency; reducing final evidence may reduce noise but may also lose the second fact needed for multi-hop questions. Adjust around evidence coverage, rather than fixed execution of "50 in, 5 out." Cross-task retrieval evaluations also show that the performance of different retrieval methods varies by domain and task. BEIR
For example, asking "Can Department A use interface X?" may require permission rules, A's role mapping, and X's version limitations simultaneously. High single-segment similarity is still not enough; the system needs to decompose the question or continue retrieving until key relationships have evidence support; explicitly state gaps when materials cannot be obtained.
How do two rankings combine?
Preserving both ranks makes every score reproducible. Smaller c emphasizes top-rank differences; larger c makes individual hit contributions closer. Fusion still requires checking evidence versions, permissions and support. RRF (Cormack et al., 2009)
RRF(d)=ℓ:d∈ℓ∑c+rℓ(d)1
Preparing the visual
How do two rankings combine?
RRF sums reciprocal rank terms, with zero contribution for missing items; it does not add differently scaled raw scores.
{"id":"ai-rrf-ranking","title":"How do two rankings combine?","summary":"RRF sums reciprocal rank terms, with zero contribution for missing items; it does not add differently scaled raw scores.","height":1000,"html":"<h2 data-i18n=\"heading\"></h2><p class=\"intro\" data-i18n=\"intro\"></p><div class=\"presets\" role=\"group\" data-i18n-label=\"mode\"><button data-mode=\"one\" data-i18n=\"one\"></button><button data-mode=\"two\" data-i18n=\"two\"></button><button data-mode=\"three\" data-i18n=\"three\"></button></div><div class=\"control\"><label for=\"c\" data-i18n=\"c\"></label><output for=\"c\" id=\"c-value\"></output><input id=\"c\" type=\"range\" min=\"0\" max=\"100\" value=\"60\" step=\"1\"></div><section class=\"panel\"><h3 data-i18n=\"scores\"></h3><div id=\"scores\"></div></section><section class=\"panel\"><h3 data-i18n=\"order\"></h3><div id=\"order\"></div></section><div class=\"copy\"><p id=\"explanation\" role=\"status\"></p><p class=\"reserve\" aria-hidden=\"true\" data-i18n=\"explain\"></p></div><details class=\"scope\"><summary data-i18n=\"scopeLabel\"></summary><p data-i18n=\"scope\"></p></details>","css":".intro{margin:8px 0 18px;color:var(--muted);font-size:14px}.control{display:grid;grid-template-columns:minmax(0,1fr) auto;gap:6px 12px;align-items:center;margin:14px 0;font-size:13px}.control output{color:var(--accent);text-align:right;font-variant-numeric:tabular-nums;min-width:4em}.control input{grid-column:1/-1;width:100%;margin:0}.presets{display:flex;gap:4px;align-items:stretch}.presets button{flex:1;min-width:0;overflow-wrap:anywhere;font-size:13px;min-height:44px}.actions{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:8px;margin:16px 0}.actions button{min-width:0;min-height:44px;font-size:13px;overflow-wrap:anywhere}.copy{display:grid;font-size:14px;margin:16px 0}.copy>*{grid-area:1/1;margin:0;overflow-wrap:anywhere}.reserve{visibility:hidden}.scope{margin-top:14px;color:var(--muted);font-size:12px}.scope summary{padding:8px 0;cursor:pointer}.scope p{margin-top:10px;overflow-wrap:anywhere}.metrics{display:grid;grid-template-columns:repeat(2,minmax(0,1fr));gap:14px;margin:18px 0}.metrics>div{border-left:2px solid var(--accent);padding:0 6px 0 10px;min-width:0}.metrics span{display:block;min-height:3.6em;color:var(--muted);font-size:12px;overflow-wrap:anywhere}.metrics strong{font-size:21px;display:block;min-height:2em;font-weight:550;overflow-wrap:anywhere;font-variant-numeric:tabular-nums}.check{display:flex;align-items:center;gap:9px;font-size:13px;margin:14px 0}.check input{width:18px;height:18px;flex-shrink:0;accent-color:var(--accent)}select,input[type=text],input[type=number]{background:var(--paper);color:var(--ink);font:inherit;font-size:14px;padding:10px;border:1px solid var(--rule);border-radius:7px;max-width:100%;min-width:0}select{width:100%;min-height:44px}.panel{padding:14px;border:1px solid var(--rule);border-radius:9px;margin:16px 0;min-width:0}.panel>header{font-size:13px;font-weight:600;margin-bottom:12px;color:var(--muted)}.panel>div{overflow-wrap:anywhere}.track-row{margin:16px 0}.track-row>span{display:block;font-size:12px;color:var(--muted);margin-bottom:8px}.track{display:flex;gap:5px;min-width:0}.cell{flex:1;min-width:0;min-height:56px;border:1px solid var(--rule);border-radius:5px;background:var(--surface);display:flex;align-items:center;justify-content:center;text-align:center;padding:6px 3px;font-size:12px;overflow-wrap:anywhere}.cell.on{background:var(--accent-soft);border-color:var(--accent)}.cell.gold{background:var(--second-soft);border-color:var(--second)}.cell.empty{border-style:dashed;color:var(--muted)}.cell.error{border-color:var(--second);text-decoration:line-through}.table{display:grid;gap:5px;font-size:12px}.tr{display:grid;gap:5px}.tr>div{background:var(--surface);border-radius:4px;padding:9px 6px;min-width:0;min-height:5em;overflow-wrap:anywhere;display:flex;align-items:center}.tr>.gold{background:var(--second-soft)}.tr>.on{background:var(--accent-soft)}.formula,.source{font:13px/1.8 ui-monospace,monospace;white-space:pre-wrap;overflow-wrap:anywhere;margin:16px 0}.formula{border-top:1px solid var(--rule);padding-top:12px;min-height:7.2em}.source{min-height:8em;background:var(--surface);padding:12px;border-radius:8px}.intervals{margin:16px 0}.interval{position:relative;height:38px;background:var(--surface);margin:6px 0;border-radius:4px;overflow:hidden}.interval i{position:absolute;height:100%;background:var(--accent-soft);border-left:2px solid var(--accent)}.interval i.gold{background:var(--second-soft);border-color:var(--second)}.interval span{position:relative;z-index:1;font:12px/38px ui-monospace,monospace;padding-left:7px}.bars{display:grid;gap:10px;margin:16px 0}.bar-row{display:grid;grid-template-columns:44px minmax(0,1fr) 62px;gap:8px;align-items:center;font:12px ui-monospace,monospace}.bar-track{height:24px;background:var(--surface);border-radius:4px;overflow:hidden}.bar-track i{display:block;height:100%;background:var(--accent);transition:width .2s}.bar-track i.gold{background:var(--second)}.bar-track i.empty{opacity:.2}.bar-row output{text-align:right}.diagram{display:block;width:100%;height:auto;margin:18px 0}.diagram text{font:13px ui-monospace,monospace;fill:var(--ink)}.node{fill:var(--surface);stroke:var(--rule)}.node.on{fill:var(--accent-soft);stroke:var(--accent)}.node.gold{fill:var(--second-soft);stroke:var(--second)}.edge{stroke:var(--rule);stroke-width:2;fill:none}.edge.on{stroke:var(--accent)}.matrix{display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:6px;margin:16px 0}.matrix button{min-width:0;min-height:44px;font:13px ui-monospace,monospace}.matrix button[aria-pressed=true]{background:var(--accent-soft);border-color:var(--accent)}.matrix button.masked{opacity:.35}.legend{font-size:12px;color:var(--muted);min-height:3.6em;overflow-wrap:anywhere}.numeric{font-variant-numeric:tabular-nums}.scope p{line-break:strict}@media(max-width:480px){.presets button,.actions button{font-size:12px}.panel{padding:12px}.metrics{gap:10px}.metrics strong{font-size:19px}.bar-row{grid-template-columns:34px minmax(0,1fr) 58px;gap:6px}}@media(prefers-reduced-motion:reduce){*,*::before,*::after{transition:none!important;animation:none!important}}\n\n.panel>h3{font-size:13px;font-weight:600;margin:0 0 12px;color:var(--muted)}\n\n.bar-row{grid-template-columns:44px minmax(0,1fr) 82px}.bar-row output{white-space:nowrap}@media(max-width:480px){.bar-row{grid-template-columns:34px minmax(0,1fr) 76px}}\n\n#order{min-height:2em;font:15px ui-monospace,monospace}.table .tr>div{min-height:3em;padding:8px 4px}","js":"const $=s=>document.querySelector(s);const set=(id,v)=>$('#'+id).textContent=v;const t=k=>viz.t(k);function cells(id,values){$('#'+id).replaceChildren(...values.map(v=>{const e=document.createElement('div');e.className='cell '+(v.cls||'');e.textContent=v.text;e.title=v.title||v.text;return e;}));}function pressed(mode){document.querySelectorAll('[data-mode]').forEach(e=>e.setAttribute('aria-pressed',String(e.dataset.mode===mode)));}const esc=s=>String(s).replace(/[&<>\"']/g,c=>({'&':'&','<':'<','>':'>','\"':'"',\"'\":'''}[c]));const fmt=s=>'{'+[...s].sort().join(', ')+'}';function graph(id,nodes,edges){const el=$('#'+id);el.setAttribute('viewBox','0 0 360 200');el.innerHTML='<defs><marker id=\"arrow-'+id+'\" viewBox=\"0 0 10 10\" refX=\"8\" refY=\"5\" markerWidth=\"5\" markerHeight=\"5\" orient=\"auto-start-reverse\"><path d=\"M 0 0 L 10 5 L 0 10 z\" fill=\"var(--muted)\"/></marker></defs>'+edges.map(e=>{const a=nodes[e[0]],b=nodes[e[1]],dx=b.x-a.x,dy=b.y-a.y,d=Math.hypot(dx,dy)||1;if(edges.some(r=>r[0]===e[1]&&r[1]===e[0])){const nx=-dy/d,ny=dx/d;return '<path class=\"edge '+(e[2]?'on':'')+'\" d=\"M '+(a.x+dx/d*22+nx*12)+' '+(a.y+dy/d*22+ny*12)+' Q '+((a.x+b.x)/2+nx*38)+' '+((a.y+b.y)/2+ny*38)+' '+(b.x-dx/d*25+nx*12)+' '+(b.y-dy/d*25+ny*12)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}return '<line class=\"edge '+(e[2]?'on':'')+'\" x1=\"'+(a.x+dx/d*25)+'\" y1=\"'+(a.y+dy/d*25)+'\" x2=\"'+(b.x-dx/d*28)+'\" y2=\"'+(b.y-dy/d*28)+'\" marker-end=\"url(#arrow-'+id+')\"/>';}).join('')+nodes.map(n=>'<g><circle class=\"node '+(n.cls||'')+'\" cx=\"'+n.x+'\" cy=\"'+n.y+'\" r=\"25\"/><text x=\"'+n.x+'\" y=\"'+(n.y+4)+'\" text-anchor=\"middle\">'+esc(n.name)+'</text>'+(n.sub?'<text x=\"'+n.x+'\" y=\"'+(n.y+43)+'\" text-anchor=\"middle\">'+esc(n.sub)+'</text>':'')+'</g>').join('');}function table(id,rows){$('#'+id).classList.add('table');$('#'+id).replaceChildren(...rows.map(r=>{const row=document.createElement('div');row.className='tr';row.style.gridTemplateColumns='repeat('+r.length+',minmax(0,1fr))';r.forEach(x=>{const c=document.createElement('div');if(typeof x==='object'){c.textContent=x.text;c.className=x.cls||'';}else c.textContent=x;row.append(c)});return row;}));}function band(id,parts,total){$('#'+id).replaceChildren(...parts.map((p,i)=>{const e=document.createElement('i');e.style.width=(100*p.value/total)+'%';e.className=i%2?'gold':'on';e.title=p.name+': '+p.value;return e}));}function intervals(id,items,total){$('#'+id).replaceChildren(...items.map(p=>{const row=document.createElement('div');row.className='interval';const bar=document.createElement('i');bar.style.left=(100*p.start/total)+'%';bar.style.width=(100*(p.end-p.start)/total)+'%';bar.className=p.cls||'';const label=document.createElement('span');label.textContent=p.label;row.append(bar,label);return row;}));}function plot(id,series,xr,yr,opts={}){const e=$('#'+id),X=x=>42+(x-xr[0])/(xr[1]-xr[0])*298,Y=y=>175-(y-yr[0])/(yr[1]-yr[0])*148;let out='<defs><clipPath id=\"clip-'+id+'\"><rect x=\"42\" y=\"27\" width=\"298\" height=\"148\"/></clipPath></defs>';for(let i=0;i<3;i++){let x=xr[0]+i*(xr[1]-xr[0])/2,y=yr[0]+i*(yr[1]-yr[0])/2;out+='<path class=\"gridline\" d=\"M '+X(x)+' 27V175M42 '+Y(y)+'H340\"/><text x=\"'+X(x)+'\" y=\"194\" text-anchor=\"middle\">'+esc(opts.xfmt?opts.xfmt(x):Number(x.toFixed(2)))+'</text><text x=\"36\" y=\"'+(Y(y)+4)+'\" text-anchor=\"end\">'+esc(opts.yfmt?opts.yfmt(y):Number(y.toFixed(2)))+'</text>';}out+='<g clip-path=\"url(#clip-'+id+')\">';for(const s of series){out+='<path fill=\"none\" stroke=\"'+(s.color||'var(--accent)')+'\" stroke-width=\"2.5\" '+(s.dash?'stroke-dasharray=\"5 4\"':'')+' d=\"'+s.pts.map((p,i)=>(i?'L':'M')+X(p[0]).toFixed(2)+','+Y(p[1]).toFixed(2)).join(' ')+'\"/>';}for(const p of opts.points||[])out+='<circle cx=\"'+X(p[0])+'\" cy=\"'+Y(p[1])+'\" r=\"4\" fill=\"var(--second)\" stroke=\"var(--paper)\" stroke-width=\"1.5\"/>';if(opts.cursor!==undefined)out+='<path d=\"M'+X(opts.cursor)+' 27V175\" stroke=\"var(--muted)\" stroke-dasharray=\"3 3\"/>';e.innerHTML=out+'</g>';e.dataset.curves=JSON.stringify(series.map(s=>s.pts));}const curve=(fn,lo,hi,n=120)=>Array.from({length:n+1},(_,i)=>{const x=lo+(hi-lo)*i/n;return[x,fn(x)]});function inputs(draw){document.querySelectorAll('input,select').forEach(e=>{e.addEventListener('input',draw);e.addEventListener('change',draw)});draw()}const n=id=>+$('#'+id).value;const show=(id,v,d=3)=>set(id,Number(v.toFixed(d)));function bars(id,values,max=1){$('#'+id).className='bars';$('#'+id).innerHTML=values.map(v=>'<div class=\"bar-row\"><span>'+esc(v.label)+'</span><div class=\"bar-track\"><i class=\"'+(v.cls||'')+'\" style=\"width:'+Math.max(0,Math.min(100,v.value/max*100))+'%\"></i></div><output>'+esc(v.text??(v.value*100).toFixed(1)+'%')+'</output></div>').join('')}function softmax(z,T=1){let max=Math.max(...z),w=z.map(x=>Math.exp((x-max)/T)),sum=w.reduce((a,b)=>a+b,0);return w.map(x=>x/sum)}let mode='one';function draw(){pressed(mode);let a=['A','B','C'],b={one:['B','D','A'],two:['D','C','B'],three:['A','B','C']}[mode],c=n('c'),s={};let rows=['A','B','C','D'].map(x=>{let i=a.indexOf(x)+1,j=b.indexOf(x)+1;s[x]=(i?1/(c+i):0)+(j?1/(c+j):0);return [x,i||'—',j||'—',s[x].toFixed(5)]});let order=Object.keys(s).filter(x=>s[x]>0).sort((a,b)=>s[b]-s[a]||a.localeCompare(b));set('c-value',c);table('scores',[['d','KW','Vec','RRF'],...rows]);set('order',order.join(' → '));set('explanation',t('explain'));document.body.dataset.scores=JSON.stringify(s);document.body.dataset.order=order.join('');}document.querySelectorAll('[data-mode]').forEach(e=>e.onclick=()=>{mode=e.dataset.mode;draw()});inputs(draw);","audio":false,"strings":{"c":"Smoothing constant c","explain":"B gets ranks 2 and 1, slightly beating A’s ranks 1 and 3. A higher score does not verify the content as true.","heading":"How do two rankings combine?","intro":"Change the vector ranking and RRF constant; inspect each score.","mode":"Vector ranking","one":"B / D / A","order":"Fused order","reset":"Reset","scope":"Keyword ranking A/B/C, three vector lists, ranks start at 1 and ties use letters; ranking fusion only, not a relevance model.","scopeLabel":"Model scope and assumptions","scores":"Ranks and scores","three":"A / B / C","two":"D / C / B"}}
The Index is a Continuously Maintained Data Product
Building an index is not a one-time task. Additions, modifications, deletions, and permission revocations all need to enter the update chain. If old snippets remain in the vector or keyword database, the model will cite deprecated rules; if two indexes are not updated synchronously, hybrid retrieval may return conflicting versions.
Change
Content to Handle
Result of Ignoring
Body Text Modification
Re-parse affected content, update index and version
Cite old rules
Document Deletion
Revoke snippets, derived indexes, and caches
Deleted content reappears
Permission Revocation
Update retrieval scope and verify on read
Unauthorized materials enter model context
Embedding Replacement
Create corresponding vectors, verify before switching
Mixed search in old and new spaces
Chunking Strategy Change
Preserve document version and original text location mapping
Old citations cannot be traced
Permission constraints should participate in retrieval and result reading; one cannot hand all tenant materials to the model first and expect it to ignore them on its own. Retrieval caches also need to bind authorization scope, document version, and query conditions; user A's hit results cannot be directly reused for user B who has no access rights.
When upgrading indexes, first build the new version, verify coverage, latency, permissions, and citation location, then switch the read entry. Retain necessary rollback information, but deletions and permission revocations cannot be resurrected due to rollback. Entries stored in the vector database are derived data; the authoritative source and update records must still be retained.
Assembling Context and Verifying Citations
Distinguish evidence from instructions, annotating document identity, location, version, and truncation status. Dynamic evidence is usually better placed after stable instructions for prefix reuse; specific location effects still need measurement, and one cannot guarantee the model uses them correctly just by "putting them at the end." Operation instructions appearing in external documents do not automatically become application authorizations.
Citations can help verify, but "answer with links" does not mean the conclusion is supported. Check item by item: do citations exist, are versions applicable, and do snippets really support the assertion? For example, if a citation only says "test environment supports," but the answer writes "production environment supports," the location is accurate but the inference overreaches.
When there is insufficient evidence, continue reading parent chapters, expand queries, or explicitly state missing information. Do not force the model to give a definite answer regardless of whether materials are sufficient. RAG can also retrieve long-term memory; there is no natural division that "RAG is only objective, memory is only subjective."
Did retrieval include the condition?
Remove B from retrieval, then from context, and try reranking. Loss occurs at different stages; reranking cannot invent a condition absent from candidates. The measure is the required evidence combination, not one relevant hit.
Preparing the visual
Did retrieval include the condition?
Reranking only handles retrieved candidates; the answer needs both retry configuration and the idempotency condition.
Annotate relevant documents or evidence units for queries, and specify which combinations are needed for multi-hop questions. For query Q, let the relevant set be R, and the set of the top k returned be S:
Metric
Definition
Blind Spot
precision@k
Number of relevant items in top k / k
Does not indicate how much evidence is still missing
recall@k
Number of hit relevant items / Total annotated relevant items
Affected by annotation completeness
reciprocal rank
Reciprocal of the rank of the first relevant result, 0 if none
Does not measure subsequent necessary evidence
MRR
Average of multiple queries' reciprocal rank
Still biased towards first hit
Evidence Combination Coverage
Whether all key conditions needed for the answer are obtained
Requires task-level annotation
For example, if R={A,C} and the return is [B,A,D], precision@3=1/3, recall@3=1/2, reciprocal rank=1/2. Simplifying recall to "is there one correct snippet" misses C; it only degenerates into hit-or-miss when exactly one relevant item is annotated.
Metric calculation can be programmatic, but relevance labels and evidence requirements may still be incomplete or divergent. Also measure rejection reasonableness, citation support, final accuracy, latency, and cost. Retrieval relevance does not equal answer derivability, and snippets entering the candidate set do not equal entering the final generation window.
Locate Layer by Layer Along the Failure Chain
Failure Location
Check Method
Adjustment Direction
Original database has no answer
Check authoritative sources
Supplement materials or explicitly state unknown
Parsing lost conditions
Compare original text with stored snippets
Fix parsing, tables, and chunking
Candidate set not recalled
Check filtering, encoding, and multi-path retrieval
Fix recall, not just adjust reranking
Candidates present but final evidence missing
Check reranking, deduplication, and budget trimming
Preserve necessary conditions and multi-hop relationships