Archive
The Daily Org

A Salesforce NewspaperCurated by Abhinav

Tracing retrieval-augmented generation accuracy failures backward through the context pipeline

The engineering team discovered that standard benchmarks masked severe accuracy drops when processing complex enterprise documents. They resolved this by treating end-to-end accuracy as an opaque metric and instead debugging the pipeline backward. By isolating each stage from source ingestion to answer generation, they identified where meaning was lost during parsing or chunking rather than blaming the large language model.

The solution required routing simple text through fast deterministic parsers while using model-based processing for pages containing tables, charts, or diagrams. Chunking strategies shifted from fixed token limits to semantic boundaries that preserve headers, diagrams, and procedural sequences. Enriching chunks with metadata and question representations, combined with the extended context capacity of SFR Embedding v3, significantly improved retrieval fidelity.

Retrieval accuracy increased by applying dynamic metadata pre-filters before semantic ranking and preparing for GraphRAG to handle multi-hop questions across connected facts. The team also replaced batch indexing with a Just-in-Time architecture to reduce latency for smaller payloads. Continuous staged evaluations allowed them to attribute accuracy gains to specific changes, ultimately raising enterprise document accuracy above ninety percent.