A new arXiv paper examines how fine-grained RAG benchmarks should be.
The authors propose a hierarchical framework for generating synthetic questions, aiming to test retrieval and answer generation at different levels of granularity rather than relying on one flat score.
That matters for production RAG systems because failures can hide inside broad metrics. A benchmark that separates where the system breaks can give teams better signals for fixing retrieval, grounding, and synthesis.