Métriques d'un RAG : Top K, recall@K, precision@K, coût par requête
Four numbers that tell you whether a RAG's retrieval is doing its job, before you even read the model's answer. **Top K**: how many passages the search returns. With K = 5 the system hands back the 5 highest-scored passages, no more. **Recall@K**: the share of useful passages that made it into those K. If 4 useful passages exist and only 2 sit in your Top K, recall is 50%. It measures what you missed. **Precision@K**: the share of the K returned passages that are actually useful. Return 5, only 2 help, precision is 40%. It measures the noise you pay for. **Cost per request**: what one single answer costs, search, any reranking and the model calls included. The first two pull in opposite directions: raising K brings back more useful passages (recall goes up) but drowns them in more noise (precision goes down), and every extra passage inflates the bill. Top P is not on this list: it is a generation setting, covered in the LLM generation parameters entry. Source: [Instagram reel DeMqr3Ijx22](https://www.instagram.com/reel/DeMqr3Ijx22/).
Strengths
- Four numeric definitions that replace "retrieval looks good" with a measurement
- Recall and precision can be computed by hand on a small set of questions whose useful passages you know
Limitations
- Recall@K and precision@K assume you know in advance which passages are useful: you have to annotate a question set
- None of the four tells you whether the model's final answer is right: they only judge retrieval and the bill
Best for
- Anyone diagnosing an off-target RAG who wants to know if the fault is in retrieval
- PMs and POs who must set a quality threshold and a per-answer budget for an AI feature