Guides9 min read
By the Obsivara team

RAG Observability: Monitoring Retrieval Quality and Cost

How to monitor a RAG pipeline: retrieval quality, grounding, latency, and cost per query — and catch the silent retrieval failures behind hallucinations.

Quick answer

RAG observability monitors the whole retrieval-augmented pipeline, not just the final LLM call. Trace the query, embedding, vector search, re-rank, and generation as one run, and watch retrieval quality, grounding, latency, and cost per query. Most RAG failures are silent: the search returns nothing relevant, the model fills the gap, and the answer looks confident and wrong — and only tracing the retrieval step surfaces that before users do.

RAG observability is the practice of monitoring a retrieval-augmented generation pipeline end to end — the query, the embedding, the vector search, the re-rank, and the generation — rather than just the final model call. It matters because RAG moves the hardest failure out of the model and into retrieval: when the search returns nothing relevant, a well-behaved model will often fill the gap with something plausible and wrong. The answer looks confident, nothing errors, and you find out from a user. This guide covers what a RAG pipeline actually looks like to an observability tool, the signals that matter, and how to instrument and operate it.

What a RAG pipeline actually looks like

From the outside, RAG is one request in and one answer out. From the inside, it is a chain: the user query is embedded, the embedding is used to search a vector store, the top matches are optionally re-ranked and trimmed to fit the context window, those chunks are stitched into a prompt, and the model generates an answer from them. Each stage can fail independently, and a failure early in the chain is invisible by the time you read the final answer.

That is why RAG observability treats the pipeline as a trace of spans — one span for the retrieval, one for the re-rank, one for the generation — not as a single opaque call. With the chain visible, a wrong answer stops being a mystery: you can see whether retrieval returned garbage, whether the right chunk was retrieved but dropped during trimming, or whether retrieval was fine and the model simply ignored it.

The signals that matter

Useful RAG observability records more than latency. These are the signals that separate a healthy pipeline from one quietly drifting:

SignalWhy it matters
Retrieval hit quality (scores, top-k)Catch queries that return nothing relevant before the model improvises on them
Grounding / citation rateMeasure how often the answer is actually supported by the retrieved context
Chunks retrieved vs usedSpot the right chunk being dropped during trimming or re-ranking
Per-stage latencyFind whether the vector search or the generation is the slow part
Cost per queryAttribute embedding + generation spend, and catch context bloat that inflates it
Empty / fallback retrievalsTrend the silent-failure rate: queries with no good match

Silent retrieval failures and hallucinations

The defining RAG failure is the silent one. A retrieval miss — the vector search returns nothing relevant, or returns stale or off-topic chunks — does not throw; it just hands the model a weak context, and the model answers anyway. The result is a hallucination with a clean 200 next to it. Because the symptom shows up in the generated text while the cause sits two stages upstream, teams that only watch the model call chase the wrong thing for days. Tracing the retrieval step is what connects the confident-but-wrong answer back to the empty search that caused it — the clustering approach in debugging LLM hallucinations in production applies directly here.

The practical defenses are cheap once you can see the pipeline: alert when the retrieval score for a query falls below a threshold, have the system say 'I do not know' or escalate rather than answer on an empty retrieval, and track the empty-retrieval rate as a first-class metric. None of these require a better model — they require seeing the retrieval step at all.

What breaks RAG quality over time

A RAG pipeline that worked at launch rarely fails all at once; it drifts. The most common cause is the knowledge base itself changing: documents are added, edited, or removed, and a re-index shifts which chunks a query matches. A query that used to retrieve the right passage now retrieves a neighbor, and answer quality slips with no code change and no error to point at. Watching the empty-retrieval rate and the average retrieval score against each pipeline's baseline turns that invisible drift into a dated, visible event — 'grounding dropped the morning after Tuesday's re-index.'

Two other slow failures are worth tracking. Embedding-model or chunking changes alter the geometry of your vector space, so retrieval behavior can shift when you upgrade an embedding model without re-embedding everything consistently. And context creep: as you add sources, prompts quietly grow until the retrieved documents exceed the useful context window, and the model starts ignoring the middle of its input — a fact that is technically present but effectively invisible. Both show up in observability before they show up in user complaints: the first as a change in retrieval scores, the second as rising cost per query with flat or falling grounding.

Instrumenting and operating RAG

How you emit the spans depends on your stack. If you build on LlamaIndex, its instrumentation hooks already break a query into retrieval, synthesis, and LLM-call events, so the pipeline traces itself with little extra code. If you build on Haystack, instrument each pipeline component — retriever, ranker, prompt builder, generator — so every stage shows up as its own span. Either way the telemetry travels out of band over the SDK or OpenTelemetry (OTLP); nothing sits in the query path, so instrumentation never slows the answer your user is waiting on.

From there it becomes an operations problem: trend the empty-retrieval rate and grounding rate against each pipeline's baseline, attribute embedding and generation spend with cost intelligence so context bloat shows up as a cost line, and alert when retrieval quality drifts after a re-index or a data change. That is the layer Obsivara adds on top of your RAG stack. If you would rather have the whole assistant designed and monitored for you, that is the RAG knowledge-assistant build practice. You can start on the free tier and scale up as usage grows — see pricing — or run a free 48-hour audit to see your retrieval quality and cost per query on real traffic.

Frequently asked questions

What is RAG observability?

It is monitoring a retrieval-augmented generation pipeline end to end — the query, embedding, vector search, re-rank, and generation — as a trace of spans rather than a single model call. It tracks retrieval quality, grounding, per-stage latency, and cost per query so you can tell whether a wrong answer came from bad retrieval, dropped context, or the model ignoring good context.

Why does RAG hallucinate even with retrieval?

Usually because retrieval quietly failed. When the vector search returns nothing relevant, or returns stale or off-topic chunks, the model fills the gap with a plausible answer rather than refusing — and nothing errors, so the hallucination ships with a clean 200. The cause sits upstream of the text you see, which is why you need to trace the retrieval step, alert on low retrieval scores, and fall back to 'I do not know' on empty results.

How do I monitor the cost of a RAG pipeline?

Attribute both the embedding calls and the generation call per query, then roll the spend up per pipeline. The biggest hidden driver is context bloat — stuffing too many or too-large chunks into the prompt — which inflates generation cost on every query. Watching cost per query alongside chunks-retrieved-versus-used catches that, and a tool like Obsivara attributes it automatically from the pipeline's spans.

See your AI operations clearly.

Connect your stack in minutes — Obsivara discovers every workflow, scores its health, and attributes every dollar.

MORE POSTS