See inside every LlamaIndex query.
LlamaIndex turns your data into answers; Obsivara tells you what each answer costs and why some go wrong. Every query traced through retrieval, re-ranking, and synthesis — each embedding and LLM call carrying tokens, latency, and dollars — so a slow or expensive RAG pipeline points to the exact step, not a guess. Instrument via the Obsivara SDK or OpenTelemetry; your pipeline keeps running unchanged, out of the request path.
LlamaIndex, fully observable.
Query-to-answer tracing
Every query engine run traced through retrieval, node post-processing, and synthesis — each step's inputs, outputs, latency, and the embedding and LLM calls it made.
RAG cost attribution
Embedding, re-ranking, and the final synthesis LLM call priced and rolled up per query, per pipeline, and per index — so you see which RAG designs are expensive by construction.
Retrieval and latency visibility
How many nodes each query retrieved and where time goes — retrieval versus synthesis — so a slow pipeline points to the exact stage instead of a vague 'RAG is slow'.
Failure attribution and health
Failed queries, embedding and LLM errors, and timeouts classified and traced to the component that broke, with health scoring and predictive alerts on degrading pipelines.
Three steps. No code changes.
Instrument with the SDK or OpenTelemetry
Wrap your LlamaIndex app with the Obsivara Python SDK, or point an OpenTelemetry / OpenInference LlamaIndex instrumentation at Obsivara's OTLP endpoint. Your pipeline keeps calling models directly — Obsivara records each run out-of-band, with no proxy in the path and no added latency.
Run your queries
Traces flow automatically as query engines and retrievers execute — no changes to your indexes, retrievers, or query logic.
See it in Obsivara
Pipelines appear in the inventory and knowledge map with per-query cost, latency, health scores, and full traces within minutes.
Signals tracked out of the box.
- →Query, retrieval, and synthesis traces
- →Tokens and cost per query and pipeline
- →Embedding vs LLM call cost split
- →Retrieval latency and nodes per query
- →Error rates and failing components
- →Health score and predictive alerts
Common questions.
Does Obsivara work with LlamaIndex query engines and agents?
Yes. Query engines, retrievers, node post-processors, and LlamaIndex agents are traced as a span tree, so you can drill from a query down into each retrieval and LLM call.
Do I need to change my LlamaIndex code?
No. You instrument via the Obsivara SDK wrapper or an OpenTelemetry/OpenInference exporter; your indexes, retrievers, and query logic stay exactly as they are.
Can Obsivara track RAG cost per query?
Yes. Embedding, re-ranking, and synthesis calls are priced and rolled up per query, pipeline, and index, with waste detection on expensive retrieval patterns.