INTEGRATION / FRAMEWORKS

See inside every LlamaIndex query.

LlamaIndex turns your data into answers; Obsivara tells you what each answer costs and why some go wrong. Every query traced through retrieval, re-ranking, and synthesis — each embedding and LLM call carrying tokens, latency, and dollars — so a slow or expensive RAG pipeline points to the exact step, not a guess. Instrument via the Obsivara SDK or OpenTelemetry; your pipeline keeps running unchanged, out of the request path.

LlamaIndexEXAMPLE
AGENT RUNS TRACED / 24H3,184
STEPS PER RUN (AVG)11.4
LOOPS FLAGGED / 24H6
ILLUSTRATIVE · EXAMPLE DATA
WHAT YOU GET

LlamaIndex, fully observable.

Query-to-answer tracing

Every query engine run traced through retrieval, node post-processing, and synthesis — each step's inputs, outputs, latency, and the embedding and LLM calls it made.

RAG cost attribution

Embedding, re-ranking, and the final synthesis LLM call priced and rolled up per query, per pipeline, and per index — so you see which RAG designs are expensive by construction.

Retrieval and latency visibility

How many nodes each query retrieved and where time goes — retrieval versus synthesis — so a slow pipeline points to the exact stage instead of a vague 'RAG is slow'.

Failure attribution and health

Failed queries, embedding and LLM errors, and timeouts classified and traced to the component that broke, with health scoring and predictive alerts on degrading pipelines.

HOW TO CONNECT

Three steps. No code changes.

1STEP 01

Instrument with the SDK or OpenTelemetry

Wrap your LlamaIndex app with the Obsivara Python SDK, or point an OpenTelemetry / OpenInference LlamaIndex instrumentation at Obsivara's OTLP endpoint. Your pipeline keeps calling models directly — Obsivara records each run out-of-band, with no proxy in the path and no added latency.

2STEP 02

Run your queries

Traces flow automatically as query engines and retrievers execute — no changes to your indexes, retrievers, or query logic.

3STEP 03

See it in Obsivara

Pipelines appear in the inventory and knowledge map with per-query cost, latency, health scores, and full traces within minutes.

WHAT WE MONITOR

Signals tracked out of the box.

  • →Query, retrieval, and synthesis traces
  • →Tokens and cost per query and pipeline
  • →Embedding vs LLM call cost split
  • →Retrieval latency and nodes per query
  • →Error rates and failing components
  • →Health score and predictive alerts
INTEGRATION FAQ

Common questions.

Does Obsivara work with LlamaIndex query engines and agents?

Yes. Query engines, retrievers, node post-processors, and LlamaIndex agents are traced as a span tree, so you can drill from a query down into each retrieval and LLM call.

Do I need to change my LlamaIndex code?

No. You instrument via the Obsivara SDK wrapper or an OpenTelemetry/OpenInference exporter; your indexes, retrievers, and query logic stay exactly as they are.

Can Obsivara track RAG cost per query?

Yes. Embedding, re-ranking, and synthesis calls are priced and rolled up per query, pipeline, and index, with waste detection on expensive retrieval patterns.

EXPLORE MORE

Other Frameworks integrations

Connect LlamaIndex in minutes.

No credit card · 5-minute setup