Guides10 min read
By the Obsivara team

OpenTelemetry for LLMs: Instrumenting GenAI with OTLP

How to instrument LLMs and GenAI apps with OpenTelemetry and OTLP — the gen_ai semantic conventions, what each span should capture, and where to send it.

Quick answer

OpenTelemetry (OTel) gives GenAI apps a vendor-neutral way to trace LLM calls: you instrument once, emit spans using the gen_ai.* semantic conventions, and export them over OTLP to any compatible backend. Capture the model, prompt and completion token counts, latency, and tool calls per span. Because OTel is out of band, nothing sits in your request path — and you can repoint the same telemetry at a new tool without re-instrumenting.

OpenTelemetry (OTel) is the open standard for application telemetry — traces, metrics, and logs — and over the last two years it has grown a dedicated vocabulary for generative AI. That matters because LLM apps are hard to see into: a single user request can fan out into a chain of model calls, retrievals, and tool executions, and a plain access log tells you none of it. Instrumenting with OpenTelemetry lets you capture that whole shape as a trace of spans, then export it over OTLP (the OpenTelemetry Protocol) to whatever backend you choose — without binding your instrumentation to one vendor. This guide covers what OTel captures for LLMs, the gen_ai semantic conventions, how to wire up popular frameworks, and where raw spans stop being enough for running AI in production.

What OpenTelemetry gives an LLM app

A trace is the record of one end-to-end operation; a span is one step inside it. For a chatbot, the trace might be a single conversation turn, with child spans for the retrieval query, the embedding call, the chat completion, and any tool the model invoked. Each span carries a start time, a duration, a status, and a bag of attributes — and it is those attributes that turn a generic trace into an LLM trace you can actually reason about.

The key design property is that OpenTelemetry is out of band. Your application still calls the model provider directly; the SDK records what happened and ships a copy of it to a collector asynchronously. Nothing sits between your code and OpenAI or Anthropic, so instrumentation adds no proxy hop to the latency your users feel and introduces no new point of failure in the request path. That is the same out-of-band principle Obsivara is built on, which is why OTLP is a first-class ingestion path for it rather than an afterthought.

The gen_ai semantic conventions

Semantic conventions are the agreed attribute names, so that a span emitted by one library means the same thing to any backend. For generative AI they live under the gen_ai.* namespace, and using them — rather than inventing your own keys — is what makes your telemetry portable across tools. These are the attributes worth capturing on every model-call span:

AttributeWhat it records
gen_ai.systemThe provider — openai, anthropic, cohere, and so on
gen_ai.operation.nameThe operation — chat, text_completion, embeddings, or execute_tool
gen_ai.request.modelThe model requested, e.g. gpt-4o or claude-sonnet
gen_ai.request.temperature / max_tokensThe sampling parameters sent with the call
gen_ai.usage.input_tokensPrompt tokens consumed — the basis for cost
gen_ai.usage.output_tokensCompletion tokens generated
gen_ai.response.model / finish_reasonsWhat the provider actually served, and why it stopped

A note on versions and privacy

The GenAI conventions are still marked experimental, so expect attribute names to shift between releases — pin your instrumentation versions and prefer libraries that track the spec closely. Capturing the full prompt and completion text is opt-in: the token counts and metadata travel by default, but most teams keep raw content off in production for privacy and payload reasons, and turn it on selectively when they need to debug a specific failing run. Decide that policy once, at the instrumentation layer, so it is consistent across every service.

How to instrument your LLM stack

There are three layers you can instrument, and most production setups combine them. The first is auto-instrumentation libraries. Projects such as OpenLLMetry and OpenInference wrap the popular provider SDKs and frameworks so that gen_ai spans appear with no manual span code — you install the package, point it at an OTLP endpoint, and calls to OpenAI, Anthropic, and others start emitting traces. This is the fastest way to broad coverage.

The second is framework-native tracing. If you build retrieval apps on LlamaIndex, its callback and instrumentation hooks emit spans for each stage of a query — retrieval, synthesis, and the LLM call — so a RAG pipeline traces itself end to end; see how to connect LlamaIndex for the wiring. On the TypeScript side, the Vercel AI SDK has built-in OpenTelemetry support: enable its experimental telemetry on a generateText or streamText call and it emits spans with token usage and model metadata you can export over OTLP.

The third is manual spans for your own logic. Wrap the code between model calls — the orchestration, the business rules, the post-processing — in your own spans, so the trace reflects your application and not just the provider calls. Tag those spans with the workflow, environment, and customer, so you can slice cost and errors by those dimensions later. Whichever layers you use, they all terminate the same way: an OTLP exporter sends spans to an OpenTelemetry Collector or straight to a backend. Point that exporter at Obsivara's OTLP endpoint and the spans you already emit become an operations view, with no second instrumentation to maintain.

Where OTLP data stops being enough

OpenTelemetry solves capture and transport brilliantly. What it does not do is operate your AI for you. Raw spans in a trace viewer answer 'what happened in this one run?' — they do not answer the questions production teams actually live with: which workflow's cost is creeping week over week, which tool call is quietly driving half your spend, whether this agent's error rate is climbing against its own baseline, or which runs finished green but wrong. Those are cross-run, trend-shaped questions, and a per-trace view cannot surface them on its own.

There is a second gap worth naming: a gen_ai span records token counts, not dollars. Pricing those tokens per model and rolling the total up per workflow is a layer you add on top of the raw telemetry. Without it, you can see that a call used 12,000 tokens but not that one workflow spent $312 last month and a third of it was retries.

From OTLP spans to AI operations

This is where an operations layer sits on top of your OpenTelemetry pipeline. Obsivara ingests gen_ai spans over OTLP — alongside its SDK, generic webhook, and native read-only n8n paths — then does the analysis OTel leaves to you: it prices every call into cost intelligence per model, agent, and workflow; scores each asset's health against its own history; lets you replay any failed run from its captured spans with trace replay; and raises predictive alerts when latency, error rate, or token usage drifts, before the degradation reaches users. Because it reads the OTLP stream out of band, nothing changes in your request path.

If you have already instrumented with OpenTelemetry, you are most of the way there: pointing the exporter at a new backend is a config change, not a re-instrumentation, which is the entire point of the standard. You can start on the free tier and move up as your estate grows — see pricing — or run a free 48-hour audit to see your real per-model cost and failure rate on your own traffic first.

Frequently asked questions

What is OpenTelemetry for LLMs?

It is the use of the OpenTelemetry standard to trace generative-AI applications: each model call, retrieval, and tool execution is recorded as a span with gen_ai.* attributes (provider, model, token counts, latency), assembled into an end-to-end trace, and exported over OTLP to a backend. Because OpenTelemetry works out of band, it adds no proxy to your request path, and the same instrumentation can send data to any compatible tool.

What are the gen_ai semantic conventions?

They are OpenTelemetry's agreed attribute names for generative AI, under the gen_ai.* namespace — for example gen_ai.system (the provider), gen_ai.request.model, gen_ai.usage.input_tokens, and gen_ai.usage.output_tokens. Using them makes your telemetry portable: a span emitted by one library means the same thing to any backend. The conventions are still experimental, so attribute names can change between releases, and capturing raw prompt and completion text is opt-in.

Do I need a proxy to use OpenTelemetry with my LLM app?

No. OpenTelemetry is out of band — your app keeps calling the model provider directly and the SDK ships a copy of the telemetry asynchronously, so nothing sits in the request path. That is different from gateway or proxy tools that route your traffic through themselves. Obsivara ingests the OTLP stream the same way, with no proxy, so instrumentation never adds latency or a new point of failure.

See your AI operations clearly.

Connect your stack in minutes — Obsivara discovers every workflow, scores its health, and attributes every dollar.

MORE POSTS