INTEGRATION / LLM PROVIDERS

Every open model on Together, accounted for.

Together AI serves open models — Llama, Mixtral, Qwen, and more — behind an OpenAI-compatible API; Obsivara tells you what each one costs and how it runs. Every completion and embedding traced with tokens, latency, and dollars, priced per hosted model and rolled up per workflow, so you can tell which open model earns its place. Instrument via the Obsivara SDK or OpenTelemetry, or POST a usage event, out of the request path.

Together AIEXAMPLE
CALLS TRACED / 24H128,412
TOKENS ACCOUNTED / 24H41.2M
SPEND ATTRIBUTED / 24H$214.60
ILLUSTRATIVE · EXAMPLE DATA
WHAT YOU GET

Together AI, fully observable.

Per-model cost and token accounting

Tokens and dollar cost for every request, priced per Together-hosted model and attributed to the workflow that made it.

Open-model comparison

Compare the open models you run on Together on cost, latency, and reliability for the same workload — so picking Llama over Mixtral is a measured call.

Latency and error monitoring

P50/P95/P99 latency per model, rate-limit tracking, and error classification across every endpoint you use.

Usage anomaly detection

Token spikes, runaway retries, and prompt-size creep flagged before they reach your invoice.

HOW TO CONNECT

Three steps. No code changes.

1STEP 01

Add your Together usage source

Instrument your app with the Obsivara SDK or an OpenTelemetry GenAI exporter, or POST one usage event per request. Together's OpenAI-compatible API is captured without special handling. Your app keeps calling Together directly — Obsivara records each call out-of-band, so nothing routes through us and no latency is added.

2STEP 02

Obsivara maps your usage

Within minutes you get a live inventory of every Together-hosted model in use, the workflows calling it, and its cost.

3STEP 03

Set budgets and alerts

Spend thresholds, latency SLOs, and error-rate alerts per workflow. Obsivara watches from there.

WHAT WE MONITOR

Signals tracked out of the box.

  • →Tokens and cost per model and call
  • →Latency P50 / P95 / P99 by model
  • →Cost per workflow and day
  • →Rate-limit and error rates
  • →Model switches between deploys
  • →Prompt size creep over time
INTEGRATION FAQ

Common questions.

Together uses an OpenAI-compatible API — does Obsivara just work?

Yes. Instrument with the Obsivara SDK, OpenTelemetry, or usage events; the OpenAI-compatible request and response shape is captured without special configuration.

Can I compare the open models I host on Together?

Yes. Every model is tracked uniformly, so you can compare Llama, Mixtral, Qwen, and others on cost, latency, and reliability on your own traffic.

Will this add latency?

No. Obsivara never sits in the request path — your app calls Together directly and reports usage afterward, so there's no proxy and no added latency.

EXPLORE MORE

Other LLM Providers integrations

Connect Together AI in minutes.

No credit card · 5-minute setup