Every open model on Together, accounted for.
Together AI serves open models — Llama, Mixtral, Qwen, and more — behind an OpenAI-compatible API; Obsivara tells you what each one costs and how it runs. Every completion and embedding traced with tokens, latency, and dollars, priced per hosted model and rolled up per workflow, so you can tell which open model earns its place. Instrument via the Obsivara SDK or OpenTelemetry, or POST a usage event, out of the request path.
Together AI, fully observable.
Per-model cost and token accounting
Tokens and dollar cost for every request, priced per Together-hosted model and attributed to the workflow that made it.
Open-model comparison
Compare the open models you run on Together on cost, latency, and reliability for the same workload — so picking Llama over Mixtral is a measured call.
Latency and error monitoring
P50/P95/P99 latency per model, rate-limit tracking, and error classification across every endpoint you use.
Usage anomaly detection
Token spikes, runaway retries, and prompt-size creep flagged before they reach your invoice.
Three steps. No code changes.
Add your Together usage source
Instrument your app with the Obsivara SDK or an OpenTelemetry GenAI exporter, or POST one usage event per request. Together's OpenAI-compatible API is captured without special handling. Your app keeps calling Together directly — Obsivara records each call out-of-band, so nothing routes through us and no latency is added.
Obsivara maps your usage
Within minutes you get a live inventory of every Together-hosted model in use, the workflows calling it, and its cost.
Set budgets and alerts
Spend thresholds, latency SLOs, and error-rate alerts per workflow. Obsivara watches from there.
Signals tracked out of the box.
- →Tokens and cost per model and call
- →Latency P50 / P95 / P99 by model
- →Cost per workflow and day
- →Rate-limit and error rates
- →Model switches between deploys
- →Prompt size creep over time
Common questions.
Together uses an OpenAI-compatible API — does Obsivara just work?
Yes. Instrument with the Obsivara SDK, OpenTelemetry, or usage events; the OpenAI-compatible request and response shape is captured without special configuration.
Can I compare the open models I host on Together?
Yes. Every model is tracked uniformly, so you can compare Llama, Mixtral, Qwen, and others on cost, latency, and reliability on your own traffic.
Will this add latency?
No. Obsivara never sits in the request path — your app calls Together directly and reports usage afterward, so there's no proxy and no added latency.