Fireworks inference, measured and priced.
Fireworks AI runs open models and your fine-tunes at speed, behind an OpenAI-compatible API; Obsivara tells you what each call costs and how fast it really is. Every completion and embedding traced with tokens, time-to-first-token, latency, and dollars, priced per model and rolled up per workflow — so a slow or pricey route points to the exact model. Instrument via the Obsivara SDK or OpenTelemetry, or POST a usage event, out of the request path.
Fireworks AI, fully observable.
Per-model cost and token accounting
Tokens and dollar cost for every request, priced per Fireworks-hosted model (including your fine-tunes) and attributed to the workflow that made it.
Speed and latency visibility
Time-to-first-token and P50/P95/P99 latency per model, so the inference speed you chose Fireworks for is measured on your real traffic.
Fine-tune vs base comparison
Compare your fine-tuned deployments against base models on cost, latency, and reliability — so a fine-tune has to earn its keep.
Reliability monitoring
Rate-limit tracking, error classification, and anomaly alerts across every endpoint you use.
Three steps. No code changes.
Add your Fireworks usage source
Instrument your app with the Obsivara SDK or an OpenTelemetry GenAI exporter, or POST one usage event per request. Fireworks' OpenAI-compatible API is captured without special handling. Your app keeps calling Fireworks directly — Obsivara records each call out-of-band, so nothing routes through us and no latency is added.
Obsivara maps your usage
Within minutes you get a live inventory of every Fireworks model and fine-tune in use, the workflows calling them, and their cost.
Set budgets and alerts
Spend thresholds, latency SLOs, and error-rate alerts per workflow. Obsivara watches from there.
Signals tracked out of the box.
- →Tokens and cost per model and call
- →Time-to-first-token and latency by model
- →Fine-tune vs base model cost
- →Rate-limit and error rates
- →Cost per workflow and day
- →Latency drift against baseline
Common questions.
Does Obsivara track my Fireworks fine-tunes separately?
Yes. Fine-tuned deployments are tracked as their own models, so you can compare them against base models on cost, latency, and reliability.
Fireworks is OpenAI-compatible — does Obsivara just work?
Yes. Instrument with the Obsivara SDK, OpenTelemetry, or usage events; the OpenAI-compatible shape is captured without special configuration.
Will this add latency?
No. Obsivara never sits in the request path — your app calls Fireworks directly and reports usage afterward, so the speed you came for is untouched.