Vertex AI spend, across every model you deploy.
Vertex AI runs Gemini, Claude, Llama, and your tuned models behind one set of Google Cloud endpoints; Obsivara tells you what the whole platform costs and how each model behaves. Every prediction traced with tokens, latency, and dollars, rolled up per endpoint, project, and workflow — so Model Garden and tuned-model spend stop being one opaque Cloud line item. Instrument via the Obsivara SDK or OpenTelemetry, out of the request path.
Vertex AI, fully observable.
Platform-wide cost attribution
Every Vertex endpoint priced and rolled up per model, project, and workflow — Model Garden calls, tuned-model endpoints, and Claude-on-Vertex in one view instead of a single Google Cloud bill.
Per-model behavior across Model Garden
Gemini, Claude, Llama, Mistral, and open models compared on tokens, latency, and reliability — measured on your real Vertex traffic, not benchmarks.
Tuned and deployed endpoint visibility
Usage and cost for your custom-tuned and self-deployed endpoints tracked next to hosted models, so you see whether a tuned endpoint earns its provisioned capacity.
Reliability and quota monitoring
Error classification, quota pressure, and latency drift per endpoint and region, with health scoring and predictive alerts.
Three steps. No code changes.
Instrument with the SDK or OpenTelemetry
Wrap your Vertex AI client with the Obsivara SDK, or point an OpenTelemetry exporter at Obsivara's OTLP endpoint. Your app calls Vertex directly; Obsivara records each prediction out-of-band, with no proxy in the path and no added latency.
Run your workloads
Traces flow as your endpoints serve predictions — no changes to your models, endpoints, or client code.
See it in Obsivara
Every endpoint, model, and project appears in the inventory with per-call cost, latency, and health within minutes.
Signals tracked out of the box.
- →Tokens and cost per model and endpoint
- →Model Garden vs tuned-model spend
- →Latency by model, endpoint, and region
- →Quota and throughput pressure
- →Error rates and failing endpoints
- →Health score and predictive alerts
Common questions.
How is this different from the Obsivara Gemini integration?
The Gemini page covers the Gemini API's model and multimodal detail. This is the Vertex platform view: cost and reliability across every model you run on Vertex, including Claude, Llama, and your tuned endpoints, rolled up per project.
Does Obsivara cover Claude and other Model Garden models on Vertex?
Yes. Any model you invoke through Vertex — Gemini, Claude, Llama, Mistral, or open models — is tracked uniformly, so you can compare cost and latency side by side.
Do I need to change my Vertex AI code?
No. You instrument via the Obsivara SDK or an OpenTelemetry exporter; your endpoints, tuned models, and client code stay exactly as they are.