Two fronts in this release: multimodal work now runs through the same gateway as text, and every request can be tagged so cost lands on the service that spent it.
New image editing endpoints, wired into the LiteLLM service integrations — same keys, budgets and logs as text calls.
Auxiliary tooling script to check processing status on long-running video jobs.
Virtual keys and request logs now carry tags, so spend traces back to the service that made the call.
Filter by service tag, break down usage metrics, and pull enhanced reporting from the console.
Tags close the loop on user-level spend: every request carries the key and tag that made it, so you can track what each user and service costs — and cap it before the invoice lands.
Load balance across OpenAI, Anthropic, Gemini, Azure, and more — with automatic fallback and retry logic.
Set monthly budgets at the global, team, API key, or model level. Block or alert when thresholds are hit.
Keyword blocking, regex filtering, PII redaction, and content policy enforcement on every request.
Every LLM request logged with token counts, cost, latency, and model used. Query and export your data.
In-memory or Redis caching with semantic similarity matching. Cut costs and latency on repeated queries.
Per-key and per-team RPM/TPM limits backed by persistent database tracking. Prevent runaway usage before it hits your bill.
Supported Providers
OpenAI · Anthropic · Azure OpenAI · Google Gemini · AWS Bedrock · Ollama
Powered by PromptCaliper
The open-source foundation powering our Compliance and Governance vertical. Deploy it yourself and get full LLM governance out of the box.
Dev Environment
Also Forging
The MAM built for teams shipping AI-generated work at volume. Private beta launching soon.