The unified control plane for your AI stack. Route to any model. Govern every tool call. Enforce hard spend limits. Audit everything.
# one endpoint, any model curl https://gateway.heimos.ai/v1/chat/completions \ -H "Authorization: Bearer $HEIMOS_KEY" \ -d '{ "model": "claude-sonnet-4-5", "messages": [{"role": "user", "content": "hello"}], "fallback": ["gpt-5", "gemini-2.5-pro"] }' # automatic failover · hard spend caps · guardrails · full audit trail 200 OK · X-Gateway-Overhead-Us: 1310 · $0.0004 · team:engineering · app:chatbot
ROUTE MODELS AND TOOLS THROUGH A SINGLE POLICY PLANE
Every gateway routes models. HeimOS is the only one that routes MCP tool calls through the same policy engine — same auth, same quotas, same guardrails, same audit trail. Your agents are governed the way your completions are.
one policy engine, both paths
One integration gives you routing, governance, observability, and security — without touching your application code.
One OpenAI-compatible endpoint for every provider. Drop in your existing SDK calls — OpenAI, Anthropic, Gemini, self-hosted Ollama, and the same models via Amazon Bedrock or Azure AI Foundry all respond through one interface. No cloud lock-in, zero code changes.
Automatic failover, latency-based routing, and weighted load balancing. Define fallback chains per model, so one provider's outage never becomes yours.
Spend is enforced atomically on the hot path — reserved before the request, reconciled after. A cap is a cap. No overshoot in the lag window, no surprise bills.
PII redaction, prompt injection defense, and content policy — enforced by fast local classifiers or the cloud engine you already pay for, including Amazon Bedrock Guardrails and Azure AI Content Safety. Redaction gates are sequential, so sensitive data never leaves the trust boundary. Fail-closed by default.
Native Model Context Protocol support with tool-level RBAC and argument scanning. Streamable HTTP transport. Your agents get the same spend controls and audit trail as your completions.
Every request logged with full metadata — model, tokens, cost, latency, team, app, user. Queryable, exportable, retained per policy. Built for teams who need to answer "who used what, when, and why."
Every other gateway makes you trade speed for control. HeimOS runs auth, quotas, guardrails, and audit on one box — and still gets out of the way.
Numbers from a pinned, reproducible k6 rig — run it yourself. See the full head-to-head benchmark →
HeimOS sits between your apps and your AI providers. Real-time config changes, zero downtime — you won't even notice it's there. It ships as containers and a config file, so it runs on AWS, Azure, GCP, or your own hardware without changing a line of your application code.
Change one base URL. Your existing OpenAI-compatible code keeps working — no SDK swap, no rewrite.
Every request passes through the gateway. Authentication, rate limiting, spend checks, and guardrails apply instantly. Requests that violate your policies are rejected before they cost you money.
The router picks the best provider based on your rules — primary model, fallback chain, latency targets, cost optimization. If a provider is down, traffic shifts automatically.
Request metadata, token counts, costs, latencies, and guardrail decisions are captured automatically. Durable, queryable, and built for compliance — not just dashboards.
Full applicability of the EU AI Act lands August 2, 2026. Logging, transparency, and data-protection obligations apply to every AI system you run. HeimOS ships the gateway-side controls — Art. 12 logging & retention, Art. 50 transparency — as configuration, not consulting.
Retention policies set per tenant, enforced in the audit store, with retention proof for your auditors.
Preconfigured guardrail bundles built around the Act's logging and transparency articles. Fail-closed by default.
Active policies, violations, retention proof, and full audit-trail export — generated on demand.
A hierarchical multi-tenant model designed for real organizations. Org admins set global policies, team leads manage their own budgets, and individual apps get their own API keys with fine-grained permissions.
No markup on model costs — you pay providers directly. HeimOS charges only for the gateway and governance layer.
Every team building with AI hits the same wall: fragmented provider APIs, no visibility into spend, no guardrails, and no audit trail. We've been there — managing keys across providers, waking up to surprise bills, wiring up custom logging for every new model.
HeimOS is the infrastructure layer we wished existed. Named after Heimdall, the all-seeing guardian of the Bifrost, HeimOS stands between your applications and the AI providers they depend on — routing traffic, enforcing policy, and keeping a complete record of everything that passes through.
Create your org, mint a key, and route your first request in under 5 minutes.