One governed route to every model
A single endpoint your coding agents and applications point at instead of a vendor. Behind it: any model you choose, credentials that never leave the platform, one policy, and every call attributed, capped and audited.
Your Tools, Unchanged
Point Claude Code, Codex, Cursor or any application speaking a standard completions API at one base URL. No plugin, no wrapper binary, no forked client.
Any Model Behind It
A frontier API today, open-weight models on your own hardware tomorrow — swapped by configuration rather than re-architecture, and never silently substituted.
Runs as a Person, Not a Key
Every turn resolves through your own identity provider to the human who made it, using their delegated grants — which is what makes spend and carbon per developer, and a monthly cap, possible at all.
AI vendors are steering large organisations off per-seat subscriptions and onto consumption contracts — the right economics for agents, and an open-ended invoice unless you hold a control point of your own. The LLM Router is that control point: one address every AI tool in the organisation talks to, so you can see what is being spent, by whom and on what, and set a ceiling before the month ends.
Read this if you're a platform lead or CTO standardising on coding agents, asking who holds the vendor keys and what the invoice will say.
The Scrydon LLM Router is a single governed endpoint that speaks the standard model APIs — Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini — so existing tools reach any model through it without modification. Every call runs under the caller's own identity against a model allowlist, is screened by data loss prevention, metered against a spend cap, and recorded in an immutable audit trail. Vendor keys stay on the platform and are never issued to a developer at all.
Every organisation standardising on coding agents arrives at the same fork, and both paths are bad: hand out raw vendor keys and lose attribution, model policy, DLP and audit in one gesture — or refuse the tools, and watch good engineers use them anyway on personal accounts with your source code as the prompt. The LLM Router is the third option. Existing tooling is pointed at one governed address and keeps working unchanged; behind that address the platform resolves the credential, enforces which models this person may reach, screens what leaves, meters the turn and writes the audit record. Because the router shares one policy and one audit chain with the platform's governed tool calls and network egress, governance does not stop at the model call — which is where an AI gateway has to stop.
LLM Router in the Scrydon platform
One integrated, sovereign architecture. Here is where LLM Router sits — highlighted against the full stack it works with.
The AI OS for Humans & AI Agents
Ontology & Semantic Layer, one connected model for your data, knowledge & processes
Combining the best of data lakes, data warehouses and search
Governed access to every model, and the agents & workflows that execute across your systems
Integrate across A2A, MCP, legacy systems and data sources
Secure domain federation, trusted data sharing, and cross-boundary intelligence
Sovereign Foundations
One route in, one policy, one audit chain
The router is a protocol edge, not a second execution path. A request in any supported dialect becomes one neutral request, and from there it takes the same governed route every other caller on the platform takes — credential resolution, the model allowlist, data loss prevention and policy, audit and metering all happen once, in the code that already owned them. The edge itself holds no credentials, so there is no path by which it can leak one.
The differentiator underneath all of it is identity. A gateway sits in front of your credentials; Scrydon is the system that holds them, so a call does not run as a key that stands for a team — it runs as a person, federated from your own identity provider, with that person's clearance deciding which models they may reach and their own delegated grants used when the agent goes on to call a tool. Attribution, revocation and least-privilege are then properties of the platform rather than conventions a caller is trusted to follow.
That shared route is what separates this from a proxy in front of your keys. The same identity, policy and audit chain also cover the tools your agents call and the network the sandbox executing them is allowed to reach — so the seams between the model, the tools and the network are not where your governance stops.
Speaks the APIs your tools already speak — Anthropic Messages, OpenAI Chat Completions, OpenAI Responses and Gemini arrive at the same endpoint and become one neutral request before anything is decided.
Identity comes from the credential, never the request — A router key is minted against a person, an organisation and an environment through your own identity provider. A request cannot name a tenant, and the platform refuses one that tries — so an agent can never be talked into acting as someone else.
Policy enforced at dispatch — Clearance-gated eligibility and an allowlist decide which models this person may reach, with data loss prevention and moderation running inline on both streaming and non-streaming calls.
Metered and audited before the answer lands — The turn is metered against the organisation's cap, and its audit record settles durably before the stream finishes — not after it, and not best-effort.
Two commands, and nothing about the work changes
The governance argument is only worth making if the developer experience is not worse than the ungoverned one. It is the same tools, the same commands, the same muscle memory — a sign-in, a minted key, a base URL in the environment, and the agent runs as it always did.

An unmodified coding agent, pointed at the router. The model is chosen by name from the catalog the organisation allows — here a GPT model served through a cloud tenancy — and switched with the client's own command. No plugin, no wrapper, no forked client.
For the administrator the corresponding gesture is equally small, and it is the one that matters at scale. Enable an integration once for the organisation and every developer's agent has those tools on its next turn, each call using that developer's own delegated grant through the agent's scoped identity: the agent can do what its human can do, and nothing more. No per-developer OAuth dance, no key distribution, and no vendor key on any laptop to be lost with it.
Sign in, mint a key — A device-authorisation flow in the browser the developer is already signed into, then an environment-scoped key. No vendor key is issued, because none is needed.
A base URL and a key — Set them in the environment and Claude Code, Codex, Gemini CLI, Cursor and Aider run governed — with prompt caching preserved, which matters when an agent re-sends its whole context every turn.
Enable an integration once — Turn a system on for the organisation and every developer's agent picks up those tools on its next turn — under that developer's own delegated credential, never a service account with organisation-wide reach.
Revoke with the same gesture — Disable the product and the next turn re-resolves. There is no per-developer rollout, and no window in which half the team still has the capability.
The shift to token pricing is happening either way
Anthropic and its peers are steering larger organisations from per-seat subscriptions onto token-based, consumption contracts. That is the right economics for agents, and it is also an open-ended invoice unless you hold a control point of your own. One governed endpoint keeps the shift on your terms: existing tooling unchanged, any model behind it, and every call measured, attributed and capped — so a token contract becomes a managed line item rather than a surprise.

Spend, calls, tokens and carbon for the month, attributed per developer and broken down by capability, model and workflow — with the on-pace projection against the ceiling you set, and an export when finance asks.
Two objections usually come next, and both have answers you can check. Governance is no longer the expensive part: allowlisting, DLP, moderation, metering and audit cost single-digit milliseconds of router time, less than the handshake a direct vendor call pays when connections are not pooled. And the router does not lock you to a vendor — it is the thing that stops one locking you in, because the model you pick today will not be the best one in a year, and swapping it should be configuration, not re-architecture.
It also stands alone. There is no ontology to build first, no process flows to map and no programme to launch: point your applications and coding agents at it, and the first week produces a spend report you did not have. When you later want governed agents or grounded retrieval, the route, the identity, the policy and the audit trail are already there.
LLM Router vs an AI gateway vs raw vendor keys
An AI gateway is a good answer to the question it asks — govern the model traffic. But watch a developer work with an agent for an afternoon and only a minority of the calls are model calls; the rest reach systems that hold your actual data. A gateway sits in front of your credentials. Scrydon is the system that holds them, alongside the tool catalog, the sandbox those tools execute in, and the boundary around that sandbox.
| Capability | Scrydon | AI gateway | Raw vendor keys |
|---|---|---|---|
| Who the call runs as | The person, resolved through your identity provider, with their own delegated grants | A key or virtual key, which stands for a team at best | Whoever is holding the key |
| Model calls governed | Allowlist, clearance, DLP, moderation, metering and audit on every call | Yes — this is what a gateway is for | None |
| Tool calls governed | Same policy, same credential model, same audit chain | Partly — where it also fronts a tool protocol | None |
| Network egress from the sandbox | Enforced outside the workload, which cannot reconfigure it | Out of scope — not in that path | None |
| Where the vendor key lives | On the platform; never issued to a developer | Vaulted by the gateway | A shell profile on a laptop |
| Cost attribution | Per developer and per turn, by model, capability and workflow | Per key or per virtual key | One undifferentiated invoice |
| Deployment | Sovereign — air-gapped, on-premises or European cloud | Varies; several are self-hostable | The vendor's cloud, by definition |
A category comparison rather than a product-by-product one: mature AI gateways govern model traffic well, and we would not pretend otherwise. The distinction we draw is one of scope, not of quality.
Frequently asked questions
What is an LLM router?+
Do developers have to change how they work?+
Who does a call actually run as?+
How is an LLM router different from an AI gateway?+
How much latency does the governance add?+
Can the router silently switch my model to a cheaper one?+
Which models can sit behind it?+
Can we start with just the router?+
Explore the platform
Prefer to write? Email hello [at] scrydon.com and we will get back to you.