Control & Trust·production

IAM.Router

Policy-driven LLM traffic routing by model, provider and perimeter.

Key capabilities

  • OpenAI-compatible inference gateway
  • Routing across local, on-premises and external models
  • Quotas, fallback, decision audit and provider health

Step-by-step guides

Use cases

Start with the outcome: open a guide, prepare prerequisites, follow the steps and verify the success signals.

01 Configure routing by model, provider and perimeterA logical model alias selects only permitted endpoints in the configured order.
Audience
AI platform engineer
Outcome
A logical model alias selects only permitted endpoints in the configured order.

Before you start

  • Provider credentials stored in a Secret
  • Verified model endpoint
  • Tenant route policy

Steps

  1. Add an endpointIn Providers, create an endpoint and set protocol, model id, perimeter and health check.
  2. Create a logical aliasIn Models, link the alias to compatible endpoints and context limits.
  3. Build the policySet tenant, sensitivity, allowed perimeter, priority and budget.
  4. Test in PlaygroundSend a request with the logical alias and confirm Decision shows the expected route.
Verify the result

Decision contains the expected tenant, policy, provider/model and perimeter.

If it does not work

For an unexpected route, check rule ordering, sensitivity label and current provider health.

02 Configure fallback and retries without duplicate billingA transient provider error is handled by bounded retries or a compatible fallback route.
Audience
SRE or AI platform engineer
Outcome
A transient provider error is handled by bounded retries or a compatible fallback route.

Before you start

  • At least two compatible endpoints
  • Client idempotency key
  • End-to-end request deadline

Steps

  1. Classify errorsRetry only timeouts, 429 and selected 5xx responses; do not retry 4xx policy errors.
  2. Set backoffSet max attempts, exponential backoff, jitter and a share of the total deadline.
  3. Add fallbackThe next endpoint must support the same contract, sensitivity and perimeter.
  4. Run a resilience testTemporarily disable the primary endpoint and confirm a single billable result.
Verify the result

Audit shows the attempt sequence, final route and a single charge record.

If it does not work

Repeated loops indicate deadline or idempotency is not propagated through the client and gateway.

Application sections

Open detailed manual

Every application screen has a separate page with controls, safe example values, CLI/API alternatives and status-specific recovery steps.

Open detailed manual →

Role in the ecosystem

IAM.Router separates an application from a specific LLM provider. The client sends OpenAI compatible request, and Router applies tenant policy, selects model and perimeter, controls the quota and records the decision.

Main scenario

POST /v1/chat/completions
Authorization: Bearer <scoped-key>
Content-Type: application/json

{"model":"policy/default","messages":[{"role":"user","content":"..."}]}

The requested model name can be a logical alias. The final provider is determined policy, health, data requirements and available budget.

Routing

Router supports several execution classes:

  • local/on-prem — data does not leave the matched circuit;
  • dedicated — dedicated endpoint for a tenant or project;
  • external — allowed cloud provider;
  • fallback — the next compatible route after retriable failure.

Retry is performed only for safe replays and is limited by deadline. Provider error should not endlessly multiply one request or bypass policy.

Integrations

IAM.Secure can inspect prompt/response before and after route. IAM.Identity specifies the tenant scope. IAM.Bot, Marketplace and other applications use Router as transport, saving own domain state regardless of the result of the LLM call.

Quotas and observability

Tenant/model/provider, tokens, latency, route decision and error are monitored class. For hosted agents, the budget is reserved before the provider call and fixed after billable result. The provider’s health is a route selection signal, not replacement for explicit policy.

Safety

  • inference keys have tenant scope;
  • administrative routes are separated from /v1/*;
  • provider secrets are not returned to the client;
  • DLP/route outcome is in the audit trail;
  • production policy is not weakened for the sake of fallback.

Before production

Check the list of models, real completion for each allowed provider, timeout/retry, budget limits, behavior at 429/5xx and security correlation id through the entire route.