Appearance
Model Routing and API Behavior
See also: Configuration Reference, Provider API Compatibility, Data Relationships, Identity and Access, Request Lifecycle and Failure Modes, Pricing Catalog and Accounting, Observability and Request Logs, ADR: Model Aliases and Provider-Only Route Config, ADR: Capability-Aware Route Gating with Strict Fail-Fast Validation, ADR: Route-Level Provider API Compatibility Profiles
This page explains how the public /v1/* surface resolves a request into one concrete route.
Source of Truth
- config parsing:
- model access and tag selection:
- alias resolution:
- route planning:
- HTTP handlers:
Public Endpoints
The live public endpoints are:
GET /v1/modelsPOST /v1/chat/completionsPOST /v1/messagesPOST /messagesPOST /v1/responsesPOST /v1/embeddings
All are authenticated.
Requested Versus Resolved Model Identity
The gateway keeps two model identities in play:
- requested model
- what the caller asked for
- resolved model
- the canonical execution target after alias resolution
That distinction is persisted into request logs.
Model Forms
Configured gateway models are either:
- provider-backed
- alias-backed
A model cannot define both routes and alias_of.
Aliases are independent gateway model keys for authorization. An allowlist on an alias does not inherit from the target model, and an allowlist on the target model does not automatically apply to the alias.
tag: Selectors
The request model field can be:
- a concrete gateway model key
- a tag selector such as
tag:fast
Tag selectors use AND semantics.
- every requested tag must exist on the chosen model
- selection only considers models in the caller's effective accessible set
- API-key grants, principal-centric restrictions, and model-level allowlists all contribute to that effective set
- blocked allowlisted models are skipped as candidates rather than selected and denied later
- candidates are ordered by model
rank, then model key
Routes, Priority, and Weight
Provider-backed models resolve to one or more routes.
Each route can define:
providerupstream_modelpriorityweightenabledcapabilities
Current planner behavior:
- lower
priorityis attempted first weightonly matters within the same priority bucket- disabled routes and routes with non-positive weight are excluded
Current runtime nuance:
- weighted routing is not multi-route fallback
- the planner produces an ordered route list
- the handler executes only the first eligible route
Weight affects selection inside a single priority bucket. It does not mean the gateway sends one request to several providers, retries the next route, or falls back after an upstream error. Configurable retry and fallback remains separate follow-up work in issue #118.
Capability-Aware Gating
Routes are filtered before provider execution based on request requirements and route capability.
Current capability dimensions:
chat_completionsresponsesstreamembeddingstoolsvisionjson_schemadeveloper_role
Capability metadata exists to fail early at the gateway edge. It is not a copy of provider marketing language.
Effective capability is the intersection of route metadata and provider runtime support.
- route capability defaults are permissive
- provider implementations can still reject unsupported API families
- partial provider routes should explicitly disable unsupported API families
For example, current Vertex routes support the chat path but not the Responses path. A Vertex chat route should keep responses: false and embeddings: false so /v1/responses or /v1/embeddings fails during capability filtering instead of later inside the provider adapter. A separate Vertex text-embedding route can set embeddings: true for supported Google embedding models.
Cloud Run OpenAI-compatible routes use the same route capability model as ordinary openai_compat routes. If the deployed vLLM service only exposes Chat Completions, keep responses: false and embeddings: false on that route even though the provider adapter can speak those OpenAI-compatible families when the upstream supports them.
Compatibility Profiles
Routes can also define provider API compatibility metadata.
Capabilities and compatibility have different jobs:
capabilitiesgates whether the route can execute a request at allcompatibilityrewrites the outbound provider request shape after a route is selected
OpenAI-compatible route profiles currently cover deterministic Chat Completions transforms such as store removal, token field renaming, developer role rewriting, reasoning_effort handling, and stream usage requests. Responses uses a separate typed request/provider path; Chat Completions transforms must not be used as Responses shims.
OpenRouter routes can also define compatibility.openrouter.provider policy. That policy is serialized into the upstream Chat Completions request body as OpenRouter's provider object after Oceans has selected one gateway route. OpenRouter order, only, ignore, zdr, latency, and price settings affect OpenRouter's upstream provider selection for that one request; they do not change Oceans route priority, route weight, or the current single-route execution behavior.
See provider-api-compatibility.md for the compatibility matrix and field-level contract.
Cloud Run vLLM/Gemma controls such as chat_template_kwargs.enable_thinking and skip_special_tokens are additive upstream request fields. Put them in route extra_body, not in a compatibility profile.
Worked Request Path
One plain path looks like this:
- request model:
tag:fast
- allowed model set:
gpt-4o-mini,claude-3-5-haiku
- selection result:
gpt-4o-mini
- alias result:
openai-gpt-4o-mini
- planned route order:
openai-primary, thenopenai-backup
- capability filter:
openai-primarystays eligible
- execution:
- the handler uses
openai-primary
- the handler uses
- request-log fields:
model_key = gpt-4o-miniresolved_model_key = openai-gpt-4o-miniprovider_key = openai-primary
Use request-lifecycle-and-failure-modes.md for the later logging, pricing, and budget effects.
/v1/models
GET /v1/models returns the gateway models visible to the authenticated API key.
Important notes:
- it reflects gateway model identity, not raw provider catalogs
- it shows grant-visible identities
- it does not promise executable routes
That last point matters. A model can be visible and still fail if route viability or capability checks remove every route.
/v1/chat/completions
Current behavior highlights:
- request IDs are assigned once at the HTTP middleware boundary and propagated through
x-request-id - budget checks run before provider execution
- successful requests write usage when usage can be normalized
- request logs store both requested and resolved model identity
/v1/messages
POST /v1/messages and the unversioned POST /messages alias accept Anthropic Messages-compatible request bodies for chat-capable routes. The gateway authenticates Anthropic-style x-api-key headers, converts the Messages request into the internal chat request boundary for model resolution and provider execution, then returns Anthropic-compatible response payloads or Messages SSE events.
For Anthropic-on-Vertex Claude routes, the Messages path supports text messages, system, max_tokens, stream, tools, tool_choice, and thinking fields that are compatible with the Vertex Anthropic mapper. OpenAI Chat Completions requests with function tools use the same Vertex tool mapping. Routes that cannot support tools should keep tools: false so capability filtering fails at the gateway edge.
/v1/responses
POST /v1/responses follows the same authentication, model resolution, route planning, budget guard, logging, and ledger flow as Chat Completions.
Important differences:
- route capability filtering requires
responses - provider execution calls the provider's Responses methods, not Chat Completions methods
- streaming preserves Responses
response.*event names and payloads instead of rewriting them into Chat Completions chunks - usage is normalized from
input_tokens,output_tokens, andtotal_tokens
/v1/embeddings
POST /v1/embeddings follows the same high-level path:
- authenticate
- resolve the requested model
- capability-filter the route set
- execute the first eligible route
- record the provider execution attempt when request logging writes a summary row
- write usage when usage can be normalized
Vertex text embeddings are supported only by explicit embedding routes for supported Google publisher models:
google/gemini-embedding-001google/gemini-embedding-2google/text-embedding-005google/text-multilingual-embedding-002
Those routes should set embeddings: true and disable unrelated API families such as chat_completions, responses, stream, tools, vision, and json_schema. Gemini chat routes should keep embeddings: false.
The OpenAI-compatible embeddings request supports string and array-of-strings input. The Vertex mapper rejects unsupported input forms before the provider call: token arrays, nested arrays, non-string values, empty arrays, empty strings, multimodal payloads, and encoding_format: "base64".
Examples:
| Request shape | Expected route behavior |
|---|---|
/v1/embeddings against an embedding-only google/gemini-embedding-001 or google/gemini-embedding-2 route | Route can pass capability filtering and execute through Vertex :predict or :embedContent, respectively. |
/v1/embeddings against a Gemini chat route with embeddings: false | Capability filtering removes the route and returns an edge error before Vertex execution. |
/v1/embeddings against google/gemini-2.0-flash with embeddings: true | Provider support checks reject the unsupported embedding model instead of treating all google/* models as embedding-capable. |
/v1/embeddings with input: [[1, 2, 3]] | The embeddings mapper rejects the unsupported OpenAI token-array/nested-array form locally. |
Route Viability Versus Capability Mismatch
| Symptom | Meaning |
|---|---|
invalid_request | the model resolved, but capability filtering removed every route |
no_routes_available | the model exists, but no usable route survived provider and route-viability checks |
That distinction is one of the fastest ways to debug a visible-but-unusable model.
Current V1 Behavior
The live runtime is intentionally narrow in this slice:
- single-route execution only
- no retry loop
- no live fallback loop
- request-attempt records describe the single provider execution attempt; configurable retry/fallback execution is tracked separately in issue #118
- strict capability filtering before provider execution
The open retry/fallback policy work must amend this section when it lands; see issue #118.
What This Page Does Not Own
- config field syntax and defaults:
- full cross-cutting request path:
- exact pricing coverage:
- budget behavior and setup:
