Overview
GoModel accepts the Anthropic Messages API request dialect atPOST /v1/messages,
in addition to its OpenAI-compatible API. Clients and SDKs that speak the Anthropic
format can point at GoModel unchanged.
The request is translated to GoModel’s canonical chat type at ingress and runs through
the same pipeline as /v1/chat/completions — so virtual models, workflow policy,
budgets, failover, the response cache, usage/cost tracking, and audit logging all
apply. Because every provider implements chat completion, an Anthropic-format request
can be routed to any configured provider (OpenAI, Gemini, Bedrock, and others),
not only Anthropic.
This differs from the passthrough API: /p/anthropic/v1/messages
forwards bytes verbatim to the Anthropic upstream only, while the managed /v1/messages
endpoint routes anywhere and is fully managed.
Native forwarding to Anthropic
When a/v1/messages request resolves to an Anthropic provider, GoModel
skips the translation round-trip and forwards the original request body
verbatim (rewriting only the model field when an alias resolved to a
different name), then relays the provider-native response or SSE stream
unchanged. This preserves everything the canonical translation cannot —
cache_control breakpoints, thinking-block signatures, anthropic-beta
headers — which coding agents like Claude Code depend on. Rate limits,
budgets, audit logging, and usage tracking still apply, for both streaming
and non-streaming responses.
Native forwarding is automatic. Requests fall back to the translated pipeline
when a feature that operates on the canonical request is in play: guardrails
request patching, the response cache, or failover routing. Requests resolving
to any non-Anthropic provider always translate.
Because the body is forwarded verbatim, none of the translation limitations
apply on this path: server-tool history (server_tool_use, web_search_tool_result, …),
container uploads, and any other block the canonical request cannot represent reach
Anthropic unchanged.
Supported endpoints
Message Batches
/v1/messages/batches shares the gateway’s native-batch pipeline with the
OpenAI-compatible /v1/batches route — the two are dialect views of the same
resource. Batch IDs are interchangeable across the two dialects (msgbatch_<uuid>
here, batch_<uuid> there). All requests in one batch must resolve to a single
provider; per-item custom_id values are required and must be unique.
Providers whose batch API is file-based (OpenAI-compatible) receive inline
requests as an automatically uploaded JSONL input file. request_counts maps the
provider’s aggregate counts: while a batch runs, unfinished requests are reported
as processing; once it ends, any remainder is attributed by the batch outcome
(canceled, expired, or errored).
Authentication
Both credential styles work, so the official Anthropic SDKs are drop-in:Authorization: Bearer <key>— GoModel’s primary scheme.x-api-key: <key>— the Anthropic-native header, accepted as a fallback when noAuthorizationheader is present.
Example
type: "message", content blocks,
stop_reason, usage). Errors use the Anthropic error envelope
({"type": "error", "error": {...}}). max_tokens is required, as in the Anthropic API.
A provider’s content filter is reported as stop_reason: "refusal". If the
provider fails partway through a stream, the stream ends with an error event
carrying the provider’s own message.
usage follows Anthropic semantics for every provider: input_tokens excludes
cached tokens, which are reported separately as cache_read_input_tokens and
cache_creation_input_tokens. For OpenAI-compatible providers, whose
prompt_tokens includes cached tokens, GoModel moves cached_tokens and
cache_write_tokens into those fields, so the three counts add up to the whole
prompt.
Streaming responses emit the Anthropic SSE event sequence (message_start,
content_block_start/content_block_delta/content_block_stop, message_delta,
message_stop).
Cost tracking and audit logs
/v1/messages requests are tracked and audited exactly like the OpenAI-compatible
routes. Cost is computed from the actual provider that served the request, and usage
is recorded under the /v1/messages endpoint so it can be filtered in the dashboard.
Limitations
These limitations apply to the translated pipeline — requests routed to a non-Anthropic provider, or to Anthropic with guardrails, response cache, or failover engaged. Requests natively forwarded to Anthropic are preserved byte-for-byte apart from themodel value when an alias
resolved to a different name, and none of the below applies.
/v1/messages translates through GoModel’s canonical chat type. Anthropic-specific
features that have no canonical equivalent are not preserved end to end:
cache_controlis preserved on the request, system/content blocks, custom tools, and tool-use/tool-result history when routed to Anthropic.{"role": "system"}messages insidemessageskeep their position andcache_controlbreakpoints when routed to a Claude 4.8+/5-family model (Claude Code appends system reminders this way, and its prompt caching depends on them staying in place). On older Claude models, which reject the role, their text is hoisted into the top-level system prompt instead.thinkingandredacted_thinkingblocks travel both ways verbatim, signatures included: responses carry thesignatureAnthropic issued (as asignature_deltawhen streaming), and an assistant turn echoed back is replayed unchanged, so thinking-enabled conversations and tool-use loops continue correctly. Other providers never see them.- Reasoning from a provider that does not sign it (DeepSeek, Fireworks, …)
is returned as a
thinkingblock withsignature: ""— the member is required by the Anthropic schema, and the empty value says the reasoning is unsigned. Echo the turn back as-is: the block replays to that provider like any other. Anthropic accepts only signatures it minted itself, so a turn carrying an unsigned block is stripped of it before reaching Claude (the rest of the turn is sent unchanged). tool_result.is_erroris preserved when routed to Anthropic. Other providers have no equivalent flag and receive the result content only.tool_use.extra_contentcarries another provider’s replay state, such as a Gemini 3 thought signature, back to that provider. GoModel sets it ontool_useblocks it returns; echo it unchanged. See Extra content, thinking blocks, and thought signatures.- Server/built-in tools (web search, code execution, …) and their history
blocks (
server_tool_use,web_search_tool_result, …) are rejected with a clear400; only custom tools (typeabsent or"custom") translate. top_kis dropped — it has no portable OpenAI-compatible equivalent, and OpenAI-family providers reject unknown request fields.temperatureandtop_pare forwarded.- Images and documents inside
tool_resultblocks (screenshots, image and PDF files read by a tool — Claude Code returns them this way) are forwarded as nativeimageanddocumentblocks when routed to Anthropic. OpenAI receives them through its Responses API, where the model can see them (see OpenAI requests served through Responses). Other providers receive only the text portion of the tool result. documentblocks (PDF, plain text, URL, or Files APIfile_idsources) translate to the OpenAI-stylefilecontent part (file_datafor inline content,file_urlfor remote URLs,file_idfor uploads). Anthropic gets the document back natively with itstitle; Gemini receives inline data only; OpenAI-compatible providers receive thefilepart as-is. On OpenAI, an untitled PDF gets a filename (document.pdf), and URL documents go through the Responses API. A plain-text document on a request that stays on Chat Completions is sent as text (headed by itstitle), because only PDF file data is accepted there; on a request served through the Responses API, such as one with a text document inside atool_result, it stays a file. Citation settings andcontextare dropped. The custom-content variant (source.type: "content") degrades to text.search_resultblocks degrade to text (title, source URL, and content); citation metadata is dropped.- Other content blocks (
container_upload,tool_reference, …) are rejected with a clear400error rather than silently dropped. stop_sequencesare honored on every provider. Providers that report the matched sequence natively (Anthropic) get the full contract back:stop_reason: "stop_sequence"plus thestop_sequencevalue, and so do OpenAI reasoning models (GPT-5 and later, o-series), where GoModel cuts the output itself because OpenAI rejectsstopfor them. Other OpenAI-family models conflate stop-parameter hits with natural stops infinish_reason, so completions there reportstop_reason: "end_turn"(output is still truncated correctly).output_config:effortsets the reasoning effort (taking precedence over the one derived fromthinking), andformat(or the older top-leveloutput_format) becomes JSON-schema structured output — strict when the schema meets OpenAI’s strict-mode rules (additionalProperties: false, every property required, and noallOf,not, or conditional keywords). Other schemas are sent withstrict: false: they guide the output but do not guarantee it. A format other thanjson_schemais rejected with a400.thinkingon a non-Anthropic model becomes a reasoning effort (budget under 10000 tokens:low, under 20000:medium, otherwisehigh; adaptive:medium), sent as each provider’s own field:reasoning_efforton OpenAI, Groq, Fireworks, and xAI. Models that reject the field (OpenAIgpt-4*, Groq models other than gpt-oss and qwen3, xAI’s-non-reasoning,grok-build,grok-2andgrok-3exceptgrok-3-mini) get the request without it instead of an error. Reasoning comes back asthinkingblocks whichever member the provider uses for it (reasoning_contentorreasoning).count_tokensis exact when the model’s provider can count tokens itself: Anthropic models are counted by Anthropic’s owncount_tokensendpoint. The forwarded request carries the resolved model (an alias is resolved first) and only the members that endpoint accepts:messages,system,tools,tool_choice,thinking,cache_control, andoutput_config;max_tokens,stream,metadata, and the sampling controls are left out because Anthropic rejects them there. OpenAI models are counted by OpenAI’s/v1/responses/input_tokensendpoint for the request as GoModel will send it. On GPT-5 and later the count matches the billed input (within a token): for a request Chat Completions serves with tools, the fixed tools preamble it bills is added. On older model families such a request can count a few tokens below the billed input (6 ongpt-4o-mini). Guardrails do not run for a count, so when the request’s workflow has prompt guardrails the prompt is never sent to a provider for counting and the estimate answers (response- and stream-only guardrails keep exact counts). For every other provider, and whenever that call fails, the gateway answers with an estimate that weights text by how densely it tokenizes (prose, code and JSON, CJK, emoji), adds per-message and per-tool framing plus the tool-use system prompt, and prices images by their pixel area. On ordinary agent traffic the estimate lands within about ten percent of a tokenizer-exact count; raw base64 in text stays under-counted, and an image whose size cannot be read is charged the largest size the provider keeps. The same estimate seedsusage.input_tokensin the streamingmessage_startevent; the authoritative counts arrive in the finalmessage_deltaevent, which SDK accumulators prefer.
/p/anthropic/v1/messages passthrough route instead.
See ADR-0007
for the design rationale and tradeoffs.