https://api.kimi.com/coding/v1.
GoModel routes chat, model listing, embeddings, and passthrough requests through the shared
OpenAI adapter. The /v1/responses endpoint is forwarded natively to the upstream Responses
API instead of being translated through chat completions, while files and batches are not
supported by the upstream endpoint.
Kimi Code retains no responses. Requests with store: true are rewritten to store: false
(the upstream rejects store: true with a 400). Chaining works only through GoModel: with a
response store configured, the gateway expands a previous_response_id chain by replaying the
stored history into the request before dispatch, and a conversation reference resolves
through the conversation store the same way. Without those stores, a request carrying
previous_response_id or conversation is rejected with an invalid-request error, because
the upstream can never resolve the referenced state.
Configure
config.yaml:
The provider
type is kimicode (no hyphen). The config key (kimicode in
the example above) is arbitrary and only identifies this entry inside your
config; it may be changed.Models
The standard chat model alias iskimi-for-coding. You can reference it in a virtual
model or call it directly through the OpenAI-compatible /v1/chat/completions endpoint.
Embeddings
Kimi Code exposesbge_m3_embed for embeddings, but this model is not documented
upstream. Treat it as experimental: test availability and dimensionality in your account
before relying on it in production workflows.
Pricing & quota
Kimi Code does not publish per-token pricing. Because GoModel’s usage-cost tracking has no price to multiply, it reportscost as zero for Kimi Code requests. The cost load-balancing
strategy cannot prefer or rank this provider by price, but it may still fall back to Kimi Code
when no priced provider is available.
Upstream quota works on a weekly refresh plus a rolling five-hour window. Plan for burst
limits and prefer conservative per-provider retry overrides: