Skip to main content
Kimi Code is an OpenAI-compatible coding assistant served at https://api.kimi.com/coding/v1. GoModel routes chat, model listing, embeddings, and passthrough requests through the shared OpenAI adapter. The /v1/responses endpoint is forwarded natively to the upstream Responses API instead of being translated through chat completions, while files and batches are not supported by the upstream endpoint. Kimi Code retains no responses. Requests with store: true are rewritten to store: false (the upstream rejects store: true with a 400). Chaining works only through GoModel: with a response store configured, the gateway expands a previous_response_id chain by replaying the stored history into the request before dispatch, and a conversation reference resolves through the conversation store the same way. Without those stores, a request carrying previous_response_id or conversation is rejected with an invalid-request error, because the upstream can never resolve the referenced state.

Configure

Or in config.yaml:
The provider type is kimicode (no hyphen). The config key (kimicode in the example above) is arbitrary and only identifies this entry inside your config; it may be changed.
You can also override the base URL and model list with:

Models

The standard chat model alias is kimi-for-coding. You can reference it in a virtual model or call it directly through the OpenAI-compatible /v1/chat/completions endpoint.

Embeddings

Kimi Code exposes bge_m3_embed for embeddings, but this model is not documented upstream. Treat it as experimental: test availability and dimensionality in your account before relying on it in production workflows.

Pricing & quota

Kimi Code does not publish per-token pricing. Because GoModel’s usage-cost tracking has no price to multiply, it reports cost as zero for Kimi Code requests. The cost load-balancing strategy cannot prefer or rank this provider by price, but it may still fall back to Kimi Code when no priced provider is available. Upstream quota works on a weekly refresh plus a rolling five-hour window. Plan for burst limits and prefer conservative per-provider retry overrides:
Last modified on October 4, 2026