Overview
GoModel exposes the OpenAI-compatible image endpoints for generation and editing. Clients and SDKs that already call OpenAI’s/v1/images/generations
or /v1/images/edits can point at GoModel unchanged.
Requests route by model through the same registry used for chat and
embeddings, so model selection, provider hints, virtual models, per-key model
access rules (user paths), budgets,
and rate limits all apply. Image generation is served by
OpenAI and the OpenAI-compatible providers that implement the endpoint
(Azure OpenAI, OpenRouter, xAI), and by
Google Gemini — Gemini image models through the native
API, plus Imagen on Vertex AI. Image edits are served by
OpenAI and Azure OpenAI, by
Gemini image models (gemini-2.5-flash-image, …), plus any
OpenAI-compatible provider that accepts OpenAI’s multipart upload on
/images/edits. A provider without the capability returns
a clear model "…" does not support image generation (or image edits) error
rather than mis-routing, and image-only models are hidden from /v1/models for
providers without image support.
Supported endpoints
Generate an image
model and prompt are required. Every other field — n, size, quality,
style, response_format, background, output_format, output_compression,
moderation, user, and any future parameter — is forwarded to the provider
unchanged, so model-specific options work without a gateway update. The
provider decides which values it accepts and returns its own error otherwise.
The exception is Gemini’s native adapter, which translates n and size to
their native equivalents — see Google Gemini.
The response is the OpenAI images envelope. data[] entries carry either a
hosted url (DALL·E, response_format: "url") or inline b64_json
(gpt-image-1 always returns base64). GoModel adds a provider field naming the
provider type that served the request; everything else is passed through,
including the usage block and echoed output parameters gpt-image-1 reports.
Edit an image
Edits upload the source image(s) and optional mask asmultipart/form-data,
exactly as OpenAI’s endpoint expects. image, prompt, and model are
required; send several source images as image[] (gpt-image-1 accepts up to
16). A mask is a PNG whose transparent pixels mark the area to change.
n, size, quality,
response_format, background, output_format, input_fidelity, user, …)
is forwarded to the provider unchanged. The uploaded files count toward the
gateway’s request body limit (BODY_SIZE_LIMIT).
Cost tracking
Image calls are recorded in usage tracking under the/v1/images/generations and /v1/images/edits endpoints:
The prompt is always priced at input_per_mtok when the provider reports input
tokens. The generated images are billed one of two ways, never both — the
reported usage decides:
- Token-billed models (
gpt-image-1, Gemini image models) reportusage. GoModel stores the token counts and prices the generated tokens with the model’soutput_image_per_mtokrate (falling back tooutput_per_mtok). The image output rate is the one providers charge here:gpt-image-1has no text output price at all, and Gemini 3 Pro Image bills image output at 12/Mtok for text. Pricing this way is exact for every size and quality. - Per-image models (DALL·E,
grok-imagine-image, Imagen) report no tokens. GoModel records the number of returned images (imagesin the raw usage data) and prices it with the model’sper_imagerate.
per_image
is the flat equivalent of one typical image’s tokens ($0.039 for a 1290-token
gemini-2.5-flash-image). A response that reported generated tokens is priced
by those tokens alone, so the two are never added together.
When the pricing cannot cost the generated images — no rate for the image output
tokens the response reported, or no reported output tokens and no
per_image/per_request rate — the usage row is flagged with a
cost-calculation caveat so it reads as “unreported usage” rather than a free
call.
Set per_image through a pricing override
or in config.yaml when the model catalog has no price for an image model:
Limitations
The image endpoints are a thin, model-routed pass to the provider and do not run through the full inference orchestrator. Compared with/v1/chat/completions:
- No failover, guardrails, or response cache — these stages are skipped. Requests are still authorized, budget-checked, rate-limited, metered, and written to the audit log.
- No streaming —
stream: trueis rejected with a400because streamed image generation is delivered as server-sent events, which these endpoints do not relay. Omitstream(or set it tofalse) to receive the complete JSON response. - Variations (
/v1/images/variations) are not exposed. Use the passthrough API (/p/{provider}/v1/images/...) to reach them on a specific provider. - OpenAI request shape in — Google Gemini’s native image APIs (Imagen
predict, Gemini image modelgenerateContent) are translated behind these endpoints; see provider notes below. Other providers whose native image API differs from OpenAI’s are not translated; use passthrough for those.
Google Gemini
With the Gemini provider in native API mode (the default), Gemini image models (gemini-2.5-flash-image, …) generate through generateContent,
and imagen-* models through Imagen’s predict API. Google retired Imagen
from the Gemini API (AI Studio) on August 17, 2026 — use a Gemini image model
there; Imagen remains available on Vertex AI.
- Gemini image models do not support multi-candidate output, so
n(up to 10) is served asnparallelgenerateContentcalls whose results are merged — each call is billed by Google. For Imagen,nmaps tosampleCount;sizemaps to the closest supported aspect ratio (1024x1024→1:1,1536x1024→3:2, …), and a raw ratio such as"16:9"passes through. Unknown JSON fields are forwarded verbatim, so native parameters (personGeneration,sampleImageSize,imageConfig, …) work unchanged. - Images always return as
b64_json. Any text a Gemini image model produces alongside the image is surfaced asrevised_prompt. - Gemini image models report token
usage(token-billed); Imagen reports none (price itper_image). - Edits work with Gemini image models only: uploads become inline image
parts ahead of the prompt.
maskis not supported — describe the region to change in the prompt. Imagen models only generate.
/openai/images/generations endpoint and edits are rejected (that surface has
no edits endpoint).
Audit logging
Image requests appear in the audit log like any other model interaction. WhenLOGGING_LOG_BODIES is enabled:
- The generation request is stored as JSON. An edit request is stored as an image body: the prompt and parameters plus one item per uploaded source image and mask (filename, content type, size).
- The response is stored as an image body too — the envelope (
created,usage,size,quality, …) plus one item per returned image. Hosted URLs are kept as links.
LOGGING_LOG_IMAGE_BODIES=true. With
it on, uploads and generated images are written to
media storage and each item carries a
media_id the dashboard renders inline through GET /admin/media/{id};
LOGGING_LOG_IMAGE_BODIES_SCOPE limits this to input (edit uploads) or
output (results). Without it, each image is a sized placeholder
(stored: false). Either way the entry stays small and complete — it never
hits the generic 1 MB body truncation even for large b64_json results, and
there is no per-entry cap on stored images. Files are kept as long as the
audit entries that reference them (LOGGING_RETENTION_DAYS).