chatgpt provider calls the Codex backend that ships with a ChatGPT
subscription (Plus, Pro, Business, or Enterprise). Usage is billed against the
subscription’s quota, not an OpenAI Platform API key.
Pair it with the Codex guide to run
Codex -> GoModel -> ChatGPT subscription.
Configure
The credential is the access token from your Codex sign-in:config.yaml:
codex login first if ~/.codex/auth.json does not exist yet. GoModel
derives the ChatGPT account ID from the token itself, so nothing else is
needed.
Models
The Codex backend has no model-listing endpoint, so GoModel ships the inventory a ChatGPT subscription can call:The '<model>' model is not supported when using Codex with a ChatGPT account.
gpt-5.4 and gpt-5.4-mini leave ChatGPT-authenticated Codex on August 31,
2026; use gpt-5.6-terra and gpt-5.6-luna instead. Both stay available to
Codex sessions authenticated with an OpenAI API key, through the openai
provider.
Supported surfaces
The Codex backend serves/responses and nothing else. GoModel translates
/v1/chat/completions onto it — request, response, and streaming — so chat
clients work too. /v1/embeddings still answers 501: the backend has no
embeddings endpoint.
Chat requests are converted to Responses requests, executed upstream, and
converted back to chat completions. The returned chatcmpl- IDs are labels,
not resource handles: GoModel pins store: false and drops
previous_response_id (the backend allows neither), so no chaining is
possible and chat clients resend full history as usual.
Chat parameters with no Responses equivalent are rejected with a 400 that
names the field, before any upstream call: n other than 1, logit_bias,
stop, seed, frequency_penalty, presence_penalty, logprobs,
top_logprobs, modalities, audio, web_search_options, and
the deprecated functions / function_call pair. Zero spellings of
logprobs, top_logprobs, and the penalties (false / 0) are tolerated:
they change nothing. prediction is dropped instead: it is a speed hint and
never changes the answer.
The backend also validates against a strict parameter allowlist. GoModel
adapts requests rather than failing them, so callers keep using the standard
Responses API:
Translated chat requests go through the same allowlist:
temperature,
top_p, and max_tokens / max_completion_tokens map to valid Responses
fields and are then silently dropped upstream, exactly as for native Responses
requests.
Because the backend streams only, a non-streaming POST /v1/responses is
served by streaming upstream and returning the final response object. Clients
see a normal non-streaming response. See
Responses compatibility for how the
gateway’s Responses surface behaves across providers.
Reported cost is not real spend
Subscription usage is flat-rate, but these model IDs also exist on the OpenAI Platform, so the model catalog attaches their per-token API prices. Usage records and dashboard totals forchatgpt show a figure that corresponds to no
actual charge, and the same collision makes GET /v1/models advertise
modes: ["chat", "responses"].
The modes are cosmetic — they drive dashboard grouping, not routing. Pricing is
not: it feeds cost tracking, budgets, and cost
load balancing.
Limits
Subscription quota is separate from API credit. When it is exhausted the gateway relays a429: