Inject custom models into Codex Desktop / Codex CLI — side by side with the native ones.
A single-file, ~300-line localhost proxy that routes by model name: your injected models (any OpenAI-compatible gateway) go to their upstream, everything else passes through to the ChatGPT backend untouched — so your ChatGPT login, account card and usage limits stay exactly like stock.
Codex Desktop ──► 127.0.0.1:8789 ──┬─ model in routes.json ─► your gateway (its API key)
└─ anything else ────────► chatgpt.com/backend-api/codex
(your OAuth headers, unmodified)
| Approach | What goes wrong |
|---|---|
Custom model_providers.* in config.toml |
Replaces the native provider — account card / usage limits disappear, native models gone |
| CC Switch / takeover proxies | Swap the whole catalog and route all traffic to one provider |
| opencode/opencodex forks | Heavy toolchain, and "native models" there are silently aliased to another subscription |
codex-inject-proxy does injection, not replacement: the picker shows native OpenAI models and your own, the account widget keeps showing your plan and usage.
Lightweight by design: one Python file, one dependency (zstandard), no
Docker, no Node, no database, no background services. Config is one JSON file
that is hot-reloaded on change.
- Python 3.9+
pip install zstandard(Codex compresses request bodies with zstd — we must unpack them to read the model name)- Codex Desktop or Codex CLI signed in with a ChatGPT account
git clone https://github.com/funnybones69/codex-inject-proxy.git
cd codex-inject-proxy
python -m venv .venv
# Windows: .venv\Scripts\pip install -r requirements.txt
# Linux/macOS: .venv/bin/pip install -r requirements.txt
cp routes.example.json routes.json # then edit it{
"_settings": { "listen_port": 8789 },
"zai/glm-5.3-flash": {
"base_url": "https://api-gateway.merge.dev/v1",
"api_key": "mg_...",
"catalog": {
"display_name": "GLM-5.3-Flash",
"description": "Z.AI GLM 5.3 Flash via merge.dev",
"context_window": 1000000,
"max_context_window": 1000000,
"default_reasoning_level": "high"
}
}
}- Key = exact model name Codex will send.
- Any number of models / providers — each with its own
base_url+api_key. - The optional
catalogblock feeds the model-picker metadata generator. - Hot-reload: edit
routes.jsonwhile the proxy runs — changes apply on the next request.
model = "gpt-5.6-sol"
openai_base_url = "http://127.0.0.1:8789/v1"
model_catalog_json = 'C:\Users\<you>\.codex\inject-model-catalog.json'Do not set model_provider and do not add [model_providers.*] —
that replaces the native provider instead of injecting next to it.
python scripts/build_model_catalog.pyThis merges the native entries from ~/.codex/models_cache.json with one entry
per model in routes.json, and sets prefer_websockets: false for all models
(the proxy speaks plain HTTPS; this skips Codex's WebSocket reconnect retries).
Restart Codex Desktop afterwards.
Windows (registry Run key):
reg add "HKCU\Software\Microsoft\Windows\CurrentVersion\Run" /v CodexInjectProxy /t REG_SZ /d "\"C:\path\to\codex-inject-proxy\.venv\Scripts\pythonw.exe\" \"C:\path\to\codex-inject-proxy\proxy.py\"" /f(pythonw.exe = no console window; paths are resolved relative to the script,
so the working directory doesn't matter.)
Linux (systemd user unit):
# ~/.config/systemd/user/codex-inject-proxy.service
[Service]
ExecStart=/path/to/codex-inject-proxy/.venv/bin/python /path/to/codex-inject-proxy/proxy.py
Restart=on-failure
[Install]
WantedBy=default.targetsystemctl --user enable --now codex-inject-proxymacOS (launchd): same idea via a LaunchAgents plist with KeepAlive.
Every request is logged to proxy.log next to proxy.py:
POST /v1/responses model=gpt-5.6-sol -> openai-passthrough status=200
POST /v1/responses model=zai/glm-5.3-flash -> inject:api-gateway.merge.dev status=200
GET http://127.0.0.1:8789/health returns {"ok":true}.
- Responses API only. Routing decision is made for
POST /v1/responses. Any other path is passed through to the ChatGPT backend. - Tool use depends on the upstream gateway. Codex's built-in web search is a hosted OpenAI backend tool — third-party gateways can't execute it. Many gateways also fail to map the model's tool calls into structured Responses-API items (the model then "narrates" fake
<tool_call>text instead). Injected models are best used for reasoning/coding chats, not agentic tool runs, unless your gateway handles tools properly. - No WebSocket transport. Codex 0.150+ prefers
wssfor responses; the proxy answers upgrades with an instant local405, and Codex falls back to HTTPS. Setprefer_websockets: falsein the catalog (the builder script does it) to skip the retry dance entirely. - Usage accounting of injected models is not shown in Codex's account widget (it only tracks the ChatGPT plan). Check your gateway's dashboard /
proxy.log. - No load balancing / retries / auth rotation — one model = one upstream + one key, by design.
- Local trust boundary. The proxy binds to
127.0.0.1only and forwards your OpenAI token exclusively tochatgpt.com; injected routes receive their own key, never your OpenAI token. Don't expose the port. - Not affiliated with or endorsed by OpenAI. Use within the terms of your OpenAI account and your gateway providers.