用 Bifrost AI 网关追踪并限制 Codex CLI 的 token 花费
Original title: Codex Token Usage: Track and Cap Spend with an AI Gateway
The title and summary in the selected language are awaiting translation.
Codex CLI 自带的 /status 和 /usage 只能看到单个开发者的 token 用量,无法按团队归因花费或在预算耗尽时拦截请求。
TL;DR
- Codex token usage grows with context size, agent turns, reasoning effort, and model choice, so similar tasks can consume very different amounts of tokens.
- Codex CLI's
/statusand/usagecommands show token activity for one developer; neither attributes spend across a team or blocks a request when a budget runs out. - Bifrost enforces budgets at the provider config, virtual key, team, and customer levels, and rejects a Codex CLI request when any applicable budget is exhausted.
- Bifrost adds 11 microseconds of overhead per request at 5,000 requests per second, so budget checks do not slow interactive Codex CLI sessions.
- Connecting Codex CLI takes a
/logout, an exported virtual key, and a namedmodel_providersentry in~/.codex/config.toml.
Codex token usage is the count of input, cached input, output, and reasoning tokens that OpenAI's Codex CLI consumes while it reads a repository, calls tools, and edits files, and it compounds quickly once a whole engineering team runs the agent daily. Bifrost, the open-source AI gateway built in Go by Maxim AI, sits between Codex CLI and the model provider to attribute every token to a developer, enforce budgets before spend happens, and log cost per request. This guide covers how to check Codex token usage, how Codex cost and usage limits work, and how to configure Bifrost so Codex CLI spend stays predictable.
What Drives Codex Token Usage
Four factors drive Codex token spend: the context the agent reads, the number of agent turns, reasoning effort, and the model's per-token rate. OpenAI's Codex pricing page states that model choice, context, reasoning, tool use, retrieval, and caching all affect usage, so prompt length alone is not a reliable estimate.
- Repository context: files, chat history, and tool results are sent as input tokens, far more than the prompt the developer typed.
- Agentic turns: plan, edit, test, and fix steps are separate requests that each replay the growing conversation.
- Reasoning effort: higher effort produces more reasoning tokens, billed as output.
- Model rate: on OpenAI's current Codex credit rate card, GPT-6 Astra costs 250 credits per million input tokens and 1,250 per million output tokens, while GPT-6 Luna costs 2.5 and 12.5. That is a 100x spread for the same token count.
Model choice is the factor a team controls most directly, which is why a shared control point matters. The pillar guide to AI gateway architecture explains that pattern for any set of clients and models.
How to Check Codex Usage
Codex CLI exposes token usage through two slash commands. /status displays session configuration and token usage, including the active model and remaining context capacity. /usage shows daily, weekly, or cumulative ChatGPT token activity. Both commands report on one developer's session or account and cannot stop a request.
OpenAI documents both on the Codex CLI slash commands page, and ChatGPT-plan users also get a usage dashboard with reset times. Codex CLI can export OpenTelemetry metrics, including a turn.token_usage metric split by input, cached input, output, and reasoning tokens; the export is off by default and enabled with an [otel] block.
| Method | What it shows | Attributes spend across a team | Blocks requests at a limit |
|---|---|---|---|
/status in Codex CLI |
Session model, config, and token usage | No | No |
/usage in Codex CLI |
Daily, weekly, or cumulative ChatGPT token activity | No | No |
| ChatGPT usage dashboard | Plan limits and reset times for one account | No | No (plan limits apply) |
| Codex OpenTelemetry export | Per-turn token counts by type | Per machine, if every machine is configured | No |
| Bifrost request logs | Tokens, cost, latency, model, and virtual key per request | Yes | Yes, through budgets and rate limits |
The difference in the last row is position. Bifrost sits on the request path, so its request logs count Codex token usage for every developer in one place, and its budgets reject the next request when a limit is reached. Client-side methods only report after the fact.
Codex Cost and Usage Limits
Codex cost depends on how a developer signs in. With a ChatGPT plan, Codex usage draws on the plan's included allowance and then on credits. With an API key, Codex usage is billed at standard API pricing for the model the key can access, with no plan allowance and no five-hour window.
OpenAI's Codex plans and rate card list the current options:
| Access option | Price (US) | How Codex usage is limited |
|---|---|---|
| Free | $0/month | Basic access on GPT-6 Luna |
| Go | $8/month | Lightweight coding tasks |
| Plus | $20/month | Local messages per five-hour period, weekly limits may also apply, then credits |
| Pro | From $100/month ($100, $200, or $500 tiers) | No five-hour limit currently; higher allowance, then credits |
| Business | $20/user/month billed annually, $25 billed monthly | Five-hour allowance on Standard Business, workspace credits |
| Enterprise / Edu | Contact sales | Flexible pricing plans have no fixed rate limits |
| API key | Standard API pricing | Billed per token; models follow what the key can access; no cloud features such as GitHub code review or Slack |
Two details matter for Codex usage limits planning. Codex credit billing has no separate cache-write charge, and GPT-5.5 retires from Codex on all ChatGPT plans on October 14, 2026, so budgets built on GPT-5.5 rates need updating.
For Codex CLI pricing in API-key mode, which is how teams route Codex through a gateway, the bill is every request's tokens at the provider's rate. The OpenAI model cost calculator gives a per-million-token baseline before rollout.
Why Codex Token Usage Needs Team-Level Controls
Codex token usage needs team-level controls because every built-in view is scoped to one developer, while the bill is scoped to the organization. Without a shared control point, a platform team cannot see which developer or model drives Codex cost, or stop a runaway agent loop before the invoice arrives. Teams typically try one of four approaches:
| Approach | Visibility | Enforcement | Limitation |
|---|---|---|---|
| ChatGPT seats per developer | Per-account dashboard | Plan limits only | No budget per project; limits are set by OpenAI, not the team |
| One shared OpenAI API key | Organization-level billing | Provider rate limits only | No per-developer attribution |
| Separate API keys per developer | Per-key billing | Manual key revocation | Key sprawl; no model or provider policy |
| AI gateway with virtual keys | Per request, per developer, per model | Budgets, rate limits, routing rules | Requires running a gateway |
Figure 1: Codex token usage becomes attributable and enforceable once every developer's traffic shares one gateway.
As Figure 1 shows, the Bifrost AI gateway gives each developer or CI job its own virtual key, so every Codex CLI request carries an identity the gateway can bill and limit. The same pattern applies to other coding agents; the guide to governing Claude Code and Codex CLI with one gateway covers running both under shared policy.
Setting Budgets and Rate Limits on Codex CLI with Bifrost
Bifrost enforces Codex CLI spend through a budget hierarchy: customer, team, virtual key, and provider config. When a request arrives, Bifrost checks every applicable budget independently, and any single exhausted budget blocks the request. The request's cost is then deducted from every level that applies.
Figure 2: One exhausted budget at any level stops the request, and its cost is deducted from every level that applies.
The budget and limits system maps onto a Codex rollout like this:
| Level | Typical Codex CLI use | Budget | Rate limits |
|---|---|---|---|
| Customer | A business unit or external client | Yes | No |
| Team | Platform or frontend team monthly allocation | Yes | No |
| Virtual key | One developer, service account, or CI job | Yes | Requests and tokens |
| Provider config | Separate OpenAI and Anthropic spend on one key | Yes | Requests and tokens |
Budgets reset on 1d, 1w, 1M, 1Q, or 1Y windows; with calendar_aligned: true, a monthly Codex budget resets on the first of the month in UTC. Rate limits use token_max_limit and request_max_limit with their own windows, such as one million tokens per hour on a provider config, which caps burst spend from an agent stuck in a retry loop.
A temporary budget override adds capacity for a set number of reset cycles without changing the base limit, which covers a release week. Bifrost Enterprise adds budget alerts to Slack, Microsoft Teams, PagerDuty, or a webhook when a budget crosses a threshold.
The Bifrost governance overview shows how these levels fit together, and the write-up on hierarchical spend controls with virtual keys covers budget design for larger organizations.
Cutting Codex Cost with Caching and Routing
Budgets cap Codex cost; caching and routing lower it. Bifrost reduces spend in three ways: it keeps provider prompt caches warm with session affinity, it can inject prompt-cache markers that Codex does not send, and it routes requests to lower-cost models when a provider budget nears its limit.
Session affinity and prompt caching
Codex CLI sends a session-id header on every request. Bifrost adopts that header for session affinity, so a session stays on the provider and API key that served it and keeps reading the same provider prompt cache, with nothing to configure. On the Codex credit rate card, cached input for GPT-6 Astra costs 25 credits per million tokens against 250 for fresh input.
Codex also sends no prompt-cache markers, so on Anthropic models nothing is cached and every turn pays full input price. Bifrost's auto prompt caching adds the marker on the first cacheable block for requests that arrive without one, so turn one writes the cache and later turns read it. Auto injection is off by default and is enabled per provider.
Semantic caching, with its limits
Bifrost semantic caching replays a stored response for an identical or similar request, so the provider is never called, and it covers the Responses API that Codex CLI uses. Caching engages only when a request carries an x-bf-cache-key header or a default_cache_key is set, and conversation_history_threshold (default 3) skips conversations with more messages than that. Long Codex sessions usually exceed it, so semantic caching suits short, repeated requests such as scripted CI prompts. The guide to semantic caching for cutting token spend covers tuning.
Cost-aware routing rules
Routing rules evaluate a CEL expression on each request. The budget_used variable reports the percentage consumed of the budget configured for the request's provider and model, including a provider config on the virtual key. A rule such as budget_used > 85, scoped to a team or virtual key, redirects Codex traffic to a lower-cost model before the hard cap is reached.
Figure 3: Shifting to a lower-cost model near the threshold keeps developers working instead of hitting a hard stop.
Routing to a different model family works because Bifrost translates Codex CLI's Responses API calls to 25+ providers and 10,000+ models in the supported providers matrix. The guide on how to route Codex CLI to any model covers provider choice and tool-calling requirements in detail.
Codex sessions that call MCP tools also pay for tool definitions on every request. The analysis of how the MCP gateway cuts token costs in Claude Code and Codex CLI quantifies that overhead and the effect of tool filtering per virtual key.
Codex CLI Config for Bifrost
Connecting Codex CLI to Bifrost takes five steps: start the gateway, run /logout in Codex CLI, export a Bifrost virtual key as OPENAI_API_KEY, add a named Bifrost provider to config.toml, and launch codex. The logout comes first because Codex CLI always prefers an existing OAuth session over a custom API key.
Figure 4: The logout step comes first because Codex CLI prefers an existing OAuth session over the gateway key.
Start the open-source Bifrost gateway with npx -y @maximhq/bifrost or docker run -p 8080:8080 maximhq/bifrost, as described in the gateway setup guide. Then, inside Codex CLI, run /logout, exit, and export the virtual key in the same terminal you will launch Codex from:
export OPENAI_API_KEY=<bifrost_virtual_key>
Add a named provider to ~/.codex/config.toml (or a project-level .codex/config.toml):
model = "openai/gpt-5.4"
model_provider = "bifrost"
[model_providers.bifrost]
name = "Bifrost"
base_url = "http://localhost:8080/openai/v1"
env_key = "OPENAI_API_KEY"
wire_api = "responses"
supports_websockets = false
Use a named model_providers entry rather than openai_base_url: with openai_base_url, Codex keeps its built-in OpenAI provider and sends a client-side web namespace tool that Amazon Bedrock rejects. supports_websockets = false keeps Codex CLI on HTTPS, which non-OpenAI models require. The full reference is on the Codex CLI integration page.
Models use the provider/model format, either at launch or mid-session with /model:
codex --model openai/gpt-5-codex
codex --model anthropic/claude-sonnet-4-5-20250929
codex --model gemini/gemini-2.5-pro
Non-OpenAI models must support tool calling, which Codex CLI relies on for file and terminal operations. To skip manual edits, the Bifrost CLI (npx -y @maximhq/bifrost-cli) sets OPENAI_BASE_URL to the gateway's /openai/v1 path, stores the virtual key in the OS keyring, and launches Codex CLI; the walkthrough on using Bifrost CLI with coding agents shows the flow.
Monitoring Codex Token Usage Across a Team
Bifrost logs every Codex CLI request with input and output tokens, cost, latency, provider, model, and virtual key. Logging runs asynchronously, so it adds no latency, and logs filter by model, provider, token range, or cost range.
The built-in request logs answer questions such as which developer ran the most expensive session this week. For dashboards, Bifrost exposes Prometheus metrics including bifrost_input_tokens_total, bifrost_output_tokens_total, and bifrost_cost_total, labeled by provider, model, and virtual_key_id, so a Grafana panel can chart Codex cost per developer. Traces can also go to any collector through OpenTelemetry.
Cost figures come from a model catalog that syncs provider pricing every 24 hours, and custom pricing applies negotiated rates. Bifrost benchmarks show 11 microseconds of overhead per request at 5,000 RPS with a 100% success rate, so none of this tracking slows a session.
Teams running both coding agents can compare this setup with the guide on monitoring Claude Code token usage, and the Claude Code gateway selection guide covers the Anthropic side of the rollout.
Frequently Asked Questions
How do I check Codex token usage?
Run /status in Codex CLI to see the session's model, configuration, and token usage, or /usage to see daily, weekly, or cumulative ChatGPT token activity. ChatGPT-plan users can also check the Codex usage dashboard for remaining limits. These views cover one developer. For team-wide Codex token usage by developer and model, route Codex CLI through Bifrost and read its per-request token logs or Prometheus metrics.
Is Codex pay to use?
Codex is included in ChatGPT Free, Go, Plus, Pro, Business, Edu, and Enterprise plans, with allowances that differ by plan, and usage beyond the allowance draws on credits. With an OpenAI API key, every request is billed at standard API pricing and cloud features such as GitHub code review are unavailable.
How much does Codex CLI cost?
Codex CLI cost follows the sign-in method. On a ChatGPT plan, it is the plan price plus any credits used past the included allowance; Plus is $20 per month and Pro starts at $100 per month. With an API key, Codex CLI cost is the per-token rate of the chosen model multiplied by the input, cached input, and output tokens each session consumes.
What happens when you hit Codex usage limits?
When a ChatGPT-plan user reaches a Codex usage limit during an active turn, the agent can finish that turn, subject to fair use. After that, Plus and Pro users can buy additional credits, and Business or Enterprise workspaces with flexible pricing can buy workspace credits. Limits reset on the schedule shown in the usage dashboard, including five-hour windows on Plus.
Can I set a spending limit on Codex CLI?
Yes. Routing Codex CLI through Bifrost lets a team attach budgets to each developer's virtual key, a team, a customer, and each provider config. When any applicable budget is exhausted, Bifrost rejects further requests until the budget resets on its daily, weekly, monthly, quarterly, or yearly window.
Can Codex CLI use non-OpenAI models through Bifrost?
Yes. Bifrost translates Codex CLI's Responses API requests to other providers, so codex --model anthropic/claude-sonnet-4-5-20250929 runs under the same virtual key and budget. The model must support tool calling, and Codex CLI must use HTTPS rather than WebSocket mode.
Start Managing Codex Token Usage with Bifrost
Codex token usage scales with every developer, agent turn, and model upgrade, while built-in Codex views report it per account. Bifrost, as the AI gateway layer in front of every model call, turns that usage into attributed, budgeted, and logged traffic through virtual keys, budgets, routing, and prompt-cache handling. For multi-model Codex setups, see the guide to running Codex CLI with Claude, Gemini, or local models.
Source: DEV Community · Codex · dev.to



