Codex 与 Claude Code 同任务实测:质量打平,Codex 成本低 2.4 倍
原文标题:Codex vs Claude Code: Same Tasks, Measured Costs
作者在 python-humanize 仓库上用四个任务(真实 issue 修 bug、按规格加功能、无行为变更重构、带 3 个植入 bug 的代码评审)各跑两遍,Claude Code 2.1.291(claude-opus-5-5)与 Codex CLI 0.160.0(gpt-6.1-sol)16 次运行全部通过检查,四个评审都找齐 3 个植入 bug 且无误报。
当前语言的正文正在等待翻译,暂时显示原文。
Most "Codex vs Claude Code" comparisons quote vendor benchmarks, or one third-party token figure that gets copied from page to page. Few of them publish the tasks, the prompts or the commit, so their numbers cannot be checked. We did the boring version instead: four real tasks on one pinned open-source repository, each run twice per CLI, headless, with the exact commands, prompts and numbers below, and the raw data to download. Measured on 2026-10-06 with Claude Code 2.1.291 (claude-opus-5-5) and Codex CLI 0.160.0 (gpt-6.1-sol), default models, default effort.
TL;DR
- Quality: a tie. All 16 runs passed their objective checks, and all four reviews caught the 3 planted bugs with no false alarm.
- Cost: Codex wins. $0.074 per passed task against $0.181 for Claude Code at API list prices, 2.4x cheaper, and cheaper on every single task.
- Speed: close. Claude Code was faster on the bug fix and the feature, the refactor was a tie, and Codex won the review because Claude Code chose to stress-test a hanging regex.
Codex vs Claude Code at a glance
| Claude Code | Codex CLI | Gemini CLI (not benchmarked) | |
|---|---|---|---|
| Maker | Anthropic | OpenAI | |
| Default model (measured) | claude-opus-5-5 | gpt-6.1-sol | – |
| Interfaces | Terminal, VS Code and JetBrains extensions, desktop app, web | Terminal, IDE extension, cloud tasks from ChatGPT | Terminal |
| Headless command |
claude -p (docs) |
codex exec (docs) |
gemini -p |
| Machine-readable output |
--output-format json or stream-json
|
--json (JSONL events) |
– |
| Permission model | Permission modes (default, acceptEdits, plan, bypassPermissions); --dangerously-skip-permissions skips every prompt |
Sandbox modes (read-only, workspace-write, danger-full-access) plus an approval policy; --dangerously-bypass-approvals-and-sandbox disables both |
– |
| Instruction file | CLAUDE.md |
AGENTS.md |
GEMINI.md |
| Entry plan | Claude Pro, $20/month | ChatGPT Plus, $20/month (Free and Go at $8 for light use) | – |
| Pay per token | API key, $4 / $20 per million input / output tokens (Opus 5.5) | API key, $2 / $10 per million input / output tokens (GPT-6.1 Sol) | – |
| License | Proprietary | Open source (Apache-2.0) | Open source (Apache-2.0) |
Prices are from the official pages on 2026-10-06: claude.com/pricing, ChatGPT plans for Codex and the OpenAI API pricing.
How we tested (and how you can rerun it)
Everything needed to check or rerun the benchmark is in the raw data archive (180 KB): the four prompts, the hidden acceptance tests, the run and check scripts, the full JSON transcript of each of the 16 runs and the computed metrics.
The repository and the four tasks
The repository is python-humanize/humanize (MIT): 1,725 lines of Python in src/, about 860 tests that run in about 7 seconds, pinned at commit 785e5dcc0d0308ad0dff3f6cc0faa7085ad0375b. Every run got a fresh git worktree of its starting commit and its own virtualenv, so no run could see another one's work.
| Task | Starting commit | Objective check |
|---|---|---|
| Bug fix from a real issue |
f971127, the parent of the fix for PR #346
|
The upstream regression test test_intword_rounding_rollover, copied in after the run, plus the full suite |
| Feature from a spec | 785e5dc |
16 hidden acceptance tests written before the runs (checked against a reference implementation), copied in after the run, plus the full suite |
| Refactor, no behavior change | 785e5dc |
tests/ untouched, old and new import paths return the same objects, docs updated, doctests and full suite green |
| Code review |
785e5dc plus one commit adding parse_size() with 3 planted bugs |
Planted bugs found, false alarms |
The review commit adds a parse_size() function, the inverse of naturalsize(), plus 12 passing tests that avoid every bug. To recreate it, append this to src/humanize/filesize.py (with import re at the top) and export it from humanize/__init__.py:
_SIZE_PATTERN = re.compile(r"^\s*(-?(?:\d+)+(?:\.\d+)?)\s*([A-Za-z]*)\s*$")
def parse_size(value: str) -> int:
"""Parse a size written by `naturalsize` back into a number of bytes."""
match = _SIZE_PATTERN.match(value)
if match is None:
msg = f"invalid size: {value!r}"
raise ValueError(msg)
number, unit = float(match.group(1)), match.group(2)
if unit in ("", "B", "Byte", "Bytes"):
return int(number)
if unit in suffixes["decimal"]:
return int(number * 1000 ** (suffixes["decimal"].index(unit) + 1))
if unit in suffixes["binary"]:
return int(number * 1024 ** suffixes["binary"].index(unit))
msg = f"unknown unit: {unit!r}"
raise ValueError(msg)
The three planted bugs:
-
Logic: the binary branch forgets the
+ 1, soparse_size("1.0 KiB")returns1. -
Edge case:
int()truncates the float product, soparse_size("8.2 MB")returns8199999. -
Security: the nested quantifier
(?:\d+)+backtracks exponentially.parse_size("1" * 24 + "!")takes 2.6 seconds on our machine and each extra digit doubles it: a denial of service on untrusted input.
The exact commands and prompts
Setup of one run:
git clone https://github.com/python-humanize/humanize && cd humanize
git worktree add --detach ../run-1 785e5dcc0d0308ad0dff3f6cc0faa7085ad0375b
cd ../run-1
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python -e '.[tests]' pytest
The two commands, same prompt text for both, default model and effort, permissionless:
# Claude Code 2.1.291
env -i HOME="$HOME" PATH="$PATH" IS_SANDBOX=1 \
claude -p "$PROMPT" --output-format stream-json --verbose \
--dangerously-skip-permissions \
--setting-sources project,local --strict-mcp-config --no-session-persistence < /dev/null
# Codex CLI 0.160.0
env -i HOME="$HOME" PATH="$PATH" \
codex exec --json --dangerously-bypass-approvals-and-sandbox \
--ignore-user-config --skip-git-repo-check -C "$WORKTREE" "$PROMPT" < /dev/null
A few flags deserve a word:
-
env -i,--setting-sources project,local,--strict-mcp-configand--ignore-user-configkeep our personal settings, plugins, MCP servers and model overrides out of the runs, so both CLIs run with their defaults. -
stream-jsongives the full transcript; its last event is the sameresultobject as--output-format json(usage,total_cost_usd,duration_ms, turns). - Our VM runs as root, and Claude Code refuses
--dangerously-skip-permissionsas root unlessIS_SANDBOX=1says the machine is disposable. That refusal is a sensible default. - These are the same headless invocations a workflow runner types for an unattended step (
claude -p,codex exec, permission prompts off), so the numbers describe CI and scheduled use, not interactive chat.
Every prompt ended with the same line: "The project virtualenv is ready in .venv: run the tests with .venv/bin/python -m pytest -q. Work only inside this repository. Do not commit." The four prompts, in full:
[Bug fix]
Fix this bug, reported against this repository (humanize):
`intword()` decides whether rounding pushed a value up into the next magnitude with:
if not largest_ordinal and rounded_value * power == powers[ordinal + 1]:
For values above ~10**22, `rounded_value * power` is evaluated in floating point and no longer
equals the exact `powers[ordinal + 1]`, so the carry is skipped and the number is rendered against
the lower magnitude:
>>> import humanize
>>> humanize.intword(10**24 - 1)
'1000.0 sextillion' # expected '1.0 septillion'
>>> humanize.intword(10**27 - 1)
'1000.0 septillion' # expected '1.0 octillion'
The same happens at 10**30 and 10**33.
Add a regression test.
[Feature]
Add two keyword-only parameters to `humanize.natural_list` (src/humanize/lists.py):
- `conjunction: str = "and"`: the word placed before the last item.
`natural_list(["a", "b", "c"], conjunction="or")` returns `"a, b or c"`.
- `max_items: int | None = None`: when set and the iterable has more than `max_items` items, keep
the first `max_items` items and replace the rest with "N more", joined with the conjunction.
`natural_list(["a", "b", "c", "d"], max_items=2)` returns `"a, b and 2 more"`;
`natural_list(["a", "b", "c", "d"], max_items=1)` returns `"a and 3 more"`. When the iterable has
`max_items` items or fewer, the output is the same as without the parameter.
- `max_items` lower than 1 raises `ValueError`.
- The default behavior must not change. Update the docstring (with examples) and add tests.
[Refactor]
Refactor, with no behavior change: move `scientific()` and `metric()` out of
src/humanize/number.py into a new module src/humanize/notation.py, together with the helpers and
constants only they use.
- `humanize.scientific`, `humanize.metric`, `humanize.number.scientific` and
`humanize.number.metric` must keep working.
- Helpers still needed by number.py must not be duplicated.
- The new module must appear in the API documentation (docs/ and mkdocs.yml) like the other modules.
- Do not modify the existing tests; the whole test suite must pass.
[Code review]
You are reviewing a pull request. The change under review is the last commit of this repository
(see `git show HEAD`).
Review it for correctness, edge cases, security and performance. Report each problem as a numbered
list: file and line, severity (high, medium or low), what is wrong, and a concrete input that
demonstrates it. Only list real problems. Do not modify any file.
The review prompt does not end with the test line, since the reviewer must not change anything.
What we measured, and what "cost" means here
- Pass or fail on the checks above, run by a script after each run, never by the agent itself.
- Wall time from launch to exit. The 16 runs went in two waves of 8 parallel runs on a 4 vCPU VM, which can add a few seconds to test-heavy runs on both sides.
-
Tokens as reported by each CLI: Claude Code's
resultevent (uncached input, cache writes, cache reads, output) and Codex'sturn.completedevents (input, cached input, output; reasoning tokens are reported as part of output, 16 to 225 per run). -
Lines changed, from
git diff --shortstatbefore the hidden tests were copied in. -
API-equivalent cost, computed from the official list prices on 2026-10-06. Claude Opus 5.5: $4 per million input tokens, $0.20 cache reads, $20 output, cache writes at $5 (5-minute) or $8 (1-hour, twice the input price). Claude Code wrote its cache with the 1-hour lifetime in every run, and our figure matched the CLI's own
total_cost_usdto the cent. GPT-6.1 Sol: $2 per million input tokens, $0.10 cached input, $10 output (cache writes are billed at $2.50, but Codex reported none).
"API-equivalent" matters. With an API key you pay exactly these amounts. On a Claude Pro or ChatGPT Plus subscription you pay a flat fee, and the same tokens are consumed from your usage limits instead.
Results
| Task | Passed (Claude Code / Codex) | Wall time, median | Cost per run, median | Codex cheaper by |
|---|---|---|---|---|
| Bug fix | 2/2 / 2/2 | 35.0 s / 42.4 s | $0.133 / $0.075 | 1.8x |
| Feature | 2/2 / 2/2 | 36.2 s / 53.9 s | $0.144 / $0.073 | 2.0x |
| Refactor | 2/2 / 2/2 | 59.3 s / 58.9 s | $0.290 / $0.101 | 2.9x |
| Code review (3 planted bugs) | 3/3 twice / 3/3 twice | 95.7 s / 39.2 s | $0.156 / $0.047 | 3.3x |
Bug fix from a real issue
| Claude Code (run 1, run 2) | Codex (run 1, run 2) | |
|---|---|---|
| Hidden test and suite | pass, pass | pass, pass |
| Wall time | 37.2 s, 32.8 s | 38.5 s, 46.2 s |
| Input tokens (of which from cache) | 75.0k (64.2k), 92.8k (84.4k) | 153.7k (129.0k), 159.6k (130.7k) |
| Tokens written to cache | 10.7k, 8.4k | none |
| Output tokens | 2,378, 1,808 | 892, 842 |
| Tool calls | 3, 4 | 8, 8 |
| Cost | $0.146, $0.121 | $0.071, $0.079 |
| Lines changed | +9 / -1, +7 / -1 | +12 / -1, +12 / -1 |
All four runs changed the same line and added a regression test. Claude Code compared against powers[ordinal + 1] / power both times; Codex did the same once and used the integer division // once, which is exactly the upstream fix. Codex explored more before editing (listing files, looking for an AGENTS.md, running the targeted test, then the suite, then git diff --check), which shows in its tool calls and input tokens. It still cost half as much.
Feature from a spec
| Claude Code (run 1, run 2) | Codex (run 1, run 2) | |
|---|---|---|
| Hidden tests and suite | pass, pass | pass, pass |
| Wall time | 37.0 s, 35.4 s | 50.6 s, 57.2 s |
| Input tokens (of which from cache) | 74.1k (63.6k), 75.3k (66.8k) | 111.6k (92.3k), 119.1k (93.4k) |
| Tokens written to cache | 10.5k, 8.4k | none |
| Output tokens | 2,852, 2,653 | 1,896, 1,864 |
| Tool calls | 3, 3 | 6, 6 |
| Cost | $0.153, $0.134 | $0.067, $0.079 |
| Lines changed | +72 / -5, +84 / -5 | +82 / -5, +76 / -5 |
The four implementations are nearly interchangeable: same keyword-only signature, same validation, same slicing for "N more". A precise spec leaves little room for a model to differ. Claude Code finished about 18 seconds sooner (median) by batching its reads and test runs into fewer tool calls.
Refactor without behavior change
| Claude Code (run 1, run 2) | Codex (run 1, run 2) | |
|---|---|---|
| Checks (tests untouched, API, docs, doctests, suite) | pass, pass | pass, pass |
| Wall time | 61.9 s, 56.6 s | 64.2 s, 53.5 s |
| Input tokens (of which from cache) | 251.9k (230.1k), 188.3k (168.7k) | 218.1k (194.0k), 196.1k (154.8k) |
| Tokens written to cache | 21.9k, 19.6k | none |
| Output tokens | 4,262, 4,136 | 2,087, 1,568 |
| Tool calls | 8, 6 | 8, 8 |
| Cost | $0.306, $0.273 | $0.088, $0.114 |
| Lines changed | +155 / -130, +154 / -130 | +151 / -133, +165 / -146 |
This task hides a trap: the helper _format_not_finite is used by both modules, so a naive move creates a circular import. The runs found three different ways out:
-
Claude Code (both runs):
notation.pyimports the helper fromnumber.py, andnumber.pyre-exportsmetricandscientificthrough a module-level__getattr__. It works, but the re-exported names are invisible to autocompletion and untyped for static checkers. -
Codex, run 1:
number.pyimports fromnotation.pyat the top, andnotation.pyimports the helper inside the function bodies. -
Codex, run 2: moves the helper into a new private module,
_number.py, imported by both. The cleanest structure of the four, at the price of one file nobody asked for.
All four pass every check. Claude Code paid 2.9x more here, mostly for twice the output and 20k tokens of cache writes per run.
Code review with planted bugs
| Claude Code (run 1, run 2) | Codex (run 1, run 2) | |
|---|---|---|
| Planted bugs found | 3/3, 3/3 | 3/3, 3/3 |
| False alarms | 0, 0 | 0, 0 |
| Extra valid findings | 2, 3 | 0, 0 |
| Wall time | 31.2 s, 160.2 s | 40.6 s, 37.8 s |
| Input tokens (of which from cache) | 94.2k (84.2k), 141.7k (130.8k) | 111.6k (97.5k), 110.7k (96.9k) |
| Tokens written to cache | 10.0k, 10.9k | none |
| Output tokens | 2,351, 2,749 | 999, 969 |
| Tool calls | 4, 6 | 5, 5 |
| Cost | $0.144, $0.168 | $0.048, $0.047 |
We expected this to be where one agent pulled ahead. It wasn't: every review found the ReDoS, the missing + 1 and the float truncation, each with a reproducing input, and none reported a bug that wasn't one.
The difference is in style. Codex returned exactly three findings, each probed with a short timeout. Claude Code verified every claim by running code, counted the damage (590 wrong results out of 99,999 values between 0.01 kB and 999.99 kB), and added valid remarks we had not planted: GNU-style suffixes like 2.9K that naturalsize() produces but parse_size() rejects, an OverflowError on very long inputs where the docstring promises ValueError, and localized output that cannot be parsed back. In its second run it also started a background ReDoS probe with 30,000 digits, waited, then killed it, which is why that run took 160 seconds. It also rated the binary bug "high" both times, where Codex said "high" once and "medium" once.
Our daily pull request review template runs this kind of review every morning; the numbers above say a Codex step does the planted-bug part at a third of the price.
Totals: tokens, time and cost per passed task
| Over 8 runs each | Claude Code | Codex |
|---|---|---|
| Passed | 8/8 | 8/8 |
| Total wall time | 452 s | 389 s |
| Input tokens | 993k (90% read from cache) | 1.18M (84% from cache) |
| Uncached input tokens | 90 | 192k |
| Tokens written to cache | 100k | none |
| Output tokens | 23.2k | 11.1k |
| Tool calls | 37 | 54 |
| Cost | $1.45 | $0.59 |
| Cost per passed task | $0.181 | $0.074 |
Where the money goes is the interesting part:
- Claude Code: 56% of its cost is cache writes ($0.80), 32% is output ($0.46) and 12% cache reads. It sends almost nothing uncached (90 tokens over eight runs) but writes 8k to 22k tokens per run into a 1-hour cache at $8 per million. For one-shot headless runs, that cache rarely pays for itself.
- Codex: 65% of its cost is uncached input ($0.38), 19% output and 17% cached input. It reads more and calls more tools, but every one of those tokens is cheaper, and it writes half the output.
Pricing and usage limits in practice
Official plans on 2026-10-06:
| Claude Code (Anthropic) | Codex (OpenAI) | |
|---|---|---|
| Entry plan | Pro: $20/month, $17/month billed annually | Plus: $20/month. Go ($8) and Free cover "lightweight" and "quick" tasks |
| Higher tiers | Max: from $100/month, "5x or 20x more usage per 5-hour session than Pro" | Pro: $100 to $500/month |
| How limits work | Rolling 5-hour window plus weekly caps, shared between Claude and Claude Code; no published message count | Estimated 15 to 160 local messages per 5 hours with GPT-6.1 Sol on Plus; local and cloud share the allowance; weekly limits may apply |
| API key | Yes, per token | Yes, "standard API pricing" per token |
Sources: claude.com/pricing, Claude Code with Pro or Max, Codex pricing.
Which is cheaper? Per task, Codex, by 1.8x to 3.3x on our four tasks. Projected at API prices, 100 tasks like these cost about $18 with Claude Code and $7.40 with Codex. An unattended job that runs 20 such tasks every weekday (about 440 a month) would cost about $80 against $33. That is the budget question for anything scheduled, like a weekly dependency update or a nightly triage.
Which hits limits first on a $20 plan? We can't say from our data: none of the 16 runs hit a limit, and neither vendor converts its limits into tokens. What we can say is that Claude Code consumed about twice the API-equivalent value per task. If both vendors size their $20 plans to a similar dollar value of compute, which is not published, Claude Code would reach its limit sooner.
Which one should you use?
When Claude Code is the better pick
- Reviews where you want a second opinion on the design, not only the bugs. It found 2 to 3 extra valid issues per review and backed each with a run.
- Short, well-specified tasks with a human waiting. It was 7 to 18 seconds faster (median) on the bug fix and the feature, by batching tool calls.
- Teams already on Claude Max, where the per-token price does not apply and limits are the only cost.
When Codex is the better pick
- Anything unattended and billed per token: CI, scheduled jobs, batch reviews. Same pass rate, 2.4x cheaper overall.
- Focused reviews where you want the bugs and nothing else: three findings, no false alarm, under $0.05.
- Refactors, cautiously: its two runs produced the most conventional module structure, but n=2 is not a trend.
Is there a "best AI coding agent" right now?
Not on this evidence. On four everyday tasks, the two leading CLIs produced code of the same quality, sometimes line for line. What separated them was price and verbosity, and both move with each model release. The honest answer is to measure on your own repository, with your own tasks, every time a model changes.
Using Codex and Claude Code together
"Most people use both" is in every comparison, usually followed by a handoff log that you paste from one terminal into the other. The cleaner split follows the results: route each task type to the agent that measured best for it, and let a workflow do the handoff.
In SideHub, a workflow is a list of steps run by an agent on your own machine, and each step can set its own CLI with provider (claude, codex, gemini or copilot). Here is a sketch that implements with Claude Code and reviews with Codex:
format: sidehub.workflow/v1
name: Implement, then cross-review
defaultProvider: claude
parameters:
- { name: issue, label: Issue number, type: string, required: true }
steps:
- name: Implement the issue
prompt: |
Read GitHub issue #{{inputs.issue}} with `gh issue view`, implement it on a new branch
and run the test suite. Do not push.
- name: Review the diff
provider: codex
inputsFrom: ["1"]
prompt: |
Review the uncommitted diff of this repository for correctness, edge cases, security and
performance. List each problem with file, line, severity and a reproducing input.
Each step runs as its own run, with its CLI, duration, exit code and the token usage read from the CLI's transcript, and you can follow it from anywhere, phone included. That means you can rerun this comparison on your own codebase and see the cost of each step, instead of trusting a blog post, this one included. The template gallery has ready-made workflows to start from, such as the nightly issue triage, a typical scheduled job where the cheaper agent per run adds up.
Limitations of this benchmark
- Small sample. Two runs per task per CLI. The review run that took 160 seconds shows how much a single run can move a median.
- One repository, one language, small tasks. A tidy Python library with fast tests, and tasks that take 30 seconds to 3 minutes. Long, ambiguous tasks in a large codebase may separate the agents more.
- Possible contamination. The bug fix comes from a public pull request committed on 2026-09-16. One Codex run reproduced the upstream fix character for character, which may be competence or memory.
- Our planted bugs. We wrote the review diff, so its bugs may be easier to spot than real ones.
-
Default settings only. Default models and effort, no project instruction files. Higher effort settings, other models or a tuned
CLAUDE.md/AGENTS.mdwould change the numbers. - Parallel runs. Eight runs at a time on 4 vCPUs can inflate wall times by a few seconds on both sides.
- Dated prices. API prices and plan limits are those of 2026-10-06, and Claude Code's choice of a 1-hour cache drives most of its cost. A change on either side would move the ratio.
The verdict
On four everyday tasks, Codex (gpt-6.1-sol) and Claude Code (claude-opus-5-5) were equally good: 16 passes out of 16, every planted bug found, no false alarm. Codex was 1.8x to 3.3x cheaper per task at API prices; Claude Code was a bit faster on small coding tasks and wrote richer reviews. Use Codex where cost per run adds up, Claude Code where you want depth or a quick answer, and rerun these four tasks on your own repository: the commands and prompts are all above.
FAQ
Is Codex better than Claude Code?
Not on quality, in our benchmark. Both CLIs passed all four tasks twice (bug fix, feature, refactor, code review) and both caught the three planted bugs in every review. Codex did it for 1.8 to 3.3 times less money at API prices; Claude Code gave longer reviews with extra valid findings.
Which is cheaper, Codex or Claude Code?
Codex, on every task we measured. At official API list prices on 2026-10-06, a passed task cost $0.074 on average with Codex (gpt-6.1-sol) against $0.181 with Claude Code (claude-opus-5-5). On a $20 plan you pay a flat fee instead, and the difference shows up as usage-limit consumption.
Which one uses fewer tokens for the same task?
It depends on which tokens. Claude Code read fewer input tokens (993k against 1.18M over eight runs) because almost all of its input came from its prompt cache, but it wrote twice as many output tokens (23.2k against 11.1k). Claude Code's output and cache writes are also priced higher, which is where the cost gap comes from.
Which one hits usage limits first on the $20 plans?
We cannot answer that from our data: none of our 16 runs hit a limit. OpenAI estimates 15 to 160 local Codex messages per 5 hours with GPT-6.1 Sol on Plus. Anthropic publishes no message count for Pro, only a rolling 5-hour window, weekly caps and limits shared between Claude and Claude Code.
Which is faster, Codex or Claude Code?
Close. Claude Code was faster on the bug fix (35 s against 42 s, median) and the feature (36 s against 54 s), the refactor was a tie (59 s each), and Codex won the review (39 s against 96 s) because one Claude Code run spent two minutes stress-testing a hanging regex on purpose.
Which is better for code review, and which for implementing features?
Both found every planted bug with no false alarm. Codex returned exactly the three bugs at a third of the cost; Claude Code verified each claim by running code and added valid design remarks (GNU suffixes, overflow, localized output). For implementation, both wrote nearly identical, passing code, so the choice comes down to price and speed.
Can you use Codex and Claude Code together?
Yes, and the useful split is by task type rather than by mood. Keep shared rules in AGENTS.md so both read the same conventions, then route each step to the CLI that measured best for it, for example implementation in one and review in the other. A workflow runner that sets the CLI per step removes the copy-paste between terminals.
Do Codex and Claude Code run headless in CI or on a schedule?
Yes: claude -p and codex exec both run a prompt to completion without a TTY and print machine-readable output (stream-json and JSONL). Unattended runs need their permission prompts disabled, with --dangerously-skip-permissions for Claude Code and --dangerously-bypass-approvals-and-sandbox (or a sandbox mode) for Codex, so run them in a disposable worktree, container or VM.
What is the difference between CLAUDE.md and AGENTS.md?
They are the project instruction files each agent loads at startup: CLAUDE.md for Claude Code, AGENTS.md for Codex (and several other agents). One repository can serve both: keep the rules in AGENTS.md and make CLAUDE.md a one-line import of it with Claude Code's @AGENTS.md syntax.
ChatGPT Codex vs OpenAI Codex CLI vs Claude Code: what is what?
Codex is OpenAI's coding agent, available as the open-source Codex CLI in your terminal, as an IDE extension and as cloud tasks from ChatGPT; all share your ChatGPT plan's allowance. Claude Code is Anthropic's equivalent, in the terminal, IDE extensions, a desktop app and the web. This benchmark compares the two CLIs in headless mode.
How does Gemini CLI compare?
Gemini CLI is Google's open-source terminal agent, with a headless gemini -p mode and a GEMINI.md instruction file. We did not include it in this benchmark; a dedicated Gemini CLI vs Claude Code comparison is planned.
What do developers on Reddit and Hacker News say?
The recurring themes are usage limits (complaints about hitting Claude limits quickly on Opus), Codex being terser and more literal, Claude Code being more thorough and talkative, and many people paying for both. These are impressions, not measurements, which is why we ran our own.
来源:DEV Community · Codex · dev.to