Juggling:用独立 Codex 会话复核答案并记录裁决
原文标题:TITLECodex answers, a separate Codex session reviews: a verdict or a HOLD, with a local record
开源项目 Juggling(Apache-2.0,alpha)让 Codex 以 Writer 身份在只读沙箱中作答,再由一个全新的独立 Codex 会话复核,返回 PASS、CHANGES_REQUESTED 或 CANNOT_REVIEW,无法验证必需步骤时以 HOLD 加原因码停止且不重试。
当前语言的正文正在等待翻译,暂时显示原文。
Problem I kept hitting with codex exec: a run finishes, and I still want a separate review and a record of what was checked. I was the one pasting answers into a second session.
So I built Juggling (Apache-2.0, alpha). You give it a self-contained text task:
- Codex answers once as the Writer, in Codex's read-only sandbox.
- A new, separate Codex session reviews that answer.
- A completed review returns PASS, CHANGES_REQUESTED or CANNOT_REVIEW. If Juggling cannot verify a required step, it stops in HOLD with a reason code. No retries; an early HOLD can stop before the Reviewer.
- juggle show re-checks a hash-chained ledger of the run.
After installing, see the flow without an account or model call:
~/.venvs/juggling/bin/juggle run --offline "Say hello" ~/.venvs/juggling/bin/juggle show latest This uses a built-in fake provider, not AI solving the task. The release includes the recommended wheel install guide:
https://github.com/picskin-inc/juggling/releases/tag/v0.1.0a1
Honest limits: Codex-only, so independence is per session, not per model. Text tasks only: unsupported command or file-read events that Juggling detects stop the run in HOLD. Real runs require macOS, Python 3.9+, Codex CLI meeting the 0.158.0 minimum floor, normal sign-in and access to the requested gpt-6-luna model. That floor is not a guarantee for every CLI build. The request is fixed to low reasoning effort; the actual serving model is not observed.
PASS is a second opinion, not proof. Juggling is not a sandbox, and HOLD cannot recall data already sent. juggle doctor checks your local setup, not sign-in or model access. Real runs dispatch at most one Writer and one Reviewer turn through your account, can incur charges, and Juggling does not measure cost.
Question for this sub: what text task would you want a second Codex session to check? And would a reviewer from a different vendor change whether you'd use it? There is no Claude or Gemini provider in this Alpha.
Repo: https://github.com/picskin-inc/juggling (I'm the author)
Security limits: https://github.com/picskin-inc/juggling/blob/v0.1.0a1/SECURITY.md
Useful feedback: doctor check IDs/statuses, HOLD reason codes, or verdict problems with a nonsensitive task. Please don't paste raw run folders or --json output.
来源:Reddit · ClaudeCode / Codex / VibeCoding · reddit.com