Skip to content
Original
AI Coder · Telegram·· 14 days agoSelectedAI score78

Role-Based Model Routing in Claude Code: One Week of Practice and Hook Enforcement

Original title: 🧠 Маршрутизация моделей в Claude Code: неделя на практикеНастроил для Claude Code маршрутизацию моделей по образцу своего Codex-конфига. Неделю на ней прожил — делюсь.Идея простая: сильнейшая модель прекрасно умеет писать код. Но гораздо полезнее занять её разбором сложной задачи, постановкой и приёмкой результата, а реализацию делегировать.🎯 Кто чем занимаетсяГлавный тред — координатор. Он читает код и готовит карточку задачи: цель, какие файлы можно менять, инварианты, критерии приёмки и команды проверки.Дальше раздаёт работу:🔎 explorer → HaikuТолько поиск по коду.🛠 worker → OpusРеализация через TDD.🧪 verifier → SonnetНезависимый прогон проверок по чужому diff.🧠 senior → Opus на highДеньги, данные, concurrency и сложные эскалации.👀 reviewer → другая модельСемантическое ревью. Обязательно другой моделью, чем у автора.🔒 Главная фича — координатору нельзя писать кодВ Codex-версии это было прописано инструкцией. Здесь запрет обеспечивается хуком PreToolUse.У субагента в JSON хука

AI overview

Drawing on his own Codex setup, the author built a role-based model routing system for Claude Code: the main thread acts as coordinator, explorer uses Haiku for code search only, worker uses Opus for TDD implementation, verifier uses Sonnet to run checks independently, senior uses high-tier Opus for money, data, and concurrency, and reviewer switches to a different model for semantic review.

Why it matters

Based on a week of hands-on testing, the author shares the configuration, Hook enforcement, and cost trade-offs of multi-model division of labor in Claude Code, which you can adapt to your own multi-agent workflows.

Full text · AI translation

🧠 Model Routing in Claude Code: A Week in Practice

I set up model routing for Claude Code, modeled on my Codex config. I ran it for a week—here's how it went.

The idea is simple: the strongest model is great at writing code. But it's far more useful to have it break down a hard problem, define the task, and accept the result, while delegating the implementation.

🎯 Who Does What

The main thread is the coordinator. It reads the code and prepares a task card: the goal, which files can be changed, invariants, acceptance criteria, and verification commands.

Then it hands out the work:

🔎 explorer → Haiku
Code search only.

🛠 worker → Opus
Implementation via TDD.

🧪 verifier → Sonnet
Independently runs checks against someone else's diff.

🧠 senior → Opus on high
Money, data, concurrency, and complex escalations.

👀 reviewer → a different model
Semantic review. Must be a different model than the author's.

🔒 The key feature: the coordinator can't write code

In the Codex version, this was enforced by an instruction. Here, the restriction is enforced by a PreToolUse hook.

A subagent has an agent_id in the hook's JSON; the main thread doesn't. As long as the .claude/routing.on marker is present, the hook rejects Edit/Write from the main thread and shell commands it recognizes as writes.

The usual "it's just one line, I'll fix it myself real quick" runs straight into the block.

It's usually on little things like this that role separation falls apart. For me, this is the key feature.

📌 How It Worked Out in Practice

• The verifier caught a mistake in my task definition that I wouldn't have noticed myself. You also need to check whether you framed the task correctly in the first place.

• The reviewer, on a different model, ran a differential test on 100 million inputs where the author had settled for unit tests.

• The coordinator's context doesn't balloon. Detailed logs, file contents, and Docker output stay with the subagents. Only results and verification evidence come back up.

💸 What About Cost

Total tokens went up.

Each subagent re-reads the files. The task card, the report, the acceptance step—extra passes over the same change. Context isolation isn't free.

On my tasks, the savings come from splitting the work: the strongest model handles task definition and acceptance, and part of the work goes to cheaper models.

• Compared to a single thread on the strongest model—cheaper.
• Compared to a single thread on a mid-tier model—more expensive.

What I'm really buying here is isolation and an independent look at the result.

🪤 Pitfalls From the Week

• A Bash wrapper with a heredoc ate the hook's stdin, and it silently let everything through. Check that the block actually fires.

• The 30-step limit for verifiers wasn't enough for Docker tests. In my scenario, 60 is the working minimum.

• The heuristic for write shell commands looks at words, not paths. That limitation needs to be kept in mind.

📦 Get the Config

The installer, ready-made configs, hook tests, and a description of the process are in the gist.

It installs with one command. Installation is idempotent, and rollback is the –restore flag.

Source: AI Coder · Telegram · t.me