Перейти к содержимому
Оригинал
DEV Community · MCP· Ahmet Zeybek·· 11 часов назадИзбранное редакциейОценка ИИ82

Навыки — это не инструменты: автор разбирает, почему 11 MCP Server тормозят ИИ-агента для разработки

Оригинальный заголовок: Skills are not tools

Краткий обзор ИИ

В феврале автор подключил к ИИ-агенту для разработки 11 MCP Server, и в пустом сеансе только список инструментов занимал 34,000 токенов; один Datadog добавил больше двухсот инструментов, из-за чего агент замедлился и стал часто выбирать не те инструменты.

Почему это важно

На своём опыте с 11 MCP Server, которые перегрузили контекст, автор показывает, что инструменты и знания должны лежать в разных «контейнерах».

Полный текст

In February I connected eleven MCP servers to a coding agent: GitHub, Linear, Postgres, Sentry, Datadog, Slack, a browser, a filesystem, two internal ones and Notion. For about a week it felt like having superpowers.

Then I looked at the token count on an empty session. Before I'd typed anything, the tool list alone was 34,000 tokens. Every tool has a name, a description and a JSON schema, and all of them get sent on every turn. Datadog by itself contributed over two hundred tools. The agent got slower. It started picking search_datadog_logs when I meant search_datadog_spans, and one afternoon it tried to answer a question about a database table by opening Notion, because the Notion server had a tool called search and my question had the word "table" in it.

Giving the agent more made it worse. That's the wall, and most teams I talk to hit it by April.

Two different problems

It happens because MCP solves one problem and gets used for two.

The problem MCP solves is access. On its own the model can't query your database, read your issue tracker or post to Slack. A server exposes those operations as tools with typed inputs, the client lists them, and the model calls them. It's a plug standard.1 That's a real win, and it's why MCP took over in about a year.

The problem it doesn't solve is knowledge: knowing that this team deploys with pnpm release, that the staging database is the one with -stg in the hostname, that a Linear ticket isn't done until the PR is merged and the label is set, that this repository's commit messages have a scope in brackets. None of that is a tool. It's procedure, the stuff a new engineer picks up in their first two weeks by asking people, and an agent doesn't have it unless you put it somewhere.

Teams tried putting it in tool descriptions, and the Datadog server's descriptions turned into small essays. Then they tried system prompts, which grew to thousands of lines loaded into every session whether the task had anything to do with deployment or databases or not. Both are the wrong container. Skills are the right one.

What a skill is

A skill is a folder with a markdown file in it, and that's nearly the whole spec. The file has a short front matter block with a name and a one line description, followed by instructions in plain prose: when to use it, how to do the thing, what the pitfalls are, which commands to run. The folder can also hold scripts, templates and reference documents the instructions point to.

---
name: release
description: "Cut a release of a package in this monorepo. Use when asked to release, publish, or tag a version."
---

# Releasing a package

1. Run `pnpm changeset status` and confirm there is a pending changeset
   for the package. If there is none, stop and ask, do not create one.
2. Run `pnpm release:dry` and paste the version bump it proposes.
3. On confirmation, `pnpm release`. This tags, publishes through the
   OIDC workflow, and opens the changelog PR.
4. Never publish with a local npm token. If `pnpm release` asks for
   one, the OIDC setup is broken, stop and report.

See ./references/changesets.md for how we write changeset entries.

What matters is how it's loaded. The agent doesn't read every skill on every turn. At startup it only sees the names and one line descriptions, which for forty skills comes to a few hundred tokens. When a task matches a description, it reads that one skill's full instructions into the context and uses them, and the rest stay on disk. The coding agents settled on this pattern this year, Claude Code first and then the others. It spread because it's the first mechanism that scales knowledge the way MCP scaled access.

Progressive disclosure is the whole trick

The design principle has a name, progressive disclosure, and it's why the two mechanisms don't compete.

MCP tool lists are flat. Every tool is described in full up front, because the model has to be able to call any of them at any time. That's correct for tools, and it's where the 34,000 tokens came from.

Skills come in layers. The index is small, the body loads on demand, and references inside the body only load if the body tells the agent to read them. A skill for a database migration workflow can be three lines in the index, two pages of instructions when it triggers, and a 40 page schema reference that only gets opened if the migration touches a particular table.

Seen that way, the fix for the eleven server problem is obvious. You don't need the Datadog server's two hundred tools in the context. You need a skill called investigate-incident that says "for latency questions use search_datadog_spans, for error rates use the RUM aggregate, here is how to read our service map, here are the three dashboards that matter", and you need the Datadog server connected. The skill holds the knowledge of which tool to use and how, and the server provides the ability to use it.

Note

You pay a tool's token cost on every turn. You pay a skill's once, when it triggers. That one difference decides which container a piece of information belongs in.

How I split it now

After the February mess I rebuilt the setup around one rule: a server only gets connected if a skill references it. If no procedure needs a tool, the tool doesn't need to be in the context.

That left four servers instead of eleven: GitHub, Linear, Postgres and Datadog. Slack, Notion and the browser went, because what I used them for turned out to be one off requests that no skill could improve, and that weren't worth 3,000 tokens a turn as tool calls. The filesystem server duplicated the agent's own file tools. The two internal servers became one, and that one got a skill.

Then I wrote skills for the procedures I kept explaining: releasing, investigating an alert, writing a migration, reviewing a pull request the way this team reviews them (which includes checking the Linear ticket is linked and that nothing in packages/site-config changed without a changeset), and onboarding a new MCP server, since that's a procedure too.

An empty session is now 6,000 tokens of tool schemas plus about 400 tokens of skill index. The agent picks the right tool nearly every time, because the skill that triggered told it which one. And when someone new joins and asks how we deploy, I send them to the same markdown file the agent reads. It's turned out to be the best documentation the team has ever had, for the boring reason that it's the only documentation that gets used every day.

What goes wrong

Skills fail in their own ways, and I've run into most of them.

The description is the trigger, and a vague one triggers on everything or nothing. "Helps with databases" fires on any question with the word in it. "Write and apply a Postgres migration in this repository. Use when asked to add, alter or drop a table, column or index" fires when it should. Write descriptions the way you'd name a function, specific enough that the wrong caller wouldn't reach for it.

Skills that repeat what the model already knows are dead weight. A skill explaining what a pull request is wastes every token it costs. One explaining that in this repository PRs are squash merged and the title becomes the changelog entry is worth every token. The test is whether a strong senior engineer from outside the company would need to be told.

Skills go stale exactly the way READMEs do, except a stale skill produces confident wrong actions where a stale README produces a confused human. When the release process changed in March, the agent ran the old command for a week, because the skill said to and the skill was all it had read. The fix was cultural. The skill lives in the repository, and a process change doesn't get merged until the skill is updated in the same PR.2

There's a security side to watch as well. A skill is instructions the agent will follow, so a skill from an untrusted source is a prompt injection you installed on purpose. Treat a shared skills directory like a shared shell profile: review it, pin it, and don't pull skills from a marketplace into an agent that can write to anything.

A quick way to audit your own setup

Open an empty session and check the token count before you type anything. If it's over 10,000, most of that is tool schemas, and for each connected server the question is whether any procedure you actually run uses it. Disconnect the ones where the answer is no.

Then look at your system prompt or project instructions file. Every paragraph that starts with "when doing X" is a skill waiting to be pulled out. Move it to a folder with a one line description, and the paragraph stays out of the context until X comes up. On my setup that turned a 3,000 token instructions file into a 400 token one plus eleven skills, and the agent got noticeably better at following what was left, because fewer instructions were competing for its attention.

The shape of it

MCP is the socket, and skills are the manual that sits next to the machine. A workshop with fifty sockets and no manuals is a room where things get plugged in wrong. A manual with no sockets is a nice read.

If your agent got slower and more confused as you gave it more servers, a bigger context window or a better model won't fix it. Disconnect the servers no procedure needs, write down the procedures you keep repeating, and let the agent load knowledge when the task calls for it, instead of carrying all of it, all the time, for everything.

Originally published at zeybek.dev.

  1. Before it, every agent had its own way of wiring in a function, and servers weren't portable. Now they are. ↩

  2. That's the rule for tests, and it should be the rule here too. ↩

Источник: DEV Community · MCP · dev.to