AI 智能体实际会怎样使用你的 MCP server:五类日志看不到的失败模式
原文标题:What AI agents actually do with your MCP server
作者用自建的 mcpspan 对演示订票 MCP server 做埋点分析,总结出五类日志里看不到的智能体调用问题:智能体猜测不存在的工具名、参数不符合 schema 被 SDK 在 handler 前拒绝、工具描述措辞导致参数类型理解错误、同一参数反复重试的循环,以及响应体积过大挤占上下文窗口。
作者用自建 MCP 分析工具实测出五类日志里看不到的智能体调用失败模式,可迁移到自己的 MCP server 排查。
当前语言的正文正在等待翻译,暂时显示原文。
An MCP server can look healthy in every log and still be hard for agents to use. Requests come in, responses go out, nothing crashes. What you don't see is how the agent got there: what it asked for that you don't have, what it got wrong before it got it right, and where it got stuck.
These are the five things that only showed up once I measured them. The screenshots are from a demo flight-booking server with simulated traffic, so the numbers are examples; the patterns are the ones agents produce.
1. Agents call tools you don't have
Agents guess tool names. Some guesses are near misses: searchFlights when the tool is search_flights, seatMap instead of seat_map. Others are a tool the agent assumed should exist, like get_flight_price.
Three things make this list useful rather than just a count of errors:
-
The closest existing name.
searchFlightsis not a missing feature, it's a naming convention one client prefers. That's a description fix, not new code. -
What the agent did next. After
get_flight_pricemost agents went on tosearch_flights, so they worked it out. Afterexport_itineraryevery session stopped: whatever the user wanted, they didn't get it. -
Which client asked.
export_itineraryonly ever came from Cursor. A tool one client expects is a different signal from one every client expects.
2. Bad arguments are the most common failure, and they hide
When an agent sends arguments that don't match the input schema, the MCP SDK rejects them before your handler runs and sends the validation error back to the model as a normal result. Your handler never ran, so your code never saw a failure, and your error rate looks better than what agents experienced.
On this server, most of book_flight's failures were refused arguments, not crashes. Which leads to the next point.
3. The tool description is code
book_flight took passengers, described as "the passengers". Agents read that as a count and sent 2. The schema wanted a list of names, so the SDK refused the call, the agent read the error and tried again with a list.
The fix was one sentence in the description. The dotted line on the chart is where it changed:
Errors drop right at the line, while the code didn't change at all. Rewording a description can change how agents use a tool more than a change to its logic, so it helps to see exactly when a description changed next to the numbers. The dashed lines are server releases, which is the other thing you'd want to rule out.
4. Agents loop, and every single call looks fine
An agent tries check_in, gets "Check-in opens 24 hours before departure", and calls check_in again with exactly the same arguments. Then again.
Each call on its own is a normal, quick response. You only see the problem when you look at the session as a sequence, or count how many calls repeat the previous one. Usually it means the answer didn't tell the agent what to do next. Here, a better message would say when check-in opens and that retrying now won't help.
5. Response size is invisible on latency charts
A tool that usually answers with 6 kB and sometimes with 650 kB looks fine on every latency chart, because it's fast. But that answer lands in the agent's context window, crowds out everything else, and costs tokens on every following turn.
The median and 95th percentile here are fine. The largest answer is a hundred times bigger. The usual cause is a query with no filters that returns everything. A default limit, or a summary with a way to ask for more, fixes it.
How I measured this
I built mcpspan for this: self-hosted analytics for MCP servers, MIT licensed. You run it with Docker and add one line to your server:
git clone https://github.com/mcpspan/mcpspan.git
cd mcpspan
docker compose up -d
import { instrument } from 'mcpspan';
instrument(server, {
apiKey: process.env.MCPSPAN_API_KEY,
endpoint: 'http://localhost:6271',
});
import os
import mcpspan
mcpspan.instrument(mcp, api_key=os.environ.get("MCPSPAN_API_KEY"), endpoint="http://localhost:6271")
There are SDKs for TypeScript, Python, Go, C#, Java, Rust, Ruby and PHP, with the same features in each. Parameter values never leave your server's process, and nothing is sent anywhere you didn't set up yourself.
It's an early version, so if there's something you'd want to see about how agents use your server, I'd like to hear it.
来源:DEV Community · MCP · dev.to




