Skip to content

MCP and tool calling

Using and building MCP servers, configuring tool calls, and keeping them secure.

Latest curated items

Items 1–20 · 59 total
Today10/6Tue
  1. DEV Community · MCP76

    How AI agents actually use your MCP server: five failure modes you won't see in logs

    The author instrumented a demo ticket-booking MCP server with a self-built tool called mcpspan and identified five kinds of agent invocation problems that never show up in logs: agents guessing tool names that don't exist; parameters that don't match the schema and get rejected by the SDK before the handler runs; parameter types misunderstood because of how the tool description is worded; retry loops that keep hitting the same parameter; and responses so large they eat into the context window.

    Why it matters: The author used a self-built MCP analysis tool to empirically surface five failure modes in agent calls that logs don't reveal, and these can be adapted to troubleshoot your own MCP server.

  2. DEV Community · MCP78

    Prompt injection is a data-plane problem: move the boundary from the model to the tool call.

    The author argues that prompt injection shouldn't be solved by making models smarter; instead, just as SQL injection is handled with parameterized queries, the boundary should be drawn where the agent executes actions.

    Why it matters: The author draws an analogy between prompt injection and SQL injection, arguing for moving the boundary from the model to the tool-call layer, and lays out a practical approach with a strategy layer and separate read and write phases.

  3. DEV Community · MCP82

    Skills Are Not Tools: Why I Gave My Coding Agent 11 MCP Servers and It Fell Apart

    In February I hooked up 11 MCP Servers to my coding agent. Just the tool list in an empty session ate up 34,000 tokens, with Datadog alone contributing over two hundred tools. The agent got slower and kept picking the wrong tools.

    Why it matters: I'm writing this up because 11 MCP Servers of my own blew up my context, and it showed me that tools and knowledge belong in different containers.

  4. DEV Community · MCP78

    Anonymous health checks on 78 registry MCP servers: 51.3% complete the full call sequence

    Pennyforge ran anonymous health checks on the 78 servers that responded to initialize out of 186 endpoints in the a–b slice of a public MCP registry. Only 40 of them (51.3%) made it through the full flow of initialize → tools/list → one safe tools/call.

    Why it matters: Anonymous health checks on 78 registry MCP servers, with reproducible data on tiered authentication and spec version migration.

  5. Hacker News · MCP78

    Flash-Agents: an MCP plugin that hands Claude Code's coding tasks to a DeepSeek Flash worker

    Flash-Agents is a Claude Code plugin that delegates bounded coding work—implementing slices, porting tests, reviewing diffs, mapping out a codebase—to a DeepSeek V4.1 Flash worker, while Claude keeps architecture, acceptance criteria, and final review.

    Why it matters: The author outsources Claude Code's coding tasks to a DeepSeek Flash worker and shares the sandbox, patches, and measured data, so you can judge the cost and safety boundaries for yourself.

  6. Hacker News · MCP78

    Spill:把超大的 MCP 返回结果移出上下文,存入本地 DuckDB

    Spill 是一个 Apache-2.0 开源工具,通过 Hook 拦截超过 32 KiB 的 MCP 工具返回,将其存为本地 DuckDB 表(~/.spill/spill.duckdb),智能体只拿到一个紧凑描述符,再用 SQL 查询而不是读入 5 万 token 的原始 JSON。

    Awaiting translation

    Why it matters: Spill 把超大 MCP 返回落到本地 DuckDB,让智能体用 SQL 取数,为上下文窗口紧张提供了一种可复用的思路。

10/5Mon
  1. Reddit · ClaudeCode / Codex / VibeCoding78

    Use CLAUDE.md plus a local MCP memory architecture to stop Claude Code from constantly needing corrections and burning through quota

    The author has open-sourced the Project Athena v9.9.9 kernel (MIT License), a local-first memory and governance framework built to solve the problems of Claude Code needing repeated corrections, sub-agents overstepping their bounds and modifying files, and losing state when it hits the quota limit mid-task.

    Why it matters: The author validated a CLAUDE.md slimming and local-memory approach across 1900 sessions, and provides a directory structure and verification rules you can reuse directly.

  2. Hacker News · MCP77

    我连续 37 天夜间测量 MCP 注册表:19.2% 的服务器 24 小时内改动了工具面

    作者搭建 mcp-transparency-log,按计划爬取官方 MCP 注册表中所有公开可达的服务器,记录其工具名、描述、JSON schema 和四个 annotation 提示,写入带签名树头的 append-only 日志。

    Awaiting translation

    Why it matters: 作者连续 37 天夜间爬取 MCP 官方注册表,用可复现的日志量化工具面变化,并公开了两次错误修正过程。

10/4Sun
  1. Hacker News · MCP76

    RugSnare: hash-pinning MCP tool descriptions to catch silent changes after approval

    RugSnare is a runtime integrity tool for MCP tool descriptions. It computes a normalized hash pin over each approved tool's { name, description, inputSchema }, and any silent change afterward triggers an alert and fails CI (exit 1).

    Why it matters: RugSnare hash-pins MCP tool descriptions, keeps watching for silent changes after approval, and shares measured data from 66 official server versions.

10/3Sat
  1. TonyBai80

    Pi 1.0 is out: the Agent engine behind OpenClaw now takes in MCP and ships Pi Durable with crash recovery

    On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.

    Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.

9/29Tue
  1. Simon Willison · Coding Agents78

    Live from OpenAI DevDay 2026: Dots personal agent, GPT-6.1 Sol, and Ultrafast announced

    Simon Willison live-blogged the OpenAI DevDay 2026 keynote from Fort Mason in San Francisco, where OpenAI announced the personal agent Dots, ChatGPT Space, GPT-6.1 Sol, Ultrafast, and more.

    Why it matters: A running, item-by-item record of what OpenAI announced at DevDay, for a quick look at what Dots, GPT-6.1 Sol, Ultrafast, and Codex Security actually look like.

  2. Tproger · Программирование88

    OpenAI 在 DevDay 发布 GPT-6.1 Sol,价格仅为 Astra 的五分之一

    OpenAI 在 9 月 29 日旧金山 DevDay 上发布 GPT-6.1 Sol,API 名为 gpt-6.1-sol,定价为每百万输入 token 2 美元、输出 10 美元,缓存输入 0.10 美元,标准价格是 GPT-6 Astra 的五分之一。

    Awaiting translation

    Why it matters: OpenAI DevDay 发布 GPT-6.1 Sol,价格降至 Astra 的五分之一,并同步更新 Codex、Agents API 与插件体系,可据此判断成本与工具链变化。

9/28Mon
  1. GitHubDaily78

    Paseo: one interface to manage Claude Code, Codex, and other AI agents

    The open-source project Paseo pulls command-line agents like Claude Code, Codex, OpenCode, and Pi into a single management interface. All agents still run locally on your own machine, and the project has already reached 18700+ stars.

    Why it matters: The original article shows how to bring multiple command-line agents into one interface and hand off tasks across agents, making it a useful reference for developers running several agents at once.

9/24Thu
  1. Hacker News · Prompt Injection78

    Can Open-Source Prompt Injection Detectors Stop Real AI Agent Attacks? Testing 629 AgentDojo Attacks in Practice

    The author tested 10 open-source prompt injection detectors against 629 AgentDojo injection attacks—each buried in real tool output—plus 97 benign samples.

    Why it matters: The author tested 10 open-source detectors against 629 real injection attacks, with full comparison data at both default thresholds and after calibration.

9/15Tue
  1. Cline · Blog62

    Cline releases the open-source desktop app Cline Desktop, aimed at open-weight models.

    Cline has released an early version of its open-source desktop app, Cline Desktop, moving the agent runtime that previously lived in the VS Code extension and CLI into a standalone workspace. It supports parallel sessions, scheduled tasks, and a Marketplace for extending tools and integrations.

    Why it matters: The official release lays out the desktop app's capabilities and open entry points, so readers can judge whether it fits their multi-agent parallel workloads.

9/9Wed
  1. Habr · Cursor87

    Turning the Cursor Agent into a Team Member: Hooks, Safe Lists, and a Telegram Bot

    A six-person mobile team used Cursor’s hooks, safe lists, and a Telegram bot to turn project conventions from prompts into runtime enforcement.

    Why it matters: The author wires Cursor’s hooks, safe lists, and a Telegram bot into a reusable team collaboration setup, so readers can judge which constraints belong in runtime enforcement rather than in prompts.

9/7Mon
  1. Vibe Code Textbook · Articles80

    编码智能体安全:提示词注入、MCP 服务器与配置中的密钥

    文章梳理了攻击者进入编码智能体工具的三条路径,即工具返回的文本、接入的 MCP 服务器和配置中的密钥,并对照 Claude Code、Codex CLI、Gemini CLI 文档在 2026-09-07 各自承诺的控制措施。

    Awaiting translation

    Why it matters: 文章梳理了编码智能体三条攻击路径,并给出一个只读配置的 Python 审计脚本,可直接用于 CI 检查。

9/5Sat
  1. Ryan Lopopolo66

    An agent platform built for inventing agents: decoupling capability interfaces from their implementations

    Author Ryan Lopopolo argues that an agent is a parameterized program built on top of a set of capabilities: models and configurations, reasoning and tool-call loops, computers, disks, context, Skills, tools, connectors, runtimes, network policies, identity, IAM, guardrails, I/O channels, and system prompts.

    Why it matters: Drawing on his experience building multiple agents, the author proposes a platform architecture that decouples capability interfaces from their implementations — a useful reference for teams building Agent platforms.

9/2Wed
  1. Cursor · Changelog66

    Cursor launches self-hosted machines, keeping tool execution within your own network

    Cursor supports self-hosted machines: code repositories, build artifacts, and secrets all stay on internal machines within your own infrastructure, and the agent handles tool calls locally. My Machines connects a single laptop or VM to a personal workflow, while Team Pools are named worker queues for teams or enterprises—scaling capacity up with requests and down when workers disconnect. Pools aren't tied to code repositories, and idle machines can sleep and then resume within a reconnection window.

    Why it matters: The official docs lay out pooled scheduling and sandbox integration for self-hosted machines, so readers can judge whether tool execution can stay within their own network.

8/31Mon
  1. Hacker News · Claude Code 高分82

    How a single website summary request hijacked Claude Code Opus 5 Auto Mode and achieved code execution

    The author used a targeted prompt injection attack chain to reach a 60-80% attack success rate in Claude Code Opus 5 Auto Mode (small sample), whereas a third-party evaluation commissioned by Anthropic had reported a 0.00% injection success rate.

    Why it matters: The author used a module-obscuring attack chain to reach a 60-80% success rate in Auto Mode, showing that the classifier is not a sandbox.