Skip to content

New tools and Skills

New tools, Skills, MCP servers, and open-source projects worth trying.

Latest curated items

Items 1–20 · 31 total
10/6Tue
  1. DEV Community · MCP76

    How AI agents actually use your MCP server: five failure modes you won't see in logs

    The author instrumented a demo ticket-booking MCP server with a self-built tool called mcpspan and identified five kinds of agent invocation problems that never show up in logs: agents guessing tool names that don't exist; parameters that don't match the schema and get rejected by the SDK before the handler runs; parameter types misunderstood because of how the tool description is worded; retry loops that keep hitting the same parameter; and responses so large they eat into the context window.

    Why it matters: The author used a self-built MCP analysis tool to empirically surface five failure modes in agent calls that logs don't reveal, and these can be adapted to troubleshoot your own MCP server.

  2. Reddit · ClaudeCode / Codex / VibeCoding76

    A local proxy spreads Claude Code requests across multiple Max accounts and switches before the quota runs out.

    The author open-sourced claudemanager, a local daemon that Claude Code points to via ANTHROPIC_BASE_URL. It only changes the request's Authorization header to route sessions to the Max account with the most remaining capacity in its 5-hour, weekly, and per-model windows, switching at custom thresholds before those windows fill up.

    Why it matters: The author also open-sourced a local proxy that automatically distributes Claude Code traffic across multiple Max accounts based on remaining quota, and logs requests along the way.

  3. Hacker News · MCP78

    Flash-Agents: an MCP plugin that hands Claude Code's coding tasks to a DeepSeek Flash worker

    Flash-Agents is a Claude Code plugin that delegates bounded coding work—implementing slices, porting tests, reviewing diffs, mapping out a codebase—to a DeepSeek V4.1 Flash worker, while Claude keeps architecture, acceptance criteria, and final review.

    Why it matters: The author outsources Claude Code's coding tasks to a DeepSeek Flash worker and shares the sandbox, patches, and measured data, so you can judge the cost and safety boundaries for yourself.

  4. Hacker News · MCP78

    Spill:把超大的 MCP 返回结果移出上下文,存入本地 DuckDB

    Spill 是一个 Apache-2.0 开源工具,通过 Hook 拦截超过 32 KiB 的 MCP 工具返回,将其存为本地 DuckDB 表(~/.spill/spill.duckdb),智能体只拿到一个紧凑描述符,再用 SQL 查询而不是读入 5 万 token 的原始 JSON。

    Awaiting translation

    Why it matters: Spill 把超大 MCP 返回落到本地 DuckDB,让智能体用 SQL 取数,为上下文窗口紧张提供了一种可复用的思路。

10/5Mon
  1. Reddit · ClaudeCode / Codex / VibeCoding78

    Use CLAUDE.md plus a local MCP memory architecture to stop Claude Code from constantly needing corrections and burning through quota

    The author has open-sourced the Project Athena v9.9.9 kernel (MIT License), a local-first memory and governance framework built to solve the problems of Claude Code needing repeated corrections, sub-agents overstepping their bounds and modifying files, and losing state when it hits the quota limit mid-task.

    Why it matters: The author validated a CLAUDE.md slimming and local-memory approach across 1900 sessions, and provides a directory structure and verification rules you can reuse directly.

  2. Hacker News · MCP77

    我连续 37 天夜间测量 MCP 注册表:19.2% 的服务器 24 小时内改动了工具面

    作者搭建 mcp-transparency-log,按计划爬取官方 MCP 注册表中所有公开可达的服务器,记录其工具名、描述、JSON schema 和四个 annotation 提示,写入带签名树头的 append-only 日志。

    Awaiting translation

    Why it matters: 作者连续 37 天夜间爬取 MCP 官方注册表,用可复现的日志量化工具面变化,并公开了两次错误修正过程。

10/4Sun
  1. Hacker News · MCP76

    RugSnare: hash-pinning MCP tool descriptions to catch silent changes after approval

    RugSnare is a runtime integrity tool for MCP tool descriptions. It computes a normalized hash pin over each approved tool's { name, description, inputSchema }, and any silent change afterward triggers an alert and fails CI (exit 1).

    Why it matters: RugSnare hash-pins MCP tool descriptions, keeps watching for silent changes after approval, and shares measured data from 66 official server versions.

10/3Sat
  1. TonyBai80

    Pi 1.0 is out: the Agent engine behind OpenClaw now takes in MCP and ships Pi Durable with crash recovery

    On October 1, the Earendil team released the Agent Harness Pi 1.0, with roughly 11.1 stars and 1.4 forks on GitHub, plus an experimental new package called Pi Durable.

    Why it matters: Pi 1.0 and Pi Durable bring distributed concepts like checkpoints, idempotent commits, and ownership trees into the Agent runtime, which you can use to weigh the engineering trade-offs of long-running Agents.

9/15Tue
  1. Lovable · Blog64

    Lovable open-sources OJ, a Rust preview engine that beats Vite on cold start and memory

    Lovable has released OJ, a preview engine written from scratch in Rust. It reads your existing vite.config.ts and runs real Vite plugins through a compatibility layer, all in a single binary, with no toolchain installed into the project.

    Why it matters: Lovable rewrote its preview engine OJ in Rust, sharing cold start and memory comparisons against Vite, plus canary data from production.

9/10Thu
  1. Vibe Code Textbook · Articles78

    编程智能体的四种提示词模式:plan mode、skills 与保存的提示词

    文章从 Claude Code、Codex 和 Gemini CLI 的官方文档中整理出四种提示词模式:先计划再编辑、给智能体一个可运行的检查、让智能体反过来访谈你、把反复重打的提示词存成文件,并给出各家对应的命令、参数和文件格式。

    Awaiting translation

    Why it matters: 横向对照 Claude Code、Codex、Gemini CLI 三家文档,给出计划模式、可运行检查、访谈式提问和保存提示词四种模式的命令与文件格式。

9/2Wed
  1. Hacker News · Agent Skills78

    mattpocock releases AI coding Agent Skills built for real engineering

    Author mattpocock has released a set of AI coding Agent Skills he uses day to day. They're aimed at real engineering rather than vibe coding, and the emphasis is on being small, easy to modify, composable, and compatible with any model.

    Why it matters: The author breaks years of engineering experience into a set of composable Skills and explains the failure mode each one targets, so readers can judge whether they fit into their own development workflow.

7/11Sat
  1. Habr · Kova13v80

    Using an evidence contract to constrain Claude Code's test conclusions: a QA Skill package

    A QA engineer distilled six months of experience doing web testing on Claude Code into an open-source Skill package called paranoid-qa. At its core is an evidence contract: Pass/Fail can only be based on actual artifacts like screenshots, request bodies, and logs; anything unverified gets marked Not tested; anything blocked by the environment gets marked Blocked; and forms must verify the real submitted payload.

    Why it matters: The author codified six months of QA experience into a testing Skill package for Claude Code, along with an evidence contract and failure checklist that can be reused directly.

7/8Wed
  1. AI Hero · Skills Updates65

    AI Hero skills repo ships v1.1: adds /wayfinder, renames /to-spec and /to-tickets

    AI Hero's skills repo ships v1.1, renaming /to-prd to /to-spec, merging /to-plan and /to-issues into /to-tickets, and adding new Skills like /wayfinder, /research, and /prototype.

    Why it matters: The author walks through the full Skill flow from grilling to deployment and gives the migration commands for the renames, the merge, and the new /wayfinder—useful for anyone building an AI development workflow.

6/25Thu
  1. Kondasamy Jayaraman · Engineering Blog78

    Ponytail:让 AI 智能体少写代码的规则集与六级决策阶梯

    Ponytail 是一个 MIT 许可的规则集,可接入 Claude Code、Codex、Cursor、Copilot、Gemini CLI 等智能体,通过六级决策阶梯让模型在写自定义代码前先判断是否该做、标准库或平台内置能否解决,从而减少生成代码量。

    Awaiting translation

    Why it matters: Ponytail 用六级决策阶梯约束 AI 智能体少写代码,并给出真实仓库上的基准与成本数据,可迁移到团队规范。

6/20Sat
5/14Thu
  1. Permission Protocol · AI Agent Incident Tracker78

    Microsoft Defender 发现 Mage AI 与 MCP 服务器未鉴权部署,可获 cluster-admin 权限执行 RCE

    Microsoft Defender for Cloud 发现生产环境中的 Mage AI 与 MCP 服务器未启用鉴权,攻击者可执行 shell 命令并获得 cluster-admin 权限,还能窃取同集群工作负载的凭据。

    Awaiting translation

    Why it matters: 材料给出 Mage AI 与 MCP 服务未鉴权部署导致 RCE 的完整链路,可据此检查自家 Helm chart 与 MCP 配置。

5/12Tue
5/11Mon
5/9Sat
  1. OpenAI · Codex Cookbook72

    OpenAI brings persistent Goals to Codex

    Starting with Codex 0.128.0, OpenAI offers Goals, turning one-off prompts into persistent objectives within a thread. Codex keeps checking evidence such as tests, benchmarks, or deliverables to decide whether the goal is done.

    Why it matters: The official docs lay out where Goals fits, how to write its six elements, and the lifecycle commands, so you can tell when a persistent objective should replace a one-off prompt.

4/30Thu
  1. AI Hero · Skills Updates62

    AI Hero 更新 Skills:/ubiquitous-language 并入 /grill-with-docs

    AI Hero 的 skills 仓库更新,把 /ubiquitous-language 废弃并合并进新 Skill /grill-with-docs,输出从 ubiquitous-language.md 改为 context.md,并支持多个限界上下文各自维护共享语言。

    Awaiting translation

    Why it matters: 作者把 /ubiquitous-language 合并为 /grill-with-docs,并给出 ADR 触发条件与多限界上下文做法,可迁移到自己的 Skill 配置。