Skip to content

New tools and Skills

New tools, Skills, MCP servers, and open-source projects worth trying.

Latest curated items

Items 21–31 · 31 total
3/31Tue
3/24Tue
  1. Permission Protocol · AI Agent Incident Tracker88

    TeamPCP 通过被污染的 Trivy GitHub Action 在 PyPI 投毒 LiteLLM 1.82.7 和 1.82.8

    TeamPCP 通过被污染的 Trivy GitHub Action 劫持 LiteLLM 的 CI/CD 流水线,窃取 PyPI 发布凭证后发布了带后门的 LiteLLM 1.82.7 和 1.82.8。

    Awaiting translation

    Why it matters: 复盘 LiteLLM 被投毒事件的三阶段攻击链与 .pth 持久化机制,可帮助排查自身 CI/CD 与 Kubernetes 风险。

2/8Sun
  1. Martin Alderson78

    Automatically improving your CLAUDE.md file with agent session logs

    The author suggests using agent session logs to improve CLAUDE.md or AGENTS.md files: Claude Code stores sessions in ~/.claude/projects, while Codex stores them in ~/.codex/sessions—both in JSONL format but with different schemas.

    Why it matters: The author works backward from agent session logs to figure out what to improve in CLAUDE.md, and has open-sourced a CLI that cuts search time from several minutes down to seconds.

12/10Wed
  1. Jesse Vincent68

    packnplay: run a coding agent in a container with one command

    Jesse Vincent has released an open-source tool called packnplay. With a single command — `packnplay run claude --dangerously-skip-permissions` — it spins up a pre-configured throwaway container to run a coding agent.

    Why it matters: The author wraps all the tedious setup for running a coding agent in a container into one command, and also shares exactly how he handles credential conflicts with Claude Code.

12/2Tue
  1. Jesse Vincent69

    Building a front-end/back-end log bridge for the coding agent to make debugging web apps easier

    When developing web apps with the coding agent, the author often runs into client-side JavaScript bugs. If the agent can't fix them by reading the code, it fires up browser MCP for interactive debugging just to see the browser console logs—burning tokens and slowing things down.

    Why it matters: The author shares a reusable front-end/back-end log bridge approach that lets the coding agent see front-end logs without browser MCP.

11/7Fri
  1. Terminal-Bench · News62

    Terminal-Bench ships version 2.0 and an optimized Harbor evaluation package

    Terminal-Bench has released version 2.0 and the Harbor package. The former is a more rigorously validated, harder benchmark for evaluating agents; the latter is for evaluating and optimizing agents. Harbor rewrites Terminal-Bench's test harness, supports deploying containers in the cloud, provides rollout interfaces for RL and SFT, and works with any agent you can put in a container.

    Why it matters: Terminal-Bench 2.0 and Harbor are released together, so readers can see how the agent evaluation benchmark is validated and how to scale it in the cloud.

10/27Mon
  1. Jesse Vincent74

    Porting Skills and Superpowers to the OpenAI Codex CLI

    Author Jesse Vincent spent an afternoon porting Superpowers and the whole SKILL.md system to the OpenAI Codex CLI, shipping it with Superpowers 3.3.0.

    Why it matters: The author ported Claude's SKILL.md system to the Codex CLI, with tool mappings and install instructions, so you can judge whether reusing Skills across models is feasible.

10/23Thu
  1. Jesse Vincent69

    Using episodic-memory to give Claude Code cross-session memory

    The author built the episodic-memory plugin for Claude Code so it can search past session logs. By default, Claude Code deletes the .jsonl session logs under ~/.claude/projects after one month; you can extend retention via cleanupPeriodDays in ~/.claude/settings.json.

    Why it matters: The author turned Claude Code's session logs into semantically searchable episodic memory, so readers can judge for themselves how long-term context is preserved across sessions.

10/19Sun
  1. Jesse Vincent71

    The author built a custom superpowers-chrome MCP that cut startup overhead from 13678 tokens to 947.

    The author built a lightweight Chrome MCP and Skill for Claude Code called superpowers-chrome. At startup, the MCP configuration takes up only 947 tokens, while Microsoft's Playwright MCP needs 13678 tokens just to be available—about 7% of the context window.

    Why it matters: The author compares the token overhead of a self-built Chrome MCP against Playwright MCP, laying out the concrete trade-offs involved in designing tool interfaces for LLMs.

10/9Thu
  1. Jesse Vincent78

    Superpowers: How the Author Used a Coding Agent in October 2025

    Author Jesse Vincent released Superpowers, a set of Skills built on Claude Code's new plugin system. Once installed, it injects a guiding prompt through the session-start hook, prompting Claude to proactively search for and use these Skills.

    Why it matters: The author packaged his own coding-agent workflow into an installable Skill plugin, so readers can directly reuse his implementation flow from brainstorming to TDD.

5/19Mon
  1. Terminal-Bench · News62

    Terminal-Bench 发布首个终端智能体评测基准

    Terminal-Bench 发布首个版本,用于量化 AI 智能体在终端中执行复杂任务的能力,首发数据集 Terminal-Bench-Core-v0 包含 80 个手工编写并人工验证的任务,每个任务配有独立 Docker 环境、人工验证的解法与测试用例。

    Awaiting translation

    Why it matters: Terminal-Bench 给出 80 个带 Docker 环境和测试用例的终端任务,可用来横向比较不同智能体在命令行中的实际表现。