Superpowers 作者用 CLAUDE.md 拦截智能体提交的低质 PR
Superpowers 作者 Jesse Vincent 称项目 GitHub star 已超 12 万,随之而来大量由智能体自动提交的低质 PR,仓库 PR 拒绝率达 94%。
Awaiting translation
Why it matters: Superpowers 作者用 94% PR 拒绝率说明智能体批量提 PR 的现状,并给出可复用的 CLAUDE.md 拦截写法。
New tools, Skills, MCP servers, and open-source projects worth trying.
Superpowers 作者 Jesse Vincent 称项目 GitHub star 已超 12 万,随之而来大量由智能体自动提交的低质 PR,仓库 PR 拒绝率达 94%。
Awaiting translation
Why it matters: Superpowers 作者用 94% PR 拒绝率说明智能体批量提 PR 的现状,并给出可复用的 CLAUDE.md 拦截写法。
TeamPCP 通过被污染的 Trivy GitHub Action 劫持 LiteLLM 的 CI/CD 流水线,窃取 PyPI 发布凭证后发布了带后门的 LiteLLM 1.82.7 和 1.82.8。
Awaiting translation
Why it matters: 复盘 LiteLLM 被投毒事件的三阶段攻击链与 .pth 持久化机制,可帮助排查自身 CI/CD 与 Kubernetes 风险。
The author suggests using agent session logs to improve CLAUDE.md or AGENTS.md files: Claude Code stores sessions in ~/.claude/projects, while Codex stores them in ~/.codex/sessions—both in JSONL format but with different schemas.
Why it matters: The author works backward from agent session logs to figure out what to improve in CLAUDE.md, and has open-sourced a CLI that cuts search time from several minutes down to seconds.
Jesse Vincent has released an open-source tool called packnplay. With a single command — `packnplay run claude --dangerously-skip-permissions` — it spins up a pre-configured throwaway container to run a coding agent.
Why it matters: The author wraps all the tedious setup for running a coding agent in a container into one command, and also shares exactly how he handles credential conflicts with Claude Code.
When developing web apps with the coding agent, the author often runs into client-side JavaScript bugs. If the agent can't fix them by reading the code, it fires up browser MCP for interactive debugging just to see the browser console logs—burning tokens and slowing things down.
Why it matters: The author shares a reusable front-end/back-end log bridge approach that lets the coding agent see front-end logs without browser MCP.
Terminal-Bench has released version 2.0 and the Harbor package. The former is a more rigorously validated, harder benchmark for evaluating agents; the latter is for evaluating and optimizing agents. Harbor rewrites Terminal-Bench's test harness, supports deploying containers in the cloud, provides rollout interfaces for RL and SFT, and works with any agent you can put in a container.
Why it matters: Terminal-Bench 2.0 and Harbor are released together, so readers can see how the agent evaluation benchmark is validated and how to scale it in the cloud.
Author Jesse Vincent spent an afternoon porting Superpowers and the whole SKILL.md system to the OpenAI Codex CLI, shipping it with Superpowers 3.3.0.
Why it matters: The author ported Claude's SKILL.md system to the Codex CLI, with tool mappings and install instructions, so you can judge whether reusing Skills across models is feasible.
The author built the episodic-memory plugin for Claude Code so it can search past session logs. By default, Claude Code deletes the .jsonl session logs under ~/.claude/projects after one month; you can extend retention via cleanupPeriodDays in ~/.claude/settings.json.
Why it matters: The author turned Claude Code's session logs into semantically searchable episodic memory, so readers can judge for themselves how long-term context is preserved across sessions.
The author built a lightweight Chrome MCP and Skill for Claude Code called superpowers-chrome. At startup, the MCP configuration takes up only 947 tokens, while Microsoft's Playwright MCP needs 13678 tokens just to be available—about 7% of the context window.
Why it matters: The author compares the token overhead of a self-built Chrome MCP against Playwright MCP, laying out the concrete trade-offs involved in designing tool interfaces for LLMs.
Author Jesse Vincent released Superpowers, a set of Skills built on Claude Code's new plugin system. Once installed, it injects a guiding prompt through the session-start hook, prompting Claude to proactively search for and use these Skills.
Why it matters: The author packaged his own coding-agent workflow into an installable Skill plugin, so readers can directly reuse his implementation flow from brainstorming to TDD.
Terminal-Bench 发布首个版本,用于量化 AI 智能体在终端中执行复杂任务的能力,首发数据集 Terminal-Bench-Core-v0 包含 80 个手工编写并人工验证的任务,每个任务配有独立 Docker 环境、人工验证的解法与测试用例。
Awaiting translation
Why it matters: Terminal-Bench 给出 80 个带 Docker 环境和测试用例的终端任务,可用来横向比较不同智能体在命令行中的实际表现。