Why it matters: It lays out a complete workflow for systematically validating Codex Skills with evals, from defining success criteria to deterministic checks and scoring.
When developing web apps with the coding agent, the author often runs into client-side JavaScript bugs. If the agent can't fix them by reading the code, it fires up browser MCP for interactive debugging just to see the browser console logs—burning tokens and slowing things down.
In the Codex Cookbook, OpenAI lays out a complete workflow for modernizing a legacy codebase with Codex CLI, using a COBOL portfolio system as the example and moving through five phases built around an ExecPlan design document.
Why it matters: Using a COBOL portfolio system as the example, it offers reusable documents and a validation workflow for modernizing legacy code in phases with Codex CLI.
Why it matters: The author ported Claude's SKILL.md system to the Codex CLI, with tool mappings and install instructions, so you can judge whether reusing Skills across models is feasible.
Anthropic rolled out its first-party Skills system simultaneously on Claude Code, Claude.ai, and the Claude API, and author Jesse Vincent quickly followed with a new version of Superpowers built on the official Skills.
Why it matters: Drawing on nearly a month of hands-on use, the author compares the official Skills with his own setup and lays out the trade-offs involved in migrating.
10/15Wed
Wednesday
Martin Fowler · Exploring Generative AISelectedAI score7474
Author Jesse Vincent released Superpowers, a set of Skills built on Claude Code's new plugin system. Once installed, it injects a guiding prompt through the session-start hook, prompting Claude to proactively search for and use these Skills.
Why it matters: The author packaged his own coding-agent workflow into an installable Skill plugin, so readers can directly reuse his implementation flow from brainstorming to TDD.
The author walks through their full workflow with Claude Code: first isolating tasks with git worktree, then using a brainstorming prompt to make Claude ask only one question at a time and confirm the design in stages, and finally using a planning prompt to break the plan into small tasks and write them into docs/plans/.
Why it matters: The author splits Claude Code into two sessions—an architect and an implementer—and shares reusable prompts plus a git worktree approach for isolating tasks.
Anthropic's applied AI team argues that context engineering is a continuation of prompt engineering, and the core idea is picking the smallest set of high-signal tokens within a limited attention budget.
Why it matters: Anthropic lays out a systematic approach to context engineering, covering the trade-offs among three long-task strategies: compression, note-taking, and sub-agents.
The author rewrote a long block of CLAUDE.md rules as a GraphViz dot flowchart, using quoted strings as node names, different shapes to distinguish decisions, commands, and warnings, and giving each flow an explicit trigger condition.
Why it matters: After rewriting the CLAUDE.md rules as a GraphViz dot flowchart, Claude followed the rules better, and this approach can be carried over to your own projects.
The Manus team shares context engineering lessons from building AI agents, centered on designing around the KV cache, managing tools by masking rather than removing them, treating the file system as context, steering attention by restating to-do items, keeping errors in context, and avoiding getting stuck on few-shot examples.
Why it matters: The Manus team distilled lessons from rewriting their agent framework four times into six context engineering principles—ready to apply directly to your own agent implementation.