Skip to content

Practice and tutorials

Best practices, reproducible workflows, and tutorials with commands and configuration examples.

Latest curated items

Items 1–20 · 117 total
Today10/6Tue
  1. DEV Community · MCP76

    How AI agents actually use your MCP server: five failure modes you won't see in logs

    The author instrumented a demo ticket-booking MCP server with a self-built tool called mcpspan and identified five kinds of agent invocation problems that never show up in logs: agents guessing tool names that don't exist; parameters that don't match the schema and get rejected by the SDK before the handler runs; parameter types misunderstood because of how the tool description is worded; retry loops that keep hitting the same parameter; and responses so large they eat into the context window.

    Why it matters: The author used a self-built MCP analysis tool to empirically surface five failure modes in agent calls that logs don't reveal, and these can be adapted to troubleshoot your own MCP server.

  2. DEV Community · MCP82

    Skills Are Not Tools: Why I Gave My Coding Agent 11 MCP Servers and It Fell Apart

    In February I hooked up 11 MCP Servers to my coding agent. Just the tool list in an empty session ate up 34,000 tokens, with Datadog alone contributing over two hundred tools. The agent got slower and kept picking the wrong tools.

    Why it matters: I'm writing this up because 11 MCP Servers of my own blew up my context, and it showed me that tools and knowledge belong in different containers.

  3. DEV Community · MCP78

    FP8 pitfall: GPU bill dropped 47%, but the model outputs “!!!!!!”

    The author ran Qwen2.5 7B/32B/72B on a single AMD MI300X with vLLM ROCm, priced at $2.99/GPU-hr. The BF16 baseline was 7B at $0.227/M, 32B at $0.77/M, and 72B at $1.67/M output tokens.

    Why it matters: The author benchmarked FP8 quantization on the MI300X and found that per-token billing can hide the model's output degrading into gibberish, then gave a reusable way to verify it.

  4. DEV Community · Claude Code78

    I tested ten Claude Code mods: when a guard crashes, the command still runs — only three held up

    I tested ten Claude Code mods on Claude Code 2.1.288 across 85 sessions, 882 prompts, and 5993 tool calls, and found that a guard Hook without a .catch gets skipped when it throws, so the command runs anyway. Only by adding a catch that returns deny does it fail closed.

    Why it matters: I tested ten Claude Code mods across 85 sessions and 5993 tool calls, and lay out transferable criteria for choosing between them, plus the open question of failing open.

10/5Mon
  1. Habr · Вайбкодинг76

    When Automated Checks Lie: Five Cases from a Project Where an AI Agent Writes the Code

    In a product project where an AI agent writes the code and the author doesn't read it, automated checks repeatedly reached the wrong conclusion. The author found 86 checks that no workflow had ever triggered, a secret scan that missed 438 of 1413 files because Git escapes Russian filenames by default, a new check that mistook WHERE for a table alias and let an injection slip through, and three false alarms from the test dashboard and the agent's replica.

    Why it matters: The author walks through five real cases to show why automated checks produce false greens or false reds, and lays out validation rules that carry over to other projects.

  2. DEV Community · Claude Code82

    How I Used Git Checkpoints to Undo Any Change Made by a Coding Agent

    For coding agents running unattended, the author built a checkpoint mechanism based on hidden git refs. Before each task starts, it snapshots the entire working tree—including untracked files—and rolls back automatically when validation fails. The restore operation itself can also be undone.

    Why it matters: With roughly 40 lines of shell, the author turned git checkpoints into rollback-capable infrastructure, laying out the concrete approach and the limits of running coding agents unattended.

  3. Reddit · ClaudeCode / Codex / VibeCoding78

    Use CLAUDE.md plus a local MCP memory architecture to stop Claude Code from constantly needing corrections and burning through quota

    The author has open-sourced the Project Athena v9.9.9 kernel (MIT License), a local-first memory and governance framework built to solve the problems of Claude Code needing repeated corrections, sub-agents overstepping their bounds and modifying files, and losing state when it hits the quota limit mid-task.

    Why it matters: The author validated a CLAUDE.md slimming and local-memory approach across 1900 sessions, and provides a directory structure and verification rules you can reuse directly.

  4. AI Hero · Skills Updates62

    AI Hero Skills v1.3 released: new /implement-spec, /pr, and /retro skills, and CONTEXT.md renamed to GLOSSARY.md

    AI Hero Skills v1.3 is out, adding three skills—/implement-spec, /pr, and /retro—and extending the main flow from /grill-with-docs → /to-spec → /to-tickets into implementation, PRs, and retros.

    Why it matters: The author rounds out the skill set into a full path from writing specs to PRs and retros, and lays out the trade-offs at each step along with the rough edges they already know about.

  5. Hacker News · MCP77

    我连续 37 天夜间测量 MCP 注册表:19.2% 的服务器 24 小时内改动了工具面

    作者搭建 mcp-transparency-log,按计划爬取官方 MCP 注册表中所有公开可达的服务器,记录其工具名、描述、JSON schema 和四个 annotation 提示,写入带签名树头的 append-only 日志。

    Awaiting translation

    Why it matters: 作者连续 37 天夜间爬取 MCP 官方注册表,用可复现的日志量化工具面变化,并公开了两次错误修正过程。

  6. DEV Community · Codex78

    How I Stopped Codex from Burning Through My Usage Quota

    The author found that two Codex browser automation tasks consumed 170,123 and 110,180 tokens respectively, so they set out to control usage through model selection, configuration files, and task splitting.

    Why it matters: Drawing on real measurements where two browser tasks burned through hundreds of thousands of tokens, the author shares a quota-saving approach: switch models and configurations based on task difficulty.

10/4Sun
  1. DEV Community · Claude Code85

    How to Stop an AI Coding Agent from Declaring a Task Done Too Early

    The author runs a fully autonomous implementation system where an orchestrator hands out tasks to parallel implementation agents (built on Claude Code). At first, agents could mark their own tasks as complete, which led to problems like tests never being run, acceptance criteria not being met, assertions loosened to make tests pass, and hardcoded return values.

    Why it matters: The author solved the problem of agents declaring completion too early with a three-layer design: checkable acceptance criteria, completion reports backed by evidence, and read-only validation agents.

  2. 宝玉82

    Drawing on a podcast episode, Baoyu walks through how Lauren Tan, who works on Grok Bot at SpaceXAI, merged 2500 PRs in a single month: at night she lets the AI check and merge on its own, then spot-checks the next morning instead of reviewing each one.

    Quotedlauren@poteto

    i had a lot of fun chatting with @mattpocockuk today about how i was able to land 2,500 PRs last month! Matt is a wonderful interviewer so i think the interview turned out really interesting both of our skill plugins work great together, so i recommend giving both a try and picking the best skills that suit your workflow https://www.youtube.com/watch?v=MN9dGgmLyso

    Why it matters: Using Lauren Tan's practice of merging 2500 PRs in a month, Baoyu explains that skipping individual reviews rests on validation Skills and rule constraints, and lays out the conditions under which he'd apply the same approach.

  3. 佬刘AI78

    Planning with GPT, Execution with DeepSeek: A Cost-Saving Two-Model Workflow

    The author has GPT (gpt-6.1-sol, reasoning tier high) handle planning, key decisions, and acceptance in Codex, and calls DeepSeek-V4.1-Flash through the official DeepSeek Harness to write code, run experiments, and fix bugs—together producing a local PDF toolkit with five working features.

    Why it matters: The author splits the work between GPT for planning and DeepSeek for execution to get a PDF toolkit running, and shares three prompts plus cache usage data that can carry over to cutting costs on long tasks.

9/30Wed
  1. Habr · Вайбкодинг80

    How Flawwow Used a Product Sandbox to Let 150 Non-Engineers Ship 262 AI-Generated Projects

    Artem Gambitsky, co-founder of the Russian e-commerce platform Flawwow, walks through the company's internal product sandbox: it lets colleagues with no engineering background push apps written by AI agents straight to production. In four months, 150 people submitted 262 projects and ran about 5000 deployments—none of that code was ever read by a developer.

    Why it matters: The author lays out the four layers of protection that let non-engineers write code with AI agents and ship it safely, plus the resource pitfalls hit along the way. All of it can be adapted to your own in-house sandbox.

9/29Tue
  1. Habr · Claude Code82

    A Product Designer Went Solo with Claude Code for a Month: What I Built Around the AI to Keep the Project from Falling Apart

    A product designer with six years of SaaS experience used Claude Code to single-handedly build Котомка, a life-planning app. Nearly all the code was written by AI; his job was to define requirements, review the results, and make decisions.

    Why it matters: Using a real repository, the author documented the pitfalls he hit while building a product on Claude Code alone, plus the rules, hooks, and testing guardrails he set up around the AI.

9/24Thu
  1. Lovable · Blog71

    Lovable ships Chats, and opens up the trajectory and inbox architecture behind its multi-agent collaboration

    Lovable launches Chats, an agent that runs at the workspace level, can hold conversations across projects and trigger builds; once changes are confirmed, it hands the task off to the project's builder agent and brings progress back into the conversation.

    Why it matters: Lovable shares the three-layer architecture behind Chats—trajectory, inbox, and activation—which you can adapt for your own multi-agent orchestration.

9/22Tue
  1. AI Coder · Telegram78

    Role-Based Model Routing in Claude Code: One Week of Practice and Hook Enforcement

    Drawing on his own Codex setup, the author built a role-based model routing system for Claude Code: the main thread acts as coordinator, explorer uses Haiku for code search only, worker uses Opus for TDD implementation, verifier uses Sonnet to run checks independently, senior uses high-tier Opus for money, data, and concurrency, and reviewer switches to a different model for semantic review.

    Why it matters: Based on a week of hands-on testing, the author shares the configuration, Hook enforcement, and cost trade-offs of multi-model division of labor in Claude Code, which you can adapt to your own multi-agent workflows.

9/21Mon
9/14Mon
  1. Addy Osmani · Blog80

    Addy Osmani on engineering methods for bringing AI agents into legacy codebases

    Addy Osmani suggests that when introducing AI agents into brownfield codebases, you should first make hidden constraints visible and make cheap changes trustworthy. He recommends dividing code into green, yellow, and red zones: green zones with solid tests let agents move in small, fast steps; yellow zones require writing characterization tests first; and red zones involving sensitive logic like authentication, billing, and permissions must have humans involved step by step. The zones are drawn by hand, and a yellow zone can only be upgraded to green once characterization tests exist and the module owner has reviewed the first batch of changes.

    Why it matters: The author turns the constraints of bringing agents into an old codebase into actionable rules—zoning, characterization tests, and migration units—and cites migration data from several companies as reference.

9/11Fri
  1. OpenAI Developer Blog · Codex66

    OpenAI on How to Rewrite Skills and Prompts for GPT-6 Astra

    In an official blog post, OpenAI lays out recommendations for adjusting Skills, AGENTS.md, and task prompts under GPT-6 Astra: Skill descriptions should be as short as possible and state clearly when they apply, and multi-flow Skills should use a root document for minimal routing instead of turning the Skill into an overly specific step-by-step checklist.

    Why it matters: OpenAI has published guidance on cleaning up Skills, AGENTS.md, and prompts under GPT-6 Astra, and it carries over to existing repository setups.