Skip to content

#Testing/verification

0 items today
10/6Tue
  1. DEV Community · MCP71

    用 100 行 Python 检查器识别链式 Skill 审批劫持

    作者用标准库 Python 写了一个约 100 行的 chain_check.py,用两条规则检测链式 Skill 审批劫持:单 Skill 规则标记同一文本中同时出现状态变更动作(upload、send、delete、transfer)和审批声明的 Skill,链式规则在已安装 Skill 间构建写读图,标记 A 写入含审批声明的文件、B 读取后执行状态变更动作的路径。

    Awaiting translation

  2. Tproger · Программирование12

    Future AGI 1.47.0 新增通话指标并改进 AI 智能体回答评估

    Future AGI 1.47.0 发布,在运行分析中新增通话与语音指标卡片,并修复模拟通话结果与运行详情的对齐问题。评估对话时,错误语言回答、涉及其他产品的回答以及未获回复的请求现统一计为未处理请求;测试环境评估还修复了完整提示词和对话参与者标识的传递。聊天构建器页面支持折叠与调整宽度,宽度在窗口缩放后保留。

    Awaiting translation

  3. DEV Community · MCP78

    Anonymous health checks on 78 registry MCP servers: 51.3% complete the full call sequence

    Pennyforge ran anonymous health checks on the 78 servers that responded to initialize out of 186 endpoints in the a–b slice of a public MCP registry. Only 40 of them (51.3%) made it through the full flow of initialize → tools/list → one safe tools/call.

    Why it matters: Anonymous health checks on 78 registry MCP servers, with reproducible data on tiered authentication and spec version migration.

  4. DEV Community · Claude Code78

    I tested ten Claude Code mods: when a guard crashes, the command still runs — only three held up

    I tested ten Claude Code mods on Claude Code 2.1.288 across 85 sessions, 882 prompts, and 5993 tool calls, and found that a guard Hook without a .catch gets skipped when it throws, so the command runs anyway. Only by adding a catch that returns deny does it fail closed.

    Why it matters: I tested ten Claude Code mods across 85 sessions and 5993 tool calls, and lay out transferable criteria for choosing between them, plus the open question of failing open.

10/5Mon
  1. Habr · Вайбкодинг76

    When Automated Checks Lie: Five Cases from a Project Where an AI Agent Writes the Code

    In a product project where an AI agent writes the code and the author doesn't read it, automated checks repeatedly reached the wrong conclusion. The author found 86 checks that no workflow had ever triggered, a secret scan that missed 438 of 1413 files because Git escapes Russian filenames by default, a new check that mistook WHERE for a table alias and let an injection slip through, and three false alarms from the test dashboard and the agent's replica.

    Why it matters: The author walks through five real cases to show why automated checks produce false greens or false reds, and lays out validation rules that carry over to other projects.

  2. Tproger · Программирование15

    Future AGI 1.46.0 为模拟分析加入决策优先级

    Future AGI 发布 1.46.0,核心改动在 Simulate 模块,为分析多次运行结果引入决策优先级,Runs 标签页现按环境中完整运行次数显示徽章。该版本还优化了运行摘要生成,以及 API v3 搜索场景时的任务处理,适合运行次数较多的团队。项目支持自行部署,云版本可免费开始使用。

    Awaiting translation

  3. Habr · Codex71

    Dan Lu 实验:测试为何在有 bug 时仍然通过,以及如何验证测试本身

    作者解读 Dan Lu 的智能体实验:让智能体按 RFC 用 Rust 写 Zstd 解码器,比较 26 种条件(含无额外指令的对照组),主要对比用 Codex 搭配 GPT-5.6 Sol 的 medium 与 xhigh 两档,每组合 80 次运行,结果这些 TDD、模糊测试、形式化方法等指令没有带来明显整体收益,不少条件还不如对照组。

    Awaiting translation

  4. Addy Osmani38

    Addy Osmani 建议为 AI 智能体配备校验产出的手段,测试是其中一部分。他推荐四类测试:模拟真实用户流程的端到端测试作为基准真相;用基于属性的测试声明绝不应发生的情况并生成数千用例尝试突破约束;替换系统时用随机输入对比新旧版本;要求测试快速且确定性运行,否则智能体可能学会重试而非修复问题。

    Awaiting translation

10/4Sun
  1. DEV Community · Claude Code85

    How to Stop an AI Coding Agent from Declaring a Task Done Too Early

    The author runs a fully autonomous implementation system where an orchestrator hands out tasks to parallel implementation agents (built on Claude Code). At first, agents could mark their own tasks as complete, which led to problems like tests never being run, acceptance criteria not being met, assertions loosened to make tests pass, and hardcoded return values.

    Why it matters: The author solved the problem of agents declaring completion too early with a three-layer design: checkable acceptance criteria, completion reports backed by evidence, and read-only validation agents.

  2. Ben Holmes58

    Ben Holmes 公开了驱动其软件工厂的多智能体系统:需求来自 Linear issue 或 Slack 讨论,分诊智能体先调研并决定是直接实现还是提问,实现智能体负责构建、必要时先与子智能体协作产出 spec,验证子智能体对实现结果做端到端测试,代码评审智能体与实现智能体循环几轮后再交人工评审上线,监控自动化则响应告警并创建 issue。

    Awaiting translation

  3. DEV Community · Vibe Coding74

    Android 上 Vibe Coding 的问题:AI 生成代码的幻觉、协程泄漏与安全数据

    作者梳理 AI 生成 Android 代码的常见问题,并引用多项研究数据:USENIX Security 2025 分析 223 万个生成代码样本、16 个模型,开源模型包名幻觉率平均 21.7%,商业模型 5.2%;CodeRabbit 分析 470 个开源 PR 发现 AI 代码缺陷率是人类代码的 1.7 倍,性能问题接近 8 倍。

    Awaiting translation

  4. Hacker News · MCP76

    RugSnare: hash-pinning MCP tool descriptions to catch silent changes after approval

    RugSnare is a runtime integrity tool for MCP tool descriptions. It computes a normalized hash pin over each approved tool's { name, description, inputSchema }, and any silent change afterward triggers an alert and fails CI (exit 1).

    Why it matters: RugSnare hash-pins MCP tool descriptions, keeps watching for silent changes after approval, and shares measured data from 66 official server versions.

10/2Fri
  1. DEV Community · Vibe Coding22

    什么是 vericoding?AI 生成代码时代的验证新范式

    vericoding(AI 引导的认证程序精化)是一种以规范而非实现为核心的软件工程方法,由定理证明器(如 Rocq、Lean)充当认证器、AI 负责猜测实现,机器负责验证代码是否符合规范。它按梯度推进:从 Gherkin 场景的 BDD,到 specsaver 的可执行运行时契约,再到 axiomander 的机械化证明。Scidonia 用该方法证明程序行为并优化慢路径。

    Awaiting translation

10/1Thu
9/30Wed
9/29Tue
  1. Habr · Claude Code82

    A Product Designer Went Solo with Claude Code for a Month: What I Built Around the AI to Keep the Project from Falling Apart

    A product designer with six years of SaaS experience used Claude Code to single-handedly build Котомка, a life-planning app. Nearly all the code was written by AI; his job was to define requirements, review the results, and make decisions.

    Why it matters: Using a real repository, the author documented the pitfalls he hit while building a product on Claude Code alone, plus the rules, hooks, and testing guardrails he set up around the AI.

9/28Mon
9/25Fri
9/23Wed