用 Codex 让 Astra、Sol 和 Luna 玩《晨风》:GPT-6 三款模型同一任务表现差异明显
开发者用 Codex 作为智能体运行时,通过 OpenMW 改造版接口 AstraBridge 让 GPT-6 Astra、GPT-6.1 Sol 和 GPT-6 Luna Max 在相同环境下完成《晨风》任务“Fargoth's Hiding Place”。
Awaiting translation
开发者用 Codex 作为智能体运行时,通过 OpenMW 改造版接口 AstraBridge 让 GPT-6 Astra、GPT-6.1 Sol 和 GPT-6 Luna Max 在相同环境下完成《晨风》任务“Fargoth's Hiding Place”。
Awaiting translation
作者用标准库 Python 写了一个约 100 行的 chain_check.py,用两条规则检测链式 Skill 审批劫持:单 Skill 规则标记同一文本中同时出现状态变更动作(upload、send、delete、transfer)和审批声明的 Skill,链式规则在已安装 Skill 间构建写读图,标记 A 写入含审批声明的文件、B 读取后执行状态变更动作的路径。
Awaiting translation
作者提出 SKILL.md 只有在 description 与用户提问匹配时才会被加载,正文决定加载后发生什么,因此值得信任的前提是先用真实输入跑一遍并读回结果。他建议把规则写成可校验的形式,并在一个更强和一个更弱的模型上分别测试,且测试要能失败。
Awaiting translation
query-inspector 是一个 Claude Code 插件,包含 tuning-report 和 inventory-report 两个 Skill,从 git 变更中提取 SQL/ORM 查询,诊断缺失索引、N+1 和反模式并生成报告。
Awaiting translation
Future AGI 1.47.0 发布,在运行分析中新增通话与语音指标卡片,并修复模拟通话结果与运行详情的对齐问题。评估对话时,错误语言回答、涉及其他产品的回答以及未获回复的请求现统一计为未处理请求;测试环境评估还修复了完整提示词和对话参与者标识的传递。聊天构建器页面支持折叠与调整宽度,宽度在窗口缩放后保留。
Awaiting translation
Pennyforge ran anonymous health checks on the 78 servers that responded to initialize out of 186 endpoints in the a–b slice of a public MCP registry. Only 40 of them (51.3%) made it through the full flow of initialize → tools/list → one safe tools/call.
Why it matters: Anonymous health checks on 78 registry MCP servers, with reproducible data on tiered authentication and spec version migration.
query-inspector 是一个 Claude Code 插件,由 tuning-report 和 inventory-report 两个 Skill 组成,从 git 变更中提取 SQL/ORM 查询,诊断缺失索引、N+1 和反模式,并给出带 file:line 的 CREATE INDEX 建议。
Awaiting translation
mcpward 更新到 1.1 并上架 GitHub Actions Marketplace,它把 MCP 服务器 tools/list 返回的工具定义存进 lockfile,工具描述、必填参数或 readOnlyHint 发生变化时让 CI 失败。
Awaiting translation
独立开发者 /u/petrucio 用 Claude Code(Opus 5.5)在 11 天内为 Unity 肉鸽卡牌游戏 Kegs of Eternity 做出免费每日谜题 Last Call,已上线其网站和 itch.io。
Awaiting translation
I tested ten Claude Code mods on Claude Code 2.1.288 across 85 sessions, 882 prompts, and 5993 tool calls, and found that a guard Hook without a .catch gets skipped when it throws, so the command runs anyway. Only by adding a catch that returns deny does it fail closed.
Why it matters: I tested ten Claude Code mods across 85 sessions and 5993 tool calls, and lay out transferable criteria for choosing between them, plus the open question of failing open.
作者用 Git worktree 让每个并行编码任务各占一个分支和检出目录,解决了多个智能体编辑同一份代码的冲突,但浏览器测试仍无法确定实际跑的是哪个 worktree 的代码。
Awaiting translation
In a product project where an AI agent writes the code and the author doesn't read it, automated checks repeatedly reached the wrong conclusion. The author found 86 checks that no workflow had ever triggered, a secret scan that missed 438 of 1413 files because Git escapes Russian filenames by default, a new check that mistook WHERE for a table alias and let an injection slip through, and three false alarms from the test dashboard and the agent's replica.
Why it matters: The author walks through five real cases to show why automated checks produce false greens or false reds, and lays out validation rules that carry over to other projects.
rashomon 发布新功能,--timeline 会标记两类模式:只有测试文件被改动后原本失败的测试立刻变绿,以及同一条测试命令在中间无改动的情况下既通过又失败。
Awaiting translation
mcpward 是一个对 MCP server 做黑盒契约与安全测试的工具,把 server 的契约快照进 lockfile,契约变化时让构建失败。
Awaiting translation
作者 Craig Solomon 分享了测试 MCP 服务端守卫的方法,核心是把守卫从工具函数体里抽成可单独调用的纯函数,让拒绝路径变成一张参数化用例表。
Awaiting translation
Future AGI 发布 1.46.0,核心改动在 Simulate 模块,为分析多次运行结果引入决策优先级,Runs 标签页现按环境中完整运行次数显示徽章。该版本还优化了运行摘要生成,以及 API v3 搜索场景时的任务处理,适合运行次数较多的团队。项目支持自行部署,云版本可免费开始使用。
Awaiting translation
开源项目 Juggling(Apache-2.0,alpha)让 Codex 以 Writer 身份在只读沙箱中作答,再由一个全新的独立 Codex 会话复核,返回 PASS、CHANGES_REQUESTED 或 CANNOT_REVIEW,无法验证必需步骤时以 HOLD 加原因码停止且不重试。
Awaiting translation
作者解读 Dan Lu 的智能体实验:让智能体按 RFC 用 Rust 写 Zstd 解码器,比较 26 种条件(含无额外指令的对照组),主要对比用 Codex 搭配 GPT-5.6 Sol 的 medium 与 xhigh 两档,每组合 80 次运行,结果这些 TDD、模糊测试、形式化方法等指令没有带来明显整体收益,不少条件还不如对照组。
Awaiting translation
Awaiting translation
逛逛GitHub 盘点了 9 月份 GitHub 上热度最高的 20 个开源项目。Archify 以约 3.34 万新增 Star 居首,它能把系统描述或代码仓库画成交互式架构图;Ponytail 和 God's Eye View 分别以约 3.12 万、3.11 万 Star 紧随其后。
Awaiting translation
RepoGuard 是一个面向 AI 辅助代码库的架构护栏工具,通过 MCP Server 接入 Cursor、Claude Desktop 等客户端,让 AI 智能体在写盘前审计代码并检查架构违规。
Awaiting translation
南洋理工大学 S-Lab、A*STAR 和 UIUC 团队提出 V-Rubrics,将 50,248 个视觉样本拆成 352,938 条可逐项检查的准则,让视觉事实、推理步骤和任务要求分别参与奖励计算。
Awaiting translation
The author runs a fully autonomous implementation system where an orchestrator hands out tasks to parallel implementation agents (built on Claude Code). At first, agents could mark their own tasks as complete, which led to problems like tests never being run, acceptance criteria not being met, assertions loosened to make tests pass, and hardcoded return values.
Why it matters: The author solved the problem of agents declaring completion too early with a three-layer design: checkable acceptance criteria, completion reports backed by evidence, and read-only validation agents.
Awaiting translation
作者梳理 AI 生成 Android 代码的常见问题,并引用多项研究数据:USENIX Security 2025 分析 223 万个生成代码样本、16 个模型,开源模型包名幻觉率平均 21.7%,商业模型 5.2%;CodeRabbit 分析 470 个开源 PR 发现 AI 代码缺陷率是人类代码的 1.7 倍,性能问题接近 8 倍。
Awaiting translation
RugSnare is a runtime integrity tool for MCP tool descriptions. It computes a normalized hash pin over each approved tool's { name, description, inputSchema }, and any silent change afterward triggers an alert and fails CI (exit 1).
Why it matters: RugSnare hash-pins MCP tool descriptions, keeps watching for silent changes after approval, and shares measured data from 66 official server versions.
作者 Ifeanyi Ejindu 发布 qarunbook,一个面向 Web 和移动应用的人工测试 runbook,用检查项(步骤加预期结果)、各平台通过/失败状态和针对检查项提出的 issue 组织测试,不跑自动化测试。
Awaiting translation
vericoding(AI 引导的认证程序精化)是一种以规范而非实现为核心的软件工程方法,由定理证明器(如 Rocq、Lean)充当认证器、AI 负责猜测实现,机器负责验证代码是否符合规范。它按梯度推进:从 Gherkin 场景的 BDD,到 specsaver 的可执行运行时契约,再到 axiomander 的机械化证明。Scidonia 用该方法证明程序行为并优化慢路径。
Awaiting translation
Rodion Mostovoy 与 Maxim Klyuchnikov 分享了在百万行级企业遗留项目上提升 AI 智能体效率的经验,包括构建术语表、处理技术债,以及用基准测试对比向量 RAG 与 GraphRAG 等上下文引擎的效果。
Awaiting translation
作者搭建了一条流水线,把 Paper Compute 导出的带标签编码智能体会话用于微调模型:会话、轮次和标签先导入 Databricks 的 Unity Catalog,再用 golden、regression、pushback、apology、no-outcome 等标签筛选训练样本和评测用例。
Awaiting translation
Anthropic 在 claude-api skill 中新增 /claude-api build-eval 和 /claude-api hillclimb 两条命令,把评估集构建与爬山调优做成引导式工作流。
Awaiting translation
A product designer with six years of SaaS experience used Claude Code to single-handedly build Котомка, a life-planning app. Nearly all the code was written by AI; his job was to define requirements, review the results, and make decisions.
Why it matters: Using a real repository, the author documented the pitfalls he hit while building a product on Claude Code alone, plus the rules, hooks, and testing guardrails he set up around the AI.
Cloudflare 发布 Worker Previews,让每个 Git 分支拥有独立且接近生产的环境,各自有稳定 URL、独立配置和状态,并带独立可观测性。
Awaiting translation
skillmem 作者发布 0.12 版本,先把十六个不变量写进 docs/INVARIANTS.md,再让 Codex 和 Claude Opus 5.5 分别寻找反例,规则是发现必须附带可复现的失败测试,由 scripts/release-gate.sh 校验。
Awaiting translation
Playwright Agent 的 Healer 能完成跑测试、诊故障、改代码、重跑一整轮,在 VS Code 的 Copilot Chat 里切到 Agent Mode、选 @playwright-test-healer 并贴一句 Run and fix failing tests 即可触发。
Awaiting translation
Rust 开发者 Nikita 认为 AI 只加速了写代码这一段,上下文准备、审查、测试和修复仍要人来做,因此该按“到可上线改动”的时间算收益,而不是按模型首次响应算。
Awaiting translation
GitHub Security Lab 于 9 月 24 日发布 Fuzzing Taskflow,这是一个面向 C/C++ 项目维护者和安全团队的自主 LLM 流水线。
Awaiting translation
Vibe Coding Production Kit(VCP)发布,这是一个零运行时依赖的 CLI,把 AI 辅助开发组织成规范、架构、有界任务、就绪门禁、验证证据和独立评审组成的工程生命周期,支持 Codex、Claude Code、Cursor、GitHub Copilot 等工具。
Awaiting translation
Cursor has released two bots, Rollouts and Security Review, both available on Team and Enterprise plans.
Why it matters: The official docs cover the monitoring and security review workflows for both bots, so readers can judge whether they fit into their existing delivery pipeline.
Playwright 官方 Generator 工具可以按已评审的测试计划逐条生成可运行的 Playwright 用例,作者用演示站 jiucaiquan.com 走了一遍完整流程。
Awaiting translation