Skip to content

#Agent

48 items today
Today10/6Tue
  1. DEV Community · Codex74

    Codex 与 Claude Code 同任务实测:质量打平,Codex 成本低 2.4 倍

    作者在 python-humanize 仓库上用四个任务(真实 issue 修 bug、按规格加功能、无行为变更重构、带 3 个植入 bug 的代码评审)各跑两遍,Claude Code 2.1.291(claude-opus-5-5)与 Codex CLI 0.160.0(gpt-6.1-sol)16 次运行全部通过检查,四个评审都找齐 3 个植入 bug 且无误报。

    Awaiting translation

  2. Reddit · ClaudeCode / Codex / VibeCoding60

    OpenAI 在欧盟为 ChatGPT 和 Codex 文本加入 textGrain 隐藏水印

    OpenAI 表示欧盟符合条件的 ChatGPT 和 Codex 文本将携带名为 textGrain 的隐藏水印,未来几周内面向所有套餐推出,目前仅限欧盟。该水印是词选择上的统计模式而非可见标签,OpenAI 称其不识别用户、账号或提示词,检测器不公开,仅获批研究人员可申请使用。

    Awaiting translation

  3. DEV Community · MCP76

    How AI agents actually use your MCP server: five failure modes you won't see in logs

    The author instrumented a demo ticket-booking MCP server with a self-built tool called mcpspan and identified five kinds of agent invocation problems that never show up in logs: agents guessing tool names that don't exist; parameters that don't match the schema and get rejected by the SDK before the handler runs; parameter types misunderstood because of how the tool description is worded; retry loops that keep hitting the same parameter; and responses so large they eat into the context window.

    Why it matters: The author used a self-built MCP analysis tool to empirically surface five failure modes in agent calls that logs don't reveal, and these can be adapted to troubleshoot your own MCP server.

  4. DEV Community · MCP71

    用 100 行 Python 检查器识别链式 Skill 审批劫持

    作者用标准库 Python 写了一个约 100 行的 chain_check.py,用两条规则检测链式 Skill 审批劫持:单 Skill 规则标记同一文本中同时出现状态变更动作(upload、send、delete、transfer)和审批声明的 Skill,链式规则在已安装 Skill 间构建写读图,标记 A 写入含审批声明的文件、B 读取后执行状态变更动作的路径。

    Awaiting translation

  5. Habr · Claude Code64

    用英文给编码智能体写提示词更省 token 吗?40 次实测只差 12 个 token

    作者用 5 个 Python 任务、Haiku 4.5 和 Sonnet 5.5 各跑两遍共 40 次,对比俄语和英语提示词的 token 消耗与成本。俄语提示词平均每任务多 12 个 token,但单次运行平均读取约 22.2 万输入 token,其中 94% 是每轮从缓存重读的系统提示词、工具说明和项目文件,语言差异只占约 0.005%。

    Awaiting translation

  6. DEV Community · MCP78

    Prompt injection is a data-plane problem: move the boundary from the model to the tool call.

    The author argues that prompt injection shouldn't be solved by making models smarter; instead, just as SQL injection is handled with parameterized queries, the boundary should be drawn where the agent executes actions.

    Why it matters: The author draws an analogy between prompt injection and SQL injection, arguing for moving the boundary from the model to the tool-call layer, and lays out a practical approach with a strategy layer and separate read and write phases.

  7. DEV Community · MCP58

    MCP 与自定义 REST 工具的安全对比及最新 MCP 架构

    文章对比了 MCP 与自定义 REST 工具在安全上的差异,指出 MCP 只标准化通信,认证、授权、最小权限、输入校验和监控仍需应用与后端自行实现。文中介绍了当前定稿的 MCP 规范 2026-07-28 的关键变化,包括无状态协议设计、server/discover、Streamable HTTP、多轮往返请求和 Tasks 扩展,并说明本地仍常用 stdio。

    Awaiting translation

  8. Tproger · Программирование12

    Future AGI 1.47.0 新增通话指标并改进 AI 智能体回答评估

    Future AGI 1.47.0 发布,在运行分析中新增通话与语音指标卡片,并修复模拟通话结果与运行详情的对齐问题。评估对话时,错误语言回答、涉及其他产品的回答以及未获回复的请求现统一计为未处理请求;测试环境评估还修复了完整提示词和对话参与者标识的传递。聊天构建器页面支持折叠与调整宽度,宽度在窗口缩放后保留。

    Awaiting translation

  9. DEV Community · MCP78

    Anonymous health checks on 78 registry MCP servers: 51.3% complete the full call sequence

    Pennyforge ran anonymous health checks on the 78 servers that responded to initialize out of 186 endpoints in the a–b slice of a public MCP registry. Only 40 of them (51.3%) made it through the full flow of initialize → tools/list → one safe tools/call.

    Why it matters: Anonymous health checks on 78 registry MCP servers, with reproducible data on tiered authentication and spec version migration.

  10. 宝玉71

    SemiAnalysis 实测 Anthropic、OpenAI 等九家 AI 订阅套餐后得出,同样 200 美元,Claude 订阅折算的 Token 用量约为 OpenAI 的 5 倍。

    Awaiting translation

    QuotedSemiAnalysis@SemiAnalysis_

    Anthropic Subscriptions Offer 5x+ More Value Than OpenAI Limit testing every AI subscription plan from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Zdotai, Cursor, and Cognition https://newsletter.semianalysis.com/p/anthropic-subscriptions-offer-5x

  11. 宝玉71

    据 The Information 10 月 5 日报道,Meta 和微软都在减少员工内部使用 Anthropic 的 Claude,转向自家模型和工具。

    Awaiting translation

    QuotedNIK@ns123abc

    🚨BREAKING: Microsoft and META are aggressively cutting employee use of Claude ahead of Anthropic's IPO Microsoft has cut internal claude spend by more than 33%, nuked per-employee token budget from $100k/month to $10k/month, and forced Copilot to auto-route to cheaper models META used Claude code to build Muse, then cut active users from 60,000 to 30,000 (50% decline) after launch, and replaced it with Muse Code Palantir and Nvidia are also scaling back claude over soaring prices and data privacy fears it’s OVER…

  12. DEV Community · MCP62

    A2A 与 MCP 在 2026 年如何分工:什么时候该让智能体互相通信而不是调用工具

    作者认为 MCP 与 A2A 不是竞争关系,而是不同层次的两类协议:MCP 面向被调用的工具与数据服务,A2A 面向拥有目标、能自行规划并回报任务状态的同级智能体。判断标准是问对方是否拥有目标并自行决策,是能力就归 MCP,是自主工作者就归 A2A。两者可以组合成栈,A2A 负责智能体之间的委派,MCP 负责智能体内部调用工具,作者建议先用工具,只有存在真正可委派的目标时才升级为智能体。

    Awaiting translation

  13. Habr · Вайбкодинг62

    一位 Rust 开发者为什么仍然害怕用 AI 写代码

    Rust 开发者 NikTimf 在 Habr 撰文说,自己仍然害怕用 AI 写代码,原因不是生成质量差,而是生成量太大、后续没人真正看懂。他列举了具体代价:多个智能体之间要反复传递上下文,答案冲突时还得自己判断谁对;公司只允许本地或自研模型时,用惯强模型的人很难退回;同事充当 meat proxy 转发 AI 答案,理解任务和推进实现的活仍落在自己身上。

    Awaiting translation

  14. Reddit · ClaudeCode / Codex / VibeCoding71

    windvane:自动起草检查点、在合适时机压缩并自行恢复的 Claude Code 插件

    作者发布 windvane,一个 MIT 许可、无依赖的 Claude Code 插件,用 Python 引擎在会话中自动看护上下文:recorder 根据任务列表、编辑、提交和上一条回复起草检查点,模型一次调用即可接受或改一个字段;插件监控上下文填充,在检查点落盘后于下一个回合边界压缩,并自行发一条提示让模型从检查点继续。

    Awaiting translation

  15. Reddit · ClaudeCode / Codex / VibeCoding58

    AI Pair 发布 VS Code 扩展:让 AI 智能体以人类节奏打字并讲解

    作者发布 VS Code 扩展 AI Pair Programmer,让编程智能体在编辑器里以可跟上的速度逐字输入并解释自己在做什么,用户可随时打断或接管。该扩展内置支持 Claude Code、Codex、OpenCode、Gemini CLI、Cursor 和 GitHub Copilot,任何支持 MCP 的智能体也可手动接入,作者称它更适合能力较强的模型,弱模型在这种工作方式下容易吃力。

    Awaiting translation