AI 编程智能体是否正在制造现有工具难以应对的安全问题?
Original title: Are AI coding agents creating a security problem that existing tools aren't solving well enough?
The title and summary in the selected language are awaiting translation.
有开发者在 Reddit 上征集 Claude Code、Codex、Cursor 等 AI 编程智能体的实际安全实践,覆盖运行前扫描恶意 skill、MCP server 与传递依赖,AI 生成 PR 的安全检查与人工审查,以及运行中的工具调用监控、文件系统/网络/凭证访问限制和沙箱隔离。提问者更关注真实事故、险情和现有工具的缺口,并追问提示词注入是否来自 README、skill 或依赖文件。
I'm trying to understand how developers are actually securing AI coding agents today — and whether there's a genuine gap or whether the existing security stack is already sufficient.
If you use Claude Code, Codex, Cursor, OpenCode, Openclaw, Hermes, etc., how do you handle these?
Before the agent runs
Do you scan repositories for malicious agent components, skills, MCP servers, or dependencies?
Do you worry about compromised or malicious transitive dependencies?
Do you verify what capabilities an agent/skill/MCP actually has?
When the agent changes code
Do you have security checks specifically for AI-generated PRs?
Do you review every AI-generated PR manually?
Have you ever had an agent introduce a vulnerability or insecure change that looked legitimate?
While the agent is running
Do you monitor what tools/actions the agent executes?
Do you restrict filesystem/network/credential access?
Do you use a sandbox or VM?
Do you have something that can detect suspicious behavior across multiple actions?
Have you experienced prompt injection from a README, skill, dependency, or other project file?
Have you had an agent perform an action that was technically permitted but clearly not what you intended?
And what are you currently using?
Sandboxing?
Network proxies?
Permission systems?
Dependency scanners?
Secret scanners?
PR security tools?
Runtime monitoring?
Manual review?
Nothing?
The bigger question:
Do you think there's a real missing security layer for AI agents, or can existing security tools + sandboxing already handle this well enough?
I'm much more interested in real incidents, close calls, frustrations, and gaps in existing tools than theoretical attacks.
And if you think this is a non-issue, I'd genuinely like to hear why.
Source: Reddit · ClaudeCode / Codex / VibeCoding · reddit.com