Перейти к содержимому
Оригинал
Hacker News · Prompt Injection· ramoz·· 12.03.2026Оценка ИИ62

Ask HN:生产环境如何缓解提示词注入

Оригинальный заголовок: Ask HN: What are you using to mitigate prompt injection?

Заголовок и краткое изложение на выбранном языке ожидают перевода.

Краткий обзор ИИ

作者在 Hacker News 分享生产环境中缓解提示词注入的多层做法。输入侧包括在系统提示词中强制角色边界、在 API 层用独立的 user 角色隔离用户内容与系统指令、对工具调用参数在执行前做 schema 校验;输出侧则用评估指标对模型输出打分,检查事实正确性、指令遵循与安全性,RAG 场景还要验证答案确实基于检索到的上下文。

Полный текст

Полный текст на выбранном языке ожидает перевода. Пока показан оригинал.

We've been dealing with this at multiple layers. Here's what actually works in production:

Input-side (preventing injection):

- Strict input sanitization with role-boundary enforcement in the system prompt. Sounds basic, but most people skip it.

- Separate "user content" from "system instructions" at the API level. Don't concatenate untrusted input into your system prompt. Use the dedicated `user` role in the messages array.

- For tool-calling agents, validate that tool arguments match expected schemas before execution. An LLM-as-judge approach for tool call safety is expensive but effective for high-stakes actions.

Output-side (catching when injection succeeds):

This is the part most people underinvest in. Even with perfect input filtering, you still need output guardrails:

- Run the LLM output through evaluation metrics that score for factual correctness, instruction adherence, and safety before it reaches the user.

- For RAG systems specifically, verify that the generated answer is actually grounded in the retrieved context, not fabricated or influenced by injected instructions.

The "defense in depth" framing matters here. Input filtering alone has a ceiling because adversarial prompts evolve faster than regex rules. Output evaluation catches the failures that slip through. We use DeepRails' Defend API for this layer - it scores outputs on correctness, completeness, and safety, then auto-remediates failures before they reach end users. But the principle applies regardless of tooling: treat output verification as a first-class concern, not an afterthought.

Simon Willison's work on dual-LLM patterns is also worth reading if you haven't: https://simonwillison.net/2023/Apr/25/dual-llm-pattern/

Источник: Hacker News · Prompt Injection · news.ycombinator.com