Prompt Injection
An attack against language-model applications in which untrusted content attempts to override instructions, redirect tool use, disclose data, or manipulate model behavior.
AI-Security Context
Prompt injection occurs when untrusted content attempts to redirect a model's instructions, tool use, data access, or output behavior. Retrieved documents, web pages, email, and tool results can all carry indirect instructions, which makes trust boundaries and the distinction between data and authority central to system design.
Security Boundary
Text filtering alone is not a complete mitigation because an attack can be semantic rather than syntactically recognizable. Tool permissions, user confirmation, least privilege, and policy enforcement must remain outside the model.
Related AI-Security Concepts
Related technical article: Artificial Intelligence: Philosophy, Theory and Practice.