Prompt Injection Attacks and Defense Mechanisms in LLM-Based AI Agents: A Systematic Review

US Journal of New Insights in Tech & Education

Kanita Haider, Oishe Al Mariz, Mahfuzulhoq Chowdhury

Chittagong University of Engineering & Technology, Chattogram, Bangladesh

US Journal of New Insights in Tech & EducationVol. 6, Issue 1August 10, 2026Online First

View Full PDF

Abstract

Large language model (LLM) based AI agents are increasingly deployed to autonomously browse the web, call external tools, retrieve documents, and act on behalf of users, yet this same autonomy exposes them to prompt injection, a vulnerability class in which adversarial instructions embedded
in user input or in externally retrieved content override the developer's original instructions and hijack the agent's behaviour. This paper presents a systematic review of the prompt injection literature published between 2022 and 2026, synthesising findings from more than thirty peer-reviewed and preprint studies to characterise the attack surface of LLM agents and the defenses proposed against it. Attacks are organised along two axes, direct versus indirect injection and manual versus optimization-based construction, while defenses are grouped into four families: prompt-level heuristics, training-time alignment, detection and filtering, and architectural isolation. Drawing on benchmark results reported across the reviewed studies, the review shows that attack success rates remain high across agent architectures, ranging from roughly twenty to over ninety percent depending on the tool environment and adversary capability, and that adaptive attackers routinely defeat single-layer defenses. Architectural approaches that separate control flow from untrusted data, such as capability-based sandboxing and dual-LLM designs, offer the strongest guarantees but at a cost to agent utility and engineering complexity. The review
identifies a persistent gap between attack sophistication and defense maturity and recommends defense-in-depth strategies, standardized cross-benchmark evaluation, and human oversight of consequential agent actions as priorities for future research and deployment.

Keywords

Prompt Injection; Large Language Models; AI Agents; LLM Security; Indirect Prompt Injection; Adversarial Robustne

Article Information

Published
August 10, 2026
Journal
US Journal of New Insights in Tech & Education
Volume / Issue
6 / 1
Article No.
USJNITE0601
Year
2026

Browse

All published articles · Journal archive