Offices in Noida · Ranchi, India admin@twaratechnologies.comCareers

AI

Securing LLM applications: a practical guide to the OWASP Top 10 for LLMs (2026)

The ten risks in the OWASP Top 10 for LLM Applications 2026, what changed from 2025, and the engineering controls that keep AI features safe in production.

By Twara TechnologiesPublished 7 min read

Why LLM features need their own threat model

Adding a large language model (LLM) to a product changes its security in ways traditional checklists do not cover. The model reads untrusted text from users, documents and tools; it produces output that other systems may act on; and its behaviour is probabilistic, so the same input does not always produce the same result. A chatbot that can only talk is one thing. An assistant that can read a mailbox, query a database or call an API is a different risk altogether.

A widely used reference for these risks is the OWASP Top 10 for Large Language Model Applications, maintained by the OWASP GenAI Security Project. This guide summarises the current edition and turns it into practical engineering controls.

The 2026 edition at a glance

OWASP published the OWASP Top 10 for LLM Applications 2026 on 3 August 2026. According to the document’s revision history, the previous edition (labelled “2025”) was released on 18 November 2024, following versions 1.0 and 1.1 in 2023. The OWASP project page points readers to the GenAI Security Project site for the current list.

The 2026 entries are:

ID Risk
LLM01:2026 Prompt Injection
LLM02:2026 Sensitive Information Disclosure
LLM03:2026 Excessive Agency
LLM04:2026 Supply Chain
LLM05:2026 Data and Model Poisoning
LLM06:2026 Unbounded Consumption
LLM07:2026 Misinformation
LLM08:2026 Hidden Context Exposure
LLM09:2026 Vector and Embedding Weaknesses
LLM10:2026 Improper Output Handling

What changed from 2025

The 2025 list contained the same themes in a different order, with System Prompt Leakage at LLM07. The project leads’ letter in the 2026 document explains the main shifts:

  • The ranking now uses incident data. Earlier editions were based on practitioner votes alone. For 2026, OWASP assembled 7,714 real incidents from public vulnerability databases and an AI-harm database, classified the 6,639 that had enough detail, and gave that evidence a quarter of the weighting. The community vote carries the other three-quarters.
  • Excessive Agency moved up to third, which the authors call the most consequential change, because agentic deployments are where damage is landing.
  • Unbounded Consumption rose four places, and Improper Output Handling fell from fifth to tenth.
  • System Prompt Leakage became Hidden Context Exposure, a broader entry covering anything placed in the model’s context that users are not meant to see.
  • Scope is clearer. The list covers the model as a component inside an application. Once a model acts on its own, with tools, memory and downstream consequences, OWASP directs readers to the separate OWASP Top 10 for Agentic Applications, announced in December 2025.

The ten risks and what to do about them

LLM01: Prompt injection

Prompt injection happens when any input (a user message, a retrieved document, a tool result, an image or stored memory) changes the model’s behaviour in ways the developer did not intend. OWASP’s central point is blunt: LLMs make no architectural distinction between instructions and data, so there is no reliable way to prevent injection today. Defence has to be architectural.

Controls: assume the instruction boundary will eventually be bypassed, and limit what a compromised model can reach. Validate every structured response in trusted application code before anything downstream acts on it. Pay particular attention to any feature that combines access to private data, exposure to untrusted content and the ability to send data out. OWASP cites this combination as the condition for high-impact exploitation, and removing any one of the three removes it.

LLM02: Sensitive information disclosure

Data can leak through answers, but also through tool-call arguments, retrieved chunks, logs, telemetry and embeddings. OWASP’s foundational controls include classifying and de-duplicating training and retrieval corpora, sending providers only the fields a task needs, enforcing document- and chunk-level authorisation inside the retrieval query, keeping secrets out of system prompts, and scrubbing logs before they reach monitoring tools.

LLM03: Excessive agency

This is the risk that a model can take damaging actions because it has too much functionality, too many permissions or too much autonomy. OWASP’s mitigations read like a least-privilege checklist:

  • offer only the tools the feature genuinely needs;
  • keep each tool narrow (a summariser that reads email should not be able to send or delete it);
  • avoid open-ended tools such as “run a shell command” or “fetch any URL”;
  • give tools the minimum database or API permissions;
  • run actions in the context of the signed-in user, not a powerful service account;
  • require human approval for high-impact actions;
  • enforce authorisation in code, not by asking the model whether an action is allowed.

LLM04: Supply chain

Models, adapters, datasets and conversion tools are third-party components. OWASP recommends vetting suppliers, keeping components patched, testing third-party models against your own use cases, and maintaining a signed inventory that extends the software bill of materials (SBOM) to models and datasets. It also advises checking that AI-suggested dependencies actually exist and are the intended package before installing them.

LLM05: Data and model poisoning

Poisoning can occur wherever data enters the system: pre-training, fine-tuning, embedding creation or retrieval. Controls include tracking dataset and model lineage, validating incoming data, filtering and scoring retrieved content in RAG systems, and using version control for datasets so you can roll back and investigate.

LLM06: Unbounded consumption

LLM calls are expensive, and an attacker can trigger costly computation cheaply. OWASP notes that simple request-rate limits are no longer enough. It recommends limits on tokens per minute and per day, pre-flight token estimation that rejects oversized requests before inference starts, and hard spending caps that stop processing rather than merely raising an alert.

LLM07: Misinformation

The risk is not only that a model is wrong, but that its fluent, confident output is trusted and acted on. OWASP recommends grounding claims in authoritative, current sources, separating generation from execution so claims are checked before action, validating tool-call arguments, and requiring approval for high-impact steps.

LLM08: Hidden context exposure

System prompts, developer instructions, tool schemas and retrieved policy text can all be extracted. OWASP’s guidance is to assume anything in the context is discoverable. Do not place credentials, tokens or connection strings in it, and never rely on hidden instructions as a security boundary. Enforce authorisation and content policies in deterministic systems outside the model.

LLM09: Vector and embedding weaknesses

Anywhere similarity search decides what the model sees (retrieval-augmented generation, vector-based memory, semantic caches), the embedding layer becomes part of the trust boundary. OWASP recommends enforcing tenant scope inside the index query and validating it on the server, applying access control at chunk level, normalising content before embedding, recording provenance for each embedding, and keeping mixed-trust content in separate indexes.

LLM10: Improper output handling

Model output passed unchecked to a browser, database, shell or terminal can lead to cross-site scripting, server-side request forgery, privilege escalation or remote code execution. OWASP’s advice is to treat the model like any other untrusted user:

  • apply context-aware output encoding for HTML, JavaScript and other targets;
  • use parameterised queries for any database operation involving model output;
  • apply a strict Content Security Policy;
  • strip control characters before writing output to terminals or logs;
  • disable automatic rendering of Markdown images and link previews in chat interfaces, which can otherwise be used to send data to an attacker’s server.

Building a secure LLM feature: an engineering checklist

The controls above overlap heavily. In practice, most teams can cover the majority of the list with a consistent set of habits.

Design

  • Write down what the feature may read, what it may do and what it must never do. Review that list for the private-data, untrusted-content, external-communication combination.
  • Prefer narrow, purpose-built tools over general ones.
  • Put a human approval step in front of payments, deletions, outbound messages and permission changes.

Build

  • Keep secrets in a secrets manager, never in prompts.
  • Enforce authorisation in application code and in retrieval queries, using the end user’s identity.
  • Validate model output against a schema, then encode or parameterise it for its destination.
  • Set token, cost and rate limits per user, per key and per account.

Operate

  • Log prompts, retrieved content, tool calls and outputs, with redaction of personal data.
  • Monitor for unusual cost, volume and tool-use patterns.
  • Red-team the feature before launch and after significant changes, including indirect injection through documents and tool results.
  • Keep an inventory of models, versions, providers and datasets in use.

Where the OWASP list fits

The Top 10 is an awareness document, not a certification standard. It is most useful as a shared vocabulary for product, engineering and security teams, and as a review checklist at design time and before release. The 2026 document also includes an appendix mapping each entry to other frameworks, including the NIST AI Risk Management Framework, NIST AI 600-1 (the Generative AI Profile), MITRE ATLAS and CWE, which helps teams that already work to one of those.

For features where the model can act independently, review the OWASP Top 10 for Agentic Applications alongside this list. OWASP itself says neither list covers that ground alone.

Sources

  1. OWASP GenAI Security Project: OWASP Top 10 for LLM Applications 2026
  2. OWASP GenAI Security Project: LLM Top 10 (2025 edition)
  3. OWASP Foundation: OWASP Top 10 for Large Language Model Applications project page

Facts in this article were checked against the linked sources on 9 October 2026. Rules, prices and standards change; check the source before relying on a detail. This article is general information, not legal or financial advice.

Related service: AI & machine learning

Have something you want to build or fix?

Tell us what you are trying to achieve. We will reply with questions, options and an honest view of what it would take, whether or not we are the right fit.