OWASP and LLMs: from a Top 10 to a complete security discipline

If your team is adding LLMs to products, internal tools or development workflows, there is a question worth answering before choosing a model: how will we control the system around the model?
OWASP has been building a concrete answer. What started in 2023 as the Top 10 for LLM Applications is now part of the OWASP GenAI Security Project, an open initiative covering the development, deployment and governance of generative AI systems.
It is not a certification or a recipe that replaces an architecture review. It is a map of risks, mitigations and practices that keeps the conversation from shrinking to “how intelligent is the model?”
OWASP’s initiative for working with LLMs
The project has several complementary tracks:
- OWASP Top 10 for LLM Applications: identifies major risks in applications that incorporate language models.
- Agentic application security: studies systems that plan, use tools, retain memory and act across multiple steps.
- MCP security: provides practical guidance for developing and using third-party MCP servers with more control.
- AI red teaming and evaluation: promotes adversarial testing before and after deployment.
- Governance and secure adoption: includes CISO checklists, governance frameworks and resources for making AI security an organizational practice.
- AIBOM and data security: works on transparency across the supply chain of AI models, data and components.
The evolution matters: an isolated LLM has one set of risks; an LLM connected to data, tools, memory and permissions has a much larger attack surface.
Recommendations you can take into a real project
The Top 10 should not be treated as a checklist to complete once. Turn its risks into technical controls and design decisions.
1. Treat prompts as untrusted input
Prompt injection can arrive from a user, a retrieved document, a web page, an email or another tool’s response. A longer system prompt is not enough.
Separate instructions from data, validate retrieved content, limit what can influence an action, and evaluate responses for relevance, groundedness and alignment with the question.
2. Give the agent the least privilege possible
An agent should not have broad access “just in case”. Each tool needs explicit permissions, limited scope and, where appropriate, human approval before impactful actions.
Reading a knowledge base, proposing a change and writing it back are different capabilities. Combining them into one operation makes errors harder to audit and contain.
3. Validate output before using it
An LLM output is not a guarantee of correctness. It may be unsafe, fabricated, incomplete or incompatible with the target system.
Validation must happen outside the model: schemas, business rules, sanitization, tests, policies and controls in the destination system. For sensitive operations, the agent should propose; a deterministic component should decide what is accepted.
4. Control data, secrets and logs
It is not enough to ask whether a provider trains on your data. You also need to know what is sent, how long it is retained, who can see it and whether logs contain prompts, documents or secrets.
The architecture should minimize what leaves your perimeter, isolate credentials, redact sensitive data and define retention and access for each log type.
5. Test attacks, not only happy paths
A system may handle a hundred normal prompts and still fail when given a poisoned document or a tool that returns malicious instructions.
AI red teaming belongs in the lifecycle: direct and indirect prompt injection, information extraction, tool abuse, privilege escalation, multimodal attacks, poisoned data and service degradation.
6. Keep inventory and traceability
In a workflow with models, prompts, tools, connectors and datasets, inventory is also a security control. You should be able to answer which model and version are used, which tools it can invoke, which data it can query and who approved each change.
Git, execution logs, prompt versioning and a component inventory help make the system auditable and reversible.
When the LLM stops answering and starts acting
The most important lesson from OWASP’s recent work is to expand the analysis when the model has agency.
An agent can plan, invoke tools, retain memory and chain decisions. That creates additional questions:
- What can it do without approval?
- What happens if an external source tries to change its instructions?
- Can memory be poisoned?
- Can a permission obtained for reading end up enabling a write?
- Are there limits on time, cost, steps and network reach?
- Is there enough telemetry to reconstruct what it did?
The answer is not necessarily to remove all autonomy. It is to surround autonomy with verifiable boundaries: least privilege, real sandboxing, tool allowlists, validation outside the model, execution limits, observability and a clear path for human intervention.
MCP needs a security review too
MCP makes it easier for a model to use external tools and data, but that connection should not be treated as a configuration detail.
Before adding an MCP server, review who maintains it, what code it runs, what data it can read, which operations it exposes, which credentials it needs and how it records actions. Pin versions, isolate third-party servers and avoid implicit permissions.
OWASP’s practical guidance for MCP servers points to the same change in mindset: the connector is part of the attack surface, not a neutral cable between the model and the system.
Where KBbridge fits
KBbridge complements these recommendations in the architecture of working with a Knowledge Base.
Externalization happens locally, and product documentation can run offline instead of turning working context into a dependency on a central service. The team chooses the LLM and can use a model hosted inside its own infrastructure when residency or confidentiality requirements demand it.
The workflow also separates proposal from application: AI works on readable, versionable text, the team reviews the diff, and synchronization validates before writing back. This does not replace an AI security program, but it makes several OWASP controls visible in the workflow: scope, traceability, human review and validation outside the model.
A practical way to start
You do not need to wait for a perfect AI committee. For a first project, run a short review with these questions:
- What data enters the model, and what should never leave the infrastructure?
- Which tools can it invoke, and with what permissions?
- Which controls validate output before an action?
- How do you detect indirect prompt injection?
- What is recorded so an execution can be reconstructed?
- How do you roll back an incorrect change?
- Which adversarial tests run before enabling the workflow?
OWASP’s central idea is simple: security is not added at the end with a filter in front of the model. It is designed around the complete system.
Sources
- OWASP — OWASP Gen AI Security Project: Mission and Charter: genai.owasp.org
- OWASP — OWASP GenAI LLM Top 10 2026: genai.owasp.org
- OWASP — Initiatives: genai.owasp.org
- OWASP — Agentic Security Initiative: genai.owasp.org
- OWASP — A Practical Guide for Secure MCP Server Development: genai.owasp.org
- OWASP — OWASP Top 10 for LLM Applications 2025: genai.owasp.org
How you try it
If you want to see how a Knowledge Base can become readable to agents, with local documentation, text in Git and validation before writing back, start with Getting Started. There is a 15-day free trial, no card required, at kbbridge.com.