The OWASP Top 10 for LLMs identifies critical security vulnerabilities in large language model applications, with prompt injection (where attackers manipulate LLM behavior through crafted inputs) and excessive agency (where LLMs perform unintended actions beyond their intended scope) being among the most common and dangerous threats. Prompt injection can be mitigated through parameterization and sanitization of inputs, while excessive agency requires strict access controls, command whitelisting, and sandboxing to prevent unauthorized system interactions.
Securing LLM Applications: OWASP Top 10 Threats and Defenses
Added:So I'm going to present on OF's top 10 for LLMs. We won't get through all 10.
We have 30 minutes. So I picked a few to highlight, show some code examples. I'll talk a little bit first on how LM apps are built and the architecture behind some of them. Uh because I think it's important to be able to conceptually look at these and know where to put the controls in from an architecture perspective.
I'll tell you guys a little bit about who I am start. This is our agenda. I'll go into the topic. I'll talk as I said on how the apps are built. And then we'll go into three of the top 10 examples. the three that I see as most common right now, but that could change tomorrow. The LLM and AI world is very rapidly changing every day. And I'll try to leave time for questions as well.
From a professional standpoint, there we go. I work for a company called JIT. We're an application security posture management company. That acronym gets thrown around a lot, but at the end of the day, we're helping businesses secure code, and there's a bunch of cool things we do. I'd been in the industry a little over 15 years, doing startups for three years now. This is my third startup. Sorry, about seven years now.
This is my third startup. Um, it's an interesting journey. If you guys are curious about either what the startup journey is like or what we do at JIT, you can see me after. I don't want to talk too much about shop. I'll talk about actually who I am because that's more interesting, I think, to people.
Uh, that's me and my partner on the right. I outside of work love everything nature and nerdy. So, I have the dungeon crawl classic there on the on the side.
That's a a variant of D&D basically if you're not familiar with the role playing world. I love tea, hiking, nature. That is my dog. Come argue with me. I think I have the cutest dog, but you can tell me you do after and show me pictures and my motorcycle up top. So, I like to get out, be in the fresh air in nature when I'm not doing these types of things.
All right, let's get into it. Every single day right now, we see another crazy headline. If you spend time on LinkedIn X, Blue Sky, wherever you get your latest LM or AI news, uh you probably think it looks a lot like this generated picture I made, which was actually pretty funny. It told chat GPT to make its own conspiracy theory about itself and generate a picture on it. And you get things like this, but it's actually not far off on what you see on LinkedIn at least and on X and some of the other social areas. I actually like this one now here. ChatGpt now sees your imagination. That's what it feels like we're we're we're doing. Um it is a crazy world. I think there's a lot of really cool applications for generative technologies and there's a lot of fluff out there. um figuring out how we implement these things in a way that's logical and makes sense is really tough.
Uh I encourage everyone to go and test and play with these things though very much. If you want to hear more about my opinions on what's real and what's fluff, I can also talk about that after.
But because of all these changes from a security perspective, this is us as security practitioners dealing with implementing these tools pretty much every day. You need a spoke break every 10 to 15 minutes and pretty exhausted.
like figuring out how to implement these things the right way securely and develop new architecture frameworks around them is pretty complicated.
Quickly we think these things aren't just hypothetical and we're not just talking about these. We see businesses implementing LLM and we see actual attacks taking place that are meaningful uh to the enterprise. So we've seen libraries that don't sanitize input which leads to code execution. It's LMS but the same type of vulnerabilities we've seen in the past and we'll talk about that a little more. Samsung actually saw chat saw their trade secrets in chat GPT's model foundationally which is really interesting from that perspective and there's other things that people have seen in copilot in other areas that it's either being abused for cyber operations or leaking data and causing heartburn for everyone in security.
Fortunately, OASP has started creating a lot of really great work around uh LM's AI technologies. The OASP LLM top 10 came out I think in 2023 I want to say, don't quote me on that. And then just recently they've kind of moved everything under this Gen AI security project which touches across everything.
So it's not only now just the top 10 from a code security perspective, but there's the working committees around the AI governance side like how do we write policy for our businesses around these things all the way to red teaming and actually implementation of secure code for these. I think everyone that's doing work around this is um really doing a great job. Uh I've spoken to some people here at the conference that are part of these committees as well. I recommend everyone jump in those Slack channels and see what everyone's working on.
So the top 10, let's see what they are real quick before I jump into some of them. They sound a little familiar to what we've done in other parts of the world, right? You have things like prompt injection, which sounds like SQL injection, right? Injection is just a common way of getting things into a system and making it do things that it's not intended to do. So we have that in LMS as well. Uh the other one I'll talk about today is excessive agency. So excessive agency, that sounds like a complicated topic, but in other words, it's just saying the system can do more than it should. So we'll talk about excessive agency. Uh we have things like misinformation, unbound consumption. I run into this one every day on my personal laptop because I test generative tools. I send things to the APIs and I forget what I sent to it, let it run, and I come back and I've spent $20 on a generative AI uh API endpoint.
So, this one's definitely very real from a threat perspective, but the threat's a little different. That's not data leakage or your business is offline because of something, but maybe your business becomes offline because you went and spent a million dollars on uh AI tokens, misinformation, vector embedding. Some of these other ones are really interesting that I'm not going to get into today, but I'm happy to do a follow-up presentation. I'm thinking about actually putting some of this content together into more of like a class format, put on YouTube or something similar. I'd love feedback if you guys think that would be helpful.
I heard it. Yes. So, let's talk about how people are building AI apps. Last year, I'd say midyear, this is pretty much what you saw from an architecture perspective from most businesses uh that were implementing large language models.
Not everyone, not the most advanced, but a large portion of people basically had a front end of some sort. There was some application code that interacted with a large language model. And to make that large language model do something for their specific business use case or their data, they used uh what we would call embedding stored in a vector database. So it's actually very straightforward architecture. There was a little more complication for scaling this, right? This is obviously very boiled down to be very simple and fit on a slide, but you'd have source data, you create embeddings, you store it in a vector database. the LM and the core code can then pull that proprietary data out to give the LLM uh more knowledge about your business, your business application and provide it to the end users. So this was basically how things were being built up until maybe mid last year when people started playing with this term that you guys hear every day in the market called agentic or agent-based AI which is really just automation with LMS is what I say uh at the end of the day. Maybe I'm doing it a little bit of a disservice but it's easy to kind of wrap your mind around it when I say automation but with an LLM. This is how most people are building LLM based applications. Now, you can't see that text down the bottom.
I'll tell you what it says in a minute, though. You have a web front end. You have some sort of control plane. You have a message bus, and then you have what we call agents typically that do a bunch of things with an LLM, take work off of that message queue, do something autonomously, and give feedback back either to a web interface, to Slack, to some other chatbased interface.
Typically um you also have other services that would run as a microservices. So this is if you're used to building event driven architectures uh or you've heard about event- driven architecture design. This is very similar. Uh there actually I don't think there really is much of a difference except the difference being you have these things called agent agents workers whatever which basically is that a handful of prompts more or less. Again, I'm oversimplifying these things to make the concept easy to to grasp, but I have a handful of prompts that figure out what to do when they get a message that's relevant to them. Um, that could be a look up a record in my customer relationship management tool. And so, a salesperson could put in to the web front end. I want to know everything about this sales deal I'm working on. uh that would kick off to the agent, picks it up out of the queue, goes uses what we have as now, model context protocol, which has only been around for a little while, goes out to the CRM, pulls the data, feeds back, the agent, then checks, did it everything right, and it ends up on the web front end. Uh this here, this box that doesn't show up here says only God knows what. So MCP servers are currently my bane of my existence um because they're so new and everyone's trying to adopt them very very rapidly.
Like we're trying to figure out how to do the thing I just talked about for my sales team, but there isn't an MCP that's published by our CRM yet. There isn't. So I can go and get one off the internet. Who wrote it? Like I I don't know what does it do? Does it every time it makes a request to my CRM, does it then just send it to another place? Like I have to go and evaluate these things very manually. I'm not going to talk a lot about that today, but I'd love to talk more about MCP threats and how we manage that as businesses because everyone's trying to implement this very rapidly.
I will talk about some of the other ones though. So let's get into prompt injection.
So prompt injection is when an attacker can design a sequence of inputs, send it to an LLM, and make the LLM do something it wasn't intended to do. Often from a malicious standpoint, this would be disclose information that the malicious user or attacker isn't supposed to have access to.
So, we saw a lot of this when the Chat GPT and some of these other foundational model companies first got their products into market. We saw people making prompts that bypassed what we would call a system prompt or essentially the way the LM is supposed to behave and give users the ability to do things that people didn't want the LM to do. We those companies have gotten very good at filtering these things now. But the first releases there was a lot a lot of publications of look what I got chat GPT to do tell me this horrible thing or give me this sensitive data etc. Those companies are much better but they had a big learning curve when they first published this. There's two CVEs that I think are very relevant for most companies um just to understand a little bit around because these were open libraries people can go get off the internet that had prompt injection vulnerabilities. uh one was last year, one was two years ago and these prompt injection vulnerabilities both allowed full remote code execution. So you put a prompt in, you gave it the right series and you could get full code execution on the system where these libraries were implemented. So pretty big deal, right?
So this is a snippet of code. It's again a lot of this is going to be very fundamental basic code to make it easy to understand anthod on a slide. This would be relatively vulnerable to prompt injection. And so the reason being just like a SQL query that you might generate, same exact concept. All the uh code is doing here is concatenating the user input directly into the prompt you're sending into the LLM. And when you concatenate directly, the LLM basically will just take that prompt and say this is all the same input. I will do whatever it says more or less. Uh and then you wait the response. The LM service sends a response back. So this is basic prompt injection. uh one technique at least where the inputs can concatenated directly into this one prompt that's sent into the system and allows you to override whatever this part might be.
So the way we can deal with this is pretty straightforward actually. You have two main techniques for securing this. You have parameterization and sanitization. You put this between the input and the outputs of the LLM and sanitize the inputs and also parameterize the inputs which allow uh the system to understand what's a system prompt that shouldn't be overrode and what is a user prompt where the user is actually trying to get something out of the system from a code perspective. Uh this is actually how you would do it with Olama. I did Lama because it's free easy to to go and interact with off the shelf. Olma actually has a system where you can tell it with control characters what is a system instruction that cannot be overrode and what is the user input for the system you're building. So parameterizing this query where you have the system instructions and the user input as separate inputs. The LM actually understands this can't be overrode. This isn't perfect by the way.
If you implemented just this literal example that I showed here, I've managed to still get it to accept prompt injection. But when I make this prompt slightly different, it won't accept most prompt injections in my testing. So, it's actually really interesting it the way this is gets implemented. There's better ways to handle this though than doing it yourself. Uh there's libraries that exist that will do filtering for these types of things. I recommend taking a look at them. I can't just speak to how well every library works. I found this one. It seems good. I haven't done a ton of testing to be fair, but I think this is a better approach and where the industry will end up going is having more standardized heristic libraries to deal with these things on the enterprise side. So, it's a heristic based filter instead of my filter that I had on that previous one, which was very little like here's the system, here's the user prompts, etc. Overall, I recommend a combination of these things.
Uh, that's what we're doing in the LM systems we're designing. Um, at JIT, we we have a combination of filtering mechanisms. It's not just one thing that we're doing. It's not just what I showed on the last slide. That's just Jacob's easy easy example. Uh and then we're using things like this as well to do that.
Any questions on on LM1 actually before I jump jump forward?
Yeah.
Yeah. Is there anything that uh when you say like the integrity of the libraries that like the rebuff that I just mentioned?
I don't know to be frank. I think that there's something different in general of AI and LLM compared to everything else we've did done like most injection type of attacks. It was 1992 uh maybe 1996 more realistically but let's just say 1992 SQL injection there wasn't great libraries to do that thing. And so maybe a couple people published it and then on some BBS early at that time someone might say like this is good or someone actually more likely at a meetup group would say oh I wrote that it and you would trust that person right that they did it and those systems kind of got trust over time that were open sourced and openly available through the community. But with AI it's like we're all sitting in offices remotely or our homes publishing and reading stuff. We don't know the people that are writing these things yet. Um, and we haven't there isn't a community reputation and a lot of this stuff over time came from community reputation in security more so than actual technical systems, right?
Uh, so I I don't have a great answer because we're so new into this. I have the same concern with like MCP because there isn't a reputation for a lot of these providers. Um, there's a lot of startups uh that are publishing it.
Yeah.
ask those questions. Number one is filter.
Yeah.
No, no, I'm not familiar with it.
Yeah. The AI act. Yeah.
describes exactly how you model going.
Yeah, I I haven't yet.
describe how built from a regulation standpoint. We often are when it comes to security. Yeah.
this way and they're mostly looking at bias, right? More so than implementation, right? Right.
Right. Into the social dynamics that we're expecting.
Sure.
Yeah. Yeah. The standard was published about a quarter ago. Yeah.
Would you agree that's supposedly I don't have an opinion on there's some differentiation between them. Yeah. Yeah. MCP is just model context protocol. I don't know if I said that earlier. I'll say it just once more in case I did. Model context protocol is basically a standardization to allow LMS to interact with other technology systems for more more simply put. All right, let's jump into into 2. And where are we on time?
Okay, so maybe we'll do one. We'll jump ahead. Uh, sensitive information disclosure. This one's very straightforward, really hypothetically.
There we go. Uh, system accidentally tricked into disclosing information.
This is like anything else with information leakage. Any system can have any sort of information leakage where sensitive data is exposed. The difference being that LLM systems are non-deterministic in their response. So when they have access to data, they often can return data much easier than a different system may, which allows disclosure. LMS are pretty good at this at disclosing information they shouldn't disclose without very strict controls around them.
Uh so this is a very simple example on what a rag uh query may look like. Rag is retrieval augmented generation where you're basically giving the LM access to data uh overall in this case you embed text embeddings is essentially the mathematic representation of language more or less that when you embed it you then can feed it into the LM it can understand it again oversimplifying a little bit when you then query that database you may get everything in it so as if I just do it directly in this implementation any user that called this function. So if you use this function generally in in a code, any user would get any data that's been embedded more or less from this perspective. Uh so how we solve that again is going to be two ways.
There's going to be passing user roles actually to the vector. So you're actually taking the sort of user role and identity data associated with it and you're passing it all the way through to the data queries. And the data structures, the vector databases themselves have roles associated with the data that's stored in them. uh and then only data that has that role associated with that user will get returned back up um from that query. So foundationally at the data level you can implement a control. I highly recommend that. Again this is I've only seen in a couple of the vector data store systems.
I haven't read all of them of course and implemented them uh in production myself just on my laptop uh testing this stuff.
Uh when you're actually creating the embeddings or storing your data in a vector database I do recommend having data filters and data permissions associated with that data. There's a bunch of different ways to do that um from that perspective and I'll talk about it very briefly. And then you should also sanitize your output as well. So you can do output sanitization which is actually a separate uh L under a separate LM top 10 um doing output sanitization. But that's quickly how you fix it is making sure that the user's role actually has access to the fundamental data that would be returned back from an embedding system and filtering data and managing the data that's put into the system to begin with. A great example is in healthcare.
I think I was talking to someone who worked in healthcare earlier where this is a big part of how do LM get implemented in healthcare so the data is properly handled that the there's no leakage that someone wouldn't get someone else's healthcare records it's a big topic in that part of the world so this is an example where now the function would take in the user's role along with the query and return only the results associated with that when it pushes that user role down to the data level. Um, so that way the database just won't return the data at all in this case. So that's how you might deal with this. I'm speeding up a little bit here to try to get a little bit more content in before before the end.
And here's a helper function as well that can be used to filter data out. So you have sensitive pattern matching.
This is similar to what you might have done in any other type of system where you didn't want sensitive data to be returned. You wanted to sanitize data.
You're doing pattern matching. You're using reaxes. Again, there's libraries that do this stuff really, really well.
So you don't have to go and try to implement it yourself.
Access controls, anonymization, data minimization are also things that to consider. We talked about access control pretty heavily there in this example, but don't forget about anonymization and just data minimization to begin with. If the data doesn't exist in the places that you don't want people to get it, they're never going to get it.
What's the time, Tom?
Sweet. sped through that one. Let's talk about excessive agency because this is the most interesting one to me right now with the way agentic systems are being developed and with what we just talked about a minute ago with a model context protocol because that's at least for this this week that's all the rage. Uh we'll see what next week holds for us.
So an excessive agency, there's essentially a mechanism that you can use LMS to do other things that aren't just return text from the models or return text from the other information you've ined fed it from an embedding system or rag or whatever it might be. So in this example here, if an LM is allowed to send emails on behalf of a user and there aren't sufficient checks, it could go and send emails it shouldn't send or for content within those emails. It might have agency that you've purposely dictated to it, but is excessive to what you expect, right? As in, it can do things on its own that you don't want it to be able to do.
This is a really great um piece of research I saw from Pillar Security on uh excessive agency. Uh I recommend going and looking at this. I'm not going to spend too much time with the amount of time we have left on it, so I can speed through here, but this was actually an attack idea that they've had it. I don't think it was in the wild that they saw it where you can have rules files published on the internet for local coders. A cursor is the big one right now you may have heard of. There's windsurf some of these others. Rue code is a VS code plugin. So these all take LLMs and do sort of agentic AI coding. And a lot of them you can essentially give it permissions to do anything. So interact with your system, reach out to other systems, automatically commit code on your behalf, etc. uh and we use rule files in these systems and these rule files tell the system how to behave and act and do the things you want it to do.
It turns out that if you put unprintable characters into these rule files, you publish that on the internet, someone goes and implements your rule file, all the printable characters look fine to you, Mr. Developer, uh everything looks great, but the hidden characters in there that aren't rendered as text are still read in by the models. and AI you can actually cause uh AI coding tools to then go and behave outside of how they should or to bypass certain things that they might. So this falls in that realm of excessive agency where uh in this case it's a really interesting attack idea or mentality or threat model.
So here on the on the uh again back to the event driven side of the house, typically what we're going to be doing for that is controlling what these agent workers actually have access to, how these MCP servers may operate or how these interact with other technologies that are both local or remote. So MCP is typically for remote. Sometimes there's local things too. Um MCP also could be applied for local.
Here's an example just this would be a local LM implementation where the way the data is presented into the system. Uh anything that follows execute colon would be executed on a system. So in other words the agency here that we've given is that any command in terminal can be executed by the lm in this format here. Uh so anytime it sees execute column col in the prompt, it'll go and execute the command and feed that back into the lm for the output. So you can see the problem with that. Um you can actually do this very directly with any of the coding tools I mentioned, cursor, windsurf, uh root code, but sometimes you purposely give it that uh excessive agency in those cases.
So that's an unsafe way to do it. Safe way would be something like this where you're actually filtering what commands are acceptable. So, it can list files or read files may be acceptable, but it can't remove files. One of the best practices for um I I don't know if I can use this term yet. I haven't decided if I'm comfortable with it because I'm not young enough, I feel like, to call it vibe coding, but in vibe coding, the way the internet's talking about using LLMs to to code recently, uh you typically are giving lots of agency for the models to be able to execute commands locally.
Um, often though everyone recommends don't give it RM. Don't give it the RM being the Linux command to delete a file or the Mac command to delete files on disk. Um, so if you don't filter it though, you don't explicitly give it these filters. It will just it could delete everything. Um, all of the code you've worked on, hopefully you're using revisioning systems, git, whatever, and it wouldn't be a big deal, but you could see the problem there. So, uh, having a allowed command list is really important and executing it within a sandbox.
Obviously, this is just a function call to a sandbox that hypothetically exists in this code example, but sandbox that execution as well. Don't just give it root permissions or pseudo permissions to everything on your machine if you're doing this locally. But same thing on the server side of the house.
So, I was able to get through everything. I hope that was informative for you guys. That's some of my learnings. I've tested this. Very open to questions and feedback, of course, on these topics. I think we're all learning together right now on the best way to go and do this and what's real and what's fluff. Um, so please feedback would be great. Other questions, I'll be here after the talk as well.
[Applause] You mind throwing the slides back up so I can go to this?
right here. 2024 55565 and 2023 29374 I'm happy to share the slides with you as well. They're not they're not proprietary. I'll give everyone a link to the Google doc if you want them.
uh like just reading about it uh news feeds uh the Vanna one someone published a research on it of one of the larger AI tool groups I can't remember who um they get all the credit for it of course yeah yeah exactly yeah I don't think that I haven't seen anything myself this year yet um it doesn't mean it doesn't exist I just haven't come across myself.
Oh, thank you. Appreciate it.
Extra. Thank you.
Up Next

Multi-Chain Prompt Injection: Bypassing LLM Security Controls
@donatocapitella
10.5K views•2024-12-09

Secure Multiparty Computation (MPC): Foundations & Challenges
@SimonsInstitute
7.3K views•2015-05-28

Bypassing Tor Censorship: Bridges and Pluggable Transport Guide
@Coding_ForEveryone
397 views•2024-06-11

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence




















![[ML News] Geoff Hinton leaves Google | Google has NO MOAT | OpenAI down half a billion](https://i.ytimg.com/vi/cjs7QKJNVYM/maxresdefault.jpg)






![[성대의 성대한 특강] LLM 보안 위협│황성재 성균관대 소프트웨어학과](https://i.ytimg.com/vi/8gdDfWr7f54/maxresdefault.jpg)











