This video explains three key AI approaches: (1) Simple AI assistants use pre-trained LLMs for basic text generation tasks; (2) Fine-tuning retrains a pre-trained model on specific domain data to create expertise in particular topics; (3) RAG (Retrieval-Augmented Generation) connects LLMs with external data sources like vector databases to provide up-to-date knowledge and context; (4) AI agents combine LLMs with tools and actions, enabling autonomous systems that can think, observe, and perform tasks. These approaches can be combined in hybrid architectures—for example, fine-tuning a model on domain data and then building RAG or agent systems on top. The choice depends on requirements: use simple LLMs for basic chat, fine-tuning for domain expertise, RAG for real-time knowledge access, and agents for autonomous task execution.
Fine-Tuning vs RAG vs AI Agents: A Comparative Analysis
Added:Hey. Hi everyone. Welcome back to my YouTube channel. My name is Sunonny Savvita and I'm back with another exciting and important video. So guys, in this particular video, we're going to understand the differences between the simple AI assistant, your fine-tuned model, rag and the AI agent. So guys this was the very mandatory video and intentionally I kept the video so your each and every concept could be clarified with respect to this fine-tuning rag AI agents and easily you can understand the differences between each other because uh so many people are having a like doubt that which one should I use should I go with a finetuning rag or AI agents or should I use the hybrid approach so inside this particular video I'm going to discuss all this thing and your each and every doubt is going to be clarified. So please make sure guys you are going to watch this video till the end because at the end I'll tell you which one you need to use when and I will uh show you some uh uh fine-tuned model so that directly you can use those particular model inside your rag architecture or inside your agents. Uh so guys as you know I started with this finetuning series and I already recorded four video inside this one. So in the first video I discussed the syllabus in the second I discussed the complete introduction of the finetuning in the next video I discuss about the transfer learning and then in the last video I discuss about all the framework and all the resources regarding the finetuning. Now in this particular video we'll see the differences and right after this one guys we going to jump to the practical.
Yes. So I'll show you the practical of each and every technique of the finetuning in my YouTube playlist itself. So if you haven't subscribed so far guys, please subscribe the channel and hit the bell icon. So whenever I'm uploading a video, you will get a notification. Now without wasting a time, uh let's start with the video guys. So here guys, I kept all the points which I am going to discuss. At the first place, we'll talk about the simple AI assistant. uh then we'll see about the then we'll see the definition of the fine tuning then we'll talk about the rag and then I'll come to the AI agents so each and every architecture I kept over here and then the differences between each one of them uh and then we'll talk about which one we should use when so here also I kept some architecture and some real time analogy so uh let's start with this AI assistant itself and I hope so you understood about the simple Simple AI assistant. So what is a simple AI assistant guys? And one more thing uh let me clarify over here. So whatever architecture I'm going to show you right whether it's a simple AI assistant or whether it's a LLM fine tuning or it's a rag or it's a agent. So one thing would be in a center and that particular thing is going to be large language model. Got it? So guys this large language models are nothing this large language model is a AI model is a deep learning based model which was trained on a very huge amount of data right so this large language model is trained on a huge amount of data which is having millions and the billions of parameter. So if I'm talking about the parameter I think you know the meaning of the parameter if you don't know you will get to know in my upcoming videos where I will discuss about the Lora and Qura. Uh parameter is nothing parameter means weight and biases. If you have a basic understanding of the deep learning or machine learning you can easily understand the meaning of this particular term. Uh now guys LLM are nothing LLM are the model which was trained on a very huge amount of data.
If you don't know about it, you can look into my previous video that I already discussed. Uh so this LLM is capable to answer any sort of a question. This LLM is capable to writing any sort of a paragraph, essay, translating something, writing code, chatting, summarizing. So in short, I can give you the example of the chat GPT. Chat GPT have been built uh on top of one LLM. The LLM is called the GPT. GPT uh there is a different different variants of the GPT which you will get now. So GPT 404 mini 03 right these are some advanced GPT based model and this was uh this was used inside the chat GPD application and it is capable of each and every for each and every task. Uh now guys whenever we talking about the simple AI assistant so simple AI assistant means what? So we have one LLM which we have trained on a very huge amount of data and to that particular LLM we are passing input and we are generating a output and if we are going if we are going to build any sort of a application using this particular LLM right uh so those uh that particular application is called simple AI assistant as you know uh in my previous video many time I explain you this particular thing so guys uh again I'm repeating so we have a llm to this LLM we are passing input and we are generating output. Now using this particular LLM we are going to be create one simple chatbot or one simple QA system. So that is nothing that is called simple AI assistant. Now this LLM could be uh my raw model. So what's the meaning of the raw model? Raw model is nothing. It's a unsupervised model.
Okay.
Unsupervised pre-trained model. So unsupervised pre-trained model means what? So if we're talking about the unsupervised pre-trained model uh so the meaning is it it was trained on a very huge amount of internet data okay huge amount of internet data and the data was not labeled. So for uh training this particular model we have used something or something which is called auto reggressive training. Auto reggressive training. In this auto reggressive training we always predict the next word. Okay. What we do guys in the auto reggressive training we always predict the next word inside the given sentence and this thing I explained you in my previous video if you don't know about it you can go and check it out so auto reggressive training and in that we are going to predict the next word and uh we are going to be train our model so this unsupervised pre-training is also called self-supervised learning in my previous video I clarify all this thing if you don't know about it you can go and check with that but with that particular video. Uh now guys uh this is clear. I hope we have a simple LLM. We have trained it on a huge amount of data and with that we are going to be generate an input. We we are going to be generate some output based on the given input.
Now coming to the next topic over here.
The next topic is a finetuning. So further what we can do we can fine-tune this particular model means we can fine-tune this model on some specific task. Uh so what does it mean? What is the meaning of that specific task? So the task could be anything. Let's say uh I want to fine-tune that on my uh own company data on my own domain specific data for what for the uh medical for the medical chatbot or what we can do so uh I have a company let's say I have a fintech company and in fintech company the requirement would be different compared to this medical company okay so what I will do here again I will collect the fintech data and again What I will do again I will train my model again I will retrain my model which retrained model which uh pre-trained model the raw model right the foundation model or which is also called the pre-trained model right so we are going to be fine tune or we are going to be retrain that pre-trained model retrain the pre-trained model okay and that pre-trained model is also called the raw model or the foundation model so we could have a multiple example and this already I discussed in my previous video if you will go and check there you will find out in a very detailed Okay. So let's uh read the definition. So what I written over here.
So here I written what we are doing we are just taking a pre-trained model and which already train on a huge amount of data set and then what we are doing we are going to be retrained again on some specific data set. Uh so here you can see imagine you have a you you hired a person who already know the English.
This is just a example. This is just an analogy uh which I mentioned over here.
So let's say you have a like person who you you hired it for the English, Hindi and for the basic science but you want them to expert in a medical science. So what you will do you will again uh train that particular person on some medical related information which is already aware about the English, Hindi and the basic science. Right? So finetuning means again we are going to be retrain the model. Again we are going to be retrain the model on some specific data.
So this architecture will clarify everything. So here we have a base model here you can see. Then we are going to be fine tune on some specific data. And now what we got? We got some custom model. So I think I no need to spend much time on this finetuning. Now let's uh look into the next architecture which is a rag one. Guys, this is the architecture of the rag. If you have seen my YouTube channel, uh I think I explained this particular architecture maybe 100 times.
I uh I recorded one very comprehensive playlist on my YouTube channel. uh if you don't know about it you can go and check you will learn a lot regarding this rag R A okay so in the RA what we have see guys u I told you one thing you know it's starting itself so whenever we are talking about any of the concept so the LLM will be the center of it okay so here also you can see the LLM this one now uh whenever we talking about let me highlight the LLM I think you can see this is the LLM guys whenever we talking about the rank so in the rag actually we are going to be connect our LLM with this here you can see this external data source okay we have this external data source or this could be anything it could be any database this could be any API this could be any web page this could be any document so whatever you can think with respect to the data this uh could be anything what we are doing we are going to be store this particular data inside the database Now this database generally is called a vector database because we are going to be vectorized the information. Now why we are doing that vectorization so that we can perform the retrieval on top of it.
Okay. So here is the retrieval pipeline this one. So we pass the query to the vector database and we fetch the relevant data. Okay. Now why we are fetching this relevant data? So that we can provide that relevant data or we can provide the appropriate context to the LLM. Okay. So the idea of the rag idea behind the rag is very simple. Uh we can connect our data external data. So you can say this is nothing. This is our external data to lm that's it. And we are storing it after doing a chunking embedding right where we are storing in the database.
And then we are going to perform the retrieval. So why we are doing it? so that we can provide this particular data the right data the contextual data and the data based on the user query to the LLM itself that's it right so that's a simple meaning of rag I think you already know so let me revise all the things so we're talking okay I I'll revise that let me complete the AI agent also and after that I'll revise it guys when we're talking about the rag so rag is stand for retrieval argument generation and here we have a retrieval we are going to be retrieve the relevant document we are passing it the llm and then we are passing it to the sorry then we are going to be generate a final answer. So we have a user question. We are retrieving the information relevant document. Then we are passing to the LLM and then final answer. And from where we are going to be retrieved the information we are going to be retrieve the information from the vector database. If you have any sort of a doubt regarding the rack, you can look into this particular architecture. It will clarify your doubt in a better way.
Okay. Now there is some more information regarding the rack. So let's say we want to connect our LLM with some realtime info. Okay. Realtime info. uh let's say my LLM was trained last year last year. Okay. And now this year I'm using this particular LLM and and and I'm asking who won the world cup who won the World Cup in 25. So my LLM will not be able to answer. So in that case what I will do I'll connect this particular LLM with some external data sources. Okay. Here I have stored some data which I fetch maybe from some web pages or maybe from some API the realtime data and then I'm providing this particular data to the LLM some context okay and then we are generating a final answer using this LLM. So why we use this rack guys to connect our LLM with a realtime information okay to uptodate for up to for updating my LLM knowledge in the form of context and one advice I can give you with respect to this rag so whenever we are talking about the rag guys okay so guys always try to use rag with some good reasoning based model in some whenever you are going to build a rag based architecture now so use some good reasoning model. Now this reasoning model will be capable to understand your complete context. Okay. So this is the one of the advice which I wanted to give you. Now let's move to the next one which is AI agent. Now guys uh whenever we are talking about the AI agent, so we should always keep couple of thing in our mind.
The first one uh in the AI agents okay the AI agents it's having a capability of the thinking it's having a capability of taking a action it's having a capability of making a observation now how so guys uh whenever we are talking about the AI agents now you need to keep one more thing in your mind which thing if you have attended my previous uh session if you have seen my previous video then you can easily answer. First is a LLM and the second is a tools. Okay. So this LLM is nothing.
This LLM actually it is a brain of the agent and this tool is nothing. It's a action. I can give you very good example. Let's say uh human human is having brain. Okay. And human is having the other part of the body, right? Let's say hands and legs. Now brain is giving a brain is giving a instruction. Okay, a brain is giving a instruction to the hand. Okay, pick something. Let's say pick this mouse. Pick this phone. Yes, I'm able to pick it, right? Let's say my brain is giving a signal to my uh legs. Okay, run, walk or stop. So, I'm able to do that particular thing. So here in the same way LLM is a brain okay that is a artificial brain which is giving a instruction to the tool okay and the tool is nothing it's called the actions getting my point so in a similar way see brain human brain it's giving a instruction okay and based on this particular based on the given instruction my hands and legs is taking a action now you can again compare so lm is nothing it's a brain it's giving some sort of a like a instruction right and based on that particular instruction we are taking a action how we are taking action via this tools so this hand and legs is nothing it's my tools okay this action is being generated by the brain okay and here the action is being generated by the llm I can give you a very good example let's say here uh uh you written a mail okay so let's say you return one mail now this mail actually you can write using the llm But if you want to send this particular mail, okay, whatever mail you have written, if you want to send it to someone, you cannot send using the LLM.
For that, you required some external API, right? So you required some external API, some external code, right? So this API, this functionality, it is a tool. Okay? Using this particular tool, using this particular API, you are performing a action. So LLM is your brain which is capable to write a mail but API and the tool okay which is performing a action for you for what for sending a mail right so that is what that is a tool right using that particular using the LLM what you are taking you are taking some instruction that instruction is nothing it is coming in the form of prompts okay you are passing a prompt you are taking that action and you are calling this appropriate tool for sending a mail that's it guys now uh here I have dro draw an uh diagram. So here see uh we have a brain in between.
So here we are giving an instruction we have a tool calling capability okay we have we can sustain the memory okay we can do the planning okay based on the multiple iteration and observation and all everything we can keep and here we have output so in short I can say uh again I'm repeating so in short I can say AI agents are nothing it's a smart AI assistant it's a smart AI assistant Okay, AI agents is nothing. It's a llm plus action. That's it guys. Okay. Uh whenever you are not able to think about this LLM and action, you can recall this particular example writing a mail and sending a mail. Guys, uh I hope we understood about the AI agents rag and all everything individually. Now let's look into the different and see uh which one we can use when. So what is the difference? I think now you easily you can uh uh write a differences. So LLM the normal LLM is nothing just a pre-trained model.
Finetuning is what a mega LLM expert in a specific topic. Okay with some uh data we are retraining the model. Here retraining would be included. Now we're talking about the rack. So in the rack what we are doing we are going to be connect our LM with some external data source.
Okay. Now here uh this external data source EDS is nothing. It is called the knowledge base also. It is a database. Got it? Now if you're talking about the agents guys. So as agent is llm plus tools. Now this tool could be anything.
Anything means anything. Okay. It could be any realtime API. It could be any database connection. Okay. It could be it could be any database connection. It could be uh maybe whatever right?
Whatever uh logic and the functionality you want to write it could be anything.
And in LLM there would be one more thing. It is capable to think.
It's capable to make a observation and how based on the prompts based on the given instruction.
So agent rag is a different where we are just going having a database vector database but agents having some intelligence thinking capability observation capability based on the given instruction and the tool calling capability and this tool calling could be anything not only the vector not only any database it could be any database any API any sort of a custom logic okay so here is again one more level of definition this is just for you if you want to make a more clarity on top of this particular topic you can do Now uh the next question which one we should use when I think you are waiting uh for this particular question only now uh let me clarify guys when we should use when which one and one more very important question is there uh can we use together can we use some hybrid approach so let me tell you all this thing with my experience what I what I feel uh because I work in a real time I did fine tuning rag everything and what I feel guys for sure I'll I I'll discuss that. So first let's look at the formal thing formal uh uses of it. So we have a normal AI assistant right normal LLM model. So we we can use for what for doing a chat text generation chatbot writing answering and all everything I told you.
Now let's say we have to train that particular model on some specific data.
Okay to make it export. So here we can fine-tune it on that specific data. Now let's say we want to connect my model with some uh knowledge base with some for some live knowledge okay we can use the rag architecture where uh data injection retrieval uh would be included now if you want to build a autonomous system uh so there we want to give the capability of the thinking search plan and all so we we are going to build the AI agents right now in the starting itself I told you whatever architecture we are going to build the one thing is going to be common that's going to be this LLM so here is also LLM Here also this LLM this one. Here also this LLM. Okay. And here also this LLM this one. Right. So the my main question is uh okay individually we can do anything mean we can do the finetuning and we can use that model. We can create a rag right and yes the rag I already told you use the good reasoning based model. Rag will work in a better way. Same goes with the AI agent also use some good model and for sure it will work in a well manner. Uh right. But guys the question is can we use both all together? Can we combine this particular thing? Yes or no. So for that also here I written some sort of a thing with a real life analogy and for sure you will be able to understand. So guys here you can see uh if we we have to do some very basic task. Okay. So in that case do use the simple LLM without like fine tuning and all if it is not required. If finetuning is required, you can fine-tune it and in straightforward manner you can do that rag. I already told you when you have to use whenever you have to connect your LLM with a with some external data sources with some live knowledge and all you can create a rag and if you want some autonom autonomy inside your application use the uh agents okay now the thing is basically uh let's say we have llm we have one llm now this llm okay so we have fine- tuned this llm on some domain specific data fine-tune on this particular LLM some domain specific data. Now, cannot we use this fine-tune model for further for the rag architecture? Yes, we can use it. Who is stopping to us? Cannot we use this fine-tuning model for a for creating our agent application, agentic based application? Yes, we can do it. Who is stopping us? So, I hope you understood now. So yes the combined solution absolutely the combined solution is possible and that could be my best system. So first we have a base model the foundation model we can fine-tune it on some specific data set and on top of it either we can build a rag or agent and nowadays guys some agentic rag is also possible. I already taught you on my YouTube channel you can go and check out that. So here you can understand with some real life analogy. So let's say uh you have a brain you have a like train your brain okay you have fine tuned the brain uh means you got a expert brain right on some knowledge and all maybe some uh someone is teaching you something or you are observing from the uh surrounding uh then basically you are going through the internet and then you are going to be retraining your brain right means sorry you are going to be connect your brain with some Google knowledge Wikipedia and books and all okay and now see you have a brain, you expertise your brain, you connect it with some uh external data sources and then agents what you are doing. So you have a expert brain with some external knowledge uh external knowledge which you connected with the internet and like which you got which sorry which you train yourself okay with your surrounding knowledge and then you connected yourself with the internet and all and then on top of it on top of it you having a capability to do something extra right with your actions. So now you have all the knowledge with that you can explore the phones, tablet, car, bike or you can do any sort of a thing.
So this is just a real life analogy which came in my mind that's why I kept it over here. You can think in a more better way or you can just look into this particular uh like example. So you have fine to the model just to change your tone tone according to the domain.
Okay. Uh on top of it you can create a rag layer. Okay. And then the same rag right. So uh not not the normal rag you can create a autonomous system some advanced rag that's called the agentic rag. So this uh kind of solution is possible. I hope you understood. Now what I'm doing guys I'm showing you some good model uh from the internet itself.
So here uh here I find out some fine-tuned model and your task would be over here. So you need to download this finetune model. See here is a DC coder 676 7B inst. So this model is specifically was uh trained on the coding data set and specifically for the cod coder and all. You can utilize this particular model and on top of it you can create your rag application. Okay where you can uh connect your this particular model with some external github repositories and repositories and all. Now on top of it itself you can create a agentic flow right you can connect your uh uh this particular model using some tools and al some real API for fetching a code or anything uh and then based on that you can generate a answer now here is one of the model from the mist only so open homes and this model have been trained on a huge amount of data guys with the mythology homes design and all. So you can utilize this model and on top of it you can create your own rag architecture and in the going forward classes we are going to retrain our own model for sure I'll show you that from a scratch. Now here is one more model. So this model actually it's a llama model again it was finetune on some external data. So you can identify on which data it it have been trained and then on top of it you can create a rag or agent or agentic rag anything. So on top of this fine-tuning model agent rag everything is possible and we can do this thing in a custom way and for sure we are going to do it in a upcoming session. So yeah this is it for this particular video. I hope your each and every doubt is clear now and you can explore more about it from your end as well and if you have any doubt you can ask me in a comment section. So until guys uh thank you. Bye-bye. Take care.
I'll see you in the next video.
Up Next

Direct Preference Optimization (DPO): Fine-Tune LLMs Without RL
@SerranoAcademy
29.4K views•2024-06-21

Secure Multiparty Computation (MPC): Foundations & Challenges
@SimonsInstitute
7.3K views•2015-05-28

Bypassing Tor Censorship: Bridges and Pluggable Transport Guide
@Coding_ForEveryone
397 views•2024-06-11

Neural Networks Explained: Math, Layers, and Learning Fundamentals
@3blue1brown
21.9M views•2017-10-05
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Artificial Intelligence







































