System calls are the fundamental mechanism by which user applications interact with the Linux kernel to access hardware resources, manage processes, and perform I/O operations. The Linux kernel is organized into five main subsystems: process scheduler (sched), memory manager (mm), virtual file system (vfs), network interface (net), and inter-process communication (ipc), which work together to provide core operating system functionality. System calls operate through software interrupts (traps) that transfer control from user mode to kernel mode, allowing applications to request services like file operations, memory allocation, and process management without directly accessing hardware. This abstraction layer provides security by preventing unauthorized hardware access and enhances portability across different Linux distributions. Tools like strace can trace system call execution to analyze application behavior and identify inefficiencies.
Linux System Calls Explained: Kernel Architecture & Syscall Tracing (Part 1)
Added:hi i'm dj ware on this episode of the cyber gizmo i'm going to continue the linux internal series and today we'll be talking about sys calls right after this [Music] so i guess uh i'm gonna call this part one even though last time we talked about the protection rings which is kind of the introduction i guess or part zero that was really the entry way into this but this is what uh i think you you folks wanted to see as some in depth and so i'll be going through all of the subsystems and talking about those and so but in order to get to those subsystems you have to do a system call to access them so that's why i thought i would start here but there's some other information some background stuff that we need to talk about and and that is the official names for what linux kernel calls their subsystems and there are five of them there's the process the process scheduler or sched the memory manager or mm the virtual file system or vfs uh the network interface or net and inter-process communication or ipc if we're looking at how the dependency map works in the kernel this would be a depiction of that dependency map so let's take it from the top let's look at memory manager so if an arrow comes out of memory manager that means that it depends on what it's pointing at so in this case memory manager is dependent upon the virtual file system the process scheduler and inter-process communication if you look at the virtual file system it's dependent upon the memory manager the process scheduler and the network because some virtual files are nfs or samba they can also be gluster and so forth and so you do have virtual file systems that sit across the network in a process communication depends only on the memory manager and the process scheduler process scheduler is only dependent upon the memory manager and then finally you have network interfaces which are only dependent upon the process scheduler to work so that's how you read that and that gives you an idea of how everything is kind of knitted together most operating systems i'm going to talk about this are monolithic all that means is that the operating system is one big executable and all of that part runs in the system space or kernel space and the kernel binary contains all of those things we just talked about the process manager the memory manager all that stuff is inside the kernel image linux is a monolithic kernel but there are others and one of them is the microkernel based operating system at the time that these were developed and mach originally was an example is an example of that but it originally was a monolithic kernel but the idea behind it was to actually drop the amount of dependency on the the kernel to do the work so in micro kernel based operating systems typically those these five functions oops let me go back out here so these these particular functions only the kernel itself would be inside the protected things like your device drivers and your file system and your inter-process communication your network they would all sit out in user land and in order to use those the the work the the program or the application would send messages and those would then be picked up uh by the individual pieces that were in user land the file system the memory manager and all that stuff and then it would go ahead and and do whatever it needed to do and of course there was things that the kernel did too and that would be to manage the process schedule it still did that inside it still does that inside i believe so yeah and mac os is based on mach it's actually the new x new and uh but it is a a derivation of mach the micro kernel also handles things like the mass message passing itself there are interrupts that it can handle uh as well because certain devices have to have interrupts in order to create a direct memory access channel to push memory to push something into memory and then they use the interrupt to inform the process that hey i'm done the memory is there you can get the data and then there's low-level process management and io as well but there's more than that um some people went uh if you look at the well let's talk about exokernel first this one was to try to crush the amount of kernel code that was inside of the system space down to the very barest of minimums the program would sit out and whenever it needed something it would just make a call to a library and the library in turn would then make the privilege call to the system so the idea was to allow multiplexing of requests at the same time to the same hardware and this would allow those requests to be queued up or to be executed if there was a way to do that in the hardware uh to allow those to occur at the same time that the thought behind this was really two-fold one by doing it this way it would free the kernel uh from the actual weighting of the i o to complete so they it would be out of the way and it could go on and do some work on some other multiplex tasks while that i o was in progress mit did some research into this this is kind of their design i don't know of any os's that currently use this but i do know that the i the second part of the idea was hoped that by doing things this way that i could use the libraries and map them for windows i could map those libraries for mac os and i could map those libraries for bsd for linux it didn't matter what os and what kernels were loaded into the system that i was running on but i wouldn't need a virtual machine to run them and i could have them all resident at the same time so it provided resources and memory on the machine to do that of course but i i mean i can't imagine how microsoft or or apple would react to something like that i'm sure it would not be pleasant uh but uh anyway that was their thought and that was their idea there's a hybrid kernel which kind of combines the the attempts to it could buy what they call the best of both worlds so attempts to combine monolithic and micro microkernels bos was one of those haiku is a modern version of bos and they do implement a a hybrid kernel as i understand dragonfly bsd is a hybrid i could now if i'm wrong about that let me know i mean things change and my slides may be a little old the data i have may be out of date uh x new uh of course is the mock based kernel that which the mac os is based on is actually a hybrid and windows nt was a hybrid it had elements of both uh the macro the micro kernel as well as the monolithic version of the kernel there's also something called a nano kernel which crushes it down even further you may even see even smaller versions called pico kernels and those are really trying to drive down to the minimum the least amount of system space and system code that's running and uh and key os was one of those i think that project's been defunct for quite some time now it uh i don't know it was i think a research project originally and uh and it was kind of exposed there was some people that played around with it when it was around but uh it didn't go anywhere as far as i know a lot of times you'll see that with operating systems you'll see them they'll come up the in the academic world it's more of hey i've got an idea i want to go try it out and they'll do that and then they'll see who who who wants it uh i mean yeah and if they don't want it it doesn't dies on the vine but anyway in order to interface with the hardware on linux you have a mechanism called the system call and that interfaces with hardware and other resources on the machine as well as the process scheduler in order to get itself into the run state which is the whole that's the whole thing that programs try to do is get themselves into the run state but that frees the user from having to learn all of the details about low-level programming having to you know allocate memory push something into memory that you want to write and then then step by step here's where i want to move it onto the i allocate the disk space i create the inode table and then i i insert the data that's in memory into that area on disk and then i write the address to it in the inode table where that file is and then i close it and i move on that just frees you up from all that low-level programming the only thing that we worry about is open read write and close i mean that's basically it there are of course positional reads and positional rights but but the basics from all that other stuff that has to happen is done by the device driver for us and the other thing that system calls do is they increase system security because the the user applications and user programs have no direct access to the hardware so only it acts as a proxy the only the linux kernel can do that and then it increases the portability of programs the syscall interfaces if it's implemented consistently across different distros or different versions of linux even different versions of linux up to a point would allow the same programs to be portable along that now portability doesn't mean it will execute it means that it can be recompiled and run so if i'm subscribing to the system calls for let's say linux 4.19 and i'm on an arm processor i can be pretty well assured that my application that runs will run in an expected way with the expected same results that it that a a normal system call would get so that's what that means it means i don't have to do anything in my program to facilitate that so what does a system call actually look like well we have a user application here that's running in user mode and it's going to issue an open command and this could be an open of a file it could be you know anything any kind of open open opens a lot of stuff but in this case there's a it calls the syst there's a mechanism that calls the system call interface and that kind of hangs out on the bridge between the user mode and the kernel mode and that will receive that call and it'll process it it'll set up a protocol and any architect architectural dependencies that are needed in order to look up and find where that open implementation is in the it could be in a library somewhere and the syscall mechanism will go and do all of that that will load the library and then find the open and then pass the control over to it the open then goes and does its stuff and it returns its result on back to the system call interface which in turn sends it back to the application kind of a a little bit different view of how that mechanism of the call triggers is that the user process makes the system call the system call executes a trap a trap is nothing more than an interrupt a software interrupt it's not a hardware interrupt it's a software interrupt that says hey i got some data wake up and it's it and uh the the kernel will say okay i got it and it'll make sure that it has a process that i think that's what the mode bit is for make sure it hasn't already processed it if if it hasn't then it goes ahead and executes the call and writes out the return and puts a bit that says one then hands it back to the system call interface on your application which then picks up that data so that's basically it it's not a very complicated thing although system call and the library functions operate similarly and that that is by design not by mistake uh so yeah it's the calls to a library and the calls to a system call will look the same uh and but the system call is implemented in the linux kernel not inside of a library and there's a special procedure that's used to transfer control to the kernel so yeah there's a special thing that has to happen a hoop that has to jump through before it'll allow it to actually actually do it so we can actually watch a system call being fired and we can see how the kernel handles it by using s trees and that traces the execution of any program and it lists any system called that that program will generate as well as any signals that it receives in response so i have an example here of sister ace on host name and each line in there will correspond to a single system call so i run that systrace and i get a line for every system call that's made for each system call the name is listed followed by the arguments and return values so let's i think probably the next thing to do is actually go set one up and let's go look at it so let me get out of there and we'll do that um let's just get out of that directory altogether and let's do uh an s trace you don't have to be root to do this because you can execute programs and the one let's take the host name one that we just did there so first of all what am i expecting to get before we do that so i want the name of the host and that's the expected return but what does it do how does this actually work so you can you can see quite a bit happening here there's the beginning of the systrace there's my command and then this is the execution so this is that actually starting up this is systrace actually starting the name application and it found it in user bin and then this is the arguments that's passed and if you've programmed at all in linux you know that argument zero is always the name of the command if there had been any additional arguments they would be listed here and then this of course is the memory address that's given to it as to the location of the of the hostname then it sets up some architectural protocols and then begins to access libraries that have that syscall in it and in an attempt to try to resolve what to do with that that hostname command so it's setting up memory it's starting to read the libraries it's starting to read code out of the libraries and then it starts to execute and write data back into memory it then does a close the architecture comes back in and does some stuff and then finally it decides to execute you name with the probably the node name bits the option and then gets a return for the msil so as you can see that hostname is kind of inefficient because uh it's gonna it's gonna do nothing more than call you name and your name is actually gonna call it back so if i did a u name minus n that would be more efficient than doing this so there you go now you know now you have some tools that start to show you g am i doing this the most efficient way or is this really just generating more work than it needs to do so uh if i did an s trace on oh i don't know uh let's i guess we could do that we could do uh you name uh that's not gonna work so so let's see it should be a lot shorter so it's set up yeah it skipped well no it didn't skip all the same protocol steps are the same but in this instance it's already ready to fire out the result so it was a lot more efficient to get it there and then this return result right here that's writing back is what ends up back on my stack for my application which would be the shell in this case that requested it so that's it there's nothing more to it than that and you can play around with all of the any of the commands that you want and and you know and and see what uh you know how those things work there's other system calls that you'll see uh if you did like a ps that's a an s trace of ps you'll see probably some of the most inefficient coding i have ever seen in my life where it actually tracks through every proc file that's out there of a process that's running to identify if it belongs to you unix doesn't linux doesn't have any mechanism to determine what you're actually running it actually has to go search through the entire process stack to find which ones you own and which ones you're running so if you just had a ps by itself that's probably the most inefficient way to find anything because it's got to go look through everything and then and then one line at a time it has to return it back to you so yeah so yeah you'll start to see things like that where you know you know if you're trying to do things in the quickest way the fastest way syscalls is an interesting way to kind of and s trace is kind of an interesting way to see how things are running if you're running an application and you want to see if you're doing things you know and you know how your how your your machine and how your application is responding s trace is a really good tool to be able to go in and analyze your application to see how well you're doing things and you can play you know if you have the time we don't always have the time when we're developing to play around with potentially new things but there's a lot of different ways to do things and maybe there's a better and a more efficient way of getting at the data that you need or pushing data back out someplace else that you need anyway next time we'll talk about process management and we'll delve into that but i just want to talk about syscalls first because it all starts with the assist call so without it nothing in the kernel will happen as far as my application programs are concerned so uh yeah anyway i i hope you i hope this was useful for you uh if it was let me know uh you know there's always things that you're going to i'm going to skip over i'm not going to cover air i don't want to get down into the weeds so deep that you fall asleep or people that have no idea what this means or what it's doing but i don't want to i don't want to bore them so basically this is just kind of a high level if you want to explore more out in the uh out in the the linux kernel where you go to get the archives for the kernel there are is a huge amount of documentation that will go through different parts of the system they'll go through usually what you'll find out there is that for every release though they'll have sections of manuals that talk about what's been introduced so if you want to piece all this together the only way to do it is to go to the the last lts release that had some major change that you were interested in and start there you know i i don't know how many times when we get to process management that's going to be a really a hard one to try to put together because linux has changed about six times on how it does its scheduling and how it does its process managing so yeah that'll be an interesting to cover hope to see you then and please like and subscribe and as always bye for now you
Up Next

Learn Go Programming: 2-Hour Beginner Crash Course Tutorial
@ZeroToMastery
16.7K views•2023-08-04

Introduction to Secure Multiparty Computation with Yehuda Lindell
@fhe_org
7.7K views•2021-02-04

HTTP Requests Explained: GET, POST, PUT, DELETE
@codecademy
103.1K views•2021-10-07

Enigma Machine Mechanics: WWII Encryption Explained
@JaredOwen
13.2M views•2021-12-11
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Computer Science












































