The Go runtime's inability to safely fork processes limits container operations on Linux, but using runtime.LockOSThread allows Go programs to manipulate thread-specific execution contexts (like file system information, mount namespaces, and network namespaces) by binding go routines to OS threads, enabling techniques such as changing the current working directory, creating isolated mount namespaces with pivot_root, and entering container network namespaces without affecting other threads.
Optimizing Docker Container Startup with Go Runtime LockOSThread
Added:uh hello everyone my name is Corey I am a software engineer at morantis and a maintainer on the Moby project the Upstream for the docker engine today I'll be sharing one Technique we use to make some Docker container operations faster and less resource intensive on Linux and explain how you could use the same techniques in your own go programs I like to start with a motivating example increasing the limit on how many container container image layers can be mounted Moby typically uses the overlay file system to compose the layers of a container image when mounting an overlay FS The Source argument is ignored and the source directory paths are passed in through the mount options there is no hard limit on the number of layers in an oci image but how many layers can we Mount at a time with overlayfs well file system specific Mount options are passed into the mount system called through that data argument note how there's no argument for the size of it of the data how then does the kernel know how much data to copy into kernel space while overlayfs in particular expects data to be a null terminated string the Linux ABI doesn't actually require that a B1 and besides even if it did the kernel can't just trust that user space is passing a valid string of a reasonably small length so in actuality the kernel copies one page of data which in practice is four kilobytes taking into account the null Terminator we have 4095 bytes of options which you know for a traditional file system where you've got like options like no exec that's fine but that's pretty restrictive for overlay FS there is a new set of file system creation context syscalls uh new and Linux 5.2 which lifts its limitation unfortunately we still have to support older kernels for now so we're stuck with the amount syscall and its limitations one way to increase the number of layers we can squeeze into one Mount option string is to reduce the redundancy so the directory paths in those upbound options can be made relative to the process's current working directory so we can change our working directory to the common prefix and squeeze a few more layers in now changing the working directory temporarily is like just fine in a single threaded process but Docker D is multi-threaded and the current working directory is global State shared by all threads so changing it would affect every open call and every other thread until it's changed back and have two threads tried to concurrently Mount uh Mount container file systems they would collide now the Mobi project historically took the approach of starting itself as a child process and then the child which changes working directory Mountain Exit now starting a whole new GoPros process from scratch dominates the time to issue the mount syscall and we have to deal with moving data and results across a process boundary I mean there's got to be a better way Let's uh first dig a bit deeper into how processes are started launching a new program as a sub-process is a multi-step procedure first you call Fork which duplicates the calling thread in a copy of its memory space now forking is fairly cheap on Linux because memory is copy on right next the child process sets up the execution context such as change in the current working directory and finally the child will Call Exec ve to replace itself with the new program well that child process doesn't have to exec it could keep running the same process same program as the parent until it exits so that's one possible solution for a child process whose only job is uh to change directory Mount and exit unfortunately Dr Damon's written and go and Fork without exec is not supported and go programs you could invoke the raw syscall if you want but the child process is going to be in a very sad State the child will inherit copies of all the mutexes in the parents but we'll start with just one thread none of the garbage collectors threads will be running and the child will likely deadlock rather quickly on one of those mutexes go programs are able to spawn new child processes but the run time makes all the arrangements to Fork an exec on your behalf it needs to a lot of Preparatory work to make it safe and reliable and the details which are deeply tied to those runtime internals so unless you are the go run time or are willing you to tie yourself to the implementation details of a particular runtime write code in an extremely limited dialect of go and pray that tool chain updates won't break you you cannot fork and go programs without also exacting now while go programs may not be able to cheaply Fork off child processes they do have an abundance of threads now how come changing the current working directory in one thread affects the current working directory of all the other threads well in a word because posic says so in a more practical sense the current working directory is shared because the threading library or in ghost case the language runtime has instructed the kernel to make it that way threads are spawned using the Clone system call which compared to Fork gives the caller more precise control over what is and is not shared between the caller and the child clone can also be used to spawn processes as there's not much distinction between processes and threads my threads just a process that shares the thread group ID virtual memory space and Signal handlers with other threads in the process other pieces of execution context can be shared but don't actually have to be under Linux uh so for instance if the Clone FS flag is passed to the Clone syscall the calling process and child process share the same file system information which encompasses the file system root the umask and the current working directory otherwise the child gets a copy now most of the process execution contexts that can be shared using clone can be unshared using the appropriately named unshare syscall thread can call unshare with the clonefest flag to then reverse the effects of Clone disassociating its file system information from that of the other threads now note that there is no way to reassociate the threads file system information afterwards you may be wondering how unshare can be used in go programs as threads aren't exposed to application code all application code runs in go routines which do not map one to one onto threads the runtime schedules go routines onto a pool of threads not entirely unlike how the kernel schedules threads onto CPU cores if a routine blocks waiting on some IO receiving on a channel requiring a mutex uh or simply if the runtime decides to preempt that go routine because it's been running for too long uh runtime may go and schedule some other go routine onto that same thread and different grow routines may run on the same thread at different times and any particular grow routine may run on different threads throughout his lifetime uh normally it does not matter that a girl routine May suddenly find itself running on a different thread as aside from having different threat IDs all the threads are practically identical well unsharing parts of a threads execution context makes the thread different from the others it would cause chaos if uh random go routines were to be scheduled onto such an unshared thread for example the go routine that wanted to change just its own working directory could unexpectedly find its working directory reverted and then some other go routine would see the change working directory all at the whims of the runtime thankfully go has a solution for this runtime.lock OS thread wires the calling go routine to its current thread until an equal number of calls are made to unlock OS thread the call and go routine will always execute in that thread exclusively since unsharing a threads file system information is irreversible no other girl routine can ever be allowed to be scheduled to run on that unchaired thread thankfully go also has a solution for this you simply return from the go routine function without unlocking it from the thread and the runtime will terminate the thread and eventually spawn a new one to replace it this is roughly what uh changing the working directory to mount looks like minus any error handling you spawn a new grow routine for this operation and lock it to a thread unsure the file system information that way you can simply go change the working directory Mount and return the ability to wire go routines to threads makes it possible to do things in go programs which could not be done in any other way I'll take you through a few other examples of how it's used within Mobi sanitization it's really hard to get right from user space the kernel can do a much better job especially because it can do it atomically the open at 2 syscall makes it easy to guard against past reversal attacks though that's only available from Linux five six in order to support older kernels Moby takes a different approach sandboxing the thread so it cannot open paths outside of where it's allowed to for use cases like hours such as untowering image layers where we don't need to sandbox arbitrary untrusted code CH root is arguably perfectly adequate when used for a memory safe language uh the root directory is part of the thread file system information so I'm sharing it makes CH root calls thread local in addition to chdir unfortunately using CH root makes Moby incompatible with gr security kernels because those kernels block chmod and make node in ch rated threads we work around this by instead using pivot root to change the root mode of the current Mountain namespace which is a much more robust uh sandboxing mechanism as well but we can't safely modify the existing mounting space as it could be shared by many other processes not to mention the other threads so we call unshare with the Clone new NS flag which moves the thread into a new Mount namespace which is initialized to a copy of the previous now we're free to mount on Mount and pivot mounts to our hearts content without affecting the amount table of any other thread or any other process another use for wiring go routines to threads is to enter a container's network namespace the only information you need to know to access the network namespace created for a container is the containers process ID Moby enters the container Network namespaces to provide a DNS resolver over the container loopback interface which can resolve the private addresses of other containers and to forward DNS queries from one container to a DNS server running in another container the set an S system call is used to move the calling thread to the namespace reference by a file descriptor unlike a more traditional process State such as the file system information I spoke about earlier uh if the thread can be moved back to its starting namespace with another call to set an s and a thread which has had its name space is restored is indistinguishable from threads which were never moved at all and so can be reused by the go runtime for other go routines a combination of unshare and set an S can also be used to cheaply create a new network namespace for example as I'm demonstrating here manipulating the execution context of threads in a language which hide threads from the application is not always going to be easy there are sharp edges and gotchas which you need to be aware of if you want to apply these techniques to your own go programs you may find unexpected and even impossible behaviors in completely unrelated parts of your application if you get things wrong the go run time and most go code assumes quite reasonably that all execution contacts are made equal that they all have the same file descriptive table view of the file system uid GID network interfaces Etc if you violate the invariant that all unlocked LS threads are fungible you're going to have a bad time make sure to always lock your go routine to a thread before manipulating its threads execution context and only unlock after you've put the thread back exactly the way you found it which may not always be possible so when in doubt keep the go routine locked to the thread and let the runtime terminate it when writing code which opens handles to the thread's original namespaces make sure to lock the go routine before opening handles to its original namespaces and open from that go routine otherwise your code might restore the thread to the wrong namespace I've done this and it was not a fun bug to chase down the initial threat of a process known as the as the thread group leader grow routines can be scheduled onto it same as any other thread this is important to keep in mind when modifying the execution context of your program's threads because the proc self-magic link refers to the thread group leader not the current thread this can trip you up in a couple of ways unless your grow routine happens to be locked to the thread group leader the files and proc itself are not going to reflect the unshared state of the threader grow routine is executing on when writing code which opens proc files for an unshared thread make sure to open the files for the for that particular thread use the proc self task directory for the current thread ID or the proc thread self magic link on Linux 317 and above remember to lock the go routine to the thread first so the thread doesn't change underneath you and to avoid avoid any surprises with code and with external processes which aren't prepared for your unshared shenanigans I recommend that you lock the main go routine to the thread group leader go guarantees that init functions will run on the thread group leader and also that main will also be executed on the leader if lock oestro has been called from an init function so long as the thread leader thread is left alone any code in your process running on unlocker routines can continue to open files through proc self and get the expected results as the threads they're running on we'll be sharing all the execution context with the leader and the last gotcha I want to talk about is the parent death signal option when starting a sub process you can instruct the kernel to send it a signal if the parent dies it's very handy for ensuring you don't leak sub process if your process crashes for example however the current the kernel considers the parent to be the thread which started the sub-process if some routine go routine which locks and exits gets scheduled onto the same thread which you had previously used to start a sub-process your sub process will get signaled seemingly at random you can guard against this by locking the grow routine you will be starting the subprocess from to its thread and not unlocking it until after the subprocess exits thank you [Applause] thank you
Up Next

How to Set CPU and Memory Limits Using Linux Control Groups (cgroups)
@onprema
616 views•2025-07-13

Introduction to Secure Multiparty Computation with Yehuda Lindell
@fhe_org
7.7K views•2021-02-04

HTTP Requests Explained: GET, POST, PUT, DELETE
@codecademy
103.1K views•2021-10-07

Enigma Machine Mechanics: WWII Encryption Explained
@JaredOwen
13.2M views•2021-12-11
Related Study Plans & Knowledge Roadmaps
Structured learning paths in Computer Science







![[Ch-03] Linux Namespace 완전 정복 | 컨테이너 격리의 핵심 기술](https://i.ytimg.com/vi/oMKB93KuPsk/maxresdefault.jpg)






























