Container Runtimes: Namespaces, Cgroups & Go

Learning Goal: Designing and implementing a minimal container runtime from scratch in Go using Linux namespaces, control groups (cgroups), and overlay filesystems to master operating-system-level virtualization.

  • Prerequisites: Basic Linux command-line familiarity and general programming logic (variables, loops, basic functions). No prior systems programming experience is required.
  • Estimated Total Study Time: 32 Hours

Module 1: Linux Systems & Go Programming Foundations

This module establishes a deep mental model of how operating systems work under the hood. You will learn about processes, execution memory spaces, and the system call interface that allows user-space programs to talk directly to the Linux kernel. Simultaneously, you will build core competency in the Go programming language to write system-level utilities.

Recommended Videos

Why this video: This video provides a clear, animated conceptual baseline of what a "process" actually is inside memory. Before attempting to isolate processes using namespaces, you must understand their structural anatomy in RAM (text, data, heap, and stack space) and how the OS manages their lifecycle.


Why this video: All container runtimes rely on system calls to request isolation mechanics from the host kernel. This video details how system calls bridge the boundary between unprivileged user space and privileged kernel space, mapping out how programs interface with resource management layers.


Why this video: A comprehensive entry point for Go (Golang) programming. Writing a custom systems container runtime requires a strong handle on Go syntax, statically typed structures, slices, maps, error handling, and project compilation patterns, all covered in this crash course.


Knowledge Checkpoint

  • Explain the difference between User Space and Kernel Space and how system calls bridge them.
  • Identify the four primary memory areas of an active process (text, data, heap, stack).
  • Compile and run a basic Go program using variables, control structures, and system package imports.

Module 2: The Core of Isolation: Chroot and Filesystems

Here, you will transition from generic system concepts to the historical ancestor of container filesystem isolation: the chroot system call. You will understand how operating systems map virtual directory paths to physical hardware and how directory jails prevent processes from traversing unauthorized file hierarchies.

Recommended Videos

Why this video: This classic explainer reframes containers away from the "mini-virtual-machine" misconception. It maps out exactly how containers act as sandboxed processes on a shared host kernel, utilizing directory isolation and kernel controls as their foundational building blocks.


Why this video: This presentation explains the mechanisms of a chroot jail. You will understand how altering the root directory / pointer restricts standard filesystem access, while discovering why chroot alone is not secure enough to create a complete container runtime boundary.


Why this video: A practical, step-by-step tutorial demonstrating how to construct a directory, copy binary components (like Bash), link core shared libraries, and execute commands inside a custom chroot environment.


Knowledge Checkpoint

  • Create a custom local rootfs folder structure and trigger a shell inside it using chroot.
  • Explain how to find shared library dependencies of any compiled binary using the ldd command.
  • Describe the limitations of chroot security and why it cannot protect network ports or process lists.

Module 3: Linux Namespaces: Process Isolation

This module covers the modern system-level mechanism for process isolation: Linux Namespaces. You will explore how namespaces partition global host resources (processes, networks, users, and mount points) so that individual execution threads believe they have exclusive access to a clean machine.

Recommended Videos

Why this video: An in-depth tech talk that outlines the system-level mechanics of different namespaces (UTS, PID, NET, IPC, MNT, USER). This conceptual framework is necessary before diving into Go-based namespace code.


Why this video: Network isolation is a core component of container environments. This tutorial demonstrates how to isolate interfaces, manipulate route tables, and establish connections between separate namespaces using virtual ethernet (veth) pairs.


Why this video: This deep-dive video addresses a crucial gap when building container runtimes in Go. Because the Go runtime runs multiple system threads under the hood, standard container execution requires locking your thread to a specific namespace scope using runtime.LockOSThread(). This video explains these challenges and how runtimes manage them.


Knowledge Checkpoint

  • Name and define the responsibilities of the six primary namespaces: UTS, PID, MNT, NET, USER, and IPC.
  • Explain why standard multi-threaded Go execution requires thread pinning (runtime.LockOSThread()) when altering process namespaces.
  • Identify which namespace hides host processes, allowing a container-bound process to see itself as PID 1.

Module 4: Control Groups (cgroups): Resource Limiting

While namespaces hide system resources, Control Groups (cgroups) limit and control resource usage (such as CPU, memory, and disk I/O). This module focuses on configuring limits directly via the virtual filesystem, detailing the differences between cgroups v1 and v2.

Recommended Videos

Why this video: This video provides a detailed, practical walkthrough of how to set up resource limits. It demonstrates how to manually navigate /sys/fs/cgroup/, write limits into control files, register processes by PID, and verify active resource limits.


Why this video: Led by renowned Linux API expert Michael Kerrisk, this session details the differences between v1 and v2 structures. Understanding the unified hierarchy design of cgroups v2 is necessary for modern container system engineering.


Practical Walkthrough: Managing Cgroups in Go

Unlike traditional configuration files, cgroup limits are managed by writing to a virtual filesystem located under /sys/fs/cgroup. To limit memory usage for a sandboxed process in Go:

  1. Create a dedicated control folder: /sys/fs/cgroup/memory/my-runtime/.
  2. Write your process limit (e.g., in bytes) directly into the memory.limit_in_bytes (cgroup v1) or memory.max (cgroup v2) control file.
  3. Write the target sandboxed process ID (PID) into the tasks or cgroup.procs file inside that folder.

package main

import ( "os" "path/filepath" "strconv" )

func LimitMemory(pid int, limitBytes int64) error { cgroupPath := "/sys/fs/cgroup/memory/my-runtime" os.MkdirAll(cgroupPath, 0755)

// Set memory ceiling limit err := os.WriteFile(filepath.Join(cgroupPath, "memory.limit_in_bytes"), []byte(strconv.FormatInt(limitBytes, 10)), 0644) if err != nil { return err } // Assign target process to the group return os.WriteFile(filepath.Join(cgroupPath, "tasks"), []byte(strconv.Itoa(pid)), 0644)

}


Knowledge Checkpoint

  • Locate the active cgroup mount path on your Linux host system.
  • Create a memory limitation control group and assign a running shell to it manually.
  • Explain the architectural differences in control hierarchies between cgroups v1 and cgroups v2.

Module 5: Overlay Filesystems and Image Layers

This module covers union filesystems, which allow you to mount multiple separate directories as a single unified system. You will explore how OverlayFS implements container layering, using copy-on-write functionality to preserve base read-only templates.

Recommended Videos

Why this video: A clear visual breakdown of how read-only layers merge with write layers. It explains the mechanics of lowerdir, upperdir, workdir, and merged layers, which is crucial for building space-efficient container runtimes.


Why this video: This deep-dive video explores the technical details of the Linux kernel's OverlayFS driver, demonstrating copy-on-write and file deletion mechanics within active mounts.


Supplemental Study: Simulating OverlayFS Mounts

To understand the layering behavior before implementing it in Go, run these commands in a Linux terminal:

mkdir lower upper work merged echo "I am from base" > lower/base.txt echo "I am a config" > lower/config.txt

Mount the layers

sudo mount -t overlay overlay -o lowerdir=./lower,upperdir=./upper,workdir=./work ./merged

Verify file visibility

cat merged/base.txt

Modify a file inside merged layer

echo "Modified!" >> merged/config.txt

Verify Copy-On-Write (COW) behavior:

The base layer remains unchanged, while changes are captured in the upper directory.

cat lower/config.txt cat upper/config.txt


Knowledge Checkpoint

  • Define the roles of the lowerdir, upperdir, workdir, and merged layers.
  • Explain the copy-on-write (COW) mechanism and why it prevents write modifications from altering base system templates.
  • Mount a dynamic OverlayFS via the command-line interface, demonstrating file creations and modifications.

Module 6: Capstone: Assembling the Go Container Runtime

In this capstone module, you will synthesize everything you have learned. You will combine namespaces, chroot execution jails, cgroup limits, and overlay mounts inside a single compilable Go application.

Recommended Videos

Why this video: Liz Rice's classic live-coding demonstration. She builds a functional container sandbox in Go from scratch in under 100 lines of code, showing exactly how namespace clone parameters and filesystem changes work in practice.


Why this video: This deep dive connects the custom runtime model with industry standards, such as the OCI Runtime Specification, runc, and modern orchestrators.


Step-by-Step Runtime Design Pattern

To build your custom container runtime in Go, structure your program around a re-execution model:

[Parent CLI Execution (run)] │ ▼ (Spawns a duplicate child process with namespace flags enabled) [Child Process Setup (child)] │ ├─► Pin executing thread using LockOSThread() ├─► Apply control limits inside /sys/fs/cgroup ├─► Mount merged OverlayFS layer ├─► Run syscall.Chroot() and PivotRoot to isolation rootfs ├─► Mount virtual filesystems (/proc) │ ▼ (Executes targeted executable command inside jail) [Target Application Execution]

  1. Re-execution Orchestration: You cannot safely apply namespaces directly to a running multi-threaded Go process. Instead, have your initial execution (run) spawn a copy of itself (/proc/self/exe) passing a hidden argument (child).
  2. Clone Configuration: Configure the spawned process with isolation options inside the SysProcAttr attributes struct, using namespace flags like CLONE_NEWUTS, CLONE_NEWPID, and CLONE_NEWNS.
  3. Sandbox Configuration: Within the isolated child execution block:
    • Mount your pre-assembled OverlayFS directories.
    • Lock execution paths to a local virtual root path using syscall.Chroot().
    • Register the child PID (os.Getpid()) inside your customized /sys/fs/cgroup limits.
    • Replace the child shell execution thread using syscall.Exec().

Knowledge Checkpoint

  • Code a complete multi-phase re-execution workflow (/proc/self/exe) in Go.
  • Safely mount /proc inside the containerized process jail, ensuring running processes are displayed correctly.
  • Test and run your custom container binary, confirming hostname, PID, and memory isolation.

Course Map


Key People Index

  • Liz Rice
    Chief Open Source Officer at Isovalent and former CNCF Governing Board Chair. Her "Containers from Scratch" talks helped demystify low-level systems programming concepts for application developers worldwide.
  • Michael Kerrisk
    Author of The Linux Programming Interface and long-time maintainer of the Linux man-pages project. He is a leading authority on namespace and cgroup implementations.
  • Solomon Hykes
    Co-founder of Docker. His early integration of namespaces, chroot, and Union Filesystems into a unified package developer workflow helped popularize modern container standards.

Final Self-Assessment

  • I can explain the fundamental differences between virtual machines and containers.
  • I can write and compile system-level utilities in Go using standard library packages.
  • I can describe how the system call interface acts as a bridge between user-space code and the Linux kernel.
  • I can build a functional, isolated application root filesystem manually using chroot.
  • I can explain the individual isolation boundaries provided by the UTS, PID, NET, and Mount namespaces.
  • I understand why Go runtimes require runtime.LockOSThread() when setting up container namespaces.
  • I can configure control limits (like CPU and memory ceilings) inside the /sys/fs/cgroup filesystem.
  • I can explain the differences in control hierarchies between cgroups v1 and cgroups v2.
  • I can configure a dynamic Union Filesystem mount using the Linux OverlayFS driver.
  • I can explain how copy-on-write (COW) preserves base image filesystems when a containerized process writes to a file.
  • I can write a multi-stage Go program that re-executes itself via /proc/self/exe to safely isolate child namespaces.
  • I have successfully designed, built, and executed a custom container runtime from scratch in Go.
Explore Further

Related Computer Science Roadmaps

View All→