teach me these concepts for the interview properly: 1. Signals, Interrupts, IVT?
SIGINT,
SIGTERM. kill sends which signal? 2. Process memory layout? 3. Stack overflow, heap
overflow? 4. How system call goes from user from kernel space?
1) Signals, Interrupts, IVT. SIGINT, SIGTERM. kill
sends which signal?
A) Signals (Unix/Linux)
What they are: lightweight asynchronous notifications delivered by the kernel to a
process/thread to indicate an event (user pressed Ctrl+C, child exited, segmentation fault, timer
fired, etc.).
Key properties
● A signal has a number + name (SIGINT, SIGTERM, SIGKILL, …).
● A process can:
○ handle it (custom handler),
○ ignore it (some signals),
○ or take the default action (terminate, stop, core dump).
● Signals can be pending and (for non-real-time signals) may be coalesced (multiple
occurrences can merge).
Common signals
● SIGINT (2): “interrupt” from terminal (usually Ctrl+C). Default: terminate.
● SIGTERM (15): “polite terminate” request. Default: terminate. Can be caught → graceful
shutdown.
● SIGKILL (9): cannot be caught/ignored. Immediate kill.
● SIGSTOP: cannot be caught/ignored. Stops process.
B) Interrupts (hardware/CPU)
What they are: events that interrupt CPU execution, usually from hardware (timer tick, NIC
packet arrival, disk completion).
Why it matters: interrupts drive kernel work like scheduling and network I/O.
Hardware interrupt vs signal
● Interrupt: hardware → CPU/kernel (low level, preempts CPU).
● Signal: kernel → process (higher level, delivered when process runs or returns to user
mode).
C) IVT (Interrupt Vector Table) / IDT
On x86:
● Modern systems use IDT (Interrupt Descriptor Table) rather than the old “IVT” concept.
● The idea: an interrupt vector number indexes a table entry telling the CPU which
handler to run (address + privilege checks).
Interview line: “Interrupts are dispatched via an interrupt vector (IDT); signals are a kernel
mechanism delivered to processes.”
2) Process memory layout (Linux user-space virtual
memory)
A typical process virtual address space (high-level):
High addresses
● Stack (grows downward)
● Memory-mapped region (shared libs, mmap() files, anonymous mappings)
● Heap (grows upward via brk/sbrk, also mmap allocations)
● .bss (zero-initialized globals)
● .data (initialized globals)
● .text (code / instructions)
Low addresses
Also important:
● Each process has its own virtual address space.
● The kernel uses page tables to map virtual pages to physical frames.
● ASLR randomizes locations of stack/heap/libs to reduce exploitability.
Interview-ready explanation:
“Text/data/bss are from the ELF binary, heap is dynamic allocation (malloc), stack is call frames,
mmap area holds shared libraries and mapped files.”
3) Stack overflow vs heap overflow
These are two different “bad writes” with different causes.
Stack overflow
What it is: stack grows beyond its limit (or you corrupt stack memory).
Common causes
● Deep or infinite recursion
● Very large local variables (big arrays on stack)
● Unbounded stack allocation (alloca misuse)
Symptoms
● Segfault (SIGSEGV) or abort
● Crash happens at/near function calls/returns due to smashed stack frames
Heap overflow
What it is: writing past the bounds of a heap allocation (allocated by malloc/new).
Common causes
● Buffer overflow on dynamically allocated arrays
● Use-after-free leading to heap metadata corruption
● Off-by-one indexing
●
Clean interview contrast
● Stack overflow: “call frames / recursion / local vars”
● Heap overflow: “malloc’ed buffers / allocator metadata / lifetime bugs”
4) How a system call goes from user space to
kernel space
A syscall is the controlled gateway from unprivileged user mode to privileged kernel mode.
Step-by-step (conceptual but correct)
1. User code calls a libc wrapper
○ e.g., read(fd, buf, n) in C calls a libc stub.
2. The stub loads:
○ syscall number (which syscall)
○ arguments into registers (ABI-defined)
3. It executes a CPU instruction to enter kernel mode:
○ On x86-64: syscall
○ (older: int 0x80)
4. CPU switches privilege level:
○ user mode → kernel mode
○ switches to kernel stack for that thread
5. Kernel syscall entry code:
○ saves registers, validates context
○ dispatches using the syscall table to sys_read/ksys_read etc.
6. Kernel performs work:
○ validates file descriptor, permissions
○ copies data safely using copy_to_user / copy_from_user
○ may block (sleep) if I/O needed
7. Kernel returns:
○ sets return value / errno in registers
○ executes sysret/iret path back to user mode
8. libc translates negative return codes to errno for the user program (depending on API).
Important interview points
● Kernel must treat user pointers as untrusted → uses safe copy helpers.
● Syscalls can sleep (block) → thread state changes (S or D).
● Context switch can happen while blocked, later resumed.
30-second “good answer” summary (memorize)
“Signals are kernel-delivered asynchronous notifications to processes; interrupts are
hardware/CPU events dispatched via the interrupt descriptor table. SIGINT is Ctrl+C, SIGTERM
is the default graceful stop; kill pid sends SIGTERM by default, kill -9 sends SIGKILL. A
process has a virtual address space: text/data/bss, heap for dynamic allocation, mmap region
for libs/files, and a stack for call frames. Stack overflow is usually deep recursion or huge locals;
heap overflow is writing past malloc’d buffers and can corrupt allocator metadata. Syscalls go
through a libc stub that loads args and syscall number, executes syscall to enter kernel
mode, kernel validates and dispatches, does work (may block), then returns to user with a
result/errno.”
teach me this now: “What happens during the boot process from the moment you turn on the
machine until you get a login prompt?”
0) Power on → CPU reset vector
● You press power.
● Power supply stabilizes, motherboard logic releases CPU from reset.
● CPU begins executing from a fixed address called the reset vector
(architecture-specific).
1) Firmware stage: BIOS/UEFI
BIOS / UEFI responsibilities
● POST (Power-On Self Test): checks RAM, CPU, basic devices.
● Initializes hardware enough to boot (memory controller, buses, storage, display).
● Builds a hardware description for the OS:
○ BIOS: legacy tables
○ UEFI: ACPI tables, SMBIOS, etc.
UEFI vs BIOS differences (interview-ready)
● UEFI can read FAT filesystem on the EFI System Partition (ESP) and load .efi
bootloaders.
● BIOS typically loads the first-stage bootloader from the MBR (very small) then chains.
At the end of this stage, firmware selects a boot device (NVMe/SSD/HDD/network) per boot
order.
2) Bootloader stage (GRUB/systemd-boot/other)
What the bootloader does
● Presents boot menu (optional).
● Loads into memory:
1. the Linux kernel image (vmlinuz)
2. the initramfs/initrd (temporary root filesystem)
● Passes kernel command line parameters (e.g., root device, debug flags).
● Transfers control to the kernel entry point.
In UEFI systems:
● Bootloader is often an .efi executable stored on the ESP.
In BIOS systems:
● Bootloader starts from MBR/boot sectors and loads the rest of itself.
3) Kernel early initialization
Once control jumps to the kernel:
A) Decompression and early setup
● Kernel decompresses itself.
● Sets up:
○ CPU modes, descriptor tables (IDT/GDT), basic paging
○ early memory allocator
○ early console (so you can see boot logs)
B) Detect hardware and drivers
● Parses firmware tables (ACPI, device tree on some platforms).
● Initializes core subsystems:
○ scheduler, timekeeping, interrupts
○ memory management (page tables, zones)
○ VFS, block layer, networking basics
● Loads/binds built-in drivers and prepares for modules.
C) Mount temporary root from initramfs
● Kernel mounts initramfs as the initial root filesystem.
● Then it runs the first user-space program:
○ historically /init from initramfs
○ on many systems it ultimately leads to systemd as PID 1
If the kernel cannot find/init /init, you get a kernel panic like “No init found”.
4) initramfs stage (early user space)
initramfs exists because the kernel may not yet have the drivers/config to mount the real root
filesystem directly.
Common tasks performed here:
● Load kernel modules required to access storage (NVMe/SATA/RAID/HBA drivers).
● Assemble RAID (mdadm) or activate LVM volume groups.
● Unlock encrypted disks (LUKS) (prompts for passphrase).
● Discover the real root filesystem location.
● Mount the real root (e.g., /dev/mapper/vg-root).
● switch_root / pivot_root from initramfs to the real root FS.
If this fails, you often drop into an initramfs shell (dracut/busybox) for rescue.
5) PID 1: init system starts (usually systemd)
After pivoting to the real root:
● Kernel executes the configured init process (commonly /sbin/init).
● On most modern distros, that is systemd, which becomes PID 1.
What PID 1 does
● Mounts/validates filesystems (often via /etc/fstab or systemd units).
● Sets hostname, time, sysctl tunables.
● Starts udev/device management (now integrated with systemd).
● Brings up networking (depends on distro: systemd-networkd, NetworkManager).
● Starts daemons and services according to the boot target:
○ [Link] (text / server mode)
○ [Link] (GUI)
Key concept: systemd starts many units in parallel, but with dependency ordering.
How to answer in an interview (a clean
45-second narrative)
“Power on triggers firmware (BIOS/UEFI) which does POST, initializes hardware, and selects a
boot device. Firmware loads a bootloader like GRUB or systemd-boot. The bootloader loads the
kernel image and initramfs into memory and passes kernel parameters. The kernel
decompresses, initializes memory, interrupts, scheduler, and device drivers, then mounts
initramfs and runs /init. Initramfs loads storage modules, assembles RAID/LVM, unlocks
encryption, mounts the real root filesystem, then switches root. Finally PID 1—usually
systemd—starts, mounts filesystems, brings up udev and networking, launches services, and
starts getty or a display manager, which gives you the login prompt.”
Common follow-ups interviewers ask (and
short answers)
“Why do we need initramfs?”
Because the kernel might not have the drivers/config to access the root filesystem yet
(RAID/LVM/encryption/network root). initramfs provides early userspace to set that up.
“What is PID 1 special?”
It’s the first userspace process. It inherits orphaned processes, handles reaping, and
orchestrates service startup/shutdown.
“Where do kernel modules load?”
Some are built-in. Others load during initramfs or later via udev/systemd when devices appear.
“How do you debug boot failures?”
● Firmware/bootloader issues: UEFI boot entries, GRUB config.
● Kernel/initramfs: kernel logs (dmesg), initramfs shell, missing storage drivers, wrong
root=.
● systemd: journalctl -b, systemctl --failed, [Link].
“If you had a program that needed 1TB of RAM and you only have 16GB, how would the Linux
system allocate memory?”“What is the swap area, regarding memory?”“Tell me something
about swap.”“What happens when a Linux machine runs out of memory?”“Explain how virtual
memory works in Linux using the example of a process trying to read data in a specific block of
memory.” Scenario question: process fails with Exception: Failed to allocate 1GB... even though
free -h shows 1GB free and swapping is disabled—explain how fragmentation and contiguous
allocation constraints can cause this.“An application is running slowly, and investigation reveals
it is heavily swapping and this is impacting performance. How would you resolve this issue?”
Below are interview-grade answers with the correct Linux mental model (virtual memory,
reclaim, swap, OOM, fragmentation), plus what to check and what to do.
1) “Program needs 1TB RAM, machine has 16GB. How
does Linux allocate memory?”
Linux does not allocate “RAM” directly. It allocates virtual memory first.
What happens
1. The process requests memory (e.g., malloc(1TB)).
2. Linux usually creates a virtual address space mapping for that size:
○ either expands the heap (brk) or creates an anonymous mapping (mmap).
3. Initially, most of that memory is not backed by physical RAM. It’s just address space.
4. Physical memory is committed later, when the process touches pages (page faults).
○ This is demand paging.
The key concept: overcommit
Linux can allow allocating more virtual memory than physical RAM+swap:
● Overcommit enabled (common default): malloc() can succeed even if you can’t
truly back it; failure might happen later when pages are touched.
● Strict accounting mode: allocation may fail earlier based on commit limits.
What decides success:
● Overcommit policy (vm.overcommit_memory)
● Available RAM
● Swap presence (adds commit backing)
● cgroup limits (in containers)
2) “What is swap area regarding memory?”
Swap is disk space reserved to hold memory pages when they can’t stay in RAM.
● It stores anonymous memory (heap/stack) pages that are not file-backed.
● File-backed pages (like program code or mmap’ed files) can be dropped and reloaded
from the file; anonymous pages must be written somewhere → swap.
Swap can be a partition or a swapfile.
3) “Tell me something about swap.”
High-signal interview points:
● Swap increases the commit limit and helps avoid immediate OOM by providing backing
store.
● Swap is orders of magnitude slower than RAM → heavy swapping causes latency
spikes.
● Linux does not “swap everything”; it tries to evict page cache first, then swap out
anonymous pages if needed.
● Swappiness ([Link]) influences preference to swap vs reclaim page cache.
● Some systems run fine with small swap; others (e.g., memory spikes) benefit from
having swap as a safety valve.
4) “What happens when a Linux machine runs out of
memory?”
There’s a sequence:
A) Reclaim starts (memory pressure)
● Kernel tries to free memory by reclaiming:
○ dropping clean file cache
○ writing dirty pages back
○ swapping out anonymous pages (if swap exists)
This can cause direct reclaim stalls → app slows down before OOM.
B) If swap exists, swapping increases
● If the system is paging hard, you see high swap-in/out and major faults.
● Latency can become terrible (thrash).
C) If it cannot free enough: OOM killer
● Kernel triggers OOM killer and kills a process to free memory.
● In containers, often it’s cgroup OOM (container limit), not global host OOM.
● Evidence is in kernel logs: dmesg / journalctl -k.
5) “Explain virtual memory using a process reading a
block of memory.”
Example: process reads address 0x7f...:
1. CPU issues a load instruction referencing a virtual address.
2. MMU consults the TLB; if not present, it walks page tables.
3. If the virtual page isn’t mapped → page fault trap to kernel.
4. Kernel checks:
○ Is this address valid in a VMA? permissions (readable)?
5. If valid:
○ If the page exists in RAM (minor fault) → map it, return.
○ If needs disk read (major fault):
■ file-backed: read from file into page cache
■ anonymous swapped out: read from swap
6. Kernel updates page tables, returns to user mode; instruction retries and succeeds.
This is why access can be fast (cache hit) or slow (major fault).
6) Scenario: “Failed to allocate 1GB even though free
-h shows 1GB free and swap is disabled” (fragmentation
+ contiguous constraints)
This is a great kernel interview trap because “free memory” is not the same as “allocatable in
the way you need”.
Two major reasons
(A) The allocation requires contiguous physical memory
Many things in Linux need physically contiguous chunks, especially in the kernel:
● DMA buffers (device I/O)
● huge pages / THP under some conditions
● certain driver allocations
● mlock/pinned pages for devices
Even if you have 1GB free total, it might be fragmented into small pieces across zones, so a
large contiguous allocation fails.
How fragmentation happens
● Long uptime + many allocations/frees
● Page cache + slab allocations + pinned pages
● No swap (less flexibility for compaction/reclaim)
Kernel behaviors involved
● Buddy allocator hands out contiguous blocks; large orders fail if memory is fragmented.
● Kernel may try memory compaction, but compaction can fail if many pages are
pinned/unevictable.
What you’d check
● dmesg -T for allocation failures: “order X allocation failed”
● /proc/buddyinfo (fragmentation by order)
● /proc/pagetypeinfo (movable/unmovable fragmentation)
● THP: cat /sys/kernel/mm/transparent_hugepage/enabled
(B) User-space “contiguous” or “commit” constraints
For normal malloc(1GB) in user space, physical contiguity usually isn’t required (it can be
many pages). But you can still fail if:
● commit limit is reached and overcommit is strict (vm.overcommit_memory=2) and
swap is disabled
● cgroup limit is hit (container memory limit) even if host shows free
● allocator cannot find a sufficiently large virtual contiguous range (rare on 64-bit, more
plausible on 32-bit or with address space fragmentation)
● the request is for hugepages (explicit hugetlbfs) which require specific availability
Interview line to use
“free -h shows aggregate free RAM, but allocations can fail due to contiguity, fragmentation,
reclaimability, cgroup limits, or strict overcommit. With swap disabled, the kernel has fewer
options to reclaim/compact.”
7) “App slow due to heavy swapping—how do you
resolve?”
First: confirm it’s truly swapping and find the root cause (leak vs working set > RAM vs
misconfig).
Step 1: Confirm swapping and its cost
● vmstat 1 → look at si/so (swap in/out)
● sar -B 1 (paging), if available
● free -h → swap used and “available”
● top / htop → processes, RES, SHR; but don’t trust it alone
● pidstat -r 1 and/or pidstat -p <pid> -r 1 for faults
Step 2: Identify who is causing it
● ps -eo pid,comm,rss,vsz --sort=-rss | head
● smem (if available) for proportional memory
● For JVM: check heap settings/GC logs
Step 3: Fix options (short-term → long-term)
Short-term mitigations
● Restart the worst offender (if leak suspected)
● Reduce concurrency / load shedding to lower working set
● Drop caches only if you understand impact (usually not the real fix)
● If safe: add swap as a pressure relief valve (even small swap can prevent OOM)
● Disable THP for some workloads if it worsens reclaim latency (workload-specific)
Tuning (be careful in interviews: explain tradeoffs)
● Adjust [Link] (lower to prefer keeping anon in RAM; but can increase risk of
OOM)
● Tune dirty writeback (vm.dirty_*) to avoid reclaim stalls
● Ensure memory limits aren’t too tight (cgroups)
Long-term fixes
● Add RAM or move workload to bigger instance
● Fix memory leak / reduce memory footprint
● Use caching properly (bounded caches)
● Use better data structures / streaming rather than loading everything
● For DB/search: tune cache sizes and working set, split/shard workloads
Key interview point
If the working set truly exceeds RAM, you either:
● reduce working set, or
● add memory, or
● accept swapping (usually unacceptable for latency-sensitive services).
“Meta PE/kernel” style 20–30 second answers you can
memorize
Why can malloc(1GB) fail with “1GB free”?
“Because free shows aggregate free RAM; allocation can fail due to fragmentation,
non-reclaimable memory, strict overcommit without swap, or cgroup limits. If a contiguous
allocation is required (kernel/DMA/hugepages), fragmentation can prevent getting a big enough
block.”
What happens when memory runs out?
“Kernel reclaims page cache, writes dirty pages, swaps anonymous pages if swap exists; if it
still can’t satisfy allocations, it triggers OOM killer (global or cgroup) and kills a process.”
teach me this topic properly for the interview. 3.8 Containers, cgroups, Namespaces, and
Services More recent guides (2024–25) explicitly mention containerization, cgroups, and
namespaces. Core types Conceptual containerization. Namespaces & cgroups roles. Turning a
process into a service. Representative questions “Explain containerization.” “It’s worth
mentioning cgroups and namespaces for containerisation,” Prepfully’s question bank
notes.“How do you make a process a service?”Expect to discuss: How containers are
implemented via namespaces (PID, NET, MNT, UTS, IPC, USER) and cgroups for resource
limits. init systems (systemd), unit files, and service management.
Thought for 26s
Below is a kernel-accurate, interview-oriented explanation of containers, namespaces,
cgroups, and services (systemd)—the exact concepts you need to answer questions like
“Explain containerization” and “How do you make a process a service?”
1) What “containerization” actually means
Containerization = OS-level virtualization.
A container is not a VM. In a VM you virtualize hardware and run a separate kernel. In a
container you share the host kernel, but give a process the illusion that it has its own machine.
So conceptually, a container is:
● a process (or group of processes)
● running with isolation (namespaces)
● and resource control (cgroups)
● plus a filesystem view (rootfs + mounts, often overlayfs)
● plus security restrictions (capabilities/seccomp/LSM) commonly used in practice.
Interview one-liner:
“A container is a regular Linux process started with namespaces for isolation and cgroups for
limits/accounting, usually with its own root filesystem and reduced privileges.”
2) Namespaces: “What you can see”
Namespaces isolate views of system resources. Think: the process sees a different world.
Core namespace types (must know)
1. PID namespace
● Process IDs are scoped; a process can be PID 1 inside the container.
● Big implication: PID 1 behavior (signal handling, zombie reaping) matters.
2. NET namespace
● Separate network stack: interfaces, routing table, iptables/nft, ports.
● That’s how containers get their own “eth0”, IP, etc.
3. MNT (mount) namespace
● Separate mount table. Containers can have different rootfs and mount points.
4. UTS namespace
● Hostname/domainname isolation (uname, hostname).
5. IPC namespace
● SysV IPC and POSIX message queues isolation.
6. USER namespace
● UID/GID mapping: a process can be “root” inside while being an unprivileged UID
outside (rootless containers).
How the kernel implements this (the detail interviewers like)
● New namespaces are created with clone() flags (e.g., CLONE_NEWPID,
CLONE_NEWNET) or unshare().
● You can join an existing namespace with setns().
● Visible at: /proc/<pid>/ns/
Useful debug commands
● lsns (show namespaces)
● readlink /proc/<pid>/ns/* (what namespaces a process is in)
3) cgroups: “What you can use”
cgroups (control groups) are about resource accounting and limits.
What cgroups do
● Limit resources: memory, CPU, pids, IO
● Account usage: how much CPU time, memory current/peak, IO
● Control behavior: throttling, OOM handling inside a group
Common controllers you should name
● CPU: shares/quotas, throttling
● Memory: hard limits, OOM within the cgroup
● IO: throttle/weight on block devices
● PIDs: cap number of processes/threads ([Link])
● cpuset: pin to CPUs/NUMA nodes (common in performance workloads)
cgroups v1 vs v2 (short, practical explanation)
● v1: multiple hierarchies, per-controller mount points, complexity.
● v2: unified hierarchy, cleaner semantics (common on modern distros).
Debug commands
● cat /proc/<pid>/cgroup
● systemd-managed: systemd-cgls, systemd-cgtop
● cgroup v2 files live under /sys/fs/cgroup/... (e.g., [Link],
[Link], [Link])
Interview gotcha (important)
● A container can OOM even if the host has free memory because the cgroup memory
limit was hit. That’s “cgroup OOM,” not global OOM.
4) How containers are actually created (the “runtime”
story)
Most modern container stacks follow the OCI model (Docker/containerd/CRI-O → runc).
A typical container start sequence:
1. Create a cgroup and configure limits ([Link], [Link], [Link]…).
2. Create namespaces (pid/net/mnt/uts/ipc/user…).
3. Set up filesystem:
○ mount the container rootfs (often overlayfs layers)
○ pivot_root()/chroot() into it
4. Set identity & security:
○ uid/gid mapping (user namespaces)
○ drop Linux capabilities
○ apply seccomp filter (syscall allowlist)
○ apply SELinux/AppArmor labels if enabled
5. execve() the container process.
Key phrase for interviews:
“Containers are processes with namespace isolation + cgroup limits, plus a prepared root
filesystem and reduced privileges.”
5) “How do you make a process a service?”
A “service” is a process that is:
● started automatically (boot / on demand)
● supervised (restarted on failure)
● has logging, environment, dependencies, resource limits
● integrated with the init system (usually systemd)
systemd basics (what to say)
● On most distros, systemd is PID 1.
● It starts and manages units (services, sockets, timers, mounts…).
● For each service, systemd:
○ starts the process
○ tracks it
○ restarts if configured
○ logs stdout/stderr to journald
○ places it in its own cgroup automatically
How systemd stops a service (signals)
● systemd typically sends SIGTERM first, waits TimeoutStopSec, then SIGKILL if it’s
still running.
● This matches what you want in production: graceful shutdown, then force.
Common follow-ups
Q: What if your process becomes PID 1 in a container?
PID 1 has special semantics: it must reap zombies and handle signals correctly. Many
container images use a tiny init like tini to do proper reaping/signal forwarding.
Q: Why mention systemd + cgroups together?
Because systemd uses cgroups to track and limit services: “every unit is a cgroup.”
6) Representative interview answers (short + strong)
“Explain containerization.”
“Containerization is OS-level virtualization: you run a normal Linux process but isolate it using
namespaces (PID, NET, MNT, UTS, IPC, USER) and control its resources using cgroups (CPU,
memory, IO, pids). It usually also has its own root filesystem and restricted privileges.”
“It’s worth mentioning cgroups and namespaces for containerisation.”
“Yes—namespaces provide isolation of what the process can see, cgroups provide limits and
accounting of what it can use. Together they are the core kernel primitives behind containers.”
“How do you make a process a service?”
“I’d write a systemd unit: define ExecStart, restart policy, dependencies, user, environment, and
resource limits. Enable it so it starts at boot, and use journald for logs. systemd supervises the
process and places it into a cgroup.”
7) What to memorize (high ROI)
● Namespaces = isolation (PID/NET/MNT/UTS/IPC/USER)
● cgroups = limits + accounting (CPU/memory/io/pids)
● Containers = process + namespaces + cgroups + rootfs + security
● systemd service = unit file + supervision + logging + cgroup tracking
“What is the difference between a process and a thread?” what does a cpu scheduler work
with?
Process vs thread (what’s different)
Process
A process is primarily an address space + resources.
● Has its own virtual memory (page tables / mm_struct)
● Own set of resources: file descriptor table, current working dir, environment, credentials,
etc.
● Isolation boundary: one process can’t directly read/write another process’s memory
without explicit mechanisms (shared memory, ptrace, etc.)
Think: “container of resources + memory map.”
Thread
A thread is primarily an execution context inside a process.
● Has its own:
○ CPU register state (PC/IP, SP, general registers)
○ stack(s) (user stack + kernel stack)
○ scheduling state, priority, CPU affinity, etc.
● Shares with other threads in the same process:
○ address space (same virtual memory mappings)
○ open files (usually), signal dispositions, etc. (depends on what’s shared)
Think: “a separate CPU ‘cursor’ running in the same program memory.”
Linux reality: “threads are processes”
In Linux, threads are implemented as tasks created with clone(). They’re represented the
same basic way as processes:
● each thread has a task_struct
● threads in the same process share the same mm_struct (address space) and often
other resources
So “process vs thread” is really about what is shared:
● Different processes: different mm_struct (different address spaces)
● Threads: same mm_struct, different stacks + registers
What does the CPU scheduler schedule?
The scheduler schedules threads (tasks), not whole processes.
More precisely: it schedules runnable entities with their own execution context:
● each has its own task_struct / PID (Linux has TIDs; userland often calls them thread
IDs)
● each can be independently runnable/blocked, prioritized, pinned to CPUs, etc.
If a process has 10 threads, the scheduler may run 0, 1, or many of them at the same time (on
different cores), depending on CPU availability and policy.
Why this matters
● A “single process” can consume multiple cores if it has multiple runnable threads.
● Context switches usually happen between threads, even within the same process.
● Switching between threads in the same process is often cheaper than between
processes because the address space (page tables) usually doesn’t change.
If you want, I can show how this appears in tools:
● ps -eLf (threads)
● top -H / htop (per-thread view)
● /proc/<pid>/task/ (Linux thread directory)
“How do you load a Linux kernel?” “How do you load a module into the Linux kernel?”
1) How do you load a Linux kernel?
At boot: the normal path
You don’t “load” a kernel from inside Linux; the firmware + bootloader loads it into RAM and
jumps to it.
Typical chain:
1. Firmware (BIOS/UEFI) initializes hardware and picks a boot device.
2. Bootloader (usually GRUB, systemd-boot, U-Boot) loads:
○ the kernel image (e.g., /boot/vmlinuz-…)
○ optionally an initramfs/initrd (early userspace image)
○ kernel command line parameters
3. Bootloader jumps to the kernel entry point.
4. Kernel initializes memory, drivers, mounts an initial root:
○ If using initramfs: kernel mounts it as temporary root, runs /init
○ initramfs loads drivers/storage/network as needed, then
switch_root/pivot_root to the real root filesystem
5. Kernel starts PID 1 (/sbin/init, often systemd)
How you “choose” which kernel boots
● Install kernel packages → they drop new vmlinuz + initramfs into /boot
● Update bootloader config (often automatic)
● Select kernel in GRUB menu at boot
Load at boot
● Add module name to:
○ /etc/modules-load.d/*.conf
● Or configure via distro tooling.
●
Tools & Troubleshooting “How do you interpret free -h/vmstat/iostat output?” “How do you find
the process using the most memory / CPU?” “How would you diagnose high CPU / slow system
/ slow web se
Thought for 11s
Below is a practical “Linux triage playbook” for free/vmstat/iostat, then how to find top
CPU/memory, then how to diagnose high CPU / slow system / slow web server.
Interpreting free -h
Example fields:
● total: installed RAM
● used: not “apps only”; includes some kernel usage
● free: completely unused RAM (usually small on healthy Linux)
● shared: shared memory (tmpfs, SHM segments)
● buff/cache: filesystem page cache + buffers (reclaimable)
● available: most important: how much RAM you can use without swapping / heavy
reclaim
Rules of thumb:
● Low free is normal if available is healthy.
● Low available (and dropping) indicates real memory pressure.
● If you see swap used climbing and available low, you’re under pressure.
Interpreting vmstat 1
Run:
vmstat 1
Key columns (most useful):
CPU
● us: user CPU (application code)
● sy: system CPU (kernel time)
● id: idle
● wa: I/O wait (CPU idle because tasks are waiting on I/O)
● st: steal (VM host stealing CPU; common on oversubscribed hypervisors)
Interpretation:
● High us → CPU-bound userland (app work)
● High sy → kernel-heavy (syscalls, networking, storage stack, interrupts, context
switching)
● High wa → I/O bottleneck or blocked storage/network FS
● High st → VM scheduling issue (host contention)
Run queue / blocking
● r: runnable tasks (want CPU)
○ If r > number_of_cores consistently → CPU contention
● b: blocked tasks (usually waiting on I/O)
○ If b is high with wa high → likely storage I/O bottleneck
Memory pressure / swapping
● si/so: swap-in / swap-out (per second)
○ Non-zero sustained = memory pressure.
● (Some vmstat versions show free, buff, cache: useful trend but “available” from free
is better.)
System activity
● in: interrupts/sec (high can indicate network storms, IRQ issues)
● cs: context switches/sec (high can indicate many threads, lock contention, or high I/O
wakeups)
Interpreting iostat -x 1
Run:
iostat -x 1
Focus on per-device (nvme0n1, xvda, etc.):
● r/s, w/s: read/write ops per second
● rkB/s, wkB/s: throughput
● await: average latency (ms) per I/O (includes queueing)
● avgqu-sz: average queue length
● %util: device busy percentage
Interpretation patterns:
● High await + high %util (~100%) → device saturated (real disk bottleneck)
● High avgqu-sz → queue building up (requests waiting)
● High await + low %util → not pure disk saturation; could be:
○ downstream storage/network issue (EBS/NFS)
○ controller/driver stalls
○ filesystem/journal contention
● Lots of I/O on one device + vmstat wa high → storage is your limiter.
Finding the process using the most CPU / memory
CPU (quick)
top
Press:
● P (or default) sort by CPU
● 1 to show per-core usage
Or one-shot:
ps -eo pid,user,comm,%cpu,%mem --sort=-%cpu | head -20
Memory (quick)
top
Press Shift+M to sort by memory.
Or one-shot (RSS in KiB):
ps -eo pid,user,comm,rss,%mem --sort=-rss | head -20
More accurate memory attribution (shared memory aware)
If installed:
smem -r -k | head
Per-process I/O (often the “real” culprit)
sudo iotop -oPa
or:
pidstat -d 1
Diagnosing: high CPU / slow system
Step 1 — Decide: CPU vs memory vs I/O vs “something else”
Run these together for 10 seconds:
uptime
free -h
vmstat 1 10
iostat -x 1 10
How to read the story:
A) CPU-bound
● vmstat: high us (or sy), low id, r > cores
● top: one/few processes pegging CPU
Next:
● Is it user CPU or system CPU?
○ high us → app logic, tight loops
○ high sy → heavy syscalls/network/disk stack, interrupts
Tools:
pidstat -u 1
perf top # if available; see hottest kernel/userspace functions
strace -p <PID> # see syscall storm (careful on prod)
B) I/O-bound (disk / storage)
● vmstat: high wa, high b
● iostat -x: high await and/or %util
Next:
● Which disk? which process?
sudo iotop -oPa
pidstat -d 1
dmesg -T | tail -n 50 # check for disk timeouts/resets
C) Memory pressure
● free -h: available low
● vmstat: si/so non-zero (swap activity)
Next:
ps -eo pid,comm,rss,%mem --sort=-rss | head
cat /proc/meminfo | egrep 'MemAvailable|Swap|Dirty|Writeback'
If no swap configured, you may see OOM kills:
dmesg -T | egrep -i 'oom|killed process' | tail
D) VM host contention
● vmstat: high st
Fix is usually outside the VM (move instance, resize, host contention).
Diagnosing: slow web server (a clean checklist)
1) Is the problem CPU, I/O, memory, or network?
Same baseline:
uptime; free -h
vmstat 1 5
iostat -x 1 5
2) Are we overloaded with connections?
ss -s
ss -tan | head
ss -tan state established | wc -l
ss -tan state time-wait | wc -l
ss -tan state close-wait | wc -l
● Many CLOSE_WAIT → app not closing sockets (bug/leak).
● Many TIME_WAIT → high churn; consider keepalive/pooling.
3) Is the web server itself the bottleneck (nginx/apache) or upstream
app/DB?
● Check worker saturation and errors:
○ nginx: stub_status (if enabled), error log
○ apache: server-status
● Look for 499/502/504 and upstream timeouts in logs.
Quick latency from the server itself:
● If localhost is fast but external is slow → network/LB/DNS/client path.
● If localhost is slow → server/app/DB.
4) Identify the hottest process behind the web server
● CPU:
top
pidstat -u 1
● I/O:
sudo iotop -oPa
pidstat -d 1
● If app is stuck:
sudo strace -p <PID>
5) Check kernel signals of trouble
dmesg -T | egrep -i
'error|warn|oom|killed|ext4|xfs|nvme|timeout|reset' | tail -n 50
Fast “symptom → likely cause” mapping
● Load avg high + vmstat r high + us high → CPU saturation
● Load avg high + wa high + iostat await/util high → disk/storage bottleneck
● available low + si/so > 0 → memory pressure/swapping
● st high → VM host contention
● close_wait rising → app socket leak / not closing
● many established + high latency + low CPU → upstream dependency (DB/API) or
lock contention
explain this properly for the interview to me now: Examples include: “How does DNS
work?”“How does TCP work?” “How does UDP work?”“Explain how traceroute works under the
hood. What part of the TCP header does traceroute modify?” (Typically ICMP/TTL
discussion.)“Explain, step by step, what happens when you enter [Link] in a
browser.”“How is a network process terminated?”These questions test understanding of: OS
socket API, kernel networking stack. ICMP, TTL, routing, ARP. How user-space apps interact
with kernel networking.
Thought for 13s
Below are interview-ready explanations with the kernel/socket API + protocol mechanics
woven in. Use these as “talk tracks”.
1) How DNS works
Goal: translate [Link] → IP address (and sometimes other records).
The moving parts
● Stub resolver: your machine (glibc/systemd-resolved) that asks a resolver.
● Recursive resolver: usually your ISP/company DNS or a public resolver ([Link],
[Link]).
● Authoritative DNS: the servers that “own” the domain’s records.
● Caches: stub + recursive + intermediate caches.
Step-by-step resolution (typical)
1. Browser/OS asks the stub resolver for A/AAAA records.
2. Stub checks local caches (and hosts file).
3. If not cached, stub sends a DNS query to the configured recursive resolver (usually
UDP/53).
4. If the recursive resolver doesn’t have it cached, it does iterative lookups:
○ Query a root server: “Who knows .com?”
○ Query a .com TLD server: “Who’s authoritative for [Link]?”
○ Query authoritative server for [Link]: “What’s [Link]?”
5. Resolver returns the final answer to the stub with a TTL (cache lifetime).
6. Stub caches it and returns it to the browser.
Practical details interviewers like
● DNS is usually UDP, but switches to TCP for large answers, truncation, or DNSSEC.
● Many sites use CNAMEs to point hostnames to CDN names; resolution may involve
multiple steps.
● Often both AAAA (IPv6) and A (IPv4) are queried; “Happy Eyeballs” may race IPv6 vs
IPv4.
2) How TCP works
TCP = reliable, ordered, congestion-controlled byte stream over IP.
Connection setup (3-way handshake)
1. Client → Server: SYN (seq = x)
2. Server → Client: SYN-ACK (seq = y, ack = x+1)
3. Client → Server: ACK (ack = y+1)
Now both sides have initial sequence numbers and can exchange data.
Data transfer mechanics
● Sequence numbers / ACKs: reliability + in-order delivery.
● Retransmission: if ACKs don’t arrive, resend after timeout or fast-retransmit.
● Flow control (receiver window): prevents overrunning receiver buffers.
● Congestion control (cwnd): adapts rate to network capacity (slow start, congestion
avoidance).
● MSS/MTU: segmentation sized to avoid fragmentation.
Connection teardown
● Graceful close is usually a FIN/ACK 4-way exchange.
● TIME_WAIT happens on the active closer to handle delayed packets safely.
● RST is an abort: “connection invalid/no longer exists”.
3) How UDP works
UDP = connectionless datagrams (message boundaries preserved).
● No handshake, no reliability, no ordering, no congestion control built-in.
● Lower overhead, good for real-time, DNS, streaming, gaming.
● Apps must handle: retries, ordering, congestion behavior if needed.
In practice: protocols like QUIC (HTTP/3) run reliability + congestion control in user space on
top of UDP.
4) How traceroute works under the hood
Key idea: manipulate IP TTL, not TCP header fields.
TTL / hop limit basics
● Each IP packet has a TTL (IPv4) or Hop Limit (IPv6).
● Every router decrements it by 1.
● If it reaches 0, router drops the packet and sends ICMP Time Exceeded back.
traceroute algorithm
1. Send probe packets with TTL = 1 → first router decrements to 0 and returns ICMP Time
Exceeded. Measure RTT.
2. Send TTL = 2 → you learn the second hop.
3. Repeat TTL = 3, 4, … until destination responds.
How it knows it reached the destination (depends on probe type)
● Classic traceroute uses UDP to a high, likely-closed port → destination replies ICMP
Port Unreachable.
● Some implementations use ICMP Echo → destination replies ICMP Echo Reply.
● Some use TCP SYN (useful through firewalls) → destination replies SYN-ACK (open) or
RST (closed).
Interview trap question
“What part of the TCP header does traceroute modify?”
Answer: Typically none. traceroute primarily changes the TTL field in the IP
header (or Hop Limit in IPv6). Some variants also choose TCP/UDP destination
ports, but TTL is the core mechanism.
5) What happens when you type [Link] in a
browser (step-by-step)
Think in layers: URL parsing → DNS → TCP/TLS → HTTP → rendering.
1) URL + policy
● Browser parses URL, checks HSTS/preload list (forces HTTPS).
● Decides proxy settings, VPN, etc.
2) DNS resolution
● As described above: stub → recursive → authoritative.
● Results cached (TTL).
3) Establish transport
● If HTTPS over TCP:
○ socket() → connect() → TCP handshake.
● If HTTP/3:
○ QUIC over UDP, different handshake mechanics.
4) TLS handshake (HTTPS)
● ClientHello (SNI = hostname, ALPN = http/1.1 or h2, key shares)
● ServerHello + certificate chain
● Key agreement → encrypted session established
● Certificate validation (CA trust, hostname match, revocation policy)
5) HTTP request/response
● Browser sends GET / with headers (Host, cookies, accept, user-agent, etc.)
● Server (often via CDN/LB) responds:
○ redirects (301/302), content, cache headers
○ sets cookies, security headers (CSP, etc.)
6) Fetch subresources + render
● Parse HTML → build DOM
● Fetch CSS/JS/images (more DNS + more connections or multiplexing in HTTP/2/3)
● Build CSSOM, render tree, layout, paint, composite
● Execute JS (may trigger more network calls)
Where kernel networking appears
● Browser calls socket API; kernel handles:
○ routing decision
○ ARP/neighbor discovery to reach gateway
○ segmentation, retransmissions, congestion control
○ NIC driver + interrupts/softirqs
6) How is a network process terminated?
There are two perspectives: process lifecycle and TCP connection lifecycle.
Process termination
● If a process exits (or gets SIGTERM/SIGKILL):
○ kernel closes its file descriptors, including sockets.
○ what happens on the wire depends on socket state/options.
What peers observe on the connection
● Graceful close (FIN):
○ process calls close() (or exits normally) → TCP sends FIN after queued data.
● Abort (RST):
○ can happen if app sets SO_LINGER with 0 timeout, or if the socket is torn down in
a way that discards unsent data → peer gets RST.
● Peer can also terminate:
○ peer sends FIN (you see CLOSE_WAIT if app doesn’t close)
○ peer sends RST
● Or termination by network conditions:
○ idle timeouts (NAT/LB)
○ retransmission timeouts
7) OS socket API ↔ kernel networking stack (what
interviewers want you to say)
Typical user-space calls
● socket() → create endpoint
● bind() → choose local IP/port (servers)
● listen()/accept() → accept connections (TCP)
● connect() → establish outgoing connection (TCP) or set default peer (UDP)
● send()/recv() / read()/write() → move data between user buffers and kernel
socket buffers
● close() → start teardown / free resources
Kernel stack path (high level)
● User buffer ↔ syscall ↔ socket buffer (sk_buff) ↔ TCP/UDP ↔ IP routing ↔
neighbor (ARP/ND) ↔ NIC driver ↔ wire
And reverse on receive, driven by interrupts/softirqs.