Linux Process Model & Signal Handling
What Every Engineer Running Containers Must Know From Linux Fundamentals to AI Platform Engineering

👋 Hey there! I'm Jeeva, a passionate DevOps and Cloud Engineer with a knack for simplifying complex infrastructures and optimizing workflows. With a strong foundation in Python and PySpark, I thrive on designing scalable solutions that leverage the power of the cloud.
🛠️ In my journey as a DevOps professional, I've honed my skills in automating deployment pipelines, orchestrating containerized environments, and ensuring robust security measures. Whether it's architecting cloud-native applications or fine-tuning infrastructure performance, I'm committed to driving efficiency and reliability at every step.
💻 When I'm not tinkering with code or diving into cloud platforms, you'll likely find me exploring the latest trends in technology, sharing insights on DevOps best practices, or diving deep into data analysis with PySpark.
📝 Join me on this exhilarating ride through the realms of DevOps, Cloud Engineering where we'll unravel the complexities of modern IT landscapes and empower ourselves with the tools to build a more resilient digital future.
Connect with me on LinkedIn: https://www.linkedin.com/in/jeevabalakrishnan
🐧 Linux Process Model & Signal Handling — What Every Engineer Running Containers Must Know
Day 01 of the Principal Engineer Learning Path — From Linux Fundamentals to AI Platform Engineering
The 2 AM Incident Nobody Could Explain
Picture this.
A senior engineer at a fast-growing SaaS company is on-call. It's 2 AM. The alerting system lights up — elevated 5xx errors, every single deployment for the past three months has caused a 30-second spike of dropped requests. The team has blamed the load balancer. They've blamed the database. They've blamed the CDN.
Nobody looked at PID 1.
The root cause was embarrassingly simple: their Python application was running as PID 1 inside a Docker container, had no SIGTERM handler, and Kubernetes was force-killing it with SIGKILL after the grace period expired — every single rolling update.
Three months of production incidents. One concept they hadn't studied deeply.
This is why we start here.
What This Series Is
This is Day 01 of the Principal Engineer Learning Path — a structured, progressive series taking you from Linux fundamentals all the way through to AI Platform Engineering, one concept per day.
Each post covers:
Why the concept matters in real production systems
Deep technical explanation (not surface level)
Architecture context with real system diagrams
Hands-on tasks you can complete in under 90 minutes
Interview questions from beginner to advanced
How it connects to AI/ML infrastructure
If you're aiming for Senior → Staff → Principal, this series is built for you.
Why Linux Process Model Is Day 01
Before containers. Before Kubernetes. Before cloud. There are Linux processes.
Every Docker container is a Linux process with namespace isolation. Every Kubernetes pod runs one or more Linux processes. Every AWS Lambda invocation spins up a Linux process. Every CI/CD job runner is a Linux process.
When things go wrong in your infrastructure — graceful shutdowns failing, zombie processes accumulating, SIGTERM being ignored, pods taking 30+ seconds to terminate — the explanation is almost always rooted in the Linux process model.
You cannot be a principal-level platform engineer without this foundation being completely solid.
Part 1 — The Linux Process Model
Every Process Has an Identity
When Linux starts a program, it creates a process — an isolated execution environment with:
| Property | Description |
|---|---|
| PID | Unique integer identifying this process |
| PPID | Parent Process ID — who spawned it |
| File descriptors | Open files, sockets, pipes |
| Memory space | Heap, stack, code segment |
| Signal handlers | Table of how to respond to signals |
| Environment variables | Inherited from parent |
| Working directory | Where the process thinks it is |
Every process on your system was created by another process using the fork() system call, followed by exec() to replace the process image with a new program. This is called the fork-exec pattern and it's how everything on Linux starts.
fork() → creates an exact copy of the parent process
exec() → replaces the copy with a new program
The Process Tree
All processes form a tree. At the root of that tree is PID 1 — the init process. On modern Linux systems, that's systemd. On Alpine Linux containers, it might be /bin/sh. On your production containers... it might be your Python app.
PID 1 (systemd)
├── PID 412 nginx worker
│ ├── PID 890 request handler
│ └── PID 891 request handler
├── PID 892 python app
│ ├── PID 993 celery worker
│ └── PID 994 [zombie] ← this is the problem
└── PID 1024 postgres
├── PID 1025 postgres: checkpointer
└── PID 1026 postgres: background writer
What Makes PID 1 Special
PID 1 has a responsibility that no other process has: it must reap zombie processes.
When a child process exits, it doesn't fully disappear. It enters a zombie state — a dead process whose exit status hasn't been collected by its parent yet. The parent is supposed to call wait() to collect it, which removes the zombie entry from the process table.
If the parent never calls wait(), the zombie stays in the process table forever. It uses no CPU or memory, but it does use a PID. And Linux has a finite PID namespace.
PID 1 (init/systemd) has special logic to adopt orphaned processes and reap them. Regular applications like Python, Node.js, or Go don't have this logic built in.
This is why running your app as PID 1 in a container is dangerous.
Part 2 — Signals: The Async Communication System
A signal is an asynchronous notification sent to a process. Think of it like a software interrupt — the process is doing its work, and suddenly it receives a tap on the shoulder saying "stop", "reload", or "something happened to your child."
The Signals You Must Know Cold
Signal Number Default Action Real-World Use
─────────────────────────────────────────────────────────────────────
SIGTERM 15 Terminate Graceful shutdown (Kubernetes, systemd)
SIGKILL 9 Force kill Last resort — cannot be caught or ignored
SIGHUP 1 Terminate Reload config without restart (nginx)
SIGINT 2 Terminate Ctrl+C in terminal
SIGCHLD 17 Ignore Notifies parent when child state changes
SIGUSR1 10 Terminate App-defined — e.g. dump stats, rotate logs
SIGUSR2 12 Terminate App-defined — e.g. trigger hot reload
SIGSTOP 19 Stop process Pause execution (cannot be caught)
SIGCONT 18 Continue Resume a stopped process
The Critical Difference: SIGTERM vs SIGKILL
This is the single most important signal distinction for platform engineers.
SIGTERM (15) is a request. You are asking the process to please shut down. The process receives this signal, and it can:
Handle it gracefully (close connections, flush buffers, finish in-flight requests, exit cleanly)
Ignore it entirely (bad practice, but possible)
Crash (if the handler throws an exception)
SIGKILL (9) is a command from the kernel. It cannot be caught, blocked, or ignored. The process is terminated immediately — mid-instruction if necessary. No cleanup. No buffer flushing. No connection draining.
Kubernetes always sends SIGTERM first. It waits for
terminationGracePeriodSeconds(default: 30s). If the process is still alive after that window, SIGKILL is sent.
If your application doesn't handle SIGTERM, every single rolling deployment force-kills your app, drops in-flight requests, and potentially corrupts state.
How to Handle SIGTERM in Your Application
Python:
import signal
import sys
import time
def handle_sigterm(signum, frame):
print("SIGTERM received — starting graceful shutdown")
# Stop accepting new requests
# Drain connection pool
# Flush metrics
# Close DB connections
print("Cleanup complete — exiting")
sys.exit(0)
signal.signal(signal.SIGTERM, handle_sigterm)
# Your app runs here
while True:
time.sleep(1)
Node.js:
process.on('SIGTERM', async () => {
console.log('SIGTERM received — graceful shutdown');
// Stop HTTP server from accepting new connections
server.close(async () => {
await db.disconnect();
await redis.quit();
process.exit(0);
});
// Force exit if cleanup takes too long
setTimeout(() => process.exit(1), 25000);
});
Go:
sigChan := make(chan os.Signal, 1)
signal.Notify(sigChan, syscall.SIGTERM, syscall.SIGINT)
go func() {
<-sigChan
log.Println("SIGTERM received — draining server")
ctx, cancel := context.WithTimeout(context.Background(), 25*time.Second)
defer cancel()
if err := server.Shutdown(ctx); err != nil {
log.Fatal("Server forced to shutdown:", err)
}
os.Exit(0)
}()
Part 3 — The PID 1 Problem in Containers
The Setup
When you write this in your Dockerfile:
FROM python:3.11-slim
COPY app.py /app.py
CMD ["python", "/app.py"]
Your Python process starts as PID 1 inside the container's PID namespace.
Here's why that's a problem:
No zombie reaping — Python doesn't call
waitpid()on child processes. If your app spawns subprocesses (PDF generation, shell commands, worker threads that fork), they become zombies when they exit.SIGTERM handling is not guaranteed — Many applications have incomplete SIGTERM handlers, or their frameworks don't wire it up properly.
Signal forwarding issues — If your CMD uses shell form (
CMD python app.pywithout brackets),/bin/shbecomes PID 1. Many shells don't forward signals to child processes.
Reproducing the Problem
# zombie_demo.py
import subprocess
import time
import os
print(f"My PID: {os.getpid()}")
# Will print 1 when run as container entrypoint — DANGER
# Spawn a child that exits — becomes zombie if not reaped
proc = subprocess.Popen(["sleep", "5"])
# NOT calling proc.wait() — this becomes a zombie after sleep exits
while True:
time.sleep(1)
# bad.Dockerfile
FROM python:3.11-slim
COPY zombie_demo.py /zombie_demo.py
CMD ["python", "/zombie_demo.py"]
docker build -f bad.Dockerfile -t pid1-demo .
docker run --name pid1-demo -d pid1-demo
# See Python running as PID 1
docker exec pid1-demo ps aux
# PID USER COMMAND
# 1 root python /zombie_demo.py ← PROBLEM
# 6 root sleep 5
# Measure how long stop takes — expect ~10s (SIGKILL timeout)
time docker stop pid1-demo
The Fix: Use tini
tini is a minimal init system designed specifically for containers. It does exactly two things:
Acts as a proper PID 1, reaping zombie child processes
Correctly forwards signals to child processes
# good.Dockerfile
FROM python:3.11-slim
RUN apt-get update && apt-get install -y --no-install-recommends tini && \
rm -rf /var/lib/apt/lists/*
COPY zombie_demo.py /zombie_demo.py
ENTRYPOINT ["/usr/bin/tini", "--"]
CMD ["python", "/zombie_demo.py"]
Now the process tree inside your container looks like this:
PID 1 tini ← init process, reaps zombies, forwards signals
└── PID 2 python app.py ← your app, receives SIGTERM properly
You can also enable this via Docker without modifying your image:
# Docker run flag
docker run --init my-image
# Docker Compose
services:
app:
image: my-image
init: true
Shell Form vs Exec Form
This is a common gotcha that bites teams regularly:
# ❌ Shell form — /bin/sh -c is PID 1, signals NOT forwarded to your app
CMD python app.py
# ✅ Exec form — your app IS the process, signals forwarded correctly
CMD ["python", "app.py"]
The shell form wraps your command in
/bin/sh -c "python app.py". The shell becomes PID 1. Many minimal shells do not forward signals to child processes. Always use exec form in production Dockerfiles.
Part 4 — Kubernetes Pod Termination: The Complete Picture
Understanding this flow end-to-end is what separates engineers who debug graceful shutdown issues in minutes from those who spend days on it.
The Full Termination Sequence
kubectl delete pod my-pod
│
▼
1. Pod marked for deletion
• Pod added to API server's deletion queue
• Pod removed from Service endpoints (no new traffic)
• ⚠️ Race condition: traffic still arrives for 2–5s during propagation
│
▼
2. preStop hook executes (if configured)
• Runs BEFORE SIGTERM is sent
• Use this to sleep and allow load balancer connection draining
• If preStop exceeds grace period, gets killed anyway
│
▼
3. SIGTERM sent to PID 1 in each container
• Your app receives SIGTERM
• Grace period countdown STARTS (default: 30 seconds)
• App should stop accepting requests and drain existing ones
│
▼
4. Grace period countdown
• terminationGracePeriodSeconds: 30 (default)
• Override per pod spec
• Extend significantly for batch jobs, ML training, Kafka consumers
│
├──────────────────────────────────┐
▼ ▼
5a. Clean exit ✅ 5b. SIGKILL ❌
App exits within grace period Grace period expired
Containers removed cleanly OS force-kills all processes
Pod deleted from etcd No cleanup possible
Potential data loss
The preStop Hook: Your Load Balancer Buffer
There is a race condition in Kubernetes that many teams don't know about.
When a pod is being terminated, Kubernetes simultaneously:
Sends SIGTERM to the pod
Removes the pod from the Service's Endpoints
These two operations happen in parallel, not sequentially. The Endpoints update needs to propagate to kube-proxy and your ingress controller. This takes a few seconds. During those seconds, your load balancer might still route traffic to a pod that's already shutting down.
The fix is a preStop sleep:
spec:
terminationGracePeriodSeconds: 60
containers:
- name: app
image: my-app:latest
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 10"]
This sleep runs before SIGTERM is sent, giving the load balancer 10 seconds to drain connections before your app starts shutting down.
Grace Period Reference by Workload Type
# Stateless API — default is usually fine
terminationGracePeriodSeconds: 30
# Kafka consumer — needs to commit offsets
terminationGracePeriodSeconds: 60
# ML inference server — needs to drain request queue
terminationGracePeriodSeconds: 120
# ML training job — needs to checkpoint model
terminationGracePeriodSeconds: 300
# Database — needs to flush WAL and close connections
terminationGracePeriodSeconds: 120
Production-Ready Deployment Manifest
apiVersion: apps/v1
kind: Deployment
metadata:
name: graceful-app
spec:
replicas: 2
selector:
matchLabels:
app: graceful-app
template:
metadata:
labels:
app: graceful-app
spec:
terminationGracePeriodSeconds: 60
containers:
- name: app
image: my-app:latest
ports:
- containerPort: 8080
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 10"]
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
Part 5 — AI Platform Engineering Connection
This is where it gets interesting for ML infrastructure engineers.
vLLM and TorchServe Signal Handling
When an LLM inference server deployed on Kubernetes receives SIGTERM, these steps must happen in the correct order:
Stop accepting new inference requests — return 503 immediately
Complete in-flight requests — potentially 10–30 second requests
Flush metrics — Prometheus counters, token counts, latency histograms
Release GPU memory —
torch.cuda.empty_cache()and model unloadingClose connections — gRPC channels, Redis, model registry
If your handler skips any step you get:
Partial responses delivered mid-token-stream
GPU memory leaks causing OOM on the next pod start
Missing billing and observability metrics
import signal
import asyncio
import torch
shutdown_event = asyncio.Event()
def handle_sigterm(signum, frame):
print("SIGTERM received — initiating graceful shutdown")
shutdown_event.set()
signal.signal(signal.SIGTERM, handle_sigterm)
async def shutdown_handler():
await shutdown_event.wait()
print("Stopping request acceptance...")
app.state.shutting_down = True # Return 503 to new requests
print("Waiting for in-flight requests to complete...")
await asyncio.sleep(30) # Wait for long inference requests
print("Releasing GPU memory...")
model.cpu()
del model
torch.cuda.empty_cache()
print("Flushing metrics...")
await metrics_client.flush()
print("Shutdown complete")
sys.exit(0)
Spot Instance Interruption in ML Training
AWS Spot Instances send SIGTERM 2 minutes before reclamation. Handling it correctly saves potentially hours of recomputation:
import signal
import torch
import sys
checkpoint_path = "/mnt/efs/checkpoints/model_latest.pt"
current_epoch = 0
global_step = 0
def save_checkpoint(signum, frame):
print(f"Spot interruption at epoch {current_epoch}, step {global_step}")
torch.save({
'epoch': current_epoch,
'model_state_dict': model.state_dict(),
'optimizer_state_dict': optimizer.state_dict(),
'loss': current_loss,
'step': global_step,
}, checkpoint_path)
print(f"Checkpoint saved to {checkpoint_path}")
sys.exit(0)
signal.signal(signal.SIGTERM, save_checkpoint)
# Training loop — checkpoint loads from file on restart
for epoch in range(start_epoch, total_epochs):
current_epoch = epoch
train_one_epoch()
RAG Pipeline Graceful Shutdown
In a RAG pipeline, drain components in dependency order:
signal.signal(
signal.SIGTERM,
lambda s, f: asyncio.create_task(graceful_rag_shutdown())
)
async def graceful_rag_shutdown():
# 1. Stop accepting new RAG queries
app.state.accepting_requests = False
# 2. Wait for active retrievals to complete
await retrieval_semaphore.drain()
# 3. Flush the embedding cache
await embedding_cache.flush()
# 4. Close vector DB connection pool
await vector_db_pool.close()
# 5. Close LLM client connections
await llm_client.close()
print("RAG service shutdown complete")
sys.exit(0)
Part 6 — Hands-On Task (90 Minutes)
Prerequisites
Docker Desktop installed
Basic Python knowledge
Optional: minikube or kind for the Kubernetes section
Step 1 — Reproduce the zombie problem (15 min)
# zombie_demo.py
import subprocess
import time
import os
import signal
print(f"PID: {os.getpid()}")
print("Danger!" if os.getpid() == 1 else "Safe — not PID 1")
# Spawn processes that will become zombies
for i in range(3):
p = subprocess.Popen(["sleep", "1"])
print(f"Spawned child PID {p.pid}")
# Deliberately NOT calling p.wait()
while True:
time.sleep(5)
print("Still running — check ps aux for zombies")
# bad.Dockerfile
FROM python:3.11-slim
COPY zombie_demo.py /zombie_demo.py
CMD ["python", "/zombie_demo.py"]
docker build -f bad.Dockerfile -t zombie-bad .
docker run --name zombie-bad -d zombie-bad
sleep 5
# Look for Z (zombie) in the STAT column
docker exec zombie-bad ps aux
Step 2 — Fix with tini (15 min)
# good.Dockerfile
FROM python:3.11-slim
RUN apt-get update && apt-get install -y --no-install-recommends tini
COPY zombie_demo.py /zombie_demo.py
ENTRYPOINT ["/usr/bin/tini", "--"]
CMD ["python", "/zombie_demo.py"]
docker build -f good.Dockerfile -t zombie-good .
docker run --name zombie-good -d zombie-good
sleep 5
# No zombie processes — tini reaps them automatically
docker exec zombie-good ps aux
Step 3 — Measure graceful shutdown timing (20 min)
# graceful_app.py
import signal
import sys
import time
import os
def handle_sigterm(signum, frame):
print("✓ SIGTERM received — starting graceful shutdown")
time.sleep(3) # Simulate real cleanup work
print("✓ Cleanup complete — exiting cleanly")
sys.exit(0)
signal.signal(signal.SIGTERM, handle_sigterm)
print(f"App running as PID {os.getpid()}")
while True:
time.sleep(1)
# Without tini — expect ~10s (Docker waits then SIGKILL)
time docker stop zombie-bad
# With tini — expect ~3s (graceful cleanup then clean exit)
time docker stop zombie-good
# The difference in those numbers is data loss risk per deployment
Step 4 — Apply to Kubernetes (30 min)
# deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: graceful-app
spec:
replicas: 2
selector:
matchLabels:
app: graceful-app
template:
metadata:
labels:
app: graceful-app
spec:
terminationGracePeriodSeconds: 60
containers:
- name: app
image: zombie-good:latest
ports:
- containerPort: 8080
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 10"]
readinessProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 5
periodSeconds: 5
# Apply and watch pod termination
kubectl apply -f deployment.yaml
kubectl delete pod -l app=graceful-app
# Watch the events
kubectl get events --watch
Interview Questions
Beginner
Q: What is the difference between SIGTERM and SIGKILL?
SIGTERM (15) is a polite request — the process can catch it, run cleanup code, and exit gracefully. SIGKILL (9) is an unconditional kernel command — it cannot be caught, blocked, or ignored, and the process is terminated immediately with no opportunity for cleanup. Always try SIGTERM first and only use SIGKILL as a last resort.
Q: What is a zombie process?
A zombie process has exited but its exit status hasn't been collected by its parent via wait(). The process is dead (no CPU, no memory), but its entry remains in the process table. They accumulate and can exhaust the PID namespace. PID 1 (init/tini) is responsible for reaping orphaned zombies.
Intermediate
Q: Why is running your application as PID 1 in a container problematic?
PID 1 has responsibilities that regular applications aren't designed for — specifically, reaping zombie child processes. If your app spawns child processes that exit, they become zombies that only PID 1 can clean up. Additionally, signal behavior is different for PID 1. The solution is using a minimal init system like tini as PID 1, with your application running as PID 2.
Q: What happens to in-flight HTTP requests when Kubernetes terminates a pod?
Kubernetes simultaneously removes the pod from Service Endpoints and sends SIGTERM. Because endpoint propagation takes 2–5 seconds through kube-proxy and ingress, traffic may still arrive at a pod that's already shutting down. Fix: (1) add a preStop sleep to buffer load balancer draining, (2) implement a SIGTERM handler that stops accepting new connections but completes existing ones, (3) set terminationGracePeriodSeconds long enough for cleanup to complete.
Advanced
Q: Walk me through the complete Kubernetes pod termination lifecycle from kubectl delete pod to process exit, identifying every point where data loss can occur.
API server marks pod for deletion
Endpoints controller removes pod from Service — propagation lag creates a 2–5s window where traffic still arrives (data loss point 1)
preStop hooks run — insufficient sleep here means traffic hits a shutting down pod (data loss point 2)
SIGTERM sent to PID 1 — if the app doesn't handle this, or handles it incompletely, in-flight requests are dropped (data loss point 3)
terminationGracePeriodSeconds countdown — if set too short for the workload, cleanup is interrupted (data loss point 4)
SIGKILL sent — no cleanup possible: open file handles corrupted, DB transactions rolled back, metrics lost, GPU memory leaked (data loss point 5)
Q: How would you design graceful shutdown for a stateful Kafka consumer that must commit offsets before exit, where shutdown could take up to 5 minutes?
Set terminationGracePeriodSeconds to 320 (5 min + 20s buffer). On SIGTERM: pause consumption immediately, finish the current message batch, commit offsets explicitly (don't rely on auto-commit), call consumer.close() to trigger rebalance before session timeout (otherwise the group waits 30–45s thinking you're slow). Use a preStop sleep for load balancer draining. Implement a hard deadline at 290s — force exit with partial commit rather than receiving SIGKILL at 320s. Instrument shutdown duration as a metric so you detect when grace period is too tight.
Common Mistakes to Avoid
| Mistake | Impact | Fix |
|---|---|---|
| Shell form CMD | SIGTERM not forwarded | Use CMD ["python", "app.py"] |
| No SIGTERM handler | Force-kill on every deploy | Implement handler in every service |
| Grace period too short | Cleanup interrupted | Measure actual shutdown time + 20% |
| No preStop hook | Dropped requests during rolling update | Add sleep 10 preStop |
| Testing shutdown without load | Handler fails under production traffic | Load-test your shutdown sequence |
| App as PID 1 | Zombie accumulation, signal issues | Use tini or --init flag |
Key Takeaways
Always use tini (or
--init) — your app should never be PID 1Always use exec form —
CMD ["python", "app.py"]notCMD python app.pyAlways implement a SIGTERM handler — every service, every language
Always add a preStop sleep — 5–15 seconds prevents dropped requests
Extend terminationGracePeriodSeconds for any workload needing more than 30 seconds to clean up
What's Next
Day 02: Linux Namespaces & cgroups
Tomorrow we go one level deeper — how do containers actually achieve isolation? What is a Linux namespace, and how does PID, network, mount, and user namespace isolation work? How do cgroups enforce CPU and memory limits? And why does all of this matter when you're debugging a Kubernetes node that's running out of memory?
Series Navigation
| Day | Topic | Phase |
|---|---|---|
| 01 | Linux Process Model & Signal Handling ← you are here | Phase 1: Foundations |
| 02 | Linux Namespaces & cgroups | Phase 1: Foundations |
| 03 | Networking Fundamentals | Phase 1: Foundations |
| 04 | DNS Deep Dive | Phase 1: Foundations |
| 05 | HTTP/HTTPS Internals | Phase 1: Foundations |
| ... | ... | ... |
Tags: linux devops kubernetes docker platform-engineering sre``containers cloud-native backend learning-in-public principal-engineer
