A night of memory chaos
I was running a 4 GB‑RAM VPS that hosts a handful of Docker containers and a personal website. A nightly backup script, scheduled with cron, suddenly started chewing through all RAM, triggering the OOM killer and bringing the whole machine to a halt. The crash happened on a Tuesday afternoon, and I had to reboot the server to get it back online. The lesson? Even a single misbehaving cron job can exhaust memory if you don’t guard against runaway processes. Below is a step‑by‑step guide on how I diagnosed the problem, added zram to give the kernel a memory cushion, and tweaked ulimit settings so that future cron jobs can’t over‑consume RAM.
The crash
The server ran a mysqldump job that dumped a 1.2 GB database nightly. The script looked like this:
#!/usr/bin/env bash
mysqldump -u root -pPASSWORD mydb > /var/backups/mydb_$(date +%F).sql
On the day of the incident, the job ran for 45 minutes, but the system’s /var/log/syslog shows:
Oct 5 14:32:07 server kernel: Out of memory: Kill process 1234 (mysqldump) score 42 or sacrifice child
Oct 5 14:32:07 server kernel: Killed process 1234, UID 0, (mysqldump)
The OOM killer had to terminate the backup process because the kernel could not find free pages. The server rebooted automatically after the kernel panicked. I had no systemd watchdog or cron‑specific alerts set up, so I only discovered the issue when the web service went down.
Diagnosing the culprit
-
Check the cron log
On Debian‑based systems, cron logs to/var/log/syslog. Search for the job’s timestamp:grep 'mysqldump' /var/log/syslog | tail -n 20 -
Inspect memory usage
Usepsandtopto see how much RAM a process consumes:ps -o pid,cmd,%mem,rss -C mysqldumpThe RSS value was ~1.5 GB, exceeding the 4 GB available after accounting for the OS and Docker containers.
-
Look for memory leaks
If a process repeatedly grows, it may have a leak.pmap -x <pid>shows the memory map. In this case, the dump command was fine; the issue was the sheer size of the dump. -
Check swap
The server had no swap configured. Without swap, the kernel had no buffer to move pages to disk, so it had to kill processes.
Understanding zram
zram is a kernel module that creates a compressed block device in RAM. It behaves like swap but is much faster because it’s in memory. The kernel can move pages to zram when RAM pressure rises, effectively giving you “compressed RAM.” On modern kernels, zram is enabled by default in many distributions, but the default size is often small (e.g., 512 MB). For a 4 GB server, you can allocate a larger zram device to act as a safety net.
Benefits
- Faster than disk swap.
- Reduces I/O bottlenecks.
- Works well with workloads that have a high working set but low I/O.
Trade‑offs
- Uses RAM for compression; if the workload is memory‑intensive, zram may not help.
- Compression overhead can increase CPU usage.
Enabling zram on Debian/Ubuntu
-
Install the module
On recent kernels,zramis built in. Verify:lsmod | grep zramIf it’s missing, install the package:
sudo apt-get install zram-tools -
Create a zram device
Usezramctlto set up a 2 GB zram device:sudo zramctl -f -s 2G -o lz4The
-o lz4flag selects the LZ4 compression algorithm, which balances speed and compression ratio. -
Format and mount
Treat the device like swap:sudo mkswap /dev/zram0 sudo swapon /dev/zram0 -
Persist across reboots
Add the following to/etc/fstab:/dev/zram0 none swap sw 0 0And create a systemd unit to set up zram at boot:
[Unit] Description=Configure zram DefaultDependencies=no Before=swap.target [Service] Type=oneshot ExecStart=/usr/bin/zramctl -f -s 2G -o lz4 ExecStartPost=/sbin/mkswap /dev/zram0 ExecStartPost=/sbin/swapon /dev/zram0 [Install] WantedBy=multi-user.targetEnable it:
sudo systemctl enable zram-setup.service -
Verify
After reboot, check:free -hYou should see the 2 GB zram listed under “Swap.”
Tuning ulimit for cron jobs
`
See also
- systemd‑resolved’s DNS Cache Stale After a Router Update: My One‑Command Fix
- Finding the 10 Most Frequent Error Lines in a 5 GB Syslog with awk in a Few Seconds
- How I Use Find and Xargs to Delete Temp Files Older Than 7 Days Without Risking rm ‑rf
- Fixing ACME DNS‑01 on a Home Lab Using Pi‑hole and Cloudflare
- How I stopped a 16‑GB microSD Raspberry Pi from dying mid‑boot because /var/log grew to 10 GB after a year of unattended cron jobs