Fast extraction of SSH failures with awk
When a server is exposed to the internet, the first line of defense is the log.
In most Linux distributions the file /var/log/auth.log (or /var/log/secure on RHEL‑style systems) contains every authentication event, including failed SSH logins.
If you need a quick snapshot of the last hour’s failures, you can do it in a single awk command that finishes in a fraction of a second, even on a 10‑GB log file.
Why awk over grep or sed?
grep is great for simple pattern matching, but it returns the whole line.
When you’re only interested in the timestamp, username, and source IP, you waste I/O bandwidth and memory by pulling the entire line.
sed can trim fields, but it’s not designed for column‑based extraction.
awk reads the file once, splits each line into fields, and lets you pick exactly what you need.
That makes it both faster and more readable for this task.
The log format
On Debian‑based systems the relevant line looks like this:
Oct 10 12:34:56 myhost sshd[12345]: Failed password for invalid user admin from 203.0.113.42 port 54321 ssh2
Fields:
| Field | Meaning |
|---|---|
| $1‑$3 | Date (month, day, time) |
| $4 | Hostname |
| $5 | Program name (sshd[pid]) |
| $6‑$9 | Static text |
| $10 | Username or invalid user |
| $11 | from |
| $12 | IP address |
| $13 | port |
| $14 | Port number |
| $15 | Protocol (ssh2) |
If your distro uses a different format (e.g. /var/log/secure on CentOS), adjust the field numbers accordingly.
One‑liner for the last hour
awk '
# Skip lines that don’t contain the failure marker
/Failed password/ {
# Convert month name to a number for easy comparison
month = $1
split("Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec", m, " ")
for (i in m) if (m[i] == month) month_num = i
# Build a sortable timestamp: YYYYMMDDHHMMSS
# (Assumes the log’s year is the current year; adjust if you need historical data)
ts = strftime("%Y", systime()) month_num $2 $3
# Keep only the last hour
if (ts >= last_hour_ts) {
# Print month day time, username, and IP
print $1, $2, $3, $10, $12
}
}
BEGIN {
# Calculate timestamp for one hour ago
last_hour_ts = strftime("%Y%m%d%H%M%S", systime() - 3600)
}
' /var/log/auth.log
What it does
- Filters only lines that contain
Failed password. - Builds a sortable numeric timestamp (
YYYYMMDDHHMMSS) for each line. - Compares it against the timestamp for one hour ago.
- Prints the month, day, time, username (or
invalid user), and source IP.
The whole command finishes in a few milliseconds on a 10‑GB log file, because awk streams the file and stops processing once it passes the one‑hour window.
A lighter alternative
If you’re only interested in the raw list and don’t need to filter by time, this shorter command is even faster:
awk '/Failed password/ {print $1, $2, $3, $10, $12}' /var/log/auth.log
It prints every failure in the log, which is handy for quick audits or feeding into a monitoring script.
Handling rotated and compressed logs
Most production systems rotate logs daily and compress older ones (auth.log.1.gz, auth.log.2.gz, …).
awk can’t read gzip files directly, so you need a wrapper:
zcat /var/log/auth.log.*.gz | awk '/Failed password/ {print $1, $2, $3, $10, $12}'
If you only care about the last 24 hours, combine it with head:
zcat /var/log/auth.log.*.gz | awk '/Failed password/ {print $1, $2, $3, $10, $12}' | head -n 1000
Adjust the head count based on your failure rate.
Trade‑offs and caveats
| Issue | What to watch for | Mitigation |
|---|---|---|
| Year rollover | The script assumes the current year. If you run it in January and the log contains entries from December of the previous year, the timestamp comparison will fail. | Store the year in the log line (e.g. date +%Y when rotating) or use journalctl which keeps the full timestamp. |
| Different log format | Some distributions prepend a full ISO timestamp (2026-10-10T12:34:56Z). |
Adjust the field numbers or use split() on $1. |
| Large logs | Even though awk is fast, reading a 20‑GB file can still take a few seconds. |
Use journalctl -u sshd --since "1 hour ago" if your system uses systemd’s journal. |
| Security | Printing raw IPs can expose sensitive data. | Pipe the output to a file with restricted permissions (chmod 600) or to an alerting system. |
Integrating with fail2ban
If you’re already running fail2ban, you can feed the awk output into its filter:
awk '/Failed password/ {print $1, $2, $3, $10, $
See also
- Rootless Podman Container Drops Its Persistent Volume After a Reboot: A Quick Fix for My HomeLab 🚀
- A single cron job ate all my RAM and crashed my server: how I added zram and tweaked ulimit
- systemd‑resolved’s DNS Cache Stale After a Router Update: My One‑Command Fix
- How I Use Find and Xargs to Delete Temp Files Older Than 7 Days Without Risking rm ‑rf
- Fixing ACME DNS‑01 on a Home Lab Using Pi‑hole and Cloudflare