journalctl -p err gives you 4,000 lines; what you actually need is the six distinct things that went wrong, ranked by severity, with counts. This script normalises PIDs, IPs and paths out of every message so identical failures collapse into one row.
Full script + README on GitHub: ruby-devops-toolkit/journald-error-digest
Step through the build below:
Every Linux box you run already has a structured log store: journald. The problem is not collecting the data, it is reading it. After an incident, journalctl -p err --since yesterday dumps thousands of lines, and 95% of them are the same three failures repeated with a different PID, client IP or inode number.
The digest script asks journald for JSON, strips the variable parts out of each message to build a signature, then groups by [unit, signature]. You get a table of distinct problems, their counts, the worst priority seen, and when each was last seen. It exits non-zero when anything is above the threshold, so it drops straight into cron or a CI health gate.
#!/usr/bin/env ruby
# frozen_string_literal: true
#
# journald_error_digest.rb - Turn a wall of journalctl noise into a ranked,
# de-duplicated digest of what actually went wrong on a Linux box.
#
# The script shells out to `journalctl -o json` (one JSON object per line),
# keeps entries at or above a priority threshold (default: err), collapses
# near-identical messages into "signatures" (numbers, PIDs, hex IDs and
# paths are normalised away), then ranks them by unit and by frequency.
#
# Usage:
# ruby journald_error_digest.rb # last 24h, priority <= err
# ruby journald_error_digest.rb --since "2 hours ago" --priority warning
# ruby journald_error_digest.rb --boot # current boot only
# ruby journald_error_digest.rb --json > digest.json
# journalctl -o json --since yesterday | ruby journald_error_digest.rb --stdin
#
# Exit codes: 0 = nothing above threshold, 1 = errors found, 2 = usage/runtime error.
#
# Tested with Ruby 3.0+ on Ubuntu 22.04 (systemd 249). Only stdlib is used.
require 'json'
require 'optparse'
require 'open3'
require 'time'
PRIORITIES = {
'emerg' => 0, 'alert' => 1, 'crit' => 2, 'err' => 3,
'warning' => 4, 'notice' => 5, 'info' => 6, 'debug' => 7
}.freeze
PRIORITY_NAMES = PRIORITIES.invert.freeze
# ---------------------------------------------------------------------------
# Options
# ---------------------------------------------------------------------------
options = {
since: '24 hours ago',
priority: 'err',
boot: false,
stdin: false,
json: false,
top: 15,
unit: nil
}
OptionParser.new do |o|
o.banner = 'Usage: journald_error_digest.rb [options]'
o.on('--since WHEN', 'journalctl --since expression (default: "24 hours ago")') { |v| options[:since] = v }
o.on('--priority LEVEL', PRIORITIES.keys, "Highest numeric priority to keep (#{PRIORITIES.keys.join('|')})") { |v| options[:priority] = v }
o.on('--boot', 'Restrict to the current boot') { options[:boot] = true }
o.on('--unit UNIT', 'Only this systemd unit (e.g. nginx.service)') { |v| options[:unit] = v }
o.on('--top N', Integer, 'Show the N most frequent signatures (default 15)') { |v| options[:top] = v }
o.on('--stdin', 'Read journalctl -o json output from STDIN instead of running journalctl') { options[:stdin] = true }
o.on('--json', 'Emit the digest as JSON for downstream tooling') { options[:json] = true }
o.on('-h', '--help') { puts o; exit 0 }
end.parse!
# ---------------------------------------------------------------------------
# Collecting entries
# ---------------------------------------------------------------------------
def build_journalctl_cmd(opts)
cmd = %w[journalctl -o json --no-pager -q]
cmd += ['-p', opts[:priority]] # journalctl filters 0..N for us
cmd += ['--since', opts[:since]] unless opts[:boot]
cmd << '-b' if opts[:boot]
cmd += ['-u', opts[:unit]] if opts[:unit]
cmd
end
def read_entries(opts)
raw = if opts[:stdin]
$stdin.read
else
out, err, status = Open3.capture3(*build_journalctl_cmd(opts))
unless status.success?
warn "journalctl failed (exit #{status.exitstatus}): #{err.strip}"
exit 2
end
out
end
raw.each_line.filter_map do |line|
line = line.strip
next if line.empty?
JSON.parse(line)
rescue JSON::ParserError
nil # journald can emit binary/blob fields; skip anything unparseable
end
end
# ---------------------------------------------------------------------------
# Normalising a message into a signature
# ---------------------------------------------------------------------------
# Two log lines that differ only by a PID, an IP, a timestamp or a hex ID are
# the same *problem*. Collapsing them is what makes the digest readable.
def signature_for(message)
sig = message.dup
sig = sig.gsub(/\b\d{1,3}(?:\.\d{1,3}){3}(?::\d+)?\b/, '<ip>') # IPv4[:port]
sig = sig.gsub(/\b0x[0-9a-f]+\b/i, '<hex>') # hex addresses
sig = sig.gsub(/\b[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}\b/i, '<uuid>')
sig = sig.gsub(%r{(/[\w.\-]+){2,}}, '<path>') # filesystem paths
sig = sig.gsub(/\b\d+(?:\.\d+)?(?:ms|s|MB|KB|GB|%)?\b/, '<n>') # bare numbers / durations
sig.squeeze(' ').strip[0, 160]
end
def field(entry, *names)
names.each do |n|
v = entry[n]
return v if v.is_a?(String) && !v.empty?
end
nil
end
def entry_time(entry)
usec = entry['__REALTIME_TIMESTAMP'].to_i
usec.zero? ? nil : Time.at(usec / 1_000_000.0)
end
# ---------------------------------------------------------------------------
# Aggregation
# ---------------------------------------------------------------------------
def aggregate(entries, max_priority)
by_sig = Hash.new { |h, k| h[k] = { count: 0, first: nil, last: nil, priority: 7, unit: nil, sample: nil } }
by_unit = Hash.new(0)
by_priority = Hash.new(0)
entries.each do |e|
prio = e['PRIORITY'].to_i
next if prio > max_priority
msg = field(e, 'MESSAGE') || '(no MESSAGE field)'
unit = field(e, '_SYSTEMD_UNIT', 'UNIT', 'SYSLOG_IDENTIFIER', '_COMM') || 'kernel'
ts = entry_time(e)
key = [unit, signature_for(msg)]
s = by_sig[key]
s[:count] += 1
s[:unit] = unit
s[:sample] ||= msg
s[:priority] = [s[:priority], prio].min
s[:first] = ts if ts && (s[:first].nil? || ts < s[:first])
s[:last] = ts if ts && (s[:last].nil? || ts > s[:last])
by_unit[unit] += 1
by_priority[PRIORITY_NAMES[prio] || prio.to_s] += 1
end
{
signatures: by_sig.map { |(unit, sig), v| v.merge(unit: unit, signature: sig) }
.sort_by { |v| [v[:priority], -v[:count]] },
units: by_unit.sort_by { |_, c| -c },
priorities: by_priority.sort_by { |name, _| PRIORITIES[name] || 99 },
total: by_sig.values.sum { |v| v[:count] }
}
end
# ---------------------------------------------------------------------------
# Rendering
# ---------------------------------------------------------------------------
def fmt_time(t)
t ? t.strftime('%m-%d %H:%M:%S') : '--'
end
def render_text(digest, opts)
puts "journald error digest (since: #{opts[:boot] ? 'this boot' : opts[:since]}, priority <= #{opts[:priority]})"
puts '=' * 78
if digest[:total].zero?
puts 'No journal entries at or above the requested priority. Nice.'
return
end
puts "Total entries: #{digest[:total]} Distinct problems: #{digest[:signatures].size}"
puts
puts 'By priority:'
digest[:priorities].each { |name, c| puts format(' %-8s %6d', name, c) }
puts
puts 'Noisiest units:'
digest[:units].first(8).each { |unit, c| puts format(' %-40s %6d', unit[0, 40], c) }
puts
puts "Top #{opts[:top]} problems (ranked by severity, then frequency):"
puts format(' %-5s %-7s %-26s %-19s %s', 'COUNT', 'PRIO', 'UNIT', 'LAST SEEN', 'MESSAGE (sample)')
digest[:signatures].first(opts[:top]).each do |s|
puts format(' %-5d %-7s %-26s %-19s %s',
s[:count], PRIORITY_NAMES[s[:priority]], s[:unit][0, 26],
fmt_time(s[:last]), s[:sample][0, 70])
end
end
def render_json(digest, opts)
out = {
generated_at: Time.now.iso8601,
since: opts[:boot] ? 'boot' : opts[:since],
max_priority: opts[:priority],
total: digest[:total],
by_priority: digest[:priorities].to_h,
by_unit: digest[:units].to_h,
problems: digest[:signatures].first(opts[:top]).map do |s|
s.merge(priority: PRIORITY_NAMES[s[:priority]],
first: s[:first]&.iso8601, last: s[:last]&.iso8601)
end
}
puts JSON.pretty_generate(out)
end
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
entries = read_entries(options)
digest = aggregate(entries, PRIORITIES[options[:priority]])
options[:json] ? render_json(digest, options) : render_text(digest, options)
exit(digest[:total].zero? ? 0 : 1)
JSON, not text. journalctl -o json emits one object per line with every field journald knows (_SYSTEMD_UNIT, PRIORITY, __REALTIME_TIMESTAMP). Parsing that is trivial compared to regexing the human format, and the -p flag lets journald do the priority filter before Ruby ever sees the data.
Signatures are the whole trick. signature_for replaces IPv4 addresses, hex values, UUIDs, filesystem paths and bare numbers with placeholders. “attempts exceeded for root from 185.220.3.203 port 19335” and the same line from another IP become one signature with a count of 41. Order matters: IPs are replaced before the generic number rule so they don’t get chopped into four <n> tokens.
Rank by severity, then frequency. A single crit EXT4 error matters more than 41 err SSH brute-force lines. The sort key is [min_priority, -count].
--stdin for testability. The sandbox this was tested in cannot read the journal (no permissions), so the script accepts journalctl -o json output on stdin. That is also how you run it against a journal exported from another host.
journald error digest (since: 24 hours ago, priority <= warning) ============================================================================== Total entries: 87 Distinct problems: 6 By priority: crit 3 err 70 warning 14 Noisiest units: sshd.service 41 nginx.service 27 [email protected] 9 docker.service 5 kernel 3 systemd 2 Top 6 problems (ranked by severity, then frequency): COUNT PRIO UNIT LAST SEEN MESSAGE (sample) 3 crit kernel 09-06 09:34:30 EXT4-fs error (device sda1): ext4_find_entry:1450: inode #303051: comm 41 err sshd.service 09-06 15:35:12 error: maximum authentication attempts exceeded for root from 185.220. 27 err nginx.service 09-06 14:40:25 connect() failed (111: Connection refused) while connecting to upstrea 2 err systemd 09-06 15:29:07 backup-nightly.service: Failed with result 'exit-code'. 9 warning [email protected] 09-06 14:37:06 checkpoints are occurring too frequently (21 seconds apart) 5 warning docker.service 09-06 15:31:19 failed to retrieve docker-runc version: exec: "docker-runc": executabl exit code: 1 (problems found)
Prerequisites
- Ruby 3.0+ (stdlib only:
json,optparse,open3,time). No gems. - A systemd-based Linux (tested on Ubuntu 22.04, systemd 249). Any distro with
journalctlworks. - Permission to read the journal: run as root, or add your user to the
systemd-journalgroup.
Full source for reference
Usage:
ruby journald_error_digest.rb # last 24h, priority <= err
ruby journald_error_digest.rb --since "2 hours ago" --priority warning
ruby journald_error_digest.rb --boot --unit nginx.service --top 5
ruby journald_error_digest.rb --json > digest.json
journalctl -o json --since yesterday | ruby journald_error_digest.rb --stdin
#!/usr/bin/env ruby
# frozen_string_literal: true
#
# journald_error_digest.rb - Turn a wall of journalctl noise into a ranked,
# de-duplicated digest of what actually went wrong on a Linux box.
#
# The script shells out to `journalctl -o json` (one JSON object per line),
# keeps entries at or above a priority threshold (default: err), collapses
# near-identical messages into "signatures" (numbers, PIDs, hex IDs and
# paths are normalised away), then ranks them by unit and by frequency.
#
# Usage:
# ruby journald_error_digest.rb # last 24h, priority <= err
# ruby journald_error_digest.rb --since "2 hours ago" --priority warning
# ruby journald_error_digest.rb --boot # current boot only
# ruby journald_error_digest.rb --json > digest.json
# journalctl -o json --since yesterday | ruby journald_error_digest.rb --stdin
#
# Exit codes: 0 = nothing above threshold, 1 = errors found, 2 = usage/runtime error.
#
# Tested with Ruby 3.0+ on Ubuntu 22.04 (systemd 249). Only stdlib is used.
require 'json'
require 'optparse'
require 'open3'
require 'time'
PRIORITIES = {
'emerg' => 0, 'alert' => 1, 'crit' => 2, 'err' => 3,
'warning' => 4, 'notice' => 5, 'info' => 6, 'debug' => 7
}.freeze
PRIORITY_NAMES = PRIORITIES.invert.freeze
# ---------------------------------------------------------------------------
# Options
# ---------------------------------------------------------------------------
options = {
since: '24 hours ago',
priority: 'err',
boot: false,
stdin: false,
json: false,
top: 15,
unit: nil
}
OptionParser.new do |o|
o.banner = 'Usage: journald_error_digest.rb [options]'
o.on('--since WHEN', 'journalctl --since expression (default: "24 hours ago")') { |v| options[:since] = v }
o.on('--priority LEVEL', PRIORITIES.keys, "Highest numeric priority to keep (#{PRIORITIES.keys.join('|')})") { |v| options[:priority] = v }
o.on('--boot', 'Restrict to the current boot') { options[:boot] = true }
o.on('--unit UNIT', 'Only this systemd unit (e.g. nginx.service)') { |v| options[:unit] = v }
o.on('--top N', Integer, 'Show the N most frequent signatures (default 15)') { |v| options[:top] = v }
o.on('--stdin', 'Read journalctl -o json output from STDIN instead of running journalctl') { options[:stdin] = true }
o.on('--json', 'Emit the digest as JSON for downstream tooling') { options[:json] = true }
o.on('-h', '--help') { puts o; exit 0 }
end.parse!
# ---------------------------------------------------------------------------
# Collecting entries
# ---------------------------------------------------------------------------
def build_journalctl_cmd(opts)
cmd = %w[journalctl -o json --no-pager -q]
cmd += ['-p', opts[:priority]] # journalctl filters 0..N for us
cmd += ['--since', opts[:since]] unless opts[:boot]
cmd << '-b' if opts[:boot]
cmd += ['-u', opts[:unit]] if opts[:unit]
cmd
end
def read_entries(opts)
raw = if opts[:stdin]
$stdin.read
else
out, err, status = Open3.capture3(*build_journalctl_cmd(opts))
unless status.success?
warn "journalctl failed (exit #{status.exitstatus}): #{err.strip}"
exit 2
end
out
end
raw.each_line.filter_map do |line|
line = line.strip
next if line.empty?
JSON.parse(line)
rescue JSON::ParserError
nil # journald can emit binary/blob fields; skip anything unparseable
end
end
# ---------------------------------------------------------------------------
# Normalising a message into a signature
# ---------------------------------------------------------------------------
# Two log lines that differ only by a PID, an IP, a timestamp or a hex ID are
# the same *problem*. Collapsing them is what makes the digest readable.
def signature_for(message)
sig = message.dup
sig = sig.gsub(/\b\d{1,3}(?:\.\d{1,3}){3}(?::\d+)?\b/, '<ip>') # IPv4[:port]
sig = sig.gsub(/\b0x[0-9a-f]+\b/i, '<hex>') # hex addresses
sig = sig.gsub(/\b[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}\b/i, '<uuid>')
sig = sig.gsub(%r{(/[\w.\-]+){2,}}, '<path>') # filesystem paths
sig = sig.gsub(/\b\d+(?:\.\d+)?(?:ms|s|MB|KB|GB|%)?\b/, '<n>') # bare numbers / durations
sig.squeeze(' ').strip[0, 160]
end
def field(entry, *names)
names.each do |n|
v = entry[n]
return v if v.is_a?(String) && !v.empty?
end
nil
end
def entry_time(entry)
usec = entry['__REALTIME_TIMESTAMP'].to_i
usec.zero? ? nil : Time.at(usec / 1_000_000.0)
end
# ---------------------------------------------------------------------------
# Aggregation
# ---------------------------------------------------------------------------
def aggregate(entries, max_priority)
by_sig = Hash.new { |h, k| h[k] = { count: 0, first: nil, last: nil, priority: 7, unit: nil, sample: nil } }
by_unit = Hash.new(0)
by_priority = Hash.new(0)
entries.each do |e|
prio = e['PRIORITY'].to_i
next if prio > max_priority
msg = field(e, 'MESSAGE') || '(no MESSAGE field)'
unit = field(e, '_SYSTEMD_UNIT', 'UNIT', 'SYSLOG_IDENTIFIER', '_COMM') || 'kernel'
ts = entry_time(e)
key = [unit, signature_for(msg)]
s = by_sig[key]
s[:count] += 1
s[:unit] = unit
s[:sample] ||= msg
s[:priority] = [s[:priority], prio].min
s[:first] = ts if ts && (s[:first].nil? || ts < s[:first])
s[:last] = ts if ts && (s[:last].nil? || ts > s[:last])
by_unit[unit] += 1
by_priority[PRIORITY_NAMES[prio] || prio.to_s] += 1
end
{
signatures: by_sig.map { |(unit, sig), v| v.merge(unit: unit, signature: sig) }
.sort_by { |v| [v[:priority], -v[:count]] },
units: by_unit.sort_by { |_, c| -c },
priorities: by_priority.sort_by { |name, _| PRIORITIES[name] || 99 },
total: by_sig.values.sum { |v| v[:count] }
}
end
# ---------------------------------------------------------------------------
# Rendering
# ---------------------------------------------------------------------------
def fmt_time(t)
t ? t.strftime('%m-%d %H:%M:%S') : '--'
end
def render_text(digest, opts)
puts "journald error digest (since: #{opts[:boot] ? 'this boot' : opts[:since]}, priority <= #{opts[:priority]})"
puts '=' * 78
if digest[:total].zero?
puts 'No journal entries at or above the requested priority. Nice.'
return
end
puts "Total entries: #{digest[:total]} Distinct problems: #{digest[:signatures].size}"
puts
puts 'By priority:'
digest[:priorities].each { |name, c| puts format(' %-8s %6d', name, c) }
puts
puts 'Noisiest units:'
digest[:units].first(8).each { |unit, c| puts format(' %-40s %6d', unit[0, 40], c) }
puts
puts "Top #{opts[:top]} problems (ranked by severity, then frequency):"
puts format(' %-5s %-7s %-26s %-19s %s', 'COUNT', 'PRIO', 'UNIT', 'LAST SEEN', 'MESSAGE (sample)')
digest[:signatures].first(opts[:top]).each do |s|
puts format(' %-5d %-7s %-26s %-19s %s',
s[:count], PRIORITY_NAMES[s[:priority]], s[:unit][0, 26],
fmt_time(s[:last]), s[:sample][0, 70])
end
end
def render_json(digest, opts)
out = {
generated_at: Time.now.iso8601,
since: opts[:boot] ? 'boot' : opts[:since],
max_priority: opts[:priority],
total: digest[:total],
by_priority: digest[:priorities].to_h,
by_unit: digest[:units].to_h,
problems: digest[:signatures].first(opts[:top]).map do |s|
s.merge(priority: PRIORITY_NAMES[s[:priority]],
first: s[:first]&.iso8601, last: s[:last]&.iso8601)
end
}
puts JSON.pretty_generate(out)
end
# ---------------------------------------------------------------------------
# Main
# ---------------------------------------------------------------------------
entries = read_entries(options)
digest = aggregate(entries, PRIORITIES[options[:priority]])
options[:json] ? render_json(digest, options) : render_text(digest, options)
exit(digest[:total].zero? ? 0 : 1)
Step-by-step walkthrough
1. Ask journald for JSON
build_journalctl_cmd assembles journalctl -o json --no-pager -q -p err --since '24 hours ago'. Passing -p means journald filters priority 0..N itself; Ruby only parses what matters. --boot swaps --since for -b, and --unit adds -u.
2. Parse defensively
read_entries runs the command with Open3.capture3 so stderr is captured separately, then JSON.parses each line inside a rescue JSON::ParserError. journald can embed binary blobs for some fields and those lines are simply skipped.
3. Normalise into a signature
signature_for is a chain of gsub calls: IPv4 (with optional port) to <ip>, 0x... to <hex>, UUIDs to <uuid>, two-or-more-segment paths to <path>, then any remaining number (optionally with ms/s/MB/% suffix) to <n>. It is truncated to 160 chars so very long messages don’t produce unique keys.
4. Aggregate
aggregate uses a Hash.new with a default block to build a record per [unit, signature]: count, first/last seen (from __REALTIME_TIMESTAMP, microseconds since epoch), the lowest (worst) priority, and the first raw message as a human-readable sample. Side tallies by unit and by priority feed the summary sections.
5. Render and exit
Text mode prints a priority breakdown, the noisiest units, and a top-N table. --json emits the same data with ISO-8601 timestamps for a dashboard or alerting hook. Exit 0 means the journal was clean, 1 means problems were found, 2 means journalctl itself failed.
What a run looks like
Troubleshooting
- “No journal files were opened due to insufficient permissions” (exit 2): you are not root and not in
systemd-journal.sudo usermod -aG systemd-journal $USER, then log in again. This is exactly what the Linux sandbox returned during testing, which is why the shown output was produced from a syntheticjournalctl -o jsonfixture piped through--stdin. - Everything shows as unit
kernel. Kernel messages have no_SYSTEMD_UNIT; that is expected. If user-space messages also land there, your journald is not recording_SYSTEMD_UNIT(containers with a shared journal do this). The fallback chain also triesUNIT,SYSLOG_IDENTIFIERand_COMM. - Counts look too low. The
--sincedefault is 24 hours and--prioritydefaults toerr, so warnings are excluded. Try--priority warning. - Two rows that are obviously the same problem. The message contains a variable token the normaliser doesn’t know about (a hostname, a username). Add one more
gsubtosignature_for. - Timestamps are off by hours. They are rendered in the local timezone of the machine running the script, not the host that produced the journal. Use
--jsonfor zone-aware ISO-8601 output.
Extending the script
- Wire it into cron:
0 7 * * * journald_error_digest.rb --since '24 hours ago' || mail -s "$(hostname) journal digest" [email protected]. Exit 1 triggers the mail only when there is something to read. - Fleet mode: run
journalctl -o json --since yesterdayover SSH on each host and pipe into--stdin; add a_HOSTNAMEcolumn to the key to keep hosts separate. - Baseline diffing: save
--jsonoutput daily and alert only on signatures that did not appear yesterday. - Push the
by_unithash to Prometheus via a textfile collector so unit error rates show up on a Grafana panel. - Add
--ignore REGEXto suppress known-noisy signatures (that SSH brute-force line, for example) without losing them from the totals.