The same alert, with and without BE AI.
Eight situations every small IT shop knows, written up as they play out with ordinary monitoring and as they play out in BE Hub. Machine names and companies are fictional; the situations are drawn from real incidents during BE Hub's development.
Disk full at 3 am
A threshold alert: D: is at 96%. Someone logs in at 7, opens Explorer, sorts by size, finds the backup folder, deletes by hand, hopes nothing needed those files.
The same alert arrives with a diagnosis: D:\Backups holds 231 GB of nightly .bak files on a 500 GB volume, growing 0.9% a day, everything else normal. A read-only listing is proposed. Sam approves it from the phone, reads the 38 files, approves the delete. BE AI verifies: 54% used. Ten minutes, no login, every step recorded.
Why it was better. The diagnosis and the listing arrive before anyone is awake. The delete is scoped to files a person has seen. Two weeks earlier the fleet row already said "full in 14 days".
Site down, or box down?
An uptime check says the firewall is unreachable. Is the site offline, the carrier, the firewall, or the ISP? Someone rings the office, then the carrier, then drives.
The alert says: no contact since 06:31, last heartbeat healthy, no reboot event, and the site router at the same address still answers. So the firewall or its uplink is down, not the site. The note says which cable to check first.
Why it was better. One alert carries the differential diagnosis a technician would have spent forty minutes assembling.
The slow morning
No alert. Logons took 40 seconds between 8 and 8:20. By the time someone looks, the machine is idle and the event log is a wall of noise.
Asked "why did logons take 40 seconds this morning", BE AI reads the samples: the domain controller sat at 97% CPU from 8:02 to 8:19. Events show the nightly backup job started late at 6:40 after a retry and was still running. Answer in under a minute, with two options and no change proposed because none is needed.
Why it was better. History plus a question beats staring at the event log. The fix is a scheduling decision, and the machine explained it.
The service that stopped after an update
IDS quietly stops after a rules update. Nobody notices for days because the firewall still passes traffic.
The service_running check fails 60 seconds later. The alert carries BE AI's reading of the log: the rules update failed to parse. It proposes a restart after the rules re-download completes; the on-call approves.
Why it was better. A check written in one sentence, an alert with the cause, a fix approved in the same thread.
The mapped drive at reception
The receptionist cannot open the shared drive. She calls at 9:05, when the phones are busiest. Someone remotes in and re-maps it.
A check "S: mapped for reception" runs every minute while she is logged in. At 7:40 it fails; the alert says the file server answers and the mapping is missing from her session. BE AI proposes the re-map; it is approved before she sits down.
Why it was better. The problem is fixed before it is reported. That is the difference between monitoring and preventative maintenance.
The machine nobody looked at
A NAS that has run for a year with nobody checking. The pool is at 78% and climbing.
The nightly pass notes it once: the pool will be full in 54 days at the current rate, a scrub runs nightly and is healthy, nothing to do yet. It says nothing again until the number changes. Seven weeks later there is time to add a disk.
Why it was better. Quiet machines cost nothing to watch and still get a date on the calendar.
The pool that lost a disk
A mirror runs degraded for weeks. df says there is plenty of space. Nobody sees the zpool status until the second disk goes.
The heartbeat carries pool health. ONLINE becomes DEGRADED, a critical alert opens with the pool name, and BE AI reads which device faulted and when. Replacing one disk is a Tuesday job instead of a restore.
Why it was better. The agent reports the thing df cannot see.
The 4G site that comes and goes
Complaints about a branch being slow at random. Nobody can reproduce it. The carrier says the signal is fine.
Every heartbeat from the router carries RSRP, SINR, operator and the data counters. The trend shows the signal dropping every afternoon and the connection type falling back to 3G. BE AI points at the pattern and the hours; the antenna gets moved.
Why it was better. Facts from the router itself, on a trend, instead of a phone call to the carrier.
What the investigator adds, every time.
Context before contact
History first: what changed, since when, how fast. A threshold has none of that.
A cause, not a number
"96%" becomes "old backups, growing 0.9% a day". "Unreachable" becomes "the firewall, not the site".
A proposal, not a login
The fix arrives as a command with a reason. Approve it, and the result comes back in plain language. No screen sharing at 3 am.
Bring your own worst alert.
Put the agent on the machine that pages you most. Ask BE AI why. See what comes back.