Infrastructure

Monitoring and Alerting: What a Small Team Actually Needs

A handful of checks covers most real incidents. What to monitor first, what to alert on, and how to keep alerts from being ignored.

Kiaanlab Engineering Updated October 4, 2026 4 min read
A control room with a desk in front of wall panels

Photo by Miha Meglic on Unsplash

Large companies run monitoring platforms with hundreds of dashboards. A small team does not need that, and trying to copy it usually produces a system nobody looks at. A short set of well-chosen checks catches most real incidents, and the goal is simple: the team should learn about a problem before customers report it.

Start outside the system

The most valuable check is the simplest. From somewhere outside your own infrastructure, request your main page every minute or two and confirm that it answers correctly. Monitoring that runs on the same server as the application goes down with it and reports nothing.

Check more than the home page. A request that goes through the application and the database, such as a login page or a health endpoint that queries the database, catches failures the home page alone would hide.

Certificates and domains

An expired TLS certificate or a lapsed domain registration takes a site down as surely as a crash, and both are entirely predictable. Monitor the expiry dates and alert weeks ahead. Turn on automatic renewal for both, and still monitor, because automatic renewal fails too.

Error tracking

When the application throws an error, capture it with its details in a tool built for that purpose, grouped by cause and with a count. Log files on a server that nobody reads do not serve this role. Error tracking shows which problems are new, which are frequent and which release introduced them.

Resources

Three measurements predict most infrastructure trouble: disk space, memory and processor load. Disk space is the classic one. A disk fills slowly for weeks, then the database stops. An alert at eighty percent gives days of warning for something that is trivial to fix early.

Background work

Scheduled jobs and queues fail quietly. A nightly job that stops running raises no error, because nothing runs. Have each job report when it completes, and alert when the report is missing. For queues, alert when the number of waiting items keeps growing.

Backups

Treat the backup like any other job: alert on failure and on absence, and check that the file size is plausible.

Alert on symptoms

An alert should mean that users are affected or will be soon, and that someone should act. "The site is not responding" and "the error rate has tripled" meet that test. "Processor load is at seventy percent" usually does not, since nothing is wrong and nothing needs doing.

Before adding an alert, ask what the person receiving it would do. If there is no action, it belongs on a dashboard, not in a notification.

Keep the number of alerts low

Every alert that turns out to need no action teaches people to ignore the channel. After a few weeks of that, the real alert is missed among the noise. Review alerts regularly. Remove or adjust any that fired without a reason to act. A team that receives three meaningful alerts a month will respond to each. A team that receives thirty a day will respond to none.

Send alerts where they will be seen

Urgent alerts should reach a person directly, by phone notification or call, and it must be clear who is responsible at any given time. Less urgent ones can go to a team channel. Email alone is too easy to overlook for anything that needs a response within the hour.

Write a short runbook

For each alert, keep a few lines on what it means and what to check first. At three in the morning, a note that says "disk full: old log files are in this folder, clear those first" saves a great deal of time.

Review after incidents

After every outage, ask one question: would monitoring have told us earlier? If the answer is no, add the missing check. Over a year this builds a set of alerts fitted to how your system really fails, which is worth more than any generic template.

An order to do it in

  • External uptime check with an alert to a phone.
  • Certificate and domain expiry.
  • Error tracking.
  • Disk, memory and load.
  • Backup and scheduled job checks.
  • Runbook notes for each alert.

Summary

Watch the service from outside, track errors, monitor the few resources that predict trouble, check silent jobs, and alert only on what needs action. Monitoring is included in every system we deliver through our cloud and DevOps service. If you find out about outages from your customers, we can fix that quickly.

KE

Kiaanlab Engineering

The engineers who design and build Kiaanlab's own AI and software systems, writing about what actually works in production.

Tell us what you're building.

A short call, no sales script, just an honest read on scope and timeline.

Discuss a similar project