An automation that crashes loudly is a nuisance. An automation that stops quietly is a real problem. Orders are not passed on, invoices are not sent, leads are not followed up, and everyone assumes the system is still doing its job. The damage grows until a customer asks where their order is.
Why it happens
Most silent failures come from a short list of causes.
- Expired credentials. An API key or token reaches its end date and every call is rejected from then on.
- A change on the other side. A connected system renames a field or changes a format. The automation still runs and now moves empty or wrong values.
- Errors that are caught and ignored. Code written to "continue on error" keeps going and tells nobody.
- A stuck queue. Work is accepted and never processed, because a worker stopped or one bad item blocks the rest.
- A trigger that stopped firing. The schedule was disabled, or the event that starts the process is no longer sent.
The last one is the hardest to see, because nothing goes wrong. Nothing happens at all.
Alert on failure
The first step is obvious and often missing: when a run fails, someone is told. The alert should go to a place people look, name the automation, say what failed and include the error. An email to a mailbox nobody reads is not an alert.
Alert on silence
Failure alerts do not cover the case where the automation never runs. For that, use a heartbeat. Each successful run reports in. If no report arrives within the expected time, an alert fires. A job that should run every hour and has not been seen for three is flagged, whether it crashed, was disabled or was never started after a server move.
Check the outcome, not only the run
A run can finish without error and still do the wrong thing. Reconciliation catches this. Once a day, compare counts between the systems: orders in the shop against orders in fulfilment, payments received against invoices marked paid. A difference means something was dropped or duplicated, whatever the logs say.
Make retries safe
Temporary problems, such as a timeout or a brief outage, should be retried automatically. For that to be safe, repeating a step must not repeat its effect. Sending an invoice twice or charging a card twice is worse than failing. Give each item a unique reference and have each step check whether it has already handled that reference.
Keep a place for items that cannot be processed
Some items will fail every time: a malformed record, a customer that no longer exists. They should not block everything behind them, and they should not vanish. Move them to a separate list with the reason, alert on it, and let a person decide what to do.
Keep a run history
For every run, record when it started, what it processed, what it changed and how it ended. When someone asks "did the system send that?" the answer should take a minute to find. A history also shows slow changes, such as a job that used to take two minutes and now takes twenty.
Test the alerts
An alert that has never fired is an assumption. Break the automation on purpose in a test environment and confirm that the right person is notified. Repeat this when staff or channels change, because alerts routed to someone who left the company are a common finding.
Summary
Alert on failure and on silence, reconcile the results, make retries safe, keep failed items visible and test that the alerts reach someone. This is the first thing we design in our AI automation projects. If you rely on automations and are not sure you would know when one stops, talk to us.