How do I know if my automation stopped working?

The automations that cost you money don't fail loudly, they stop running, and your error alerts are watching the wrong thing.

An isometric row of automation stages with only the last one lit and the earlier stages dark

Say you run a six-van commercial glazing firm in Birmingham. Quote requests arrive through a form on the site, land in Pipedrive, and fire a WhatsApp message to whoever is on call. You built it in Make about eighteen months ago. It has been fine. Then a quantity surveyor rings on a Tuesday to ask why nobody got back to him, and you find eleven quote requests sitting unread in the form's own storage, going back nine days.

No red dots. No error emails. The automation stopped working and the dashboard shows a clean fortnight.

Most automations don't fail loudly, they stop, and you find out from a customer. Make, Zapier and n8n all alert you when a run fails, and the failures that cost real money are the ones where no run happens at all, so there's nothing to alert on. The fix is to monitor for absence rather than for errors: make every automation report in on a schedule, and treat silence as the alarm.

Not with a crash, then. With a shrug.

Why don't I get an alert when my automation stops working?

Because the alert is wired to the wrong event. An error means a step ran and failed, and a step running means the trigger fired. When the trigger never fires there's no run, no error, and nothing to send you an email about. It was a clean fortnight in the sense that a stopped clock keeps perfect time.

Three things cause it, roughly in order of how often they turn up.

The connection expired. Most tools talk to each other over OAuth 2.0, the standard that lets one app act on your account without holding your password. It issues a short-lived access token plus a refresh token that renews it in the background.

Plenty of platforms renew forever and you never think about it. Others enforce a hard expiry, or revoke everything when someone changes a password or an admin deletes a user. Then the trigger quietly stops returning data.

The thing it watches changed shape. Someone in marketing renames a form field from Phone to Contact number. Someone in sales renames the Pipedrive stage from New enquiry to Enquiry. The automation keeps running on its schedule and keeps matching nothing, because the filter still looks for the old value. A filter that excludes every record looks exactly like a quiet week.

The platform switched it off. Make deactivates a scheduled scenario after three consecutive errors, and Zapier will disable a Zap that errors on nearly every run for long enough. That behaviour is correct. A broken automation hammering someone's API every minute is worse than a stopped one. It also turns a loud failure into an absence, and absences are invisible.

The automation that runs perfectly and does nothing

The nastier version passes every check. A webhook (a URL one system calls to push data into another the moment something happens) gets re-pointed during a website rebuild, and now hits an endpoint that returns 200 and throws the body away. The sender records a success. The receiver never sees a thing. Every log is green.

It looks healthier than the automations that are genuinely working, because working ones pick up the odd retry and the odd timeout in their history. A spotless log isn't always good news. The same blind spot catches contact forms: the form says thank you, the mail server accepts the message, and the email lands in spam where nobody reads it.

How do I know if my automation is working right now?

Give it a heartbeat, and alert on the heartbeat going missing. The pattern is called a dead man's switch: instead of waiting for something to report a failure, you decide in advance what a healthy day looks like, and you get told when it does not arrive.

Two pieces, and both are free at small-business volumes.

  • A ping on every successful run. Add a final step that makes an HTTP request to a monitoring URL. Healthchecks.io is the plainest version: you tell it how often to expect a ping and how much lateness to forgive, and it notifies you when the ping doesn't come. Cronitor does the same job.
  • A daily receipt in business units. One line, written where a human will read it: 14 quote requests routed, 3 invoices filed, 0 WhatsApp messages sent. Not run counts. Business units.

The receipt is the half most people skip and the half that catches the 200-and-discard failure. Write it to a Google Sheet, one row a day, and the sheet becomes a baseline you can read at a glance.

Fourteen, eleven, sixteen, twelve, zero. Your eye finds the zero before any monitoring tool decides it counts as a fault, because a run that happened and did nothing isn't a fault to the platform at all.

A ping proves the automation ran. Only the receipt proves it did something.

A route through a chain of automation steps, with several nodes left dark and unreachable off the main path.
A route through a chain of automation steps, with several nodes left dark and unreachable off the main path.

What should I check first when an automation stopped working?

Check whether the trigger fired at all, before anything else. Most of the time you can stop there. The order below runs from most likely to least and takes about ten minutes.

  1. Look at the toggle, and at who moved it. Make, Zapier and n8n all record a deactivation.
  2. Count the runs, not the errors. Zapier keeps a per-run history for every Zap, Make has execution history, n8n has an executions list. Nothing since the ninth is a different problem from runs that failed.
  3. Open the connected account and look for a reconnect prompt. If one app needs reauthorising, check every other automation using that same login.
  4. Compare a filter's value against a record created today, character for character. Trailing spaces and renamed picklists cause most of these.
  5. Go and look at the destination record. A 2xx response (the success range of HTTP status codes) isn't proof the data arrived in a usable field, which is the same reason a CRM integration can look real and not be.

One nuance on counting runs. Most triggers poll rather than listen, so the platform checks for new data on an interval set by your plan, somewhere between one and fifteen minutes. If your automation looks slow rather than dead, you're probably reading a polling gap as a fault.

What is going to break a working automation in the next two years?

Authentication, mostly, and the dates are already published. Platforms retire the old way of logging in, and anything still using it stops on a known day.

The clearest one on the calendar: Salesforce is retiring the SOAP API login() call, the endpoint that swaps a username and password for a session, in API versions 31.0 through 64.0, effective with the Summer '27 release, which lands in June 2027. Any integration that authenticates by posting a username and password to that endpoint stops authenticating. Anything already on OAuth 2.0 carries on untouched. Salesforce notified customers in April 2026, so there is roughly a year of warning, and that warning goes to an IT contact who probably doesn't know the little Make scenario exists.

So do the boring thing now. List every automation you run, which accounts it authenticates with, and how. That list takes an afternoon, and it's what turns a deprecation notice into a ten-minute job instead of a week of forensics. Picking what deserves that care is its own question, and our post on the first process worth automating is the place to start.

Where absence monitoring does not help

Heartbeat monitoring tells you an automation ran. It doesn't tell you the automation is right. If your lead router sends every enquiry to the wrong estimator, it will ping cheerfully for a year and the receipt will show a healthy number next to a wrong outcome. Absence is what this catches. Bad logic needs a human reading the thing, quarterly.

It also won't get your data back. When a webhook is dropped, the payload is usually gone: the sender retried, failed, and moved on. Sometimes we can rebuild a window from the source system's own records. Sometimes we cannot, and then the only thing left to do is ring the people who filled in the form and admit it.

And the part that costs us something. If you run two automations and take ten enquiries a week, don't hire anyone for this. Open the CRM every Monday at 9am, count the new records, compare it to your gut. That habit is free and catches the same failures.

Monitoring earns its keep once nobody can hold the expected numbers in their head any more. There's a simple test: if you can't say out loud roughly how many records each automation should create this week, you've passed the line. If you can, you haven't, and we'd be selling you plumbing for a house with one tap.

Set up one automation alert this week

Pick the automation that would hurt most if it stopped, which is almost always the one moving new enquiries. Add a final step that pings a free Healthchecks.io check on success, set the expected period to match your real volume, and send the alert to a phone instead of an inbox. Twenty minutes, and your worst silent failure becomes a text message.

If you would rather not be the person who remembers to look, that is the job we do. We build workflow automation with the monitoring inside it, and we audit what other people built. Tell us what you are running and we will tell you where it is most likely to go quiet.

Common questions

Still wondering

Can I get a text message when an automation fails?

Yes, and you should, but wire it to absence rather than to errors. Point a monitoring service like Healthchecks.io or Cronitor at an SMS or push channel, set the expected ping period to match your real volume, and you get told when a run does not happen. Error-only alerts into an inbox get filtered, batched and ignored, which is how a nine-day gap goes unnoticed.

Why did my Zap work when I tested it but not on its own?

Testing runs the steps with data you hand it, which skips the part that usually breaks: the trigger finding new records by itself. A manual test passes while the polling query returns nothing, the webhook points somewhere stale, or the connected account needs reauthorising. Always confirm a real record created end to end, from the form or the inbox, not from the editor's test button.

How often should I check my automations manually?

Monthly for the logs, quarterly for the logic. The monthly pass looks for runs that stopped, connections showing a reconnect prompt, and filters matching nothing. The quarterly pass is harder and more valuable: read what the automation actually does and confirm it still matches how the business works now. Renamed stages, new staff and changed pricing break correct automations without breaking them technically.

Does self-hosted n8n make silent failures better or worse?

Better on control, worse on defaults. Self-hosting gives you the execution database, server logs and a reusable error workflow you assign once, so you can see more than a hosted plan shows you. Nothing notifies you out of the box, though, and an out-of-memory kill or a container restart leaves no application-level error at all. Heartbeat monitoring matters more here, not less.