- automation
- reliability
- error-handling
- operations
Your Automation Works Fine — Until the One Day It Doesn't
A practical guide to building error-handling and failure alerts into your automations, with a cost breakdown showing why silent failures are more expensive than slow processes.
Kamal Farooqi4 min read
The 2am Problem Nobody Plans For
Most automation projects get built for the happy path. The API responds, the data is formatted correctly, the third field in the spreadsheet isn't blank, and everything flows through exactly like the demo you built. Then, three weeks later, a vendor changes their API response format, or a customer submits a form with an emoji in the name field, and the whole thing breaks quietly in the background.
Nobody notices for four days. By the time someone does, you've got 40 orders that never synced to the warehouse system, or 15 leads that never made it into the CRM. The automation didn't fail loudly — it failed silently, which is worse.
This is the part of automation that doesn't show up in sales pitches: what happens when it breaks, and how fast you find out.
Why This Gets Skipped
Error handling isn't visible in a demo. Nobody asks to see it in the sales call. It takes extra build time, and it's tempting to skip it to hit a deadline or a lower price point. The math seems to favor skipping it — until the first real failure.
Here's a rough comparison based on a typical small-business workflow (order processing, 200 orders/month):
| No error handling | Basic error handling | |
|---|---|---|
| Build time | 8 hours | 11-12 hours |
| Extra build cost | $0 | roughly $300-450 |
| Time to detect a failure | 2-7 days (someone notices manually) | under 15 minutes (alert fires) |
| Orders affected per incident | 20-50 | 1-3 |
| Cost to manually fix a backlog | 3-6 hours of staff time | 15-30 minutes |
| Customer-facing impact | Delayed orders, refund requests, lost trust | Minimal to none |
The extra few hundred dollars up front is cheap insurance. The real cost isn't the build — it's the multi-day gap between "something broke" and "someone knows."
What Good Error Handling Actually Looks Like
You don't need anything exotic. Most reliability problems in small-business automation come down to four habits.
1. Validate inputs before you act on them. If a webhook sends you a record missing a required field, don't let it flow through and create a broken entry downstream. Catch it, log it, and stop it there.
2. Retry transient failures automatically. APIs time out. Rate limits get hit. A retry with a short delay (say, three attempts over two minutes) quietly resolves a large share of failures without anyone ever knowing there was an issue.
3. Alert a human when retries don't work. After the automatic retries are exhausted, send a message somewhere a person will actually see it — a Slack channel, an email, a text. Not a log file nobody opens. The alert should say what failed and with which record, not just "error occurred."
4. Keep a dead-letter queue. Failed items shouldn't vanish. Park them somewhere (a dedicated Airtable table, a database table, a tagged folder) so they can be reprocessed once the root cause is fixed, instead of manually hunting through logs to figure out what got missed.
None of this requires a dedicated engineering team. In tools like Make or n8n, this is a matter of adding error-handling branches and a notification step — typically 20-30% more build time, not double.
A Concrete Example
Take a generic example: a home services company automates appointment booking from their website form into their scheduling system and CRM. Average volume: 300 bookings a month.
Without error handling, a malformed phone number format breaks the sync for any booking submitted that way. Over one bad week, 22 bookings silently fail to create CRM records. The office manager notices when a customer calls asking why nobody confirmed their appointment. It takes her an afternoon — about 3 hours — to reconstruct the missing bookings from form submission emails, at a loaded cost of roughly $60-75. Two of the 22 customers don't get followed up in time and book with a competitor instead. Call that $400-600 in lost revenue, conservatively.
With basic validation and an alert, the same bad phone number format triggers a Slack message within minutes: "Booking failed — invalid phone format — record attached." Someone fixes it in under five minutes, and the other 21 similarly formatted entries from that week get caught and corrected before any customer notices.
Same bug. Completely different outcome. The difference wasn't more sophisticated automation — it was 3 extra hours of build work addressing failure cases up front.
The Real Takeaway
If you're evaluating an automation build — your own or a vendor's — ask one question before anything else: "What happens when this fails, and how would we find out?" If the answer is "we'd notice eventually," that's the gap worth closing before you scale the volume running through it.
Reliability isn't a nice-to-have feature you add later. It's the difference between an automation that saves you money and one that quietly creates a new problem you haven't discovered yet.
If you want a second set of eyes on how your current automations handle failure — or you're building something new and want it done right the first time — get in touch and we'll walk through it together.