- automation
- error-handling
- reliability
- operations
Your Automation Will Fail. Here's How to Make Sure You Find Out Before Your Client Does
Most small businesses build automations that work great on day one and fail silently three weeks later. Here's how to build error-handling in from the start, with real numbers on what silent failures cost.
Kamal Farooqi5 min read
The automation that worked for six weeks
Here's a scenario we see constantly: a business builds a workflow that pulls new leads from a web form, enriches them, and drops them into the CRM with a task assigned to a rep. It runs perfectly for six weeks. Then an API key expires, or a field gets renamed in a third-party app, or a rate limit gets hit during a busy afternoon. The workflow doesn't crash with an alarm bell — it just stops, or worse, it keeps running and silently drops records.
Nobody notices for 11 days. That's not a made-up number — it's roughly the average gap we see between a workflow breaking and someone noticing, in businesses that don't have monitoring built in. By the time it's caught, there are 40-60 leads sitting in limbo, a client complaint has already landed, or an invoice never got sent.
The automation itself wasn't the problem. The lack of a plan for when it breaks was.
Why automations fail silently
Most small business automation tools (Zapier, Make, n8n) are built to run quietly in the background. That's the appeal — no one wants to babysit a workflow. But "quiet" cuts both ways. If you don't explicitly tell the system what to do when something goes wrong, the default behavior is usually one of these:
- The run errors out and just... stops. No notification unless you set one up.
- The run "succeeds" technically but passes along bad or empty data, because the upstream API returned something unexpected and nothing caught it.
- The run retries forever on a loop, burning task quota or API calls until someone notices a bill.
None of these are dramatic. That's exactly why they're dangerous — a loud failure gets fixed in an hour, a quiet one gets fixed when a client calls asking why they never got an invoice.
What reliable automation actually looks like
Reliability isn't about building something that never breaks. Nothing automated runs forever without hiccups — APIs change, services have outages, someone edits a spreadsheet column by hand. Reliability is about shrinking the time between "something broke" and "someone competent knows about it."
That comes down to four habits, and none of them are expensive to add:
1. Alert on failure, not just on success
Every workflow should have a path for errors that sends a Slack message, email, or text to a real person — not just a log entry buried in a dashboard nobody opens. This is usually a 10-minute addition per workflow.
2. Validate data before it moves, not after
If a workflow expects an email field and gets a blank one, don't let it pass through to the CRM. Add a check: if a required field is missing or malformed, route it to a "needs review" queue instead of pretending it succeeded.
3. Log what happened, with enough detail to debug fast
A line that says "error" tells you nothing. A line that says "Lead ID 4821 — HubSpot API returned 429 (rate limit) at 2:14pm" tells you exactly where to look. Good logging is the difference between a 5-minute fix and a 2-hour investigation.
4. Build in retries, with limits
Transient issues (a slow API, a momentary timeout) shouldn't trigger a full failure alert. Retry two or three times with a short delay first. But cap it — a retry loop with no limit is how you rack up thousands of wasted task executions.
What this actually costs, in both directions
| No error-handling | Basic error-handling built in | |
|---|---|---|
| Setup time | 0 extra hours | 1-3 extra hours per workflow |
| Time to detect a failure | Days, sometimes weeks | Minutes |
| Cost of a missed lead batch (est.) | $500-$3,000 in lost pipeline, depending on deal size | Near zero — caught same day |
| Client-facing failures (missed invoice, no confirmation email) | Happens eventually, damages trust | Rare, caught before client notices |
| Ongoing maintenance load | Higher — fixes happen reactively, often urgently | Lower — fixes happen on your schedule |
The upfront cost is real but small: on a typical client project, adding proper error-handling to five or six core workflows adds maybe half a day of work. Compare that to the cost of a single missed lead batch during a busy sales month, or an automated invoicing workflow that quietly stops sending invoices for two weeks. The math isn't close.
The questions worth asking about your current setup
If you already have automations running, it's worth checking a few things today:
- If this workflow failed right now, would anyone know within the hour?
- Does it alert on failure, or just stop?
- What happens if the data coming in is missing a field or formatted wrong?
- Is there a human-readable log, or just a black box?
If the honest answer to any of these is "I'm not sure," that's the gap to close first — before adding more automations on top of a foundation that doesn't tell you when it's cracking.
If you want a second set of eyes on what's already running in your business — or want error-handling built in from day one on something new — book a free call and we'll walk through it together.