Use four layers. 1) Turn on Retry On Fail for nodes that call APIs. 2) Choose the right On Error behavior per node. 3) Build one error workflow with the Error Trigger node and assign it to every production workflow. 4) Catch silent failures, meaning runs that "succeed" but produce nothing, with item-count checks and a heartbeat.
The most expensive n8n problem I fix for clients isn't a workflow that crashes. It's a workflow that broke quietly three weeks ago while everyone assumed it was working. Leads stopped reaching the CRM. Orders stopped syncing. Nobody got an alert, because nobody set one up.
Here's the error-handling setup I add to every production n8n system, step by step.
Layer 1: Retry On Fail (for Temporary Errors)
Most API errors are temporary: a timeout, a rate limit, a brief outage. Retrying fixes them automatically.
- Open the node that calls an external service (HTTP Request, CRM, Google Sheets and so on).
- Click the Settings tab.
- Turn on Retry On Fail.
- Set Max Tries (up to 5) and Wait Between Tries (ms) (up to 5000).
Three tries with a 2,000 to 3,000 ms wait is a good default. Don't retry nodes that create records or send messages unless the action is safe to repeat, or you may get duplicates.
Layer 2: Choose the Right "On Error" Behavior
In the same Settings tab, On Error decides what happens when a node still fails:
| Option | What happens | Use it when |
|---|---|---|
| Stop Workflow (default) | Execution stops and is marked as failed | The step is essential, like creating the CRM contact |
| Continue | The error is ignored and the workflow moves on | The step is optional, like an enrichment lookup |
| Continue (using error output) | The node gets a second, red error branch | You want to handle failures yourself (log, alert, fallback) |
Be careful with plain Continue: the execution counts as successful, so your error workflow will never fire. Use the error output instead when you still need to know.
Layer 3: One Error Workflow for Everything
This is the most important layer. One workflow catches failures from all your other workflows and alerts you.
Step 1: Create the error workflow
- Create a new workflow called Error Alerts.
- Add the Error Trigger node as the first node.
- Add a Slack, Gmail, Telegram or WhatsApp node to send the alert.
- Save it. Error workflows don't need to be activated.
Step 2: Write a useful alert message
The Error Trigger gives you details about the failure. A good alert tells you what broke, where and why, and links straight to the failed run:
🚨 n8n workflow failed
Workflow: {{ $json.workflow.name }}
Failed at: {{ $json.execution.lastNodeExecuted }}
Error: {{ $json.execution.error.message }}
Open run: {{ $json.execution.url }}
Step 3: Connect every production workflow
- Open a production workflow.
- Open the workflow menu (three dots) and click Settings.
- In Error workflow, select Error Alerts.
- Save, and repeat for every live workflow.
Layer 4: Catch Silent Failures
A silent failure is an execution that finishes "successfully" but does nothing useful. Your order sync returns zero orders because a token expired, or a filter is wrong and every lead is skipped. No error, so no alert.
Fail on empty results
- After the key step (for example "Get new orders"), add an IF node that checks whether any items came back.
- On the "no items" branch, add a Stop and Error node with a clear message like "Order sync returned 0 orders".
- Because the execution now fails, your error workflow sends the alert.
Only do this where zero results is truly abnormal. A lead form at 3 a.m. with no new leads is normal.
Add a heartbeat for workflows that stop running
If a workflow is deactivated, or its trigger stops firing, nothing runs and nothing fails. Add a small scheduled workflow that uses the n8n API to check when each important workflow last ran successfully, and alerts you if it's been too long.
Bonus: Log Every Important Run
For business-critical workflows, write one row per run to Google Sheets or a database: time, records processed and status. When a client asks "did my lead get through on Tuesday?", you can answer in seconds. It also makes debugging much faster. See how to fix a broken webhook integration.
Want This Without Building It?
I built ProofMyAI for exactly this. It connects to n8n or Make with an API key and alerts you on failed runs, error spikes, slow runs, workflows that stop running and silent zero-output runs. Alerts go to Slack, Discord, Teams or email. More background in my guide to AI chatbot and workflow monitoring.
Or grab my ready-made n8n Error Alerts & Silent-Failure Monitor template.
Production Checklist
- ☐ Retry On Fail on every node that calls an external API
- ☐ On Error set deliberately per node (no accidental Continue)
- ☐ One error workflow assigned to every production workflow
- ☐ Alert includes workflow name, failed node, error and a link
- ☐ Zero-result checks on key steps
- ☐ Heartbeat for workflows that must run on a schedule
- ☐ Run log for business-critical workflows
Frequently Asked Questions
Create a separate workflow that starts with the Error Trigger node and sends a message to Slack, email or WhatsApp. Then open each production workflow, go to Settings and choose that workflow in the Error workflow field. Every failed automatic execution will trigger an alert.
The most common reasons are that the Error workflow field is empty in the failing workflow's settings, the failure happened during a manual test run (error workflows only fire for automatic executions), or the node was set to Continue on error so the execution never actually failed.
Continue ignores the error and passes the input data on as if nothing happened. Continue (using error output) adds a second, red error branch to the node so you can handle failures explicitly, for example logging them or sending an alert, while the workflow keeps running.
With Retry On Fail enabled in a node's settings, Max Tries can be set up to 5 and Wait Between Tries up to 5,000 milliseconds. Three tries with a 2 to 3 second wait is a sensible default for most APIs.
Silent failures are runs that succeed but produce nothing. Add an IF node after the key step that checks the item count, and use the Stop and Error node to fail the execution when it's zero, which then triggers your error workflow. For workflows that stop running entirely, use a scheduled heartbeat check or a monitoring tool.
Need Your n8n Workflows Made Reliable?
I audit existing n8n setups, add error handling and alerts, and fix what's broken. Send me a WhatsApp message with what you're running.
Get a Free Automation Audit → 💬 WhatsApp me