Summary
A customer's Monday Board displayed twice as many employees per customer. Investigation revealed that an automation had duplicated client records in the Accounts board approximately 5 days prior to the report. When the N8N workflow ran to sync data, it executed twice due to the duplicated records, causing the inflated employee counts.
No n8n Code node ran for about 63 days, from 27 Jul 2026 13:01 UT C to 28 Sep 2026 13:19
UT C. The JavaScript runner on the worker lost its connection to the worker and gave up
without retrying. The container still showed as "running"
, so nothing restarted it and
nothing alerted anyone.
Symptom: every workflow step using a Code node f ailed after 60 seconds with "T ask
request timed out after 60 seconds – not matched to a runner"
. This hit scheduled runs
and manual "Execute step" runs alike.
Impact: every workflow containing a Code node f ailed on every run in that window ,
including the Monday CRM syncs (Org Sync New Orgs, Sync Deals to Accounts). The
full list of affected workflows is still to be pulled from the Executions view .
Detection: found by hand on 28 Sep, while investigating a single f ailing Monday sync.
No alert fired.
Status: resolved on 28 Sep by restarting the n8n-worker-runner container , and Code
nodes have been running since. A canary and runner-f ailure alerts to #service-alerts
have been built. Automatic restart is covered in the playbook below .Environment
N8N runners with Monday Board integration; workflow URL: https://webhooks.prpl.tools/workflow/6G6lXWb3sQ2R1swn
n8n 2.13.2 runs self-hosted in queue mode. The editor and triggers run on the main
container , and all executions run on a separate worker . Because OFFLOAD MANUAL EXECUTIONS TO
WORKERS=true is set, even a manual "Execute step" in the editor runs on the worker . So when the worker's runner stops, every Code node on the instance stops.
Item |
Value |
Compose project |
int-purplecomputing-n8n-qdggo0 (Dokploy , /etc/dokploy/compose/int-purplecomputing- n8n-qdggo0/code/docker-compose.yml) |
Host |
Public IP 82.215.64.58. Believed to be VM01 at Binary Racks, going by Michael's note in #accounts |
Management UI |
Portainer , environment local (ID 3), user team@purplecomputing.com |
Runner container |
int-purplecomputing-n8n-qdggo0-n8n-worker- runner-1, image n8nio/runners:2.13.2, health port 5680 |
Runner mode |
N8N_RUNNERS_MODE=external, N8N_RUNNERS_ENABLED=true |
Second stack |
int-purplecomputing-n8nrunnerpostgresollama-b3ly6t (n8n:latest, runners:nightly) is still running; its purpose is unknown |
Timeline
The outage began at 13:01 UT C on 27 Jul and ended at 13:19 UT C on 28 Sep. Container times
are UT C (BST is UT C+1). The worker log we pulled only covers 08:04–13:05 on 28 Sep.
| When (UTC) | Event |
| 28 Sep, afternoon |
Canary workflow and runner-f ailure error workflow built and tested; alerts go to #service-alerts |
| 28 Sep 14:01 |
Hourly scheduled workflows complete their Code nodes for the first time since July |
| 28 Sep 13:28 |
First Code tasks picked up by the runner ("Runner process exited on idle timeout" lines resume) |
| 28 Sep 13:19:34 |
n8n-worker-runner-1 restarted in Portainer; JS and Python launchers reconnect cleanly . Outage ends |
| 28 Sep 11:57 |
W orker container restarted; timeouts continue, because the runner is untouched |
| 28 Sep 08:31–13:01 |
21 "T ask request timed out after 60 seconds" errors in the worker log (earliest in the log window) |
| 28 Sep, morning |
Pieter investigates a f ailing Monday Org Sync. The creator has left and there is no documentation |
| 27 Jul 13:01:28 |
Runner log: ERROR [launcher:js] Failed to execute launch command: handshake failed: websocket: bad handshake. The JS launcher stops and logs nothing further . Outage begins |
| 27 Jul 12:38 BST |
CSA T received on ticket #85050 "Issue with N8N runners – F AO Denis" (resolved) |
| 21 Jul 15:56 | Runner container last restarted before the incident |
| 26 Jun | Ticket #85050 raised: an earlier runner issue |
| 23 Apr |
Current stack ( qdggo0) deployed with n8n 2.13.2 and a runner sidecar |
Symptoms
Monday Board displayed twice as many employees per customer. The Accounts board contained duplicate client records (approximately 5 days old at time of investigation).
Root cause
An unidentified automation had duplicated every Monday company record in the Accounts board. The workflow logic itself was sound, but when the sync workflow ran, it executed against both the original and duplicate records, resulting in doubled employee counts. Execution history was insufficient to determine the initial cause of the duplication.
Root cause: at 13:01 UT C on 27 Jul, the JavaScript launcher inside n8n-worker-runner
f ailed one websocket handshake with the worker's task broker . It then stopped for good,
with no retry . The container process stayed up, so Docker's restart: always never
kicked in. Every Code task sent to the worker afterwards waited 60 seconds for a runner
and timed out.
Page 3 of 6n8n T ask Runner Outage – RCA & Auto-Restart Playbook
Likely trigger (not confirmed): the worker or its broker was briefly unavailable, probably a
restart or redeploy . That fits the timing: ticket #85050 was closed 23 minutes earlier .
W orker logs from 27 Jul are no longer available.
Contributing factors
No health check or monitoring on the runner. Portainer showed it as "running" the
whole time, and nothing tested whether Code nodes actually ran.
Failures went unseen for 63 days. The error workflow ( QEL3x5SXwqcK1uFa) most likely
notifies someone who has left, and nobody else was watching execution errors.
One runner for all executions. With manual runs also offloaded to the worker , a single
stuck sidecar stopped every Code node on the instance.
Knowledge loss. The builder left without handing over or documenting where the
stack runs, how it is deployed (Dokploy), or how to reach it (Portainer , Cloudflare
Access).
Early misdirection. The error text points at timeouts and capacity , and the first guess
was a data change in monday .com. Neither was the cause.
Also found (not causal) Secrets in plain text. An OpenAI API key , a Cloudflare tunnel token, and one password reused for Postgres, the runner auth token and the runner secret all sit in container
environment variables. An IT Glue API key and a monday .com token are hard-coded in workflow nodes.
A second n8n stack ( b3ly6t ) is still running. If it runs the same Monday workflows, it could explain the duplicate records noted in an earlier KB draft. Zendesk OAuth credential fails to refresh. Expir
Resolution
Service was restored by restarting one container . Everything else below either made it
possible to find that container or reduces the risk of a repeat.
1. Worked around the error in two workflows. Rebuilt Monday CRM Org Sync New Orgs
(v3) and Monday CRM – Sync Deals to Accounts (v2) with no Code nodes (Set, Filter , IF ,
Stop and Error and Summarize nodes). Both were tested on a local n8n 2.13 against the
originals, and outputs matched exactly . That testing caught and fixed a grouping bug
before release.
Page 4 of 6n8n T ask Runner Outage – RCA & Auto-Restart Playbook
2. Located the stack. Used 1P assword ("MT Synology Authenticate Behind Cloudflare
Access"), Cloudflare Zero T rust tunnels ( purple
vm01
n8n) and Portainer , then
_
_
identified the Dokploy compose project qdggo0.
3. Mapped the setup. Inspected the main, worker and runner containers. This confirmed
queue mode, the external runner , and that manual executions are offloaded to the
worker .
4. Found the cause. The worker log showed every Code task timing out, and the runner
log showed the 27 Jul handshake error with nothing logged after it.
5. Fixed it. Restarted n8n-worker-runner-1 in Portainer at 13:19 UT C on 28 Sep, then
checked that tasks were being picked up and that the 14:01 scheduled runs succeeded.
6. Added monitoring. T wo workflows alert #service-alerts: Monitor – Code Runner
Canary (every 15 min) and Monitor – Runner Failure Alerts (an error workflow that
filters for runner errors). Both were tested with the runner healthy and deliberately
broken.
7. Credentials. Documented moving monday .com to a service-account token. Moved the
Zendesk nodes from OAuth to an API-token credential, which now authenticates
successfully .
Verification
The customer was notified that the issue had been resolved and counts should now be accurate. They were advised to monitor for recurrence and report if the issue happens again, as the root cause of the initial duplication remains unclear.
Comments
0 comments
Please sign in to leave a comment.