Skip to main content
Kevin Kirui
Backend Engineering · 8 min

When a Timeout Doesn't Cancel the Request: Fixing Cancellation Propagation in Node.js

How I traced a workflow deadline through an LLM request, retry layer, and backoff sleep, then used AbortController to propagate cancellation.

A timeout can stop a workflow from waiting without stopping the work underneath it.

While working on Biomedical Maintenance Triage, a backend service that turns unstructured hospital equipment maintenance reports into structured routing decisions, I traced a gap in the workflow's deadline handling.

Each workflow can involve a model request, transport retries, and a model-output repair step. The outer deadline could return a timeout response without communicating cancellation to those lower layers.

The problem was not simply the timeout. It was cancellation propagation.

Tracing the failure

The execution path crossed several layers:

Workflow deadline
        ↓
Model client
        ↓
Retry layer
        ↓
Backoff sleep

The old deadline raced an already-started operation against a timeout promise. If the timeout won, the workflow could return while the underlying operation remained active.

That created a distinction between:

The workflow has stopped waiting.

and:

The work belonging to the workflow has received a cancellation signal.

Without that signal, a request that later failed with a retryable error could enter another retry, even though the workflow had already returned a timeout.

This was an implementation-level failure mode, not a claim that I had observed a production incident or measured additional API charges.

Propagating the deadline

The deadline now creates an AbortController and passes its signal into the operation before starting it.

The important change was moving from an already-created promise to a promise factory:

const controller = new AbortController();

const promise = promiseFactory(controller.signal);

When the deadline fires, it aborts the controller and rejects the timeout promise with WorkflowDeadlineError.

The workflow still uses Promise.race(). The difference is that the timeout now also communicates cancellation to the underlying work.

The same signal reaches both the model request and its retry wrapper. The model request receives it through the SDK's request options:

await getClient().chat.completions.create(
  {
    model: process.env.LLM_MODEL,
    temperature: 0,
    messages,
  },
  { signal }
);

The repair path receives the signal as well. Cancellation therefore remains available if the workflow reaches the second model call used to repair invalid output.

Cancellation takes precedence

A transport timeout and a workflow cancellation are different decisions:

SDK timeout, without caller cancellation
        ↓
Apply the normal retry policy

Workflow cancellation
        ↓
Do not retry the failed attempt

In the retry layer, the cancellation check happens before normal error classification.

If the signal is aborted when an attempt rejects, the wrapper throws a TransportError with these properties:

{
  kind: "cancelled",
  retryable: false,
}

This prevents the resulting error from being treated as an ordinary retryable transport failure.

Backoff also needs cancellation

Passing a signal to the model request addresses only part of the execution path.

Between attempts, the retry layer waits:

Attempt fails
        ↓
Backoff sleep
        ↓
Next attempt

If cancellation arrives during backoff, the waiting period also needs to respond. Otherwise, cancellation handling may be delayed until the timer finishes.

The default sleep() helper now accepts the same signal. It rejects immediately if the signal is already aborted. If cancellation arrives while the timer is active, it clears the timer and rejects.

The retry wrapper catches that rejection and throws the same non-retryable cancellation error shape used for cancellation during an attempt.

The principle is straightforward:

Every asynchronous layer that can keep the operation alive needs a cancellation path.

What the tests cover

The commit adds three tests using Node's built-in test runner.

1. Ordinary timeout still retries

The first operation throws APIConnectionTimeoutError on its first attempt and succeeds on its second.

No cancellation signal is supplied. The assertions check that the result is "ok" and that two attempts occurred.

This checks that introducing cancellation did not break the existing timeout-retry behavior.

2. Caller cancellation prevents retries

The second test simulates an aborted SDK call: the operation aborts the controller and throws APIUserAbortError.

It checks that the wrapper rejects with a TransportError classified as "cancelled", marks it non-retryable, and makes only one attempt:

assert.equal(attempts, 1);

This tests the retry wrapper's response to simulated cancellation. It does not make a live model request.

3. Cancellation interrupts backoff

The third operation throws a retryable timeout and enters backoff. The controller is then scheduled to abort after approximately 20 milliseconds.

For this test, the retry layer receives an injected abortable sleep based on Node's node:timers/promises API:

const realSleepFn = (ms, signal) =>
  timersSetTimeout(ms, undefined, { signal });

The assertions check that cancellation is classified consistently, no second attempt starts, and execution finishes in less than 500 milliseconds rather than waiting out the expected 1,000-millisecond backoff.

This exercises the retry wrapper with an abortable timer. It does not directly test the separate custom sleep() implementation used by default.

Scope of the fix

The change gives the workflow deadline a cancellation path through the initial request, repair request, retry handling, and backoff.

It does not establish that an upstream provider immediately stops server-side processing or avoids billing once a request is aborted. Those are separate provider-side behaviors.

The scope is narrower: communicate cancellation to the client-side operation and prevent the retry layer from continuing through the tested cancellation paths.

The engineering lesson

A timeout is not necessarily cancellation.

When an operation spans network calls, retries, and timers, its deadline needs to reach the layers that can continue doing work.

For this workflow, the useful question was not just:

Does the timeout response arrive?

It was:

After that response arrives, what work could still be alive underneath?

Tracing that question through the implementation exposed the missing cancellation path.

The code and project

The fix and tests are available in the cancellation propagation commit on GitHub.

This work is part of Biomedical Maintenance Triage, an AI-assisted pipeline for turning equipment maintenance reports into structured routing decisions.

View the Maintenance Triage project.

Get in touch

I'm currently open to backend, AI engineering, and healthcare systems opportunities in Nairobi or remotely.

Get in touch.