Skip to main content
Kevin Kirui
← Back to Work

Healthcare AI

Biomedical Maintenance Triage

A bounded LLM workflow that classifies biomedical equipment maintenance reports into controlled routing categories for hospital operations.

Biomedical Maintenance Triage

Project Overview

2026
Year
Completed
Status
7 min
Reading Time

Technologies & Topics

Healthcare AIClinical WorkflowsBiomedical EngineeringWorkflow AutomationLLM SystemsReliability Engineering

Context

Biomedical equipment maintenance in a hospital setting depends on maintenance reports submitted by clinical staff, most of whom are not biomedical engineers and are not expected to describe equipment failures in technical language.

These reports arrive through whatever channel is fastest in the moment, often a short written note or a verbal handoff transcribed after the fact. They vary widely in specificity, terminology, and urgency markers.

Hospitals typically route these reports to operational teams, biomedical engineering, ICT support, facilities, or vendor escalation, based on manual reading and judgment. This works, but it does not scale cleanly or leave an auditable trace of routing decisions.

This project sits at that translation point: between an unstructured report written under time pressure and the structured operational routing decision a hospital's maintenance systems actually need.

Objectives

The project was developed around four objectives:

  1. Define a closed output contract for equipment type, issue type, urgency, assigned team, confidence, and reason, so every response is structurally predictable regardless of how the source report is written.

  2. Separate probabilistic model output from deterministic application policy, so no policy decision, what happens at low confidence, what counts as a retryable failure, is ever made by the model; these decisions belong to the application.

  3. Build explicit failure handling for invalid model output, transport failures, timeouts, and caller cancellation, rather than treating these as edge cases to patch in later.

  4. Create reproducible evaluation and usage logging so system behaviour could be measured against real runs and real dependency behaviour, not just described from the implementation.

My Approach

I approached the project as a bounded healthcare workflow rather than a general-purpose conversational AI application.

The model proposes a classification. The application validates it. The application then applies the routing policy. That separation creates a clear boundary between probabilistic model behaviour and deterministic application decisions.

I also treated reliability as part of the product design rather than an afterthought. Model calls sit behind bounded retries, explicit timeout handling, cancellation, schema validation, and a single route-level deadline.

POST /maintenance-triage

Request body:
{
  "text": "The ICU ventilator in bed 3 keeps alarming."
}

Success response (200):
{
  "equipment_type": "ventilator",
  "issue_type": "alarm_fault",
  "urgency": "high",
  "assigned_team": "biomedical_engineering",
  "confidence": 0.91,
  "reason": "..."
}

Failure responses map to specific, distinguishable causes rather than a single generic error:

400  invalid_input           malformed request body
422  invalid_model_output    model output failed validation after one repair attempt
429  upstream_rate_limit     provider rate limit exhausted after retries
500  model_configuration_error / model_provider_error   provider rejected credentials or request
502  model_transport_error / upstream_failure            provider unreachable or failed after retries
503  llm_disabled            kill switch active, system refuses to guess
504  upstream_timeout        model request timed out
504  workflow_timeout        60-second workflow deadline exceeded

Distinguishing these matters operationally: a 503 tells an operator the system is deliberately refusing to guess, while a 504 tells them the request genuinely exceeded its time budget. Collapsing them into a single "error" would obscure what happened and why.

System Design

The request lifecycle moves through a fixed sequence of stages, each with a single responsibility:

Input validation
       ↓
Kill switch check
       ↓
Initial LLM call
       ↓
JSON parsing + schema validation
       ↓
One repair attempt if validation fails
       ↓
Quarantine if repair also fails
       ↓
Deterministic confidence policy
       ↓
Final routing classification

Two concerns run orthogonally to this pipeline rather than as steps within it: transport reliability and cancellation. Every model call, the initial classification and the single repair attempt, runs inside bounded retries and a cancellation boundary.

That separation is what let the two reliability bugs described below be fixed and verified independently, without touching the validation logic, the confidence policy, or the route's public response contract.

Technical Implementation

Reliability boundaries around the model call

Every model request is wrapped in a bounded-retry layer with three explicit boundaries:

  • A 15-second provider timeout on each individual request attempt.
  • Bounded transport retries (2 attempts) with backoff, applied only to genuinely retryable failure classes (timeouts, rate limits, 5xx responses), never to validation failures or authentication errors.
  • A 60-second workflow deadline owned by the route, independent of any single request's timeout.

Two verified bugs in that boundary layer

Two defects surfaced during this work, both invisible without testing against the real error shapes the pinned SDK produces rather than the shapes a reasonable implementation would assume.

Timeout misclassification. The original isTimeoutError() check matched on error.name === "AbortError" and similar string comparisons. Installing the pinned openai SDK and inspecting its actual error types revealed that timeouts do not always produce an AbortError; some arrive as network exceptions with different error signatures. The fix: test against the real SDK's error classes, not assumed shapes.

Dangling in-flight work after the workflow deadline. The 60-second deadline was implemented as Promise.race([classification, timeout]). When the timeout won the race, the route returned promptly, but the in-flight model request continued running in the background. If an abort landed during the retry backoff sleep (a gap in the original wiring), it would be ignored. The fix: wire the AbortController through the entire retry loop and the backoff sleep, so an abort takes effect immediately.

Both fixes are on main at commits 9fc3e01 (timeout classification) and 53074e1 (cancellation wiring), with regression tests at e28c474 and alongside 53074e1.

Output validation and safety

The model's output is constrained to a closed schema; the application, not the model, controls what happens next:

confidence >= 0.50 (application routing threshold)
    → accept model classification

confidence < 0.50 (application routing threshold)
    → equipment_type = "other"
    → assigned_team = "biomedical_engineering"
    → urgency preserved
    → issue_type preserved
    → confidence preserved
    → reason preserved

Important: The 0.50 threshold is an application routing boundary, not a statistically calibrated probability of correctness. It is a policy decision: below this reported confidence, the application defaults to the safest routing (biomedical engineering) and escalates urgency decisions to human review. This boundary may shift based on production experience; it is not derived from statistical analysis of the evaluation suite.

If the model's output fails schema validation, the system makes exactly one repair attempt, providing the validation error back to the model. If that repair also fails, the failure is quarantined—routed to a human queue rather than guessed-at with a fallback heuristic. For a routing system where an incorrect classification cascades into operational decisions, deliberate abstention is safer than introducing an unevaluated secondary classification mechanism.

Cost and usage logging

Every successful model completion records structured usage metadata, prompt version, call type, model, token counts, and duration, without logging the maintenance report text, the prompt, or model output. This preserves operational tracing while protecting report content.

Results

The project has two distinct kinds of evidence, and they answer two different questions.

Does the workflow produce the intended classification?

An eight-case evaluation suite covers clear malfunction, calibration, connectivity, facilities/power failure, consumable supply, ambiguous equipment, low-confidence urgency preservation, and prompt repair scenarios.

8/8 assigned_team correct   (100%)
8/8 fully matched

Every case matched on every checked field, including Case 7's low-confidence override (a 0.4-confidence classification correctly triggered the fallback policy while preserving urgency rather than guessing).

What this result demonstrates: The workflow produces the intended classification across these eight selected scenarios.

What this result does not demonstrate: Statistical accuracy or precision across the broader distribution of hospital maintenance reports, or guaranteed consistency across re-runs. This project uses openrouter/free, which routes each request to a different underlying free model. That means this score reflects behaviour demonstrated in one documented run under one provider routing strategy, not a guaranteed invariant. Re-running the same suite against a different provider or a different pinned model would produce different results.

For productionization, this evaluation would need to expand significantly: run the suite against a pinned model multiple times to establish variance, extend the suite to cover edge cases and adversarial inputs, measure real-world accuracy against actual hospital maintenance reports, and track routing drift over time.

Does the reliability engineering around the model behave correctly?

A regression suite of 15 tests, built against the real SDK's error classes rather than mocked shapes, passes in full:

15/15 tests passing

These tests cover: real SDK timeout and abort errors being correctly classified (and real rate-limit and generic errors correctly not misclassified as timeouts), caller cancellation taking precedence over retries, backoff timing, and deadline enforcement.

Separately, the cancellation fix was verified once, manually, against a real hung network connection: a dummy server that accepted a connection and never responded, with the workflow deadline temporarily set to 5 seconds. The route returned a workflow_timeout error and the in-flight request was properly aborted.

Important: These two test suites answer different questions. The reliability tests verify transport, timeout, cancellation, and retry behaviour. The evaluation suite tests routing behaviour. A passing reliability suite does not guarantee routing accuracy; a strong routing evaluation does not guarantee that reliability defects won't appear under production load. Both kinds of evidence are necessary and insufficient on their own.

Lessons Learned

The most consequential lesson was that a plausible-looking timeout check can be silently wrong. isTimeoutError() looked correct by inspection; it checked the error properties a timeout error would be expected to have. Only testing against the real SDK's error types revealed it was incomplete. The takeaway: test against real dependencies, not assumed shapes.

The second lesson came from fixing my own fix. Wiring the AbortController correctly for the in-flight request case initially left a gap: an abort landing during the retry backoff sleep would not take effect. The fix worked for some timing windows and silently failed in others. The takeaway: cancellation boundaries need to run orthogonally through the entire pipeline, including sleeps and backoff logic, not just the main request path.

The third lesson was more about the product than the code: for a system whose output routes real operational decisions, the boundary between what the model decides and what the application decides is not just an architectural preference—it is a reliability requirement. A model that makes both the classification and the policy decision about what to do when confidence is low can fail silently. An application that clearly separates these two concerns can fail explicitly and be operated on.

What I'd Do Next

The first priority would be pinning to a specific model rather than continuing to rely on openrouter/free's auto-routing. That would remove the run-to-run variance currently documented as a limitation and establish a reproducible baseline for accuracy and latency.

The second priority would be extending the timeout-and-cancellation verification beyond the two SDK error classes and native AbortController shapes currently tested. The current fix is verified against specific conditions; a production system would need to verify across a wider range of failure modes and network conditions.

The third priority would be establishing observability and operational runbooks: quarantine rate monitoring, latency distribution tracking, routing drift detection, provider dependency controls, and on-call playbooks for common failure modes.

The fourth priority would be extending the evaluation suite to establish statistical confidence in accuracy across a representative sample of hospital maintenance reports, rather than a curated set of eight scenarios. This would require access to real or synthetic maintenance data at scale.

Finally, I'd want to generalize the retry-and-cancellation layer itself, the classification logic, the abort-aware backoff, and the deadline wiring, into a reusable module rather than code specific to this route. Other bounded LLM workflows would benefit from these same reliability patterns.

Summary

I designed a bounded LLM workflow for operational routing, separated model output from deterministic application policy, and found and fixed real reliability defects at the model-provider boundary. The project demonstrates how careful system design and testing against real dependencies can surface and prevent subtle failures that plausible-looking implementations would miss.

This is a prototype. It shows how a bounded healthcare workflow can be built, not how one would be deployed. Productionization requires provider strategy, observability, statistical validation at scale, and operational ownership. This work remains ahead.

Continue Exploring

More Case Studies

Explore additional healthcare AI and clinical workflow design projects.

Healthcare AI · 2026

Clinical Workflow Signal Audit

My role: Scoped, designed, and built the workflow audit and analysis.

Modeled signal-to-action latency across 500 synthetic ICU workflow events. Found a 36-minute median response time and 63.8% SLA compliance, with a full audit trail from signal to action.

Clinical Workflow Signal Audit

Public Health · 2026

Kenya Health Facilities Dashboard

My role: Designed and built the county-level planning and analytics tool.

Built a county-level planning tool on 10,483 health facility records from Kenya's public facility data, surfacing service and ownership gaps by region.

Kenya Health Facilities Dashboard