Skip to content

fix(health): bound Redis and DB checks with a timeout - #1599

Open
Nife-tanny wants to merge 6 commits into
LabsCrypt:mainfrom
Nife-tanny:fix/health-redis-timeout
Open

Nife-tanny wants to merge 6 commits into
LabsCrypt:mainfrom
Nife-tanny:fix/health-redis-timeout

Conversation

@Nife-tanny

@Nife-tanny Nife-tanny commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Closes #1492

Summary

GET /health now runs all of its dependency probes in parallel, each bounded by HEALTHCHECK_TIMEOUT_MS (default 800 ms), and actually pings Redis. A hung Redis now answers in ~0.8 s with redis: "timeout" instead of stalling.

⚠️ /health is currently broken on main (fixed here in a separate commit)

Merge 38a3e2d (Sept 26) kept the checks block of the route but dropped the code that defines indexerLagDegraded, indexerFailureDegraded, redisStatus and sorobanRpcOk. On main, every request throws ReferenceError: redisStatus is not defined and returns 500. Render's healthCheckPath, the docker-compose healthcheck and the CI boot check would all treat the service as down. This is also why 8 tests in tests/health.test.ts fail on main.

Commit 5954b30 is a pure merge repair. It restores those definitions and the event-counter body fields from main's side of the merge (45e956d) and keeps the ledger-lag additions from 87b656a, with no behaviour changes. It is separate so it can be reviewed or cherry-picked on its own. Note that #1579, #1530, #1570 and #1590 also edit this route and partly repair the same lines.

Also note: on main the route never pinged Redis. Its status came from isRedisAvailable(), a flag that is set once at connect and never cleared on disconnect, so it stayed ok after Redis died.

Changes

  • backend/src/lib/with-timeout.ts (new): withTimeout(promise, ms, label) and TimeoutError. The timer is always cleared in finally, and the abandoned promise gets a no-op catch so a late rejection is never unhandled.
  • backend/src/routes/health.routes.ts:
    • The DB ping, indexer-state lookup, Redis ping, Soroban RPC check and (when the indexer is enabled) the network-tip lookup run in parallel with Promise.all. Each probe catches its own errors, so none reject.
    • Redis fast path: when REDIS_URL is set but the ioredis client isn't ready (or never connected), the check returns unavailable immediately, without pinging.
    • checkRpcHealth() defaults to a 3 s timeout, so it now gets the same bound.
    • Failures are logged at warn once when a component starts failing and at info on recovery; per-probe errors go to debug. No error text, hosts or connection strings appear in the response.
    • Responses carry Cache-Control: no-store.
  • backend/src/config/swagger.ts: HealthResponse gains the redis field, adds timeout to the database/redis enums, and documents the Redis-only 200 case.
  • backend/.env.example: documents HEALTHCHECK_TIMEOUT_MS.

Response per scenario

Only the relevant fields are shown. Existing fields are unchanged and only redis is added.

Scenario HTTP status db redis (= checks.redis.status) checks.database.status
DB and Redis OK 200 ok connected ok ok
Redis not configured (REDIS_URL unset) 200 ok connected not_configured ok
Redis doesn't answer within the timeout 200 degraded connected timeout ok
Redis ping fails / client not ready 200 degraded connected unavailable ok
DB query fails 503 degraded disconnected as measured down
DB query times out 503 degraded disconnected as measured timeout
DB and Redis both failing 503 degraded disconnected unavailable / timeout down / timeout
Indexer lag > 60 s or failure spike (unchanged) 503 degraded connected as measured ok

Example (Redis hung, DB OK):

{"status":"degraded","db":"connected","redis":"timeout","indexerEnabled":false,"indexerLag":null,"indexerLedgerLag":null,"eventsProcessed":0,"eventsFailed":0,"lastErrorAt":null,"indexerDegraded":false,"uptime":12.3,
 "checks":{"database":{"status":"ok"},"indexer":{"status":"disabled","enabled":false,"lagSeconds":null},"redis":{"status":"timeout"},"sorobanRpc":{"status":"ok"}}}

Redis keeps the pre-existing unavailable value rather than down for backward compatibility with the documented enum. timeout is new.

Timeout budget

  • The per-probe timeout is 800 ms (HEALTHCHECK_TIMEOUT_MS), not 1,000 ms. Probes run in parallel, so the worst case is about 800 ms plus request handling. That leaves ~200 ms of headroom for event-loop lag, JSON serialisation and the network under the issue's 1,000 ms target.
  • Measured worst case with a hung Redis: 0.81–0.83 s (see below).
  • The timeout only stops waiting. The underlying query or ping keeps running until its own client gives up (see the suggestions below).

❓ Question for the maintainer: HTTP status when only Redis is degraded

The issue asks for both a "fast 503" and { status: "degraded", redis: "timeout" }. This PR returns 200 with status: "degraded" when only Redis is down or timing out, for these reasons:

  • The route's documented contract is that "Liveness (200 vs 503) is determined by DB reachability alone", and the pre-merge code said "Redis is optional … never affects the top-level isHealthy verdict". Redis only powers cross-instance SSE fan-out; single-instance mode works without it.
  • /health is the only health route and is used as a liveness signal: render.yaml sets healthCheckPath: /health, docker-compose uses wget --spider (fails on non-2xx), and the CI boot check uses curl --fail. Returning 503 on a Redis outage would make the platform mark every API instance unhealthy (and potentially restart them) for a Redis-only problem, turning a partial degradation into a full outage.

Side effect: status: "degraded" used to always mean 503. Now it can also appear with 200, and the HTTP code plus checks tell the two cases apart. If you'd prefer Redis failures to return 503, it's a one-line change (isHealthy would include !redisDegraded). A cleaner long-term split would be separate liveness and readiness routes (see suggestions).

Acceptance criteria

  • Healthcheck responds within 1,000 ms even if Redis is completely unresponsive.
    • Tests: answers at the timeout with redis "timeout" when the ping never resolves (fake clock: responds at 800–900 ms), responds in under 1,000 ms of real time when Redis never answers (real timers: ~820 ms), runs the checks in parallel: two 700 ms probes take ~700 ms, not 1400 ms, does not wait on a hanging Soroban RPC health check.
    • Manual: 0.829 s and 0.813 s against a hung Redis, vs a 10 s curl timeout without the bound.
  • Correctly distinguishes between database and Redis degradation.
    • Tests: returns 503 with database "down" when the DB query fails and Redis is fine, returns 503 with database "timeout" when the DB query hangs and Redis is fine, reports both failures when the DB and Redis are down, and the Redis-only cases (200 + degraded).
  • Unit test added for Redis timeout simulation.
    • Tests: the Redis timeout tests above, plus tests/with-timeout.test.ts (6 tests).

Tests

tests/with-timeout.test.ts (new, 6 tests) uses fake timers and vi.getTimerCount() to prove no timer is left behind. It covers:

  • resolving before the timeout
  • propagating the original rejection
  • TimeoutError exactly at the deadline (799 ms: pending, 800 ms: rejected)
  • no unhandledRejection when the abandoned promise rejects later
  • ignoring a late resolution
  • thenables (Prisma queries)

tests/health.test.ts keeps all 8 existing tests, which failed on main and now pass, and adds 18. The Redis client (getPublisher) and Soroban RPC are mocked, so the tests are offline. The fake-clock tests advance only setTimeout in 10 ms steps while yielding real event-loop turns, which measures how much simulated time the endpoint needed. The new tests cover:

  • the Redis timeout at the default and at HEALTHCHECK_TIMEOUT_MS=300
  • Redis ping rejected → unavailable
  • client reconnecting/connecting/end/close → unavailable without calling ping
  • REDIS_URL set but not connected → unavailable
  • REDIS_URL unset → not_configured, 200 ok
  • DB down and timeout → 503
  • both failing
  • both OK → 200 ok with Cache-Control: no-store
  • parallel execution (~700 ms, not 1,400 ms)
  • a hanging RPC doesn't block the response
  • the body never contains error text, hosts or credentials
  • a failing component is logged once, not on every probe

Verification

Manual timing: the real app and the real ioredis client ran against a small fake Redis TCP server. The fake completes ioredis's ready check (INFO) and then either answers PING or never does, which is the "connected but unresponsive" hang. Docker wasn't available locally, so Postgres was unreachable, and every manual response below is 503 because of the DB. The Redis field and the timing are what these runs show; DB-up combinations are covered by the tests above.

### Redis responsive
"redis":"ok"            HTTP 503  time_total=0.349355s
### Redis hung (connected, never answers PING)
"redis":"timeout"       HTTP 503  time_total=0.828881s
### Redis hung, second request
"redis":"timeout"       HTTP 503  time_total=0.812528s
### Redis process killed after connect (client no longer ready -> fast path, no ping)
"redis":"unavailable"   HTTP 503  time_total=0.020048s
"redis":"unavailable"   HTTP 503  time_total=0.014376s
### Redis not configured
"redis":"not_configured" HTTP 503 time_total=0.111736s
### Control: hung Redis with HEALTHCHECK_TIMEOUT_MS=60000 (no effective bound)
HTTP 000  time_total=10.010948s   (curl --max-time 10 gave up, exit 28)

Checks: the backend has no lint script, and CI does not lint the backend.

Check Result
tsc -p tsconfig.json 76 errors on main → 72, 0 new. The 4 removed are the undefined names in health.routes.ts; the rest are pre-existing (below)
vitest run (full backend suite) main: 71 failed / 485 passed. This branch: 62 failed / 518 passed. No new failures (diffed by name); 8 health tests fixed; 1 flaky test (below) passed
npm run codegen:openapi runs and picks up the HealthResponse change. The committed swagger/flowfi.openapi.json was already out of date on main (missing /metrics, /v1/streams/simulate and dead-letter routes), so I didn't regenerate it here to avoid mixing in unrelated drift
Forbidden patterns (any, ts-ignore, eslint-disable, .skip/.only) in changed files none

⚠️ Pre-existing failures on main (unrelated to this PR)

All reproducible on upstream/main at 72ec8ef:

  • npm ci fails: package-lock.json is out of sync with package.json. This is why Frontend CI, Backend CI and the preview workflow fail at Install dependencies.
  • prisma generate fails (P1012): schema.prisma defines model IndexerDeadLetterEvent twice (lines 71 and 164), so npm run build/npm test and the Docker image build fail. I verified locally with a client generated from a temporary, untracked de-duplicated copy of the schema; no schema change is included here.
  • The backend can't start under Node: pg-pool.js has no export named getPoolMetrics. The manual check ran through vite-node.
  • 72 remaining tsc errors, in sorobanService, soroban-event-worker, admin.routes, stream.controller, prisma.ts and tests.
  • 62 remaining failing tests in indexer-service, soroban.service, soroban-event-worker, stream-simulation, pg-pool, worker-correlation-id and integration/admin-* / reset-replay-race.
  • Flaky: stream.test.ts › withdraws the claimable amount for the recipient failed once under full-suite load on main and passes in isolation (3/3).
  • Clippy also fails on main.

Suggestions (not changed here)

  • ioredis commandTimeout: lib/redis.ts sets enableOfflineQueue: false, so commands fail fast while disconnected, but has no commandTimeout. Any command to a connected-but-hung Redis (e.g. SSE publishes) waits indefinitely. Setting commandTimeout (and connectTimeout) on the clients would bound every command, not just this probe.
  • isRedisAvailable() goes stale: _available is set to true on connect and never reset when the connection drops. sse.service.ts (lines 212 and 220) uses it to decide whether to publish through Redis, so during an outage it keeps trying Redis instead of falling back to local delivery.
  • The Redis client stops reconnecting: retryStrategy returns null after 3 attempts, so after a short Redis outage the clients end permanently and don't recover until the process restarts. /health will then report unavailable until a restart.
  • Separate liveness and readiness: e.g. /health/live (process only, for restarts) and /health/ready (DB/Redis/indexer, for traffic routing). That would resolve the 200-vs-503 question above cleanly.
  • Abandoned probes keep running: a timed-out DB probe keeps its query running until Postgres's statement_timeout (PG_STATEMENT_TIMEOUT_MS, default 30 s). Consider a shorter dedicated timeout for health queries if probes are frequent.

Nife-tanny and others added 6 commits September 30, 2026 08:05
Merge 38a3e2d kept the `checks` block of GET /health from main but dropped
the code that defines indexerLagDegraded, indexerFailureDegraded,
redisStatus and sorobanRpcOk, so every request threw a ReferenceError and
/health answered 500 (Render and Docker health checks see it as down).

Restore those definitions and the event-counter body fields from main's
side of the merge (45e956d), keeping the ledger-lag additions from
87b656a. This is a pure merge repair with no behaviour changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Races a promise against a timer and rejects with a TimeoutError. The
timer is always cleared, and the abandoned promise gets a no-op catch so
a late rejection never becomes an unhandled rejection.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GET /health could stall indefinitely on a connected but unresponsive
Redis (ioredis has no command timeout configured), and the restored
Soroban RPC check waits up to 3 s. Run every probe in parallel, each
bounded by HEALTHCHECK_TIMEOUT_MS (default 800 ms), so the endpoint
answers in under 1 s even when a dependency hangs.

- Redis is actually pinged; a client that is not `ready` is reported as
  unavailable without pinging. The stale isRedisAvailable() flag is no
  longer used here.
- Components report ok / down|unavailable / timeout, with a new top-level
  `redis` field. Redis-only degradation sets status "degraded" but keeps
  HTTP 200, so a Redis outage does not fail liveness probes.
- Failures are logged once on state change; no error details are
  returned. Responses are Cache-Control: no-store.

Closes LabsCrypt#1492

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Cover a hanging, failing and not-ready Redis, DB failure and timeout,
both failing, parallel execution, HEALTHCHECK_TIMEOUT_MS, no error
leakage and log de-duplication, using a fake clock plus one real-time
check. Mock the Redis client and Soroban RPC so tests stay offline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sync the generated spec and frontend API types with the /health schema
changes in swagger.ts so the OpenAPI drift check passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Backend] Add explicit timeout and circuit check to Redis ping in /api/v1/health

1 participant