Skip to content

DF-26076: fix hourly GSR WebSocket data stalls - #5381

Open
cl-efornaciari wants to merge 1 commit into
mainfrom
fix/DF-26076/gsr-token-refresh-v2
Open

cl-efornaciari wants to merge 1 commit into
mainfrom
fix/DF-26076/gsr-token-refresh-v2

Conversation

@cl-efornaciari

@cl-efornaciari cl-efornaciari commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Fixes DF-26076. Supersedes #5252 — see What changed from #5252 below, including a correction to its stated root cause.

Symptom

Every GSR EA instance — Chainlink-run and node-operator-run — stops serving data for ~2 minutes, once an hour, and returns a burst of 504s. Reported 2026-07-23, still occurring, and re-reported via NTI-251.

Root cause

A GSR WebSocket session delivers usable data for exactly 60 minutes after it opens, then silently stops. GSR does not close the socket and sends no notice.

The adapter had no way to detect this. It read the token's validUntil into a type and then discarded it, and had no refresh timer. So the only thing that noticed was the framework's generic staleness check, WS_SUBSCRIPTION_UNRESPONSIVE_TTL, at its 120s default. That fires 120s after the last usable message, closes with code 1000, and reconnects successfully on the first try.

The damage is the blind window, not the reconnect. CACHE_MAX_AGE is 90s, so cached prices expire ~30s before the framework even notices, and every caller gets a 504 until the new connection is serving.

Evidence from production

62.0-minute period, 12 consecutive cycles — {app="gsr", env="production"} |= "exceeding the threshold":

05:12:06  06:14:06 (+62.0)  07:16:08 (+62.0)  08:18:09 (+62.0)  09:20:09 (+62.0)
10:22:10 (+62.0)  11:24:11 (+62.0)  12:26:13 (+62.0)  ...  18:30:19 (+62.0)

62.0 min = 60 min of data + 120s to detect.

The adapter's own log, the 18:30 cycle:

18:30:19.508  The connection is unresponsive (last message 120898ms ago)
18:30:19.518  Last message was received 120898 ago, exceeding the threshold of 120000ms, closing connection...

The data gap — rate(ws_message_total{app="gsr",env="production",direction="received"}[30s]) at 10s step: 346/s at 18:28:40 → 0.2/s from 18:29:20 through 18:30:40 → 311/s at 18:30:50.

504s land on the cycle — baseline 0.1–0.9/s, spikes at 14:23, 15:25, 16:27, 17:29, 18:31 of 16–26/s, each 62 min apart.

GSR never closes the socket — ws_connection_active was 1 at all 360 one-minute samples over 6h, and the only closure code in production is 1000, the framework's own close.

The fix

Two changes, one for each half of the 62 minutes.

1. Don't wait to be cut off (authutils.ts, price.ts, tokenRefresh.ts)

Parse validUntil into an expiry, cache the token, and five minutes ahead of expiry renew in place via GSR's PUT /token — the provider's documented renewal path, removed from this adapter in #2459 and restored here. If renewal is refused, reconnect immediately instead, still ahead of expiry. Either way there is no stall.

2. Make the safety net land inside the cache's lifetime (config/index.ts)

envDefaultOverrides: {
  WS_SUBSCRIPTION_UNRESPONSIVE_TTL: 30_000,
}

The token travels in the WebSocket handshake, so a successful REST renewal is not proof the session survived. Rather than hand-roll a second liveness check, this lowers the framework's existing one so it fires while cached prices are still fresh. It runs every background pass, watches the same lastMessageReceivedAt, and catches a stall from any cause — not only one following a renewal.

Precedence is env var ?? envDefaultOverrides ?? framework default, so this is a shipped default that infra and node operators can still override.

Shipping it in the image rather than as deployed config is deliberate: the issue affects node-operator-run instances too, and a Helm value would only reach Chainlink-run pods.

Why 30s

Bounded above by CACHE_MAX_AGE (90s) minus reconnect time; bounded below by the longest normal gap between cache-writing messages. GSR streams 200–350 msg/s on this deployment, so 30s of silence is unambiguous. Framework range is 1 000–180 000.

Also included

Subscription messages are now batched via customSubscriptionMessages, which requires a framework bump. Carried over from #5252, rebased onto the current framework release: 2.17.1 → 2.19.1.

Testing

Test Suites: 4 passed    Tests: 18 passed    Snapshots: 3 passed
tsc (build + test configs), eslint, prettier — clean

New unit coverage for token expiry parsing, PUT /token renewal and its failure path, refresh scheduling, and the TTL override — the last pinned against both the CACHE_MAX_AGE constraint and the framework's accepted range so it can't silently drift back.

No snapshot changes: the existing integration snapshots pass unmodified.

Not verified — needs a soak

PUT /token is exercised only against mocks. Whether GSR still supports it, and whether a renewal actually extends a live session, are unproven without provider credentials. If it is refused, refreshOrReconnect falls back to a proactive reconnect ~5 min before expiry, which still removes the outage — so this degrades safely, but a stage soak should confirm the intended path. Watch for Token renewed, expires at ... versus Token renewal failed (...).

What changed from #5252

  • Corrected the stated root cause. DF-26076: Fix GSR adapter hourly WebSocket disconnections #5252 says "GSR's server closed the connection at the 1-hour mark, reconnection attempts used the expired token and failed." Production contradicts both halves: GSR never closes the socket, and the framework re-invokes the options callback on every reconnect (websocket.ts:553-554), so reconnects always mint a fresh token and succeed first try. The outage is detection latency.
  • Implemented the liveness verification it promised but did not contain. Its changeset claimed the adapter "verifies that data is still arriving shortly after the old expiry"; livenessTimer was declared and cleared, but never assigned. Replaced by the framework's own check (see fix Testing #2), which is continuously armed rather than only after a renewal.
  • Fixed a version regression — its package.json rolled the adapter back 2.6.3 → 2.5.8.
  • Dropped .github/workflows/publish-internal.yml (+246), unrelated to the fix.
  • Rebased onto current main (it branched 2026-07-31) and squashed 16 commits into one.

Follow-ups (not in this PR)

  • gsr-data-streams is in a permanent 2-minute reconnect loop in production — 100 unresponsive-closes in 6h, 504 baseline ~9.7/s vs gsr's ~0.4/s. Different period, likely different cause: lastMessageReceivedAt only advances on messages that produce a cache write, so a connection receiving only error frames looks permanently unresponsive.
  • The framework's WS_SUBSCRIPTION_UNRESPONSIVE_TTL default (120s) exceeds the CACHE_MAX_AGE default (90s), so any adapter whose provider stalls silently inherits this same bug shape out of the box. Worth revisiting upstream.

🤖 Generated with Claude Code

@changeset-bot

changeset-bot Bot commented Sep 10, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: c451e7c

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 1 package
Name Type
@chainlink/gsr-adapter Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@cl-efornaciari
cl-efornaciari force-pushed the fix/DF-26076/gsr-token-refresh-v2 branch 2 times, most recently from 3412f02 to b37d1b9 Compare September 11, 2026 23:13
A GSR session stops delivering data one hour after it is opened, without
closing the socket or sending any notice. The adapter had no way to know,
so the only thing that noticed was the framework's staleness check at
WS_SUBSCRIPTION_UNRESPONSIVE_TTL (120s) - by which point cached prices had
aged out at CACHE_MAX_AGE (90s), ending every cycle in a burst of 504s.

Production showed this on a 62 minute period: 60 minutes of data plus 120
seconds to detect, repeating for 12 consecutive cycles.

Two changes, one for each half:

- Read validUntil from the token response and, five minutes ahead of
  expiry, renew in place via GSR's PUT /token (removed in #2459, restored
  here) rather than waiting to be cut off. A refused renewal reconnects
  immediately instead, still ahead of expiry.

- Default WS_SUBSCRIPTION_UNRESPONSIVE_TTL to 30s for this adapter. The
  token travels in the WebSocket handshake, so a successful renewal is not
  proof the session survived. Keeping the framework's own liveness check
  inside CACHE_MAX_AGE means any stall - from a renewal that did not take,
  or any other cause - is caught and reconnected while cached prices are
  still being served. Shipping it as an adapter default rather than as
  deployed config means node operators get it too. Still overridable via
  the environment.

Also batches subscription messages via customSubscriptionMessages, which
requires the framework bump to 2.19.1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant