DF-26076: fix hourly GSR WebSocket data stalls - #5381
Open
cl-efornaciari wants to merge 1 commit into
Open
cl-efornaciari wants to merge 1 commit into
cl-efornaciari wants to merge 1 commit into
Conversation
🦋 Changeset detectedLatest commit: c451e7c The changes in this PR will be included in the next version bump. This PR includes changesets to release 1 package
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
cl-efornaciari
force-pushed
the
fix/DF-26076/gsr-token-refresh-v2
branch
2 times, most recently
from
September 11, 2026 23:13
3412f02 to
b37d1b9
Compare
A GSR session stops delivering data one hour after it is opened, without closing the socket or sending any notice. The adapter had no way to know, so the only thing that noticed was the framework's staleness check at WS_SUBSCRIPTION_UNRESPONSIVE_TTL (120s) - by which point cached prices had aged out at CACHE_MAX_AGE (90s), ending every cycle in a burst of 504s. Production showed this on a 62 minute period: 60 minutes of data plus 120 seconds to detect, repeating for 12 consecutive cycles. Two changes, one for each half: - Read validUntil from the token response and, five minutes ahead of expiry, renew in place via GSR's PUT /token (removed in #2459, restored here) rather than waiting to be cut off. A refused renewal reconnects immediately instead, still ahead of expiry. - Default WS_SUBSCRIPTION_UNRESPONSIVE_TTL to 30s for this adapter. The token travels in the WebSocket handshake, so a successful renewal is not proof the session survived. Keeping the framework's own liveness check inside CACHE_MAX_AGE means any stall - from a renewal that did not take, or any other cause - is caught and reconnected while cached prices are still being served. Shipping it as an adapter default rather than as deployed config means node operators get it too. Still overridable via the environment. Also batches subscription messages via customSubscriptionMessages, which requires the framework bump to 2.19.1. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
cl-efornaciari
force-pushed
the
fix/DF-26076/gsr-token-refresh-v2
branch
from
September 11, 2026 23:18
b37d1b9 to
c451e7c
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes DF-26076. Supersedes #5252 — see What changed from #5252 below, including a correction to its stated root cause.
Symptom
Every GSR EA instance — Chainlink-run and node-operator-run — stops serving data for ~2 minutes, once an hour, and returns a burst of 504s. Reported 2026-07-23, still occurring, and re-reported via NTI-251.
Root cause
A GSR WebSocket session delivers usable data for exactly 60 minutes after it opens, then silently stops. GSR does not close the socket and sends no notice.
The adapter had no way to detect this. It read the token's
validUntilinto a type and then discarded it, and had no refresh timer. So the only thing that noticed was the framework's generic staleness check,WS_SUBSCRIPTION_UNRESPONSIVE_TTL, at its 120s default. That fires 120s after the last usable message, closes with code 1000, and reconnects successfully on the first try.The damage is the blind window, not the reconnect.
CACHE_MAX_AGEis 90s, so cached prices expire ~30s before the framework even notices, and every caller gets a 504 until the new connection is serving.Evidence from production
62.0-minute period, 12 consecutive cycles —
{app="gsr", env="production"} |= "exceeding the threshold":62.0 min = 60 min of data + 120s to detect.
The adapter's own log, the 18:30 cycle:
The data gap —
rate(ws_message_total{app="gsr",env="production",direction="received"}[30s])at 10s step: 346/s at 18:28:40 → 0.2/s from 18:29:20 through 18:30:40 → 311/s at 18:30:50.504s land on the cycle — baseline 0.1–0.9/s, spikes at 14:23, 15:25, 16:27, 17:29, 18:31 of 16–26/s, each 62 min apart.
GSR never closes the socket —
ws_connection_activewas 1 at all 360 one-minute samples over 6h, and the only closure code in production is1000, the framework's own close.The fix
Two changes, one for each half of the 62 minutes.
1. Don't wait to be cut off (
authutils.ts,price.ts,tokenRefresh.ts)Parse
validUntilinto an expiry, cache the token, and five minutes ahead of expiry renew in place via GSR'sPUT /token— the provider's documented renewal path, removed from this adapter in #2459 and restored here. If renewal is refused, reconnect immediately instead, still ahead of expiry. Either way there is no stall.2. Make the safety net land inside the cache's lifetime (
config/index.ts)The token travels in the WebSocket handshake, so a successful REST renewal is not proof the session survived. Rather than hand-roll a second liveness check, this lowers the framework's existing one so it fires while cached prices are still fresh. It runs every background pass, watches the same
lastMessageReceivedAt, and catches a stall from any cause — not only one following a renewal.Precedence is
env var ?? envDefaultOverrides ?? framework default, so this is a shipped default that infra and node operators can still override.Shipping it in the image rather than as deployed config is deliberate: the issue affects node-operator-run instances too, and a Helm value would only reach Chainlink-run pods.
Why 30s
Bounded above by
CACHE_MAX_AGE(90s) minus reconnect time; bounded below by the longest normal gap between cache-writing messages. GSR streams 200–350 msg/s on this deployment, so 30s of silence is unambiguous. Framework range is 1 000–180 000.Also included
Subscription messages are now batched via
customSubscriptionMessages, which requires a framework bump. Carried over from #5252, rebased onto the current framework release: 2.17.1 → 2.19.1.Testing
New unit coverage for token expiry parsing,
PUT /tokenrenewal and its failure path, refresh scheduling, and the TTL override — the last pinned against both theCACHE_MAX_AGEconstraint and the framework's accepted range so it can't silently drift back.No snapshot changes: the existing integration snapshots pass unmodified.
Not verified — needs a soak
PUT /tokenis exercised only against mocks. Whether GSR still supports it, and whether a renewal actually extends a live session, are unproven without provider credentials. If it is refused,refreshOrReconnectfalls back to a proactive reconnect ~5 min before expiry, which still removes the outage — so this degrades safely, but a stage soak should confirm the intended path. Watch forToken renewed, expires at ...versusToken renewal failed (...).What changed from #5252
optionscallback on every reconnect (websocket.ts:553-554), so reconnects always mint a fresh token and succeed first try. The outage is detection latency.livenessTimerwas declared and cleared, but never assigned. Replaced by the framework's own check (see fix Testing #2), which is continuously armed rather than only after a renewal.package.jsonrolled the adapter back 2.6.3 → 2.5.8..github/workflows/publish-internal.yml(+246), unrelated to the fix.main(it branched 2026-07-31) and squashed 16 commits into one.Follow-ups (not in this PR)
gsr-data-streamsis in a permanent 2-minute reconnect loop in production — 100 unresponsive-closes in 6h, 504 baseline ~9.7/s vsgsr's ~0.4/s. Different period, likely different cause:lastMessageReceivedAtonly advances on messages that produce a cache write, so a connection receiving only error frames looks permanently unresponsive.WS_SUBSCRIPTION_UNRESPONSIVE_TTLdefault (120s) exceeds theCACHE_MAX_AGEdefault (90s), so any adapter whose provider stalls silently inherits this same bug shape out of the box. Worth revisiting upstream.🤖 Generated with Claude Code