In OKD deployments (okd-project/okd) the nodes run CentOS Stream CoreOS (SCOS), whose NFD labels report feature.node.kubernetes.io/system-os_release.ID=centos and …VERSION_ID=10. getOSTag() in internal/state/nodepool.go therefore renders the driver image suffix as centos10 → nvcr.io/nvidia/driver:<version>-centos10, a tag that is not published, and the driver pod fails with ImagePullBackOff. RHEL and its clones are already special-cased there (ol, rocky, rhel → major-only suffix); centos falls through to the default branch.
The documented workaround (confirmed in #1533) is pinning spec.driver.version to a digest, which is used verbatim. That works, but freezes the image generation: rebuilds never flow under the same pin, leaving no tag-based upgrade path for the pinned reference. In practice we also hit the pinned tag being removed from nvcr.io (gpu-driver-container#985), and the newer images we would otherwise move to were blocked by a separate issue (gpu-driver-container#984).
Expected behavior: users running OKD/SCOS should be able to select a published, compatible driver image suffix (or otherwise prevent the operator from constructing an unavailable -centos10 tag).
One shape that would solve it: a CRD-level override of the derived OS tag — e.g. spec.driver.osTag: rhel10.0 — taking precedence over the os_release-derived suffix. A smaller alternative in the same spirit: extending getOSTag's existing ol/rocky/rhel normalization to map centos → rhel. For context on compatibility: OKD is the community OpenShift distribution and the operator's OpenShift/DTK integration is active on these nodes (not a standalone CentOS host), and we build against the published -rhel10 driver images on SCOS today (digest-pinned 580.126.20-rhel10.0, kernel 6.12.0-233.el10) with the kmod building and loading fine. Happy with either shape — whichever better fits the project's support model.
Related: #1533 (original ask, resolved via the digest workaround, closed stale) · #2811 (downstream: repoConfig in DTK container) · gpu-driver-container#984 / #985
Environment: GPU Operator v26.3.3 · OKD 4.22.0-okd-scos.6 · CentOS Stream CoreOS 10, kernel 6.12.0-233.el10.x86_64 · NVIDIA L4 · amd64
In OKD deployments (okd-project/okd) the nodes run CentOS Stream CoreOS (SCOS), whose NFD labels report
feature.node.kubernetes.io/system-os_release.ID=centosand…VERSION_ID=10.getOSTag()in internal/state/nodepool.go therefore renders the driver image suffix ascentos10→nvcr.io/nvidia/driver:<version>-centos10, a tag that is not published, and the driver pod fails with ImagePullBackOff. RHEL and its clones are already special-cased there (ol,rocky,rhel→ major-only suffix);centosfalls through to the default branch.The documented workaround (confirmed in #1533) is pinning
spec.driver.versionto a digest, which is used verbatim. That works, but freezes the image generation: rebuilds never flow under the same pin, leaving no tag-based upgrade path for the pinned reference. In practice we also hit the pinned tag being removed from nvcr.io (gpu-driver-container#985), and the newer images we would otherwise move to were blocked by a separate issue (gpu-driver-container#984).Expected behavior: users running OKD/SCOS should be able to select a published, compatible driver image suffix (or otherwise prevent the operator from constructing an unavailable
-centos10tag).One shape that would solve it: a CRD-level override of the derived OS tag — e.g.
spec.driver.osTag: rhel10.0— taking precedence over the os_release-derived suffix. A smaller alternative in the same spirit: extendinggetOSTag's existingol/rocky/rhelnormalization to mapcentos→rhel. For context on compatibility: OKD is the community OpenShift distribution and the operator's OpenShift/DTK integration is active on these nodes (not a standalone CentOS host), and we build against the published-rhel10driver images on SCOS today (digest-pinned580.126.20-rhel10.0, kernel6.12.0-233.el10) with the kmod building and loading fine. Happy with either shape — whichever better fits the project's support model.Related: #1533 (original ask, resolved via the digest workaround, closed stale) · #2811 (downstream: repoConfig in DTK container) · gpu-driver-container#984 / #985
Environment: GPU Operator v26.3.3 · OKD 4.22.0-okd-scos.6 · CentOS Stream CoreOS 10, kernel 6.12.0-233.el10.x86_64 · NVIDIA L4 · amd64