Repository navigation
Fix race conditions caused by missing fences in OpenCL kernels for cholesky_decompose, tridiagonalization and diag_inv - #3432
Open
jachymb wants to merge 5 commits into
Conversation
added 5 commits
October 4, 2026 12:56
Contributor
Jenkins Console Log Machine informationDistributor ID: Ubuntu Description: Ubuntu 20.04.3 LTS Release: 20.04 Codename: focal CPU: Architecture: x86_64 CPU op-mode(s): 32-bit, 64-bit Byte Order: Little Endian Address sizes: 43 bits physical, 48 bits virtual CPU(s): 256 On-line CPU(s) list: 0-255 Thread(s) per core: 2 Core(s) per socket: 64 Socket(s): 2 NUMA node(s): 2 Vendor ID: AuthenticAMD CPU family: 23 Model: 49 Model name: AMD EPYC 7742 64-Core Processor Stepping: 0 Frequency boost: enabled CPU MHz: 1497.090 CPU max MHz: 3416.0681 CPU min MHz: 1500.0000 BogoMIPS: 4491.56 Virtualization: AMD-V L1d cache: 4 MiB L1i cache: 4 MiB L2 cache: 64 MiB L3 cache: 512 MiB NUMA node0 CPU(s): 0-63,128-191 NUMA node1 CPU(s): 64-127,192-255 Vulnerability Gather data sampling: Not affected Vulnerability Indirect target selection: Not affected Vulnerability Itlb multihit: Not affected Vulnerability L1tf: Not affected Vulnerability Mds: Not affected Vulnerability Meltdown: Not affected Vulnerability Mmio stale data: Not affected Vulnerability Old microcode: Not affected Vulnerability Reg file data sampling: Not affected Vulnerability Retbleed: Mitigation; untrained return thunk; SMT enabled with STIBP protection Vulnerability Spec rstack overflow: Mitigation; Safe RET Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization Vulnerability Spectre v2: Mitigation; Retpolines; IBPB conditional; STIBP always-on; RSB filling; PBRSB-eIBRS Not affected; BHI Not affected Vulnerability Srbds: Not affected Vulnerability Tsa: Not affected Vulnerability Tsx async abort: Not affected Vulnerability Vmscape: Mitigation; IBPB before exit to userspace Flags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_good nopl xtopology nonstop_tsc cpuid extd_apicid aperfmperf rapl pni pclmulqdq monitor ssse3 fma cx16 sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand lahf_lm cmp_legacy svm extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw ibs skinit wdt tce topoext perfctr_core perfctr_nb bpext perfctr_llc mwaitx cpb cat_l3 cdp_l3 hw_pstate ssbd mba ibrs ibpb stibp vmmcall fsgsbase bmi1 avx2 smep bmi2 cqm rdt_a rdseed adx smap clflushopt clwb sha_ni xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local clzero irperf xsaveerptr rdpru wbnoinvd amd_ppin arat npt lbrv svm_lock nrip_save tsc_scale vmcb_clean flushbyasid decodeassists pausefilter pfthreshold avic v_vmsave_vmload vgif v_spec_ctrl umip rdpid overflow_recov succor smca sev sev_es G++: g++ (Ubuntu 9.4.0-1ubuntu1~20.04) 9.4.0 Copyright (C) 2019 Free Software Foundation, Inc. This is free software; see the source for copying conditions. There is NO warranty; not even for MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. Clang: clang version 10.0.0-4ubuntu1 Target: x86_64-pc-linux-gnu Thread model: posix InstalledDir: /usr/bin |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Introduction
On my machine a noticed that the GPU kernel for
cholesky_decomposeis buggy and actually produces wrong results. I confirmed with Oclgrind that this is due to a fencing race condition. The attached test fails on my Radeon GPU against current develop and is fixed with this patch. I then tested other kernels in the repo with Oclgrind and found possible fencing race conditions in the kernels fortridiagonalizationanddiag_inv. For these two, I was unable to construct a failing test so maybe they are just latent issues. However, this kind of problems may be machine-specific, so I believe it's better fixed.This fixes #3429. See also #3430.
The rest of this commentary was made by Fable5.1 which helped me debug this.
Summary
In three kernels, work items read or overwrite buffer entries that other work items access on the other side of a
barrier(CLK_LOCAL_MEM_FENCE). That flag orders operations on local memory only; such accesses needCLK_GLOBAL_MEM_FENCE.cholesky_decompose: both barriers now fence global memory. This kernel returns wrong results on an AMD GPU. With more than 64 work items, some of them read entries of the matrix as they were before other work items overwrote them, socholesky_decomposeof amatrix_clwith more than 64 rows silently returns a wrong factor.tridiagonalization_householder: the barrier after the first loop now fences local and global memory.diag_inv: abarrier(CLK_GLOBAL_MEM_FENCE)now separates the reads of the matrix from the final copy into it.The last two were found with the race detector of Oclgrind (
oclgrind --data-races <test>); no wrong result from them has been seen. On the small cases of their tests, Oclgrind reports 14, 62 and more than 1000 races in the three kernels on develop and none with this PR.Tests
MathMatrixOpenCL.cholesky_decompose_cpu_vs_cl_one_work_groupfactors a 200Ă—200 matrix, which is one call of the kernel with 200 work items. On the AMD GPU it fails on develop in 5 of 5 runs and passes with the fix.Side Effects
Release notes
Fix race conditions caused by missing fences in OpenCL kernels for
cholesky_decompose,tridiagonalizationanddiag_invChecklist
Copyright holder: Jáchym Barvínek jachymb@gmail.com
The copyright holder is typically you or your assignee, such as a university or company. By submitting this pull request, the copyright holder is agreeing to the license the submitted work under the following licenses:
- Code: BSD 3-clause (https://opensource.org/licenses/BSD-3-Clause)
- Documentation: CC-BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
the basic tests are passing
./runTests.py test/unit)make test-headers)make test-math-dependencies)make doxygen)make cpplint)the code is written in idiomatic C++ and changes are documented in the doxygen
the new changes are tested