Repository navigation
Conversation
ScalarOperations -> AutoVectorizedSimdOperations: the default is C2 auto-vectorized loops, and "Scalar" collides with Vortex's scalar values. VectorSupport -> SimdOperationsSupport, operations() -> preferred() (JDK ArraysSupport / VectorSpecies.ofPreferred vocabulary). The selector drops nothing else: it only ever exposed operations(). Pure rename, no behavior change. Package stays core.simd. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… ptypes (#516) The writer and the reader each carried their own copies of the same loops: min/max in ArrayStats and again in PrimitiveEncodingEncoder, three widening loops, four all-values-equal checks, the run counter, the zone-map sums, the dictionary code bound. They now live in SimdOperations (minMax, sum, sumFloating, runs, allEqual, widenInto and widenArrayInto, narrowInto and narrowArrayInto, maxUnsigned) and the callers use them: ArrayStats, PrimitiveEncodingEncoder, PrimitiveArrays, ConstantEncodingEncoder, MaskedEncodingEncoder, AlpRdEncodingEncoder, FrameOfReferenceEncodingEncoder, ZigZagEncodingEncoder, ZoneMapStats and DictLayoutDecoder. Every kernel accepts all eleven ptypes: floats widen as their raw bits and narrow back into float carriers, the 8-byte types narrow as a copy, maxUnsigned covers U64, and minMax covers each type in its natural order. Comparing with the Rust source turned up two statistics differences, fixed here: float min/max now order -0.0 before 0.0 and skip NaN (Rust's total_compare with skip_nans; the old first-zero rule was a divergence), and U64 min/max and ArrayStats' dense span are unsigned. Rust's zone Sum also skips NaN and ours does not; that is #523, documented in docs/compatibility.md and left alone. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JbquncyA7mTMJbDJo1s5HD
VectorApiSimdOperations implements every SimdOperations kernel and width as an explicit Vector API loop over the CPU's preferred species with a scalar tail, using the API's own reductions (reduceLanes); only F16, which has no half-float lanes, stays a plain loop. SimdOperationsSupport selects it when the JVM is launched with --add-modules jdk.incubator.vector and the CPU has 128-bit vectors, and otherwise keeps the auto-vectorized implementation, also if the Vector API cannot be linked. The flag is the opt-in (ADR 0005 rules out a per-scan switch). The shape follows Hardwood's VectorOperations (Apache-2.0), credited in the class. It departs from Rust in two reductions, as a recorded decision (docs/compatibility.md, CLAUDE.md, the class javadoc): float sums accumulate lane-wise, so their last bits can differ from Rust's sequential f64 sum, and I64/U64 sum overflow is detected per lane instead of per prefix. The default implementation stays Rust-parity, and 131 differential tests pin every other kernel of the two implementations to each other. Measured with JMH on NEON, AVX2 and AVX-512, bytes and shorts are 2x to 98x faster than the C2 loop (min/max, allEqual, maxUnsigned, runs, sum), delta/undelta/pack 1.3x to 3x, float sums 2x to 8x; widening and narrowing, and 32-bit-lane min/max, still lose to C2. Narrow sums and widening are built without cross-width Vector API conversions, which are not intrinsified (about 40x slower) or are slow on AVX2. Build wiring: core compiles against the incubating module and every test JVM runs with it; a second surefire pass runs the SIMD tests without it to cover the fallback. JDK 26 reports the incubating module as a compiler warning and its javadoc cannot see it, so core silences that lint category and a JDK 26+ profile adds the module for javadoc. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JbquncyA7mTMJbDJo1s5HD
SimdVectorApiBenchmark runs each kernel with both implementations in one JVM (it lives in the core.simd package to reach the package-private classes), and scripts/simd-benchmark.sh runs the valid (kernel, ptype) groups and prints a table of how many times faster the Vector API is, with JMH's confidence interval and a verdict: a row is a win or a loss only when the whole range clears 1, otherwise it says "within noise". The SIMD benchmark workflow runs it on demand and when the SIMD code changes, on x86-64 (as allocated and capped at AVX2 with -XX:MaxVectorSize=32), 64-bit ARM Linux and Apple silicon, one fork by default, and writes the table to the job summary. Runners with the same label are not the same CPU (an ubuntu-latest run came back AVX2 once and AVX-512 the next), which is why the x86 width is pinned. It gates nothing: single-fork numbers on shared runners move by tens of percent. It is the second-architecture evidence ADR 0005 asks for. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JbquncyA7mTMJbDJo1s5HD
dfa1
force-pushed
the
refactor/516-simd-naming
branch
from
October 10, 2026 20:33
8232bd1 to
3f8b2ea
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Consolidates the primitive-array loops of the reader and writer behind
SimdOperationsand adds an opt-in Vector API implementation. Tracks #516.1. Consolidation (
SimdOperations,core.simd)minMax,sum,sumFloating,runs,allEqual,widen*/narrow*(segment and heap array) andmaxUnsignednow live in one place and cover all eleven ptypes. Callers that carried their own copies of these loops use the kernels:ArrayStats,PrimitiveEncodingEncoder(min/max, sums),PrimitiveArrays,ConstantEncodingEncoder,MaskedEncodingEncoder,AlpRdEncodingEncoder,FrameOfReferenceEncodingEncoder,ZigZagEncodingEncoder,ZoneMapStats,DictLayoutDecoder. Renames:ScalarOperations->AutoVectorizedSimdOperations,VectorSupport->SimdOperationsSupport#preferred().2. Vector API implementation (
VectorApiSimdOperations)Selected only when the JVM is launched with
--add-modules jdk.incubator.vectorand the CPU has 128-bit vectors; otherwise the default stays (ADR 0005 rules out a per-scan flag). Every kernel and width is an explicit vector loop and the reductions use the Vector API's own (reduceLanes). Plain loops remain only forF16(no half-float lanes) and scalar tails. Pattern follows Hardwood'sVectorOperations(Apache-2.0), credited in the class.3. Documented differences from Rust
The Vector API implementation is a deliberate departure in two places, recorded in
docs/compatibility.md,CLAUDE.mdand the class javadoc: float sums accumulate lane-wise (last-bit differences from Rust's sequentialf64sum), andI64/U64sum overflow is detected per lane rather than per prefix. The default implementation stays Rust-parity.Parity fixes found on the way: float min/max now order
-0.0before0.0and skip NaN (Rust'stotal_comparewithskip_nans), andU64min/max/dense-span are unsigned. Found and not fixed: Rust's zoneSumskips NaN and ours does not (#523).Measured: how many times faster than the C2 loop (JMH, 262144 elements, 1 fork of 5 iterations, CI runners)
Above 1x the Vector API is faster. These are single-fork numbers on shared runners; the benchmark job marks a row a win or a loss only when JMH's 99.9% interval is clear of 1 and the full tables (with intervals) are in each job's summary.
minMaxI8minMaxI16minMaxI32minMaxI64allEqual_constantI8allEqual_constantI16allEqual_constantI32maxUnsignedI8maxUnsignedI32runs_noRunsI8runs_noRunsI32sumI8sumI16sumI32deltaI32undeltaI32packI32sumFloatingF64widenArrayI8widenArrayI32narrowArrayI8Notes:
-XX:MaxVectorSize(AVX2). The AVX-512 leg also moves a lot between runs of the same code (sumon bytes read 3.2x and 1.9x in two runs), so treat single-fork numbers on shared runners as +-40%.delta/undelta/pack(1.3x to 3x), and float sums (2x to 7x). Narrowsumreached its numbers by splitting wider lanes with shifts instead of converting lane widths, which is cheap on NEON and was not on AVX2 (bytes 0.19x -> 1.0x, shorts 0.41x -> 1.8x there).minMax/maxUnsignedon 32-bit lanes, andsumonints on x86 (0.6x to 0.7x). A lane-splitintsum was tried and measured worse on all four legs, so it was reverted. The 64-bit rows are not consistent across hardware (minMaxonlongs: 2.0x on AVX2, 0.8x on AVX-512), so those are not a claim.intlanes.CI
coreruns the SIMD tests without it to cover the fallback. JDK 26 needed a lint exclusion and a javadoc profile for the incubating module.SIMD benchmarkworkflow (manual, or when the SIMD code changes) runsscripts/simd-benchmark.shon x86-64 (as allocated and capped at AVX2), ARM Linux and Apple silicon, and writes a speedup table with JMH confidence intervals to the job summary. Informational only. It is the second-architecture evidence ADR 0005 asks for.Not in this PR
ADR 0005's status line (still "Deferred") should be updated once these numbers are agreed.
PrimitiveArrays.compactand the codec-specific loops (Pco, ALP, FSST) are untouched.Test plan
./mvnw verify(all modules) passes locally on JDK 25 and 26JavaWritesRustReads,RustWritesJavaReads) green🤖 Generated with Claude Code