Skip to content

SIMD kernels in one place, plus an opt-in Vector API implementation (#516) - #524

Open
dfa1 wants to merge 4 commits into
mainfrom
refactor/516-simd-naming
Open

dfa1 wants to merge 4 commits into
mainfrom
refactor/516-simd-naming

Conversation

@dfa1

@dfa1 dfa1 commented Oct 10, 2026 •

Copy link
Copy Markdown
Owner

What

Consolidates the primitive-array loops of the reader and writer behind SimdOperations and adds an opt-in Vector API implementation. Tracks #516.

1. Consolidation (SimdOperations, core.simd)

minMax, sum, sumFloating, runs, allEqual, widen*/narrow* (segment and heap array) and maxUnsigned now live in one place and cover all eleven ptypes. Callers that carried their own copies of these loops use the kernels: ArrayStats, PrimitiveEncodingEncoder (min/max, sums), PrimitiveArrays, ConstantEncodingEncoder, MaskedEncodingEncoder, AlpRdEncodingEncoder, FrameOfReferenceEncodingEncoder, ZigZagEncodingEncoder, ZoneMapStats, DictLayoutDecoder. Renames: ScalarOperations -> AutoVectorizedSimdOperations, VectorSupport -> SimdOperationsSupport#preferred().

2. Vector API implementation (VectorApiSimdOperations)

Selected only when the JVM is launched with --add-modules jdk.incubator.vector and the CPU has 128-bit vectors; otherwise the default stays (ADR 0005 rules out a per-scan flag). Every kernel and width is an explicit vector loop and the reductions use the Vector API's own (reduceLanes). Plain loops remain only for F16 (no half-float lanes) and scalar tails. Pattern follows Hardwood's VectorOperations (Apache-2.0), credited in the class.

3. Documented differences from Rust

The Vector API implementation is a deliberate departure in two places, recorded in docs/compatibility.md, CLAUDE.md and the class javadoc: float sums accumulate lane-wise (last-bit differences from Rust's sequential f64 sum), and I64/U64 sum overflow is detected per lane rather than per prefix. The default implementation stays Rust-parity.

Parity fixes found on the way: float min/max now order -0.0 before 0.0 and skip NaN (Rust's total_compare with skip_nans), and U64 min/max/dense-span are unsigned. Found and not fixed: Rust's zone Sum skips NaN and ours does not (#523).

Measured: how many times faster than the C2 loop (JMH, 262144 elements, 1 fork of 5 iterations, CI runners)

Above 1x the Vector API is faster. These are single-fork numbers on shared runners; the benchmark job marks a row a win or a loss only when JMH's 99.9% interval is clear of 1 and the full tables (with intervals) are in each job's summary.

kernel / type Apple silicon (NEON 128) ARM Linux (NEON 128) x86 AVX2 (256, pinned) x86 AVX-512 (512)
minMax I8 11.9x 9.7x 43.7x 64.6x
minMax I16 5.7x 4.9x 21.4x 43.7x
minMax I32 0.7x 0.8x 0.7x 0.8x
minMax I64 1.1x 1.1x 2.0x 0.9x
allEqual_constant I8 12.7x 9.5x 29.3x 40.8x
allEqual_constant I16 6.6x 4.8x 14.8x 15.6x
allEqual_constant I32 1.1x 1.0x 1.2x 1.5x
maxUnsigned I8 6.7x 5.8x 16.1x 27.5x
maxUnsigned I32 1.3x 0.8x 0.8x 0.9x
runs_noRuns I8 4.3x 4.4x 5.7x 12.5x
runs_noRuns I32 0.7x 0.8x 1.6x 2.4x
sum I8 6.9x 9.3x 1.0x 3.2x
sum I16 2.8x 4.1x 1.8x 2.4x
sum I32 2.0x 2.0x 0.7x 0.8x
delta I32 1.8x 2.2x 2.7x 2.5x
undelta I32 2.0x 2.1x 2.4x 2.9x
pack I32 1.4x 1.5x 2.5x 2.9x
sumFloating F64 2.0x 1.9x 4.0x 7.9x
widenArray I8 0.8x 0.5x 0.5x 0.8x
widenArray I32 0.4x 0.4x 0.6x 0.9x
narrowArray I8 0.4x 0.4x 0.4x 0.8x

Notes:

  • All four columns are one commit's run (the final code). The ubuntu-latest label does not pin the CPU: runners have come back with 256-bit and 512-bit vectors, so the x86 job runs twice, as allocated (AVX-512 here) and capped at 32 bytes with -XX:MaxVectorSize (AVX2). The AVX-512 leg also moves a lot between runs of the same code (sum on bytes read 3.2x and 1.9x in two runs), so treat single-fork numbers on shared runners as +-40%.
  • Wins: bytes and shorts (2x to 98x), delta/undelta/pack (1.3x to 3x), and float sums (2x to 7x). Narrow sum reached its numbers by splitting wider lanes with shifts instead of converting lane widths, which is cheap on NEON and was not on AVX2 (bytes 0.19x -> 1.0x, shorts 0.41x -> 1.8x there).
  • Still behind C2: widening and narrowing (0.4x to 1.0x; at best parity), minMax/maxUnsigned on 32-bit lanes, and sum on ints on x86 (0.6x to 0.7x). A lane-split int sum was tried and measured worse on all four legs, so it was reverted. The 64-bit rows are not consistent across hardware (minMax on longs: 2.0x on AVX2, 0.8x on AVX-512), so those are not a claim.
  • A direct Vector API conversion across a lane ratio of 4+ into 64-bit lanes is not intrinsified (about 40x slower); widening and narrowing go through int lanes.

CI

  • Every test JVM runs with the module (the suite exercises the Vector API); a second surefire pass in core runs the SIMD tests without it to cover the fallback. JDK 26 needed a lint exclusion and a javadoc profile for the incubating module.
  • New SIMD benchmark workflow (manual, or when the SIMD code changes) runs scripts/simd-benchmark.sh on x86-64 (as allocated and capped at AVX2), ARM Linux and Apple silicon, and writes a speedup table with JMH confidence intervals to the job summary. Informational only. It is the second-architecture evidence ADR 0005 asks for.

Not in this PR

ADR 0005's status line (still "Deferred") should be updated once these numbers are agreed. PrimitiveArrays.compact and the codec-specific loops (Pco, ALP, FSST) are untouched.

Test plan

  • ./mvnw verify (all modules) passes locally on JDK 25 and 26
  • CI matrix green on Linux, Windows and macOS, Java 25 and 26
  • 131 differential tests pin every kernel of the Vector API implementation to the default one; the SIMD tests also pass without the module
  • Rust interop (JavaWritesRustReads, RustWritesJavaReads) green
  • Reviewed: benchmark summaries on the AVX2-pinned and default x86 legs

🤖 Generated with Claude Code

@dfa1 dfa1 closed this Oct 10, 2026
@dfa1 dfa1 reopened this Oct 10, 2026
dfa1 and others added 4 commits October 10, 2026 22:30
ScalarOperations -> AutoVectorizedSimdOperations: the default is C2
auto-vectorized loops, and "Scalar" collides with Vortex's scalar values.
VectorSupport -> SimdOperationsSupport, operations() -> preferred() (JDK
ArraysSupport / VectorSpecies.ofPreferred vocabulary). The selector drops
nothing else: it only ever exposed operations().

Pure rename, no behavior change. Package stays core.simd.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… ptypes (#516)

The writer and the reader each carried their own copies of the same loops: min/max in
ArrayStats and again in PrimitiveEncodingEncoder, three widening loops, four
all-values-equal checks, the run counter, the zone-map sums, the dictionary code bound.
They now live in SimdOperations (minMax, sum, sumFloating, runs, allEqual, widenInto and
widenArrayInto, narrowInto and narrowArrayInto, maxUnsigned) and the callers use them:
ArrayStats, PrimitiveEncodingEncoder, PrimitiveArrays, ConstantEncodingEncoder,
MaskedEncodingEncoder, AlpRdEncodingEncoder, FrameOfReferenceEncodingEncoder,
ZigZagEncodingEncoder, ZoneMapStats and DictLayoutDecoder.

Every kernel accepts all eleven ptypes: floats widen as their raw bits and narrow back into
float carriers, the 8-byte types narrow as a copy, maxUnsigned covers U64, and minMax
covers each type in its natural order.

Comparing with the Rust source turned up two statistics differences, fixed here: float
min/max now order -0.0 before 0.0 and skip NaN (Rust's total_compare with skip_nans; the
old first-zero rule was a divergence), and U64 min/max and ArrayStats' dense span are
unsigned. Rust's zone Sum also skips NaN and ours does not; that is #523, documented in
docs/compatibility.md and left alone.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JbquncyA7mTMJbDJo1s5HD
VectorApiSimdOperations implements every SimdOperations kernel and width as an explicit
Vector API loop over the CPU's preferred species with a scalar tail, using the API's own
reductions (reduceLanes); only F16, which has no half-float lanes, stays a plain loop.
SimdOperationsSupport selects it when the JVM is launched with --add-modules
jdk.incubator.vector and the CPU has 128-bit vectors, and otherwise keeps the
auto-vectorized implementation, also if the Vector API cannot be linked. The flag is the
opt-in (ADR 0005 rules out a per-scan switch). The shape follows Hardwood's
VectorOperations (Apache-2.0), credited in the class.

It departs from Rust in two reductions, as a recorded decision (docs/compatibility.md,
CLAUDE.md, the class javadoc): float sums accumulate lane-wise, so their last bits can
differ from Rust's sequential f64 sum, and I64/U64 sum overflow is detected per lane
instead of per prefix. The default implementation stays Rust-parity, and 131 differential
tests pin every other kernel of the two implementations to each other.

Measured with JMH on NEON, AVX2 and AVX-512, bytes and shorts are 2x to 98x faster than the
C2 loop (min/max, allEqual, maxUnsigned, runs, sum), delta/undelta/pack 1.3x to 3x, float
sums 2x to 8x; widening and narrowing, and 32-bit-lane min/max, still lose to C2. Narrow
sums and widening are built without cross-width Vector API conversions, which are not
intrinsified (about 40x slower) or are slow on AVX2.

Build wiring: core compiles against the incubating module and every test JVM runs with it;
a second surefire pass runs the SIMD tests without it to cover the fallback. JDK 26 reports
the incubating module as a compiler warning and its javadoc cannot see it, so core silences
that lint category and a JDK 26+ profile adds the module for javadoc.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JbquncyA7mTMJbDJo1s5HD
SimdVectorApiBenchmark runs each kernel with both implementations in one JVM (it lives in
the core.simd package to reach the package-private classes), and scripts/simd-benchmark.sh
runs the valid (kernel, ptype) groups and prints a table of how many times faster the
Vector API is, with JMH's confidence interval and a verdict: a row is a win or a loss only
when the whole range clears 1, otherwise it says "within noise".

The SIMD benchmark workflow runs it on demand and when the SIMD code changes, on x86-64
(as allocated and capped at AVX2 with -XX:MaxVectorSize=32), 64-bit ARM Linux and Apple
silicon, one fork by default, and writes the table to the job summary. Runners with the
same label are not the same CPU (an ubuntu-latest run came back AVX2 once and AVX-512 the
next), which is why the x86 width is pinned. It gates nothing: single-fork numbers on
shared runners move by tens of percent. It is the second-architecture evidence ADR 0005
asks for.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JbquncyA7mTMJbDJo1s5HD
@dfa1
dfa1 force-pushed the refactor/516-simd-naming branch from 8232bd1 to 3f8b2ea Compare October 10, 2026 20:33

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant