Skip to content

Use intoBitSet to speed up ReqExclBulkScorer with dense exclusions - #16669

Merged
romseygeek merged 8 commits into
apache:mainfrom
kkewwei:optimize_reqEx
Sep 24, 2026
Merged

romseygeek merged 8 commits into
apache:mainfrom
kkewwei:optimize_reqEx

Conversation

@kkewwei

@kkewwei kkewwei commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

Description

Solved: #16668

@kkewwei

kkewwei commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor Author

I added a benchmark task set in luceneutil with wikimediumall:
kkewwei/luceneutil@2b77ca0

After 19 benchmark rounds, the final results are:

                            TaskQPS baseline      StdDevQPS my_modified_version      StdDev                Pct diff p-value
              ReqExclSparse1TopN      140.97      (2.6%)      138.99      (3.3%)   -1.4% (  -7% -    4%) 0.133
                        PKLookup      184.94      (9.2%)      182.53     (10.2%)   -1.3% ( -18% -   19%) 0.670
               ReqExclDense3TopN        2.80      (6.0%)       12.70     (28.3%)  353.6% ( 301% -  412%) 0.000
               ReqExclDense7TopN        1.29      (8.3%)        8.38     (32.3%)  550.7% ( 470% -  644%) 0.000
              ReqExclDense11TopN        0.76      (8.0%)        6.47     (39.1%)  749.2% ( 650% -  865%) 0.000
              ReqExclDense16TopN        0.55      (7.9%)        5.09     (42.7%)  833.0% ( 725% -  958%) 0.000

For the dense path, we can see a 300%+ improvement.

Comment thread lucene/core/src/java/org/apache/lucene/search/ReqExclBulkScorer.java Outdated
@kkewwei

kkewwei commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor Author

I also do the benchmark with MustNotBooleanQueryBenchmark

  • For dense path:
mustNotTermCount Baseline Candidate improvement
3 41.67 ± 1.71 ops/s 2897.34 ± 80.36 ops/s 69.54x
7 44.17 ± 2.15 ops/s 400.85 ± 9.78 ops/s 9.07x
11 32.03 ± 2.06 ops/s 365.19 ± 11.18 ops/s 11.40x
  • For spare path:
mustNotTermCount Baseline Candidate improvement
3 5245.66 ± 123.27 ops/s 5238.78 ± 150.57 ops/s -0.13%
7 4598.34 ± 109.50 ops/s 4692.22 ± 79.55 ops/s +2.04%
11 3790.53 ± 117.27 ops/s 3842.41 ± 88.07 ops/s +1.37%

We can also see a 300%+ improvement in dense path.

@parkertimmins

Copy link
Copy Markdown
Contributor

Nice PR! I recently opened a PR which did something similar, but in a less general way (It only applied to non-scoring queries). This is a better approach. That PR was to speed up not-equals queries, where the excluded filter matches relatively few values. It includes a benchmark showing that must_not performs poorly when it's excluded query is selective, as compared to filter, which perform well when it's included query is selective. The benchmark uses a single range query and varies it's selectivity. I ran the benchmark on this PR and got good results!

Inner query selectivity Filter default MUST_NOT default Filter Panama MUST_NOT Panama
0.01 219.2 218.9 492.5 494.5
0.1 191.4 189.2 327.6 331.3
0.5 187.7 204.0 429.0 430.8
0.9 540.4 513.1 472.7 459.1
0.99 690.7 695.9 487.4 468.4

For comparison, these are the baseline results from a few weeks ago:

Inner query selectivity Filter default MUST_NOT default Filter Panama MUST_NOT Panama
0.01 236.786 ± 5.855 42.162 ± 5.477 545.464 ± 15.000 44.303 ± 2.776
0.10 200.210 ± 21.402 43.665 ± 4.309 337.842 ± 5.168 44.643 ± 0.734
0.50 202.529 ± 1.772 56.723 ± 5.886 470.476 ± 24.148 56.038 ± 0.550
0.90 549.659 ± 69.162 130.545 ± 1.347 517.654 ± 24.159 136.902 ± 8.890
0.99 778.650 ± 75.649 205.053 ± 21.086 523.284 ± 8.663 204.756 ± 14.010

@kkewwei

kkewwei commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

@parkertimmins Thanks for running the benchmark and sharing the results!

I'm exploring whether we can remove the sparse path and use the dense path universally, which would simplify the implementation.

@kkewwei

kkewwei commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

I remove the sparse path and use the dense path universally, test twice, the final results are:
test1:

                            TaskQPS baseline      StdDevQPS my_modified_version      StdDev                Pct diff p-value
              ReqExclSparse1TopN      143.18      (2.9%)      138.73      (2.4%)   -3.1% (  -8% -    2%) 0.000
                        PKLookup      192.52      (8.6%)      197.90      (5.3%)    2.8% ( -10% -   18%) 0.227
               ReqExclDense3TopN        2.83      (7.0%)       13.19     (15.1%)  366.4% ( 321% -  417%) 0.000
               ReqExclDense7TopN        1.29      (8.4%)        8.70     (19.1%)  572.7% ( 502% -  655%) 0.000
              ReqExclDense11TopN        0.77      (8.4%)        6.74     (26.5%)  777.0% ( 684% -  886%) 0.000
              ReqExclDense16TopN        0.55      (8.5%)        5.32     (31.1%)  871.2% ( 766% -  995%) 0.000

test2:

                            TaskQPS baseline      StdDevQPS my_modified_version      StdDev                Pct diff p-value
              ReqExclSparse1TopN      138.79      (2.6%)      134.85      (2.6%)   -2.8% (  -7% -    2%) 0.001
                        PKLookup      182.40      (9.8%)      181.24      (9.1%)   -0.6% ( -17% -   20%) 0.830
               ReqExclDense3TopN        2.80      (6.2%)       12.83     (17.6%)  357.6% ( 314% -  406%) 0.000
               ReqExclDense7TopN        1.32      (8.0%)        8.41     (21.9%)  536.6% ( 469% -  615%) 0.000
              ReqExclDense11TopN        0.78      (7.9%)        6.47     (28.8%)  728.9% ( 641% -  831%) 0.000
              ReqExclDense16TopN        0.56      (7.9%)        5.07     (34.5%)  807.4% ( 708% -  923%) 0.000

It looks like the situation hasn't worsened; perhaps the spare path can be deleted.

@kkewwei

kkewwei commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

@parkertimmins I also run the benchmark ReqExclNumericRangeBenchmark mentioned in #16568 about default MUST_NOT and Panama MUST_NOT with no spare path. The benchmark results show that the current implementation(only dense path) also delivers significant performance improvements for sparseMUST_NOT queries.

Provider Ratio Native (ops/s) Dense (ops/s) Speedup
Default 0.01 20.064 143.334 7.14×
Default 0.10 20.643 123.017 5.96×
Default 0.50 28.372 117.931 4.16×
Default 0.90 56.810 272.285 4.79×
Default 0.99 77.982 389.228 4.99×
Panama 0.01 18.511 254.776 13.76×
Panama 0.10 19.192 202.275 10.54×
Panama 0.50 27.917 246.289 8.82×
Panama 0.90 59.247 267.438 4.51×
Panama 0.99 78.942 263.249 3.33×

@kkewwei

kkewwei commented Sep 22, 2026

Copy link
Copy Markdown
Contributor Author

@romseygeek Sorry to bother you — would you please have a look in your spare time? It appears to bring a decent performance improvement to the must_not scenario. Thanks in advance!

@romseygeek romseygeek left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks great, thanks @kkewwei. I left a couple of comments.

Comment thread lucene/core/src/java/org/apache/lucene/search/ReqExclBulkScorer.java Outdated
Comment thread lucene/core/src/test/org/apache/lucene/search/TestReqExclBulkScorer.java Outdated
Comment thread lucene/CHANGES.txt Outdated
@kkewwei

kkewwei commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

@romseygeek Thank you very much for taking the time to review my code in detail. Your suggestions were very valuable, and I’ve addressed them all.

@romseygeek
romseygeek merged commit c5fa32c into apache:main Sep 24, 2026
12 checks passed
@romseygeek

Copy link
Copy Markdown
Contributor

Thanks @kkewwei!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants