Repository navigation
perf(sparse): trim the stages around the CSR solver hand-off - #992
Conversation
Merging this PR will improve performance by 44.5%
|
| Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|
| ⚡ | test_to_lp[nodal_balance_sparse-severity=0] |
3.9 MB | 2.5 MB | +54.8% |
| ⚡ | test_to_lp[nodal_balance_sparse-severity=100] |
3.9 MB | 2.5 MB | +54.71% |
| ⚡ | test_to_lp[nodal_balance_sparse-severity=50] |
3.9 MB | 2.5 MB | +54.71% |
| ⚡ | test_to_lp[knapsack-n=10000] |
2.1 MB | 1.4 MB | +52.92% |
| ⚡ | test_to_lp[storage-n=250] |
38.5 MB | 25.8 MB | +49.11% |
| ⚡ | test_to_lp[merge_balance-severity=0] |
2.7 MB | 1.8 MB | +48.43% |
| ⚡ | test_to_lp[rolling-severity=0] |
2 MB | 1.4 MB | +48.07% |
| ⚡ | test_to_lp[nodal_balance-severity=50] |
3.8 MB | 2.6 MB | +45.33% |
| ⚡ | test_to_lp[nodal_balance-severity=0] |
3.8 MB | 2.6 MB | +45.27% |
| ⚡ | test_to_lp[storage-n=10] |
1.6 MB | 1.1 MB | +44.9% |
| ⚡ | test_to_lp[piecewise-n=1000] |
1.9 MB | 1.4 MB | +37.3% |
| ⚡ | test_to_lp[basic-n=250] |
29.9 MB | 21.8 MB | +37.15% |
| ⚡ | test_to_lp[expression_arithmetic-n=250] |
46.6 MB | 34.9 MB | +33.6% |
| ⚡ | test_to_lp[sos-n=1000] |
1.4 MB | 1.1 MB | +21.12% |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing perf/sparse-pipeline (908a59b) with master (7e007ff)
Footnotes
-
181 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
Build cost — v1 vs legacyv1 build peak & time relative to legacy, on this commit — not a comparison against master (that is CodSpeed).
Full table (time + peak, mean)📊 Interactive plots + CSV: download the semantics-report-v1-vs-legacy artifact from this run. Report-only · not a gate · refreshed on every push · obsolete once legacy is dropped. |
|
AI-assisted benchmark proposal from our PyPSA/Linopy/HiGHS work. We use PyPSA 1.3.0 and Linopy 0.9.1 with the direct HiGHS API. On six reconstructed 168-hour power-system LPs, median PyPSA model construction was 3.249 s, extension construction 0.197 s, and to_highspy 1.077 s. These are release-path timings, not measurements of this PR or sparse=True. We would like to contribute a reproducible peak-memory benchmark covering model construction, constraint freezing, matrix assembly, direct solver hand-off, and solution/dual read-back. The benchmark would report process peak RSS alongside wall time and model rows/columns/nonzeros, and separate Python/NumPy allocations from native solver memory. We would compare identical LP inputs and independently check primal/dual residuals. Would extending an existing benchmark here be useful, and which benchmark location/model size would you prefer? We have not measured a RAM saving from this PR yet. Our application also has full-network-per-worker loading overhead, which we are addressing separately so that it does not get attributed to Linopy. |
|
Note The benchmark results and chart below were generated with AI assistance. Following up on our earlier comment: we have now measured head Three fresh-process, interleaved construction repeats used the same 168-hour BirdFlow network, with float64 coefficients throughout. Medians:
The head reduced construction-process peak RSS by 4.5% against the sparse base, or 18.4% against 0.9.1. Matrix dimensions were 2583288 × 1658758, with 5035943 nonzeros. Independent dumps of A, c, b, lb, ub, senses and variable/constraint labels were exactly equal across all three variants. Dumps ran after measured stages and peak capture. We also ran one sequential 1008-hour A/B between 0.9.1 and the PR head: six 168-hour chunks, rebuilding each model, from the same sealed input network and BirdFlow commit
The single solve A/B is indicative only: its wall-time difference cannot be assigned exclusively to this PR. The base arm was construction-only, and this was representative-window validation, not an annual run. Both solve logs contained intermediate native factorization/imprecision warnings. We did not independently measure primal/dual residuals. For reproducibility, our construction/comparison harness and representative solve runner are published. The construction memory reduction is useful for our workload, even though absolute build-time savings are small relative to solving.
|
…es via getModel MOSEK bound keys and names are built vectorised. The polars LP writer skips identity scaling, avoids re-sorting frozen rows and builds fewer string branches; output is byte-identical. HiGHS names, when requested, are attached through getModel/passModel so the Hessian survives.
…ntraction chunks Merge concatenates all CSR operands once instead of folding pairwise. Expression.solution maps labels directly and evaluates CSR rows without densifying or rebuilding the constraint matrix. Contraction chunks are sized by nonzeros instead of a fixed row count.
… read-back, one freeze gather Constraint blocks are gathered into preallocated CSR buffers without vstack or eliminate_zeros; scaling is skipped when identity and label lookups use the label dtype. Frozen constraints receive duals on active rows directly. Mask and non-empty filtering share one gather and sense conversion is vectorised.
2e33cec to
908a59b
Compare
) * perf(csr): one-pass n-ary added, row gather for permuting reindex (#1010) Make CSRLinearExpression.added n-ary in one COO pass (as in #992), gather rows directly when reindexed only permutes cells, and index aggregated rows in the store's index dtype. * test: reuse densified helper for mypy; add release note (#1010) * test(csr): drop reindex test superseded by test_reindex_matches_dense
Note
The following content was generated by AI.
Changes proposed in this Pull Request
Follow-up to #985/#987/#988/#990, rebased onto #1012 and #1015. An assessment of the sparse (
Model(sparse=True)) data flow from expression construction to the solver found the core hand-off close to optimal (one CSR build per operation, one label remap, a no-copy pass to the direct APIs) but avoidable cost in the stages around it. This PR removes that cost.model.matricesand the LP file stay bit-identical to master; a CSR-backedsolutionmatches the dense path up to floating-point summation order. Each change was implemented, independently reviewed and revised in its own round.Solver hand-off
test_optimization.py -k mosekon a licensed machine.getModel()/passModel()instead ofgetLp(). Re-passing the LP-only object drops the Hessian, so a QP solved with names on got a wrong objective. A regression test is added.LP writer
Expressions
linopy.mergeof CSR-backed expressions concatenates all operands once instead of folding pairwise, so an N-way merge is linear in total nonzeros rather than quadratic in operand count. The concatenation holds all operands' terms before duplicates are summed, so merging many operands over the same variables has a higher peak memory (see benchmarks).LinearExpression.solutionon a CSR-backed expression is evaluated on its sparse backing without densifying. Like the dense path since perf(matrices): trim the Model.matrices build and drop redundant builds #1012, it raises when the expression references variables removed from the model.Matrix assembly and read-back
model.matricesgathers constraint blocks into preallocated CSR buffers in one pass, withoutvstackor a per-block scaling copy; scaling is skipped when identity. Label lookups use the model's label dtype, somatrices.clabelsandmatrices.indicator_binvarareint32by default (behaviour change, noted in the release notes).sanitize_zerosare vectorised without large temporaries.The chunk sizing of the sparse
@contraction was also tried but is redundant after #990 and was reverted.Benchmarks against current master (Model(sparse=True), 8.16M rows, 9.86M variables, ~19.7M nonzeros, highspy 1.15, numpy 2.4)
Best of 3 runs; peak is the tracemalloc peak of one extra run.
model.matricesassign_result)sanitize_zerosper solvemergeof 8 CSR operands on the same variables, 6.8M nnz eachexpr.solution, CSR expressionexpr.solution, dense expressionWith frozen constraints on a dense model the LP write (2.96 to 2.34 s),
model.matrices,assign_resultandsanitize_zerosgains are the same.model.matricesand the LP file hash identically to master in both modes.Verification
model.matricescompared value-for-value against master over 25 models (masks, explicit zeros, NaN rhs, scaling, removed containers, integer/binary, indicator, empty).solutioncompared against the dense path with absent cells, NaN solution values and auxiliary coordinates.Checklist
AGENTS.md).doc.doc/release_notes.rstof the upcoming release is included.