Conversation
af53945 to
3815ccf
Compare
There was a problem hiding this comment.
Review in progress, I pressed ctrl-enter by mistake*
I don't fully understand the performance gains here.
I'm looking at the assembly of the following:
use heapless::HistoryBuffer;
#[unsafe(no_mangle)]
pub fn write_some_value(buf: &mut HistoryBuffer<u8, 31>, value: u8) {
buf.write(value);
}|
Ok, I think I see some weirdness that needs further exploration. In #598 (comment) I tested the benchmark but I think it's only because of the missing I do reproduce your benchmark results, but there seems to be variations depending on the |
3815ccf to
d63eec8
Compare
Under Miri, `RUSTC` is Miri itself, which doesn't do codegen, so the probe succeeds on any target and the pool tests then hit ARM inline assembly. Assisted-by: Claude Opus 5.5 (claude-opus-5-5)
The old one predates the upcoming MSRV bump to 1.95. Assisted-by: Claude Opus 5.5 (claude-opus-5-5)
Prep for using `core::hint::cold_path`. Assisted-by: Claude Opus 5.5 (claude-opus-5-5)
Covers the sliding-window search from issue rust-embedded#598 over several window sizes. Assisted-by: Claude Fable 5.1 (claude-fable-5-1) Assisted-by: Claude Opus 5.5 (claude-opus-5-5)
Like `new`. A zero-capacity buffer makes `recent_index` underflow. Originally implemented by wyf-777 in PR rust-embedded#687. Assisted-by: Claude Fable 5.1 (claude-fable-5-1) Assisted-by: Claude Opus 5.5 (claude-opus-5-5)
Without the hint, LLVM turns the reset of `write_at` into conditional moves, making each write depend on the previous one. This recovers most of the throughput lost since 0.9.0 (rust-embedded#598). Assisted-by: Claude Fable 5.1 (claude-fable-5-1) Assisted-by: Claude Opus 5.5 (claude-opus-5-5)
73fa65b to
761d12f
Compare
Done, along with other changes you asked for. Here are the results from my machine:
So very weird indeed. 🤷 |
Supersedes #687 and addresses #598.
Since the view-types refactor in 0.9.0, LLVM compiles the wrap-around reset of
write_atinHistoryBuf::writeto conditional moves, so each write depends on the previous one. Callingcore::hint::cold_pathin that branch brings back a real branch.cold_pathneeds Rust 1.95,hence the MSRV bump. That in turn needs a newer pinned nightly in CI, and with newer Miri the build
script's LL/SC probe succeeds on x86_64, so two small prep commits deal with those.
This also rejects zero capacity in
HistoryBuf::new_withat compile time, asnewalready does(from #687, thanks @wyf-777). The compile-fail tests for both constructors are
cfg(doctest)doctests rather than
cfailUI tests: trybuild only runscargo check, which never evaluatesthese post-monomorphization assertions, so UI tests for them compile successfully.
Measurements
Median
writethroughput in GB/s, 8 MiB input, x86_64, rustc 1.98. All three versions run thesame benchmark code as
benches/history_buf.rs, built out of tree so that 0.8.0 can be included.LTO means
lto = "fat"withcodegen-units = 1.With LTO, this PR is 2.4 to 2.8 times faster than main at every size. It matches or beats 0.8.0
except at sizes 9 and 31, where it is 6 to 11% slower. 0.8.0 itself swings between about 2 and
4 GB/s depending on the size. With the default profile, this PR is about 20% faster than main up
to size 32 and unchanged from 64 up.
For
write_and_search, this PR is within 6% of main at every size with LTO. With the defaultprofile, individual sizes move between 14% slower and 40% faster than main. Main shows the same
two plateaus at other sizes, and forcing 64-byte loop alignment reshuffles which sizes are fast,
so this looks like code layout rather than a change in the work done.
Credit to @BVollmerhaus for the report and bisect, and @sgued for the bounds-check analysis.
Generated by Claude Fable 5.1 and Claude Opus 5.5.
🤖 Generated with Claude Code