perf(evolution): ⚡ halve the follower-marking epoch stamp - #259
perf(evolution): ⚡ halve the follower-marking epoch stamp#259diagonal-hamiltonian wants to merge 2 commits into
Conversation
|
Docs preview: https://pr-259.monoprop-docs.pages.dev |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #259 +/- ##
=======================================
Coverage 97.70% 97.70%
=======================================
Files 14 14
Lines 742 742
Branches 98 98
=======================================
Hits 725 725
Misses 12 12
Partials 5 5
Flags with carried forward coverage won't be shown. Click here to find out more. |
f1f3067 to
b01e6fa
Compare
b01e6fa to
2be9966
Compare
`detail::MatchedEpochSet` carried a `uint32_t` stamp per term for a counter that never leaves the struct -- never serialised, never exchanged, only ever compared to `cur_`. This narrows it to `uint16_t`, keeps the wrap discipline explicit, and does NOT shrink the array anywhere. Two mechanisms: - the stamp width, with the wrap branch now live (once per 65535 gates) and the `std::fill` the thing that keeps it correct -- `epoch_` retains stale stamps for row indices reused after a truncation, so deleting the fill aliases marks; - `matched_scratch_bytes` joins `MPOperatorMemoryBreakdown` inside `total_bytes()`, with the stamp array being propagator-owned, so only `MonomialPropagator::operator_memory_usage()` can fill it in. The partitioned path sums per-partition breakdowns, so each picks up its own array. NO `shrink_to_fit()` ANYWHERE, not even an uncalled method. An earlier revision called `matched_scratch_.shrink_to_fit()` from `initialize_operator_caches_()` on the premise that the array "grows only when the term count does". That is false on Hubbard: the function runs after EVERY `build_graph()` and `propagate()`, so 29 Trotter steps reach it 29 times against a term count that grows at each step. Its A/B read +0.2821 GiB peak RSS (+3.12 B/term), 6/6, non-overlapping, at the 1x16 hubbard cell; removing it read -0.24 GiB, 6/6. And the shrink cannot free resident bytes at all -- `resize(n, 0)` never writes past `n`, so the capacity it releases was never faulted in. A retained-but-uncalled method plus a comment asking the next reader not to call it is a trap; the rejected experiment lives in the PR body instead. Re-cut twice and rebased onto main: the original was authored on a base carrying the profiling instrument, then briefly on the graph-memory PR whose key-set test is not on main. The mechanism never changed across any of it. It contains no profiling code: monoprop_PROF, profile:: and Profile.h appear zero times in the diff. MEASURED, campaign pr7v2off, both arms ENABLE_PROFILE=OFF, 12 cells x 10 interleaved reps in one allocation with order flipped per cell: kernel peak RSS (`/usr/bin/time -v`, node-sum) falls in 16 of 16 memory cells, 14 at unanimous 10/10, 15 of 16 clearing Holm across the memory family at largest adjusted p = 0.043. Hubbard `propagate` 0.97-0.99x, the hubbard `build_graph` rung 0.975-0.980x, pauli 0.989-0.999x -- hubbard beating pauli is what a per-term saving must do. Timing is a null: 1 of 24 tests resolved (`gradient[pauli]`, A 1x128, N=1, 0.9836x, 10/10, adjusted p = 0.0469, on an operation this diff cannot reach), the other 23 inside 0.964-1.014x. 2 B/term is a floor on the ARRAY's saving, not on the peak-RSS delta: peak RSS is a maximum over time and the stamp array need not be at its own maximum then. Hubbard `propagate` at N=1 implies 0.54 and 1.05 B/term of node-sum delta, below the array's own floor, which says something about WHEN the peak lands and nothing about the width. SPLIT: an earlier cut of this branch also released `init_op_map` in `get_operator()`. That is now a separate change so this PR carries one story, which means the measured binaries contain a release this branch does not. It was worth 1,189 B (8357 -> 7168 B in this campaign's own ledger) against cells of 9.5-40 GiB -- order 1e-7, far below the smallest resolved effect -- so the figures stand. Re-gated, not re-measured. An earlier instrumented pair (both arms ENABLE_PROFILE=ON) agrees on the `build_graph` rung and on pauli but read hubbard `propagate` at 0.96-0.97x -- it overstated exactly the number being advertised. The uninstrumented pair above is what certifies the shipping code, and is what is quoted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2be9966 to
8dbe2a5
Compare
|
🤖 AI text below 🤖 Measured at Full timing breakdown — all 24 tests, OFF/OFFCampaign
Holm across all 24 leaves one survivor, Worst-rank peak RSS, for the OOM questionNode-sum is the only cross-layout figure and is what the body tables; worst-rank is what an OOM is decided by. At layout A the two are equal (1 rank/node). The direction holds on worst-rank everywhere except two pauli cells, which reverse: |
|



🤖 AI text below 🤖
Halve the follower-marking epoch stamp (u32 → u16)
Narrows
detail::MatchedEpochSet's per-term stamp touint16_t— 2 B/term of operator row storage — with wrap handled explicitly and no shrink of the array anywhere.perf/epoch-stamp-noshrink-v28dbe2a5, one commit on412315e, 7 files +105/−12.cur_. The wrap branch is now live (once per 65535 gate applications instead of once per 2^32) and thestd::fillis what keeps it correct;matched_epoch_stamp_wrap_reached_by_gate_countcycles a whole period rather than assigning tocur_, which is the only way to prove the fill runs rather than that the branch is reachable.matched_scratch_bytesjoins the memory breakdown, insidetotal_bytes(). The stamp array is propagator-owned, soestimate_memory_usage()cannot see it and onlyMonomialPropagator::operator_memory_usage()can fill it in — the partitioned path sums per-partition breakdowns, so each picks up its own array.Memory: main
48cadcbvs this branch, bothENABLE_PROFILE=OFFKernel
/usr/bin/time -vpeak RSS, summed over the ranks on the node, produced outside the code under test. 16 cells × 10 interleaved reps in one allocation, order flipped per (rep, cell); paired per-rep ratios, median of ratios;agree= reps pointing the same way.1x128· N=1 ·fresh1x128· N=2 ·fresh8x16· N=1 ·fresh8x16· N=2 ·fresh1x128· N=1 ·fresh1x128· N=2 ·fresh8x16· N=1 ·fresh8x16· N=2 ·fresh1x128· N=1 ·fresh1x128· N=1 ·graph1x128· N=2 ·fresh1x128· N=2 ·graph8x16· N=1 ·fresh8x16· N=1 ·graph8x16· N=2 ·fresh8x16· N=2 ·graphPeak RSS falls in 16 of 16 cells, 14 at unanimous 10/10, and 15 of 16 clear Holm across the memory family — largest adjusted p = 0.043 (
pauli-B-N1 fresh, 9/10); the only failure ispauli-B-N2 fresh(8/10), the smallest effect present. That unanimity is node-sum only: on worst-rank RSS two cells reverse direction (pauli-A-N2 fresh1.0000,pauli-B-N2 fresh1.0027). Hubbard beating pauli is what a per-term saving must do — but 2 B/term is a floor on the array's saving, not on the peak-RSS delta, because peak RSS is a maximum over time and the stamp array need not be at its own maximum at that instant: hubbardpropagateat N=1 implies only 0.54 and 1.05 B/term of node-sum delta, below the array's own floor, which is a statement about when the peak lands and not about the width.Timing: a null
1 of 24 tests resolved; the other 23 span 0.964–1.014x. The one survivor is
gradient[pauli], layout A1x128, N=1: 11355.2 → 11172.4 ms, 0.9836x, 10/10, Holm-adjusted p = 0.0469 — exactly at the boundary, and on an operation that reaches only the allreduce and never touches a table this diff changes, so it reads as a null either way. Absolute medians at N=1 below; the N=2 half and the full 24-row breakdown will follow in a comment. Grid: layout A = 1 rank/node × 128 partitions, B = 8 × 16, N = 1 and 2 nodes; hubbard--hubbard-cutoff=10 --hubbard-lower-atol=1.25e-05(96,981,051 terms), pauli--pauli-cutoff=14 --pauli-lower-atol=5e-05(91,273,861 terms), plus abuild_graph[hubbard]rung at--hubbard-trotter-steps=2with--hubbard-lower-atol=1e-04, not the main grid's1.25e-05.1x128propagate[hubbard]8x16propagate[hubbard]1x128build_graph[hubbard]8x16build_graph[hubbard]1x128build_graph[pauli]1x128propagate[pauli]1x128energy[pauli]1x128gradient[pauli]8x16build_graph[pauli]8x16propagate[pauli]8x16energy[pauli]8x16gradient[pauli]Caveats and scope
total_bytes()is not comparable across this commit: a build withoutmatched_scratch_bytesreports a total lower by roughly the stamp array while holding the same or more resident memory, so subtract the field or re-measure the baseline.1250a27e…, port6f14ecb1…).init_op_mapinget_operator(); that mechanism has since been split out, soget_operator()here ismain's. The release was worth 1,189 B (8357 → 7168 B in this campaign's own ledger) against cells of 9.5–40 GiB — order 1e-7, far below the smallest resolved effect — so the figures above stand. Re-gated, not re-measured.Gates
ctest -L unit219/219 and-L serial218/218 green on theENABLE_PROFILE=OFFbuild of this tree (87ac7147…), the four Python MPI layouts 1×1, 1×16, 2×8 and 8×16 at 592 passed each, andclang-format --output-replacements-xml0 replacements on all 7 changed C++/binding files. Rebased ontoorigin/main412315e; the four commits since48cadcbare docs and build config only — zero files undercpp/,src/monoprop/, CMake or compiler flags — and their one build change (nanobind>=2.13.0,<3→==2.15.0) is the version both measured arms were already built with, so the numbers stand.