Optimise builder execution loop - #9923
robert3005 wants to merge 3 commits into
5 benchmarks regressed
⚠️ Unknown Walltime execution environment detected
Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.
For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.
⚠️ Different runtime environments detected
Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.
⚡ 14 improved benchmarks
❌ 5 regressed benchmarks
✅ 2186 untouched benchmarks
⏩ 218 skipped benchmarks1
Warning
Please fix the performance issues or acknowledge them on CodSpeed.
Performance Changes
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | Simulation | decode_primitives[f32, (1000, 512)] |
42.5 µs | 63.1 µs | -32.55% |
| ❌ | Simulation | random_i16[0.95] |
81.2 µs | 98.2 µs | -17.31% |
| ❌ | Simulation | decompress[u64, (4000, 1024)] |
72.2 µs | 85 µs | -15.09% |
| ❌ | WallTime | dict_canonicalize_gt_u8_neon[1000000] |
487.9 µs | 560.1 µs | -12.89% |
| ❌ | WallTime | dict_canonicalize_gt_u8_neon[16000000] |
8.4 ms | 9.3 ms | -10.09% |
| ⚡ | Simulation | random_i8[0.5] |
95.7 µs | 71.1 µs | +34.62% |
| ⚡ | Simulation | chunked_opt_bool_into_canonical[(10, 100)] |
279 µs | 210.2 µs | +32.71% |
| ⚡ | Simulation | chunked_varbinview_into_canonical[(10, 100)] |
371.3 µs | 301.9 µs | +23% |
| ⚡ | Simulation | take[duplicates/repeated/primitive/nonnull/chunks=16/indices=1000] |
243.3 µs | 202.5 µs | +20.17% |
| ⚡ | Simulation | chunked_varbin_into_canonical[(10, 100)] |
434.6 µs | 364.9 µs | +19.13% |
| ⚡ | Simulation | chunked_opt_bool_into_canonical[(100, 50)] |
246.1 µs | 207 µs | +18.88% |
| ⚡ | WallTime | filtered_owned_i64_avx2[OneNullInEight] |
25.7 µs | 22 µs | +16.92% |
| ⚡ | Simulation | chunked_varbinview_opt_into_canonical[(10, 100)] |
558.8 µs | 484.6 µs | +15.31% |
| ⚡ | WallTime | dict_canonicalize_gt_u8_avx512[16000000] |
7.7 ms | 6.8 ms | +14.39% |
| ⚡ | Simulation | chunked_opt_bool_into_canonical[(1000, 10)] |
101.8 µs | 90.9 µs | +11.93% |
| ⚡ | Simulation | decompress[u64, (4000, 4)] |
140 µs | 125.4 µs | +11.71% |
| ⚡ | Simulation | chunked_varbinview_into_canonical[(100, 50)] |
363.5 µs | 326.9 µs | +11.19% |
| ⚡ | WallTime | compare_int_eq_neon |
5 µs | 4.6 µs | +10.42% |
| ⚡ | Simulation | chunked_varbinview_opt_into_canonical[(100, 50)] |
457.7 µs | 415.1 µs | +10.27% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing claude/vibrant-hawking-jl94oh (160a4f3) with develop (90bee5f)
Footnotes
-
218 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩