Skip to content

LATX: streamline AVX audio hot paths - #439

Open
luzeng87 wants to merge 8 commits into
lat-opensource:masterfrom
luzeng87:gb702-avx-audio
Open

LATX: streamline AVX audio hot paths#439
luzeng87 wants to merge 8 commits into
lat-opensource:masterfrom
luzeng87:gb702-avx-audio

Conversation

@luzeng87

@luzeng87 luzeng87 commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Summary

  • remove redundant AVX moves and temporary registers in compare, move, shuffle, and arithmetic translations
  • forward scalar arithmetic results into following FMA instructions when register and control-flow checks prove it safe
  • preserve VPSHUFD sources when destination and source registers differ
  • retain VCOMIS flag correctness in translated and LBT fallback paths

This PR is stacked on #437 and #438; the additional vector-lowering and scalar-FMA changes are the final commits.

Validation

  • instruction-level native x86, JIT, cold AOT, and hot AOT tests: passed
  • scalar FMA, VPSHUFD distinct-register, VCOMIS, and upper-half-state fixtures: passed
  • fast test suite: passed

Signed-off-by: Lu Zeng luzeng87@gmail.com

Track VEX.128 destinations whose architectural YMM high halves are known to be zero, and materialize those clears only when a 256-bit operation can observe them or before leaving the TB. This removes repeated LASX clear instructions while preserving signal, JIT, TU, and AOT-visible state.

Add a standalone JIT, cold-AOT, and hot-AOT semantic test for the deferred state.

Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Extend high-bit reduction to VEX scalar arithmetic, FMA, and move instructions, and emit scalar results directly when the preserved lanes are dead.

Remove redundant vector temporaries and copies from compare, bitwise, shuffle, min/max, and packed FMA translations. Use LSX bit selection for packed min/max results.

On Geekbench 7 Audio Encoder, these changes are part of the hot-AOT improvement from 557 to 596 before scalar FMA forwarding.

Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Recognize scalar add, subtract, or multiply results that are immediately inserted into an XMM destination and then consumed by scalar FMADD or FMSUB.

Keep the scalar value in its temporary until the fused operation and remove the intermediate VEXTRINS.W. Require exact operand and insertion matches so unrelated IR2 sequences remain unchanged.

Add a standalone test for FMADD, FMSUB, NaN payloads, and preserved XMM upper lanes. The optimization raised Geekbench 7 Audio Encoder hot AOT from 596 to a 600 median.

Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant