Skip to content

LATX: optimize scalar pixel updates and VROUNDPS truncation - #442

Closed
luzeng87 wants to merge 11 commits into
lat-opensource:masterfrom
luzeng87:gb403-hdr-pr
Closed

LATX: optimize scalar pixel updates and VROUNDPS truncation#442
luzeng87 wants to merge 11 commits into
lat-opensource:masterfrom
luzeng87:gb403-hdr-pr

Conversation

@luzeng87

@luzeng87 luzeng87 commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Summary

  • fuse a scalar pixel-update instruction sequence after exact register, memory, and constant checks
  • implement VROUNDPS SAE truncation with IEEE-754 bit operations, avoiding serialized FCSR reads and writes
  • preserve memory order, intermediate XMM state, exceptional values, and signed zero

This PR is stacked on #437, #438, and #439; the scalar pixel-update and VROUNDPS changes are the final commits. The CPU feature-reporting change remains separate in #440.

Validation

  • scalar update fixture checks outputs, general registers, memory order, and intermediate XMM state
  • VROUNDPS fixture covers XMM/YMM, signed zero, subnormal values, fractions, the 2^23 boundary, infinity, qNaN, and sNaN payloads
  • fixtures pass in JIT, cold AOT, and hot AOT modes
  • fast test suite: 24 passed, 0 failed

Signed-off-by: Lu Zeng luzeng87@gmail.com

Track VEX.128 destinations whose architectural YMM high halves are known to be zero, and materialize those clears only when a 256-bit operation can observe them or before leaving the TB. This removes repeated LASX clear instructions while preserving signal, JIT, TU, and AOT-visible state.

Add a standalone JIT, cold-AOT, and hot-AOT semantic test for the deferred state.

Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Extend high-bit reduction to VEX scalar arithmetic, FMA, and move instructions, and emit scalar results directly when the preserved lanes are dead.

Remove redundant vector temporaries and copies from compare, bitwise, shuffle, min/max, and packed FMA translations. Use LSX bit selection for packed min/max results.

On Geekbench 7 Audio Encoder, these changes are part of the hot-AOT improvement from 557 to 596 before scalar FMA forwarding.

Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Recognize scalar add, subtract, or multiply results that are immediately inserted into an XMM destination and then consumed by scalar FMADD or FMSUB.

Keep the scalar value in its temporary until the fused operation and remove the intermediate VEXTRINS.W. Require exact operand and insertion matches so unrelated IR2 sequences remain unchanged.

Add a standalone test for FMADD, FMSUB, NaN payloads, and preserved XMM upper lanes. The optimization raised Geekbench 7 Audio Encoder hot AOT from 596 to a 600 median.

Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
@luzeng87 luzeng87 changed the title LATX: optimize Geekbench HDR hot paths LATX: optimize scalar pixel updates and VROUNDPS truncation Sep 2, 2026
@luzeng87

luzeng87 commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

This stacked branch includes the unsafe deferred YMM-state change from #437. Keeping its unrelated instruction changes in one branch also makes correctness attribution difficult. Closing it; only independently verified functional changes should remain open.

@luzeng87 luzeng87 closed this Sep 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant