LATX: optimize scalar pixel updates and VROUNDPS truncation - #442
Closed
luzeng87 wants to merge 11 commits into
Closed
LATX: optimize scalar pixel updates and VROUNDPS truncation#442luzeng87 wants to merge 11 commits into
luzeng87 wants to merge 11 commits into
Conversation
Track VEX.128 destinations whose architectural YMM high halves are known to be zero, and materialize those clears only when a 256-bit operation can observe them or before leaving the TB. This removes repeated LASX clear instructions while preserving signal, JIT, TU, and AOT-visible state. Add a standalone JIT, cold-AOT, and hot-AOT semantic test for the deferred state. Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Extend high-bit reduction to VEX scalar arithmetic, FMA, and move instructions, and emit scalar results directly when the preserved lanes are dead. Remove redundant vector temporaries and copies from compare, bitwise, shuffle, min/max, and packed FMA translations. Use LSX bit selection for packed min/max results. On Geekbench 7 Audio Encoder, these changes are part of the hot-AOT improvement from 557 to 596 before scalar FMA forwarding. Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Recognize scalar add, subtract, or multiply results that are immediately inserted into an XMM destination and then consumed by scalar FMADD or FMSUB. Keep the scalar value in its temporary until the fused operation and remove the intermediate VEXTRINS.W. Require exact operand and insertion matches so unrelated IR2 sequences remain unchanged. Add a standalone test for FMADD, FMSUB, NaN payloads, and preserved XMM upper lanes. The optimization raised Geekbench 7 Audio Encoder hot AOT from 596 to a 600 median. Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Signed-off-by: Lu Zeng <luzeng87@gmail.com>
Contributor
Author
|
This stacked branch includes the unsafe deferred YMM-state change from #437. Keeping its unrelated instruction changes in one branch also makes correctness attribution difficult. Closing it; only independently verified functional changes should remain open. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR is stacked on #437, #438, and #439; the scalar pixel-update and VROUNDPS changes are the final commits. The CPU feature-reporting change remains separate in #440.
Validation
Signed-off-by: Lu Zeng luzeng87@gmail.com