-
Notifications
You must be signed in to change notification settings - Fork 610
Pull requests: AI-Hypercomputer/maxtext
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
enable return final hidden logits in maxtext
#5299
opened Sep 20, 2026 by
copybara-service
Bot
Loading…
[Qwen3.5] Mask left-padding tokens in Qwen3NextGatedDeltaNet and guar…
#5294
opened Sep 19, 2026 by
jiantanfig
Collaborator
Loading…
8 tasks done
Support keeping MoE router projection unquantized for stability
gemini-review
#5293
opened Sep 19, 2026 by
shuningjin
Collaborator
Loading…
4 tasks done
Add m3 package scaffolding with architecture rules and layout test
#5291
opened Sep 19, 2026 by
bvandermoon
Collaborator
Loading…
4 tasks done
Add free_kv_cache_during_weight_sync to the RL config and forward it to tunix
#5289
opened Sep 18, 2026 by
gutianyu-google
Collaborator
Loading…
3 tasks done
fix(attention test): give the DSV4 compressed mask test the cp axis rule
#5287
opened Sep 18, 2026 by
gulsumgudukbay
Collaborator
Loading…
4 tasks done
fix(router replay test): size the data batch by device count
#5286
opened Sep 18, 2026 by
gulsumgudukbay
Collaborator
Loading…
4 tasks done
Fix issues with trainable_parameters_mask in MaxTextTrainingEngine
gemini-review
#5285
opened Sep 18, 2026 by
niting
Collaborator
Loading…
4 tasks done
Overhaul DeepSeek-V4-Flash 3-stage E2E test suite
#5283
opened Sep 18, 2026 by
AusarYao28
Collaborator
•
6/6
•
Draft
4 tasks done
Evaluation & Autoregressive Decode Infrastructure Fixes
#5282
opened Sep 18, 2026 by
AusarYao28
Collaborator
•
5/6
•
Draft
4 tasks done
NNX from_pretrained Deep-Merge & tid2eid Routing Table Sideloading
#5281
opened Sep 18, 2026 by
AusarYao28
Collaborator
•
3/6
•
Draft
4 tasks done
Disallow use_indexer config combinations that zero the entire training objective
#5280
opened Sep 18, 2026 by
AusarYao28
Collaborator
•
4/6
•
Draft
4 tasks done
Apply YaRN RoPE rescaling to DeepSeek-V4 compressed attention
#5279
opened Sep 18, 2026 by
AusarYao28
Collaborator
•
2/6
•
Draft
4 tasks done
Prevent XLA from fusing the MoE backward pass into the optimizer update
#5278
opened Sep 18, 2026 by
AusarYao28
Collaborator
•
1/6
•
Draft
4 tasks done
Cosmos 2.1: Unified 3D mRoPE & Timestep Embedding Primitives
#5277
opened Sep 18, 2026 by
hengtaoguo
Collaborator
•
Draft
4 tasks
Resolve kv_tp_size and moe_mlp_tp_size with fallback to rollout TP.
#5276
opened Sep 18, 2026 by
niting
Collaborator
Loading…
4 tasks done
Add self-contained Lineage DeepSeek-V3 fast path (
models/deepseek_lineage/ and models/lineage_adapter.py) to MaxText for TPU v7x (deepseek3-671b-lineage), including MLA splash attention, router, batch-split MoE expert layers, and unit tests.
#5275
opened Sep 18, 2026 by
copybara-service
Bot
Loading…
[Stacked PR 8/8] Optimize GDN backward kernel with native BF16 MXU, algebraic GEMM reductions, and multi-group grid
#5274
opened Sep 18, 2026 by
Rohan-Bierneni
Collaborator
•
8/8
Loading…
4 tasks done
[Stacked PR 7/8] Optimize GDN forward pass with fast transpose broadcasts and TPU hardware dot safeguards
#5273
opened Sep 18, 2026 by
Rohan-Bierneni
Collaborator
•
7/8
Loading…
4 tasks done
Fix pipeline_utils mutating axes_to_remove and add unit tests
#5272
opened Sep 18, 2026 by
harshaa765
Loading…
4 tasks done
[Fix] MoE DP reduce-scatter and logical axis sharding in vLLM serving path
#5271
opened Sep 18, 2026 by
sierraisland
Contributor
Loading…
4 tasks done
Add SFT training script using MaxTextTrainingEngine to experimental
#5269
opened Sep 17, 2026 by
SurbhiJainUSC
Collaborator
•
Draft
4 tasks done
Derive fallback segment positions from an exclusive real-token count
#5266
opened Sep 17, 2026 by
A9isha
Collaborator
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.