Parallelize partial output-tile stores across the store warp - #17
Open
morluto wants to merge 1 commit into
Open
Conversation
morluto
marked this pull request as ready for review
July 27, 2026 17:43
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TL;DR
Keep output-pipeline and TMA ownership on the elected STORE-warp lane while distributing the boundary-safe partial-tile copy across all 32 lanes.
Closes #14.
Motivation
Full 16-row output tiles use TMA, but the final partial tile uses a manual scalar copy to avoid writing beyond the current sequence.
The manual path currently runs only on the lane selected by
elect_one_sync(). ForD=128, one lane can therefore perform up to:This preserves sequence boundaries but serializes the entire tail copy.
Store-warp ownership
Pipeline wait and release remain elected-lane operations. Only the boundary-safe manual copy is distributed across the warp. Full-tile and final-state TMA stores are unchanged.
Changes
tail-storebenchmark mode for future measurements.Suggested review order
csrc/smxx/fwd_kernel2.cuh— STORE-warp ownership, synchronization, and lane-strided copy.tests/test_fwd_full.py— sequence-boundary and tail-length coverage.benchmarks/bench_fwd.py— isolated benchmark case.The synchronization points deliberately surround only shared-memory consumption: after the elected-lane wait and before the elected-lane release.
Test coverage
The new variable-length case places tail lengths 1 through 15 next to one another, exercising both the manual store and sequence-boundary indexing.
Checks
python -m py_compile benchmarks/bench_fwd.py tests/test_fwd_full.pygit diff --check origin/master...HEAD