Skip to content

[blackwell][draft] sm100: QKV projection + RoPE Triton kernel (#228) - #236

Draft
Andrewxu313 wants to merge 2 commits into
mainfrom
tairan/blackwell-02-05-qkv-rope-triton
Draft

[blackwell][draft] sm100: QKV projection + RoPE Triton kernel (#228)#236
Andrewxu313 wants to merge 2 commits into
mainfrom
tairan/blackwell-02-05-qkv-rope-triton

Conversation

@Andrewxu313

Copy link
Copy Markdown
Contributor

Closes #228. Part of #204 (Blackwell Phase 2).

Summary

New triton/qkv_proj_rope.py: tl.dot BF16 matmul + fused split epilogue (Q/K/V separate buffers) + optional RoPE on Q+K. F.linear not acceptable (loses split+RoPE epilogue).

Spec

blackwell-kernel-port-v1.md § Sub-task 5

v-tairan Copilot user and others added 2 commits May 30, 2026 09:24
Placeholder commit for draft PR. Implementation tracked in:
batchgen-agent-metadata/batchgen_design/blackwell/blackwell-kernel-port-v1.md

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Fills the SM100 placeholder for the fused QKV WGMMA kernel. On Blackwell
the WGMMA .cu is not built, so cuda_qkv_wgmma() dispatches to a pure-Triton
GEMM+split kernel with a standard rotate_half RoPE epilogue matching the
SM90a CUDA convention (out[n1]=x1*c-x2*s, out[n2]=x2*c+x1*s).

- batchgen_kernels/triton/qkv_proj_rope.py: qkv_proj_split, apply_rope_qk,
  qkv_proj_rope
- qkv_wgmma.py: sm100 dispatch branch + availability via Triton
- triton/__init__.py: export new kernels

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions github-actions Bot added the ci:run Trigger build + GPU regression on H20 label May 31, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci:run Trigger build + GPU regression on H20

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[blackwell] sm100: QKV projection + RoPE Triton kernel

1 participant