Skip to content

[blackwell][draft] sm100: recompile 9 generic CUDA kernels (#225) - #233

Draft
Andrewxu313 wants to merge 1 commit into
mainfrom
tairan/blackwell-02-02-recompile-generic
Draft

[blackwell][draft] sm100: recompile 9 generic CUDA kernels (#225)#233
Andrewxu313 wants to merge 1 commit into
mainfrom
tairan/blackwell-02-02-recompile-generic

Conversation

@Andrewxu313

Copy link
Copy Markdown
Contributor

Closes #225. Part of #204 (Blackwell Phase 2).

Summary

Populate _sm100_extensions with 9 generic CUDA / SM80-mma.sync kernels (marlin_grouped_gemm, fp8_blockwise_ops, marlin_transform, 4x routing generics, dispatch_scatter_3d, fused_kv_norm_rope, fused_q_absorb, fused_q_split). fused_gate.cu EXCLUDED (WGMMA).

Spec

blackwell-kernel-port-v1.md § Sub-task 2

Placeholder commit for draft PR. Implementation tracked in:
batchgen-agent-metadata/batchgen_design/blackwell/blackwell-kernel-port-v1.md

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions github-actions Bot added the ci:run Trigger build + GPU regression on H20 label May 31, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci:run Trigger build + GPU regression on H20

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[blackwell] sm100: recompile 9 generic CUDA kernels

1 participant