【Fork】CUDA Templates and Python DSLs for High-Performance Linear Algebra
C++ · active 2023-08-15 → 2026-07-24 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 1075 days from 2023-08-15 to 2026-07-24. Pushes: 384 total, peak 9 in a day. Pull requests: 835 total, peak 11 in a day. Issues: 2,022 total, peak 32 in a day. Comments: 5,657 total, peak 37 in a day. Stars: 5,369 total, peak 247 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| hwu36 | 1.4K | 323 | 273 | 540 |
| github-actions[bot] | 1.2K | 0 | 5 | 1K |
| thakkarV | 704 | 1 | 9 | 615 |
| mnicely | 285 | 0 | 2 | 148 |
| ccecka | 238 | 0 | 0 | 208 |
| manishucsd | 203 | 0 | 12 | 129 |
| Junkai-Wu | 153 | 32 | 30 | 59 |
| jackkosaian | 141 | 2 | 1 | 132 |
| alexsamardzic | 140 | 0 | 8 | 100 |
| ziyuhuang123 | 106 | 0 | 0 | 36 |
| mhoemmen | 94 | 0 | 1 | 66 |
| HanGuo97 | 91 | 0 | 0 | 44 |
| d-k-b | 89 | 2 | 4 | 58 |
| fengxie | 88 | 4 | 1 | 67 |
| reed-lau | 65 | 0 | 22 | 36 |
| zekunf-nv | 62 | 10 | 13 | 38 |
| alihassanijr | 61 | 0 | 8 | 35 |
| IzanCatalan | 59 | 0 | 0 | 39 |
| jeromeku | 59 | 0 | 1 | 26 |
| brandon-yujie-sun | 58 | 0 | 0 | 38 |
Recent activity
Latest issues, pull requests and releases
- Issue#3406zkyue2026-07-24 02:37[CuTe DSL] setmaxregister_* can be silently dropped by ptxas (C7508) — nothing surfaces to Python; min_blocks_per_mp=1 prevents it
- Issue#3395yunweili32026-07-19 20:19[BUG] CUTLASS DSL
- Pull request#3366hwu362026-07-07 02:06
- Issue#3365tridao2026-07-03 04:29[BUG] `CUTE_DSL_KEEP=sass` is ignored in cutlass-dsl 4.6.0
- Issue#3365tridao2026-07-03 04:29[BUG] `CUTE_DSL_KEEP=sass` is ignored in cutlass-dsl 4.6.0
- Issue#329016bit-ykiko2026-07-02 15:00[BUG] `cute.gemm` rejects valid SM90 WGMMA m64n8k32 in SS mode — degenerate partition layout `(1,1):(0,0)`
- Issue#3358ankutalev2026-06-30 10:49[QST] Difference between Fp4 GEMM and Fp4 Ultra Gemm?
- Pull request#3345NVIDIA-JerryChen2026-06-23 09:00
- Pull request#3337Soojal18072026-06-20 18:09
- Pull request#3335dukallis2026-06-19 18:47
- Issue#3327github-actions[bot]2026-06-17 22:49[BUG]
- Issue#2570linear[bot]2026-06-16 17:51[QST][Cute-DSL] TMA Copy on a slice of an array
- Issue comment#3033github-actions[bot]2026-06-14 16:36[QST] [CuTeDSL] Alignment dropped by partition_S.
- Issue#3002github-actions[bot]2026-06-07 09:08[BUG] segmentation fault when calling `cute.print_tensor` on fp8 shared-memory tensor
- Pull request#3299ANIKET-SHIVAM2026-06-04 23:25
- Issue#2534linear[bot]2026-06-04 17:57[QST] confusion about various shapes encountered
- Issue comment#3177github-actions[bot]2026-05-31 05:36[BUG] Discrepancy between CuTe C++ and pycute
- Issue comment#3124hwu362026-05-30 23:12[CuTeDSL] Add SM103 grouped block-scaled GEMM kernel and tests
- Issue comment#3227idonati2026-05-30 19:12`nvidia-cutlass-dsl` 4.5.0: `nvvm.mma.block_scale` lowering produces PTX rejected by ptxas (sm_120/120f/121a)
- Issue#3170github-actions[bot]2026-05-30 05:22[CuTe DSL] libs-base and libs-cu13 4.4.x ship divergent _cutlass_ir.so for the same path; libs-base emits malformed _mma PTX for SM120 mxf4nvf4 mma
- Issue#3190zcbenz2026-05-29 22:14[QST][CuTe] `make_coord(int(m_coord), int(n_coord), _, int(l_coord))`
- Issue#3284infinitron2026-05-29 01:11[QST] Using 2xSM load in MMA kernels with input transform stage
- Issue#3284infinitron2026-05-29 01:11[QST] Using 2xSM load in MMA kernels with input transform stage
- Issue#3284infinitron2026-05-29 01:11[QST] Using 2xSM load in MMA kernels with input transform stage
- Issue comment#3276anakinxc2026-05-29 00:47[CuTeDSL] Fix the incorrect import of from_dlpack from docs
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 5,369 stars here means stars gained during the window, not the repo's star count.