sglang的后端FlashInfer: Kernel Library for LLM Serving
Python · active 2025-04-06 → 2026-08-10 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 492 days from 2025-04-06 to 2026-08-10. Pushes: 1,008 total, peak 14 in a day. Pull requests: 1,345 total, peak 21 in a day. Issues: 961 total, peak 171 in a day. Comments: 6,186 total, peak 62 in a day. Stars: 1,260 total, peak 33 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| yzh119 | 2.5K | 468 | 350 | 813 |
| gemini-code-assist[bot] | 2K | 0 | 0 | 1.3K |
| coderabbitai[bot] | 1.9K | 0 | 0 | 1K |
| flashinfer-bot | 706 | 25 | 110 | 568 |
| aleozlx | 627 | 110 | 73 | 191 |
| bkryu | 495 | 52 | 75 | 208 |
| yongwww | 304 | 45 | 76 | 102 |
| kahyunnam | 174 | 25 | 36 | 49 |
| saltyminty | 147 | 43 | 8 | 35 |
| nvmbreughe | 140 | 6 | 10 | 71 |
| Edenzzzz | 125 | 0 | 11 | 75 |
| github-actions[bot] | 114 | 0 | 3 | 0 |
| yyihuang | 110 | 14 | 45 | 26 |
| cyx-6 | 109 | 5 | 26 | 24 |
| jimmyzho | 97 | 10 | 24 | 40 |
| sricketts | 94 | 1 | 9 | 43 |
| nv-yunzheq | 93 | 11 | 22 | 37 |
| weireweire | 86 | 0 | 8 | 50 |
| IwakuraRein | 76 | 15 | 9 | 39 |
| nvpohanh | 68 | 0 | 4 | 55 |
Recent activity
Latest issues, pull requests and releases
- Issue#4396aleozlx2026-08-07 19:56[Bug][v0.6.17][gb300]tests/gemm/test_groupwise_scaled_gemm_fp8.py:201: Mismatched elements: 124 / 8192 (1.5%)
- Pull request#4409aleozlx2026-08-07 19:42
- Issue comment#4406aleozlx2026-08-07 19:36fix(monomoe): restore CUDA 12.0+ compatibility in tma_load_2d
- Pull request#4406aleozlx2026-08-07 18:04
- Pull request#4405flashinfer-bot2026-08-07 17:38
- Issue comment#4331bkryu2026-08-07 17:27perf(moe): persist b12x MoE CuTe-DSL kernels to the disk cache
- Issue comment#3849flashinfer-bot2026-08-07 16:56feat: Grouped-token MLA support for the TRTLLM-Gen FMHA backend.
- Issue comment#4081flashinfer-bot2026-08-07 16:51feat(gdn): u/d cache spec-decode kernels for replayssm
- Pull request#4352flashinfer-bot2026-08-04 23:21
- Issue comment#4344bkryu2026-08-04 17:50Revert "test: Add sharding support to scripts/task_run_unit_tests.sh"
- Issue comment#4027flashinfer-bot2026-08-04 00:35MoE monokernel Bug fix, barrrier remove and kernel rewrite.
- Issue comment#3454murphymatt2026-08-01 01:23[Feature request] TRT-LLM-Gen fused MoE cubins for GeGLU activation (mixed-precision formats)
- Issue comment#4261gemini-code-assist[bot]2026-07-30 03:35deep_gemm: arch-based split for DEEPGEMM_RUBIN cubins
- Issue comment#4226gemini-code-assist[bot]2026-07-29 00:49Make CuTe-DSL arch guard env-aware and gate norm's DSL dispatch on it
- Issue comment#4122coderabbitai[bot]2026-07-24 02:15Adds SM107 support
- Issue comment#4111gemini-code-assist[bot]2026-07-23 18:34fix(tests): gate grouped_mm cudnn MOE tests on the runtime min version constant
- Issue comment#4017gnovack2026-07-22 20:42skip invalid experts in dispatch
- Pull request#3975ishovkun2026-07-21 18:39
- Issue comment#3697flashinfer-bot2026-07-18 00:34Port the TensorRT-LLM one-sided A2A optimizations to Flashinfer
- Pull request#4013yichengj02026-07-17 01:41
- Issue comment#3960kahyunnam2026-07-16 18:43fix(gdn): compile SM12x CuteDSL kernels as sm_121a on DGX Spark
- Issue comment#3975flashinfer-bot2026-07-16 02:49mamba checkpointing SSU: two-kernel split + ring-buffer cache for checkpointing SSU
- Issue comment#3987jiahanc2026-07-16 02:24feat(moe): enable BiasType::Mn (LoRA delta) for nvfp4/mxfp4 MoE
- Issue comment#3874flashinfer-bot2026-07-15 17:15feat(jit): JitSpec ABC + disk cache for JIT-compiled CuTe-DSL kernels
- Issue comment#3971nvpohanh2026-07-15 02:48SM103 serving hang caused by TRTLLM_GEN_BMM artifact regeneration in #3708 (batched_gemm-dd6d23e-721ae60, 0.6.14): FP4 batched-GEMM clusters stuck in mbarrier phase wait
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 1,260 stars here means stars gained during the window, not the repo's star count.