Skip to content

sglang的后端FlashInfer: Kernel Library for LLM Serving

Python · active 2025-04-062026-08-10 (UTC)

Partial coverage11,164 / 11,801 hourly files (95%) · 2 absent upstream · 634 failed, retryable2025-04-062026-08-10 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
14.9K
Pushes
1K
Pull requests
1.3K
Issues
961
Stars
1.3K
Forks
235

Activity over time

Daily event counts in the loaded window

Line chart, 492 days from 2025-04-06 to 2026-08-10. Pushes: 1,008 total, peak 14 in a day. Pull requests: 1,345 total, peak 21 in a day. Issues: 961 total, peak 171 in a day. Comments: 6,186 total, peak 62 in a day. Stars: 1,260 total, peak 33 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
yzh1192.5K468350813
gemini-code-assist[bot]2K001.3K
coderabbitai[bot]1.9K001K
flashinfer-bot70625110568
aleozlx62711073191
bkryu4955275208
yongwww3044576102
kahyunnam174253649
saltyminty14743835
nvmbreughe14061071
Edenzzzz12501175
github-actions[bot]114030
yyihuang110144526
cyx-610952624
jimmyzho97102440
sricketts941943
nv-yunzheq93112237
weireweire860850
IwakuraRein7615939
nvpohanh680455

Recent activity

Latest issues, pull requests and releases

  • Issue#4396aleozlx2026-08-07 19:56
    [Bug][v0.6.17][gb300]tests/gemm/test_groupwise_scaled_gemm_fp8.py:201: Mismatched elements: 124 / 8192 (1.5%)
  • Pull request#4409aleozlx2026-08-07 19:42
  • Issue comment#4406aleozlx2026-08-07 19:36
    fix(monomoe): restore CUDA 12.0+ compatibility in tma_load_2d
  • Pull request#4406aleozlx2026-08-07 18:04
  • Pull request#4405flashinfer-bot2026-08-07 17:38
  • Issue comment#4331bkryu2026-08-07 17:27
    perf(moe): persist b12x MoE CuTe-DSL kernels to the disk cache
  • Issue comment#3849flashinfer-bot2026-08-07 16:56
    feat: Grouped-token MLA support for the TRTLLM-Gen FMHA backend.
  • Issue comment#4081flashinfer-bot2026-08-07 16:51
    feat(gdn): u/d cache spec-decode kernels for replayssm
  • Pull request#4352flashinfer-bot2026-08-04 23:21
  • Issue comment#4344bkryu2026-08-04 17:50
    Revert "test: Add sharding support to scripts/task_run_unit_tests.sh"
  • Issue comment#4027flashinfer-bot2026-08-04 00:35
    MoE monokernel Bug fix, barrrier remove and kernel rewrite.
  • Issue comment#3454murphymatt2026-08-01 01:23
    [Feature request] TRT-LLM-Gen fused MoE cubins for GeGLU activation (mixed-precision formats)
  • Issue comment#4261gemini-code-assist[bot]2026-07-30 03:35
    deep_gemm: arch-based split for DEEPGEMM_RUBIN cubins
  • Issue comment#4226gemini-code-assist[bot]2026-07-29 00:49
    Make CuTe-DSL arch guard env-aware and gate norm's DSL dispatch on it
  • Issue comment#4122coderabbitai[bot]2026-07-24 02:15
    Adds SM107 support
  • Issue comment#4111gemini-code-assist[bot]2026-07-23 18:34
    fix(tests): gate grouped_mm cudnn MOE tests on the runtime min version constant
  • Issue comment#4017gnovack2026-07-22 20:42
    skip invalid experts in dispatch
  • Pull request#3975ishovkun2026-07-21 18:39
  • Issue comment#3697flashinfer-bot2026-07-18 00:34
    Port the TensorRT-LLM one-sided A2A optimizations to Flashinfer
  • Pull request#4013yichengj02026-07-17 01:41
  • Issue comment#3960kahyunnam2026-07-16 18:43
    fix(gdn): compile SM12x CuteDSL kernels as sm_121a on DGX Spark
  • Issue comment#3975flashinfer-bot2026-07-16 02:49
    mamba checkpointing SSU: two-kernel split + ring-buffer cache for checkpointing SSU
  • Issue comment#3987jiahanc2026-07-16 02:24
    feat(moe): enable BiasType::Mn (LoRA delta) for nvfp4/mxfp4 MoE
  • Issue comment#3874flashinfer-bot2026-07-15 17:15
    feat(jit): JitSpec ABC + disk cache for JIT-compiled CuTe-DSL kernels
  • Issue comment#3971nvpohanh2026-07-15 02:48
    SM103 serving hang caused by TRTLLM_GEN_BMM artifact regeneration in #3708 (batched_gemm-dd6d23e-721ae60, 0.6.14): FP4 batched-GEMM clusters stuck in mbarrier phase wait

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 1,260 stars here means stars gained during the window, not the repo's star count.