Skip to content

A high-throughput and memory-efficient inference and serving engine for LLMs (Windows build & kernels).

active 2025-03-152026-08-07 (UTC)

Partial coverage11,579 / 12,666 hourly files (91%) · 2 absent upstream · 1,083 failed, retryable2025-03-012026-08-10 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
398
Pushes
43
Pull requests
3
Issues
45
Stars
176
Forks
20

Activity over time

Daily event counts in the loaded window

Line chart, 511 days from 2025-03-15 to 2026-08-07. Pushes: 43 total, peak 8 in a day. Pull requests: 3 total, peak 2 in a day. Issues: 45 total, peak 3 in a day. Comments: 89 total, peak 17 in a day. Stars: 176 total, peak 5 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

Recent activity

Latest issues, pull requests and releases

  • Issue comment#59sergeng2026-07-02 05:26
    [Bug]: RuntimeError: Sparse Attention Indexer CUDA op requires DeepGEMM to be installed
  • Issue#62rogeroberg2026-05-10 07:05
    [Bug]: Machete generation fails on native Windows/Hopper because CMake uses POSIX PYTHONPATH separator
  • Issue comment#53tengcaicai2026-05-05 23:26
    [Bug]: 2 GPUs: Crash loading Qwen 3.5 0.8B with tensor parallel 2: access violation in all_reduce
  • Issue comment#53tengcaicai2026-05-04 04:10
    [Bug]: 2 GPUs: Crash loading Qwen 3.5 0.8B with tensor parallel 2: access violation in all_reduce
  • Issue comment#58wzgrx2026-05-03 06:15
    [Installation]: Latest v0.20.0 Windows build failed: metadata-generation-failed (pyproject.toml PEP621 error)
  • Issue comment#58SystemPanic2026-05-02 03:13
    [Installation]: Latest v0.20.0 Windows build failed: metadata-generation-failed (pyproject.toml PEP621 error)
  • Issue comment#58wzgrx2026-05-02 03:03
    [Installation]: Latest v0.20.0 Windows build failed: metadata-generation-failed (pyproject.toml PEP621 error)
  • Issue#59SystemPanic2026-05-02 02:16
    [Bug]: RuntimeError: Sparse Attention Indexer CUDA op requires DeepGEMM to be installed
  • Issue comment#59SystemPanic2026-05-02 02:16
    [Bug]: RuntimeError: Sparse Attention Indexer CUDA op requires DeepGEMM to be installed
  • Issue comment#56SystemPanic2026-04-30 22:32
    [Feature]: vllm v0.20.0 released
  • Issue comment#53lostmsu2026-04-25 12:29
    [Bug]: 2 GPUs: Crash loading Qwen 3.5 0.8B with tensor parallel 2: access violation in all_reduce
  • Issue#55lostmsu2026-04-25 12:25
    [Installation]: https://download.pytorch.org/whl/nightly/cu126 - nightly bit is gone
  • Issue comment#43wzgrx2026-04-04 05:35
    [Installation]: CMake compilation error with cu130 torch2.10
  • Issue#49Code4SAFrankie2026-03-26 21:08
    [Bug]: ModuleNotFoundError: No module named 'llguidance'
  • Issue comment#43venscn2026-03-26 04:27
    [Installation]: CMake compilation error with cu130 torch2.10
  • Issue#48stevenxuxin2026-03-26 01:41
    [Bug]: a ')' lost in line#367 of csrc/quantization/fp4/nvfp4_experts_quant.cu in vllm-for-windows branch
  • Issue comment#43nguyen9x2026-03-22 23:28
    [Installation]: CMake compilation error with cu130 torch2.10
  • Issue#43venscn2026-03-08 05:01
    [Installation]: CMake compilation error with cu130 torch2.10
  • Issue#40SystemPanic2026-03-06 17:59
    [Installation]: nvfp4_experts_quant.cu: compile error: "FLOAT" has already been declared in the current scope
  • Issue#38SystemPanic2026-03-06 16:44
    [Installation]: An error occurred when I tried to build from source in the cuda12.8, pytorch2.10.0, win10 environment
  • Issue#40sqweek2026-03-04 15:35
    [Installation]: nvfp4_experts_quant.cu: compile error: "FLOAT" has already been declared in the current scope
  • Issue comment#39wuwenthink2026-03-02 16:25
    [Installation]: 你好作者,我在win10环境下编译sm=120的blackwell的时候,使用cuda12.8相关的torch环境无法安装成功。作者如果可以安装成功的话,有条件可以打包一个50系显卡可以用的whl吗?
  • Issue comment#38User-Clb2026-03-02 07:02
    [Installation]: An error occurred when I tried to build from source in the cuda12.8, pytorch2.10.0, win10 environment
  • Issue comment#38SystemPanic2026-03-02 02:50
    [Installation]: An error occurred when I tried to build from source in the cuda12.8, pytorch2.10.0, win10 environment
  • Pull request#37mzyil2026-03-01 22:09

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 176 stars here means stars gained during the window, not the repo's star count.