开源本地大模型使用 LLM inference in C/C++
C++ · active 2025-03-19 → 2026-08-10 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 510 days from 2025-03-19 to 2026-08-10. Pushes: 3,423 total, peak 26 in a day. Pull requests: 5,009 total, peak 51 in a day. Issues: 3,758 total, peak 46 in a day. Comments: 19,618 total, peak 131 in a day. Stars: 12,806 total, peak 121 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| ggerganov | 3.6K | 1.1K | 460 | 1.2K |
| CISC | 3.3K | 489 | 320 | 1.6K |
| github-actions[bot] | 3K | 0 | 1.3K | 703 |
| ngxson | 2.3K | 288 | 235 | 1.1K |
| JohannesGaessler | 1.2K | 106 | 143 | 671 |
| 0cc4m | 1.2K | 316 | 97 | 506 |
| pwilkin | 932 | 58 | 44 | 534 |
| slaren | 850 | 107 | 78 | 389 |
| jeffbolznv | 782 | 18 | 120 | 503 |
| am17an | 772 | 89 | 85 | 412 |
| allozaur | 492 | 83 | 40 | 217 |
| ServeurpersoCom | 456 | 30 | 34 | 301 |
| taronaeo | 417 | 68 | 45 | 186 |
| gabe-l-hart | 391 | 4 | 20 | 251 |
| aldehir | 337 | 19 | 41 | 206 |
| danbev | 331 | 80 | 93 | 84 |
| compilade | 311 | 53 | 21 | 132 |
| max-krasnyansky | 273 | 91 | 29 | 92 |
| ericcurtin | 273 | 89 | 23 | 104 |
| IMbackK | 259 | 20 | 22 | 142 |
Recent activity
Latest issues, pull requests and releases
- Issue#24473github-actions[bot]2026-08-10 01:08Feature: Compact Conversation Action
- Issue#26759sriharshapy2026-08-08 07:43Misc. bug: (ggml-hexagon) FLASH_ATTN_EXT produces nondeterministic wrong results on the HMX path (v75/SM8650)
- Issue comment#25731F3zz1k2026-08-08 03:01Add TML Inkling architecture
- Issue comment#26689arthw2026-08-08 02:22SYCL: TILE for quantized KV decode
- Issue comment#24709github-actions[bot]2026-08-08 01:08Misc. bug: Windows <->Linux different behavior
- Releasegithub-actions[bot]2026-08-07 21:23b10326
- Issue#26741wolfpld2026-08-07 20:03Eval bug: deepseek4 produces garbled output when parallel processing is enabled and speculation is in use
- Issue comment#26735CISC2026-08-07 19:29Misc. bug: convert_hf_to_gguf.py crashes on gemma-4-31B-it.
- Issue#26738bokrosbalint2026-08-07 19:12Page fault at depth with -fa 0 and MoE expert offload on HIP
- Issue#26738bokrosbalint2026-08-07 19:12Page fault at depth with -fa 0 and MoE expert offload on HIP
- Pull request#26717allozaur2026-08-07 18:40
- Issue comment#26734ngxson2026-08-07 18:26tests : speed-up server test suite 3x
- Issue comment#25664khimaros2026-08-07 18:03Eval bug: vk::DeviceLostError within a few turns on DeepSeekv4-Flash (RADV_STRIXHALO)
- Issue comment#24176patrickzel2026-08-07 17:34server: improve user message detection and create checkpoints at every user message
- Issue comment#26563miltos222026-08-07 17:30Expert caching that greatly increases performance. Self contained, off by default, use -ehs N to activate
- Pull request#26734ggerganov2026-08-07 17:10
- Pull request#26733ServeurpersoCom2026-08-07 17:00
- Releasegithub-actions[bot]2026-08-07 16:58b10313
- Issue comment#26731grafail2026-08-07 16:54CUDA: fix thread/block count in quantized cpy kernel launches
- Issue comment#26294ggml-gh-bot[bot]2026-08-07 16:50CUDA: fix duplicate expert id compaction in mul_mat_id (#24591)
- Issue comment#26294am17an2026-08-07 16:50CUDA: fix duplicate expert id compaction in mul_mat_id (#24591)
- Issue comment#26563miltos222026-08-07 16:21Expert caching that greatly increases performance. Self contained, off by default, use -ehs N to activate
- Pull request#26727sfallah2026-08-07 14:41
- Issue#26694hailanlan05772026-08-07 03:19Eval bug: DeepSeek-V4-Flash degenerates into repetition and leaks special tokens in long agentic chats (Metal, b10289)
- Issue comment#26566fairydreaming2026-08-05 19:38ggml-webgpu: try to fix the new flash_attn error on nvidia machine
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 12,806 stars here means stars gained during the window, not the repo's star count.