Skip to content

开源本地大模型使用 LLM inference in C/C++

C++ · active 2025-03-192026-08-10 (UTC)

Partial coverage11,359 / 12,221 hourly files (93%) · 2 absent upstream · 860 failed, retryable2025-03-192026-08-10 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
55.7K
Pushes
3.4K
Pull requests
5K
Issues
3.8K
Stars
12.8K
Forks
2.3K

Activity over time

Daily event counts in the loaded window

Line chart, 510 days from 2025-03-19 to 2026-08-10. Pushes: 3,423 total, peak 26 in a day. Pull requests: 5,009 total, peak 51 in a day. Issues: 3,758 total, peak 46 in a day. Comments: 19,618 total, peak 131 in a day. Stars: 12,806 total, peak 121 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
ggerganov3.6K1.1K4601.2K
CISC3.3K4893201.6K
github-actions[bot]3K01.3K703
ngxson2.3K2882351.1K
JohannesGaessler1.2K106143671
0cc4m1.2K31697506
pwilkin9325844534
slaren85010778389
jeffbolznv78218120503
am17an7728985412
allozaur4928340217
ServeurpersoCom4563034301
taronaeo4176845186
gabe-l-hart391420251
aldehir3371941206
danbev331809384
compilade3115321132
max-krasnyansky273912992
ericcurtin2738923104
IMbackK2592022142

Recent activity

Latest issues, pull requests and releases

  • Issue#24473github-actions[bot]2026-08-10 01:08
    Feature: Compact Conversation Action
  • Issue#26759sriharshapy2026-08-08 07:43
    Misc. bug: (ggml-hexagon) FLASH_ATTN_EXT produces nondeterministic wrong results on the HMX path (v75/SM8650)
  • Issue comment#25731F3zz1k2026-08-08 03:01
    Add TML Inkling architecture
  • Issue comment#26689arthw2026-08-08 02:22
    SYCL: TILE for quantized KV decode
  • Issue comment#24709github-actions[bot]2026-08-08 01:08
    Misc. bug: Windows <->Linux different behavior
  • Releasegithub-actions[bot]2026-08-07 21:23
    b10326
  • Issue#26741wolfpld2026-08-07 20:03
    Eval bug: deepseek4 produces garbled output when parallel processing is enabled and speculation is in use
  • Issue comment#26735CISC2026-08-07 19:29
    Misc. bug: convert_hf_to_gguf.py crashes on gemma-4-31B-it.
  • Issue#26738bokrosbalint2026-08-07 19:12
    Page fault at depth with -fa 0 and MoE expert offload on HIP
  • Issue#26738bokrosbalint2026-08-07 19:12
    Page fault at depth with -fa 0 and MoE expert offload on HIP
  • Pull request#26717allozaur2026-08-07 18:40
  • Issue comment#26734ngxson2026-08-07 18:26
    tests : speed-up server test suite 3x
  • Issue comment#25664khimaros2026-08-07 18:03
    Eval bug: vk::DeviceLostError within a few turns on DeepSeekv4-Flash (RADV_STRIXHALO)
  • Issue comment#24176patrickzel2026-08-07 17:34
    server: improve user message detection and create checkpoints at every user message
  • Issue comment#26563miltos222026-08-07 17:30
    Expert caching that greatly increases performance. Self contained, off by default, use -ehs N to activate
  • Pull request#26734ggerganov2026-08-07 17:10
  • Pull request#26733ServeurpersoCom2026-08-07 17:00
  • Releasegithub-actions[bot]2026-08-07 16:58
    b10313
  • Issue comment#26731grafail2026-08-07 16:54
    CUDA: fix thread/block count in quantized cpy kernel launches
  • Issue comment#26294ggml-gh-bot[bot]2026-08-07 16:50
    CUDA: fix duplicate expert id compaction in mul_mat_id (#24591)
  • Issue comment#26294am17an2026-08-07 16:50
    CUDA: fix duplicate expert id compaction in mul_mat_id (#24591)
  • Issue comment#26563miltos222026-08-07 16:21
    Expert caching that greatly increases performance. Self contained, off by default, use -ehs N to activate
  • Pull request#26727sfallah2026-08-07 14:41
  • Issue#26694hailanlan05772026-08-07 03:19
    Eval bug: DeepSeek-V4-Flash degenerates into repetition and leaks special tokens in long agentic chats (Metal, b10289)
  • Issue comment#26566fairydreaming2026-08-05 19:38
    ggml-webgpu: try to fix the new flash_attn error on nvidia machine

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 12,806 stars here means stars gained during the window, not the repo's star count.