Skip to content

DFlash & TurboQuant in llama.cpp with up to 3x faster generation and 7.5x more KV cache in same VRAM

active 2026-05-092026-08-10 (UTC)

Partial coverage12,768 / 14,597 hourly files (87%) · 2 absent upstream · 1,823 failed, retryable2024-12-102026-08-11 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
172
Pushes
50
Pull requests
1
Issues
17
Stars
74
Forks
1

Activity over time

Daily event counts in the loaded window

Line chart, 94 days from 2026-05-09 to 2026-08-10. Pushes: 50 total, peak 8 in a day. Pull requests: 1 total, peak 1 in a day. Issues: 17 total, peak 2 in a day. Comments: 18 total, peak 3 in a day. Stars: 74 total, peak 12 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
Anbeeld6950012
Ezzz-dev2000
ethernidee2000
Bino51501000
naquad1000
Ponkipon1001
l3ateman1000
gdeyoung1001
yundddd1000
esmail-mkh1001
The-Dude-20201000
aamsellem1001
cryptopsy01000
dandenkijin1001
dani-nagy1001
github-actions[bot]1010

Recent activity

Latest issues, pull requests and releases

  • Issue#69yundddd2026-06-21 01:55
    Eval bug: crash during long agentic session
  • Issue#74ethernidee2026-06-16 19:07
    Misc. bug: extra 200+ MB VRAM consumption on start
  • Issue#60Anbeeld2026-06-14 23:54
    Eval bug: Sm tensor doesn't work
  • Issue comment#60Anbeeld2026-06-14 23:54
    Eval bug: Sm tensor doesn't work
  • Issue#61Anbeeld2026-06-14 21:24
    Eval bug: 0.3.2 compiled from latest source crashes
  • Issue#59Ezzz-dev2026-06-06 15:05
    Eval bug: Vulkan/HIP is way too slow in comparison to llama.cpp
  • Issue#57Ezzz-dev2026-06-05 15:42
    Eval bug: prompt is re-processing for no reason
  • Issue comment#45Anbeeld2026-06-04 22:58
    Eval bug: llama.cpp (AMD GPU): partial layers run on CPU when loading dflash model, cannot fully offload to GPU
  • Issue#52Anbeeld2026-06-04 21:28
    Eval bug: Multimodal model loading fails for Gemma 4 12B (unknown projector type: gemma4uv)
  • Issue#49Anbeeld2026-06-02 22:49
    Misc. bug: xmfp6 and turbo3 sharing same ID:42 this preventing from loading xmfp6 models
  • Issue comment#30Ponkipon2026-06-02 10:41
    Compile bug: Metal library fails to compile — block_turbo4_0 missing signs / rnorm fields referenced by quantize_turbo4_0
  • Issue comment#39Anbeeld2026-05-27 00:31
    Eval bug: low DFlash tps/AR for multi-GPU setups
  • Issue comment#41Anbeeld2026-05-26 23:22
    Eval bug: Segfault at slot initialization with CUDA on SM75 (Turing) — ngl > 0
  • Issue#41Bino51502026-05-26 22:59
    Misc. bug: Segfault at slot initialization with CUDA on SM75 (Turing) — ngl > 0
  • Issue#40cryptopsy02026-05-26 22:04
    profiling radeon 7900xt (20gb vram)
  • Issue comment#33gdeyoung2026-05-26 01:19
    DFlash segfault (exit 139) with Qwen3.6-27B on RTX PRO 4500 Blackwell (SM120)
  • Issue#37The-Dude-20202026-05-25 20:12
    Eval bug: When trying to WRITE a large content file I see the below error in the logs and the write timesout...
  • Issue comment#22Anbeeld2026-05-25 00:36
    No speed gain on 20gb vram with 27b q4
  • Issue#35ethernidee2026-05-24 23:56
    Eval bug: --ctx-size with --parallel 1 is divided by 2
  • Issue comment#32Anbeeld2026-05-24 20:34
    spec-dec: support device-aware recurrent GPU tape placement on ROCm multi-GPU
  • Issue comment#15dani-nagy2026-05-23 19:06
    Eval bug: Qwen3.6 35B A3B Q4 crash with "longer" context using turbo3 + dflash, cpu-offloading | Win10
  • Issue comment#22Anbeeld2026-05-23 01:06
    No speed gain on 20gb vram with 27b q4
  • Issue comment#28Anbeeld2026-05-23 00:31
    Eval bug: Multi GPU with draft errors: - the tokens for sequence 0 in the input batch have a starting position of Y = 1360 it is required that the sequence positions remain consecutive: Y = X + 1
  • Issue#27l3ateman2026-05-22 22:14
    invalid reduced-logits token -1
  • Issue#4Anbeeld2026-05-22 17:46
    Eval bug: llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'dflash'

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 74 stars here means stars gained during the window, not the repo's star count.