DFlash & TurboQuant in llama.cpp with up to 3x faster generation and 7.5x more KV cache in same VRAM
active 2026-05-09 → 2026-08-10 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 94 days from 2026-05-09 to 2026-08-10. Pushes: 50 total, peak 8 in a day. Pull requests: 1 total, peak 1 in a day. Issues: 17 total, peak 2 in a day. Comments: 18 total, peak 3 in a day. Stars: 74 total, peak 12 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| Anbeeld | 69 | 50 | 0 | 12 |
| Ezzz-dev | 2 | 0 | 0 | 0 |
| ethernidee | 2 | 0 | 0 | 0 |
| Bino5150 | 1 | 0 | 0 | 0 |
| naquad | 1 | 0 | 0 | 0 |
| Ponkipon | 1 | 0 | 0 | 1 |
| l3ateman | 1 | 0 | 0 | 0 |
| gdeyoung | 1 | 0 | 0 | 1 |
| yundddd | 1 | 0 | 0 | 0 |
| esmail-mkh | 1 | 0 | 0 | 1 |
| The-Dude-2020 | 1 | 0 | 0 | 0 |
| aamsellem | 1 | 0 | 0 | 1 |
| cryptopsy0 | 1 | 0 | 0 | 0 |
| dandenkijin | 1 | 0 | 0 | 1 |
| dani-nagy | 1 | 0 | 0 | 1 |
| github-actions[bot] | 1 | 0 | 1 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue#69yundddd2026-06-21 01:55Eval bug: crash during long agentic session
- Issue#74ethernidee2026-06-16 19:07Misc. bug: extra 200+ MB VRAM consumption on start
- Issue#60Anbeeld2026-06-14 23:54Eval bug: Sm tensor doesn't work
- Issue comment#60Anbeeld2026-06-14 23:54Eval bug: Sm tensor doesn't work
- Issue#61Anbeeld2026-06-14 21:24Eval bug: 0.3.2 compiled from latest source crashes
- Issue#59Ezzz-dev2026-06-06 15:05Eval bug: Vulkan/HIP is way too slow in comparison to llama.cpp
- Issue#57Ezzz-dev2026-06-05 15:42Eval bug: prompt is re-processing for no reason
- Issue comment#45Anbeeld2026-06-04 22:58Eval bug: llama.cpp (AMD GPU): partial layers run on CPU when loading dflash model, cannot fully offload to GPU
- Issue#52Anbeeld2026-06-04 21:28Eval bug: Multimodal model loading fails for Gemma 4 12B (unknown projector type: gemma4uv)
- Issue#49Anbeeld2026-06-02 22:49Misc. bug: xmfp6 and turbo3 sharing same ID:42 this preventing from loading xmfp6 models
- Issue comment#30Ponkipon2026-06-02 10:41Compile bug: Metal library fails to compile — block_turbo4_0 missing signs / rnorm fields referenced by quantize_turbo4_0
- Issue comment#39Anbeeld2026-05-27 00:31Eval bug: low DFlash tps/AR for multi-GPU setups
- Issue comment#41Anbeeld2026-05-26 23:22Eval bug: Segfault at slot initialization with CUDA on SM75 (Turing) — ngl > 0
- Issue#41Bino51502026-05-26 22:59Misc. bug: Segfault at slot initialization with CUDA on SM75 (Turing) — ngl > 0
- Issue#40cryptopsy02026-05-26 22:04profiling radeon 7900xt (20gb vram)
- Issue comment#33gdeyoung2026-05-26 01:19DFlash segfault (exit 139) with Qwen3.6-27B on RTX PRO 4500 Blackwell (SM120)
- Issue#37The-Dude-20202026-05-25 20:12Eval bug: When trying to WRITE a large content file I see the below error in the logs and the write timesout...
- Issue comment#22Anbeeld2026-05-25 00:36No speed gain on 20gb vram with 27b q4
- Issue#35ethernidee2026-05-24 23:56Eval bug: --ctx-size with --parallel 1 is divided by 2
- Issue comment#32Anbeeld2026-05-24 20:34spec-dec: support device-aware recurrent GPU tape placement on ROCm multi-GPU
- Issue comment#15dani-nagy2026-05-23 19:06Eval bug: Qwen3.6 35B A3B Q4 crash with "longer" context using turbo3 + dflash, cpu-offloading | Win10
- Issue comment#22Anbeeld2026-05-23 01:06No speed gain on 20gb vram with 27b q4
- Issue comment#28Anbeeld2026-05-23 00:31Eval bug: Multi GPU with draft errors: - the tokens for sequence 0 in the input batch have a starting position of Y = 1360 it is required that the sequence positions remain consecutive: Y = X + 1
- Issue#27l3ateman2026-05-22 22:14invalid reduced-logits token -1
- Issue#4Anbeeld2026-05-22 17:46Eval bug: llama_model_load: error loading model: error loading model architecture: unknown model architecture: 'dflash'
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 74 stars here means stars gained during the window, not the repo's star count.