TurboQuant KV Cache Compression for llama.cpp — 5.2x memory reduction with near-lossless quality | Implementation of Google DeepMind's TurboQuant (ICLR 2026)
C++ · active 2026-03-29 → 2026-08-08 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 133 days from 2026-03-29 to 2026-08-08. Pushes: 80 total, peak 10 in a day. Pull requests: 1 total, peak 1 in a day. Issues: 4 total, peak 2 in a day. Comments: 23 total, peak 6 in a day. Stars: 11 total, peak 2 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| AmesianX | 100 | 80 | 0 | 18 |
| jarkevithwlad | 2 | 0 | 0 | 1 |
| primoco | 1 | 0 | 0 | 1 |
| alrunan | 1 | 0 | 0 | 1 |
| richardokonicha | 1 | 0 | 1 | 0 |
| zekrom-vale | 1 | 0 | 0 | 1 |
| infinity7117 | 1 | 0 | 0 | 0 |
| github-actions[bot] | 1 | 0 | 0 | 1 |
Recent activity
Latest issues, pull requests and releases
- Issue comment#17github-actions[bot]2026-05-26 01:24Eval bug: gemma-4-26b-a4b-it-turbo using tbq4_0 is insane (and gemma3 fails to load correctly)
- Issue comment#18AmesianX2026-04-11 17:12Eval bug: llama-server crashes without error if using any turboquant cache type. works normally with q4_0
- Issue comment#17AmesianX2026-04-11 17:11Eval bug: gemma-4-26b-a4b-it-turbo using tbq4_0 is insane (and gemma3 fails to load correctly)
- ReleaseAmesianX2026-04-11 17:06TurboQuant v1.5.3 — Double WHT Per-Head for D=64
- Issue#18infinity71172026-04-08 11:10Eval bug: llama-server crashes without error if using any turboquant cache type. works normally with q4_0
- Issue comment#17AmesianX2026-04-08 00:43Eval bug: gemma-4-26b-a4b-it-turbo using tbq4_0 is insane (and gemma3 fails to load correctly)
- Issue#5AmesianX2026-04-08 00:19rocm build failed
- Issue comment#5AmesianX2026-04-08 00:19rocm build failed
- Issue comment#14AmesianX2026-04-08 00:19Compile bug: warning: no usable GPU found Build When DGGML_HIP=ON
- Issue comment#17zekrom-vale2026-04-07 18:06Eval bug: gemma-4-26b-a4b-it-turbo using tbq4_0 is insane (and gemma3 fails to load correctly)
- ReleaseAmesianX2026-04-06 17:33v1.5.2 — PPL 21%→8%, Attention Sharpening + V Rotation Bugfix
- Issue#15jarkevithwlad2026-04-06 11:01Eval bug: gemma output <unused24><unused24><unused24><unused24>
- Issue comment#11AmesianX2026-04-06 10:23V-cache precision bug persists in v1.4.1 on head_dim=128 models (Qwen3-14b) — KV recall degradation from ctx ~1100t
- Issue comment#11primoco2026-04-06 09:47V-cache precision bug persists in v1.4.1 on head_dim=128 models (Qwen3-14b) — KV recall degradation from ctx ~1100t
- Issue comment#15AmesianX2026-04-06 08:49Eval bug: gemma output <unused24><unused24><unused24><unused24>
- Issue#8AmesianX2026-04-06 08:39Windows Support
- Issue comment#4AmesianX2026-04-06 08:39pre-allocated tensor (cache_k_l3 (view)) in a buffer (Vulkan0) that cannot run the operation (SET_ROWS)
- ReleaseAmesianX2026-04-06 08:13TurboQuant v1.5.1 — Exceeds f16 Quality (4.2x compression)
- Issue comment#13AmesianX2026-04-05 15:26feat: CPU flash attention support for TurboQuant head_dim=256 (_0 types)
- Issue comment#11AmesianX2026-04-04 11:45V-cache precision bug persists in v1.4.1 on head_dim=128 models (Qwen3-14b) — KV recall degradation from ctx ~1100t
- Pull request#10richardokonicha2026-04-04 08:55
- Issue comment#6AmesianX2026-04-04 08:11Support gemma 4
- Issue comment#6jarkevithwlad2026-04-03 07:43Support gemma 4
- Issue comment#4AmesianX2026-04-03 02:14pre-allocated tensor (cache_k_l3 (view)) in a buffer (Vulkan0) that cannot run the operation (SET_ROWS)
- Issue comment#5AmesianX2026-04-03 02:10rocm build failed
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 11 stars here means stars gained during the window, not the repo's star count.