Skip to content

TurboQuant KV Cache Compression for llama.cpp — 5.2x memory reduction with near-lossless quality | Implementation of Google DeepMind's TurboQuant (ICLR 2026)

C++ · active 2026-03-292026-08-08 (UTC)

Partial coverage13,981 / 16,761 hourly files (83%) · 2 absent upstream · 2,779 failed, retryable2024-09-122026-08-11 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
143
Pushes
80
Pull requests
1
Issues
4
Stars
11
Forks
2

Activity over time

Daily event counts in the loaded window

Line chart, 133 days from 2026-03-29 to 2026-08-08. Pushes: 80 total, peak 10 in a day. Pull requests: 1 total, peak 1 in a day. Issues: 4 total, peak 2 in a day. Comments: 23 total, peak 6 in a day. Stars: 11 total, peak 2 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
AmesianX10080018
jarkevithwlad2001
primoco1001
alrunan1001
richardokonicha1010
zekrom-vale1001
infinity71171000
github-actions[bot]1001

Recent activity

Latest issues, pull requests and releases

  • Issue comment#17github-actions[bot]2026-05-26 01:24
    Eval bug: gemma-4-26b-a4b-it-turbo using tbq4_0 is insane (and gemma3 fails to load correctly)
  • Issue comment#18AmesianX2026-04-11 17:12
    Eval bug: llama-server crashes without error if using any turboquant cache type. works normally with q4_0
  • Issue comment#17AmesianX2026-04-11 17:11
    Eval bug: gemma-4-26b-a4b-it-turbo using tbq4_0 is insane (and gemma3 fails to load correctly)
  • ReleaseAmesianX2026-04-11 17:06
    TurboQuant v1.5.3 — Double WHT Per-Head for D=64
  • Issue#18infinity71172026-04-08 11:10
    Eval bug: llama-server crashes without error if using any turboquant cache type. works normally with q4_0
  • Issue comment#17AmesianX2026-04-08 00:43
    Eval bug: gemma-4-26b-a4b-it-turbo using tbq4_0 is insane (and gemma3 fails to load correctly)
  • Issue#5AmesianX2026-04-08 00:19
    rocm build failed
  • Issue comment#5AmesianX2026-04-08 00:19
    rocm build failed
  • Issue comment#14AmesianX2026-04-08 00:19
    Compile bug: warning: no usable GPU found Build When DGGML_HIP=ON
  • Issue comment#17zekrom-vale2026-04-07 18:06
    Eval bug: gemma-4-26b-a4b-it-turbo using tbq4_0 is insane (and gemma3 fails to load correctly)
  • ReleaseAmesianX2026-04-06 17:33
    v1.5.2 — PPL 21%→8%, Attention Sharpening + V Rotation Bugfix
  • Issue#15jarkevithwlad2026-04-06 11:01
    Eval bug: gemma output <unused24><unused24><unused24><unused24>
  • Issue comment#11AmesianX2026-04-06 10:23
    V-cache precision bug persists in v1.4.1 on head_dim=128 models (Qwen3-14b) — KV recall degradation from ctx ~1100t
  • Issue comment#11primoco2026-04-06 09:47
    V-cache precision bug persists in v1.4.1 on head_dim=128 models (Qwen3-14b) — KV recall degradation from ctx ~1100t
  • Issue comment#15AmesianX2026-04-06 08:49
    Eval bug: gemma output <unused24><unused24><unused24><unused24>
  • Issue#8AmesianX2026-04-06 08:39
    Windows Support
  • Issue comment#4AmesianX2026-04-06 08:39
    pre-allocated tensor (cache_k_l3 (view)) in a buffer (Vulkan0) that cannot run the operation (SET_ROWS)
  • ReleaseAmesianX2026-04-06 08:13
    TurboQuant v1.5.1 — Exceeds f16 Quality (4.2x compression)
  • Issue comment#13AmesianX2026-04-05 15:26
    feat: CPU flash attention support for TurboQuant head_dim=256 (_0 types)
  • Issue comment#11AmesianX2026-04-04 11:45
    V-cache precision bug persists in v1.4.1 on head_dim=128 models (Qwen3-14b) — KV recall degradation from ctx ~1100t
  • Pull request#10richardokonicha2026-04-04 08:55
  • Issue comment#6AmesianX2026-04-04 08:11
    Support gemma 4
  • Issue comment#6jarkevithwlad2026-04-03 07:43
    Support gemma 4
  • Issue comment#4AmesianX2026-04-03 02:14
    pre-allocated tensor (cache_k_l3 (view)) in a buffer (Vulkan0) that cannot run the operation (SET_ROWS)
  • Issue comment#5AmesianX2026-04-03 02:10
    rocm build failed

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 11 stars here means stars gained during the window, not the repo's star count.