Layer-Condensed KV cache w/ 10 times larger batch size, fewer params and less computation. Dramatic speed up with better task performance. Accepted to ACL 2024.
active 2024-05-20 → 2025-11-10 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 540 days from 2024-05-20 to 2025-11-10. Pushes: 55 total, peak 9 in a day. Pull requests: 9 total, peak 5 in a day. Issues: 19 total, peak 2 in a day. Comments: 44 total, peak 8 in a day. Stars: 146 total, peak 35 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| why-in-Shanghaitech | 86 | 55 | 9 | 21 |
| SpoSer23 | 11 | 0 | 0 | 5 |
| ChenHong30 | 9 | 0 | 0 | 5 |
| Mostafa-Emad77 | 8 | 0 | 0 | 6 |
| pigdogbaby | 6 | 0 | 0 | 6 |
| 311dada | 3 | 0 | 0 | 1 |
| Draconis98 | 2 | 0 | 0 | 0 |
| alvi75 | 1 | 0 | 0 | 0 |
| 123shivanshukumar | 1 | 0 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue#17Draconis982025-07-13 12:06Missing torch dependency for flash-attn in requirements.txt
- Issue#17Draconis982025-07-13 05:33Missing torch dependency for flash-attn in requirements.txt
- Issue#16123shivanshukumar2025-06-21 14:03KV computation during inference
- Issue#12ChenHong302025-02-21 07:02Question about supporting other settings of cross-layer KV cache sharing
- Issue comment#15why-in-Shanghaitech2024-12-25 10:58Layer Types Configuration Intuition
- Issue#14SpoSer232024-12-12 13:42Is training from scratch essential?
- Issue comment#14why-in-Shanghaitech2024-12-10 01:11Is training from scratch essential?
- Issue comment#14SpoSer232024-12-09 13:12Is training from scratch essential?
- Issue comment#14why-in-Shanghaitech2024-12-09 00:28Is training from scratch essential?
- Issue#14SpoSer232024-12-08 21:07Is training from scratch essential?
- Issue#13SpoSer232024-12-05 22:30use_sequential flag in training
- Issue comment#13pigdogbaby2024-12-05 03:13use_sequential flag in training
- Issue#13SpoSer232024-12-04 11:49use_sequential flag in training
- Issue comment#11Mostafa-Emad772024-12-03 11:38Guidance on Fine-Tuning Llama in LCKV Framework
- Issue#11Mostafa-Emad772024-12-03 11:38Guidance on Fine-Tuning Llama in LCKV Framework
- Issue comment#11why-in-Shanghaitech2024-12-02 02:03Guidance on Fine-Tuning Llama in LCKV Framework
- Issue comment#11why-in-Shanghaitech2024-12-01 11:08Guidance on Fine-Tuning Llama in LCKV Framework
- Issue comment#11why-in-Shanghaitech2024-12-01 10:59Guidance on Fine-Tuning Llama in LCKV Framework
- Issue comment#10Mostafa-Emad772024-12-01 10:36What is the specific version of the Python this project running on?
- Issue comment#11Mostafa-Emad772024-12-01 09:11Guidance on Fine-Tuning Llama in LCKV Framework
- Issue comment#11why-in-Shanghaitech2024-11-28 14:10Guidance on Fine-Tuning Llama in LCKV Framework
- Issue comment#11why-in-Shanghaitech2024-11-21 02:01Guidance on Fine-Tuning Llama in LCKV Framework
- Issue comment#11Mostafa-Emad772024-11-20 15:03Guidance on Fine-Tuning Llama in LCKV Framework
- Issue comment#11why-in-Shanghaitech2024-11-20 12:19Guidance on Fine-Tuning Llama in LCKV Framework
- Issue comment#11Mostafa-Emad772024-11-20 11:00Guidance on Fine-Tuning Llama in LCKV Framework
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 146 stars here means stars gained during the window, not the repo's star count.