Skip to content

TransMLA: Multi-Head Latent Attention Is All You Need

active 2025-03-052026-03-04 (UTC)

Partial coverage11,520 / 12,566 hourly files (92%) · 2 absent upstream · 1,042 failed, retryable2025-03-052026-08-10 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
188
Pushes
32
Pull requests
0
Issues
24
Stars
100
Forks
0

Activity over time

Daily event counts in the loaded window

Line chart, 365 days from 2025-03-05 to 2026-03-04. Pushes: 32 total, peak 4 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 24 total, peak 8 in a day. Comments: 30 total, peak 6 in a day. Stars: 100 total, peak 4 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

Recent activity

Latest issues, pull requests and releases

  • Issue comment#38fxmeng2025-10-16 02:30
    Support Qwen3, e.g Qwen3-32B(GQA) can be tranasformed to MLA
  • Issue comment#38xueliangyang-oeuler2025-10-16 02:12
    Support Qwen3, e.g Qwen3-32B(GQA) can be tranasformed to MLA
  • Issue comment#37fxmeng2025-10-16 01:17
    Question: Applicability of TransMLA to DLM
  • Issue comment#38fxmeng2025-10-16 01:04
    Support Qwen3, e.g Qwen3-32B(GQA) can be tranasformed to MLA
  • Issue comment#38GiorgioZacharo2025-10-15 15:37
    Support Qwen3, e.g Qwen3-32B(GQA) can be tranasformed to MLA
  • Issue#28Stanleytowne2025-09-23 11:58
    论文中kv cache压缩如何复现
  • Issue comment#34Stanleytowne2025-09-23 11:57
    No performance difference with vLLM: qwen2.5-7b vs qwen2.5-7b-MLA
  • Issue#29Stanleytowne2025-09-23 05:21
    SpeedUp Test and KV Cache is not actually Used?
  • Issue comment#15Stanleytowne2025-09-23 05:16
    torchtune related code open source
  • Issue#15Stanleytowne2025-09-23 05:16
    torchtune related code open source
  • Issue#13Stanleytowne2025-09-23 05:16
    转换qwen-7b模型时,vllm推理报‘Cannot use FlashAttention-2 backend for head size 512.’错误
  • Issue#12Stanleytowne2025-09-23 05:16
    Inference results are not consistent
  • Issue#10Stanleytowne2025-09-23 05:10
    rank r
  • Issue#5Stanleytowne2025-09-23 05:10
    vLLM/TensorRT
  • Issue#2Stanleytowne2025-09-23 05:09
    rope compatible
  • Issue comment#32Stanleytowne2025-09-23 05:07
    Can't run the converted model with vllm.LLM
  • Issue comment#28Stanleytowne2025-09-23 05:06
    论文中kv cache压缩如何复现
  • Issue comment#35Stanleytowne2025-09-23 04:49
    Error with running qwen-3b-instruct converted
  • Issue comment#36Stanleytowne2025-09-23 04:35
    Question: Details of KV cache compression ratios and settings for all models
  • Issue comment#34nikita-yatchenko2025-09-07 15:16
    No performance difference with vLLM: qwen2.5-7b vs qwen2.5-7b-MLA
  • Issue comment#24nikita-yatchenko2025-09-04 15:38
    qwen2-7b and llama2-7b unable to deploy
  • Issue comment#32ItGirls2025-09-04 04:06
    Can't run the converted model with vllm.LLM
  • Issue comment#33JSYRD2025-09-04 01:32
    Qwen 转换后丢失中文能力
  • Issue#33wym422025-09-03 03:47
    Qwen 转换后丢失中文能力
  • Issue comment#32wym422025-09-03 03:07
    Can't run the converted model with vllm.LLM

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 100 stars here means stars gained during the window, not the repo's star count.