Skip to content

Official implementation of TransNormerLLM: A Faster and Better LLM

active 2023-08-152025-10-29 (UTC)

Complete coverage26,386 / 26,386 hourly files (100%) · 2 absent upstream2023-08-152026-08-18 (UTC)
Events
187
Pushes
24
Pull requests
2
Issues
12
Stars
112
Forks
8

Activity over time

Daily event counts in the loaded window

Line chart, 807 days from 2023-08-15 to 2025-10-29. Pushes: 24 total, peak 11 in a day. Pull requests: 2 total, peak 1 in a day. Issues: 12 total, peak 2 in a day. Comments: 29 total, peak 6 in a day. Stars: 112 total, peak 5 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
XuyangShen211407
Doraemonzzz162012
liddalidd4400
XintianHan4003
weigao2664400
Hanshifancoder4003
Leopold23333001
relic-yuexi3001
redbrain2020
janEbert1000
waneon1000
iminfine1000
wangyuxin871000
OpenNLPLab1231001
ChuanhongLi1001

Recent activity

Latest issues, pull requests and releases

  • Issue comment#12Leopold23332024-11-29 01:33
    Confusions about the decay parameter λ
  • Issue#12Leopold23332024-11-29 01:33
    Confusions about the decay parameter λ
  • Issue comment#12Doraemonzzz2024-11-28 10:34
    Confusions about the decay parameter λ
  • Issue#12Leopold23332024-11-28 10:09
    Confusions about the decay parameter λ
  • Issue comment#11Hanshifancoder2024-11-01 06:29
    Differences between Lightning Attention1 and Lightning Attention2 code implementations
  • Issue comment#11Doraemonzzz2024-11-01 06:25
    Differences between Lightning Attention1 and Lightning Attention2 code implementations
  • Issue comment#11Hanshifancoder2024-11-01 06:19
    Differences between Lightning Attention1 and Lightning Attention2 code implementations
  • Issue comment#11Doraemonzzz2024-11-01 03:33
    Differences between Lightning Attention1 and Lightning Attention2 code implementations
  • Issue comment#11Hanshifancoder2024-11-01 03:28
    Differences between Lightning Attention1 and Lightning Attention2 code implementations
  • Issue comment#11Doraemonzzz2024-10-31 15:35
    Differences between Lightning Attention1 and Lightning Attention2 code implementations
  • Issue#11Hanshifancoder2024-10-31 10:43
    Differences between Lightning Attention1 and Lightning Attention2 code implementations
  • Issue#10Doraemonzzz2024-04-01 07:27
    This is not a linear attention transformer.
  • Issue comment#10Doraemonzzz2024-04-01 07:16
    This is not a linear attention transformer.
  • Issue#10iminfine2024-04-01 07:07
    This is not a linear attention transformer.
  • Issue comment#9Doraemonzzz2024-01-25 13:22
    Bugs in Triton operator?
  • Pull request#5redbrain2024-01-24 15:08
  • Issue comment#8Doraemonzzz2024-01-24 08:22
    Benchmark results can not be reproduced
  • Issue#8waneon2024-01-24 08:16
    Benchmark results can not be reproduced
  • Issue comment#7XuyangShen2024-01-23 05:00
    你好,请问各个参数量的模型默认加载使用会占用多少显存?不同max_new_tokens大概会要占用多少显存?
  • Issue comment#7ChuanhongLi2024-01-23 03:36
    你好,请问各个参数量的模型默认加载使用会占用多少显存?不同max_new_tokens大概会要占用多少显存?
  • Issue#7relic-yuexi2024-01-12 09:44
    你好,请问各个参数量的模型默认加载使用会占用多少显存?不同max_new_tokens大概会要占用多少显存?
  • Issue comment#7OpenNLPLab1232024-01-12 03:58
    你好,请问各个参数量的模型默认加载使用会占用多少显存?不同max_new_tokens大概会要占用多少显存?
  • Issue comment#7relic-yuexi2024-01-12 01:58
    你好,请问各个参数量的模型默认加载使用会占用多少显存?不同max_new_tokens大概会要占用多少显存?
  • Issue comment#7XuyangShen2024-01-11 13:25
    你好,请问各个参数量的模型默认加载使用会占用多少显存?不同max_new_tokens大概会要占用多少显存?
  • Issue comment#7XuyangShen2024-01-11 07:06
    你好,请问各个参数量的模型默认加载使用会占用多少显存?不同max_new_tokens大概会要占用多少显存?

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 112 stars here means stars gained during the window, not the repo's star count.