Skip to content

Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization 帮助用户计算在训练或推理大模型时所需的GPU内存。该项目不仅可以计算GPU内存的占用情况,还可以提供详细的内存分布情况、评估量化方法、处理的最大上下文长度等信息,帮助用户选择适合自己的GPU配置。

active 2023-09-122026-08-26 (UTC)

Complete coverage26,631 / 26,631 hourly files (100%) · 2 absent upstream2023-08-152026-08-28 (UTC)
Events
1.6K
Pushes
98
Pull requests
0
Issues
25
Stars
1.3K
Forks
75

Activity over time

Daily event counts in the loaded window

Line chart, 1080 days from 2023-09-12 to 2026-08-26. Pushes: 98 total, peak 12 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 25 total, peak 4 in a day. Comments: 26 total, peak 5 in a day. Stars: 1,323 total, peak 132 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
RahulSChand12098014
LaniakeaS7001
Geministudents4003
Anindyadeep4003
HuaYZhao2000
fishiu2001
RaccoonOnion2000
ChloeL191001
bver1001
flei20191000
AaronZLT1001
jag891000
01lin1000
hunter-xue1001
RabbitTwin1000

Recent activity

Latest issues, pull requests and releases

  • Issue#16jag892024-12-10 14:28
    How to get list of model names and architecture details programmatically? How did you populate all_configs.json
  • Issue comment#8bver2024-11-08 15:09
    compute in gpu_configs.json meaning
  • Issue comment#9RahulSChand2024-11-02 02:15
    why batch size does not effect to memory usage in inference mode
  • Issue#9RahulSChand2024-11-02 02:15
    why batch size does not effect to memory usage in inference mode
  • Issue comment#13RahulSChand2024-11-02 02:03
    Can you add the A100 gpus?
  • Issue#13RahulSChand2024-11-02 02:03
    Can you add the A100 gpus?
  • Issue comment#13hunter-xue2024-11-01 07:25
    Can you add the A100 gpus?
  • Issue#1401lin2024-09-12 14:41
    How to calculate token_per_second_by_GPU, there are specific formula instructions?
  • Issue comment#13ChloeL192024-08-26 04:41
    Can you add the A100 gpus?
  • Issue comment#1AaronZLT2024-07-29 09:30
    Results are inconsistent and is not reliable enough
  • Issue#12HuaYZhao2024-03-28 07:48
    Activation Memory
  • Issue#12HuaYZhao2024-03-20 13:04
    Activation Memory
  • Issue#11flei20192024-01-16 01:42
    What's the meaning of magic numbers?
  • Issue#10RabbitTwin2024-01-13 07:54
    Missing License
  • Issue#9LaniakeaS2023-12-13 02:54
    why batch size does not effect to memory usage in inference mode
  • Issue#8RaccoonOnion2023-12-04 07:54
    compute in gpu_configs.json meaning
  • Issue comment#7LaniakeaS2023-11-09 03:20
    Name and size from same model can cause different result
  • Issue#7LaniakeaS2023-11-09 03:20
    Name and size from same model can cause different result
  • Issue comment#6RahulSChand2023-11-09 03:02
    DeepSpeed support
  • Issue#6RahulSChand2023-11-09 03:02
    DeepSpeed support
  • Issue comment#7RahulSChand2023-11-09 02:36
    Name and size from same model can cause different result
  • Issue#7LaniakeaS2023-11-09 02:22
    Name and size from same model can cause different result
  • Issue#6LaniakeaS2023-11-09 01:49
    DeepSpeed support
  • Issue#5LaniakeaS2023-11-08 12:13
    The memory usage in LoRA finetuning
  • Issue comment#5RahulSChand2023-11-08 12:11
    The memory usage in LoRA finetuning

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 1,323 stars here means stars gained during the window, not the repo's star count.