Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization 帮助用户计算在训练或推理大模型时所需的GPU内存。该项目不仅可以计算GPU内存的占用情况,还可以提供详细的内存分布情况、评估量化方法、处理的最大上下文长度等信息,帮助用户选择适合自己的GPU配置。
active 2023-09-12 → 2026-08-26 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 1080 days from 2023-09-12 to 2026-08-26. Pushes: 98 total, peak 12 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 25 total, peak 4 in a day. Comments: 26 total, peak 5 in a day. Stars: 1,323 total, peak 132 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| RahulSChand | 120 | 98 | 0 | 14 |
| LaniakeaS | 7 | 0 | 0 | 1 |
| Geministudents | 4 | 0 | 0 | 3 |
| Anindyadeep | 4 | 0 | 0 | 3 |
| HuaYZhao | 2 | 0 | 0 | 0 |
| fishiu | 2 | 0 | 0 | 1 |
| RaccoonOnion | 2 | 0 | 0 | 0 |
| ChloeL19 | 1 | 0 | 0 | 1 |
| bver | 1 | 0 | 0 | 1 |
| flei2019 | 1 | 0 | 0 | 0 |
| AaronZLT | 1 | 0 | 0 | 1 |
| jag89 | 1 | 0 | 0 | 0 |
| 01lin | 1 | 0 | 0 | 0 |
| hunter-xue | 1 | 0 | 0 | 1 |
| RabbitTwin | 1 | 0 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue#16jag892024-12-10 14:28How to get list of model names and architecture details programmatically? How did you populate all_configs.json
- Issue comment#8bver2024-11-08 15:09compute in gpu_configs.json meaning
- Issue comment#9RahulSChand2024-11-02 02:15why batch size does not effect to memory usage in inference mode
- Issue#9RahulSChand2024-11-02 02:15why batch size does not effect to memory usage in inference mode
- Issue comment#13RahulSChand2024-11-02 02:03Can you add the A100 gpus?
- Issue#13RahulSChand2024-11-02 02:03Can you add the A100 gpus?
- Issue comment#13hunter-xue2024-11-01 07:25Can you add the A100 gpus?
- Issue#1401lin2024-09-12 14:41How to calculate token_per_second_by_GPU, there are specific formula instructions?
- Issue comment#13ChloeL192024-08-26 04:41Can you add the A100 gpus?
- Issue comment#1AaronZLT2024-07-29 09:30Results are inconsistent and is not reliable enough
- Issue#12HuaYZhao2024-03-28 07:48Activation Memory
- Issue#12HuaYZhao2024-03-20 13:04Activation Memory
- Issue#11flei20192024-01-16 01:42What's the meaning of magic numbers?
- Issue#10RabbitTwin2024-01-13 07:54Missing License
- Issue#9LaniakeaS2023-12-13 02:54why batch size does not effect to memory usage in inference mode
- Issue#8RaccoonOnion2023-12-04 07:54compute in gpu_configs.json meaning
- Issue comment#7LaniakeaS2023-11-09 03:20Name and size from same model can cause different result
- Issue#7LaniakeaS2023-11-09 03:20Name and size from same model can cause different result
- Issue comment#6RahulSChand2023-11-09 03:02DeepSpeed support
- Issue#6RahulSChand2023-11-09 03:02DeepSpeed support
- Issue comment#7RahulSChand2023-11-09 02:36Name and size from same model can cause different result
- Issue#7LaniakeaS2023-11-09 02:22Name and size from same model can cause different result
- Issue#6LaniakeaS2023-11-09 01:49DeepSpeed support
- Issue#5LaniakeaS2023-11-08 12:13The memory usage in LoRA finetuning
- Issue comment#5RahulSChand2023-11-08 12:11The memory usage in LoRA finetuning
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 1,323 stars here means stars gained during the window, not the repo's star count.