Nitro is an C++ inference server on top of TensorRT-LLM. OpenAI-compatible API. Run blazing fast inference on Nvidia GPUs. Used in Jan
active 2024-04-25 → 2025-03-08 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 318 days from 2024-04-25 to 2025-03-08. Pushes: 194 total, peak 17 in a day. Pull requests: 55 total, peak 6 in a day. Issues: 16 total, peak 6 in a day. Comments: 28 total, peak 8 in a day. Stars: 14 total, peak 1 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| hiento09 | 113 | 88 | 21 | 0 |
| vansangpfiev | 82 | 32 | 25 | 11 |
| nguyenhoangthuan99 | 47 | 39 | 3 | 1 |
| CameronNg | 33 | 30 | 1 | 0 |
| gabrielle-ong | 14 | 0 | 0 | 7 |
| jan-service-account | 10 | 5 | 5 | 0 |
| github-actions[bot] | 8 | 0 | 0 | 8 |
| imtuyethan | 2 | 0 | 0 | 1 |
| Van-QA | 1 | 0 | 0 | 0 |
| tikikun | 1 | 0 | 0 | 0 |
| irfanpena | 1 | 0 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue comment#33gabrielle-ong2024-11-28 07:23feat: TensorRT-LLM load multiple models
- Issue#33gabrielle-ong2024-11-28 07:23feat: TensorRT-LLM load multiple models
- Issue comment#32gabrielle-ong2024-11-28 07:20feat: TensorRT-LLM Unload Model
- Issue#32gabrielle-ong2024-11-28 07:20feat: TensorRT-LLM Unload Model
- Issue comment#31gabrielle-ong2024-11-28 07:20feat: TensorRT-LLM Request Interruption
- Issue#31gabrielle-ong2024-11-28 07:19feat: TensorRT-LLM Request Interruption
- Issue comment#30gabrielle-ong2024-11-28 07:19feat: TensorRT-LLM InferenceRequest and stop_words_list
- Issue#30gabrielle-ong2024-11-28 07:19feat: TensorRT-LLM InferenceRequest and stop_words_list
- Issue comment#29gabrielle-ong2024-11-28 07:19feat: TensorRT-LLM Inflight batching
- Issue#29gabrielle-ong2024-11-28 07:19feat: TensorRT-LLM Inflight batching
- Issue comment#55gabrielle-ong2024-11-28 07:19feat: tensorrt-llm-engine README.md file
- Issue#55gabrielle-ong2024-11-28 07:19feat: tensorrt-llm-engine README.md file
- Issue comment#54gabrielle-ong2024-11-28 07:19feat: TensorRT-LLM Support for logits_prob
- Issue#74gabrielle-ong2024-09-27 06:00Llama3.2-11b-Vision
- Pull request#73hiento092024-09-26 08:36
- Pull request#73hiento092024-09-26 08:00
- Pull request#72nguyenhoangthuan992024-09-04 02:55
- Issue comment#27imtuyethan2024-08-29 03:55bug: tensorRT - Switching between model is causing error satisfyProfile Runtime dimension does not satisfy any optimization profile
- Issue#27imtuyethan2024-08-29 03:55bug: tensorRT - Switching between model is causing error satisfyProfile Runtime dimension does not satisfy any optimization profile
- Issue comment#51vansangpfiev2024-08-28 10:12feat: use batch-manager instead of gpt-runtime
- Issue#51vansangpfiev2024-08-28 10:12feat: use batch-manager instead of gpt-runtime
- Pull request#72nguyenhoangthuan992024-08-27 08:28
- Releasegithub-actions[bot]2024-08-06 05:240.0.9
- Pull request#71vansangpfiev2024-08-06 03:47
- Pull request#71vansangpfiev2024-08-05 06:52
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 14 stars here means stars gained during the window, not the repo's star count.