Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning.
active 2025-04-27 → 2026-08-18 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 479 days from 2025-04-27 to 2026-08-18. Pushes: 993 total, peak 163 in a day. Pull requests: 1 total, peak 1 in a day. Issues: 9 total, peak 2 in a day. Comments: 3 total, peak 1 in a day. Stars: 532 total, peak 48 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| cfregly | 999 | 994 | 0 | 1 |
| cr7258 | 2 | 0 | 1 | 0 |
| Rustem | 2 | 0 | 0 | 2 |
| AnDongLi | 1 | 0 | 0 | 0 |
| 17Swagat | 1 | 0 | 0 | 0 |
| abdullah-athar | 1 | 0 | 0 | 0 |
| tunglinwood | 1 | 0 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue comment#10Rustem2026-03-19 00:21ch4/dist_all_reduce.py - Fix NCCL hang in dist_allreduce.py: assign correct GPU per rank
- Issue#8cr72582026-03-15 07:55ch03/numa_topology_script.sh: nvidia-smi topo -m does not support -i flag
- Pull request#7cr72582026-03-14 16:49
- Issue#6AnDongLi2026-03-09 21:05The file names are misleading in chapter 10
- Issue comment#4Rustem2026-02-28 17:03Unable to run chapter 01 benchmark
- Issue#517Swagat2026-02-28 12:13How to utilize this repo for learn about cuda?
- Issue comment#4cfregly2026-02-17 05:59Unable to run chapter 01 benchmark
- Issue#3tunglinwood2026-01-20 01:37Meetup Group Recording
- Issue#2abdullah-athar2025-09-22 17:27Starter code for Auto-optimize Pytorch Code
- Issue#1cfregly2025-05-31 14:08Upgrade everything to new PyTorch 2.7, CUDA 12.8, Triton 3.3 (Blackwell support)
- Issue#1cfregly2025-05-22 06:03Upgrade everything to new PyTorch 2.7, CUDA 12.8, Triton 3.3 (Blackwell support)
- Issue#1cfregly2025-05-22 06:03Upgrade everything to new PyTorch 2.7, CUDA 12.8, Triton 3.3 (Blackwell support)
- Issue#1cfregly2025-04-28 19:07Upgrade everything to new PyTorch 2.7, CUDA 12.8, Triton 3.3 (Blackwell support)
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 532 stars here means stars gained during the window, not the repo's star count.