Skip to content

cfregly/ai-performance-engineering

View on GitHub ↗Related repositories →

Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning.

active 2025-04-272026-08-18 (UTC)

Complete coverage26,444 / 26,444 hourly files (100%) · 2 absent upstream2023-08-152026-08-20 (UTC)
Events
1.6K
Pushes
993
Pull requests
1
Issues
9
Stars
532
Forks
47

Activity over time

Daily event counts in the loaded window

Line chart, 479 days from 2025-04-27 to 2026-08-18. Pushes: 993 total, peak 163 in a day. Pull requests: 1 total, peak 1 in a day. Issues: 9 total, peak 2 in a day. Comments: 3 total, peak 1 in a day. Stars: 532 total, peak 48 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
cfregly99999401
cr72582010
Rustem2002
AnDongLi1000
17Swagat1000
abdullah-athar1000
tunglinwood1000

Recent activity

Latest issues, pull requests and releases

  • Issue comment#10Rustem2026-03-19 00:21
    ch4/dist_all_reduce.py - Fix NCCL hang in dist_allreduce.py: assign correct GPU per rank
  • Issue#8cr72582026-03-15 07:55
    ch03/numa_topology_script.sh: nvidia-smi topo -m does not support -i flag
  • Pull request#7cr72582026-03-14 16:49
  • Issue#6AnDongLi2026-03-09 21:05
    The file names are misleading in chapter 10
  • Issue comment#4Rustem2026-02-28 17:03
    Unable to run chapter 01 benchmark
  • Issue#517Swagat2026-02-28 12:13
    How to utilize this repo for learn about cuda?
  • Issue comment#4cfregly2026-02-17 05:59
    Unable to run chapter 01 benchmark
  • Issue#3tunglinwood2026-01-20 01:37
    Meetup Group Recording
  • Issue#2abdullah-athar2025-09-22 17:27
    Starter code for Auto-optimize Pytorch Code
  • Issue#1cfregly2025-05-31 14:08
    Upgrade everything to new PyTorch 2.7, CUDA 12.8, Triton 3.3 (Blackwell support)
  • Issue#1cfregly2025-05-22 06:03
    Upgrade everything to new PyTorch 2.7, CUDA 12.8, Triton 3.3 (Blackwell support)
  • Issue#1cfregly2025-05-22 06:03
    Upgrade everything to new PyTorch 2.7, CUDA 12.8, Triton 3.3 (Blackwell support)
  • Issue#1cfregly2025-04-28 19:07
    Upgrade everything to new PyTorch 2.7, CUDA 12.8, Triton 3.3 (Blackwell support)

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 532 stars here means stars gained during the window, not the repo's star count.