PostTrainBench measures how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
Python · active 2025-11-28 → 2026-08-31 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 277 days from 2025-11-28 to 2026-08-31. Pushes: 66 total, peak 4 in a day. Pull requests: 11 total, peak 2 in a day. Issues: 12 total, peak 2 in a day. Comments: 31 total, peak 4 in a day. Stars: 89 total, peak 8 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| rank-and-file | 56 | 43 | 0 | 9 |
| hrdkbhatnagar | 47 | 22 | 8 | 8 |
| max-andr | 8 | 1 | 0 | 7 |
| shiraeisenberg | 3 | 0 | 1 | 2 |
| nikil-ravi | 2 | 0 | 1 | 1 |
| justinwangx | 2 | 0 | 1 | 1 |
| nevasinisasikumar-cmyk | 2 | 0 | 0 | 1 |
| answers111 | 1 | 0 | 0 | 0 |
| echo-yiyiyi | 1 | 0 | 0 | 1 |
| wise-east | 1 | 0 | 0 | 1 |
Recent activity
Latest issues, pull requests and releases
- Issue comment#41rank-and-file2026-05-15 08:09Simplify aggregation scripts (+ more quality of life changes)
- Issue#38rank-and-file2026-05-02 04:32Support eval of LoRA adapters
- Pull request#31hrdkbhatnagar2026-04-04 17:52
- Issue#28hrdkbhatnagar2026-04-03 11:34Log time and other system details in the agent traces
- Issue comment#10rank-and-file2026-04-02 14:49feat: add basic ReAct agent
- Issue comment#25hrdkbhatnagar2026-03-04 09:54Improve contamination judge
- Issue#26rank-and-file2026-03-04 07:47Extract traces
- Issue comment#25rank-and-file2026-03-04 07:23Improve contamination judge
- Issue comment#25rank-and-file2026-03-04 07:21Improve contamination judge
- Issue#25rank-and-file2026-03-04 06:42Improve contamination judge
- Issue#15hrdkbhatnagar2026-02-23 15:14Add GLM 5
- Issue#16hrdkbhatnagar2026-02-22 19:44update readme with v1 results
- Pull request#23hrdkbhatnagar2026-02-22 19:44
- Pull request#23hrdkbhatnagar2026-02-22 19:43
- Issue#22answers1112026-02-22 06:38Question about resources.json
- Issue comment#10max-andr2026-02-22 00:45feat: add basic ReAct agent
- Issue#20hrdkbhatnagar2026-02-21 16:41add extra logging for terminated and killed runs
- Pull request#21hrdkbhatnagar2026-02-21 16:40
- Issue comment#10justinwangx2026-02-20 07:08feat: add basic ReAct agent
- Issue comment#18wise-east2026-02-17 20:23how to compute weighted average performance for individual model scores?
- Issue comment#19hrdkbhatnagar2026-02-15 14:37`huggingface-hub` Version Conflict with current `transformers` and `inspect-ai` dependency requirement
- Issue comment#13echo-yiyiyi2026-02-14 20:29When will you support `slurm`?
- Issue#15hrdkbhatnagar2026-02-14 20:01Add GLM 5
- Issue comment#18hrdkbhatnagar2026-02-13 00:21how to compute weighted average performance for individual model scores?
- Issue#17hrdkbhatnagar2026-02-11 21:41add ability to evaluate agents without API (such as Codex 5.3)
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 89 stars here means stars gained during the window, not the repo's star count.