A repo for RLHF training and BoN over LLMs, with support for reward model ensembles.
active 2024-03-09 → 2025-11-04 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 606 days from 2024-03-09 to 2025-11-04. Pushes: 4 total, peak 3 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 31 total, peak 5 in a day. Comments: 28 total, peak 11 in a day. Stars: 42 total, peak 4 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| tlc4418 | 25 | 4 | 0 | 13 |
| RylanSchaeffer | 20 | 0 | 0 | 8 |
| ZixuanLiu4869 | 3 | 0 | 0 | 1 |
| tsWen0309 | 3 | 0 | 0 | 2 |
| JohannesAck | 3 | 0 | 0 | 2 |
| cassidylaidlaw | 2 | 0 | 0 | 0 |
| xueyongfu11 | 2 | 0 | 0 | 0 |
| georgao35 | 2 | 0 | 0 | 1 |
| yiwan-rl | 1 | 0 | 0 | 0 |
| sheikhshafayat | 1 | 0 | 0 | 1 |
| zetian1025 | 1 | 0 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue comment#18ZixuanLiu48692025-05-18 19:43Command for reward training
- Issue#18ZixuanLiu48692025-05-18 19:41Command for reward training
- Issue#18ZixuanLiu48692025-05-18 19:40Command for reward training
- Issue comment#17tlc44182025-01-16 11:11Model loading error when conducting PPO experiment
- Issue comment#15tlc44182025-01-16 11:08BoN sampling: ArrowInvalid: offset overflow while concatenating arrays
- Issue#15tlc44182025-01-16 11:08BoN sampling: ArrowInvalid: offset overflow while concatenating arrays
- Issue comment#14tlc44182025-01-16 11:06AttributeError: 'Namespace' object has no attribute 'residual_dropout_lima'
- Issue#14tlc44182025-01-16 11:06AttributeError: 'Namespace' object has no attribute 'residual_dropout_lima'
- Issue comment#13tlc44182025-01-16 11:04My BoN result is different from the result (Figure 3a) shown in the paper.
- Issue#13tlc44182025-01-16 11:04My BoN result is different from the result (Figure 3a) shown in the paper.
- Issue#6tlc44182025-01-16 10:46Why is reward model training not logged to W&B?
- Issue comment#10tlc44182025-01-16 10:45Unable to Run PPO Training Using HuggingFace Path of SFT'd language model
- Issue#10tlc44182025-01-16 10:45Unable to Run PPO Training Using HuggingFace Path of SFT'd language model
- Issue#17tsWen03092024-12-10 07:57Model loading error when conducting PPO experiment
- Issue comment#10tsWen03092024-12-10 07:40Unable to Run PPO Training Using HuggingFace Path of SFT'd language model
- Issue comment#15tsWen03092024-11-22 06:32BoN sampling: ArrowInvalid: offset overflow while concatenating arrays
- Issue#16xueyongfu112024-10-25 12:40How to reproduce the experiments related to PPO
- Issue#16xueyongfu112024-10-23 08:48How to reproduce the experiments related to PPO
- Issue#15cassidylaidlaw2024-09-23 18:07BoN sampling: ArrowInvalid: offset overflow while concatenating arrays
- Issue#14cassidylaidlaw2024-09-23 16:55AttributeError: 'Namespace' object has no attribute 'residual_dropout_lima'
- Issue#13yiwan-rl2024-08-06 21:05My BoN result is different from the result (Figure 3a) shown in the paper.
- Issue#9tlc44182024-08-05 09:21Best-of-n Pipeline takes ages - how to accelerate?
- Issue#8tlc44182024-08-05 09:21When (not) to use Flash Attention?
- Issue#1tlc44182024-08-05 09:21How to re-implement the score-KL curve?
- Issue comment#5sheikhshafayat2024-08-01 14:04Training Reward Model: AttributeError: 'Namespace' object has no attribute 'residual_dropout_lima'. Did you mean: 'residual_dropout'?
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 42 stars here means stars gained during the window, not the repo's star count.