Extrapolating RLVR to General Domains without Verifiers
active 2025-06-23 → 2026-04-24 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 306 days from 2025-06-23 to 2026-04-24. Pushes: 3 total, peak 2 in a day. Pull requests: 3 total, peak 3 in a day. Issues: 18 total, peak 3 in a day. Comments: 23 total, peak 5 in a day. Stars: 103 total, peak 10 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| yiranyyu | 11 | 3 | 1 | 3 |
| jibo27 | 8 | 0 | 2 | 6 |
| qxzha | 4 | 0 | 0 | 3 |
| hcx-stu | 3 | 0 | 0 | 3 |
| JefferyChen453 | 3 | 0 | 0 | 1 |
| Leon-Francis | 3 | 0 | 0 | 3 |
| banjiuyufen | 2 | 0 | 0 | 1 |
| zyandtom | 1 | 0 | 0 | 1 |
| biburger | 1 | 0 | 0 | 0 |
| RLHF-V | 1 | 0 | 0 | 0 |
| m22f | 1 | 0 | 0 | 0 |
| Lauorie | 1 | 0 | 0 | 0 |
| Trae1ounG | 1 | 0 | 0 | 0 |
| lu-jun-yu | 1 | 0 | 0 | 0 |
| maydaygmail | 1 | 0 | 0 | 0 |
| deleteeeee | 1 | 0 | 0 | 0 |
| PangziZhang523 | 1 | 0 | 0 | 1 |
Recent activity
Latest issues, pull requests and releases
- Issue#23lu-jun-yu2025-10-02 11:12about reward or ppl curve
- Issue comment#16zyandtom2025-09-25 02:33Question about the runtime of main experiments in the RLPR paper
- Issue#17yiranyyu2025-09-05 12:39奖励去偏的问题
- Issue comment#16hcx-stu2025-09-05 09:47Question about the runtime of main experiments in the RLPR paper
- Issue comment#20jibo272025-09-05 09:12Question about ground truth label
- Issue comment#19jibo272025-09-05 09:07generalize to reasoning model
- Issue#20Trae1ounG2025-08-18 10:44Question about ground truth label
- Issue comment#16PangziZhang5232025-08-11 03:36Question about the runtime of main experiments in the RLPR paper
- Issue comment#12Leon-Francis2025-08-08 14:48Question on max_length computation in ray_trainer.py
- Issue comment#12Leon-Francis2025-08-08 13:34Question on max_length computation in ray_trainer.py
- Issue comment#12jibo272025-08-08 11:56Question on max_length computation in ray_trainer.py
- Issue comment#16hcx-stu2025-08-08 10:27Question about the runtime of main experiments in the RLPR paper
- Issue comment#16jibo272025-08-08 08:38Question about the runtime of main experiments in the RLPR paper
- Issue comment#12Leon-Francis2025-08-07 09:22Question on max_length computation in ray_trainer.py
- Issue comment#16hcx-stu2025-08-06 10:12Question about the runtime of main experiments in the RLPR paper
- Issue#14yiranyyu2025-07-31 10:27请问训练过程中保存的模型参数怎么使用vllm进行解码预测?
- Issue#14maydaygmail2025-07-28 15:27请问训练过程中保存的模型参数怎么使用vllm进行解码预测?
- Issue#13CaptainEven2025-07-25 03:17能够用于VLM模型的后训练微调吗?需要修改、适配的多吗?
- Issue#12JefferyChen4532025-07-24 05:41Question on max_length computation in ray_trainer.py
- Issue comment#12JefferyChen4532025-07-22 09:14Question on max_length computation in ray_trainer.py
- Issue#12JefferyChen4532025-07-22 06:14Question on max_length computation in ray_trainer.py
- Issue comment#10JefferyChen4532025-07-16 08:34Question on reward debiasing
- Issue#10JefferyChen4532025-07-16 08:34Question on reward debiasing
- Issue comment#10yiranyyu2025-07-16 08:29Question on reward debiasing
- Issue comment#11banjiuyufen2025-07-16 08:25Regarding the adaptation issue of Qwen3
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 103 stars here means stars gained during the window, not the repo's star count.