Code for "Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate"
active 2025-01-29 → 2026-02-08 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 376 days from 2025-01-29 to 2026-02-08. Pushes: 64 total, peak 36 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 13 total, peak 2 in a day. Comments: 15 total, peak 4 in a day. Stars: 162 total, peak 21 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
Recent activity
Latest issues, pull requests and releases
- Issue comment#10yyyhz2025-06-10 08:16Unable to reproduce results
- Issue comment#9yucc-leon2025-06-06 03:02Really interesting! But why this worked?
- Issue comment#9wenhuchen2025-06-05 19:50Really interesting! But why this worked?
- Issue#7Wyyyb2025-06-04 21:04CFT data of MetaMath & NuminaMath
- Issue comment#7Wyyyb2025-06-04 21:04CFT data of MetaMath & NuminaMath
- Issue comment#8Wyyyb2025-06-04 20:57base model with chat template
- Issue#8yangzhch62025-05-13 05:53base model with chat template
- Issue#7yangzhch62025-05-08 10:17CFT data of MetaMath & NuminaMath
- Issue#5Wyyyb2025-04-15 20:33The model Qwen2.5-Math-7B without fine-tuning cannot reproduce the ratings in the paper
- Issue comment#5Wyyyb2025-04-09 15:09The model Qwen2.5-Math-7B without fine-tuning cannot reproduce the ratings in the paper
- Issue#6Wyyyb2025-04-09 15:08Unfair results due to using MATH as the validation set
- Issue comment#6wenhuchen2025-03-14 03:30Unfair results due to using MATH as the validation set
- Issue comment#6wenhuchen2025-03-14 03:17Unfair results due to using MATH as the validation set
- Issue comment#6wenhuchen2025-03-14 03:01Unfair results due to using MATH as the validation set
- Issue comment#6wenhuchen2025-03-14 02:52Unfair results due to using MATH as the validation set
- Issue#4wenhuchen2025-03-11 13:27about inference
- Issue#4zhuyglx2025-03-11 05:55about inference
- Issue#3lijianshe022025-02-22 05:03critique loss function
- Issue#3lijianshe022025-02-22 03:58critique loss function
- Issue comment#2DripNowhy2025-02-11 11:52Cannot reproduce
- Issue#2DripNowhy2025-02-11 11:51Cannot reproduce
- Issue comment#2Wyyyb2025-02-11 09:55Cannot reproduce
- Issue comment#2DripNowhy2025-02-11 04:59Cannot reproduce
- Issue comment#2wenhuchen2025-02-11 04:42Cannot reproduce
- Issue#2DripNowhy2025-02-11 04:06Cannot reproduce
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 162 stars here means stars gained during the window, not the repo's star count.