Skip to content

The scripts for MMLU-Pro, using a smaller IRT-tuned dataset

active 2024-05-182026-05-19 (UTC)

Complete coverage26,585 / 26,585 hourly files (100%) · 2 absent upstream2023-08-152026-08-26 (UTC)
Events
664
Pushes
73
Pull requests
24
Issues
114
Stars
294
Forks
44

Activity over time

Daily event counts in the loaded window

Line chart, 732 days from 2024-05-18 to 2026-05-19. Pushes: 73 total, peak 7 in a day. Pull requests: 24 total, peak 4 in a day. Issues: 114 total, peak 4 in a day. Comments: 114 total, peak 8 in a day. Stars: 294 total, peak 4 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
Wyyyb11244638
wenhuchen7427126
EwoutH380014
chigkim160211
billbradley9006
ubergarm5003
NSbuilder5001
tech-jun-jones5041
eldarkurtic4003
johns2s4001
Vyryn4030
mujtabaasif3030
RodriMora3001
TheTinyTeddy3002
MXueguang3201
gnalbandyan2000
emanuelevivoli2001
mrconter12000
L1aoXingyu2020
jing-P2000

Recent activity

Latest issues, pull requests and releases

  • Issue#79Linux-Server2026-05-13 11:31
    Benchmark Result Mismatch
  • Pull request#78lavdnone22026-03-17 18:21
  • Issue#76Qifeng-Wu992025-10-17 21:37
    Typo report
  • Issue comment#73TheTinyTeddy2025-08-29 09:30
    Reproducing the score of 0.9149 in the math category for GPT-OSS 20B on the Leaderboard
  • Issue comment#73MXueguang2025-08-22 03:28
    Reproducing the score of 0.9149 in the math category for GPT-OSS 20B on the Leaderboard
  • Issue comment#73wenhuchen2025-08-21 13:02
    Reproducing the score of 0.9149 in the math category for GPT-OSS 20B on the Leaderboard
  • Issue comment#73TheTinyTeddy2025-08-21 06:53
    Reproducing the score of 0.9149 in the math category for GPT-OSS 20B on the Leaderboard
  • Issue comment#73wenhuchen2025-08-21 05:29
    Reproducing the score of 0.9149 in the math category for GPT-OSS 20B on the Leaderboard
  • Issue#73TheTinyTeddy2025-08-21 01:40
    Reproducing the score of 0.9149 in the math category for GPT-OSS 20B on the Leaderboard
  • Issue#63Wyyyb2025-04-15 20:31
    Redundant line in evaluate_from_local.py
  • Issue#72Wyyyb2025-04-15 20:31
    Why are there repeated questions and answers?
  • Issue comment#21billbradley2025-03-31 12:07
    OpenAI o1-preview and o1-mini
  • Issue#72zhejunliux2025-03-25 07:28
    Why are there repeated questions and answers?
  • Issue comment#21billbradley2025-03-21 20:38
    OpenAI o1-preview and o1-mini
  • Issue#71Ulov8882025-03-19 12:07
    how to run 5-shot instead of 0-shot test?
  • Issue#71Ulov8882025-03-19 12:05
    how to run 5-shot instead of 0-shot test?
  • Issue#70jing-P2025-03-17 08:47
    the func: update_result in evaluate_from_api did not store model_answer
  • Issue#70jing-P2025-03-17 08:39
    the func: update_result in evaluate_from_api did not store model_answer
  • Issue#69EwoutH2025-03-17 07:46
    Validate small high-scoring models
  • Pull request#68L1aoXingyu2025-03-13 15:47
  • Pull request#68L1aoXingyu2025-03-13 15:47
  • Issue#64Wyyyb2025-03-03 05:06
    Random seed bug
  • Pull request#65Wyyyb2025-02-28 17:39
  • Pull request#66Wyyyb2025-02-28 17:39
  • Pull request#67Wyyyb2025-02-28 17:38

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 294 stars here means stars gained during the window, not the repo's star count.