Skip to content

huggingface/evaluation-guidebook

View on GitHub ↗Related repositories →

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

active 2025-03-042026-06-29 (UTC)

Partial coverage11,541 / 12,628 hourly files (91%) · 2 absent upstream · 1,083 failed, retryable2025-03-022026-08-10 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
481
Pushes
3
Pull requests
2
Issues
3
Stars
455
Forks
15

Activity over time

Daily event counts in the loaded window

Line chart, 483 days from 2025-03-04 to 2026-06-29. Pushes: 3 total, peak 2 in a day. Pull requests: 2 total, peak 1 in a day. Issues: 3 total, peak 1 in a day. Comments: 3 total, peak 1 in a day. Stars: 455 total, peak 41 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
clefourrier5301
kerker7772020
fabthebest1000
heuristicwave1001
derekszen1000
kofboy20001001

Recent activity

Latest issues, pull requests and releases

  • Issue#40fabthebest2026-03-02 07:13
    [Contribution Proposal] New section: Evaluating LLMs on low-resource languages and non-Western cultural contexts
  • Pull request#37kerker7772025-11-19 01:27
  • Pull request#36kerker7772025-11-18 15:26
  • Issue#35derekszen2025-10-19 04:41
    VLM evals?
  • Issue comment#14clefourrier2025-09-18 08:31
    [TOPIC] How to design a good benchmark depending on your eval goals
  • Issue#14clefourrier2025-09-18 08:31
    [TOPIC] How to design a good benchmark depending on your eval goals
  • Issue comment#34heuristicwave2025-05-06 02:47
    Add Korean translation
  • Issue comment#34kofboy20002025-05-01 05:23
    Add Korean translation

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 455 stars here means stars gained during the window, not the repo's star count.