Skip to content

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

active 2026-01-102026-02-27 (UTC)

Complete coverage26,474 / 26,474 hourly files (100%) · 2 absent upstream2023-08-152026-08-22 (UTC)
Events
91
Pushes
33
Pull requests
18
Issues
20
Stars
0
Forks
0

Activity over time

Daily event counts in the loaded window

Line chart, 49 days from 2026-01-10 to 2026-02-27. Pushes: 33 total, peak 10 in a day. Pull requests: 18 total, peak 7 in a day. Issues: 20 total, peak 7 in a day. Comments: 0 total, peak 0 in a day. Stars: 0 total, peak 0 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
pjavanrood4625100
Raywnh25880

Recent activity

Latest issues, pull requests and releases

  • Issue#57Raywnh2026-01-27 05:27
    [BUG] Registry incorrect repo searching
  • Issue#57Raywnh2026-01-27 05:27
    [BUG] Registry incorrect repo searching
  • Pull request#42Raywnh2026-01-25 20:54
  • Pull request#44Raywnh2026-01-25 20:54
  • Issue#45Raywnh2026-01-25 20:53
    med_qa benchmark fails: Dataset scripts no longer supported
  • Pull request#46Raywnh2026-01-25 20:53
  • Issue#53Raywnh2026-01-25 01:39
    Arithmetic benchmark fails: Dataset scripts no longer supported
  • Pull request#52Raywnh2026-01-25 01:37
  • Pull request#50Raywnh2026-01-25 01:36
  • Issue#45Raywnh2026-01-25 01:33
    med_qa benchmark fails: Dataset scripts no longer supported
  • Pull request#44Raywnh2026-01-25 01:31
  • Issue#43Raywnh2026-01-25 01:31
    SIQA benchmark fails: Dataset scripts no longer supported
  • Pull request#42Raywnh2026-01-25 01:30
  • Pull request#40Raywnh2026-01-24 22:28
  • Issue#39Raywnh2026-01-24 22:27
    [Benchmark] Fix tinyGSM8k: add bounds checking in compute_corpus
  • Issue#37Raywnh2026-01-24 21:58
    [Benchmark] Fix real_toxicity_prompts: remove incorrect exact_match metric
  • Issue#7Raywnh2026-01-24 20:33
    Fix CoQA metric and multi-doc loading
  • Issue#33pjavanrood2026-01-23 00:53
    Fix KeyError in truthful_qa_generative_prompt
  • Issue#27pjavanrood2026-01-23 00:52
    Fix subset names in StoryCloze
  • Pull request#28pjavanrood2026-01-23 00:52
  • Issue#25pjavanrood2026-01-23 00:52
    Fix column mismatch and metric in SimpleQA
  • Issue#23pjavanrood2026-01-23 00:52
    Fix TypeError in real_toxicity_prompts due to None choices
  • Issue#21pjavanrood2026-01-23 00:52
    Fix key mismatch and context access in PubMedQA
  • Pull request#22pjavanrood2026-01-23 00:52
  • Issue#13pjavanrood2026-01-23 00:51
    Fix non-existent evaluation splits in lextreme

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.