Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
active 2026-01-10 → 2026-02-27 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 49 days from 2026-01-10 to 2026-02-27. Pushes: 33 total, peak 10 in a day. Pull requests: 18 total, peak 7 in a day. Issues: 20 total, peak 7 in a day. Comments: 0 total, peak 0 in a day. Stars: 0 total, peak 0 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| pjavanrood | 46 | 25 | 10 | 0 |
| Raywnh | 25 | 8 | 8 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue#57Raywnh2026-01-27 05:27[BUG] Registry incorrect repo searching
- Issue#57Raywnh2026-01-27 05:27[BUG] Registry incorrect repo searching
- Pull request#42Raywnh2026-01-25 20:54
- Pull request#44Raywnh2026-01-25 20:54
- Issue#45Raywnh2026-01-25 20:53med_qa benchmark fails: Dataset scripts no longer supported
- Pull request#46Raywnh2026-01-25 20:53
- Issue#53Raywnh2026-01-25 01:39Arithmetic benchmark fails: Dataset scripts no longer supported
- Pull request#52Raywnh2026-01-25 01:37
- Pull request#50Raywnh2026-01-25 01:36
- Issue#45Raywnh2026-01-25 01:33med_qa benchmark fails: Dataset scripts no longer supported
- Pull request#44Raywnh2026-01-25 01:31
- Issue#43Raywnh2026-01-25 01:31SIQA benchmark fails: Dataset scripts no longer supported
- Pull request#42Raywnh2026-01-25 01:30
- Pull request#40Raywnh2026-01-24 22:28
- Issue#39Raywnh2026-01-24 22:27[Benchmark] Fix tinyGSM8k: add bounds checking in compute_corpus
- Issue#37Raywnh2026-01-24 21:58[Benchmark] Fix real_toxicity_prompts: remove incorrect exact_match metric
- Issue#7Raywnh2026-01-24 20:33Fix CoQA metric and multi-doc loading
- Issue#33pjavanrood2026-01-23 00:53Fix KeyError in truthful_qa_generative_prompt
- Issue#27pjavanrood2026-01-23 00:52Fix subset names in StoryCloze
- Pull request#28pjavanrood2026-01-23 00:52
- Issue#25pjavanrood2026-01-23 00:52Fix column mismatch and metric in SimpleQA
- Issue#23pjavanrood2026-01-23 00:52Fix TypeError in real_toxicity_prompts due to None choices
- Issue#21pjavanrood2026-01-23 00:52Fix key mismatch and context access in PubMedQA
- Pull request#22pjavanrood2026-01-23 00:52
- Issue#13pjavanrood2026-01-23 00:51Fix non-existent evaluation splits in lextreme
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 0 stars here means stars gained during the window, not the repo's star count.