Skip to content

Data and code for L-Eval, a comprehensive long context language models evaluation benchmark

active 2023-10-182026-03-22 (UTC)

Partial coverage19,249 / 24,809 hourly files (78%) · 2 absent upstream · 5,558 failed, retryable2023-10-142026-08-12 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
202
Pushes
25
Pull requests
0
Issues
14
Stars
139
Forks
1

Activity over time

Daily event counts in the loaded window

Line chart, 887 days from 2023-10-18 to 2026-03-22. Pushes: 25 total, peak 7 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 14 total, peak 3 in a day. Comments: 23 total, peak 3 in a day. Stars: 139 total, peak 3 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
ChenxinAn-fdu4125014
coo00ookie4002
chunniunai220ml4003
zhimin-z3000
cizhenshi3001
ylsung2002
Ocean-6272001
sheryc1000
ogencoglu1000
Desdemonaqipi1000

Recent activity

Latest issues, pull requests and releases

  • Issue comment#19ChenxinAn-fdu2024-08-13 04:49
    hotpotwikiqa_mixup、multifieldqa_zh_mixup等都没有没有answer_keywords
  • Issue#19ChenxinAn-fdu2024-08-13 04:49
    hotpotwikiqa_mixup、multifieldqa_zh_mixup等都没有没有answer_keywords
  • Issue#19Desdemonaqipi2024-08-13 03:00
    hotpotwikiqa_mixup、multifieldqa_zh_mixup等都没有没有answer_keywords
  • Issue comment#18ChenxinAn-fdu2024-08-11 04:32
    Clarification regarding multidoc_qa
  • Issue#18ogencoglu2024-08-10 20:33
    Clarification regarding multidoc_qa
  • Issue comment#14ylsung2024-07-19 19:33
    failed reproduce llama3-8b result
  • Issue comment#14ChenxinAn-fdu2024-07-18 06:27
    failed reproduce llama3-8b result
  • Issue comment#14ylsung2024-07-18 06:12
    failed reproduce llama3-8b result
  • Issue comment#16ChenxinAn-fdu2024-07-03 02:53
    Question about the leaderboard
  • Issue comment#16ChenxinAn-fdu2024-07-03 02:52
    Question about the leaderboard
  • Issue comment#16cizhenshi2024-07-03 02:52
    Question about the leaderboard
  • Issue#16cizhenshi2024-07-03 02:51
    Question about the leaderboard
  • Issue#16cizhenshi2024-07-03 02:39
    Question about the leaderboard
  • Issue comment#14chunniunai220ml2024-06-17 04:16
    failed reproduce llama3-8b result
  • Issue comment#14ChenxinAn-fdu2024-06-17 01:51
    failed reproduce llama3-8b result
  • Issue#14chunniunai220ml2024-06-17 01:01
    failed reproduce llama3-8b result
  • Issue comment#13chunniunai220ml2024-06-16 08:13
    How to Reproduce Results on Llama3-8b?
  • Issue comment#13chunniunai220ml2024-06-15 14:14
    How to Reproduce Results on Llama3-8b?
  • Issue#13Ocean-6272024-06-08 09:15
    How to Reproduce Results on Llama3-8b?
  • Issue comment#13Ocean-6272024-06-08 09:14
    How to Reproduce Results on Llama3-8b?
  • Issue comment#13ChenxinAn-fdu2024-06-06 02:18
    How to Reproduce Results on Llama3-8b?
  • Issue#11sheryc2024-01-18 06:17
    Problems with the sci_fi evaluation
  • Issue comment#11ChenxinAn-fdu2024-01-18 04:20
    Problems with the sci_fi evaluation
  • Issue comment#11ChenxinAn-fdu2024-01-18 04:05
    Problems with the sci_fi evaluation
  • Issue#10zhimin-z2023-12-12 02:39
    Except GSM100, other datasets are evaluated in 0-shot?

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 139 stars here means stars gained during the window, not the repo's star count.