Skip to content

简化且可自定义的框架,用于有效的大型模型评估和性能基准测试

Python · active 2024-07-292026-08-27 (UTC)

Complete coverage26,686 / 26,686 hourly files (100%) · 2 absent upstream2023-08-152026-08-30 (UTC)
Events
7.6K
Pushes
1.2K
Pull requests
539
Issues
934
Stars
1.6K
Forks
180

Activity over time

Daily event counts in the loaded window

Line chart, 760 days from 2024-07-29 to 2026-08-27. Pushes: 1,167 total, peak 18 in a day. Pull requests: 539 total, peak 11 in a day. Issues: 934 total, peak 22 in a day. Comments: 2,366 total, peak 43 in a day. Stars: 1,569 total, peak 21 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
Yunnglin2.7K9253591.1K
wangxingjun77849422411083
gemini-code-assist[bot]36900268
Copilot17000107
jackqdldd790054
yingdachen371125
charliedream1220015
ZHAOFEGNSHUN190012
winni0190014
maidouxiaozi170014
Haruka130716007
LeoCeasar160015
Bigfishering160010
XYZliang150013
Devliang2415008
shell-nlp15002
copilot-pull-request-reviewer[bot]14002
simonqian13009
penguinwang9682513037
ShaohonChen12039

Recent activity

Latest issues, pull requests and releases

  • Pull request#1616Yunnglin2026-08-21 16:10
  • Pull request#1616Yunnglin2026-08-21 10:13
  • Issue comment#1415Yunnglin2026-06-16 02:27
    Run evalscope with --datasets swe_bench_verified_mini, i got error.
  • Issue comment#1405Yunnglin2026-06-11 09:52
    fix: add retry mechanism for jsonl file writing
  • Pull request#1402haoruilee2026-06-08 01:49
  • Issue#766Yunnglin2026-06-02 09:04
    在evalscope如何中快速自定义RAGAS其他测评指标,想要自定义更多指标
  • Issue#1378wenruihua2026-05-28 02:41
    容器内评测terminal_bench_v2数据集,创建容器时添加了-v /var/run/docker.sock:/var/run/docker.sock:ro ,但是还是有报错 No such file or directory: 'docker'
  • Issue comment#1368wenruihua2026-05-26 01:42
    evalscope测试minimax2.5在swe_bench_verified数据集上的精度,每次老是有一个输入长度过了模型的最大上下文长度,导致测不出来具体的精度值。
  • Issue#1361zcr19971102026-05-25 09:03
    评测HLE测试集,结果都是0
  • Issue#1369fate083010172026-05-25 02:19
    希望支持bigcodebench数据集的评测支持
  • Issue comment#1356Yunnglin2026-05-22 07:14
    容器内评测swe_bench_verified_mini 数据集报错ConnectionRefusedError: [Errno 111] Connection refused
  • Issue comment#1313Yunnglin2026-05-18 14:41
    关于对视频的支持
  • Issue comment#1150Yunnglin2026-05-18 14:41
    评测base模型
  • Pull request#1346Yunnglin2026-05-18 03:07
  • Pull request#1345Yunnglin2026-05-13 04:35
  • Pull request#1345Yunnglin2026-05-13 04:31
  • Issue#1340llc-kc2026-05-12 04:46
    性能压测自定义数据集能否支持openai数据格式
  • Issue comment#1338Yunnglin2026-05-12 03:08
    swe-smith预构建数据集的意义
  • Issue comment#1331Yunnglin2026-05-11 03:35
    FGA-BLIP2 强绑定了 CUDA,希望能根据实际设备更灵活地运行
  • Pull request#1336Yunnglin2026-05-11 03:19
  • Issue comment#1328Yunnglin2026-05-11 02:04
    如何使用 EvalScope 测试 RPM
  • Pull request#1329Yunnglin2026-05-09 08:07
  • Pull request#1326Yunnglin2026-05-07 04:40
  • Issue#1312twilighgt2026-04-27 02:53
    【tau2-bench】用例个数与官方对不上 && retail领域数据偏低
  • Pull request#1309Yunnglin2026-04-23 05:02

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 1,569 stars here means stars gained during the window, not the repo's star count.