Skip to content

AISBench Benchmark is a model evaluation tool built on OpenCompass, compatible with OpenCompass’s configuration system, dataset structure, and model backend implementation, while extending support for service-based models.

active 2025-11-122026-08-14 (UTC)

Complete coverage26,322 / 26,322 hourly files (100%) · 2 absent upstream2023-08-152026-08-15 (UTC)
Events
1.2K
Pushes
87
Pull requests
125
Issues
120
Stars
30
Forks
6

Activity over time

Daily event counts in the loaded window

Line chart, 276 days from 2025-11-12 to 2026-08-14. Pushes: 87 total, peak 5 in a day. Pull requests: 125 total, peak 9 in a day. Issues: 120 total, peak 5 in a day. Comments: 546 total, peak 57 in a day. Stars: 30 total, peak 3 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
gemini-code-assist[bot]21900176
SJTUyh179473545
Copilot16911116
GaoHuaZhang140232846
github-actions[bot]12303056
wenba01123944
Keithwwa4111213
zhongzhouTan-coder401614
Libotry22157
ivanbao97838031
ivan-zouming7003
SoulPainter-zhang5020
Donyzai5005
Liccol4001
muqing-li4003
zhangmuzhibangde4000
Dawn9523001
shenchuxiaofugui3003
yuanhechen2002
liuyang-20262000

Recent activity

Latest issues, pull requests and releases

  • Issue#466lxr882026-08-14 08:22
    [文档] MMLU-Pro数据集文档中,mmlu_pro_gen_5_shot_str任务的few-shot值错误
  • Issue#460jschen0692026-08-12 11:41
    【需求】【资料】【AISBench】支持数据集geometry3k精度测评
  • Pull request#400ivanbao97832026-07-15 09:02
  • Issue comment#387github-actions[bot]2026-07-02 03:16
    [Bug] aisbench后台多进程的情况下会在Summarizing performance results...时卡死
  • Issue#375SJTUyh2026-06-30 03:10
    [疑问] AisBench和Evalscope在性能测试指标计算上有什么差异吗?相同的服务配置,测试相同的场景,性能结果差异较大。
  • Issue#380github-actions[bot]2026-06-30 02:20
    [需求] 优化SWE-Bench_Pro数据集资源清理逻辑
  • Issue#380ivanbao97832026-06-30 02:20
    [需求] 优化SWE-Bench_Pro数据集资源清理逻辑
  • Issue#380ivanbao97832026-06-30 02:20
    [需求] 优化SWE-Bench_Pro数据集资源清理逻辑
  • Issue#220SJTUyh2026-06-27 06:17
    [Bug] longbenchv2精度测评无结果输出
  • Pull request#373ivanbao97832026-06-27 01:21
  • Issue#236zhongzhouTan-coder2026-06-26 02:17
    [Bug] GLM5-W4A8性能测试
  • Issue comment#233zhongzhouTan-coder2026-06-26 02:16
    [Bug] 测试gpqa数据集,发现有些预测结果为空,但是通过网关抓包,发现模型是返回了数据了的
  • Issue#125wenba02026-06-26 01:42
    [Bug] mmlu_pro数据集评测精度结果都是0
  • Issue#125wenba02026-06-26 01:42
    [Bug] mmlu_pro数据集评测精度结果都是0
  • Issue#125wenba02026-06-26 01:42
    [Bug] mmlu_pro数据集评测精度结果都是0
  • Issue comment#233github-actions[bot]2026-06-26 01:38
    [Bug] 测试gpqa数据集,发现有些预测结果为空,但是通过网关抓包,发现模型是返回了数据了的
  • Issue#227SJTUyh2026-06-25 10:54
    [Bug] longbench跑性能测试时数据翻倍了
  • Issue comment#347F0undLinks2026-06-18 01:58
    Update answer pattern for COT chat prompt
  • Issue#342zhongzhouTan-coder2026-06-15 02:06
    [Bug] SWE BENCH 使用 mini-swe-agent 测试一台机器同时跑两台任务会互相干扰
  • Pull request#335github-actions[bot]2026-06-11 01:26
  • Pull request#335github-actions[bot]2026-06-11 01:26
  • Pull request#335github-actions[bot]2026-06-11 01:26
  • Issue#311zhongzhouTan-coder2026-05-29 01:18
    [Bug] vllm_api_general_chat.py配置中URL末尾未加/导致访问地址不对
  • Issue#311zhongzhouTan-coder2026-05-29 01:18
    [Bug] vllm_api_general_chat.py配置中URL末尾未加/导致访问地址不对
  • Pull request#312Copilot2026-05-28 11:05

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 30 stars here means stars gained during the window, not the repo's star count.