AISBench Benchmark is a model evaluation tool built on OpenCompass, compatible with OpenCompass’s configuration system, dataset structure, and model backend implementation, while extending support for service-based models.
active 2025-11-12 → 2026-08-14 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 276 days from 2025-11-12 to 2026-08-14. Pushes: 87 total, peak 5 in a day. Pull requests: 125 total, peak 9 in a day. Issues: 120 total, peak 5 in a day. Comments: 546 total, peak 57 in a day. Stars: 30 total, peak 3 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| gemini-code-assist[bot] | 219 | 0 | 0 | 176 |
| SJTUyh | 179 | 47 | 35 | 45 |
| Copilot | 169 | 1 | 1 | 116 |
| GaoHuaZhang | 140 | 23 | 28 | 46 |
| github-actions[bot] | 123 | 0 | 30 | 56 |
| wenba0 | 112 | 3 | 9 | 44 |
| Keithwwa | 41 | 11 | 2 | 13 |
| zhongzhouTan-coder | 40 | 1 | 6 | 14 |
| Libotry | 22 | 1 | 5 | 7 |
| ivanbao9783 | 8 | 0 | 3 | 1 |
| ivan-zouming | 7 | 0 | 0 | 3 |
| SoulPainter-zhang | 5 | 0 | 2 | 0 |
| Donyzai | 5 | 0 | 0 | 5 |
| Liccol | 4 | 0 | 0 | 1 |
| muqing-li | 4 | 0 | 0 | 3 |
| zhangmuzhibangde | 4 | 0 | 0 | 0 |
| Dawn952 | 3 | 0 | 0 | 1 |
| shenchuxiaofugui | 3 | 0 | 0 | 3 |
| yuanhechen | 2 | 0 | 0 | 2 |
| liuyang-2026 | 2 | 0 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue#466lxr882026-08-14 08:22[文档] MMLU-Pro数据集文档中,mmlu_pro_gen_5_shot_str任务的few-shot值错误
- Issue#460jschen0692026-08-12 11:41【需求】【资料】【AISBench】支持数据集geometry3k精度测评
- Pull request#400ivanbao97832026-07-15 09:02
- Issue comment#387github-actions[bot]2026-07-02 03:16[Bug] aisbench后台多进程的情况下会在Summarizing performance results...时卡死
- Issue#375SJTUyh2026-06-30 03:10[疑问] AisBench和Evalscope在性能测试指标计算上有什么差异吗?相同的服务配置,测试相同的场景,性能结果差异较大。
- Issue#380github-actions[bot]2026-06-30 02:20[需求] 优化SWE-Bench_Pro数据集资源清理逻辑
- Issue#380ivanbao97832026-06-30 02:20[需求] 优化SWE-Bench_Pro数据集资源清理逻辑
- Issue#380ivanbao97832026-06-30 02:20[需求] 优化SWE-Bench_Pro数据集资源清理逻辑
- Issue#220SJTUyh2026-06-27 06:17[Bug] longbenchv2精度测评无结果输出
- Pull request#373ivanbao97832026-06-27 01:21
- Issue#236zhongzhouTan-coder2026-06-26 02:17[Bug] GLM5-W4A8性能测试
- Issue comment#233zhongzhouTan-coder2026-06-26 02:16[Bug] 测试gpqa数据集,发现有些预测结果为空,但是通过网关抓包,发现模型是返回了数据了的
- Issue#125wenba02026-06-26 01:42[Bug] mmlu_pro数据集评测精度结果都是0
- Issue#125wenba02026-06-26 01:42[Bug] mmlu_pro数据集评测精度结果都是0
- Issue#125wenba02026-06-26 01:42[Bug] mmlu_pro数据集评测精度结果都是0
- Issue comment#233github-actions[bot]2026-06-26 01:38[Bug] 测试gpqa数据集,发现有些预测结果为空,但是通过网关抓包,发现模型是返回了数据了的
- Issue#227SJTUyh2026-06-25 10:54[Bug] longbench跑性能测试时数据翻倍了
- Issue comment#347F0undLinks2026-06-18 01:58Update answer pattern for COT chat prompt
- Issue#342zhongzhouTan-coder2026-06-15 02:06[Bug] SWE BENCH 使用 mini-swe-agent 测试一台机器同时跑两台任务会互相干扰
- Pull request#335github-actions[bot]2026-06-11 01:26
- Pull request#335github-actions[bot]2026-06-11 01:26
- Pull request#335github-actions[bot]2026-06-11 01:26
- Issue#311zhongzhouTan-coder2026-05-29 01:18[Bug] vllm_api_general_chat.py配置中URL末尾未加/导致访问地址不对
- Issue#311zhongzhouTan-coder2026-05-29 01:18[Bug] vllm_api_general_chat.py配置中URL末尾未加/导致访问地址不对
- Pull request#312Copilot2026-05-28 11:05
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 30 stars here means stars gained during the window, not the repo's star count.