评估大模型的框架[ACL 2024 Demo] Official GitHub repo for UltraEval: An open source framework for evaluating foundation models.
active 2023-11-23 → 2026-02-08 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 809 days from 2023-11-23 to 2026-02-08. Pushes: 35 total, peak 6 in a day. Pull requests: 21 total, peak 6 in a day. Issues: 26 total, peak 4 in a day. Comments: 35 total, peak 13 in a day. Stars: 237 total, peak 13 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| Hothan01 | 54 | 25 | 6 | 12 |
| R0k1e | 14 | 0 | 0 | 14 |
| max-yue | 12 | 7 | 5 | 0 |
| Ranger-Los | 10 | 2 | 8 | 0 |
| zhsky2017 | 4 | 0 | 0 | 3 |
| kijlk | 3 | 0 | 0 | 1 |
| fan2goa1 | 3 | 0 | 0 | 1 |
| renjie-ranger | 3 | 1 | 2 | 0 |
| nkfnn | 2 | 0 | 0 | 1 |
| AIR-hl | 2 | 0 | 0 | 1 |
| stepbystepcode | 2 | 0 | 0 | 1 |
| SefaZeng | 1 | 0 | 0 | 0 |
| ericxsun | 1 | 0 | 0 | 0 |
| xh-yuan | 1 | 0 | 0 | 0 |
| Ming0310 | 1 | 0 | 0 | 1 |
| ExcitingYi | 1 | 0 | 0 | 0 |
| Abigail61 | 1 | 0 | 0 | 0 |
| xinformatics | 1 | 0 | 0 | 0 |
| Hongcheng-Gao | 1 | 0 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue comment#23AIR-hl2024-10-11 03:57TypeError: unsupported operand type(s) for +: 'int' and 'dict'
- Issue comment#23Ming03102024-10-11 03:23TypeError: unsupported operand type(s) for +: 'int' and 'dict'
- Issue#25Hothan012024-10-09 03:09Raw dataset 404 not found
- Issue#18Hothan012024-09-17 15:46PIQA等QA类任务测评结果差
- Issue#17Hothan012024-09-17 15:46HumanEval的评测结果与官方HumanEval的评测结果不同
- Issue#8Hothan012024-09-17 15:46加载OpenBMB/MiniCPM-2B-dpo-fp32失败
- Issue#20Hothan012024-09-17 15:45可以在不安装 vllm 的情况下使用吗?
- Issue comment#24Hothan012024-09-17 15:45How to evaluate models on my own custom MCQA dataset?
- Issue#25ExcitingYi2024-09-15 15:31Raw dataset 404 not found
- Issue#24xinformatics2024-09-10 09:58How to evaluate models on my own custom MCQA dataset?
- Issue#23AIR-hl2024-08-03 13:36TypeError: unsupported operand type(s) for +: 'int' and 'dict'
- Issue#22xh-yuan2024-07-08 12:26ProSparse can not be reproduced
- Issue#21Abigail612024-06-14 08:27minicpm-2b-sft-bf16在gsm8k的复现结果与论文不一致
- Issue comment#20Hothan012024-05-24 06:49可以在不安装 vllm 的情况下使用吗?
- Issue#20SefaZeng2024-05-21 02:26可以在不安装 vllm 的情况下使用吗?
- Issue comment#19fan2goa12024-04-25 20:07What should I write if I want to cite this work?
- Issue#19fan2goa12024-04-25 20:07What should I write if I want to cite this work?
- Issue comment#19Hothan012024-04-24 02:22What should I write if I want to cite this work?
- Issue#19fan2goa12024-04-23 21:46What should I write if I want to cite this work?
- Issue comment#18zhsky20172024-04-17 04:21PIQA等QA类任务测评结果差
- Issue comment#18zhsky20172024-04-17 03:52PIQA等QA类任务测评结果差
- Issue comment#18zhsky20172024-04-17 03:50PIQA等QA类任务测评结果差
- Issue comment#18Hothan012024-04-17 03:27PIQA等QA类任务测评结果差
- Issue#18zhsky20172024-04-17 02:27PIQA等QA类任务测评结果差
- Issue comment#17Hothan012024-04-11 09:55HumanEval的评测结果与官方HumanEval的评测结果不同
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 237 stars here means stars gained during the window, not the repo's star count.