Skip to content

OpenCompass is an LLM evaluation platform, supporting a wide range of models (LLaMA, LLaMa2, ChatGLM2, ChatGPT, Claude, etc) over 50+ datasets.

active 2023-08-152023-08-29 (UTC)

Partial coverage20,502 / 26,258 hourly files (78%) · 2 absent upstream · 5,753 failed, retryable2023-08-152026-08-13 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
581
Pushes
44
Pull requests
74
Issues
50
Stars
99
Forks
14

Activity over time

Daily event counts in the loaded window

Line chart, 15 days from 2023-08-15 to 2023-08-29. Pushes: 44 total, peak 11 in a day. Pull requests: 74 total, peak 14 in a day. Issues: 50 total, peak 13 in a day. Comments: 161 total, peak 39 in a day. Stars: 99 total, peak 11 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
gaotongxiao131212934
YuanLiuuuuuu87131228
Leymore365814
tonysy291213
yingfhu24367
wellcasa230015
yyk-wew210311
Ezra-Yu170211
Sweetclover13007
fangyixiao1811025
liushz8122
KaiLv695001
jodie22353374002
simonjoe2464001
zycheiheihei3001
cdpath3021
Quehry3001
tiansiyuan2000
franztao2001
Luodian2011

Recent activity

Latest issues, pull requests and releases

  • Issue comment#312yyk-wew2023-08-29 12:19
    [Feat] Support Qwen-VL-Chat on MMBench.
  • Pull request#312yyk-wew2023-08-29 11:12
  • Issue#283Leymore2023-08-29 10:08
    [Bug] Eval HumanEval dataset when k > 1 in pass@k
  • Issue comment#283Leymore2023-08-29 10:08
    [Bug] Eval HumanEval dataset when k > 1 in pass@k
  • Issue#283KaiLv692023-08-29 08:04
    [Bug] Eval HumanEval dataset when k > 1 in pass@k
  • Issue comment#283Leymore2023-08-29 07:32
    [Bug] Eval HumanEval dataset when k > 1 in pass@k
  • Issue comment#284Leymore2023-08-29 07:29
    [Bug] TypeError: cross_entropy_loss(): argument 'ignore_index' (position 5) must be int, not NoneType
  • Pull request#287Leymore2023-08-29 07:19
  • Issue comment#286Leymore2023-08-29 05:10
    [Feature] Add qwen & qwen-chat support
  • Issue comment#285tonysy2023-08-29 03:46
    [Feature] Support llava on seed-bench, mme and more multi-modal benchmarks.
  • Issue#281zycheiheihei2023-08-29 03:36
    [Feature] Support of InstructBlip on MME Dataset
  • Issue comment#281zycheiheihei2023-08-29 03:36
    [Feature] Support of InstructBlip on MME Dataset
  • Issue comment#281yyk-wew2023-08-29 03:29
    [Feature] Support of InstructBlip on MME Dataset
  • Issue#275tonysy2023-08-29 03:13
    [Bug] RuntimeError: "addmm_impl_cpu_" not implemented for 'Half'
  • Issue comment#275tonysy2023-08-29 03:13
    [Bug] RuntimeError: "addmm_impl_cpu_" not implemented for 'Half'
  • Issue#234tonysy2023-08-29 03:11
    race的评测分数异常[Bug]
  • Issue comment#234tonysy2023-08-29 03:11
    race的评测分数异常[Bug]
  • Issue#284simonjoe2462023-08-28 09:48
    [Bug] TypeError: cross_entropy_loss(): argument 'ignore_index' (position 5) must be int, not NoneType
  • Pull request#280gaotongxiao2023-08-28 09:35
  • Issue comment#283Leymore2023-08-28 08:18
    [Bug] Eval HumanEval dataset when k > 1 in pass@k
  • Issue#283KaiLv692023-08-28 08:17
    [Bug] Eval HumanEval dataset when k > 1 in pass@k
  • Issue comment#275gaotongxiao2023-08-28 06:53
    [Bug] RuntimeError: "addmm_impl_cpu_" not implemented for 'Half'
  • Issue#281zycheiheihei2023-08-28 05:44
    [Feature] Support of InstructBlip on MME Dataset
  • Pull request#280Leymore2023-08-28 04:20
  • Issue comment#277gaotongxiao2023-08-28 03:33
    [WIP] Support GSM8k evaluation with tools by Lagent and LangChain

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 99 stars here means stars gained during the window, not the repo's star count.