Skip to content

This repo contains the code for "MEGA-Bench Scaling Multimodal Evaluation to over 500 Real-World Tasks" [ICLR2025]

active 2024-10-112025-09-11 (UTC)

Complete coverage26,587 / 26,587 hourly files (100%) · 2 absent upstream2023-08-152026-08-26 (UTC)
Events
238
Pushes
113
Pull requests
0
Issues
12
Stars
71
Forks
9

Activity over time

Daily event counts in the loaded window

Line chart, 336 days from 2024-10-11 to 2025-09-11. Pushes: 113 total, peak 13 in a day. Pull requests: 0 total, peak 0 in a day. Issues: 12 total, peak 1 in a day. Comments: 31 total, peak 4 in a day. Stars: 71 total, peak 7 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
woodfrog8365017
eric-zqwang262600
wenhuchen171601
Violettttee14008
TianhaoLiang20006600
shipengai4003
zjwu05222001
wizyoung2001
Luodian1000
Liuziyu771000

Recent activity

Latest issues, pull requests and releases

  • Issue comment#10woodfrog2025-07-07 07:56
    应该看macro_mean还是micro_mean
  • Issue comment#10shipengai2025-06-23 02:11
    应该看macro_mean还是micro_mean
  • Issue comment#10shipengai2025-06-20 11:08
    应该看macro_mean还是micro_mean
  • Issue comment#10shipengai2025-06-20 11:02
    应该看macro_mean还是micro_mean
  • Issue comment#10woodfrog2025-06-20 09:50
    应该看macro_mean还是micro_mean
  • Issue#10shipengai2025-06-20 08:01
    应该看macro_mean还是micro_mean
  • Issue#2woodfrog2025-03-29 19:21
    Integrations into popular evaluation frameworks like lmms_eval or vlmevalkit
  • Issue#9Liuziyu772025-01-20 19:18
    How to use GPT-4o for evaluation?
  • Issue comment#8woodfrog2025-01-03 21:18
    Question about answer parse
  • Issue comment#8Violettttee2025-01-03 12:17
    Question about answer parse
  • Issue comment#8woodfrog2025-01-03 12:11
    Question about answer parse
  • Issue#8Violettttee2025-01-03 09:19
    Question about answer parse
  • Issue comment#7Violettttee2024-12-27 09:08
    About "Application" category score
  • Issue comment#7woodfrog2024-12-27 06:28
    About "Application" category score
  • Issue#7Violettttee2024-12-27 05:20
    About "Application" category score
  • Issue comment#6woodfrog2024-12-26 23:14
    Request to evaluate O1 model series (Multimodality + API now available)
  • Issue comment#5woodfrog2024-12-16 17:42
    Wrong aspect_ratio param in llava-onevision evaluation
  • Issue comment#5woodfrog2024-12-16 17:38
    Wrong aspect_ratio param in llava-onevision evaluation
  • Issue#5Luodian2024-12-16 09:38
    Wrong aspect_ratio param in llava-onevision evaluation
  • Issue comment#4woodfrog2024-11-23 09:57
    About subsets evaluation
  • Issue comment#4Violettttee2024-11-21 07:34
    About subsets evaluation
  • Issue comment#4woodfrog2024-11-10 04:15
    About subsets evaluation
  • Issue comment#4Violettttee2024-11-09 10:41
    About subsets evaluation
  • Issue#4Violettttee2024-11-09 10:41
    About subsets evaluation
  • Issue comment#4woodfrog2024-11-08 22:53
    About subsets evaluation

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 71 stars here means stars gained during the window, not the repo's star count.