Skip to content

Official repo of VLABench, a large scale benchmark designed for fairly evaluating VLA, Embodied Agent, and VLMs.

Python · active 2025-02-192026-05-25 (UTC)

Partial coverage11,680 / 12,902 hourly files (91%) · 2 absent upstream · 1,219 failed, retryable2025-02-192026-08-10 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
297
Pushes
27
Pull requests
11
Issues
59
Stars
131
Forks
9

Activity over time

Daily event counts in the loaded window

Line chart, 461 days from 2025-02-19 to 2026-05-25. Pushes: 27 total, peak 4 in a day. Pull requests: 11 total, peak 5 in a day. Issues: 59 total, peak 7 in a day. Comments: 57 total, peak 10 in a day. Stars: 131 total, peak 7 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

Recent activity

Latest issues, pull requests and releases

  • Issue#87robosciyzq2026-04-23 06:05
    can not generate episode in many task
  • Issue#84JuilieZ2026-02-15 15:18
    each episodes faild when pi05_ft_vlabench_primitive
  • Issue comment#78YunOh212026-01-18 11:13
    How to know specific prompt used for results video
  • Issue comment#19nansun54102026-01-12 11:55
    Abs EEF Action Space or Abs Joint Action Space
  • Issue#78YunOh212025-12-28 17:00
    How to know specific prompt used for results video
  • Issue#76hemin08062025-12-02 05:23
    how to evaluate pi0-fast-ft-primitive-10task-deltachunk ckpt
  • Issue#75Xj15675893542025-11-21 01:45
    PDDL file writing issues for scenarios
  • Issue#74Saberlve2025-11-15 09:27
    How to evaluate other VLA model in VLABench?
  • Issue#73Jiang-Yuxuan1232025-11-12 07:46
    Data generation ends early
  • Issue#73Jiang-Yuxuan1232025-11-12 07:38
    Data generation ends early
  • Issue#73Jiang-Yuxuan1232025-11-12 07:35
    Data generation ends early
  • Issue comment#56Shiduo-zh2025-11-11 03:48
    The task video result is only one frame
  • Issue#44Shiduo-zh2025-11-11 03:41
    Question about evaluating models like VoxPoser and CoPa in VLABench
  • Issue comment#26Shiduo-zh2025-11-11 03:40
    Experimental reproduction of Table 2
  • Issue#24Shiduo-zh2025-11-11 03:39
    pi0 inference in VLABench's mujoco environment
  • Issue comment#24Shiduo-zh2025-11-11 03:38
    pi0 inference in VLABench's mujoco environment
  • Issue#22Shiduo-zh2025-11-11 03:37
    Six Benchmark results
  • Issue comment#69Shiduo-zh2025-11-11 03:36
    Inquiry about the release timeline for VLABench evaluation pipeline
  • Issue#66Shiduo-zh2025-11-11 03:33
    About Task Generation
  • Issue comment#66Shiduo-zh2025-11-11 03:33
    About Task Generation
  • Issue comment#55Shiduo-zh2025-11-11 03:30
    Question about the implementation of progress score
  • Issue comment#65Shiduo-zh2025-11-11 03:26
    about the results of Interactive Evaluation
  • Issue comment#62Shiduo-zh2025-11-11 03:23
    Integrate CLIP-RT policy into VLABench
  • Issue comment#67Shiduo-zh2025-11-11 03:21
    questions about the lora checkpoint of openvla
  • Issue comment#49Jiang-Yuxuan1232025-11-11 01:44
    environment incompatibility issue

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 131 stars here means stars gained during the window, not the repo's star count.