Skip to content

bigcode-project/bigcodebench

View on GitHub ↗Related repositories →

[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI

active 2024-12-182026-08-07 (UTC)

Partial coverage12,748 / 14,564 hourly files (88%) · 2 absent upstream · 1,810 failed, retryable2024-12-122026-08-11 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
207
Pushes
22
Pull requests
10
Issues
16
Stars
109
Forks
11

Activity over time

Daily event counts in the loaded window

Line chart, 598 days from 2024-12-18 to 2026-08-07. Pushes: 22 total, peak 6 in a day. Pull requests: 10 total, peak 2 in a day. Issues: 16 total, peak 2 in a day. Comments: 31 total, peak 4 in a day. Stars: 109 total, peak 3 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

Recent activity

Latest issues, pull requests and releases

  • Issue#118GhostDog982026-02-22 23:02
    [Feat] Give estimates for exact cost/token count for hard
  • Issue comment#109bitkira2025-09-30 13:40
    Question: Why do sequential generation on OpenAIChatDecoder
  • Issue#114BradKML2025-09-12 02:08
    🤗 [REQUEST] - Qwen3-Next series
  • Pull request#113KMasaki02102025-09-03 04:45
  • Pull request#113KMasaki02102025-09-03 04:43
  • Pull request#104terryyz2025-09-02 20:06
  • Pull request#112alexazhou2025-08-30 15:31
  • Issue comment#106xhinini2025-08-09 21:16
    Gradio does not work properly
  • Issue#110BradKML2025-08-06 07:24
    🤗 [REQUEST] - GPT-OSS 20B and 120B
  • Issue comment#93rsmith492025-07-31 05:51
    Clarification about running the canonical solution
  • Issue#107p81sunshine2025-07-14 04:13
    Is it normal for the ground truth accuracy not to be 100%?
  • Issue comment#102BradKML2025-07-14 01:37
    🤗 [REQUEST] - Arsh-llm-0.7b
  • Issue comment#102arsh-team2025-06-28 13:04
    🤗 [REQUEST] - Arsh-llm-0.7b
  • Issue comment#103KedarnathKC2025-06-26 02:02
    TypeError: pass_k int not iterable in evaluate.py
  • Pull request#104KedarnathKC2025-06-26 02:01
  • Issue#102arsh-team2025-06-12 06:45
    🤗 [REQUEST] - Arsh-llm-0.7b
  • Issue#101BradKML2025-06-08 05:28
    🤗 [REQUEST] - RedNote Dots.LLM1
  • Issue comment#96terryyz2025-06-03 22:51
    Are there any standard instructions for the prompts?
  • Issue comment#97BradKML2025-05-28 04:28
    🤗 [REQUEST] - <MODEL_NAME>
  • Issue comment#98BradKML2025-05-26 02:53
    Add Qwen3
  • Issue#98tzurtutjuzrtzurtzurtz2025-05-25 17:26
    Add Qwen3.
  • Issue comment#96terryyz2025-05-17 22:54
    Are there any standard instructions for the prompts?
  • Issue comment#95terryyz2025-05-17 22:51
    Lots of importErrors
  • Issue#95terryyz2025-05-17 22:51
    Lots of importErrors
  • Issue comment#91terryyz2025-05-17 22:51
    What is the prompt format?

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 109 stars here means stars gained during the window, not the repo's star count.