Skip to content

Ayanami0730/deep_research_bench

View on GitHub ↗Related repositories →

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents (Tickr Modified)

active 2025-06-132026-06-18 (UTC)

Partial coverage11,302 / 12,123 hourly files (93%) · 2 absent upstream · 818 failed, retryable2025-03-232026-08-10 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
417
Pushes
19
Pull requests
2
Issues
25
Stars
326
Forks
23

Activity over time

Daily event counts in the loaded window

Line chart, 371 days from 2025-06-13 to 2026-06-18. Pushes: 19 total, peak 4 in a day. Pull requests: 2 total, peak 1 in a day. Issues: 25 total, peak 3 in a day. Comments: 21 total, peak 3 in a day. Stars: 326 total, peak 22 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
Ayanami073023608
imlrz151302
SauravP973002
YiLIU0012001
ydai-hub2001
icecola122001
NielsRogge1000
licheer1111000
jbdatascience1000
alcholiclg1001
CocoPrince1000
lwyBZss8924d1000
miguelrios1001
NickFryXX1000
igor9silva1000
mbeato1001
belambert1010
mfkhalil1001
husimple1000
lisaalaz1000

Recent activity

Latest issues, pull requests and releases

  • Issue comment#46imlrz2026-04-24 07:31
    Gemini 2.5-pro to be shut down on June 17th
  • Issue comment#44mbeato2026-04-02 04:24
    Listed on Awesome MPP 🎉
  • Issue comment#41imlrz2026-02-28 13:38
    Leaderboard Submission: Bodhi Deep Research Team — Overall Score 54.45
  • Issue comment#36ydai-hub2025-12-31 05:56
    Request for entry in the Leaderboard - iFlow-deepResearch achieves overall score of 51.62
  • Issue#36ydai-hub2025-12-31 03:44
    Request for entry in the Leaderboard - iFlow-deepResearch achieves overall score of 51.62
  • Issue comment#28SauravP972025-12-21 16:55
    Deprecation of judge models
  • Issue comment#34SauravP972025-12-21 16:43
    Updated the deprecated gemini model names with the corresponding stable models
  • Issue comment#31miguelrios2025-12-21 03:08
    Request for entry in the Leaderboard - Static-DRA achieves overall score of 34.72
  • Pull request#33belambert2025-12-20 16:29
  • Issue comment#28mfkhalil2025-12-10 11:10
    Deprecation of judge models
  • Issue#31SauravP972025-12-05 15:24
    Request for entry in the Leaderboard - Static-DRA achieves overall score of 34.72
  • Issue comment#26Ayanami07302025-11-18 02:32
    Request to add `salesforce-air-deep-research` agent to leaderboard
  • Issue#19Ayanami07302025-11-07 08:51
    Deer-flow evaluation
  • Issue#25Even03042025-10-19 07:23
    Question Regarding FACT Evaluation for Agents Using End-of-Report Source Lists
  • Issue#24licheer1112025-09-22 08:19
    question about your article baseline
  • Pull request#23straeter2025-09-21 12:54
  • Issue comment#22alcholiclg2025-09-09 11:30
    Clarification needed on RACE evaluation metric results
  • Issue comment#16YiLIU0012025-08-03 15:10
    How to mitigate length bias in the benchmark?
  • Issue#15Ayanami07302025-08-03 14:38
    position preference
  • Issue#13Ayanami07302025-08-03 14:38
    Potential for bias towards Gemini deep research
  • Issue#12Ayanami07302025-08-03 14:38
    Retry times when scraping URL in FACT Evaluation
  • Issue#17lisaalaz2025-07-29 11:24
    Human task duration
  • Issue comment#16Ayanami07302025-07-29 04:01
    How to mitigate length bias in the benchmark?
  • Issue#16YiLIU0012025-07-29 03:36
    How to mitigate length bias in the benchmark?
  • Issue comment#14Ayanami07302025-07-28 16:14
    Splitting results between Chinese and English examples

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 326 stars here means stars gained during the window, not the repo's star count.