[ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Python · active 2025-10-28 → 2026-08-01 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 278 days from 2025-10-28 to 2026-08-01. Pushes: 138 total, peak 14 in a day. Pull requests: 6 total, peak 1 in a day. Issues: 17 total, peak 2 in a day. Comments: 21 total, peak 5 in a day. Stars: 120 total, peak 11 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| jxhe | 95 | 94 | 1 | 0 |
| lockon-n | 56 | 44 | 0 | 7 |
| cursor[bot] | 16 | 0 | 0 | 11 |
| williamIIliu | 3 | 0 | 0 | 1 |
| luccabb | 2 | 0 | 2 | 0 |
| KaeHyun | 2 | 0 | 0 | 1 |
| bugmaker00 | 2 | 0 | 1 | 0 |
| jitrc | 1 | 0 | 0 | 0 |
| Skripkon | 1 | 0 | 0 | 0 |
| LLLeoLi | 1 | 0 | 1 | 0 |
| Xubqpanda | 1 | 0 | 1 | 0 |
| superrrrrr1995 | 1 | 0 | 0 | 0 |
| onestardao | 1 | 0 | 0 | 0 |
| gc641533-stack | 1 | 0 | 0 | 1 |
| cwyoon-99 | 1 | 0 | 0 | 0 |
| WassupWuK | 1 | 0 | 0 | 0 |
| TomLucidor | 1 | 0 | 0 | 0 |
| razvangabdumitru-spec | 1 | 0 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
- Issue#45cwyoon-992026-05-29 04:28poste & woocommerce containers fail again (with "port is already allocated")
- Pull request#40LLLeoLi2026-05-04 11:54
- Pull request#38bugmaker002026-04-26 11:50
- Issue#36WassupWuK2026-03-31 04:37The `git-repo` task's evaluation script may be overly strict
- Issue comment#36lockon-n2026-03-31 04:11The `git-repo` task's evaluation script may be overly strict
- Issue comment#32KaeHyun2026-03-13 05:32Gap for DeepSeek v3.2-exp (OpenRouter)
- Issue#32KaeHyun2026-03-13 05:21Gap for DeepSeek v3.2-exp (OpenRouter)
- Issue comment#29lockon-n2026-02-18 02:16Evaluate claude with extended thinking mode
- Issue#31razvangabdumitru-spec2026-02-12 23:26Investigation: dev vs main diverged — commits and potential conflicts
- Issue#26onestardao2026-02-06 09:00Question: using an S-class TXT "tension field" as a high-pressure context for Toolathlon agents?
- Pull request#22Xubqpanda2025-12-30 11:37
- Issue#21lockon-n2025-12-27 02:51Dedicated Server Fails Due to Hardcoded Port in WebSocket Proxy
- Issue comment#21lockon-n2025-12-27 02:51Dedicated Server Fails Due to Hardcoded Port in WebSocket Proxy
- Issue#21jitrc2025-12-26 22:57Dedicated Server Fails Due to Hardcoded Port in WebSocket Proxy
- Issue comment#19gc641533-stack2025-12-25 09:20Upgrade mcp-remote to support proxy
- Issue#18lockon-n2025-12-24 10:37Public Evaluation Service seems not working now.
- Issue comment#16lockon-n2025-12-22 08:49global deployment
- Issue#16bugmaker002025-12-22 05:15global deployment
- Issue#15TomLucidor2025-12-18 06:04Inclusion of SLMs into the leaderboard
- Issue#14superrrrrr19952025-12-17 09:04Notion MCP Timeout Error
- Issue#12Skripkon2025-12-09 07:12Contradiction in the paper: do frequent tool call errors negatively affect model overall performance?
- Pull request#13luccabb2025-12-08 21:38
- Pull request#11luccabb2025-12-05 05:07
- Issue comment#4lockon-n2025-11-30 04:01运行测试脚本的时候出现python的问题
- Issue#8lockon-n2025-11-30 03:56task label
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 120 stars here means stars gained during the window, not the repo's star count.