A dynamically reusable benchmark designed to evaluate Large Language Models’ abstract reasoning capabilities through continuously generated “Guess the Rule” games. This repository features an automated rule generation system, customizable APIs for on-demand dataset creation, and an interactive game interface for robust and scalable LLM assessment.
active 2024-09-27 → 2025-01-10 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 106 days from 2024-09-27 to 2025-01-10. Pushes: 163 total, peak 32 in a day. Pull requests: 11 total, peak 3 in a day. Issues: 0 total, peak 0 in a day. Comments: 0 total, peak 0 in a day. Stars: 5 total, peak 2 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| alishazal | 70 | 65 | 5 | 0 |
| m1chae11u | 56 | 55 | 1 | 0 |
| XiangZheng2002 | 29 | 26 | 3 | 0 |
| JunoLee128 | 13 | 11 | 2 | 0 |
| arihantchoudhary | 6 | 6 | 0 | 0 |
Recent activity
Latest issues, pull requests and releases
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 5 stars here means stars gained during the window, not the repo's star count.