Skip to content

Open-Superintelligence-Lab/5-dollar-llm

View on GitHub ↗Related repositories →

Train LLM from scratch for $5 USD - Research.

Python · active 2025-12-132026-03-21 (UTC)

Complete coverage26,715 / 26,715 hourly files (100%) · 2 absent upstream2023-08-152026-09-01 (UTC)
Events
378
Pushes
140
Pull requests
34
Issues
14
Stars
73
Forks
15

Activity over time

Daily event counts in the loaded window

Line chart, 99 days from 2025-12-13 to 2026-03-21. Pushes: 140 total, peak 25 in a day. Pull requests: 34 total, peak 5 in a day. Issues: 14 total, peak 4 in a day. Comments: 53 total, peak 9 in a day. Stars: 73 total, peak 20 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
vukrosic2061421932
RohanKhanBD14058
shehab-ashraf5014
Ffinnis4021
marc-shade3003
bigwolfeman3003
github-actions[bot]3003
Tan9uy2020
VishnuTejaJ2020
toheedakhtar2020
code2tan1000
HosLak1000
Roma-Potapov1001
batuhanozkose1010

Recent activity

Latest issues, pull requests and releases

  • Pull request#94batuhanozkose2026-01-03 16:25
  • Issue comment#89vukrosic2026-01-02 12:57
    Change tokenizer to mistral 7b
  • Issue comment#89vukrosic2026-01-02 11:58
    Change tokenizer to mistral 7b
  • Issue comment#89RohanKhanBD2026-01-02 10:57
    Change tokenizer to mistral 7b
  • Issue comment#88vukrosic2026-01-02 10:22
    merge QKVO projections
  • Pull request#88shehab-ashraf2026-01-02 09:36
  • Issue#25vukrosic2025-12-27 14:59
    What can we implement from Karpathy's "Best ChatGPT $100 can buy"
  • Issue#87vukrosic2025-12-27 14:57
    Does replacing some layers with linear attention (GQA or KDA) from FLA improves the model training?
  • Issue comment#84vukrosic2025-12-26 07:51
    Implement drop-muon optimizer
  • Issue comment#84Ffinnis2025-12-26 07:23
    Implement drop-muon optimizer
  • Pull request#70vukrosic2025-12-25 11:34
  • Issue comment#70vukrosic2025-12-25 11:34
    Fix hardcoded sequence length limitation for RoPE
  • Issue comment#83vukrosic2025-12-25 08:52
    Fix Windows multiprocessing error in DataLoader
  • Pull request#83VishnuTejaJ2025-12-25 08:52
  • Issue comment#84Roma-Potapov2025-12-25 03:29
    Implement drop-muon optimizer
  • Issue comment#84vukrosic2025-12-24 21:38
    Implement drop-muon optimizer
  • Pull request#83VishnuTejaJ2025-12-24 14:24
  • Pull request#82RohanKhanBD2025-12-24 14:14
  • Issue#12vukrosic2025-12-24 09:43
    Token Embedding "Smearing" ⏩
  • Issue#49vukrosic2025-12-24 09:41
    ReLU² Research
  • Issue comment#49vukrosic2025-12-24 09:41
    ReLU² Research
  • Issue comment#76vukrosic2025-12-23 18:58
    Token-smear
  • Pull request#76RohanKhanBD2025-12-23 18:58
  • Pull request#73Ffinnis2025-12-23 17:17
  • Issue comment#73vukrosic2025-12-23 17:16
    Integrated Parallel Transformer Block +10% on training speed

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 73 stars here means stars gained during the window, not the repo's star count.