Skip to content

[ICLR2025] DiffuGPT and DiffuLLaMA: Scaling Diffusion Language Models via Adaptation from Autoregressive Models

active 2025-03-122026-06-08 (UTC)

Partial coverage11,482 / 12,480 hourly files (92%) · 2 absent upstream · 994 failed, retryable2025-03-082026-08-10 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
184
Pushes
1
Pull requests
2
Issues
13
Stars
141
Forks
6

Activity over time

Daily event counts in the loaded window

Line chart, 454 days from 2025-03-12 to 2026-06-08. Pushes: 1 total, peak 1 in a day. Pull requests: 2 total, peak 2 in a day. Issues: 13 total, peak 4 in a day. Comments: 21 total, peak 3 in a day. Stars: 141 total, peak 4 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
summmeer131012
Facico5001
iloverdl4000
Wiselnn5704021
yair-schiff2001
ctseng7772001
lmmlzn1000
tara-tan1001
mywang441000
duterscmy1001
astradzhao1001

Recent activity

Latest issues, pull requests and releases

  • Issue comment#26summmeer2025-11-18 03:59
    How can I reproduce GSM8K result based on given LoRA adapter? What are the hyperparameters?
  • Issue comment#19tara-tan2025-08-03 11:46
    Model degeneration issue?
  • Issue comment#23nitsanluke2025-07-24 12:57
    diffugpt-s (127M) generations aren't fluent and repeats
  • Issue comment#23summmeer2025-07-24 03:39
    diffugpt-s (127M) generations aren't fluent and repeats
  • Issue comment#22duterscmy2025-07-17 03:50
    No evaluation script for GSM8K
  • Issue#21Facico2025-06-27 03:00
    Question on Initial Grad Norm Behavior in DiffuLLaMA Training
  • Issue comment#21summmeer2025-06-26 21:58
    Question on Initial Grad Norm Behavior in DiffuLLaMA Training
  • Issue#21Facico2025-06-18 11:12
    Question on Initial Grad Norm Behavior in DiffuLLaMA Training
  • Issue comment#20summmeer2025-06-14 00:31
    Qwen model import error
  • Issue#20mywang442025-06-12 09:32
    qwen模型导入失败
  • Issue comment#19summmeer2025-05-31 07:19
    Model degeneration issue?
  • Issue comment#19astradzhao2025-05-30 19:13
    Model degeneration issue?
  • Pull request#18Wiselnn5702025-05-30 02:58
  • Pull request#18Wiselnn5702025-05-30 02:58
  • Issue comment#10Wiselnn5702025-05-30 02:30
    [Bug Report?]RuntimeError: The size of tensor a (1024) must match the size of tensor b (1048576) at non-singleton dimension 3
  • Issue comment#10summmeer2025-05-29 18:22
    [Bug Report?]RuntimeError: The size of tensor a (1024) must match the size of tensor b (1048576) at non-singleton dimension 3
  • Issue#17Facico2025-05-11 15:23
    Should dsigma be 1/t in DiffuLLaMA-training?
  • Issue comment#16Facico2025-05-09 02:04
    Why did the authors not adopt the mask annealing strategy for LLaMA?
  • Issue#16Facico2025-05-09 02:04
    Why did the authors not adopt the mask annealing strategy for LLaMA?
  • Issue comment#15summmeer2025-05-04 08:33
    Cannot download DiffuLLama-gsm
  • Issue comment#14summmeer2025-05-04 08:27
    Potential Issue with Lambada evaluation
  • Issue comment#12summmeer2025-05-04 08:26
    Does it support beam search generation strategy?
  • Issue#15iloverdl2025-05-02 20:22
    Cannot download DiffuLLama-gsm
  • Issue#14iloverdl2025-05-02 17:44
    Potential Issue with Lambada evaluation
  • Issue#13iloverdl2025-05-02 06:33
    [Potential bug] Errors when loading DiffuGPT-M

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 141 stars here means stars gained during the window, not the repo's star count.