Skip to content

Bringing BERT into modernity via both architecture changes and scaling, for nlp lm

active 2024-12-192026-06-07 (UTC)

Partial coverage19,418 / 25,022 hourly files (78%) · 2 absent upstream · 5,601 failed, retryable2023-10-052026-08-12 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
1.3K
Pushes
28
Pull requests
17
Issues
57
Stars
961
Forks
79

Activity over time

Daily event counts in the loaded window

Line chart, 536 days from 2024-12-19 to 2026-06-07. Pushes: 28 total, peak 9 in a day. Pull requests: 17 total, peak 7 in a day. Issues: 57 total, peak 6 in a day. Comments: 150 total, peak 12 in a day. Stars: 961 total, peak 83 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
NohTow380033
warner-benjamin2914110
bclavie171043
neavo8006
davedgd7004
znsoftm5004
ohmeow5230
stefan-it4003
BramVanroy4004
rangehow4001
QXGeraldMo4003
JunyiZhu-AI4002
umarbutler4003
eanson0234004
tjasmin1114003
ahxxm3003
amishparekh3002
turtle0x13002
staghado3201
tomaarsen3021

Recent activity

Latest issues, pull requests and releases

  • Issue comment#238review-notebook-app[bot]2026-03-21 22:56
    Fix notebook
  • Issue comment#173geraldstanje12026-02-12 22:03
    Request for code to quantize and convert ModernBERT Model to ONNX
  • Issue comment#246staghado2026-02-10 10:27
    Resuming takes too much time with `load_path` activated
  • Issue comment#246NohTow2026-01-12 08:52
    Resuming takes too much time with `load_path` activated
  • Issue comment#246Rijgersberg2026-01-09 16:43
    Resuming takes too much time with `load_path` activated
  • Issue comment#252jsrozner2026-01-05 16:15
    Huggingface - modernbert checkpoints
  • Issue comment#163warner-benjamin2025-12-12 03:54
    MaskedLM nan training loss
  • Issue comment#163beapirate2025-12-12 00:18
    MaskedLM nan training loss
  • Issue comment#153jmcmanus152025-12-05 23:39
    checkpoints for further pretraining
  • Issue comment#149jwijffels2025-11-28 14:15
    On the performance of token classification
  • Issue comment#149BramVanroy2025-11-26 11:28
    On the performance of token classification
  • Issue comment#149BramVanroy2025-11-26 10:06
    On the performance of token classification
  • Issue comment#149NohTow2025-11-26 07:40
    On the performance of token classification
  • Issue comment#149BramVanroy2025-11-26 07:21
    On the performance of token classification
  • Issue comment#149stefan-it2025-11-25 16:07
    On the performance of token classification
  • Issue comment#149BramVanroy2025-11-25 08:54
    On the performance of token classification
  • Issue#249mmichall2025-10-23 10:07
    StopIteration error during training with modernbert-large-context-extension.yaml configuration
  • Issue comment#248warner-benjamin2025-10-22 19:18
    Why is the code truncating the samples?
  • Issue#247dtamayo-nlp2025-10-17 14:44
    `global_train_batch_size` doesn't affect the number of steps, bug in gradient accumulation?
  • Issue comment#236ahxxm2025-09-20 14:59
    pre-training speed, data loader and gpu utilization
  • Issue comment#236ahxxm2025-09-20 09:31
    pre-training speed, data loader and gpu utilization
  • Issue comment#242NohTow2025-09-17 07:37
    Question about masking - why mask on unpadded sequence?
  • Issue#245Keramatfar2025-08-23 07:26
    Continue pretraining on my dataset
  • Issue comment#241ceferisbarov2025-08-22 18:05
    `_pad_token` attribute?
  • Issue comment#174eanson0232025-08-21 17:19
    ModernBertModel works on the CPU but fails on the GPU

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 961 stars here means stars gained during the window, not the repo's star count.