Bringing BERT into modernity via both architecture changes and scaling, for nlp lm
active 2024-12-19 → 2026-06-07 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 536 days from 2024-12-19 to 2026-06-07. Pushes: 28 total, peak 9 in a day. Pull requests: 17 total, peak 7 in a day. Issues: 57 total, peak 6 in a day. Comments: 150 total, peak 12 in a day. Stars: 961 total, peak 83 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| NohTow | 38 | 0 | 0 | 33 |
| warner-benjamin | 29 | 14 | 1 | 10 |
| bclavie | 17 | 10 | 4 | 3 |
| neavo | 8 | 0 | 0 | 6 |
| davedgd | 7 | 0 | 0 | 4 |
| znsoftm | 5 | 0 | 0 | 4 |
| ohmeow | 5 | 2 | 3 | 0 |
| stefan-it | 4 | 0 | 0 | 3 |
| BramVanroy | 4 | 0 | 0 | 4 |
| rangehow | 4 | 0 | 0 | 1 |
| QXGeraldMo | 4 | 0 | 0 | 3 |
| JunyiZhu-AI | 4 | 0 | 0 | 2 |
| umarbutler | 4 | 0 | 0 | 3 |
| eanson023 | 4 | 0 | 0 | 4 |
| tjasmin111 | 4 | 0 | 0 | 3 |
| ahxxm | 3 | 0 | 0 | 3 |
| amishparekh | 3 | 0 | 0 | 2 |
| turtle0x1 | 3 | 0 | 0 | 2 |
| staghado | 3 | 2 | 0 | 1 |
| tomaarsen | 3 | 0 | 2 | 1 |
Recent activity
Latest issues, pull requests and releases
- Issue comment#238review-notebook-app[bot]2026-03-21 22:56Fix notebook
- Issue comment#173geraldstanje12026-02-12 22:03Request for code to quantize and convert ModernBERT Model to ONNX
- Issue comment#246staghado2026-02-10 10:27Resuming takes too much time with `load_path` activated
- Issue comment#246NohTow2026-01-12 08:52Resuming takes too much time with `load_path` activated
- Issue comment#246Rijgersberg2026-01-09 16:43Resuming takes too much time with `load_path` activated
- Issue comment#252jsrozner2026-01-05 16:15Huggingface - modernbert checkpoints
- Issue comment#163warner-benjamin2025-12-12 03:54MaskedLM nan training loss
- Issue comment#163beapirate2025-12-12 00:18MaskedLM nan training loss
- Issue comment#153jmcmanus152025-12-05 23:39checkpoints for further pretraining
- Issue comment#149jwijffels2025-11-28 14:15On the performance of token classification
- Issue comment#149BramVanroy2025-11-26 11:28On the performance of token classification
- Issue comment#149BramVanroy2025-11-26 10:06On the performance of token classification
- Issue comment#149NohTow2025-11-26 07:40On the performance of token classification
- Issue comment#149BramVanroy2025-11-26 07:21On the performance of token classification
- Issue comment#149stefan-it2025-11-25 16:07On the performance of token classification
- Issue comment#149BramVanroy2025-11-25 08:54On the performance of token classification
- Issue#249mmichall2025-10-23 10:07StopIteration error during training with modernbert-large-context-extension.yaml configuration
- Issue comment#248warner-benjamin2025-10-22 19:18Why is the code truncating the samples?
- Issue#247dtamayo-nlp2025-10-17 14:44`global_train_batch_size` doesn't affect the number of steps, bug in gradient accumulation?
- Issue comment#236ahxxm2025-09-20 14:59pre-training speed, data loader and gpu utilization
- Issue comment#236ahxxm2025-09-20 09:31pre-training speed, data loader and gpu utilization
- Issue comment#242NohTow2025-09-17 07:37Question about masking - why mask on unpadded sequence?
- Issue#245Keramatfar2025-08-23 07:26Continue pretraining on my dataset
- Issue comment#241ceferisbarov2025-08-22 18:05`_pad_token` attribute?
- Issue comment#174eanson0232025-08-21 17:19ModernBertModel works on the CPU but fails on the GPU
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 961 stars here means stars gained during the window, not the repo's star count.