LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Python · active 2026-01-13 → 2026-08-08 (UTC)
Activity over time
Daily event counts in the loaded window
Line chart, 208 days from 2026-01-13 to 2026-08-08. Pushes: 684 total, peak 25 in a day. Pull requests: 106 total, peak 6 in a day. Issues: 261 total, peak 15 in a day. Comments: 507 total, peak 26 in a day. Stars: 2,451 total, peak 264 in a day.
- Pushes
- Pull requests
- Issues
- Comments
- Stars
Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.
Top contributors
Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity
| Contributor | Contributions | Pushes | PRs | Comments |
|---|---|---|---|---|
| jundot | 957 | 682 | 9 | 162 |
| kuanjames | 31 | 0 | 0 | 23 |
| TipKnuckle | 14 | 0 | 3 | 9 |
| ethannortharc | 11 | 0 | 1 | 6 |
| fry69 | 9 | 0 | 0 | 9 |
| thornad | 9 | 0 | 7 | 2 |
| jasonpaulso | 9 | 0 | 3 | 3 |
| beamivalice | 8 | 0 | 3 | 3 |
| craii | 6 | 0 | 0 | 6 |
| mdevk | 6 | 0 | 1 | 2 |
| Copilot | 6 | 0 | 0 | 4 |
| blightbow | 6 | 0 | 3 | 2 |
| zulufoxtrot | 5 | 0 | 0 | 2 |
| crazyi | 5 | 0 | 0 | 4 |
| jimmylzt188 | 5 | 0 | 0 | 4 |
| zviratko | 5 | 0 | 0 | 4 |
| dependabot[bot] | 5 | 0 | 5 | 0 |
| richgoodson | 5 | 0 | 4 | 1 |
| Neo2025new | 5 | 0 | 2 | 2 |
| dpbattaglia-gps | 5 | 0 | 0 | 4 |
Recent activity
Latest issues, pull requests and releases
- Issue comment#2553DiscoStew60822026-08-07 19:10perf(deepseek-v4): speed up long-prompt prefill
- Pull request#2552tannerdsilva2026-08-07 18:09
- Issue comment#2378kuanjames2026-07-28 03:26Laguna-S-2.1-oQ4e-fast infinite thinking
- Issue comment#2325allenzha2026-07-22 05:26Web Dashboard crashes when switching to ModelScope tab; HuggingFace tab subsequently fails to load (oMLX 0.5.2, macOS 26.5)
- Issue#2301thaiduong04032026-07-20 10:18Dflash cannot run from source
- Issue comment#2221richgoodson2026-07-17 00:33Tool call output as raw XML instead of internal tool invocation with qwen3.6-27b
- Issue#2259oneadr0667-hub2026-07-16 01:21gemma4 (Gemma4ForConditionalGeneration) image input is silently dropped — per-layer multimodal fusion not wired up
- Pull request#2243aidiffuser2026-07-14 19:31
- Issue#2203f7d74rm2n4-lab2026-07-11 19:39No speed boost on M5 Pro with version 0.5.0
- Pull request#2094dependabot[bot]2026-07-05 02:23
- Pull request#2085ethannortharc2026-07-04 01:39
- Issue comment#2052zwcf52002026-07-03 07:50Fix embedding MLX resource release on unload
- Issue#2073aosama2026-07-02 23:25Support Laguna models
- Issue comment#2049DarkSwoop2026-07-02 11:15Token sampling crashes with "Thread group size (1024) is greater than maximum allowed threads per threadgroup (896)" on M2 Ultra
- Issue comment#2053zwcf52002026-07-02 00:22fix: prevent Instruct model first-turn memory residency via engine cleanup
- Issue comment#1841wintercharm2026-06-28 06:38Attempting to quantize NEX-mini fails on missing MTP headers
- Pull request#2023dependabot[bot]2026-06-28 02:22
- Pull request#2023dependabot[bot]2026-06-28 02:22
- Pull request#2023dependabot[bot]2026-06-28 02:22
- Issue comment#2016dahai802026-06-28 00:48TOKEN GENERATION is very slow
- Issue#2019bthemouth2026-06-27 21:24Problems with context length
- Issue comment#1961Ramcode642026-06-27 02:23fix: allow uppercase letters in profile and template names
- Issue#2011daveattard-cell2026-06-27 02:19cohere2_moe chat streaming can crash with UnicodeDecodeError in mlx_vlm BPE detokenizer
- Issue comment#1965jimmyken2026-06-27 01:42gemma-4-e2b-it-4bit VLM loading regression: "Received 140 parameters not in model" (worked previously)
- Issue#1908jundot2026-06-22 01:49Bug: /v1/responses adapter prepends instructions without deduplicating system messages, triggering strict chat template validation
Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 2,451 stars here means stars gained during the window, not the repo's star count.