Skip to content

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Python · active 2026-01-132026-08-08 (UTC)

Partial coverage10,976 / 11,449 hourly files (96%) · 2 absent upstream · 470 failed, retryable2025-04-202026-08-10 (UTC)— sampled evenly across the window, so rankings and trends hold; absolute counts scale up.
Events
4.3K
Pushes
684
Pull requests
106
Issues
261
Stars
2.5K
Forks
217

Activity over time

Daily event counts in the loaded window

Line chart, 208 days from 2026-01-13 to 2026-08-08. Pushes: 684 total, peak 25 in a day. Pull requests: 106 total, peak 6 in a day. Issues: 261 total, peak 15 in a day. Comments: 507 total, peak 26 in a day. Stars: 2,451 total, peak 264 in a day.

  • Pushes
  • Pull requests
  • Issues
  • Comments
  • Stars

Stars, PRs, issues and forks are under-captured in the later part of this window. GH Archive progressively stopped capturing non-push events during 2026 — −95% or worse by the end of the window. Every series here except Pushes fades for that reason, so a decline above reflects the archive, not this repository. Pushes stay reliable throughout, so read them, and the contributor counts derived from them, as the real signal. Data health has the measurements.

Top contributors

Pushes, PRs, issues, reviews and comments — stars and forks excluded, so this is contribution rather than popularity

ContributorContributionsPushesPRsComments
jundot9576829162
kuanjames310023
TipKnuckle14039
ethannortharc11016
fry699009
thornad9072
jasonpaulso9033
beamivalice8033
craii6006
mdevk6012
Copilot6004
blightbow6032
zulufoxtrot5002
crazyi5004
jimmylzt1885004
zviratko5004
dependabot[bot]5050
richgoodson5041
Neo2025new5022
dpbattaglia-gps5004

Recent activity

Latest issues, pull requests and releases

  • Issue comment#2553DiscoStew60822026-08-07 19:10
    perf(deepseek-v4): speed up long-prompt prefill
  • Pull request#2552tannerdsilva2026-08-07 18:09
  • Issue comment#2378kuanjames2026-07-28 03:26
    Laguna-S-2.1-oQ4e-fast infinite thinking
  • Issue comment#2325allenzha2026-07-22 05:26
    Web Dashboard crashes when switching to ModelScope tab; HuggingFace tab subsequently fails to load (oMLX 0.5.2, macOS 26.5)
  • Issue#2301thaiduong04032026-07-20 10:18
    Dflash cannot run from source
  • Issue comment#2221richgoodson2026-07-17 00:33
    Tool call output as raw XML instead of internal tool invocation with qwen3.6-27b
  • Issue#2259oneadr0667-hub2026-07-16 01:21
    gemma4 (Gemma4ForConditionalGeneration) image input is silently dropped — per-layer multimodal fusion not wired up
  • Pull request#2243aidiffuser2026-07-14 19:31
  • Issue#2203f7d74rm2n4-lab2026-07-11 19:39
    No speed boost on M5 Pro with version 0.5.0
  • Pull request#2094dependabot[bot]2026-07-05 02:23
  • Pull request#2085ethannortharc2026-07-04 01:39
  • Issue comment#2052zwcf52002026-07-03 07:50
    Fix embedding MLX resource release on unload
  • Issue#2073aosama2026-07-02 23:25
    Support Laguna models
  • Issue comment#2049DarkSwoop2026-07-02 11:15
    Token sampling crashes with "Thread group size (1024) is greater than maximum allowed threads per threadgroup (896)" on M2 Ultra
  • Issue comment#2053zwcf52002026-07-02 00:22
    fix: prevent Instruct model first-turn memory residency via engine cleanup
  • Issue comment#1841wintercharm2026-06-28 06:38
    Attempting to quantize NEX-mini fails on missing MTP headers
  • Pull request#2023dependabot[bot]2026-06-28 02:22
  • Pull request#2023dependabot[bot]2026-06-28 02:22
  • Pull request#2023dependabot[bot]2026-06-28 02:22
  • Issue comment#2016dahai802026-06-28 00:48
    TOKEN GENERATION is very slow
  • Issue#2019bthemouth2026-06-27 21:24
    Problems with context length
  • Issue comment#1961Ramcode642026-06-27 02:23
    fix: allow uppercase letters in profile and template names
  • Issue#2011daveattard-cell2026-06-27 02:19
    cohere2_moe chat streaming can crash with UnicodeDecodeError in mlx_vlm BPE detokenizer
  • Issue comment#1965jimmyken2026-06-27 01:42
    gemma-4-e2b-it-4bit VLM loading regression: "Received 140 parameters not in model" (worked previously)
  • Issue#1908jundot2026-06-22 01:49
    Bug: /v1/responses adapter prepends instructions without deduplicating system messages, triggering strict chat template validation

Totals cover only the window loaded into ClickHouse and count events, not GitHub's lifetime totals — 2,451 stars here means stars gained during the window, not the repo's star count.