signalGitHub Trending2026-10-06
antirez/ds4
DwarfStar is a self-contained inference engine for DeepSeek V4 Flash and PRO, GLM 5.2/5.3, and Qwen3.8 Flash, targeting Metal, CUDA, and ROCm. It runs on consumer hardware like Macs with 96 GB or more, DGX Spark, and Strix Halo, with SSD streaming for smaller systems. An eight-L40S setup reaches 126 t/s aggregate generation across 16 sessions.
- for who
- Developers and hobbyists wanting to run large open-weight models locally on personal hardware.
- why now
- New engine runs latest DeepSeek V4 and GLM 5 on consumer hardware.
- what changes
- They can run capable models on Macs, DGX Spark, or Strix Halo without cloud dependency, even with limited RAM via SSD streaming.
- to do
- Download the project and use its GGUF files to run supported models on Metal, CUDA, or ROCm.
key points
- DwarfStar supports DeepSeek V4 Flash/PRO, GLM 5.2/5.3, Qwen3.8 Flash on Metal, CUDA, ROCm
- Eight-L40S setup achieves 126 t/s aggregate generation with 16 sessions
- SSD streaming enables large models on machines with insufficient RAM
#local inference#DeepSeek#open source tool
score
score 7 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source