DwarfStar 4 (ds4) runs DeepSeek V4 Flash locally with 2-bit experts

DwarfStar 4, shortened to ds4, is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. According to its project site, it supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack. The site lists an MIT license and C, Metal, CUDA and ROCm backends.

The pitch starts from a constraint. DeepSeek V4 Flash is a large mixture-of-experts model, and the usual way to run it is remote serving. ds4 starts from the opposite end: making it practical on one local machine. The core technique is asymmetric 2-bit quantization. It compresses the routed experts while keeping critical shared paths precise, and the site says this is how the supported routed-MoE builds fit their target machines.

The project is deliberately not a generic GGUF runner. It follows a small, opportunistic set of model families and validates each supported layout end to end. Its own GGUF files are the target; generic GGUF files are not. The architecture is described as project GGUFs, a self-contained engine and agent-facing interfaces, checked against official model outputs.

A second design choice is prefix caching on disk: long prefixes can be saved to SSD and resumed by prompt hash, so a restart does not have to mean a full re-prefill. The CLI, HTTP APIs and native agent all share the same model state and cache.

Three binaries cover the use cases: ./ds4 for chat, ./ds4-server for local APIs and ./ds4-agent for persistent coding sessions. The server speaks OpenAI and Anthropic-style APIs, so local coding agents can connect to your own machine with a base URL.

Setup has three steps: fetch the project GGUF, build for your backend, then talk to it. The example given is ./download_model.sh ds4f-q2 && make. A fit check on the site says that at 128 GB, V4 Flash Q2 is the baseline, GLM 5.3 Q2 and Qwen Q4 also fit, and V4.1 Q2 streams from SSD. The page also labels Qwen as running on 64GB.

The one benchmark figure quoted is a reference row for an M5 Max with 128GB at 32K context: 34.4 tokens per second generation and 557 tokens per second prefill. The site calls these estimates from the ds4 benchmark table, says the rows are reference rows from upstream, and advises reading prefill and generation separately, especially for long-context agent workloads.

Key facts

  • ds4 (DwarfStar 4) is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines, under an MIT license.
  • It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, using asymmetric 2-bit quantization of the routed experts while keeping critical shared paths precise.
  • Reference row on an M5 Max with 128GB at 32K context: 34.4 tokens per second generation, 557 tokens per second prefill.
  • ds4-server speaks OpenAI and Anthropic-style APIs; long prefixes can be saved to SSD and resumed by prompt hash.
  • It runs the project's own GGUFs only: generic GGUF files are not the target.

Why it matters

Large mixture-of-experts models are usually served remotely. ds4 argues for the opposite: compress only the routed experts to 2 bits, keep the critical shared paths precise, and the model becomes practical on a high-memory machine. It bundles a CLI, HTTP APIs and a coding agent over one shared model state and cache, so one local install covers chat, API access and agent sessions.

Who it affects

People with high-memory Mac, CUDA or ROCm machines who want to run DeepSeek V4 Flash, GLM 5.x or Qwen3.8 Flash Next themselves. The site's fit check puts V4 Flash Q2 as the baseline at 128 GB and lists Qwen as running on 64GB. Developers who use local coding agents are a second audience, since the server accepts OpenAI and Anthropic-style API calls.

How to use it

Download the project GGUF, build for your backend, then start the CLI or server. The site's example is ./download_model.sh ds4f-q2 && make. Use ./ds4 for chat, ./ds4-server for local APIs and ./ds4-agent for persistent coding sessions. To connect an editor, agent or API client, point it at the local server with a base URL. The license is MIT. Generic GGUF files will not work as a target; use the project's own.

How solid is it

Everything here comes from the project's own site, so these are the maintainers' claims, not independent tests. The single benchmark figure (34.4 tokens per second generation, 557 prefill on an M5 Max with 128GB at 32K context) is described by the site as an estimate from its benchmark table. The page does not name the author or creator of ds4, so the submission headline's mention of the creator of Redis is not confirmed by the page itself.

Risks and caveats

The page gives no comparison against llama.cpp or other engines, and no output-quality or accuracy figures for the 2-bit quantization. It gives no details on which layers count as critical shared paths or how the quantization is implemented. Support is limited to a small, opportunistic set of model families, so other models are out of scope. Hardware needs are high: V4.1 Q2 streams from SSD even at 128 GB. The site also warns to read prefill and generation speeds separately for long-context agent workloads.

“Compress the routed experts, keep critical shared paths precise.”

— ds4 project site