AI NewsOpen sourceAnnouncement

DwarfStar is a new tool from the Redis creator for running AI models on your own Mac, Linux or AMD computer

Salvatore Sanfilippo, the creator of Redis, released DwarfStar, a local inference engine that runs DeepSeek V4 and Qwen 3.8 on a Mac or Linux computer with 64 GB or more of memory.

AI News

Editorial3 min read

LinkedInX

Why it mattersA model running on your own machine does not count against your API budget and does not send your prompts to a provider, which changes what work can safely use a frontier model.

A big AI model that used to need a data centre now fits on a well-specified Mac. Salvatore Sanfilippo, the person who wrote Redis, posted DwarfStar 4 (ds4) to the DwarfStar site on 2026-09-17, and it reached the top of Hacker News on 2026-10-02 with 254 points at the time of writing. Sanfilippo describes it on his page as "a narrow C inference engine for high-memory Mac, CUDA and ROCm machines".

The project's own GitHub repository had 23,048 stars and 2,219 forks when checked today, under an MIT licence. The code was first pushed in May 2026 and has been receiving commits since. It supports DeepSeek V4 Flash and V4.1 Flash, GLM 5.x and Qwen 3.8 Flash Next, both text and vision models.

What is different about it

Sanfilippo's page says this is "not a generic GGUF runner" and that "ds4 follows a small, opportunistic set of model families". So it does not try to run every model that llama.cpp runs. It targets a short list and tunes for them.

Two choices follow from that. First, the quantization. The page says ds4 uses "asymmetric 2-bit quantization" that targets the routed experts in a mixture-of-experts model while preserving critical shared paths. That is how a 400-billion-parameter class model can fit on a 128 GB machine.

Second, the key-value cache can live on disk. The page calls this "KV cache as a disk citizen" and says it writes long prefixes to the SSD and resumes by prompt hash. For an agent that reads the same project files on every call, that is the difference between re-processing 60,000 tokens of context each time and loading them from disk once.

What it costs to run

The page lists four reference machines. On an M5 Max with 128 GB and a 2-bit quantization, a 2,048-token prompt reads at 790.2 tokens a second and generates at 39.4. At a 65,536-token prompt the same machine reads at 398.5 and generates at 27.6. On an NVIDIA DGX Spark with 128 GB, prefill holds at 823 tokens a second across both prompt lengths and generation is 13.8 to 18.1. The V4.1 Q4 quantization, which keeps more precision, needs a Mac Studio with 512 GB of unified memory.

The supported hardware list is Apple Silicon Macs with 64 GB or more running Metal, NVIDIA CUDA Linux boxes including the DGX Spark, and AMD Strix Halo Framework Desktops on ROCm. All three paths are in the same C codebase.

A model you run yourself has no per-token price, no rate limit and no outbound request. For a coding agent that reads the same codebase on every turn, those three things decide whether a task that would cost 20 dollars on a hosted API can run at all. The trade is the machine: a Mac Studio or a DGX Spark with 128 GB is a serious capital purchase, so the saving only exists if you use the model enough to pay the hardware back. The Hacker News comments and the project page both say ds4 is designed to be forked and tuned to a specific setup, which is a different offer from a hosted model that charges per request.

Source

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX