Pimp My IDE / garage dispatch
Back to garage
October 2, 2026 | local inference / hardware fit

Local is a memory route.

DwarfStar 4 narrows the model list so it can fit large open weights onto specific high-memory machines. The interesting part is not "runs locally." It is what stays in memory, what streams from disk, and what you measured.

Choose the model layout, backend, memory class, storage route, and context length before you call local inference practical.
Memory gravity bay32 to 512 GB
Resident weightsDisk is part of the machine

The word "local" hides the load path.

DwarfStar 4, or ds4, is a C inference engine for a small set of DeepSeek, GLM, and Qwen model families. It supports Metal, CUDA, and ROCm, but the same model layout does not run on every backend. The project chooses specific GGUF layouts instead of trying to load every GGUF file.[1]

That is a useful constraint. A local-model picker should not stop at the model name. It should name the exact quantization, backend, resident memory, disk-backed data, context setting, and unsupported paths.

"Local" names the address. It does not tell you whether the model fits, streams, stalls, or survives a restart.

Memory and storage form one route.

The ds4 hardware guide starts Apple Silicon at 64 GB for a Qwen 3.8 Q2 path with 41.73 GiB of resident weights and a 95.37 GiB n-gram table on disk. It names 96 to 128 GB as the practical DeepSeek V4 Flash class. At 128 GB, the guide also lists a resident GLM 5.3 Flash Q2 path and a streamed DeepSeek V4.1 Q2 path.[2]

Those are project recommendations, not our measurements. The useful split is still clear. Resident weights consume memory. Routed experts or auxiliary tables may stream from storage. Context state adds another load. A machine can start a model and still miss the latency target for daily work.

Context changes the speed sheet.

The project benchmark page reports separate prefill and generation rates. Its M5 Max 128 GB DeepSeek V4 Flash Q2 rows list 39.4 generated tokens per second at 2,048 tokens of context and 27.6 at 65,536. The DGX Spark rows list faster prefill but slower generation for the same published model class. These are upstream measurements on named machines and settings.[3]

Do not carry one tokens-per-second number into procurement. Save machine, backend, quantization, context, prompt shape, prefill rate, generation rate, and power mode. Measure your own workload after the model starts.

Durable cache is a behavior to test.

Ds4 stores prompt-prefix state on disk and keys it by the prefix hash. Its architecture notes say tool-call mappings also persist in those cache files. That can avoid repeated prefill work after a restart, but it creates a storage contract. You need to test cache reuse, invalidation, disk growth, and deletion with the model and client you plan to use.[4]

The repository is a real artifact. We cloned revision 0aaea5a238fb41a35106a551e73c8409dfb751ac, built its CPU targets, and ran ./ds4 --help. The build produced the CLI, server, benchmark, evaluation, and agent executables. We did not download weights or run inference, so this smoke check proves source retrieval and CPU compilation only.[1]

Buy the route, not the parameter count.

  1. Pick one supported model layout and one backend.
  2. Separate resident bytes from streamed files and cache growth.
  3. Choose the context length before reading a speed result.
  4. Run one cold prompt, one warm prefix, and one restart.
  5. Keep output quality and task completion beside latency.

Large local models can trade recurring provider calls for hardware, storage, setup, and test work. That can be a good trade. The receipt needs more than a successful startup.

Interactive makeover / memory gravity bay

Route the load before launch.

Traditional purpose replaced: a generic compatibility badge. Better version: choose a platform and memory class, then draft the checks for layout, backend, storage, and a local benchmark.

Hardware route

The native controls own the state. The verdict repeats the limits stated in the project hardware guide.

Platform
Installed memory

Use the slider or a labeled detent.

128 GB
Project-documented 128 GB classDeepSeek V4 Flash Q2 is the baseline. GLM 5.3 Flash Q2 can fit resident. DeepSeek V4.1 Q2 uses SSD streaming.
Fit-ticket sections
Generated fit ticket

Keep the load path visible

Selection drafts required fields. It does not inspect hardware, download weights, or run inference.

128 GB Apple Silicon route selected. Ticket incomplete.No fit-ticket section is selected.
All four sections selected means the ticket structure is ready. It does not prove the model fits, the storage route is fast enough, the cache is correct, or the workload result is acceptable.

Sources read

Source log and evidence boundary
  1. DwarfStar 4 repository, read and cloned October 2, 2026. The repository is the artifact source for the C engine, supported interfaces, build targets, tests, and MIT license. We cloned revision 0aaea5a238fb41a35106a551e73c8409dfb751ac, ran make cpu, and exercised ./ds4 --help in a disposable directory.
  2. Ds4 hardware requirements, updated September 17 and read October 2, 2026. This project page supplies the named memory classes, backend limits, resident and streamed routes, and unsupported platform-model combinations.
  3. Ds4 benchmarks, read October 2, 2026. This project page supplies the M5 Max, DGX Spark, context, prefill, generation, and distributed rows. They are upstream results, not our measurements.
  4. Ds4 architecture notes, read October 2, 2026. This project page describes the narrow model list, model-specific layouts, asymmetric quantization, disk-backed prefix cache, tool-call mapping, and fixture-based checks.
  5. Hacker News item 49936575, "From the creator of Redis; run LLM locally with ds4", resolved through the official Hacker News API on October 2, 2026. It was the discovery signal and supports none of the technical claims.

Evidence boundary. The compatibility, memory, architecture, and speed statements come from project sources. We verified repository retrieval, the exact revision, the license, CPU compilation, the five produced executables, and CLI help. We did not download model weights, test a GPU backend, run inference, measure speed, inspect quantization quality, test cache persistence, or compare model output.