ESC
开源 5 分钟阅读

Deltafin 在 MacBook Pro 上以约 1 token/s 运行 2.8T 参数的 Kimi K3:从四块 SSD 流式加载

开源项目 Deltafin 宣布在 MacBook Pro 上以 0.29 token/s 运行完整的 2.8 万亿参数 Kimi K3 模型,专家权重从四块 SSD 流式读取,吞吐量较首版提升约 20 倍。项目坚持不做量化,保留全部 16 个路由专家与 Moonshot 原版权重,目标是证明约 1.5 万美元的家用设备也能运行原本需要 200 万美元基础设施的前沿模型。

来源:Hacker News

  • 0.2901 token/s (3.447 s/token) — 1.9% higher throughput than last update

Upstream’s historical M1 benchmarks:

  • 0.2847 token/s (August 2, 2026) — 7.0% higher throughput
  • 0.2660 token/s (July 30, 2026) — 102.9% higher throughput
  • 0.1311 token/s (July 28, 2026) — 829.8% higher throughput
  • 0.0141 token/s (July 27, 2026)

model platforms accelerators experts runtime license

Pure raw uncut K3 quality, as fast as possible. Speed must never come from reducing model quality. Deltafin keeps all 16 routed experts and the full K3 target as the sole authority for every single token.

Our goal is to squeeze out every last drop of efficiency possible when running a huge model like K3, with all options on the table… except for reducing quality.

Deltafin is not a product pitch. It is an experiment in how far consumer hardware can be pushed, and what we can learn by attempting something so challenging.

Kimi K3 targets infrastructure on the scale of 16 nodes and roughly 4.8 TB of aggregate VRAM. That means the full 2.8T parameters and the 1M-token context window, with the expert bank never pruned. On any home setup, this is an extreme constraint. Every 1% improvement is very hard-won. But each gain can teach something.

Research and exploration is the point. That is our mission. Not everything has to be a “minimum viable product” to impress venture capitalists. If Deltafin helps make frontier models usable on a $15,000 home setup, instead of a $2,000,000 infrastructure like Kimi recommends, we believe that is worthwhile progress on our self-hosted AI journey. Plus everything learned along the way could even benefit other projects in unexpected ways.

“We choose to run the full 2.8-trillion-parameter model locally, and do the other things, not because they are easy, but because they are hard.” — John F. Kennedy probably

Other projects appear to run full K3, somehow faster. But look closer: they’ve re-encoded K3’s expert bank down to ~3 bits. Clever engineering toward a different goal: the smallest K3 that fits and is “close enough.” Those weights are no longer the ones Moonshot released, and nobody, including them, has measured what those compromises cost.

Deltafin is the other experiment: every expert byte exactly as Moonshot shipped it, made as fast as physics allows.

Deltafin installs almost everything it needs. See Requirements if you’re missing anything.

1. Get it

git clone https://github.com/gavamedia/deltafin.git cd deltafin

2. Build it

cargo build –locked –release

3. Download the FULL 1.7 TB K3 model to disk (optional, but fastest)

./target/release/deltafin setup –full

Or, if you don’t have enough disk space:

3. Stream K3 as you use it (slower, but 215 GB to start)

./target/release/deltafin setup –stream

setup --stream installs the resident model and fetches exact experts on-demand only, initially running far more slowly when routes have no local cache yet. As you build up your cache over time, this can be a way to save space, storing only the parts of the model you use, running entirely off cache on disk.

The normal setup includes Inferact’s Kimi-K3-DSpark. It takes 6.635 GiB on disk and approximately 4.49 GiB when admitted at runtime. Deltafin avoids materializing DSpark’s redundant copy of K3’s embedding. Chat and server requests use DSpark automatically when beneficial; any failures, insufficient headroom, or bad live economics simply leaves full K3 running by itself.

Qwen is a separate add-on for faster raw text continuation:

4. Optionally install qwen later

./target/release/deltafin setup-qwen

Qwen speeds up raw completion only: the small models guess what comes next, K3 checks the guess, and you get identical output in less time. That helps code autocomplete and other /v1/completions traffic, plus deltafin run --prompt ... — one measured 17-token completion ran 2.7× faster with the same output IDs.

This adds 4.337 GiB on disk, and because Qwen will not improve chat speed, we make it an optional add-on. You can add it to a fresh install with deltafin setup --full --include-qwen, or add it later with the command above.

./target/release/deltafin upgrade

upgrade gets what you need, and rebuilds the binary. Models, converted weights, and caches are left alone. It never re-runs setup or re-downloads K3.

git status --short

# ⬆️ Continue only when that returns nothing

git pull --ff-only
cargo build --locked --release
./target/release/deltafin upgrade

Continue only when git status --short is empty. If it lists files, preserve or commit that work yourself, rather than allowing an upgrade procedure to guess. Existing model data remains in the same repository-root directories.

upgrade needs a clean, non-diverged branch, and it remembers how the binary was built, so an NVIDIA/CUDA build stays a CUDA build rather than quietly falling back to CPU. Anything unexpected safely stops the upgrade.

upgrade ignores build environment variables — it reuses whatever the binary was already built with. So to switch configuration (CPU to CUDA, say, or a moved LibTorch tree), run cargo build --locked --release yourself once with the new variables set; see Requirements. That becomes the recorded setup, and later upgrades keep it.

# Chat: apply K3's audited chat template and stop at the model's end marker.
./target/release/deltafin run --chat \
 --prompt "What are the three largest moons of Saturn?"

# Raw continuation: cap output because raw text has no chat end boundary.
./target/release/deltafin run \
 --prompt "The capital of France is" --max-new 17

# Add cumulative throughput and native transaction statistics.
./target/release/deltafin run \
 --prompt "The largest planet in our solar system is" --max-new 17 --stats

Without --stats, generated text streams normally instead of printing one diagnostic line per token. --max-new N limits new tokens; it does not alter the prompt or context. Chat output stops at K3’s control boundary, while raw completion should normally use a bound.