Models / Run locally

Run DeepSeek V4 locally

DeepSeek · ≈671B MoE (≈37B active) · open weights

Not a single-GPU model: even at 4-bit the weights alone need a multi-GPU server (≈400GB). Realistic local paths are a distilled/smaller variant, a Mac Studio cluster, or renting GPUs — most teams run the API and self-host only at sustained scale.

By quantization

Hardware requirements

Quantization≈ File size≈ Memory neededRuns on
Q4_K_M (4-bit)389 GB447 GBMulti-GPU servers (4–8× 80GB) or 512GB Mac Studio clusters
Q8_0 (8-bit)718 GB826 GBMulti-node GPU clusters — datacenter serving only
FP16 (full)1409 GB1620 GBMulti-node GPU clusters — datacenter serving only

Sizes are computed from the ≈671B parameter count with standard GGUF math and ~15% runtime overhead — verify against the model card before buying hardware. As a mixture-of-experts model only ≈37B parameters are active per token, which helps speed — but the full weights still have to fit in memory.

Local setup

How to run it

The shortest path is Ollama — it downloads a sensible default quantization and serves an OpenAI-compatible endpoint on localhost:

ollama run deepseek-v4

For serious throughput on server hardware, vLLM with tensor parallelism is the standard serving stack; llama.cpp and MLX (Apple silicon) win below that line.

The API route

Or skip the GPU entirely

The same model is one API call away at $0.3 / M input · $0.9 / M output — no download, no VRAM math, metered to the exact call. Free starter credits cover your first runs, and because the weights are open, nothing locks you in: start metered, move to your own hardware if sustained volume ever justifies it.

FAQ

Frequently asked questions

Can my GPU run DeepSeek V4?

DeepSeek V4 is ≈671B MoE (≈37B active). At 4-bit quantization you need roughly 447GB of memory (multi-gpu servers (4–8× 80gb) or 512gb mac studio clusters). Not a single-GPU model: even at 4-bit the weights alone need a multi-GPU server (≈400GB). Realistic local paths are a distilled/smaller variant, a Mac Studio cluster, or renting GPUs — most teams run the API and self-host only at sustained scale.

Is there a DeepSeek V4 GGUF?

Open-weight releases in this family get community GGUF conversions on Hugging Face shortly after release — search the model name plus "GGUF" and pick the quantization your memory allows from the table above.

Is running DeepSeek V4 locally cheaper than the API?

Only at sustained volume. The API price is $0.3 / M input · $0.9 / M output with no hardware, electricity, or ops cost — a machine that can serve this model costs more per month idle than most teams' entire inference bill. Prototype metered, self-host when utilization justifies it.

Can I use the outputs commercially?

Open weights; commercial use permitted (MIT-style model license).

Explore

Run other models locally

One key. Every model. Exact prices.

Prototype on DeepSeek V4 with free starter credits while the weights download. Free starter credits included — no subscription required.