Lab · Local models · ~2 min
Local VRAM estimator
Pick a model size and quantization to estimate VRAM before you download weights. Then compare local stacks like Ollama and MLC LLM.
How to read this
- Weights plus a KV/overhead allowance — not a measured llama.cpp dump.
- Long context and image models need more than this table shows.
- Compare Ollama, LM Studio, and MLC LLM after you know the memory budget.