Local AI Model Directory
Last updated: September 2026
Filter local AI models by what your hardware can run and what you'll use it for. Every model links to more detail on picking and running it. For an exact memory check on any of these, use the hardware calculator.
Showing 17 of 17 models
Llama 3.1 8B
~8GB min • Q4 ~5.5GB
The most broadly supported starting model, with the widest tool compatibility.
See run guide →Mistral 7B
~6GB min • Q4 ~4.7GB
Fastest model at this size, with a permissive license.
See hardware-tier guide →Qwen3 8B
~8GB min • Q4 ~5.5GB
Strong general-purpose pick, especially outside English.
See run guide →Phi-4-mini
~4GB min • Q4 ~2.6GB
Small footprint, strong reasoning for its size class.
See hardware-tier guide →Qwen2.5-Coder 7B
~6GB min • Q4 ~4.7GB
Coding-focused variant of Qwen sized for entry hardware.
See coding guide →Qwen3 14B
~14GB min • Q4 ~9.6GB
Community favorite at this tier for code-adjacent work.
See coding guide →DeepSeek-R1-Distill 14B
~16GB min • Q4 ~9.6GB
Distilled from DeepSeek's reasoning model; strong at math and logic chains.
See run guide →Gemma 3 12B
~14GB min • Q4 ~8.2GB
Handles image input alongside text; well-rounded general use.
See hardware-tier guide →Phi-4
~14GB min • Q4 ~9.6GB
Smaller footprint than most peers at similar reasoning quality.
See hardware-tier guide →Qwen3 32B
~24GB min • Q4 ~21.9GB
The clear step up from the 14B tier in general reasoning and coding.
See run guide →Gemma 3 27B
~24GB min • Q4 ~18.5GB
Largest Gemma 3 size that comfortably fits a 24GB card.
See hardware-tier guide →DeepSeek-R1-Distill 32B
~32GB min • Q4 ~21.9GB
The strongest local reasoning most people can actually run.
See run guide →Mistral Small 3.1
~20GB min • Q4 ~16.4GB
Fast execution and strong translation, multimodal.
See hardware-tier guide →Qwen2.5-Coder 32B
~24GB min • Q4 ~21.9GB
The strongest coding-dedicated pick most people can realistically run locally.
See coding guide →Devstral
~20GB min • Q4 ~16.4GB
Purpose-built for agentic, multi-step coding workflows rather than single-shot completion.
See coding guide →Llama 4 Scout
~60GB min • Q4 ~60GB
High-memory unified-memory Macs only. Full parameter set loads regardless of active-parameter count; see the tier guide for why.
See hardware-tier guide →Llama 3.3 70B
~64GB min • Q4 ~49.1GB
Meaningful step up from the 8B version; needs a Mac Studio-class machine or a high-VRAM multi-GPU setup at Q4.
See run guide →Model specs (parameter counts, licenses, approximate memory needs) are cross-checked against multiple current local-AI model guides and, where available, official model cards (see the fact-check log). Memory figures use the same Q4_K_M formula as the hardware calculator. This directory gets rechecked as new model versions ship, since local model releases move fast.
Frequently asked questions
How current is this directory?
It reflects widely-recommended local models as of this writing, cross-checked against multiple current sources rather than one static ranking. New models ship constantly in this space; this page gets rechecked periodically rather than left to age.
Why do the memory figures here differ slightly from what I've seen elsewhere?
Every source uses slightly different assumptions about quantization overhead and context length. The figures here use this site's own documented formula, visible on the hardware calculator page, so you can see exactly how they're computed rather than trusting a black-box number.
Can I filter by more than tier and use case?
Not yet. License and context window are shown as tags on each card so you can scan for them, even though they're not filterable dimensions yet.