Local AI Model Directory

By , Ready Utilities

Last updated: September 2026

Filter local AI models by what your hardware can run and what you'll use it for. Every model links to more detail on picking and running it. For an exact memory check on any of these, use the hardware calculator.

Showing 17 of 17 models

Llama 3.1 8B

8B paramsLlama Community License

~8GB min • Q4 ~5.5GB

The most broadly supported starting model, with the widest tool compatibility.

See run guide →

Mistral 7B

7B paramsApache 2.0

~6GB min • Q4 ~4.7GB

Fastest model at this size, with a permissive license.

See hardware-tier guide →

Qwen3 8B

8B paramsApache 2.0Multilingual

~8GB min • Q4 ~5.5GB

Strong general-purpose pick, especially outside English.

See run guide →

Phi-4-mini

3.8B paramsMIT

~4GB min • Q4 ~2.6GB

Small footprint, strong reasoning for its size class.

See hardware-tier guide →

Qwen2.5-Coder 7B

7B paramsApache 2.0Coding

~6GB min • Q4 ~4.7GB

Coding-focused variant of Qwen sized for entry hardware.

See coding guide →

Qwen3 14B

14B paramsApache 2.0

~14GB min • Q4 ~9.6GB

Community favorite at this tier for code-adjacent work.

See coding guide →

DeepSeek-R1-Distill 14B

14B paramsMIT

~16GB min • Q4 ~9.6GB

Distilled from DeepSeek's reasoning model; strong at math and logic chains.

See run guide →

Gemma 3 12B

12B paramsGemma LicenseImage input

~14GB min • Q4 ~8.2GB

Handles image input alongside text; well-rounded general use.

See hardware-tier guide →

Phi-4

14B paramsMIT

~14GB min • Q4 ~9.6GB

Smaller footprint than most peers at similar reasoning quality.

See hardware-tier guide →

Qwen3 32B

32B paramsApache 2.0

~24GB min • Q4 ~21.9GB

The clear step up from the 14B tier in general reasoning and coding.

See run guide →

Gemma 3 27B

27B paramsGemma LicenseImage input

~24GB min • Q4 ~18.5GB

Largest Gemma 3 size that comfortably fits a 24GB card.

See hardware-tier guide →

DeepSeek-R1-Distill 32B

32B paramsMIT

~32GB min • Q4 ~21.9GB

The strongest local reasoning most people can actually run.

See run guide →

Mistral Small 3.1

24B paramsApache 2.0128K context

~20GB min • Q4 ~16.4GB

Fast execution and strong translation, multimodal.

See hardware-tier guide →

Qwen2.5-Coder 32B

32B paramsApache 2.0Coding

~24GB min • Q4 ~21.9GB

The strongest coding-dedicated pick most people can realistically run locally.

See coding guide →

Devstral

24B paramsApache 2.0Agentic coding

~20GB min • Q4 ~16.4GB

Purpose-built for agentic, multi-step coding workflows rather than single-shot completion.

See coding guide →

Llama 4 Scout

109B total (MoE, ~17B active)Llama 4 Community License10M context

~60GB min • Q4 ~60GB

High-memory unified-memory Macs only. Full parameter set loads regardless of active-parameter count; see the tier guide for why.

See hardware-tier guide →

Llama 3.3 70B

70B paramsLlama Community License

~64GB min • Q4 ~49.1GB

Meaningful step up from the 8B version; needs a Mac Studio-class machine or a high-VRAM multi-GPU setup at Q4.

See run guide →

Model specs (parameter counts, licenses, approximate memory needs) are cross-checked against multiple current local-AI model guides and, where available, official model cards (see the fact-check log). Memory figures use the same Q4_K_M formula as the hardware calculator. This directory gets rechecked as new model versions ship, since local model releases move fast.

Sources

  1. Hugging Face (opens in new tab)
  2. daily.dev (opens in new tab)
  3. iRoyal (opens in new tab)

Frequently asked questions

How current is this directory?

It reflects widely-recommended local models as of this writing, cross-checked against multiple current sources rather than one static ranking. New models ship constantly in this space; this page gets rechecked periodically rather than left to age.

Why do the memory figures here differ slightly from what I've seen elsewhere?

Every source uses slightly different assumptions about quantization overhead and context length. The figures here use this site's own documented formula, visible on the hardware calculator page, so you can see exactly how they're computed rather than trusting a black-box number.

Can I filter by more than tier and use case?

Not yet. License and context window are shown as tags on each card so you can scan for them, even though they're not filterable dimensions yet.