Ranked guide

Local / Private AI — Your Brain, Your Machine, Your Rules

Here's a radical idea: what if you could run a genuinely smart AI on your own hardware, and nothing you told it would ever leave your machine? No cloud servers. No data collection. No subscription fees. Just you, your laptop, and an intelligence that respects your privacy by design. Welcome to the open-weight revolution.

Decision first

Our ranking

Start with the winner, then compare the trade-offs that might change the answer for you.

#1 Local / Private AI

Qwen3.8 — 27B

Alibaba (Qwen Team)

Qwen3.8-27B is the rare local model whose compromises line up with hardware people actually own: one dense 27B checkpoint for text, images, video, coding, and tool-using agents. A 4-bit build fits on a 24 GB-class GPU, while the Apache 2.0 license keeps the work private and commercially usable.

Why It Wins

Qwen reports 61.7 on SWE-bench Pro, 73.0 on Terminal Bench 2.1, 42.2 on DeepSWE 1.1, 84.3 on OSWorld-Verified, and 70.7 on CoWorkBench. It has 262,144-token native context, optional YaRN scaling toward 1M, adjustable reasoning effort, preserved thinking across turns, and native image and video understanding.

The Catch

This is launch-day evidence, not settled truth. Most headline scores come from Qwen's own table; some use the Claude Code harness, corrected tasks, special prompts, or Qwen-created benchmarks. The common Q4_K_M package is roughly 18 GB including its vision projector, but runtime overhead and KV cache still limit how much context fits on a 24 GB card.

9.2 Editorial score
Read review
Best for

Qwen3.8-27B is the rare local model whose compromises line up with hardware people actually own: one dense 27B checkpoint for text, images, video, coding, and tool-using agents. A 4-bit build fits on a 24 GB-class GPU, while the Apache 2.0 license keeps the work private and commercially usable.

Why It Wins

Qwen reports 61.7 on SWE-bench Pro, 73.0 on Terminal Bench 2.1, 42.2 on DeepSWE 1.1, 84.3 on OSWorld-Verified, and 70.7 on CoWorkBench. It has 262,144-token native context, optional YaRN scaling toward 1M, adjustable reasoning effort, preserved thinking across turns, and native image and video understanding.

Watch out

This is launch-day evidence, not settled truth. Most headline scores come from Qwen's own table; some use the Claude Code harness, corrected tasks, special prompts, or Qwen-created benchmarks. The common Q4_K_M package is roughly 18 GB including its vision projector, but runtime overhead and KV cache still limit how much context fits on a 24 GB card.

#2

A phenomenally efficient 284B/13B active parameter MoE model that dominates agentic and coding workflows while fitting comfortably on mid-tier hardware.

9.1 Editorial score
Read review
#3

GLM-5.3

Z.ai (Zhipu AI)

The most interesting local model you cannot download yet. GLM-5.3 keeps GLM-5.2's 744B mixture-of-experts base and extracts a large coding-and-agent leap through post-training alone. Z.ai promises weights about two weeks after launch; until they arrive, this is a hosted model with an unusually credible self-hosted future.

9.2 Editorial score
Read review
#4

Kimi K3

Moonshot AI

The first open-weight model that looks like a closed frontier brain. 2.8 trillion parameters of mixture-of-experts, native vision, a full million-token context, and a license that lets you run it commercially — all downloadable to a machine you control. The catch is the machine: it needs a datacenter, not a desk.

8.5 Editorial score
Read review
#5

Gemma 4

Google DeepMind

Not one model — five. Google DeepMind's Gemma 4 is a family spanning everything from a 2-billion-parameter sliver that runs on your phone to a 31-billion-parameter powerhouse for servers. Each member has different architecture, different strengths, and different hardware requirements. The E2B fits in 1 GB of RAM. The 12B Unified runs a full multimodal AI on a laptop GPU. The 26B MoE activates only 3.8B parameters per token. All Apache 2.0, all open weights. This guide walks through each one so you know exactly which Gemma fits your hardware and your workflow.

8.2 Editorial score
Read review
Questions, answered

Frequently Asked Questions