← All articles
2026-08-1410 min read

How to run Qwen 3.8 27B locally, on a PC or a Mac

Run Qwen 3.8 27B locally: real file sizes for every quant, what it needs in VRAM or unified memory, the Mac route through MLX, and the setting that decides it.

Qwen 3.8 27B landed on Hugging Face on 14 August 2026 under Apache 2.0, and it is the model most people mean when they say they want something good running on their own machine. It is dense rather than a mixture of experts, so it fits on one graphics card. It reads images and video. Its context window is a quarter of a million tokens out of the box. And it is free to use commercially, which the licences on several of its rivals are not.

Below: what you are downloading, what it needs from your hardware, the Mac route, four ways to install it, and the setting that decides whether you get a great model or something that looks broken.

What Qwen 3.8 27B actually is

27.78 billion parameters across 64 layers, all trained in BF16. What makes it unusual is the attention layout. Most models this size run full attention on every layer, which is where the memory goes when a conversation gets long. Qwen 3.8 27B runs a cheaper linear kind, Gated DeltaNet, on three layers out of four, and saves proper attention for every fourth layer. Sixteen of its 64 layers do the expensive work. The config file spells this out as full_attention_interval: 4, and the consequence shows up further down in the memory arithmetic.

Vision is built in rather than bolted on. There is no separate VL variant to hunt for: the single repository carries a vision tower, and the model takes images and video as input. In GGUF form that tower travels in a separate file, which is the mmproj line in the table below.

Context is 262,144 tokens natively, and the model card documents stretching it to a full million with YaRN scaling. Two hundred and sixty thousand tokens is somewhere around a 600 page book in one sitting, which is more than most people will ever fill.

Qwen's own benchmark table puts it at 90.3 on LiveCodeBench, 89.2 on GPQA Diamond, 61.7 on SWE-bench Pro and 94.6 on MathVision. Those are the vendor's numbers on the vendor's harness, as launch benchmarks always are, so treat them as a shape rather than a ranking.

What it costs in disk and memory

Nobody runs the raw weights at home. You run a quantised build, which is the same model with the numbers stored more coarsely. The sizes below are the actual file sizes in the unsloth/Qwen3.8-27B-GGUF repository, read on 7 September 2026.

QuantFile sizeWho it is for
UD-IQ2_XXS7.3 GB8 GB cards, and quality takes a visible hit
UD-Q2_K_XL9.8 GB12 GB cards, almost no room left for context
UD-Q3_K_XL13.1 GB16 GB cards with room for context
UD-IQ4_XS14.3 GBthe largest one that fits a 16 GB card whole
UD-Q4_K_M16.5 GBthe default, start here
UD-Q5_K_M19.8 GB24 GB cards, slightly sharper
UD-Q6_K22.0 GB24 GB cards, close to lossless
Q8_029.0 GB32 GB and up, effectively lossless
BF1654.7 GBthe original, split across two files
mmproj-BF160.93 GBthe extra file that gives it eyes

Two traps in that table worth naming. First, Unsloth publishes almost everything as UD- dynamic quants, so there is no plain Q2_K or Q3_K_M in that repository and asking for one gets you nothing. Second, the same nominal quant is not the same size everywhere: ggml-org's plain Q4_K_M is 18.97 GB against Unsloth's 16.46 GB for UD-Q4_K_M. If you sized a 16 GB card against one number and downloaded the other, that is where your evening went.

UD-Q4_K_M at 16.5 GB is the one to start with. On a 24 GB card it loads whole with room for a long conversation. A 16 GB card is the awkward size: Q4_K_M spills into system RAM, which works but costs speed, and UD-IQ4_XS at 14.3 GB is the largest build that stays entirely on the card. Add the 0.93 GB projector on top if you want image input, and add the conversation cache on top of that. If you are not sure where your machine lands, our breakdown of how much RAM you need for local AI does the arithmetic properly.

On a Mac

Apple Silicon handles this model well, because unified memory means the GPU can reach all of it. The number that matters is how much unified memory your Mac has, and the honest tiers look like this: 24 GB is tight but workable at 4-bit if you close everything else, 32 GB is comfortable, and 48 GB or more lets you run a heavier quant or keep a second model warm alongside it.

There is no official Qwen statement about how much memory you need, so treat any specific figure you read as somebody's arithmetic rather than a spec. The nearest thing to a vendor number is LM Studio's own catalogue page, which lists a minimum of 16 GB of system memory for this model.

Two Mac-native routes exist beyond plain GGUF. MLX builds run on Apple's own framework and tend to be quicker on M-series chips: mlx-community/Qwen3.8-27B-4bit is 16.05 GB, lmstudio-community/Qwen3.8-27B-MLX-6bit is 22.78 GB, and mlx-community/Qwen3.8-27B-8bit is 29.50 GB. LM Studio lists MLX builds at 4, 5, 6 and 8 bits next to the GGUF ones, and Ollama now publishes MLX tags of its own. Which chip holds what, tier by tier, is in what a Mac can run locally in 2026.

The context window is cheaper than you think

This is where that attention layout pays. Every model keeps a cache of what it has read, and the cache grows with every token. The full attention layers here carry 4 key-value heads at a head dimension of 256, which is 4 KB per layer per token at 16-bit. A conventional 64 layer model with the same heads would pay that on all 64 layers: about 256 KB per token, so a 100,000 token document would want roughly 25 GB of memory on top of the model itself. That is why long context is so often theoretical.

Qwen 3.8 27B pays it on 16 layers, which comes to 64 KB per token. The same 100,000 token document costs about 6 GB. That is the difference between a spec sheet number and something you use on a Tuesday afternoon.

The one setting that decides everything

Qwen 3.8 ships with its own chat template, and the template is not decoration. It marks where your message ends and the model's answer begins, and it carries the switch for the reasoning.

Load the model without it and you get one of two failure modes, both of which look like broken weights rather than a broken setting. Either it rambles and never stops, or it answers in a strange clipped voice and forgets the conversation between turns. People report this as "the quant is bad" or "27B is overrated". It is neither.

In llama.cpp the fix is one flag, --jinja, which tells it to use the template packed inside the GGUF file. In LM Studio and Ollama the template comes along with the model, so this mostly bites people who fetched a bare GGUF by hand. If the answers look wrong, check the template before you blame the weights. Which tool handles this for you is part of our comparison of Ollama and LM Studio.

The sampling settings Qwen actually recommends

The model card publishes two sets, and they are not close to each other. Getting the second one wrong is the usual cause of a model that repeats itself.

SettingThinking onThinking off
temperature1.00.7
top_p0.950.80
top_k2020
presence_penalty0.01.5

That presence penalty of 1.5 with thinking off is the surprising one, and it is deliberate: without the reasoning pass in front of the answer, the model needs the nudge away from repeating itself.

Four ways to install it

In LU Labs, on Windows or Linux, the Model Manager carries the official 27B in Unsloth's dynamic quants, a 9B distill for smaller cards, and the Ollama tags. Picking a vision build downloads the projector next to the model file and the LU Engine starts llama.cpp with --mmproj pointed at it, so image input works without you ever meeting the flag. The whole first run, from hardware check to first answer, is in our guide to how to run AI locally.

Ollama is the shortest route on any platform. ollama pull qwen3.8 gets you the 27B at 18 GB, and there is no other size under that name. The library lists twelve tags at 256K context, all accepting text and images, including 27b-q8_0 at 30 GB, 27b-bf16 at 56 GB, and MLX tags for Apple Silicon.

LM Studio is worth it for the model browser. The catalogue entry is qwen/qwen3.8-27b, with GGUF and MLX builds at several bit depths, and the app tells you whether a build fits before you spend the download.

If you would rather not have an app in the way, Unsloth's card gives the llama.cpp one-liner: llama serve -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M. Add --jinja and you are done.

Giving it eyes

Vision lives in the mmproj file, and it is under a gigabyte. Point your runtime at both files and the model can look at screenshots, photos, diagrams and video frames. Skip it and you have a very good text model that will politely tell you it cannot see the image you just pasted. Ollama's tags carry the projector already, so this only comes up if you assemble the files yourself.

Thinking, and turning it off

Thinking is on by default here. That is right for a hard question and wasteful for "what is the capital of Portugal", so Qwen made it a switch: pass enable_thinking: false in the chat template arguments and the reasoning pass goes away. There is also a reasoning_effort control with low, medium and xhigh, where xhigh is the default.

Locally the cost of leaving it on is not money, it is the minutes you spend watching tokens arrive, which for most people is the more annoying currency. Turn it down for quick questions and leave it up for the ones that deserve it.

When 27B is not the Qwen you wanted

The two Qwen 3.8 flagships are trillion-parameter models and no quantisation brings them to a desk. If what you came for is Qwen 3.8 Max or A95B rather than the 27B, they run on our GPUs and answer in a browser tab: run Qwen 3.8 without a GPU has the prices and what a credit pack buys. Every plan and every credit pack reaches the whole chat catalogue, so there is no tier to climb for them. Which of the two flagships suits which job is in Qwen 3.8 Max vs A95B.

The split most people settle into is straightforward. Run the 27B at home when the material should not leave your machine, when you want it offline, and when you would rather pay once for a card than per token forever. Reach for the big ones when the problem is genuinely harder than your hardware. For where the 27B sits among everything else you could download, our roundup of the best open-weight LLMs of 2026 puts it in context.

Sources