{"id":1000039,"date":"2026-10-11T13:43:00","date_gmt":"2026-10-11T12:43:00","guid":{"rendered":"https:\/\/siyaz.tech\/?p=1000039"},"modified":"2026-10-11T13:43:00","modified_gmt":"2026-10-11T12:43:00","slug":"local-ai-models-gguf-quantization-guide","status":"publish","type":"post","link":"https:\/\/siyaz.tech\/index.php\/2026\/10\/11\/local-ai-models-gguf-quantization-guide\/","title":{"rendered":"GGUF, Q4_K_M and Other Things That Look Like Wi-Fi Passwords: A Field Guide to Running AI Locally"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">SARCASM WARNING! (Also: maths warning. There are numbers in here, and annoyingly, they&#8217;re the useful part.)<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Well hello there, my RAM-anxious friends! In <a href=\"https:\/\/siyaz.tech\/index.php\/2026\/10\/11\/kaggle-free-ai-server-ollama-cloudflare-tunnel\/\">my last post<\/a> I pulled a model ending in <code>Q4_K_M<\/code>, called it &#8220;the model&#8217;s travel-size shampoo&#8221;, and sprinted off to talk about tunnels. That was lazy of me. Today we pay that debt.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Because here&#8217;s what happens to everyone the first time. You find the model the whole internet is excited about. You click through to download it. And instead of a nice friendly button, you get this:<\/p>\n\n\n\n<pre class=\"wp-block-code\"><code>model-Q2_K.gguf        3.2 GB\nmodel-Q3_K_M.gguf      4.0 GB\nmodel-Q4_0.gguf        4.7 GB\nmodel-Q4_K_M.gguf      4.9 GB\nmodel-IQ4_XS.gguf      4.5 GB\nmodel-Q5_K_M.gguf      5.7 GB\nmodel-Q6_K.gguf        6.6 GB\nmodel-Q8_0.gguf        8.5 GB<\/code><\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Same model. Eight files. Names that look like Wi-Fi passwords from a hotel in 2009. Nobody explains the difference, so you do what every sensible adult does when faced with a wine list in a language they don&#8217;t speak: pick the one in the middle and hope.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That works more often than it should. But by the end of this post you&#8217;ll read those filenames like a sommelier reads a label, and you&#8217;ll know exactly why the middle one is usually right, and when it very much isn&#8217;t.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83c\udfe0 Why Run It Yourself When the Cloud Is Right There?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Let&#8217;s get the uncomfortable truth out of the way first. The big hosted models are better at hard problems. Anyone telling you the model on your gaming PC beats them at serious reasoning is selling you a course.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But &#8220;the smartest possible answer&#8221; isn&#8217;t the only thing that matters. Running a model locally buys you four things no API will ever sell you:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Your data stays home.<\/strong> Contracts, client code, incident logs, that spreadsheet called <code>FINAL_v7_REALLY_FINAL.xlsx<\/code>. None of it leaves the machine. Your DPO might actually smile.<\/li>\n\n\n<li><strong>No meter running.<\/strong> Once you own the hardware, a million tokens costs you some electricity and a slightly warmer room.<\/li>\n\n\n<li><strong>Nobody changes it behind your back.<\/strong> The file you download today behaves exactly the same in three years. No surprise &#8220;improvements&#8221;, no deprecation emails, no new rate limits.<\/li>\n\n\n<li><strong>It works offline.<\/strong> On a plane. In an air-gapped lab. During the next big cloud outage, while everyone else refreshes a status page.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The catch? You give up some raw brainpower, and you become your own infrastructure team. Congratulations on the promotion. There is no pay rise.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83e\udde0 A Model Is Just a Very, Very Large Spreadsheet<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Strip away the hype and a language model is, physically, a giant pile of numbers called <strong>weights<\/strong>. An &#8220;8B&#8221; model has about 8 billion of them. A &#8220;70B&#8221; has 70 billion. That&#8217;s it. No soul. No consciousness. Just the world&#8217;s most expensive spreadsheet.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Those weights normally ship in 16-bit precision: two bytes each. So the maths for the original, untouched model is brutal:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>8 billion \u00d7 2 bytes = 16 GB<\/strong><\/li>\n\n\n<li><strong>70 billion \u00d7 2 bytes = 140 GB<\/strong><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Most consumer graphics cards have between 8 and 24 GB of memory. A full-fat 70B model doesn&#8217;t fit on any of them. It doesn&#8217;t fit on most workstations. It barely fits in the budget meeting.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The fix is <strong>quantization<\/strong>: store each weight with fewer bits (8, 5, 4, even 2) and accept a little rounding error in exchange for a model a quarter of the size.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Think of it as <strong>JPEG compression for brains<\/strong>. A RAW photo is huge and perfect. A high-quality JPEG is a fraction of the size and nobody can tell. Crank the compression and you start getting blocky smudges round the edges. Crank it further and your holiday photos look like they were painted by a potato.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Model weights behave the same way. Light compression: invisible. Heavy compression: the model gets quietly dumber. It forgets a fact here, mangles a code block there, wanders off-topic in long answers. Extreme compression: fluent, confident nonsense.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">And here&#8217;s where the analogy breaks, which is the important bit. A badly compressed JPEG <em>looks<\/em> bad. You notice instantly. A badly compressed model doesn&#8217;t look blurry at all. It looks like a consultant: articulate, confident, and wrong. You only find the damage when you check its work, which is exactly why you need to know what you&#8217;re downloading.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83d\udce6 GGUF: Flat-Pack Furniture, Allen Key Included<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>GGUF<\/strong> is the file format from the <a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\">llama.cpp<\/a> project, introduced in August 2023 to replace its older GGML format. It&#8217;s what sits underneath Ollama, LM Studio, koboldcpp, Jan, and most of the &#8220;just run the thing&#8221; tools.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Its superpower isn&#8217;t compression. It&#8217;s that <strong>everything comes in one box<\/strong>:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>the quantized weights;<\/li>\n\n\n<li>the architecture (how many layers, how attention is wired);<\/li>\n\n\n<li>the tokenizer (how text gets chopped into pieces the model understands);<\/li>\n\n\n<li>the chat template (how a conversation is laid out so the model knows who&#8217;s talking);<\/li>\n\n\n<li>metadata like the recommended context length.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Before GGUF, running a model meant juggling a weights file, a tokenizer file, a config file, and a prayer that they all matched. Last time, IKEA delivered the wardrobe without the Allen key. GGUF is the version where the key, the screws, and the instructions are all taped inside the box.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Its other killer feature is <strong>partial GPU offload<\/strong>. A GGUF model can run on the CPU, the GPU, or <em>both at once<\/em>: put as many layers as fit on the graphics card and run the rest from ordinary RAM. That&#8217;s why GGUF will run (slowly) on hardware that &#8220;can&#8217;t fit&#8221; the model, while most GPU-only formats just throw an out-of-memory error and walk off.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">\ud83d\uddc2\ufe0f The Other Formats You&#8217;ll Trip Over<\/h3>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Format<\/th><th>What it is<\/th><th>Use it when<\/th><th>Vibe<\/th><\/tr><\/thead><tbody><tr><td><strong>Safetensors<\/strong> (FP16\/BF16)<\/td><td>The original, uncompressed weights<\/td><td>Fine-tuning, research, datacentre budgets<\/td><td>RAW photo<\/td><\/tr><tr><td><strong>GGUF<\/strong><\/td><td>llama.cpp&#8217;s all-in-one quantized file<\/td><td>CPU, Apple Silicon, mixed CPU+GPU, &#8220;I just want it to work&#8221;<\/td><td>Swiss Army knife<\/td><\/tr><tr><td><strong>MLX<\/strong><\/td><td>Apple&#8217;s native format<\/td><td>Apple Silicon, where it&#8217;s often a bit faster than GGUF<\/td><td>Turtleneck<\/td><\/tr><tr><td><strong>GPTQ \/ AWQ<\/strong><\/td><td>GPU-only 4-bit methods<\/td><td>Serving with vLLM on NVIDIA cards<\/td><td>Server room<\/td><\/tr><tr><td><strong>EXL2 \/ EXL3<\/strong><\/td><td>ExLlama&#8217;s flexible-bitrate formats<\/td><td>Model fits entirely in NVIDIA VRAM and you want maximum speed<\/td><td>Racing stripes<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">One person, one machine? GGUF is the right default. Serving a whole team from one GPU box? Look at vLLM with AWQ, GPTQ or FP8 instead. GGUF was built for your desk, not for a hundred users at once.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83d\udd2c Q4_K_M: A Letter-by-Letter Autopsy<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Back to that wall of files. Every suffix is a recipe. Let&#8217;s dissect one.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">The number: bits per weight<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><code>Q4<\/code> means each weight is stored in <em>roughly<\/em> 4 bits. <code>Q8<\/code> means roughly 8. <code>Q2<\/code> means roughly 2 and a lot of regret. Smaller number, smaller file, more rounding error. Hold on to the word &#8220;roughly&#8221;. It comes back to bite later.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><code>_0<\/code> and <code>_1<\/code>: the boomer formats<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\"><code>Q4_0<\/code>, <code>Q4_1<\/code>, <code>Q5_0<\/code> and <code>Q8_0<\/code> are the original llama.cpp recipes. They chop the weights into blocks of 32 and give each block one shared scale (<code>_0<\/code>), or a scale plus an offset (<code>_1<\/code>).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Picture a group photo where everyone has to share one exposure setting. Simple, fast, fine on a good day. But if one person in the block is an outlier (wearing a high-vis vest, say), everyone else gets rounded badly to make room for them. Just like a shared office thermostat.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><code>Q8_0<\/code> is still excellent, because at 8 bits there&#8217;s precision to spare. But <code>Q4_0<\/code> and <code>Q4_1<\/code> are retirement material. The next family beats them at the same size.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><code>_K<\/code>: the K-quants<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The K-quants (<code>Q2_K<\/code> to <code>Q6_K<\/code>), added to llama.cpp in mid-2023, are smarter. They group weights into <strong>super-blocks of 256<\/strong>, split those into smaller sub-blocks, give each sub-block its own scale, and then compress <em>the scales themselves<\/em> to claw back space. Compression inside compression. Very Inception.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Back to the group photo: now every small cluster of people gets its own exposure. One person in high-vis no longer ruins it for everyone else. Same file size, measurably better accuracy.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><code>_S<\/code>, <code>_M<\/code>, <code>_L<\/code>: coffee sizes, except medium is actually the right answer<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">This is the bit almost nobody explains. A model isn&#8217;t one uniform blob. It&#8217;s hundreds of separate weight matrices (tensors), and <strong>some of them matter far more than others<\/strong>. Certain attention and feed-forward tensors are drama queens: round them a little too hard and the whole model suffers.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">So the last letter tells you how generous the recipe is with the drama queens:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong><code>_S<\/code> (small):<\/strong> nearly everything at the base bit-width. Smallest file, most damage.<\/li>\n\n\n<li><strong><code>_M<\/code> (medium):<\/strong> the sensitive tensors get bumped up a tier. In <code>Q4_K_M<\/code>, some of them get 6 bits. This is the sweet spot.<\/li>\n\n\n<li><strong><code>_L<\/code> (large):<\/strong> even more tensors upgraded. Only offered on low-bit quants like <code>Q3_K_L<\/code>, where every extra bit is worth fighting for.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">That&#8217;s why <code>Q4_K_M<\/code> became the community default, and the one Ollama usually hands you: it spends the extra bits exactly where they buy the most brain.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\"><code>IQ<\/code>: the i-quants (for when you&#8217;re desperate)<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Then there&#8217;s the newer family with names that sound like rejected Star Wars droids: <code>IQ1_S<\/code>, <code>IQ2_XXS<\/code>, <code>IQ3_XXS<\/code>, <code>IQ3_M<\/code>, <code>IQ4_XS<\/code>, <code>IQ4_NL<\/code>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Instead of rounding each weight to the nearest mark on a ruler, i-quants snap <em>groups<\/em> of weights to the nearest entry in a carefully designed codebook, a lattice of patterns chosen to match what real weights look like. It&#8217;s cleverer packing, and it shines exactly where K-quants start falling apart: <strong>below 4 bits<\/strong>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Two catches, because there are always catches:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Speed depends on your hardware.<\/strong> Unpacking a codebook is more work per weight. On a GPU you&#8217;ll barely notice. On an older CPU, an i-quant can be noticeably slower than a K-quant of the same size.<\/li>\n\n\n<li><strong>They really want an imatrix<\/strong> (next section). An i-quant made without one is a wasted opportunity.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\">imatrix: free quality you didn&#8217;t know you had<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">You&#8217;ll see repos labelled &#8220;imatrix quants&#8221;, or files with <code>i1<\/code> in the name. An <strong>importance matrix<\/strong> is made by running sample text through the full-precision model and recording which weights actually do the heavy lifting. The quantizer then protects those and squashes the freeloaders.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Same file size, less damage, biggest payoff at low bit-widths. It&#8217;s like discovering your credit card came with airport lounge access all along. When you have the choice, take the imatrix version.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83d\udcc9 So How Much Dumber Does It Actually Get?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Enough metaphors. Receipts. Here&#8217;s llama.cpp&#8217;s own reference table for <strong>Llama 3 8B<\/strong>, straight from the source code of its quantization tool. &#8220;Perplexity increase&#8221; measures how much more <em>surprised<\/em> the compressed model is by real text than the original. Lower is better. Zero means no measurable damage.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Quant<\/th><th>File size<\/th><th>Real bits\/weight<\/th><th>Perplexity increase<\/th><th>Verdict<\/th><\/tr><\/thead><tbody><tr><td><strong>Q8_0<\/strong><\/td><td>8.5 GB<\/td><td>8.5<\/td><td>+0.003<\/td><td>Identical twin<\/td><\/tr><tr><td><strong>Q6_K<\/strong><\/td><td>6.6 GB<\/td><td>6.6<\/td><td>+0.022<\/td><td>Effectively lossless<\/td><\/tr><tr><td><strong>Q5_K_M<\/strong><\/td><td>5.7 GB<\/td><td>5.7<\/td><td>+0.057<\/td><td>You&#8217;d need a lab to tell<\/td><\/tr><tr><td><strong>Q4_K_M<\/strong><\/td><td>4.9 GB<\/td><td>4.9<\/td><td>+0.175<\/td><td>The sweet spot<\/td><\/tr><tr><td>Q4_K_S<\/td><td>4.7 GB<\/td><td>4.7<\/td><td>+0.269<\/td><td>Fine if every MB counts<\/td><\/tr><tr><td>Q4_0<\/td><td>4.7 GB<\/td><td>4.6<\/td><td>+0.469<\/td><td>Same size as Q4_K_S, twice the damage. Why?<\/td><\/tr><tr><td><strong>Q3_K_M<\/strong><\/td><td>4.0 GB<\/td><td>4.0<\/td><td>+0.657<\/td><td>Starting to forget things<\/td><\/tr><tr><td>Q2_K<\/td><td>3.2 GB<\/td><td>3.2<\/td><td>+3.520<\/td><td>Lights on, nobody home<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Source: <code>tools\/quantize\/quantize.cpp<\/code> in llama.cpp. Sizes converted from the GiB figures listed there.<\/em><\/p>\n\n\n\n<figure class=\"szt-diagram\" style=\"margin:2em 0\"><svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" viewBox=\"0 0 760 360\" role=\"img\" aria-label=\"Quality loss by quantization level for Llama 3 8B: tiny from Q8_0 to Q4_K_M, then a cliff at Q3 and Q2\" style=\"width:100%;height:auto;font-family:system-ui,-apple-system,'Segoe UI',Roboto,sans-serif\" font-size=\"13\">\n<rect x=\"0\" y=\"0\" width=\"760\" height=\"360\" rx=\"10\" fill=\"#ffffff\"\/>\n<text x=\"24\" y=\"34\" font-size=\"15\" font-weight=\"600\" fill=\"#1f2328\">Quality holds, holds, holds\u2026 then falls off a cliff<\/text>\n<text x=\"24\" y=\"54\" font-size=\"11.5\" fill=\"#57606a\">Perplexity increase vs the original model, Llama 3 8B (lower is better)<\/text>\n<g fill=\"#1f2328\" font-weight=\"600\" text-anchor=\"end\">\n<text x=\"100\" y=\"91\">Q8_0<\/text>\n<text x=\"100\" y=\"125\">Q6_K<\/text>\n<text x=\"100\" y=\"159\">Q5_K_M<\/text>\n<text x=\"100\" y=\"193\">Q4_K_M<\/text>\n<text x=\"100\" y=\"227\">Q4_0<\/text>\n<text x=\"100\" y=\"261\">Q3_K_M<\/text>\n<text x=\"100\" y=\"295\">Q2_K<\/text>\n<\/g>\n<line x1=\"112\" y1=\"70\" x2=\"112\" y2=\"306\" stroke=\"#8a8f98\" stroke-width=\"1.25\"\/>\n<rect x=\"112\" y=\"78\" width=\"2\" height=\"18\" fill=\"#8a8f98\"\/>\n<rect x=\"112\" y=\"112\" width=\"3\" height=\"18\" fill=\"#8a8f98\"\/>\n<rect x=\"112\" y=\"146\" width=\"8\" height=\"18\" fill=\"#8a8f98\"\/>\n<rect x=\"112\" y=\"180\" width=\"25\" height=\"18\" rx=\"2\" fill=\"#dbe7fb\" stroke=\"#2f6fde\" stroke-width=\"2\"\/>\n<rect x=\"112\" y=\"214\" width=\"67\" height=\"18\" fill=\"#f3f4f6\" stroke=\"#8a8f98\" stroke-width=\"1.25\" stroke-dasharray=\"5 4\"\/>\n<rect x=\"112\" y=\"248\" width=\"94\" height=\"18\" fill=\"#8a8f98\"\/>\n<rect x=\"112\" y=\"282\" width=\"500\" height=\"18\" fill=\"#8a8f98\"\/>\n<g font-size=\"11.5\" fill=\"#57606a\">\n<text x=\"122\" y=\"91\">+0.003  (8.5 GB)<\/text>\n<text x=\"123\" y=\"125\">+0.022  (6.6 GB)<\/text>\n<text x=\"128\" y=\"159\">+0.057  (5.7 GB)<\/text>\n<text x=\"145\" y=\"193\" fill=\"#1f2328\" font-weight=\"600\">+0.175  (4.9 GB)  \u2190 the sweet spot<\/text>\n<text x=\"187\" y=\"227\">+0.469  (4.7 GB)  legacy format, same size as Q4_K_S<\/text>\n<text x=\"214\" y=\"261\">+0.657  (4.0 GB)<\/text>\n<text x=\"620\" y=\"295\">+3.520  (3.2 GB)<\/text>\n<\/g>\n<text x=\"24\" y=\"336\" font-size=\"11.5\" fill=\"#57606a\">Bars drawn to scale. Data: llama.cpp, tools\/quantize\/quantize.cpp.<\/text>\n<\/svg><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Three things jump out, and they&#8217;re the whole point of this post:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>One: the damage is not a straight line.<\/strong> From 8 bits down to 5 costs almost nothing. From 4 to 3 costs more than everything above it put together. From 3 to 2, the damage goes up five-fold. Quality doesn&#8217;t leak out slowly. It holds, holds, holds, then falls off a cliff. Much like a project timeline.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Two: look at <code>Q4_0<\/code> next to <code>Q4_K_S<\/code>.<\/strong> Practically the same size. The old format does almost twice the damage. That single row is the entire case for K-quants.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Three: remember &#8220;roughly&#8221;?<\/strong> Here it is. <code>Q4_K_M<\/code> actually lands at 4.9 bits per weight, not 4. The upgraded tensors, the stored scales and the higher-precision embedding and output layers all add up. When you&#8217;re working out whether something fits, use the real number, not the label. The label is marketing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">One honest caveat before anyone tattoos that table on their arm. Perplexity is a blunt instrument. It averages over a lot of text, and it can hide damage that&#8217;s concentrated in one skill: maths, code, long documents, non-English languages. Serious quantizers increasingly also report <strong>KL divergence<\/strong>, which measures how far the compressed model&#8217;s predictions drift from the original&#8217;s, token by token. If one task matters to you, test that task. Don&#8217;t trust any single number. Including mine.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83d\udc18 Big and Squashed Beats Small and Pristine<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Here&#8217;s the decision people actually face: <em>I&#8217;ve got 24 GB. Do I run a 14B model at Q8, or a 32B model at Q4?<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Instinct says &#8220;the higher-quality file wins&#8221;. Instinct is usually wrong. A bigger model at a sensible quant generally beats a smaller model at high precision, because the bigger model simply <em>knows more and reasons better<\/em>, and <code>Q4_K_M<\/code> only shaves a sliver off that. A slightly blurry photo of the Mona Lisa still beats a crystal-clear photo of a stick figure.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The rule of thumb:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>Pick the biggest model that fits at Q4_K_M or better. Only go below Q4 when the jump in model size is large.<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Where the rule breaks: below about 3 bits the cliff is so steep that a crushed big model can lose to a clean small one. And a model from a newer generation often beats an older one twice its size outright. &#8220;Bigger wins&#8221; means bigger <em>within the same family and era<\/em>, not &#8220;grab the 2023 70B because the number is larger&#8221;. That&#8217;s how you end up with a very large, very confident dinosaur.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">\ud83e\udde9 Mixture-of-Experts: the plot twist<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">More and more open models are <strong>Mixture-of-Experts (MoE)<\/strong>: Mixtral, DeepSeek, the Qwen3 MoE models, OpenAI&#8217;s gpt-oss. They come with two numbers, like <em>30B-A3B<\/em>: 30 billion parameters in total, but only about 3 billion <strong>active<\/strong> for any single token. It&#8217;s a company of 30 specialists where only three turn up to each meeting. Which, to be fair, is also how most companies work.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Memory<\/strong> scales with the <em>total<\/em>. You still have to store all 30B somewhere.<\/li>\n\n\n<li><strong>Speed<\/strong> scales with the <em>active<\/em> count. Each token only touches about 3B of them.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">So an MoE model needs the memory of a 30B model but writes at roughly the speed of a 3B one. That makes them brilliant for machines with lots of ordinary RAM and a modest GPU. llama.cpp can even keep the bulky expert weights in system RAM while the always-used layers sit on the GPU.<\/p>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83e\uddee Will It Fit? Maths You Can Do on a Napkin<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Step 1: the weights<\/h3>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>Size in GB \u2248 billions of parameters \u00d7 real bits per weight \u00f7 8<\/strong><\/p>\n<\/blockquote>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Model size<\/th><th>Q4_K_M (~4.9 bits)<\/th><th>Q8_0 (~8.5 bits)<\/th><\/tr><\/thead><tbody><tr><td>8B<\/td><td>~4.9 GB<\/td><td>~8.5 GB<\/td><\/tr><tr><td>14B<\/td><td>~8.6 GB<\/td><td>~15 GB<\/td><\/tr><tr><td>27B<\/td><td>~16.5 GB<\/td><td>~29 GB<\/td><\/tr><tr><td>32B<\/td><td>~19.6 GB<\/td><td>~34 GB<\/td><\/tr><tr><td>70B<\/td><td>~43 GB<\/td><td>~74 GB<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Spot the 27B row? That&#8217;s where last post&#8217;s &#8220;about 17 GB&#8221; came from: 27 billion \u00d7 4.9 bits \u00f7 8 \u2248 16.5 GB. And that&#8217;s why one 15 GB Kaggle T4 couldn&#8217;t hold it and two could. See? Napkin maths. Not witchcraft.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Step 2: the KV cache (the bit that ambushes everyone)<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">While the model reads and writes, it keeps a running memory of every token in the conversation. That&#8217;s the <strong>KV cache<\/strong>, and it grows with every token. On long conversations it can rival the model itself.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For Llama 3 8B at the default 16-bit precision, the cache costs <strong>128 KB per token<\/strong>. Sounds tiny. Now multiply:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>8,000 tokens \u2192 <strong>about 1 GB<\/strong><\/li>\n\n\n<li>32,000 tokens \u2192 <strong>about 4 GB<\/strong><\/li>\n\n\n<li>128,000 tokens \u2192 <strong>about 16 GB<\/strong>, more than three times the size of the Q4_K_M model itself<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">This is the number one reason a model that &#8220;definitely fits&#8221; crashes or crawls the moment you paste in a long document. It&#8217;s also why Ollama ships with a small default context: a bigger window costs real memory. Last time we bumped it to 32K with <code>OLLAMA_CONTEXT_LENGTH<\/code>. Now you know what that bump costs.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Pro tip: llama.cpp can compress the cache too. <code>--cache-type-k q8_0 --cache-type-v q8_0<\/code> roughly halves it for very little quality loss. Ollama exposes the same thing through <code>OLLAMA_KV_CACHE_TYPE=q8_0<\/code>, with flash attention switched on.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Step 3: leave headroom<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Add 1 to 2 GB for the runtime, its scratch buffers and, if it&#8217;s also your display card, your desktop. A model that fits with 100 MB to spare will find a way not to fit. It&#8217;s like packing a suitcase to exactly 23.0 kg: the airport scale will disagree.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">\ud83d\udda5\ufe0f Where to start, by hardware<\/h3>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>What you&#8217;ve got<\/th><th>Comfortable choice<\/th><\/tr><\/thead><tbody><tr><td>8 GB VRAM \/ 16 GB laptop<\/td><td>7\u20139B at Q4_K_M to Q5_K_M<\/td><\/tr><tr><td>12\u201316 GB VRAM<\/td><td>12\u201314B at Q4_K_M to Q6_K<\/td><\/tr><tr><td>24 GB VRAM<\/td><td>~30B-class at Q4_K_M, or 14B at Q8_0 with a long context<\/td><\/tr><tr><td>32\u201364 GB Mac (unified memory)<\/td><td>30B-class at Q6_K to Q8_0; 70B-class at Q3\u2013Q4 on the bigger configs<\/td><\/tr><tr><td>64 GB+ system RAM, modest GPU<\/td><td>Large MoE models with the experts parked in RAM<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83d\udc0c Why Is It So Slow? (It&#8217;s the Memory, Not the Maths)<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Local AI has two speeds, and they&#8217;re limited by completely different things.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Prompt processing<\/strong> (reading your input) is limited by raw compute. GPUs flatten it; CPUs suffer. That&#8217;s the awkward silence after you paste a 20-page PDF into a CPU-only setup.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Token generation<\/strong> (writing the answer) is limited by <strong>memory bandwidth<\/strong>. To produce each token, the machine has to read essentially every active weight from memory, once. Every. Single. Token. So the speed limit is roughly:<\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p class=\"wp-block-paragraph\"><strong>Max tokens per second \u2248 memory bandwidth \u00f7 size of the model in memory<\/strong><\/p>\n<\/blockquote>\n\n\n\n<p class=\"wp-block-paragraph\">Plug in some real hardware for a 4.9 GB Q4_K_M 8B model:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Dual-channel DDR5 desktop (~80 GB\/s): ceiling \u2248 <strong>16 tokens\/s<\/strong><\/li>\n\n\n<li>Apple M-series Max chip (~400\u2013550 GB\/s): ceiling \u2248 <strong>80\u2013110 tokens\/s<\/strong><\/li>\n\n\n<li>RTX 4090 (~1,000 GB\/s): ceiling \u2248 <strong>200 tokens\/s<\/strong><\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Real numbers land below these ceilings, but the ratios hold. And remember those two Kaggle T4s that &#8220;type slowly&#8221;? Each one moves about 320 GB\/s, and with the model split across both cards they take turns, not work in parallel. 320 \u00f7 16.5 GB gives a ceiling of roughly 19 tokens a second before anything else slows it down. Mystery solved. It was never lazy. It was reading 16.5 GB for every word.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That one formula explains three things that confuse almost everyone:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Smaller quants are faster.<\/strong> Fewer bytes to read per token. Q4 writes faster than Q8 on the same machine, full stop.<\/li>\n\n\n<li><strong>Apple Silicon punches above its weight.<\/strong> Unified memory gives a laptop chip the kind of bandwidth that normally needs a separate graphics card, and lets it hold models no consumer GPU can.<\/li>\n\n\n<li><strong>Partial offload hurts more than you&#8217;d expect.<\/strong> Every layer left in system RAM is read at system-RAM speed. Put 90% of a model on the GPU and you don&#8217;t get 90% of GPU speed. The slow 10% sets the pace, like the one person in the group project who hasn&#8217;t opened the document yet.<\/li>\n<\/ul>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83e\udea4 Five Ways to Make a Great Model Look Stupid<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Most &#8220;this local model is rubbish&#8221; complaints aren&#8217;t about the model at all. They&#8217;re about one of these:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>The wrong chat template.<\/strong> Every model family expects a conversation laid out a specific way. Get it wrong and a brilliant model rambles, answers its own questions, or never shuts up. GGUF carries the right template inside the file, so use tools that read it, and be suspicious of hand-rolled prompt formats.<\/li>\n\n\n<li><strong>A base model instead of an instruct model.<\/strong> A base model just continues text; it doesn&#8217;t follow instructions. Ask it a question and it may reply with three more questions. For chat, look for <code>Instruct<\/code>, <code>Chat<\/code> or <code>it<\/code> in the name.<\/li>\n\n\n<li><strong>The goldfish context window.<\/strong> We covered this last time and it&#8217;s still the classic. Small default context, long document, and the beginning silently falls off the edge. Set the context length yourself, and budget the KV cache for it.<\/li>\n\n\n<li><strong>Over-squashing a small model.<\/strong> Small models have less spare capacity to absorb rounding errors. A 3B model at Q2_K is rarely worth the disk space. A 70B at Q3_K_M can be.<\/li>\n\n\n<li><strong>Trusting one benchmark.<\/strong> Test on <em>your<\/em> work: your codebase, your documents, your language. Ten minutes comparing two quants on real prompts beats an hour of scrolling leaderboards. Leaderboards are the LinkedIn of AI: everyone looks great on there.<\/li>\n<\/ol>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83d\udccb The Cheat Sheet (For Those Who Scrolled Straight Here)<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">I see you. I respect you. Here&#8217;s the whole post:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li><strong>Start with <code>Q4_K_M<\/code>.<\/strong> Best quality per gigabyte for most people.<\/li>\n\n\n<li><strong>Memory to spare? Go up to <code>Q5_K_M<\/code> or <code>Q6_K<\/code>.<\/strong> <code>Q8_0<\/code> is essentially lossless, but rarely worth the extra size over <code>Q6_K<\/code>.<\/li>\n\n\n<li><strong>Need to go under 4 bits? Use <code>IQ<\/code> quants with an imatrix,<\/strong> not <code>Q2_K<\/code> or <code>Q3_K_S<\/code>.<\/li>\n\n\n<li><strong>Skip <code>Q4_0<\/code> and <code>Q4_1<\/code><\/strong> unless a specific tool or chip needs them.<\/li>\n\n\n<li><strong>Bigger model at Q4 beats smaller model at Q8,<\/strong> within the same family and generation.<\/li>\n\n\n<li><strong>Memory budget = weights + KV cache + 1\u20132 GB of headroom.<\/strong><\/li>\n\n\n<li><strong>Speed \u2248 memory bandwidth \u00f7 model size.<\/strong> Want it faster? Shrink the model or buy faster memory. There is no third option.<\/li>\n<\/ol>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<h2 class=\"wp-block-heading\">\ud83c\udfc1 Final Rant: Read the Label<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">So, that wall of files on the download page? It isn&#8217;t noise any more. It&#8217;s a menu. Every entry is a deliberate trade between how much the model knows and how much machine you&#8217;re willing to give it.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Will your desk ever beat the biggest model in the cloud? No. But it doesn&#8217;t need to. For private documents, offline work, and endless tinkering without a bill at the end of the month, a well-chosen local model is plenty. And &#8220;well-chosen&#8221; is the whole game.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But all sarcasm aside, the real lesson here isn&#8217;t <code>Q4_K_M<\/code>. It&#8217;s that the tools now make running a model a one-line command, and that&#8217;s exactly why people stop asking what the command actually downloaded. The download is the easy part. Knowing what&#8217;s in it is your job.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Because at the end of the day, no number of bits fixes not reading the label.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Happy (informed) downloading!<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Same model, eight files, names like hotel Wi-Fi passwords. Here&#8217;s what GGUF, Q4_K_M, K-quants and i-quants actually mean, how much dumber each one makes your model, and the napkin maths to know if it&#8217;ll fit before you hit download.<\/p>\n","protected":false},"author":2,"featured_media":1000041,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":""},"categories":[7,50],"tags":[66,67,64,69,61,68],"class_list":["post-1000039","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-tech","category-tutorials","tag-gguf","tag-llama-cpp","tag-llm","tag-local-ai","tag-ollama","tag-quantization"],"jetpack_sharing_enabled":true,"jetpack_featured_media_url":"https:\/\/siyaz.tech\/wp-content\/uploads\/2026\/10\/local-ai-q4-k-m-suitcase-scaled.jpg","_links":{"self":[{"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/posts\/1000039","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/comments?post=1000039"}],"version-history":[{"count":1,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/posts\/1000039\/revisions"}],"predecessor-version":[{"id":1000040,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/posts\/1000039\/revisions\/1000040"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/media\/1000041"}],"wp:attachment":[{"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/media?parent=1000039"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/categories?post=1000039"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/siyaz.tech\/index.php\/wp-json\/wp\/v2\/tags?post=1000039"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}