uncategorized

On AI

9 posts

Latest

The Tax Is Read Speed, Not Intelligence: Qwen 3.8 27B on a Mac Studio, Within Noise of a 5090

Same model as the last post — Qwen 3.8 27B at 4-bit — moved from an RTX 5090 to a four-year-old M1 Ultra. Decode more than tripled once the model's own MTP head did the drafting (10.4 → 35.0 tok/s), the heavyweight diffusion drafter everyone benchmarks on Blackwell never broke even, and the capability benchmarks came back within noise of the rig: HoF-Bench 48.9 vs 51.6, HumanEval+ 90.9 vs 92.1, AIME 2025 at 90.0%. The one number that refuses to move is prefill — 16× slower — and the post is honest about why that's physics, not tuning.

llm · benchmarks · apple-silicon · 21 Aug 2026 · 25 min read
  1. The Quantization Tax Nobody Charged: Qwen 3.8's Q4 Quant Matches Its Published Scores on a Single 5090

    The model everyone was waiting for finally dropped, and the cheapest sensible quant reproduces Qwen's published BF16 scores on my 32 GB card — GPQA Diamond 89.4 against a published 89.2, within error, single run. Four independent instruments agree that Q4 costs nothing you can measure, and it's the only quant that holds 262K at four slots. Also inside: the 16,384-token cap that manufactured a fake 72%, and why 'indistinguishable' is not 'beat'.

    llm · benchmarks · quantization · 16 Aug 2026 · 40 min read
  2. Glimmer of a Comeback: Meta's New 30B Tops 95 Real CVEs on Launch Day

    Meta's first open-weight model in years dropped this morning. By tonight it had edged out six local models at rediscovering real CVEs on my 5090. The win is real but not apples-to-apples — a 26× token budget it couldn't avoid, and one task that didn't fit its window. Half this post is about that fine print.

    llm · benchmarks · security · 10 Aug 2026 · 19 min read
  3. Q8 Stopped Costing Speed: How Speculative Decoding Rewrites Quantization Math on Apple Silicon

    The bandwidth math says a 4-bit quant should decode 1.7x faster than 8-bit. On an M1 Ultra it's 1.22x — and once you stack llama.cpp's MTP head with n-gram self-speculation, the gap disappears entirely. A benchmark story about which hardware limit you're actually sitting on, with one 211 GB mistake included.

    llm · benchmarks · apple-silicon · 19 Jul 2026 · 19 min read
  4. Case File DS4 — A 284B Model on a 2021 Mac: The Local Frontier Verdict

    antirez's ds4 squeezes DeepSeek V4 Flash — 284B parameters — into 80.8 GiB that runs fully resident on a 2021 Mac Studio. We benchmarked it to 262k context, ran an independent HumanEval, AIME 2025, and GPQA Diamond suite it was never tuned for (92.7%, 10/13, and ~92% of the official GPQA score at 2 bits), A/B'd the 90.9 GiB hybrid quant, and kept the one negative result: speculative decoding makes this machine slower.

    llm · local-inference · apple-silicon · 6 Jul 2026 · 14 min read
  5. Case File VibeThinker-3B — A 3B vs. the Giants: The Reasoning Verdict

    VibeThinker-3B is Qwen2.5-Coder-3B with a reasoning post-train bolted on. We re-ran it under one harness — caught two benchmark bugs — and it ties a 27B reasoner on AIME and HMMT, beats a small frontier model, and leaps +90 over its own base. It's also the weakest model here on broad science. Both are true.

    llm · benchmarks · distillation · 20 Jun 2026 · 12 min read
  6. Case File Gemma-4 — Distilled vs. Stock: The Reasoning Verdict

    A 12B distilled from frontier teachers (Composer 2.5 + Fable 5) for Python coding, benchmarked head-to-head against stock gemma-4-12B across 249 items. It lost off-domain reasoning — and then lost its own home turf too.

    llm · benchmarks · distillation · 17 Jun 2026 · 8 min read
  7. The Chaos Engine: How Chunked Prefill Unmasked the Non-Linear Nature of Text Diffusion

    A 26 GB text-diffusion model demanded 523 GB of RAM. The fix was one parameter — and the parameter changed what the model said. Running diffusiongemma at its full 262K context on a Mac.

    llm · benchmarks · diffusion · 12 Jun 2026 · 24 min read
  8. Turbo3 vs q8_0: What KV Cache Quantization Really Costs on an RTX 5090

    A benchmark of two KV cache formats on Gemma-4-31B, and what the numbers actually mean for people who use long context.

    llm · benchmarks · 14 Apr 2026 · 26 min read