uncategorized
llm · benchmarks · apple-silicon · speculative-decoding · local-inference
On AI · 21 Aug 2026

Local models · Apple Silicon · Qwen 3.8 27B

The Tax Is Read Speed, Not Intelligence: Qwen 3.8 27B on a Mac Studio, Within Noise of a 5090

Listen instead — AI narration in the author's voice
The Tax Is Read Speed, Not Intelligence: Qwen 3.8 27B on a Mac Studio, Within Noise of a 5090

The previous post established that Qwen 3.8 27B loses nothing measurable at 4-bit on an RTX 5090. This one moves the same model to the other machine on my desk — a Mac Studio M1 Ultra from 2022 — and asks two questions: how fast can Apple Silicon actually run it, and how much intelligence does the platform move cost? The answers: 35 tok/s using a drafter Qwen ships inside the model that most runtimes ignore, and nothing a benchmark can detect. The tax shows up somewhere else entirely.

35.0
tok/s decode — 3.4× the plain bf16 starting point, on the model's own MTP head
90.0%
AIME 2025 pass@1, thinking on, all 30 problems run locally
48.9/51.6
HoF-Bench strict, Mac vs 5090 — ~2.5 tasks apart on 95 real CVEs
16×
the prefill gap vs the 5090 — the one tax that stays on the bill

Same model, different physics

The 5090 post ended with a 17 GB Q4 quant matching Qwen’s published scores. The obvious follow-up: what happens on the machine most people actually own — no discrete GPU, unified memory, MLX instead of llama.cpp? Plain bf16 decoding on the M1 Ultra starts at a sobering 10.4 tok/s. But this model ships with a multi-token-prediction head baked into its weights, and MTPLX uses it: draft with the model’s own head, verify in one pass, commit through exact rejection sampling. Output is bit-identical to plain decoding. With a 4-bit quant and the right compute dtype, decode lands at 35.0 tok/s — about 26 words a second, comfortably faster than reading pace.

The drafter that pays rent is the free one

I also ran DFlash 2 — the block-diffusion drafter behind the 15×-on-Blackwell headlines. On the M1 Ultra it never breaks even: its external 3.6 GB drafter costs a flat ~430 ms per speculation step against a ~96 ms plain decode step, so its best case is +14% at bf16, and at 4-bit it’s a net slowdown despite accepting 4 of every 5 drafted tokens. The model’s native MTP head costs almost nothing to consult, and at draft depth 2 it turns the best plain baseline’s 25.2 tok/s into 35.0. On this class of hardware, speculation only pays when the drafter is free.

The Mac gives up read speed, not intelligence

Capability, benchmarked against the 5090 rig running the same model at the same 4-bit class: HoF-Bench (95 real CVEs) 48.9% vs 51.6% — about 2.5 tasks, inside the pack the rig post described. HumanEval+ 90.9 vs 92.1 — within the ±2.1 error bar that post itself cites. Then AIME 2025, thinking on, all 30 problems: 27/30 = 90.0%, where two of the three misses were 64K-token truncations rather than wrong answers — a 128K rerun of just those two scores 96.7%, with one honest asterisk documented inside. What actually separates the machines is reading: decode is ~2× slower, and prefill at 100K context is ~123 tok/s against the rig’s ~1,900–2,100 — a 16× gap. Decode streams weights, and 800 GB/s of unified memory competes; prefill is batched compute, and that’s where a 5090’s FLOPS live.

What to take away

If your workload is generation-heavy — writing, code, chat, agents that think more than they read — a Mac you already own is a legitimate inference box for this model, with the full 262,144-token window and an SSD session cache that turns a repeated 100K-token prefill from 834 s into 21 s. If your workload is long-context read-heavy, the datacenter GPU keeps its moat, and no amount of tuning changes that: it’s compute-bound physics. The whole setup is four commands (pip install mtplx → setup --download → tune --retune → quickstart), and the tuner measures your chip rather than assuming mine. The deep dive has every run, the fixed-overhead math that explains why the fancy drafter loses, and the caveats each headline number carries.

The question this post answers

The 5090 post settled the quantization question for Qwen 3.8 27B: Q4 costs nothing you can measure. But it settled it on a machine with a 32 GB Blackwell card, and the most common question local inference actually gets is a different one: what about a Mac? Unified memory means even a mid-tier Apple Silicon machine can hold weights that would need a multi-GPU rig elsewhere — the M1 Ultra on my desk has 128 GB of it at ~800 GB/s. What nobody hands you is an honest account of what that memory buys and what it doesn’t.

So this post runs the same model through two phases. Phase one is a speculative-decoding shootout: plain autoregressive decoding, DFlash 2 (the external block-diffusion drafter), and MTPLX (which drives the multi-token-prediction head Qwen ships inside the model — the same head the 5090 post used via llama.cpp’s --spec-type draft-mtp, worth +54% there). Phase two is capability: HoF-Bench, EvalPlus, a throughput probe at 100K context, and AIME 2025 — each chosen because it fits this hardware’s budget, with the ones that don’t fit named rather than quietly skipped.

Every number here was measured on this machine and audited against the run log and raw result files while writing, except where a figure is explicitly labeled as published or leaderboard-reported — and those labels survive into every table they touch. That’s the rule the last two posts established, and it holds.

The machine, the contenders, one naming trap

The setup: Qwen 3.8 27B on an M1 Ultra in four commands, via the model’s own MTP head

host Mac Studio, Apple M1 Ultra, 128 GB unified memory (~800 GB/s), macOS 26.5.2
stack Python 3.14.6, MLX 0.32.0, mlx-lm 0.31.3
tools dflash 0.1.0, mtplx 2.8.3 (both PyPI, both vetted before install)
model Qwen 3.8 27B — bf16 (54 GB) and 4-bit-class checkpoints

A useful sanity anchor before any speculation: bf16 decode is bandwidth-bound, so the ceiling is roughly memory bandwidth over model size — 800 GB/s ÷ 54 GB ≈ 14.8 tok/s theoretical. The measured plain-MLX baseline of 10.4 tok/s is ~70% of that ceiling, which says the baseline is already reasonably efficient and anything dramatic has to come from either smaller weights or fewer forward passes. Those are, in fact, the only two levers this post finds.

The contenders:

method drafting mechanism extra download
Plain AR none — the baseline —
DFlash 2 external 3.6 GB block-diffusion drafter; drafts a whole block in one pass drafter checkpoint
MTPLX the model’s own MTP head; near-free draft, exact rejection sampling none beyond the model

Both speculative methods are lossless by construction — DFlash by verification, MTPLX by exact rejection sampling — so this entire phase is a speed comparison. Quality is phase two’s job.

One naming trap, spelled out once because the vocabulary collision is genuinely unlucky: MTPLX’s “D2” means MTP draft depth 2 — two tokens drafted ahead. It has nothing to do with the “2” in “DFlash 2”. Everywhere below, D1/D2/D3 are MTPLX depths and “block 4/5/8” are DFlash block widths.

And one trap for anyone reproducing this: mlx-community/Qwen3.8-27B-MTP-bf16 is a 0.9 GB head-only repo — no trunk weights. It will not run anything by itself; the runnable MTPLX checkpoints are the Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed pair.

DFlash 2: the heavyweight that never breaks even

The shootout: every configuration tested, plain baselines vs DFlash 2 vs MTPLX

DFlash 2 is the drafter behind the loudest headline in speculative decoding this year — up to 15× on Blackwell hardware. The workload for all dflash runs: gsm8k, 4 samples, max 1,024 new tokens, reasoning xhigh, temperature 1.0, top-p 0.95, top-k 20 — a speed measurement with realistic generation, not a quality eval. Each run re-measures its own block-size-1 baseline in the same harness, which is how you catch a drafter that only looks good against a stale baseline.

Three runs, one story:

run precision block baseline tok/s DFlash tok/s speedup
1 bf16 8 (drafter’s trained max) 10.35 11.83 1.14×
2 bf16 4 10.42 8.42 0.81× — net slowdown
3 4-bit target + 4-bit draft 5 (README recipe) 25.16 23.63 0.94× — net slowdown

The acceptance numbers make the failure interesting rather than dumb: at block 4, the model accepted an average of 3.52 of 4 drafted tokens and took the whole block 68.8% of the time — excellent drafting — and still lost. At 4-bit with the README-recommended settings, 4.00 of 5 accepted, 57.6% full blocks, still a slowdown. A drafter can be accurate and uneconomical at the same time, and on this chip DFlash is exactly that.

Note the two baselines in that table, because they’re the quantization lever measured cleanly: plain 4-bit decoding is 25.2 tok/s against bf16’s 10.4 — ~2.4× from weight size alone, no speculation involved. Bandwidth-bound decoding scales with bytes streamed per token, and the 5090 post already established the quality side of that trade is free.

The measurement that explains everything: flat per-step overhead

Working backwards from time-per-output-token gives the number that reframes the whole shootout: each DFlash speculation step — one drafter forward plus one verify pass — costs ~440 ms at block 8 and ~420 ms at block 4. Essentially flat as the block gets wider. A plain decode step on the same setup: ~96 ms.

That flatness is the finding. If the cost per step barely moves with width, the expensive part is not the wide verify matmul — it’s a large fixed overhead: the external diffusion drafter’s own forward pass plus MLX dispatch and sync. Three consequences fall out:

  • Wider is strictly better for DFlash here. The fixed cost amortizes over more drafted tokens, so block 8 — the drafter’s trained maximum — is its best case, and it still only reaches 1.14×.
  • The classic “draft length 4–8” intuition does not transfer. That sweet spot comes from autoregressive drafters whose cost grows per drafted token. A diffusion drafter with flat step cost inverts those economics entirely.
  • The faster your baseline, the worse the heavyweight drafter looks. The 4-bit run confirms the theory from a second angle: the quantized drafter’s step cost dropped to ~169 ms, but the baseline step dropped to ~40 ms — the overhead ratio stayed above 4×, and DFlash lost even at 80% acceptance. A drafter that can’t pay rent against a 10 tok/s baseline has no chance against a 25 tok/s one.

Which is precisely the argument for the drafter that costs nothing: the MTP head reuses the trunk’s own hidden state, so the draft is nearly free and the per-step overhead should collapse to roughly one verify pass. Shallow depths become viable again. That prediction is testable, and the next section tests it.

MTPLX: the drafter Qwen already paid for

MTPLX serves the multi-token-prediction head that ships inside Qwen 3.8’s weights — the same head llama.cpp exposed on the 5090. Its tune command benchmarks every draft depth on your machine and saves the winner. Two checkpoints exist, and the difference between them taught me something about Apple Silicon I didn’t know:

Run 4 — Optimized-Speed (4-bit, 21.3 GB), bf16 compute:

depth tok/s vs AR
AR 19.7 1.00×
D1 27.2 1.38×
D2 27.8 1.41× — best, saved
D3 23.8 1.21×

Run 5 — Optimized-Speed-FP16 (21.3 GB), fp16 compute:

depth tok/s vs AR
AR 22.3 1.00×
D1 33.8 1.52×
D2 35.0 1.57× — best, saved
D3 29.5 1.32×

First, the discovery hiding in the second checkpoint’s name: the “FP16” build is not an fp16 trunk. It’s 21.3 GB — the same size as the 4-bit build — because it is the 4-bit quant, shipped with an fp16 compute path. It exists because M1 and M2 have no native bf16 support; every bf16 operation on this chip is emulated. Same quant, same head, same acceptance rates (D2: 99.2% on the first drafted token, 96.0% on the second — identical across both runs), and yet every stage is ~25–30% faster: the compute dtype alone is worth 1.26× end-to-end (35.0 vs 27.8 tok/s). mtplx hardware also reports “M5 TensorOps eligible: false” here — MTPLX’s own 2.24× M5 Max headline benefits from hardware this chip doesn’t have, so treat their marketing numbers and mine as different machines, because they are.

Second, the depth curve is exactly the shape the fixed-overhead theory predicted. With a near-free draft, speedup ≈ acceptance at shallow depth: D1 accepts 100% and delivers 1.52×; D2 accepts 99.2%/96.0% and squeezes out a bit more; D3’s acceptance decays to 93.4/86.9/78.1% while verify cost grows (71.4 → 96.1 → 107.4 ms) and it already loses to D2. The M1’s wide-verify cost curve is the same one that killed DFlash — it just starts from a base cost so much lower that depth 2 clears it easily.

Two caveats that belong next to these tables. The bf16 apples-to-apples against DFlash’s Run 1 is not possible with published checkpoints — no true fp16-trunk MTPLX build exists in the catalog, so MTPLX is only ever measured here at 4-bit-class. And mtplx tune generates with its own prompt set and sampling, not the temp-1.0 gsm8k workload of the dflash runs — acceptance rates across the two harnesses aren’t directly comparable, but each method’s speedup is computed against its own AR baseline in its own harness, which is the ratio that’s fair.

The ranking, and three stacked lessons

Everything tested, decode tok/s on the M1 Ultra, one list:

# configuration tok/s
1 MTPLX 4-bit, fp16 compute, D2 35.0
2 MTPLX 4-bit, fp16 compute, D1 33.8
3 MTPLX 4-bit, bf16 compute, D2 27.8
4 Plain AR 4-bit (mlx-lm) 25.2
5 DFlash 2 4-bit, block 5 23.6 — slower than its own baseline
6 MTPLX 4-bit, bf16 compute, AR 19.7
7 DFlash 2 bf16, block 8 11.8
8 Plain AR bf16 10.4
9 DFlash 2 bf16, block 4 8.4

The winner is 3.4× the bf16 starting point, 1.4× the best plain quantized baseline, and 3.0× DFlash’s best result. The gap decomposes into three independent, stacked lessons — each one a real multiplier on this hardware:

  1. Quantize first (10.4 → 25.2, ~2.4×). Decode streams weights; fewer bytes is more tokens. The 5090 post is the other half of this lesson: at Q4, the quality cost is nothing you can measure.
  2. Speculate with a free drafter, not a heavyweight one (25.2 → 35.0, ~1.4×). The external diffusion drafter carries a >4× fixed per-step overhead on M1-class silicon and never breaks even at 4-bit. The native MTP head is effectively free and wins at shallow depth.
  3. Match the compute dtype to the silicon (27.8 → 35.0, 1.26×). No native bf16 on M1/M2 means the fp16-compute build is faster at every depth with identical acceptance — a free 26% most people leave on the table by downloading the default checkpoint.

And one anti-lesson: the “draft length 4–8” tuning folklore transfers to neither method. Diffusion drafters have flat step cost, so you go as wide as the drafter was trained (block 8). MTP heads are near-free, so shallow wins (D2, with D3 already losing). The folklore encodes the economics of a third kind of drafter neither of these is.

Why prefill is the tax: bandwidth vs FLOPS

Speculation only accelerates decode. The other axis — how fast the model reads — got its own probe: mac_bench_decode.py, an OpenAI-compatible port of the rig’s harness, against the served MTPLX winner at D2.

context prefill decode wall clock (300 tok out)
short (37 tok) — 36.6 tok/s 13.3 s
~100K, cold ~123 tok/s ~14.7 tok/s 833.9 s
~100K, warm repeat skipped — SSD session cache hit 14.7 tok/s 21.0 s

Against the 5090 running the same model (Q4 + MTP, llama.cpp): decode 36.6 vs 75.6 tok/s short-context — about 2×, respectable. Prefill at 100K: ~123 vs ~1,900–2,100 tok/s — ~16×. And decode itself degrades ~2.5× from short to 100K context (36.6 → 14.7) as the attention layers’ cache grows.

The 2× and the 16× are not the same kind of gap, and the difference is the durable lesson about Apple Silicon. Decode is bandwidth-bound: one token at a time, stream the weights, and 800 GB/s of unified memory genuinely competes with a discrete card. Prefill is compute-bound: big batched matmuls that want raw FLOPS, and the 5090 has roughly an order of magnitude more of them. Speculation attacks only the bandwidth-bound side. Nothing attacks the compute-bound side except smaller weights — and caching.

Which is why the least glamorous row of that table might be the most important one: MTPLX’s SSD session cache is on by default, and it turned a repeated 100K prefill from 834 s into 21 s. llama.cpp’s prompt cache does something similar on the rig, but on hardware where prefill is the weak axis, prefill-you-never-repeat is worth more than any decode multiplier in this post. Scope every “3.4× faster” claim accordingly: it’s a claim about generation-heavy work.

Capability vs the 5090: HoF-Bench and EvalPlus

Mac vs 5090: capability within noise, throughput split by axis

Phase two serves the winner (mtplx quickstart, D2, context_length 262,144 confirmed via /v1/models, reasoning off to match every rig run’s enable_thinking:false) and asks what the platform move costs in capability. The slate was chosen to fit ≤4 h per benchmark on this hardware: HoF-Bench for security, EvalPlus HumanEval+ for code, the throughput probe above, and AIME 2025 in the next section. Skipped, with reasons rather than silence: RULER (a single 262K prefill cost the 5090 ~237 s; at this machine’s measured ~123 tok/s it’s ~35 minutes per sample — the full sweep is not a Mac-sized job), GPQA Diamond and IFBench (7.0 and 5.7 h on the 5090 → 15 h+ here; the rig post’s measured 89.4 and 78.0 stand as this quant-family’s numbers), and KL-divergence (needs the Q8/Q6/Q4 trio side by side, which was cleaned up after the last post).

HoF-Bench v1 — 95 real, disclosed CVEs, the harness and grading methodology from the Glimmer post: the model gets the pre-fix source files and must name the most serious vulnerability; strict scoring requires the right vulnerability family in the right file. Temperature 0, seed 42, single shot. The Mac run: 95/95 tasks completed, 48.9% strict (94/95 scored), 100% file localization. Inserted into the rig leaderboard:

model runtime strict %
Qwen 3.8 Q4, MTP off (5090) llama.cpp 51.6
Qwen 3.8 Q4, MTP on (5090) llama.cpp 50.5
Muse Glimmer 30B (5090) llama.cpp 50.0
Qwen 3.8 4-bit MTPLX D2 (M1 Ultra) MLX 48.9
Tess-4 27B (5090) llama.cpp 48.4
Gemma-26 A4B (5090) llama.cpp 47.4

That’s 2.7 points below the rig’s best arm — about 2.5 tasks at 1.05 points each — inside the “tight pack” the rig post warned against ranking within. The failure signature is identical to every model on that board: perfect file localization, all points lost on vulnerability category. And the comparison is honestly not apples-to-apples: different quant recipe (MTPLX’s 4-bit vs UD-Q4_K_XL), different runtime (MLX vs llama.cpp), and MTP D2 active — the rig’s own MTP-on arm cost ~1.1 points versus off, consistent with what the Mac row shows. The paired per-task test that would settle “same or different” needs the rig’s per-task outputs, which aren’t synced here; until then, the honest reading is indistinguishable within this benchmark’s resolution.

EvalPlus HumanEval+ — greedy pass@1 against the same server, all 164 problems in 18 minutes 51 seconds of wall clock, decode-dominated — exactly the workload shape where D2 pays.

dataset base tests + extra tests 5090 Q4 (+ extra)
HumanEval 92.7 90.9 92.1

1.2 points below the rig on HumanEval+ — under the ±2.1 binomial standard error the rig post itself cites for n=164: statistically indistinguishable. (MBPP+ wasn’t run — it’s another 2–3 h on this hardware, and HumanEval+ already answers the “is the quant broken” question. Noted so the omission is a decision, not an oversight.)

One macOS trap worth documenting because it fails silently as a zero: evalplus’s reliability_guard skips RLIMIT_STACK on Darwin but still sets RLIMIT_AS, which macOS rejects — every test errors and the report reads a clean pass@1 of 0.000. The fix is EVALPLUS_MAX_MEMORY_BYTES=-1, which disables the memory rlimit; the rig script’s 8 GB value only works on Linux. If you ever see a competent model score exactly zero on EvalPlus on a Mac, it didn’t; your sandbox did.

AIME 2025: near-frontier math on a desktop Mac

AIME 2025 on the M1 Ultra: 27/30 pass@1, the two truncations, and the 128K rerun

The Artificial Analysis Intelligence Index can’t be run locally — it’s their composite of ~10 hosted evals, and GPQA alone blows this machine’s budget. But one component fits: AIME 2025, thirty problems, exact-match on the boxed answer. Conditions matched to Qwen’s published protocol: thinking on at xhigh, temperature 1.0, top-p 0.95, max 65,536 thinking tokens. A three-problem timing probe (27 s / 1,082 s / 79 s — that spread is chain-length variance, and it never got smaller) preceded the full run, which ultimately took 8.08 hours and 667,448 completion tokens, median chain 17,874 tokens, mean decode 23.0 tok/s with reasoning on.

Score: 27/30 = 90.0% pass@1, measured.

The three misses are where the last post’s biggest lesson recurs at a larger scale. Two of them (problems 13 and 29) are not wrong answers — they’re 65,536-token cap-hits with no answer emitted. The rig post’s fake 72% came from a 16K cap; here the same trap reappears at 64K. Among chains that finished reasoning, accuracy is 27/28 = 96.4%. A rerun of just the two cap-hits at a 128K budget produced two correct answers — 29/30 = 96.7% at 128K — but that figure carries an asterisk the raw logs force on it: problem 13 (83,727 tokens) genuinely needed the bigger budget, while problem 29’s successful chain came in at 58,379 tokens, under the old cap — that one is temperature-1.0 resampling luck, not budget. So the primary number stays 27/30 at 64K, with the 128K figure reported alongside and both misses’ original outputs archived.

For scale — and these rows are leaderboard- or vendor-reported, under an easier protocol (hosted, typically multi-sample averages, uncapped thinking budgets), shown for context, never as a controlled comparison:

model AIME 2025 source, conditions
Claude Opus 4.6 99.79 Anthropic system card; avg of 5 runs, no tools — with Anthropic’s own note that contamination may inflate it
GLM-4.7 95.7 published leaderboard figure
Kimi K2 Thinking 94.5 model card, no tools (the widely quoted 99.1 is the with-Python figure)
DeepSeek V3.2 93.1 technical report, pass@1
Qwen 3.8 27B, 4-bit, this Mac 90.0 (96.7 @128K) measured; single-pass pass@1, 64K cap

Each of those rows was re-verified against its primary source while writing this post, which is how the Kimi correction surfaced: the 99.1 circulating in threads (including, briefly, mine) is Kimi’s with-Python-tools number; its no-tools score is 94.5. With the tool-free numbers lined up, a 4-bit 27B on a four-year-old desktop Mac — under a harsher single-pass, capped protocol — sits directly below the hosted frontier pack, not in a different league. The quantization tax remains where the last post left it: unmeasurable.

One operational note for anyone reproducing an 8-hour benchmark on a Mac: the run only survived repeated session-level kill events after detaching the harness with nohup/disown — caffeinate alone did not help, because the kills targeted session-tracked processes, not the OS’s power state.

Spot check: the uncensored fine-tune

A gated community fine-tune (orcarouter/Qwen3.8-27B-Uncensored-MLX, 4-bit, 16.1 GB) claims the same base model with restrictions removed — worth a cheap better-or-worse signal, not a full evaluation. First finding, before any benchmark: the fine-tune strips the MTP head (mtplx inspect reports zero MTP tensors), so it serves AR-only — no D2, roughly half the decode speed, before any quality question.

Design: a deterministic every-3rd-task subset of HoF-Bench (n=32, preserving the repo mix), paired against the baseline’s verdicts on the same tasks — powered only to detect large gaps, by design. Result: baseline 16/32, uncensored 15/32; three discordant tasks; McNemar exact p = 1.0. Statistically indistinguishable, numerically a hair worse, and strictly worse operationally: no MTP head, one 900-second timeout needing a retry, and gated-download friction. No reason to switch, and no reason to spend the other 63 tasks finding that out more precisely.

The verdict, assembled

Phase one’s answer: on M1-class Apple Silicon, the winning stack for this model is 4-bit weights + the model’s own MTP head at depth 2 + an fp16-compute build — 35.0 tok/s, every multiplier measured separately (2.4× quant, 1.4× speculation, 1.26× dtype), all three lossless or measured-lossless. The heavyweight external drafter never breaks even here, and the reason is a flat ~430 ms fixed per-step cost that no acceptance rate can outrun.

Phase two’s answer, in one table:

axis 5090 rig (Q4) M1 Ultra (MTPLX 4-bit, D2) gap
HoF-Bench strict 51.6% 48.9% ~2.5 tasks, within the pack
HumanEval+ 92.1 90.9 within ±2.1 SE
AIME 2025 — (not run there) 90.0% measured near the reported frontier pack
decode, short ctx 75.6 tok/s 36.6 tok/s ~2×
prefill @100K ~1,900–2,100 tok/s ~123 tok/s ~16×

Capability survives the platform move; what doesn’t is read speed. HoF-Bench took ~14 minutes on the rig and ~4 hours here, and that ratio is prefill, not intelligence. For generation-heavy work the Mac is a legitimate stand-in with the full 262K window; for long-context read-heavy work the datacenter GPU keeps a moat that is physics, not software. The reproduction path, in full:

pip install mtplx
mtplx setup --download --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed-FP16
mtplx tune --retune     # measures YOUR chip, saves the best depth
mtplx quickstart --port 8000

What I’d trust and what I wouldn’t

What I’d trust: every measured number traces to the run log or the raw result files (HoF grade report, EvalPlus results, the AIME per-problem outputs including the archived 64K truncations), and was audited against them while writing. The speculative-decoding comparisons are each computed against a same-harness baseline. Both speculative methods are lossless by construction, so the speed tables carry no hidden quality trade.

What I wouldn’t generalize without care:

  • The shootout speed runs are 4-sample gsm8k measurements — throughput instruments, not quality evals. Their baselines re-measured consistently (10.35 / 10.42 tok/s) which is the check that they’re stable, but the tok/s figures are workload-dependent.
  • This is one chip. The M1 Ultra has no native bf16 and no M5-generation tensor hardware; MTPLX’s own numbers on an M5 Max are a different machine, and the DFlash economics that fail here are exactly what succeeds on Blackwell. The durable part is the method — measure the per-step overhead on your silicon — not my constants.
  • Mac-vs-rig capability rows are not paired. Different quant recipe, different runtime, MTP on vs off; the per-task McNemar that would settle it needs artifacts not synced here. “Within noise” is the strongest claim the data supports — and the weakest it forces.
  • AIME is a single pass at temperature 1.0 on 30 problems: one resampling of the run could plausibly move it a problem in either direction, and the 128K rerun’s problem-29 success is explicitly luck, which is why 90.0% stays the headline.
  • The frontier AIME rows are reported, not run — different budgets, different sampling, likely multi-sample; they’re on the table for scale, and any conclusion finer than “same neighborhood” over-reads them.
  • The uncensored spot check is powered for large gaps only (n=32). It answers “is it obviously better?” (no), not “is it equal?”.

The one-line version: the second machine confirms what the first one claimed — the quantization tax on this model rounds to zero everywhere a benchmark can look — and adds the Mac-specific fine print: your unified memory buys you the writing speed and the whole 262K window, but the reading speed belongs to FLOPS, and those still live in the datacenter.