uncategorized
llm · benchmarks · security · local-inference
On AI · 10 Aug 2026

Local models · Security benchmarks · Muse Glimmer 30B

Glimmer of a Comeback: Meta's New 30B Tops 95 Real CVEs on Launch Day

Listen instead — AI narration in the author's voice
Glimmer of a Comeback: Meta's New 30B Tops 95 Real CVEs on Launch Day

Meta’s first open-weight model in years dropped this morning. By tonight it had edged out six local models at rediscovering real, disclosed CVEs — and the story of getting it to run at all is at least as interesting as the score.

50.0%
strict-match on debut, best of seven
7 × 95
local models × real CVEs, one 5090
0/5
the httpd blind spot, unanimous across all seven
26×
the token budget it needed just to answer

The comeback

Meta released Muse Glimmer 30B today. It’s their return to shipping open weights after a long stretch where that conversation belonged to Qwen, DeepSeek, and the rest of the labs that took the baton Llama handed off in 2023 — a 30B dense model, Apache 2.0, distilled from their larger Muse Spark, small enough to run on an 18 GB card.

I have a standing benchmark for exactly this occasion: HoF-Bench, ninety-five real, disclosed CVEs across eight real codebases — curl, OpenSSL, GnuTLS, Apache httpd, OpenEMR, pearweb, FOGProject, WeKan — where the model gets the pre-fix source files and has to name the most serious vulnerability. Six local models had already been through it on my RTX 5090. So the question wrote itself: can the comeback model actually hunt?

The answer, and the fine print

Yes. Muse Glimmer scored the best strict-match of all seven models — beating Tess-4, the 27B fine-tune that previously held the lead, and both of the mixture-of-experts models that beat everything before it.

Two caveats, stated plainly because they matter. First, Muse Glimmer answers every question through a chain of reasoning it will not skip — the server flags that normally suppress thinking do nothing, I checked in the source — so it ran with twenty-six times the token budget of the other six models. Not a tuning choice; at the standard budget it returns empty answers on even the easiest task in the corpus. Second, one task literally does not fit: its context window is half everyone else’s, and the biggest curl bundle overflows it before the first output token. So it won on ninety-four tasks, not ninety-five. Its raw hit count still beats the runner-up’s, so the ranking survives both — but this is a directional win, not a controlled one.

What the benchmark actually taught me

The score table is the least interesting part. Three findings travel:

Finding the file is a solved problem; naming the bug is not. Six of seven models — including the smallest, a 9B — located the right file every single time. Diagnosing what’s wrong in it landed between 27% and 50%. That gap is the whole benchmark.

Models have label habits, and calibration is everything. The Gemma family leans on “Broken Access Control” and is right about 84% of the time it does — the dataset really is skewed that way. Ornith-9B pattern-matches the same code to “Server-Side Request Forgery” and went 0-for-13. Same reflex, opposite calibration, and that difference alone is most of the spread between last place and the pack.

All seven models share one blind spot — which means it isn’t the models. Every model scored zero on Apache httpd. Not similar wrong answers: locally defensible reads of the code that the disclosed advisories classify differently, because the real impact depends on code the task never shows you. A seventh model from a different lab hitting the same wall is the benchmark’s task shape talking, not model quality.

What to take away

  • Meta came back with a genuinely strong local model. Day-one llama.cpp support, day-one top score on a real-CVE benchmark — with caveats that are about infrastructure, not intelligence.
  • A local scanner’s hard problem is classification, not localization. Read the deep dive before trusting any single-number verdict on that.
  • An unsuppressible reasoning channel is a real operational cost. Slower wall-time, bigger budgets, and no server flag will save you. Check before you deploy this model in a pipeline.
  • When a benchmark says a model failed, check the benchmark’s shape first. The httpd zero is a limitation of file-scoped review, and it took seven models agreeing to prove it.

Why this benchmark exists

I already had a rig, six models, and a finished run when Meta’s release landed this morning. The rig: HoF-Bench v1, a purpose-built vulnerability-detection dataset — 95 real CVE detections across eight repos (curl, OpenSSL, GnuTLS, httpd in C; OpenEMR, pearweb, FOGProject in PHP; WeKan in JavaScript), each pinned to a pre-fix commit, ground truth carrying category, CWE, file path, line range, and fix commit. The scanner side sees repo, commit, and target files — never the CVE identity; an automated leak check enforces that split.

One skew to keep in mind for everything below: WeKan alone contributes 24 of the 95 tasks — a quarter of the corpus — and its CVEs are overwhelmingly access-control-shaped. That matters when the per-repo numbers arrive.

The task itself is single-shot: all of a task’s target files concatenated and labeled by path (149 unique files fetched at pinned commits; the largest task is a seven-file curl bundle around 528K characters), one call, temperature 0, and a fixed answer format — a VULN: line naming file, category, and location, then a short justification. max_tokens 600 for six of the seven models. The seventh gets its own section.

Worth being precise about the grading: HoF-Bench ships a reference harness with a genuinely strict methodology — five detector passes, skeptical triage, an arbiter, and a blinded LLM judge that requires agreement on path, root cause, attacker-controlled condition, and impact. I didn’t use it: it only speaks to hosted APIs, it’s far heavier than single-shot, and the judge costs real money. My grader is a deterministic local script checking two proxies per task, independently: file-match (did the model name the right file?) and family-match (does its free-text category regex-classify into the same coarse family — memory, XSS, SQL injection, authorization, DoS, and so on — as the ground truth?). Strict-match is both at once. This is explicitly weaker than the official judge: it will pass some “right file, wrong bug” and “right bug, wrong file” cases the blinded judge would catch. Treat every number here as a lighter signal, not a HoF-Bench-official score.

Hardware is the usual suspect: an RTX 5090 with 32 GB of VRAM, llama.cpp served through llama-swap, one single-shot profile per model, no speculative decoding anywhere, most models at their full 262,144-token context.

Launch day: an architecture llama.cpp had never heard of

Muse Glimmer’s GGUF reports general.architecture = "muse-glimmer", and my three-week-old llama.cpp build said exactly what you’d expect:

llama_model_load: error loading model: unknown model architecture: 'muse-glimmer'

This is a trap with a tempting wrong exit. The tensor names (attn_gate, attn_q_norm, attn_k_norm) and hyperparameters (sliding_window, final_logit_softcapping, logit_scale) look close enough to the Gemma/afmoe family that “it’s secretly architecture X, just patch the metadata string” was a real thought — this rig already hosts Tess-4, a Qwen3.5 fine-tune under its own name, so rebranded architectures are a live hypothesis. It would have been the wrong move. An arch string patched to something that merely loads risks silent numerical corruption — wrong shape assumptions, wrong norm epsilon, wrong RoPE application — with no crash to warn you. A clean failure is strictly better than a wrong success.

Checked properly instead: no trace of the architecture anywhere in the installed llama.cpp source, and a web search confirming Muse Glimmer is a genuinely new model — 30B dense, its own perception encoder, distilled from Meta’s Muse Spark, released today under Apache 2.0.

The fix was fifty commits of git pull away. Same-day upstream support had already landed — a single PR authored by a Hugging Face engineer with co-authors from both Meta and Hugging Face: 22 files, 877 insertions, covering the model graph, the converter, chat-template handling, and full multimodal support. It even carried its own bugfix commit for an image-input regression, validated with byte-identical greedy output. That’s a tested, cross-company launch patch, not a rushed hack — and it’s worth pausing on how far the ecosystem has come, because this pipeline — the GGUF format, llama.cpp, the whole local-inference world this blog runs benchmarks on — is downstream of a Meta release. Open-weight models existed before Meta: GPT-2 in 2019, EleutherAI’s GPT-J and GPT-NeoX in 2021–22, BLOOM in 2022. But it was Llama, in February 2023, that made competitive-scale open weights mainstream — months before Qwen’s first open weights in August 2023 and DeepSeek’s in November 2023. The lab that started the wave spent a couple of years out of the leaderboard conversation. Today it walked back in, and the ecosystem its old release created had support ready the same morning.

The architecture, read straight from the merged source

Rebuilding llama.cpp bought me the model graph source, which is more precise than any launch blog:

  • Dense, 52 layers, grouped-query attention with per-head Q- and K-norm. No mixture-of-experts, no state-space layers — unlike two of the six incumbents it was about to face.
  • An attention output gate: a learned sigmoid(gate) scaling on the attention output before projection. Not an MoE router — a per-token volume knob on attention itself.
  • Positional embeddings the unusual way round: RoPE runs only on the sliding-window layers; the full-attention layers get no positional embedding at all. If you assumed the global layers carry position, this model is the counterexample.
  • Sliding-window period 4 — one global layer per four windowed ones, same ratio family as Gemma — and the window metadata is mandatory: the loader hard-fails without it.
  • Gemma-family tells everywhere: final-logit softcapping, a logit scale, plus a hardcoded separate epsilon for the post-attention and post-FFN norms. Shared heritage or convergent design; the source doesn’t say.
  • Native context: 131,072 tokens. Half of what the other six models run at. Ask for more and the server silently caps you, so the profile pins it explicitly. At full context with an 8-bit KV cache it sits at 17.4 GB of VRAM — comfortable on a 32 GB card.
  • Vision encoder and a bundled speculative-decode drafter landed in the same PR; both stayed off for this run, matching every other model’s profile.

The reasoning tax no flag can waive

This is the finding I’d most want another operator to read before deploying this model.

Muse Glimmer always reasons before answering, through a harmony-style channel where the model first addresses itself, then the user. The server’s parser splits this correctly into reasoning and content, so a casual test looks fine. The problem appears under a budget. At this benchmark’s standard 600-token ceiling, the model burned the entire budget reasoning and returned empty content — on the single smallest task in the corpus, 445 characters of code. Not a tail case. The easiest task in the set.

Two models in the original six had already taught me a version of this lesson — their chat templates default a thinking mode to on, and both initially burned their whole budget on hidden chain-of-thought. Nemotron’s first pass scored 1.1% because of it. But those had a Jinja toggle to flip. Muse Glimmer is the harder case:

  • Both of llama-server’s standard mitigations — --reasoning off and --reasoning-budget 0 — had zero effect. Byte-identical reasoning, byte-identical content, flag on or off.
  • The parser source explains why: the generation prompt handed to the model is unconditionally bare — the model decides whether to reason. The reasoning flags only control how the server labels text that has already been generated. They do not, and structurally cannot, change what gets generated.

So the run used a 16,000-token ceiling for this model, chosen after measuring the real requirement: the median-size task needed about 4,200 completion tokens; across the full run the range was 870–12,371. Nothing hit the new ceiling, and since generation stops naturally at the model’s own end token, the tall ceiling costs nothing on easy tasks. But it is a genuine methodology deviation — twenty-six times the budget the other six models got — made under necessity, not preference. Every comparison below carries that asterisk.

It also costs wall-clock: 10.5 to 178.9 seconds per task, median 52 — 94 minutes for the full run, meaningfully slower than every other model here, because it narrates its way to every answer. That’s an architectural property, not a configuration failure, and nothing at the server layer can currently change it.

The one task it cannot run

Before committing to a full run I tokenized all 95 prompts against Muse Glimmer’s own tokenizer. The biggest task — that seven-file curl bundle — comes to 140,346 prompt tokens against a 131,072 hard ceiling. Over before the first output token; confirmed live as an HTTP 400. Every other task fits with real margin (the second-largest leaves ~34,800 tokens of headroom), and the other six models, at double the context, ran all 95. So Muse Glimmer is scored out of 94, and that denominator difference is stated everywhere it matters.

Results

model params (active) file-match family-match strict-match
Muse Glimmer 30B 30B (dense) 100.0% 50.0% 50.0%*
Tess-4-27B 27B (dense, Qwen3.5 fine-tune) 100.0% 48.4% 48.4%
Gemma-4-26B-A4B 26B (4B active, MoE) 100.0% 47.4% 47.4%
Nemotron 3 Nano 30B-A3B 30B (~3B active, hybrid MoE) 95.8% 47.4% 44.2%
qwen-3.6-27B 27B (dense) 100.0% 45.3% 45.3%
Gemma-4-12B 12B (dense) 100.0% 45.3% 45.3%
Ornith-1.0-9B 9B (dense) 100.0% 27.4% 27.4%

* Out of 94 tasks, not 95, at max_tokens 16,000 rather than 600 — both out of necessity, both detailed above. Raw hit count: 47, against Tess-4’s 46 out of 95 — the ranking holds on raw hits with one fewer scoreable task, but treat it as directionally real rather than strictly controlled. (Gemma-4-31B was deliberately skipped: the 26B-A4B covers the family, scored better on its smoke test, and runs four times faster.)

The headline pattern isn’t the winner — it’s the file-match column. Localization is saturated. Six of seven models, from 9B to 30B, found the right file every single time; the seventh missed four times for formatting-adjacent reasons. Diagnosis — the family-match column — is where the entire spread lives, from 27.4% to 50.0%. On this benchmark, “can the model find where to look?” stopped being a differentiator at sizes smaller than anything I tested. “Can it name what’s wrong?” is the open problem.

A concrete pair to make the two columns vivid, from the very first curl task: ground truth is an improper-authentication bug in lib/vssh/wolfssh.c. Gemma names exactly that file — and calls it an integer overflow. File-match pass, family-match fail. Almost every interesting failure in this corpus has that shape: right neighborhood, wrong diagnosis, against a file that plausibly contains more than one bug-shaped pattern.

Per-repo, where the stories live

Family-match by repo:

model curl (10) fogproject (5) gnutls (8) httpd (5) openemr (12) openssl (22) pearweb (9) wekan (24)
Muse Glimmer 30B 4/9* 0 3 0 6 8 7 19
Tess-4-27B 4 3 2 0 8 7 5 17
Gemma-4-26B-A4B 3 1 1 0 8 6 6 20
Nemotron 3 Nano 3 4 1 0 7 4 6 20
qwen-3.6-27B 3 2 2 0 8 6 6 16
Gemma-4-12B 3 0 1 0 8 6 6 19
Ornith-1.0-9B 2 0 1 0 4 6 6 7

* out of 9 — the context-overflow task was a curl task.

Muse Glimmer’s win is broad, not one lucky cluster: best OpenSSL score of all seven (8/22 — nobody else above 7), best pearweb (7/9), tied-best curl on the nine tasks it could attempt. Its OpenEMR is middling and its FOGProject is a story unto itself (below). Tess-4 keeps the honor of being top-or-tied across the widest spread of repos — its previous lead was earned by being evenly competent, not by riding WeKan’s access-control cluster the way the Gemma variants partly do.

The httpd blind spot: zero for five, seven models, no exceptions

Every model scored 0/5 family-match on Apache httpd. My first instinct was a grading bug. It isn’t. The ground truth categories are denial-of-service (three), memory safety (one), and cross-site scripting (one). What the models reported instead — independently, across seven architectures from five labs — were command injection, buffer overflow, LDAP injection, certificate validation, race conditions: all locally defensible reads of the same code, all classified differently by the disclosed advisory.

The FTP proxy task is the cleanest example. Every model reads the unsanitized FTP response handling as a command-injection or buffer-overflow primitive. The actual CVE is filed as cross-site scripting — because the vulnerable path is that FTP response getting reflected into an HTML directory listing elsewhere in the request pipeline, outside the files the task hands the model. The models found real-looking primitives; the advisory classified by end-to-end impact through code the models never saw.

That’s a limitation of file-scoped review — mine, and by extension any scanner that reviews files instead of request paths — not a model-quality gap. Six models agreeing could have been correlated bias; a seventh from a different lab, different architecture family, released years after the others, hitting the identical wall is about as clean a signal as this kind of benchmark can produce. Muse Glimmer even reproduces the pattern in miniature on GnuTLS and OpenSSL: four tasks where the advisory’s disclosed impact is denial-of-service and it reports an out-of-bounds read — very plausibly a correct read of the underlying memory bug, scored wrong because it classified by primitive rather than by the advisory’s chosen impact.

Label habits: one reflex, three calibrations

The most human finding in the corpus. Models develop favorite labels for recurring code shapes, and whether that helps or destroys them is purely a question of calibration.

Well-calibrated: Gemma-4-26B-A4B reached for “Broken Access Control” 25 times — and was right 21 of them, 84%. Suspicious on its face (a quarter of its verdicts, one label), but the corpus really is shaped that way: WeKan’s 24 tasks are nearly all access-control CVEs. The 12B does the same at 84%. And here’s the independent evidence that this is calibration rather than grader-gaming: Muse Glimmer, an unrelated architecture from a different lab, converges on the same habit — 21 of its 94 verdicts, correct 85.7% of the time. Two labs, one dataset skew, same well-fitted reflex. Across the 26 ground-truth authorization tasks, the Gemmas caught 21, Muse Glimmer 20, Tess-4 and Nemotron 19, qwen 17.

Miscalibrated: Ornith-1.0-9B looked at largely the same WeKan code — endpoints that take an identifier and fetch data by it — and pattern-matched it to “Server-Side Request Forgery” thirteen times. It was wrong all thirteen. The confusion is legible: SSRF is attacker-controlled destination, IDOR is attacker-controlled identity, and both look like “endpoint takes a reference and fetches something.” Its authorization recall: 6 of 26. That one miscalibrated reflex is most of the distance between its 27.4% and everyone else’s 44-plus.

And the winner isn’t immune: Muse Glimmer went 0/5 on FOGProject, and three of those five misses are the same verdict, verbatim — “Unauthenticated Mass Assignment” against the same recurring PHP foreach ($_REQUEST …) idiom — on three different tasks whose ground truth is stored XSS every time. Not three unrelated mistakes: one code shape, read identically and identically wrong, three times. Structurally the same phenomenon as Gemma’s habit and Ornith’s habit, just landing on the wrong label for this particular idiom. A per-repo breakdown makes that visible; an aggregate score would have buried it. The other two FOGProject misses are the httpd shape again — plausible primitives (an authentication bypass, an RCE-flavored class instantiation) that the advisories classify as something else.

What I’d trust and what I wouldn’t

What I’d trust: every number came from a deterministic local grader over saved raw responses, recomputed from the grading report for this post rather than quoted from memory; temperature 0 and fixed seeds throughout; and the two proxy metrics are defined narrowly enough to be reproducible.

What I wouldn’t generalize without care:

  • This is my lightweight grader, not HoF-Bench’s blinded judge. It says nothing about root cause or impact, and it will pass some half-right answers the official methodology would fail. The absolute percentages would all drop under the real judge; I’d expect the ordering to be more durable, but that’s an expectation, not a measurement.
  • Muse Glimmer’s win carries both asterisks. A 16,000-token budget against everyone else’s 600 — forced by an unsuppressible reasoning channel, verified at the parser-source level — and a 94-task denominator against everyone else’s 95, forced by a context ceiling half the size. On raw hits it still wins, 47 to 46. Cite it as “best, directionally” or don’t cite it.
  • The corpus skews. A quarter of all tasks are one JavaScript project’s access-control CVEs. Models fluent in that pattern gain more than their general skill warrants; the per-repo table is the honest view.
  • File-match saturation is partly by construction. Tasks hand the model a small, explicitly scoped file list. Real-world triage across a whole repo is a different, harder localization problem — the workflow post is about exactly that gap.
  • Single-shot is the floor, not the ceiling. No retrieval, no multi-pass triage, no tool use. These numbers measure raw pattern recognition under a fixed format, which is one narrow — if honest — slice of what a security workflow does with these models.

The one-line version: Meta came back, the ecosystem its own 2023 release built had day-one support waiting, and the comeback model took the top of my chart within hours — through a reasoning tax nothing can switch off, minus one task it physically cannot read, ahead of six models that all agree with it about exactly one thing: on Apache httpd, everybody’s wrong in a locally defensible way.