← Back to Articles Directory
AI Models August 10, 2026 8 min read

Meta's Muse Glimmer 30B Released: Hands-On Review, Q2 Quant Tests & Benchmarks

On August 10, 2026, Meta released Muse Glimmer, a 29.6B parameter open-source model distilled from Muse Spark under Apache 2.0. We tested the Unsloth GGUF IQ2_XXS quant on an RTX 3060 Ti, evaluated its full benchmark suite, and ran a 6,352-token mathematical rollback stress test.

Mohid Mirza

Co-Founder & Lead Programmer of AcceleratedLogic AI

On August 10th, 2026, Meta released their new **Muse Glimmer** model, and here are my hands-on thoughts and benchmark breakdowns after running it through real-world testing.
## Model Specs & Architecture First, let's look at the specifications. Muse Glimmer features a **131,072 token context window** and is released fully open-source under the **Apache 2.0 License**. Because it is open-source, we know its exact parameter count: **29.6 Billion parameters**. It is basically a distilled, high-efficiency version of Meta Muse Spark. Notably, it is a dense model rather than a Mixture-of-Experts (MoE)—meaning you cannot offload just a small active expert fraction to your GPU, but it does come in many quantized flavors, going all the way down to a Q2 quant.
## Local Performance: RTX 3060 Ti & Unsloth Q2 Quant In this test, I evaluated the **Muse Glimmer 30B GGUF IQ2_XXS by Unsloth** to see how feasible it is to run on consumer hardware. - **Memory Footprint:** With its full context window loaded, it takes around **15GB total** split between VRAM and system RAM to run on an **NVIDIA RTX 3060 Ti**, making it remarkably memory efficient for a 30B-class model. - **Inference Speed:** With GPU offload, it runs at an average speed of **2.33 tokens per second**. While not lightning fast, it is steady and sufficient for longer-horizon reasoning tasks.
## Comprehensive Benchmark Suite Here is how Muse Glimmer-30B compares against Gemma4-31B Thinking Mode and Qwen3.6-27B Thinking Mode across high reasoning, agentic coding, multimodal, safety, and general capabilities:
Category Benchmark Muse Glimmer-30B Gemma4-31B Thinking Mode Qwen3.6-27B Thinking Mode
**High Reasoning** MCP Atlas (Public) **75.5** 54.2 62.5
DeepSearch QA **74.6** 61.7 71.1
𝛕3-Banking **23.5** 15.1 16.7
WildClawBench **47.6** 37.6 43.2
GDPVal-AA v2 953 811 **1141**
Gaia2 **43.3** 36.4 40.0
SkillsBench (with skills) 44.3 32.4 **46.6**
OSWorld-Verified 65.9 58.5 **75.6**
**Agentic Coding** SWE-Bench Pro **51.2** 36.9 50.2
SWE-Bench Verified 76.0 66.6 **77.2**
TerminalBench 2.1 (with terminus2) 51.7 43.4 **60.7**
SciCode **43.6** 43.4 39.8
**Multimodal** Charxiv Reasoning **78.8** 77.7 78.4
ScreenSpot Pro 75.4 **75.9** 76.1
OmniDocBench v1.5 75.8 72.5 **77.8**
MMMU Pro 74 73 **75**
**Safety** CI Memories Violation (↓): 26.4 / Coverage: 64.8 Violation (↓): 12.1 / Coverage: 53.0 Violation (↓): 53.4 / Coverage: 66.9
Siren AgentDojo Attack Success Rate (↓): 28.4 / Utility: 94.2 Attack Success Rate (↓): 25.6 / Utility: 90.8 Attack Success Rate (↓): 40.3 / Utility: 92.7
**General Capabilities** IFBench **77.0** 76.0 70.8
AIME 2026 **94.7** 89.2 94.1
GPQA Diamond (AA) 83.5 **85.7** 84.2
HLE Text (AA) 22.0 **23.6** 23.1
AA-LCR **80.0** 68.3 73.3
Beam128K **65.1** 58.2 63.0
Overall, I'm somewhat disappointed that Muse Glimmer doesn't completely blow away competing models—given that some of them are a little dated at this point—but it demonstrates solid performance in general coding and high-reasoning tasks.
## Complex Math & Logic Test: The Memory Trap Challenge I wanted to test how accurate the Q2 quantization is under heavy reasoning constraints. I gave it a complex discrete math problem: > **Problem Statement:** > Let a State Vector $S = (x, y, k)$ represent a point on a 2D grid where $x, y \in \mathbb{Z}$ and $k \in \mathbb{Z}^+$ is a tracking weight. > Start at origin state $S_0 = (1, 1, 7)$. The system evolves in discrete steps $n = 1, 2, 3 \dots$ according to $T(S)$: > - **Path A (if $k$ is even):** $S_n = (x + y, \; y \cdot \gcd(x, k), \; k / 2)$ > - **Path B (if $k$ is odd):** $S_n = (x - y, \; \lfloor \sqrt{x^2 + y^2} \rfloor, \; 3k + 5)$ > > **The Constraints & Triggers:** > - **The Memory Trap:** If at any point $x$ becomes strictly negative, you must immediately roll back exactly two steps in your history, invert the sign of $y$ at that past step, and resume calculation along the opposite path choice for that step (forcing a branch change). > - **The Halting Condition:** System terminates the moment $k > 500$ OR $k = 1$. > > **Task:** Calculate the exact final State Vector $S_{\text{final}}$ at termination, and provide the total number of step evaluations (including rollbacks).
**Result:** After **6,352 output tokens**, Muse Glimmer actually arrived at the exact correct answer! This proves that even under extreme Q2 quantization (IQ2_XXS), its step-by-step mathematical reasoning and recursive history tracking remain remarkably coherent.
## Speculative Decoding & Cloud Endpoints - **Speculative Decoder:** I also tested the speculative decoder shipped alongside it. In practice, it didn't dramatically boost throughput—increasing speed by only ~1 token per second. - **Cloud Endpoints:** For those who don't want to run it locally, the only free cloud endpoint currently available is on **NVIDIA NIM**.
## Final Verdict: 8 / 10 Overall, Meta Muse Glimmer gets a **8 out of 10**. It is surprisingly capable for a Q2 quantization on consumer hardware (RTX 3060 Ti), though I do hope Meta continues pushing performance boundaries given the length of time since their last major open-source release.
${relatedPostsHtml}