Meta has released Muse Glimmer, a 30-billion-parameter open-weight AI model designed for agentic tasks such as organizing a schedule, writing code, and conducting research.
The model can run locally on a Mac or PC equipped with a single consumer graphics card without using the cloud. It is available for free download on Hugging Face under an Apache 2.0 license.
Muse Glimmer's size, emphasizing its balance between performance and consumer-hardware compatibility.
Muse Glimmer was developed using distillation from the Muse Spark model family and uses a tiny companion model called the DFlash drafter model to speed up response times.
The DFlash drafter model allows for 3.1x faster responses on an NVIDIA RTX 5090, 1.8x on an M5 Max, and 1.5x on an M4 Max.
Using 4-bit quantization reduces memory needs from over 55 GB to under 20 GB, enabling local execution.
At full precision (fp16), the model requires over 55 GB of memory. With quantization, the 4-bit version reduces requirements to under 20 GB, allowing it to run on systems with 24 GB or 32 GB of GPU or unified memory.
In agentic tests, Muse Glimmer scored 75.5 on MCP Atlas, compared to 54.2 for Gemma and 62.5 for Qwen. It also scored 74.6 on DeepSearch QA and 43.3 on GAIA2, according to Meta. The company claims Muse Glimmer outperforms Gemma4-31B and Qwen3.6-27B in several benchmarks.
Meta announced the model in a video by Mark Zuckerberg, who described his vision of AI as a personalized 'superintelligence'. He stated, 'If the power of superintelligence is held by a small number of individuals, businesses, governments, or AI itself, then that will naturally lead to outcomes that are less favorable for everyone else.'
Any policy that slows American model releases – even by a month – could add significant risk to American leadership while letting foreign models race ahead.
Zuckerberg argued that open source is a positive force against centralization, emphasizing that 'distillation' is an important principle to protect. He also released a 6,500-word essay titled 'The Future is for Everyone' outlining his philosophy, and announced a $1 billion fund to support communities affected by datacenter installations.
The release comes amid a surge in Chinese AI usage. According to OpenRouter data, Chinese LLMs reached 34.25 trillion weekly tokens, while US models contributed 9.17 trillion tokens for the week beginning August 03. Chinese models have led the global AI model usage count for fifteen consecutive weeks.