Which open models of August 2026 are worth running on your own hardware
Table of contents
- Key takeaways
- What this review covers and what it does not
- What actually shipped with downloadable weights
- The names circulating with no weights behind them
- What each licence actually permits
- Muse Glimmer and the 24 GB ceiling
- What to pick for the memory you have
- What to do about the upgrade treadmill
- Frequently asked questions
- How do I check for myself whether a model has open weights?
- Is Nemotron 3.5 Lightning worth it with 24 GB?
- Can a model with open weights be used in a paid product?
- Conclusion
- Sources
August 2026 produced about a dozen models with genuinely downloadable weights, and several names that exist only as a hosted API. The ones that change what fits on your own machine are Muse Glimmer 30B and Qwen3.8-27B, both Apache 2.0, plus Ling-3.0-tiny for smaller hardware.
August 2026 shipped more open weights than fit on a graphics card. The volume is not the problem. The problem is that a good share of the names circulating in this month’s round-ups have no downloadable weights at all: they are API services wearing an open-model name. This review separates the two, states the licence on each model, and gives the memory each one actually needs.
Key takeaways
- The release that most changes what fits on a home machine is Muse Glimmer 30B, from Meta Superintelligence Lab, Apache 2.0 licensed and designed from the start to fit in 24 GB.
- Qwen3.8-27B is the month’s other complete release: 27 billion dense parameters, image and video understanding, and Apache 2.0 with no small print.
- Muse Spark 1.2 has no published weights. Meta announced on 10 August that it would open them; as of 30 August they had not appeared.
- GLM-5.2 Turbo, Seed 2.1 Turbo and DeepSeek V4 Flash Vision Exp exist only as APIs: none has a repository in its own vendor’s organisation.
- Qwen3.8-Max is not a closed model. It is the hosted product name for Qwen3.8-2.4T-A95B, whose weights are published, though at 2.4 trillion parameters you will not be running them at home.
What this review covers and what it does not
It covers what was published between 1 and 30 August 2026, and it was written on 30 August 2026. Everything here was checked that day by opening the vendor’s repository and reading the model card and the licence file. None of it comes from someone else’s summary.
Worth saying, because an article with a date in the title ages fast. In three months there will be better models and some of these licences will have been revised. What will not age is the method: before believing a model is open, look for its repository in the vendor’s official organisation. If it is not there, there are no weights.
What actually shipped with downloadable weights
These are the August models whose weights you can download today. The memory column is the real size of the quantised download, not the parameter count:
| Model | Date | Parameters | Licence | Fits on your own machine? |
|---|---|---|---|---|
| Muse Glimmer 30B | 9 Aug | 29.6B dense | Apache 2.0 | Yes, 17 GB with the official quantisation |
| Qwen3.8-27B | 5 Aug | 27.8B dense | Apache 2.0 | Yes, 18 GB at q4 |
| Granite 4.2 (3B / 8B / 30B) | 7 Aug | 3.66B / 8.79B / 29.28B | Apache 2.0 | Yes, from 2.2 GB |
| Ling-3.0-tiny | 10 Aug | 7.9B total, 1.3B active | MIT | Yes, 8.34 GiB at FP8 |
| LFM2.5-VL-3B | 11 Aug | 3.12B | LFM Open License 1.0 | Yes |
| Nemotron 3.5 Lightning | 1 Aug | 31.6B total, 3B active | OpenMDW 1.1 | Only just, 25 GB |
| GLM-5.3-Flash | 25 Aug | 320B total, 18B active | MIT | No, server class |
| Qwen3.8-Flash-Next | 24 Aug | 180B | Qwen Community 1.0 | No |
| Qwen3.8-2.4T-A95B | 8 Aug | 2.4T total, 95B active | Qwen3.8-Max | No |
| DeepSeek V4-Pro-0813 | 13 Aug | 1.65T | MIT | No |
| Hy4-preview (Tencent) | 27 Aug | 780B | Apache 2.0 | No |
The first five rows are the ones that matter to somebody with a consumer card in front of them. The rest are genuinely open weights, under serious licences, that need a rack to start.
The names circulating with no weights behind them
This is where checking it yourself pays. Four of the month’s most repeated names are not open weights:
Muse Spark 1.2 is the most confusing case. Meta announced on 10 August that it would open the weights, reversing the closed policy of the three previous Muse Spark releases. A promise is not a file: as of 30 August there is no Muse Spark repository on Hugging Face. What Meta did publish, on 9 August, was Muse Glimmer 30B, a model distilled from Muse Spark and licensed Apache 2.0. If you read that "Meta opened Muse Spark 1.2", the model you can download today is called Glimmer.
GLM-5.2 Turbo, Seed 2.1 Turbo and DeepSeek V4 Flash Vision Exp appear in the month’s lists with a price per million tokens, which is precisely the clue. None has a repository in zai-org, ByteDance-Seed or deepseek-ai. They are hosted variants. Two date corrections that keep getting repeated: GLM-5.2 is a June release, not August, and DeepSeek V4-Flash is from 31 July.
Qwen3.8-Max deserves a different nuance, because it is a naming confusion rather than a closure. The Qwen3.8-2.4T-A95B card says it plainly: "Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools". Max is the service; the weights of that same model are published. The card adds a line that is genuinely news: "For the first time, Qwen3.8 brings a Qwen-Max-class model to open release".
On the "fourteen new models in August" count doing the rounds: I found no primary source supporting it, so I am not repeating it. The ones I could verify individually are the eleven in the table above.
What each licence actually permits
Downloadable weights do not mean you can sell the result. The licences are worth reading, because the differences are wide:
- Apache 2.0 (Muse Glimmer, Qwen3.8-27B, Granite 4.2, Hy4) and MIT (GLM-5.3-Flash, Ling-3.0, DeepSeek V4-Pro): commercial use with no conditions. These are the clean ones.
- Qwen3.8-Max Licence: free, but it requires the model name displayed in your interface above 100 million monthly users or US$20 million monthly revenue, and a separate licence if you run a model-as-a-service business above US$50 million over twelve months.
- Qwen Community License 1.0 (Qwen3.8-Flash-Next): same idea, but its second clause has no threshold. Any model-as-a-service business needs permission, whatever your revenue. Oddly, it is stricter than the licence on the bigger model.
- GLM-5.3 Licence: it asks for a security review only from businesses above US$10 billion. For everyone else it is permissive.
- LFM Open License 1.0 (Liquid AI): derived from Apache, but its section 5 conditions commercial use on not exceeding a US$10,000,000 annual revenue threshold. Above that figure, commercial use, in the licence’s own words, "is not licensed under this Agreement". It is the most restrictive in this review.
- OpenMDW 1.1 (Nemotron 3.5 Lightning), NVIDIA’s licence for weights and data, with public text[1].
None of these models is free software in the strict sense: none published its training data. Open weights and open source remain different things.
Muse Glimmer and the 24 GB ceiling
Of everything in August, this is the release that changes the arithmetic. Muse Glimmer has 29.6 billion parameters, of which roughly 1.8 billion are a ViT-G/14 perception encoder that lets it read screenshots, charts and documents. Context of 131,072 tokens, more than a hundred languages, and knowledge up to 4 January 2026.
What matters is that Meta built it to fit a consumer card. Its card puts it this way: "We use quantization techniques to compress the model’s weights to approximately 4-bit precision, shrinking the language model to under 20 GB. This leaves enough headroom for the model’s KV cache, the perception encoder for image understanding, and the speculative decoding drafter to run simultaneously within a 24 GB or 32 GB envelope".
They also give the degradation figure, which almost nobody publishes. The 17 GB variant, aimed at 24 GB of memory, loses 1.0% on average across fifteen benchmarks; the dynamic variant for 32 GB loses 0.2%. If you have wrestled with model quantisation and llama.cpp, you will know how rare it is for a vendor to put that number in writing.
The model also ships a speculative-decoding drafter that proposes blocks of sixteen tokens at once. Meta’s measured figures, at batch size one with greedy decoding:
| Machine | No speculation | With speculation | Speedup |
|---|---|---|---|
| NVIDIA RTX 5090 | 74.9 tokens/s | 233.4 tokens/s | 3.1x |
| Apple M4 Max | 23.7 tokens/s | 37.8 tokens/s | 1.5x |
| Apple M5 Max | 26.6 tokens/s | 50.2 tokens/s | 1.8x |
A 5090 producing 233 tokens per second from a thirty-billion-parameter model with vision is something that a year ago needed a rented machine.
What to pick for the memory you have
The number that governs is not the parameter count, it is how much memory the quantised file plus its context cache occupies. With that in mind:
8 GB of VRAM. Granite 4.2 3B takes 2.2 GB and Granite 4.2 8B at q4 comes in at 5.1 GB, both Apache 2.0 with 128K native context. This is the tier where a clean licence matters most, because it is usually learning hardware or a small project that might grow.
16 GB of VRAM. Ling-3.0-tiny lands here, and it is my pick of the month for power against consumption: 7.9 billion total parameters but only 1.3 billion active per token, MIT licensed, with a memory peak of 8.34 GiB at 8K context. That leaves plenty of room for a long context. For image work, LFM2.5-VL-3B fits comfortably, with the caveat of its licence.
24 GB of VRAM. The tier for Muse Glimmer with its 17 GB quantisation, and for Qwen3.8-27B, which at q4 and in NVFP4 format comes to 18 GB. Granite 4.2 30B also fits in 18 GB. If your work is agentic and involves screenshots, Glimmer; if you need enormous context, Qwen3.8-27B starts at 262,144 tokens and stretches to a million.
Apple Silicon with unified memory. The allocation is more generous because memory is shared, but bandwidth governs speed. Ling-3.0-tiny reaches 86 to 90 tokens per second on an M4 Pro MacBook. Muse Glimmer sits at 37.8 on an M4 Max and 50.2 on an M5 Max, usable for conversation but not for batch work. If you go this way, the guide to installing Ollama on macOS saves you the first afternoon.
Nearly all of it downloads with one command:
ollama pull granite4.2:3b # 2.2 GB, 8 GB tier
ollama pull granite4.2:8b # 5.3 GB, 8 GB tier
ollama pull qwen3.8:27b-nvfp4 # 18 GB, 24 GB tier
ollama pull granite4.2:30b # 18 GB, 24 GB tier
ollama pull nemotron-3.5-lightning:30b-a3b # 25 GB, wants 32 GB
If this is your first time, start with installing Ollama. And before picking a model for an agent, check which ones handle tool calls well, because the differences among these eleven are wider there than in any general-knowledge benchmark.
What to do about the upgrade treadmill
Eleven models in one month is dizzying, and the dizziness has a cost: every model swap forces a pass over your prompts, your tools and your evaluations. My position after this review is unheroic.
Change model when a hard limit breaks, not when a table moves by a point. A hard limit is something now fitting on your card that did not fit before, a capability appearing that you did not have (vision, million-token context, reliable tool calls), or a licence that stops working for you. August offers two of those real jumps: a thirty-billion-parameter agentic multimodal model in 24 GB, and an MIT reasoning model running at 90 tokens per second on a laptop.
The rest is version noise. One practical note: pin the model version in your configuration. Ollama tags get recut and rebuilt, and discovering that your agent changed quantisation overnight is a lost afternoon. If you want the format detail, it is in Gemma 4 running with Ollama.
Frequently asked questions
How do I check for myself whether a model has open weights?
Look for the repository in the vendor’s official Hugging Face organisation, not in the general search. If the only result is a third-party copy, there is no official release. Then open the repository’s LICENSE file, because the label the site shows sometimes reads "other" and hides revenue conditions.
Is Nemotron 3.5 Lightning worth it with 24 GB?
It is tight. There are 31.6 billion total parameters with only 3 billion active, so it generates fast, but the download is 25 GB and the full weights have to be resident in memory. With 32 GB it is comfortable; with 24 GB you will be trimming context and still running close to the edge.
Can a model with open weights be used in a paid product?
It depends on the licence, and in August the range is wide. With Apache 2.0 or MIT, yes, with no conditions. With the LFM Open License, only while you invoice less than 10 million dollars a year. With the Qwen licences you can sell a product, but offering the model as a service to third parties requires separate permission.
Conclusion
August 2026 was a good month for anyone running models on their own machine, though not quite for the reasons the round-ups give. The headline is not the 2.4-trillion-parameter model: it is that Meta published, under Apache 2.0, an agentic multimodal model of thirty billion parameters measured at 233 tokens per second on a consumer card, and that an MIT-licensed MoE gives reasoned answers at 90 tokens per second on a laptop. The other thing the month leaves behind is a cheap reminder: check the repository before believing the word "open". The Spanish version of this article is at Qué modelos abiertos de agosto de 2026 merece la pena ejecutar.
Sources
- public text
- Meta Superintelligence Lab, Muse Glimmer 30B model card
- Qwen, Qwen3.8-27B model card
- Qwen, Qwen3.8-2.4T-A95B model card and licence
- Z.AI, GLM-5.3-Flash model card
- IBM, Granite 4.2 8B model card
- inclusionAI, Ling-3.0-tiny model card
- NVIDIA, Nemotron 3.5 Lightning 30B-A3B model card
- Liquid AI, LFM2.5-2.6B and the LFM Open License 1.0
- Ollama, qwen3.8 model library