A Deep Dive into the Traits of Seven Open-Weight Models
A2Agent Team · 2026-10-09T00:00:00Z
Reviewed on 2026-10-09 · 7 models (5 from China, 2 from elsewhere) · about 10 minutes to read
The short answer
For coding and agent development, look at Qwen 3.8-Max and GLM-5.3-Flash. For long context and enterprise workloads, use DeepSeek V4.1 Flash. If you want to self-host something small and use it commercially without license headaches, pick Gemma 4. For multimodal and multilingual work, look at MiniMax M3 and Mistral Large 4.
API prices across the seven run from about $0.15 per million input tokens to about $15 per million output tokens. Compare input, output and cache-hit prices separately, because the cache-hit price is often where you save the most.
1. Seven models at a glance
If you're skimming: Qwen and GLM for coding, DeepSeek for long context, Gemma 4 for lightweight commercial use. Every model below gets the same three spec columns (scale, architecture, usage), followed by what the model is for.
Reading the specs
- Look at active parameters, not total parameters. The total tells you how much disk the weights need; the active count tells you how much compute each generated token costs. Kimi K3 has 2.8 trillion parameters and GLM-5.3-Flash has 320B, nearly a 9× difference, yet GLM is cheap to run because it activates only 18B per step.
- Weight memory ≈ total parameters × bytes per parameter. That's about 2 bytes at FP16 and about 0.5 bytes with INT4 quantization. Inference memory depends mostly on active parameters and the KV cache, and has little to do with the total.
- A 1M-token window costs nothing until you fill it. The KV cache grows linearly with the tokens you actually put in, and that's what you pay for.
DeepSeek V4.1 Flash: a 552B MoE with the KV cache cut to 1/4 of the previous generation
Long context · enterprise workloads · MIT license
Of the flagships here, this one needs the least memory for long-context and enterprise work.
| Scale | Architecture | Usage |
|---|---|---|
| 552B total; 8B / 16B active (input / output) | MoE + CED; 1M context; native vision | MIT license; weights released; launched 2026-09-10 |
DeepSeek built it for complex reasoning, code generation and agent tasks, and says it beats the company's own flagship, DeepSeek-V4-Pro, on performance, cost, speed and total runtime.
- The model is asymmetric: 8B parameters are active on the input side and 16B on the output side, so reading long documents takes less compute while answer quality holds up.
- At 1M context it needs a quarter of the previous generation's GPU memory and an eighth of its storage, which puts long-context work within reach of much smaller budgets.
- It takes text and images directly, with a 1M-token context and up to 384K output tokens.
- Since 2026-09-14, requests to
deepseek-v4-proare billed at Flash rates and old model names keep working, so existing DeepSeek API users don't have to change any code.
In practice, 552B is the size of the weights (about 1.1 TB on disk) and 8B / 16B is the compute that runs each step. Because so few parameters wake up, each token is cheap and throughput is high, but self-hosting means finding a lot of storage first. MoE, the asymmetric layout and KV cache compression together let the same memory hold a longer context, and the 1M window is big enough for whole-repository retrieval and long-running agents. The MIT license lets you use, modify and redistribute it commercially. The API model name is deepseek-flash, old names are routed automatically so upgrading costs nothing, and the Hugging Face repository includes a technical report.
Source: DeepSeek announcement
Qwen 3.8-Max: the first Qwen-Max-class flagship with open weights
Agentic coding · long-horizon autonomous tasks · Qwen License (conditional commercial use)
This is the model to hand a multi-day engineering project and let it run on its own.
| Scale | Architecture | Usage |
|---|---|---|
| 2.4T total; 95B active | Qwen3.5 architecture; DeltaNet + MoE; 1M context; native vision | Qwen3.8-Max License; weights opened 2026-08-12; launched 2026-08-03 |
It's the most capable Qwen model so far and covers coding, real-world work, research and long-horizon tasks. Qwen says the goal is to finish complex tasks end to end and hand back results you can use directly.
- Qwen demonstrated more than 10 days of autonomous development with no human intervention. Over 16 days the model made 265 commits, opened 127 PRs and handled 151 issues, writing, testing and fixing code in a closed loop.
- It reproduced a paper in 5 days with 7,600 lines of code, then came up with its own method that beat the original result by 2.71 points on AIME24. Its PaperBench score of 93.0 is the highest in Qwen's official benchmark table.
- It can write code while operating a software interface the way a person would, then check its own work for misaligned layouts and broken animations and fix what it finds.
- It speaks both the OpenAI and Anthropic API protocols, so Claude Code, Codex and OpenClaw connect directly. The model name is
qwen3.8-max.
The 2.4T figure is the size of the weights (about 4.8 TB at FP16); 95B is the compute per step, which is flagship territory, so most people will reach it through the API. Self-hosting starts at multiple nodes (see section 4). DeltaNet linear attention, interleaved with MoE layers, makes long inputs cheaper to process: the 1M context fits more than 200 pages of documents or 100 hours of video, and Qwen showed full structured extraction over inputs that size.
Ordinary commercial use is free. Products with more than 100 million monthly active users or more than $20M in monthly revenue must give attribution, and a model-serving business with $50M in annual revenue needs a separate license. reasoning_effort has three levels (xhigh / medium / low) for trading depth against cost, and existing agent toolchains work without changes.
Source: Qwen blog
GLM-5.3-Flash: a 320B hybrid-attention model with only 18B active
Agentic coding · native multimodal · MIT license
It improves on its predecessor across the board at a tenth of the cost, which makes it the best value here for high-frequency calls.
| Scale | Architecture | Usage |
|---|---|---|
| 320B total; 18B active | Sparse + linear attention; IndexPool; 1M context; native multimodal | MIT license; weights released; launched 2026-08-26 |
This is the first natively multimodal model in the GLM-5 series. Z.ai says it beats GLM-5.2 everywhere at a tenth of the price, and that its coding and agent benchmark scores are close to Claude Opus 4.8.
- Z.ai's own line is that it "outperforms GLM-5.2 at one-tenth the price."
- Linear attention handles local information and sparse attention handles global retrieval. IndexPool cuts index memory to a quarter, so the 1M context runs fast and cheaply.
- Compared with the GLM-4.5 series, total parameters are similar (320B vs. 355B), but active parameters drop from 32B to 18B and layers from 92 to 45. Inference cost falls by roughly half.
- Z.ai already serves it on large clusters of Chinese AI chips, with end-to-end performance 3× the initial baseline and per-token cost comparable to mainstream NVIDIA GPUs.
Only 18B of the 320B parameters wake up per step, the lightest load of the seven, so per-token cost is very low. The weights take about 640 GB on disk, which puts self-hosting in the middle tier (see section 4). Because the attention design doesn't compute the full context every time, retrieval stays precise across the 1M window at much lower cost, and document work and high-frequency agents gain the most. The MIT license allows commercial use, modification and redistribution. Weights are on Hugging Face and run on SGLang, vLLM and TokenSpeed.
Source: Z.ai blog
Kimi K3: the first open-source flagship at the 3-trillion-parameter scale
Long context · long-horizon coding · custom license
Reach for it when you need the highest capability and cost isn't a concern.
| Scale | Architecture | Usage |
|---|---|---|
| 2.8T total; 104B active per token | KDA + AttnRes; LatentMoE, 16 of 896 experts; 1M context; native vision | Kimi K3 License; weights released 2026-07-27; technical report pending |
It's Moonshot AI's most capable model so far and the first open-source model at the 3-trillion scale. It targets long-horizon coding, knowledge work and deep reasoning, and Moonshot puts its overall scaling efficiency at about 2.5× that of K2.
- KDA linear attention plus attention residuals make very long contexts faster to compute and lighter on memory.
- Only 16 of its 896 experts wake up for each token. That sparsity is what keeps a 2.8T model practical to run.
- Moonshot says it can keep going on long engineering tasks with very little human supervision: working through large codebases, coordinating terminal tools, and using screenshots and visual feedback to refine games, front ends and CAD models.
- Thinking mode is always on.
reasoning_effortsupports low / high / max (default max) to control reasoning depth and cost.
With 104B parameters active per step, it carries the heaviest compute of the seven, and both quality and latency sit at the top. The weights are about 5.6 TB, so self-hosting starts at multiple nodes (see section 4). Native vision accepts images and video, the 1M context covers long codebase tasks, and uploaded video files are referenced with ms://.
Research, deployment, fine-tuning and internal use are free. A model-serving business with more than $20M in revenue over 12 consecutive months needs a separate agreement, and very large products must give attribution. The API model name is kimi-k3. The flagship has to be unlocked with a top-up (minimum ¥10), and new-user vouchers don't apply to it.
Sources: Moonshot AI blog · Kimi platform docs
MiniMax M3: the first flagship to open-source three frontier capabilities at once
Multimodal · 1M context · community license (conditional commercial use)
Pick it if you need coding, long context and multimodality in one open-weight model.
| Scale | Architecture | Usage |
|---|---|---|
| 428B total; 23B active (MoE) | MSA sparse attention; 1M context, 512K guaranteed; text + image + video | MiniMax Community License; weights released; launched 2026-06 |
MiniMax describes it as the first open-weight model with frontier coding, million-token context and native multimodality. It's aimed at AI coding assistants and long-running automation; until now, only a few closed models offered all three.
- Compared with M2 at 1M context, prefill is 9× faster, decode is 15× faster and per-token compute drops to 1/20, which makes long documents and long videos much cheaper to process.
- Text, images and video were mixed in from the first training step, so multimodality is part of the base model.
- It reproduced an ICLR outstanding paper on its own in 12 hours (18 commits, 23 experiment figures). In another run it optimized a CUDA kernel over 147 iterations and raised hardware utilization from 7.6% to 71.3%, a 9.4× speedup, with no human intervention.
- The
thinkingparameter accepts enabled / adaptive / disabled; adaptive decides for itself when to reason. API caching is automatic and needs no setup.
At 23B active out of 428B, it sits between GLM and the big flagships, and long agent runs stay affordable. The weights are about 856 GB, upper-middle tier for self-hosting (see section 4). MSA is a sparse attention operator designed for million-token contexts. The API guarantees at least 512K of context, and long-video understanding and long-horizon agents benefit most.
Non-commercial use is free. Commercial use requires the credit "Built with MiniMax M3", and annual revenue above $20M requires written authorization. It has the widest framework support of the seven: SGLang, vLLM, Transformers, KTransformers, unsloth and ATOM. The technical report is arXiv:2606.13392.
Sources: MiniMax model page · Hugging Face model card
Gemma 4: an Apache 2.0 family that runs on everything from a Raspberry Pi to a workstation
Lightweight self-hosting · on-device · Apache 2.0
Pick it to run a model locally on a phone or a Raspberry Pi, or to use one commercially with no strings attached.
| Scale | Architecture | Usage |
|---|---|---|
| 5 sizes: E2B, E4B, 12B, 26B A4B, 31B Dense | Built on Gemini 3 research; 256K context (128K for small sizes); MTP multi-token prediction | Apache 2.0; weights released; launched 2026-03 |
This is Google DeepMind's most capable Gemma family so far. Google says every size performs at the frontier for its class, from phones and edge devices (E2B, E4B) to consumer GPUs and workstations (26B, 31B), across reasoning, agent workflows, coding and multimodal understanding.
- Multi-token prediction (MTP) speeds up decoding by up to 2.2× on mobile GPUs and up to 1.5× on CPUs, with hardware such as the M4 MacBook seeing a clear gain. Google reports no quality loss.
- The E2B model in on-device format is 2.58 GB. It runs at about 8 tokens per second on a Raspberry Pi 5 and decodes at 56 tokens per second on the iPhone 17 Pro's GPU. The LiteRT-LM toolchain converts fine-tuned models to that format.
- E2B and E4B are for phones and edge devices, 12B sits in the middle, the 26B A4B MoE favors efficiency and the 31B Dense favors quality. No other family among the seven spans devices to workstations.
- Yale's Gemma-based Cell2Sentence-Scale 27B found a potential new pathway for cancer treatment, DolphinGemma is used to study dolphin communication, Living Models uses Gemma 4 to decode plant DNA, and Syngenta uses it to identify plants more accurately.
The 31B Dense weights are about 61 GB at FP16, while E2B in on-device format is 2.58 GB, two orders of magnitude apart within one family. That gives Gemma 4 the lowest self-hosting bar of the seven (see section 4) and makes offline use practical. Its 256K context (128K on small sizes) is short next to the others, but MTP makes generation fast. It handles reasoning, coding and agent tasks, and the quick on-device decoding makes it the strongest fit for offline local apps.
Apache 2.0 allows commercial use, modification and redistribution with no extra conditions. That makes it the easiest of the seven to customize and redistribute, and it's a common base for governments and companies building sovereign AI systems. You can try it free in AI Studio; Vertex hosted pricing covers only 26B A4B.
Sources: Gemma 4 model card · LiteRT on-device guide · Gemmaverse
Mistral Large 4: the strongest open-weight flagship for cybersecurity
Cybersecurity · European sovereignty · weights pending
Pick it for offensive and defensive security work where a closed model's refusals would get in the way.
| Scale | Architecture | Usage |
|---|---|---|
| 1.05T total; 52B active | MoE; 1M context; native multimodal; trained on 160+ languages | Preview, API only; weights due at month-end; license not yet announced |
It's Mistral's largest and most capable flagship so far, nicknamed "Le Chonk". Mistral says it competes with the strongest open-source models in the world, clearly beats any US or European open-weight model, and leads open models on enterprise workloads such as cybersecurity, finance and law.
- It ranks in the global top five on the AA cyber index and first among open-weight models from outside China. It scored 82% on a test of reproducing and patching vulnerabilities, the highest of any model, and solved 93% of Cybench challenges. Claude Opus 5.5 and GPT-6 Astra scored close to zero on the same test because they refuse to carry out the tasks, and defenders need a system that won't.
- On Dense 200 it scored 42%, ahead of GPT-6 Astra at 41%. Mistral demonstrated it inspecting gigapixel satellite imagery for disaster relief, zooming in to check engineering drawings and pulling evidence out of PDFs.
- In vals.ai's evaluation it beat GPT-6 Astra on both legal and financial tasks. It outperformed every open-source model on Harvey's legal agent benchmark, and its 59.9% on AutomationBench is ahead of Kimi K3 and DeepSeek V4 Pro.
- Mistral trained it from scratch on 3,800 Grace Blackwell GPUs in its own European data centers. European deployments run end to end under EU law, independent of other cloud providers, and the training data covers more than 160 languages, including every official EU language.
At 1.05T parameters with 52B active per step, it's Mistral's largest model. For now it's only available through the preview API; once the weights are out, self-hosting will fall in the giant tier (see section 4). Mistral calls its image understanding "a step change", and perception-heavy agent work (satellite imagery, engineering drawings, evidence retrieval from PDFs) is where it does best.
The weights are due at the end of the month and the license hasn't been announced. You can try it in Mistral Studio, but check the current status of the weights and license before using it commercially. Mistral says it will publish more about the architecture and post-training methods when the weights come out.
Source: Mistral AI announcement
2. Price comparison
These numbers come from the Vals Index, which measures how often each model completes agent tasks in finance, coding, law and other fields, along with what each task costs. The cost differences run to an order of magnitude and beyond.
| Model | Type | Vals Index (completion rate) | Cost per task (USD) | vs. cheapest |
|---|---|---|---|---|
| Claude Fable 5.1 | Closed | 68.83% | $28.92 | 95× |
| GPT-6 Astra | Closed | 66.61% | $19.09 | 63× |
| Gemini 3.8 Flash | Closed | 62.25% | $5.39 | 18× |
| DeepSeek V4.1 Flash | Open | 57.86% | $0.30 | 1× (baseline) |
| Kimi K3 | Open | 57.81% | $6.47 | 21× |
| GLM 5.3 | Open | 56.97% | $6.73 | 22× |
| Qwen 3.8 Max | Open | 51.84% | $4.00 | 13× |
| MiniMax-M3 | Open | 42.72% | $2.50 | 8× |
On the same set of real tasks, the most expensive closed model costs $28.92 per task and the cheapest open model $0.30. The price gap is 95×, and the completion rates are only about 11 percentage points apart. Data checked on 2026-10-09 against the official Vals Index page.
Further reading: Closed vs. Open-Source LLMs in 2026
3. Pick by use case
The picks below start from the job you need done. Each of the four common scenarios gets a top pick and the reasoning behind it.
Scenario 1: coding and agent development
Writing code, running tool calls and building workflows. Top picks: Qwen 3.8-Max and GLM-5.3-Flash.
By their makers' accounts, both lead in agentic coding. GLM-5.3-Flash suits high-frequency calls in particular: with a $0.15 input price and an MIT license, an agent loop can burn through tokens without the bill hurting. Switch to Qwen 3.8-Max when coding quality matters most.
Scenario 2: long context and enterprise workloads
Long documents, knowledge bases and enterprise automation. Top pick: DeepSeek V4.1 Flash.
KV cache compression brings the memory and storage cost of long context down to a quarter of the previous generation, and the 1M context fits long documents, knowledge bases and long-running agents. If you need stronger long-horizon engineering, Kimi K3 is the step up, but read its commercial license terms first.
Scenario 3: lightweight self-hosting and commercial compliance
Running locally, using the model commercially and keeping GPU spend down. Top picks: Gemma 4 and Kimi K3.
Gemma 4's Apache 2.0 license allows commercial use with no extra conditions, and the 31B Dense model runs on consumer hardware once quantized. Kimi K3 is the largest model here and can do the most, but its license adds conditions you'll need to check before commercial use. Teams on a tight budget can also start with free tiers.
Scenario 4: multimodal and multilingual
Understanding images, working across languages and serving several markets. Top picks: MiniMax M3 and Mistral Large 4.
MiniMax M3 is natively multimodal and already well proven in coding and automation. Mistral Large 4 supports more than 160 languages, which matters for products serving several markets. It's still in preview, though, and its pricing and weight status may change, so confirm the latest official details before you commit to a launch.
4. Self-hosting cost estimates
You can also deploy these models yourself. The figures in the table are rough orders of magnitude, meant only to tell you whether self-hosting is within reach.
| Tier | Representative models | Memory (rough) | GPUs (reference) |
|---|---|---|---|
| Lightweight, ≤31B | Gemma 4 31B Dense | About 60-80 GB at FP16; 24-40 GB after INT4 quantization | 1× 80 GB or a consumer GPU |
| Mid-size MoE, 320-550B | GLM-5.3-Flash | Weights in the hundreds of GB; inference memory set by active parameters (18B) | 4-8× 80 GB |
| Giant, 1T+ | Qwen 3.8-Max / Kimi K3 / Mistral Large 4 | Weights in the TB range; needs multi-node storage and loading | Multi-node, 16+× 80 GB |
These are estimates, not official hardware recommendations. Real requirements depend on the quantization scheme, batching and inference framework.
The memory column covers two separate costs: weight memory and inference memory. Weight memory is roughly total parameters × bytes per parameter. FP16 uses 2 bytes per parameter and INT4 uses 0.5, which is where the lightweight tier's 60 to 80 GB (24 to 40 GB quantized) comes from; Gemma 4 31B is about 61 GB at FP16 and about 15 GB at INT4. Inference memory depends mainly on active parameters and the KV cache. GLM-5.3-Flash has 640 GB of weights, but it wakes only 18B per step, so running it doesn't take 640 GB of GPU memory. For the giant tier, "TB range" refers to the weights themselves: about 4.8 TB for Qwen 3.8-Max and 5.6 TB for Kimi K3. Storing them alone takes several machines, which is why people mostly use this tier through APIs.
In the GPU column, "80 GB" means enterprise cards such as the H100 or A100, each several times the price of a consumer card. "1× 80 GB" means a single card is enough, "4-8× 80 GB" means four to eight cards in one machine, and "16+× 80 GB" means a cluster of several machines. A consumer card with around 24 GB is enough only for a quantized lightweight model or a small on-device model such as Gemma E2B or E4B.
Some practical advice:
- Don't rush to buy GPUs. Get your product working on an API, measure real usage, then work out whether self-hosting would pay for itself.
- If you want to try self-hosting, start with Gemma 4. One card, even a consumer one, is enough to begin.
- The mid-size tier (models like GLM) only makes sense for teams with high, steady usage. Even quantized it needs four to eight 80 GB cards, so compare a month of GPU rental against your API bill first.
- For the giant tier, start on the API.
- If your team won't keep the hardware busy over the long run, the API will cost less.
5. Wrap-up
A few patterns hold across all seven models. Total parameters keep growing while active parameters stay small, and the active count is what drives cost. The price gap between open and closed models is measured in multiples: $0.30 against $28.92 for the same task, or 95×. Licenses fall into three groups: MIT, Apache 2.0, and licenses with commercial conditions. Choosing a model comes down to your use case, how you'll deploy it and which license terms you can accept. The table maps each model to the teams it suits.
| Model | Best for | Why |
|---|---|---|
| DeepSeek V4.1 Flash | Teams working with long documents, knowledge bases and enterprise agents; high-volume callers sensitive to cost per task | 1M context plus KV cache compression (1/4 the memory); at $0.30 per task it's the cheapest in the Vals data, and the best value for long-context work |
| Qwen 3.8-Max | Platform teams that need multi-day autonomous engineering; heavy agent orchestrators | Official demo of 10+ days of unattended coding, plus a paper reproduction that beat the original; dual-protocol support lets Claude Code and Codex connect without changes |
| GLM-5.3-Flash | Product teams making high-frequency calls; agentic coding and office automation | Only 18B active and, per Z.ai, a tenth of its predecessor's cost, with an MIT license on top, so token spend in agent loops is easiest to control |
| Kimi K3 | Deep-reasoning and long-horizon coding teams that put quality first and don't mind cost | 2.8T parameters and 104B active per token, the most of the seven; 57.81% Vals completion rate, second among open models. It also has the strictest commercial license terms and the highest compute cost |
| MiniMax M3 | Teams that need coding, long context and multimodality together; long-video understanding | First open model with all three frontier capabilities; MSA speeds up prefill/decode at 1M context by 9×/15×; supported by six deployment frameworks |
| Gemma 4 | On-device, offline and private-deployment teams; enterprises sensitive to compliance and data sovereignty | E2B on-device format is 2.58 GB and runs on a Raspberry Pi; Apache 2.0 with no extra commercial conditions; the lowest self-hosting bar of the seven |
| Mistral Large 4 | Security and red teams; enterprises that require European data sovereignty | 93% on Cybench and 82% on vulnerability reproduction and patching, the highest of any model; beats GPT-6 Astra on legal and financial tasks in third-party evaluation; EU-local deployment. Still in preview, with weights not yet released |
The table draws on the earlier sections: official release information (section 1), Vals Index completion rates and costs (section 2) and the self-hosting estimates (section 4). Data checked on 2026-10-09.