DeepSeek API
DeepSeek sells two models here, V4 Pro and V4 Flash. Both run a 1M context and both are tagged for reasoning, so the only split is price: Flash lists at roughly a third of Pro on input and output alike.
2 models
Side-by-side per-1M-token rates for every supported model
| Model | Platform | Context | Type | Input /1M | Output /1M | Discount |
|---|---|---|---|---|---|---|
| Z.ai | 200K | chatagent | $1.00 | $3.20 | … | |
| Z.ai | 200K | chatcoding | $1.15 | $4.00 | … | |
| Z.ai | 1M | chatcoding | $1.40 | $4.40 | … | |
| Z.ai | 1M | chatcoding | $1.40 | $4.40 | … | |
| Z.ai | 1M | chatvision | $0.150 | $0.500 | … | |
| Moonshot | 256K | chatvision | $0.600 | $3.00 | … | |
| Moonshot | 256K | chatvision | $0.950 | $4.00 | … | |
| Moonshot | 256K | chatcoding | $0.950 | $4.00 | … | |
| Moonshot | 1M | chatvision | $3.00 | $15.00 | … | |
| DeepSeek | 1M | chatreasoning | $0.435 | $0.870 | … | |
| DeepSeek | 1M | chatreasoning | $0.140 | $0.280 | … | |
| Qwen | 1M | chatvision | $0.500 | $3.00 | … | |
| Qwen | 1M | chatvision | $0.200 | $1.20 | … | |
| Qwen | 1M | chatvision | $0.200 | $0.800 | … | |
| Qwen | 1M | chatvision | $1.20 | $4.80 | … | |
| Qwen | 1M | chatreasoning | $1.70 | $5.10 | … | |
| Qwen | 1M | chatvision | $0.150 | $0.470 | … | |
| Qwen | 1M | chatreasoning | $2.00 | $6.00 | … | |
| MiniMax | 1M | chatcoding | $0.300 | $1.20 | … | |
| MiniMax | 200K | chatcoding | $0.300 | $1.20 | … | |
| MiniMax | 1M | chatagent | $0.300 | $1.20 | … |
Prices synced with platform billing (USD per 1M tokens).
The five model families, and what each one covers
DeepSeek sells two models here, V4 Pro and V4 Flash. Both run a 1M context and both are tagged for reasoning, so the only split is price: Flash lists at roughly a third of Pro on input and output alike.
2 models
Z.ai ships five GLM models. GLM-5 and GLM-5.1 stop at 200K, tagged for agent work and coding. GLM-5.2 and GLM-5.3 reach 1M at the same rate as each other. GLM-5.3 Flash keeps that 1M window, reads images and costs about a tenth of GLM-5.3.
5 models
Kimi K2.5, K2.6 and K2.7 Code all run 256K. The first two read images, the third is the coding-tagged one. K3 is the outlier: a 1M window at roughly three times the K2.6 input rate and nearly four times the output rate.
4 models
Qwen is the widest family here, seven models all running a 1M context. The three Flash releases are the cheap tier and get cheaper with each version. Qwen3.7 MAX and Qwen3.8 MAX carry the reasoning tag and the highest rates. Plus sits in between.
7 models
All three MiniMax models list at the same price, so window and tag decide. M2.5 and M3 run 1M, M2.7 stops at 200K. M2.5 and M2.7 are tagged for coding, M3 for agent work.
3 models