Cheapest AI Models for Code in 2026: Price Comparison on A2Agent

A2Agent Team ยท 2026-09-29T00:00:00Z

Say a coding agent runs 500 times a month. Each run reads about 200,000 tokens of repository context and writes about 40,000 tokens of plans, diffs and test fixes. On Claude Opus 4.8 at list price, that month costs $1,000. The same 500 runs on DeepSeek V4 Flash through A2Agent cost $11.76.

Most of that gap is the price per million tokens. Some of it comes from the way coding agents use tokens, which looks very different from a chat session. The prices below are live A2Agent prices from September 29, 2026, for every model we sell.

Price tags for AI coding models: DeepSeek V4 Flash $0.168, DeepSeek V4 Pro $0.52, MiniMax M2.5 $0.60, Kimi K2.7 Code $2.80, GLM-5.3 $3.08, Claude Opus 4.8 $25, GPT-5.5 $30 per million output tokens

Quick answer

The cheapest model on A2Agent is DeepSeek V4 Flash: $0.084 per million input tokens and $0.168 per million output, with a 1M-token context window. DeepSeek's own price is $0.14 / $0.28, so that is 40% off.

Other prices worth knowing:

  • MiniMax M2.5 and M2.7 are the cheapest models tagged for coding, at $0.15 / $0.60 (half of MiniMax's official price).
  • DeepSeek V4 Pro is the cheapest strong reasoning model, at $0.261 / $0.522 with a 1M context.
  • GLM-5.2 ($0.84 / $2.64) and GLM-5.3 ($0.98 / $3.08) are tagged for coding and have a 1M context. Kimi K2.7 Code ($0.665 / $2.80) is also tagged for coding and has 256K.
  • Claude Opus 4.8 lists at $5 / $25 and GPT-5.5 at $5 / $30 on their vendors' own APIs. We don't sell either. Per output token, GPT-5.5 costs about 180 times what DeepSeek V4 Flash does.

Full price table: 21 models on A2Agent

Prices are in US dollars per million tokens, sorted by output price. "Official" is the vendor's own list price. "Tag" is the category shown in the A2Agent model catalog.

# Model Tag A2Agent input A2Agent output Official input / output Discount Context
1 DeepSeek V4 Flash reasoning $0.084 $0.168 $0.14 / $0.28 40% 1M
2 GLM-5.3 Flash vision $0.105 $0.35 $0.15 / $0.50 30% 1M
3 Qwen3.8 Flash vision $0.135 $0.423 $0.15 / $0.47 10% 1M
4 DeepSeek V4.1 Flash โ€” $0.136 $0.504 $0.17 / $0.63 20% โ€”
5 DeepSeek V4 Pro reasoning $0.261 $0.522 $0.435 / $0.87 40% 1M
6 Qwen3.7 Flash vision $0.14 $0.56 $0.20 / $0.80 30% 1M
7 MiniMax M2.5 coding $0.15 $0.60 $0.30 / $1.20 50% 1M
8 MiniMax M2.7 coding $0.15 $0.60 $0.30 / $1.20 50% 200K
9 MiniMax M3 agent $0.15 $0.60 $0.30 / $1.20 50% 1M
10 Qwen3.6 Flash vision $0.16 $0.96 $0.20 / $1.20 20% 1M
11 GLM-5 agent $0.60 $1.92 $1.00 / $3.20 40% 200K
12 Kimi K2.5 vision $0.42 $2.10 $0.60 / $3.00 30% 256K
13 Qwen3.5 Plus vision $0.35 $2.10 $0.50 / $3.00 30% 1M
14 GLM-5.2 coding $0.84 $2.64 $1.40 / $4.40 40% 1M
15 GLM-5.1 coding $0.805 $2.80 $1.15 / $4.00 30% 200K
16 Kimi K2.7 Code coding $0.665 $2.80 $0.95 / $4.00 30% 256K
17 GLM-5.3 coding $0.98 $3.08 $1.40 / $4.40 30% 1M
18 Qwen3.7 Plus vision $0.96 $3.84 $1.20 / $4.80 20% 1M
19 Qwen3.7 MAX reasoning $1.36 $4.08 $1.70 / $5.10 20% 1M
20 Qwen3.8 MAX reasoning $1.80 $5.40 $2.00 / $6.00 10% 1M
21 Kimi K3 vision $2.40 $12.00 $3.00 / $15.00 20% 1M
โ€” Claude Opus 4.8 (not on A2Agent) โ€” โ€” โ€” $5.00 / $25.00 โ€” 1M
โ€” GPT-5.5 (not on A2Agent) โ€” โ€” โ€” $5.00 / $30.00 โ€” 1M

A tag says what a model is mainly built for, and models do fine outside their tag. DeepSeek V4 Pro and V4 Flash are tagged for reasoning, and both work well as coding agents. Kimi K3 is tagged for vision, and it recently topped the Frontend Code Arena.

The spread is about 180ร—

Every model on A2Agent costs between $0.168 and $12 per million output tokens. Twenty of the 21 cost less than $6. The two closed frontier models cost $25 and $30.

Log-scale number line of output price per 1M tokens: all 21 A2Agent models between $0.168 and $12, Claude Opus 4.8 at $25 and GPT-5.5 at $30, about 180 times apart

Ten models charge less than $1 per million output tokens: three from DeepSeek, four Qwen and GLM flash models, and three MiniMax models that share one price. The coding-tagged models fall into two groups. MiniMax M2.5 and M2.7 cost $0.60, and the GLM-5.x coding models and Kimi K2.7 Code cost $2.64 to $3.08.

The seller matters too. The model behind a DeepSeek V4 Flash call on A2Agent is the same one DeepSeek serves, and the bill is 40% lower. Across the catalog our discount runs from 10% to 50% below each vendor's own price, applied equally to input and output.

Why coding agents spend tokens differently

A chat request is small on both sides. You paste a snippet, ask a question, and get a paragraph back. A coding agent works in a loop. It reads files, plans, edits, runs the tests, reads the test log and tries again. Each turn usually sends the whole conversation back to the model: the system prompt, tool definitions, every file already read and every test log already seen.

Diagram comparing a chat request with a coding agent loop of read files, think, edit and run tests, read test log, where input tokens grow each turn until the PR is ready

So input tokens make up most of the bill. By turn 14 the agent may be sending 200K tokens to change 30 lines. Most price comparisons lead with output prices, but for a coding model the input price deserves at least as much attention.

Thinking tokens are billed as output. In reasoning mode a model can write thousands of hidden tokens before its first line of code, so two models with the same output price can produce very different bills if one of them thinks three times as long.

Caching helps a lot here. The repeated prefix (system prompt, tool schemas, files already read) can often be billed at the provider's cheaper cache-read rate. Keep that prefix stable, and keep timestamps and random IDs out of the top of your prompt, because they break the cache.

When you compare models, the number to track is what each merged PR costs you, which the per-token price only partly predicts.

What a month of coding costs

We priced three workloads at A2Agent rates, with official prices alongside. None of these numbers include caching, so a setup with good cache hit rates will pay less.

1. Solo developer, agent in the IDE

This one assumes 40 agent tasks a day for 30 days, at about 30K input and 3K output tokens per task. That comes to 36M input and 3.6M output tokens a month, a typical load for someone working in Cline, Roo Code or Claude Code.

Model A2Agent / month Official / month
DeepSeek V4 Flash $3.63 $6.05
MiniMax M2.5 $7.56 $15.12
DeepSeek V4 Pro $11.28 $18.79
Kimi K2.7 Code $34.02 $48.60
GLM-5.3 $46.37 $66.24
Claude Opus 4.8 โ€” $270
GPT-5.5 โ€” $288

2. PR review bot for a team

A team bot reviews 2,000 pull requests a month. Each review reads about 50K input tokens (the diff plus surrounding files) and writes 4K. That comes to 100M input and 8M output tokens a month.

Model A2Agent / month Official / month
DeepSeek V4 Flash $9.74 $16.24
MiniMax M2.5 $19.80 $39.60
DeepSeek V4 Pro $30.28 $50.46
Kimi K2.7 Code $88.90 $127.00
GLM-5.3 $122.64 $175.20
Claude Opus 4.8 โ€” $700
GPT-5.5 โ€” $740

3. Autonomous coding agent in production

Here the agent takes an issue all the way to a PR, 500 times a month, at about 200K input and 40K output tokens per run. That comes to 100M input and 20M output tokens a month.

Model A2Agent / month Official / month
DeepSeek V4 Flash $11.76 $19.60
MiniMax M2.5 $27.00 $54.00
DeepSeek V4 Pro $36.54 $60.90
Kimi K2.7 Code $122.50 $175.00
GLM-5.3 $159.60 $228.00
Kimi K3 $480.00 $600.00
Claude Opus 4.8 โ€” $1,000
GPT-5.5 โ€” $1,100

In the third scenario Opus costs about 85 times as much as V4 Flash, and GPT-5.5 about 94 times as much. That is well under the 180ร— gap in output prices, because this workload is mostly input and input prices sit closer together.

Is the cheap model good enough?

For most coding work, yes. In our comparison of closed and open-source LLMs we looked at two leaderboards that measure different things.

Vals Index scores how often a model completes real tasks in finance, law, coding and agent workflows. The top closed model completes 68.83%. Three models you can call on A2Agent reach about 83% to 84% of that: DeepSeek V4.1 Flash (57.86%), Kimi K3 (57.81%) and GLM-5.3 (56.97%). DeepSeek V4.1 Flash was also the cheapest model on the board per completed task, at about $0.30.

Arena's Image-to-WebDev leaderboard rates front-end code built from a design. Qwen3.8 MAX scores 1,639 there and GLM-5.3 Flash 1,588, about 95% and 92% of the leader's 1,733. On the Frontend Code Arena, Kimi K3 took first place in July with 1,679.

The closed frontier still leads, by roughly 16% to 25% on completed tasks and by single digits on user preference. Per output token, Opus and GPT-5.5 cost about 2 times (against Kimi K3) to 86 times (against GLM-5.3 Flash) as much as those models. We'd pay that premium for a tricky migration or a subtle concurrency bug, or for a change nobody will review closely. Renaming, writing tests, fixing lint errors, small features and diff review make up most coding work, and there the extra quality rarely changes the result.

A leaderboard score won't tell you how reliably a model calls tools, whether it follows your repo's conventions, or how long it likes to think. Every model on A2Agent sits behind the same API key, so it takes an afternoon to run your own tasks through two or three candidates before you commit.

Route by task instead of picking one model

The cheapest setup is usually a mix of models. On A2Agent one API key covers all of them, and switching means changing the model string.

Coding job Suggested model Why
Commit messages, docstrings, log triage, "which test failed and why" DeepSeek V4 Flash, Qwen3.8 Flash, GLM-5.3 Flash Output under $0.50 per million; 1M context
Everyday edits in Cline, Roo Code or Kilo Code MiniMax M2.5 / M2.7 Cheapest coding-tagged models at $0.60 output
Long multi-step agent runs MiniMax M3, GLM-5 Tagged for agents; M3 has a 1M window at $0.60 output
Hard debugging, careful reasoning DeepSeek V4 Pro Strong reasoning at $0.522 output, 1M context
Whole-repository prompts GLM-5.3, GLM-5.2 Coding-tagged with a 1M window
Coding agent inside a 256K budget Kimi K2.7 Code Coding-tagged, with a lower input price than the GLM-5.x coding models
Frontend and UI generation Kimi K3 Led the Frontend Code Arena at 1,679 in July 2026

Switching takes two settings

A2Agent accepts both the OpenAI and the Anthropic API formats, so most coding tools only need a base URL and a key.

OpenAI-compatible tools and SDKs (Cline, Roo Code, Kilo Code, Cursor, Codex CLI):

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_A2AGENT_KEY",
    base_url="https://api.a2agent.me/v1",
)

resp = client.chat.completions.create(
    model="minimax-m2.5",  # or deepseek-v4-flash, glm-5.3, kimi-k2.7-code ...
    messages=[{"role": "user", "content": "Write a unit test for parse_date()."}],
)
print(resp.choices[0].message.content)

Claude Code, through the Anthropic Messages API:

export ANTHROPIC_BASE_URL="https://api.a2agent.me"
export ANTHROPIC_AUTH_TOKEN="YOUR_A2AGENT_KEY"

There are step-by-step guides for Claude Code, Codex CLI, Cline, Roo Code, Kilo Code and Cursor.

Check the context window before you switch. Our models range from 200K to 1M, and if your agent packs whole repositories into the prompt, a 200K model will cut them short. Use exact model IDs, since kimi-k2.7-code and kimi-k2.5 are different models at different prices. Reasoning modes and effort controls also differ from vendor to vendor. Cap them where you can, because thinking tokens are billed as output.

Which one should you pick?

  • Under $10 a month: DeepSeek V4 Flash. It handled the solo-developer workload for $3.63.
  • Under $50 a month: MiniMax M2.5 as the main coding model, with V4 Flash for small calls and V4 Pro when something needs more careful reasoning.
  • Around $150 a month for a team: Kimi K2.7 Code or GLM-5.3 for agent runs, with the cheaper models handling everything else.
  • If one task really needs the closed frontier, send that task to Opus or GPT-5.5 and keep the rest on cheaper models.

A2Agent is pay-as-you-go with no subscription, and your balance doesn't expire. Current prices are always in the model catalog.

Sources