Pricing

Pay per token.

Top up, create a key, and go.

ModelInput$ per 1M tokensOutput$ per 1M tokensCache read$ per 1M tokens
DeepSeek V4 FlashFeatured

Our fastest model for agentic and coding ⚡ A 1M context window keeps long agentic runs in one pass. The default for new chats and CLI setups.

$0.14$0.28$0.028
Kimi K3

The first open 3T-class model, neck-and-neck with the closed frontier on agentic coding. 1M context, native vision.

$3.00$15.00$0.30
DeepSeek V4 ProDeprecated · discontinued Sep 14, 2026

The coding flagship: the official 0813 release of DeepSeek V4 Pro, with a 1M context window and high-effort reasoning by default. Replaced by GLM 5.3.

$1.32$3.96$0.044
GLM 5.3

Z.ai's flagship coding model: the GLM 5.2 successor on the same base, with large gains on complex, long-horizon tasks and a 1M context window. Always thinks (dial low to max).

$1.40$4.40$0.26
GLM 5.3 Flash

Z.ai's fast coding model: a 320B MoE with 18B active parameters, the first natively multimodal GLM-5 series release (images and video), on a 1M context window. Always thinks (dial low to max).

Umans Coder is an alias for this model today: same rates, routed automatically.

$0.15$0.50$0.03
Umans Flash

The light, fast complement for the roles around the main coder: context gathering, scout subagents, summaries, quick edits.

$0.15$1.00$0.05

USD per 1M tokens. Top up the wallet, create a key, and every request debits exactly what the model used.

Top up & start

Labs experiments are free for seat holders while they run, and founding users have priority on seats. Watch the status page for openings and retirements.