Pricing
Pay per token.
Top up, create a key, and go.
| Model | Input$ per 1M tokens | Output$ per 1M tokens | Cache read$ per 1M tokens |
|---|---|---|---|
DeepSeek V4 FlashFeatured Our fastest model for agentic and coding ⚡ A 1M context window keeps long agentic runs in one pass. The default for new chats and CLI setups. | $0.14 | $0.28 | $0.028 |
Kimi K3 The first open 3T-class model, neck-and-neck with the closed frontier on agentic coding. 1M context, native vision. | $3.00 | $15.00 | $0.30 |
DeepSeek V4 ProDeprecated · discontinued Sep 14, 2026 The coding flagship: the official 0813 release of DeepSeek V4 Pro, with a 1M context window and high-effort reasoning by default. Replaced by GLM 5.3. | $1.32 | $3.96 | $0.044 |
GLM 5.3 Z.ai's flagship coding model: the GLM 5.2 successor on the same base, with large gains on complex, long-horizon tasks and a 1M context window. Always thinks (dial low to max). | $1.40 | $4.40 | $0.26 |
GLM 5.3 Flash Z.ai's fast coding model: a 320B MoE with 18B active parameters, the first natively multimodal GLM-5 series release (images and video), on a 1M context window. Always thinks (dial low to max). Umans Coder is an alias for this model today: same rates, routed automatically. | $0.15 | $0.50 | $0.03 |
Umans Flash The light, fast complement for the roles around the main coder: context gathering, scout subagents, summaries, quick edits. | $0.15 | $1.00 | $0.05 |
USD per 1M tokens. Top up the wallet, create a key, and every request debits exactly what the model used.
Top up & startLabs experiments are free for seat holders while they run, and founding users have priority on seats. Watch the status page for openings and retirements.