Models / Use cases
Cheapest AI models
Cheap only counts if the output is usable. Every model here clears a quality bar on extraction, classification, and short-form generation, and the prices shown are the exact rates your calls are metered at.
Ranked
The picks
GLM-5.3-Flash
Z.ai
The cheapest model here, and it still carries a million tokens of context, vision, and tool calling.
Context
1M tokens
Max output
131K tokens
Tools
Yes
Vision
Yes
View GLM-5.3-Flash pricing and playground →
Gemini 3.1 Flash-Lite
Google's cheapest, for classification and extraction at volume where you still want vision.
Gemini 3.5 Flash-Lite
A step up in capability at flash-lite pricing, for high-volume agentic work and translation.
DeepSeek V4.1 Flash
The reference point for price-performance among open models.
Replaces DeepSeek V4 Flash, retired by DeepSeek on 2026-09-10.
Gemini 3 Flash (preview)
PreviewFlash-tier prices with a giant context, cheap and long-context at once.
Retired
Previously recommended here
The reference point for price-performance among open models.
Use DeepSeek V4.1 Flash instead, same endpoint, same key, one line changes.
FAQ
Frequently asked questions
What does CompanyFabric add to the provider's own price?
A 5% platform fee: you pay it once when you buy credit, and metered usage is charged at the provider's published rate plus 5%. BYOK routes on your own provider key at 0% platform fee, so at sustained volume that is the cheaper path.
Is the cheapest model ever the right choice?
More often than people expect. Extraction, classification, routing and short-form generation are solved problems at this tier, and the quality gap only opens on multi-step reasoning. Route by task, not by reputation.
Why is the cheapest model not always what companyfabric/auto picks?
It usually is. Auto picks the cheapest model that can actually serve the request, so it will skip a cheaper model that lacks vision, lacks tool calling, or cannot fit the prompt, rather than failing the call.
Do these prices change?
Providers reprice regularly, and every price here carries the date we last verified it against the provider's own rate card. Where a provider has published a future change, such as Google's 2027 increase, we bill the new rate automatically on the day.
How far does a dollar go at these prices?
At GLM-5.3-Flash rates, one dollar buys more than six million input tokens. That is the reason to route volume work here and reserve the frontier tier for the calls that actually need it.
One key. Every model. Exact prices.
Route every pick on this page through one key. Free starter credits included, no subscription required.