Models / Use cases

Best reasoning models, where the frontier tier earns its price

Reasoning-heavy tasks reward the frontier tier: extended thinking, reliable self-correction, and long chains of tool calls. Here is where each frontier model earns its price, and where a cheaper model is enough.

Ranked

The picks

FAQ

Frequently asked questions

Is a reasoning model worth the price for everyday tasks?

Usually not. Route everyday extraction and summarization to a fast tier and reserve reasoning models for planning and verification steps. One balance across both makes the split trivial.

Do reasoning models bill for thinking tokens?

Yes, and every provider on this page counts them as output tokens. That is why a reasoning model's real cost per task can be several times its headline output rate, and why the cheaper tiers win on volume work.

Does a bigger context window make a model reason better?

No. Context is capacity, not capability. Gemini 3.1 Pro's million-token window is the reason to use it over source material you cannot summarise, not evidence it thinks harder than a model with 200K.

How do I compare reasoning quality without trusting benchmarks?

Run the same prompt against two or three of these in the chat and read the outputs side by side. Benchmarks measure a distribution; your task is one point in it, and the comparison takes a minute.

Can I cap what a reasoning model spends on one request?

Yes. Set max_tokens, and the gateway refuses any request whose projected cost exceeds your balance before it reaches the provider, so a runaway thinking loop cannot quietly drain an account.

One key. Every model. Exact prices.

Route every pick on this page through one key. Free starter credits included, no subscription required.