Most people are overpaying for the wrong AI model tier.
Not every AI model labelled flagship is actually best in class. Five tiers, mapped end to end — and the one where most teams quietly overpay.

The AI model landscape splits into five distinct categories. Most people only ever use one of them — the expensive one — and then form an opinion about cost based on running every task through it.
“Flagship” is a price tier. It is not a promise. Several mid-tier and open-source models beat flagships on specific jobs, once you know which job. The skill is not picking the best model; it is knowing which category a task belongs to before you send it.
Here is the full map — category by category, with what each one is actually for.
The short version
- Flagship = the hard problems. Mid-tier = the actual work. Light = the volume.
- Specialized = the one job. Local = your data, your rules.
- Roughly 80% of daily work belongs in the mid-tier. Most overpaying happens because it is run on flagships instead.
1. Flagship models
Reserve these for the hardest reasoning. Top of the benchmark, top of the invoice — and the two facts are related in a way that is easy to forget when you are choosing a default.
- Claude Fable 5 is the smartest, but slow and expensive. Claude Opus 5 is the more honest alternative for most flagship-class work.
- Gemini 3.1 Pro owns multimodal and video.
- GPT-5.6 Soul bundles coding, browsing and image generation in one place.
- Grok 4.5 is cost-effective with real-time data.
- Kimi K3 is open source — near-flagship capability, and self-hostable.
The test for whether a task belongs here is not how important it feels. It is whether the reasoning is genuinely hard: multi-step, ambiguous, expensive to get wrong. Drafting an email is important and is not hard reasoning.
2. Mid-tier workhorses
Run roughly 80 percent of your daily work here. This tier has the best balance of speed, cost and capability, and it is where most people should live by default.
- Claude Sonnet 5 handles casual chat and coding.
- GPT-5.6 Terra is the cheaper flagship spin-off.
- Gemini 3.6 Flash is surprisingly strong at multimodal for the price.
- Open source is closing fast here: GLM 5.2, MiniMax M3, and DeepSeek V4 Pro for dirt-cheap math and reasoning.
This is the tier where the overpaying happens. If your default is a flagship, every routine draft, rewrite and code change is billed at reasoning rates for work that did not require reasoning. Moving the default down one tier and promoting individual tasks upward when they need it is almost always cheaper and rarely worse.
3. Light models
Deploy these for automations, sub-agents and bulk processing. Tiny, fast, nearly free — and completely adequate for work where the answer is largely mechanical.
- Haiku 4.5 is built for sub-agent work.
- GPT-5.6 Luna handles high-volume chat.
- Gemini 3.5 Flash-Lite is the fastest measured model available today.
- DeepSeek V4 Flash and MiMo V2.5 dominate high-volume coding jobs.
The category that benefits most from this tier is agent architectures. A sub-agent doing a file rename, a classification or a lookup inherits whatever model the parent was using unless told otherwise — which quietly bills chores at heavy-reasoning rates. Pointing that work at a light model is usually the single biggest cost reduction available.
4. Specialized models
Pick the job first, then pick the model built only for that job. General models will do most of these adequately; specialists do them properly.
- Code: Kimi K2.7 Code, Qwen 3 Coder.
- Image: GPT Image 2, Nano Banana, Meta Muse.
- Video: Veo 3.1, Runway Gen-4.5.
- Audio: Suno and Udio for music, ElevenLabs for voice.
- Enterprise retrieval: NVIDIA Nemotron 3 and Cohere for search pipelines.
The ordering matters. People tend to pick a model and then ask what it can do. In this category the productive direction is the reverse: define the single job, then find the model whose entire existence is that job.
5. Local and self-hosted
Run on your own hardware when privacy, latency or cost makes the call — and note that any one of those three is sufficient on its own.
- Qwen 3 635B fits on a Mac Studio.
- Hermes 4 is tuned for agent work.
- GPT-OSS from OpenAI is free to run.
- Step 3.7 Flash flies on Apple silicon if you have the RAM.
Local is the only category where the deciding factor is often not capability at all. If the data cannot leave the building, the best cloud model in the world is not an option, and the question becomes which local model is good enough.
The part most people get wrong
“Flagship” is a price tier, not a promise.
Several mid-tier and open-source models beat flagships on specific jobs — once you know which job. Gemini 3.5 Flash-Lite is the fastest measured model available and sits in the light tier. Kimi K3 is open source and near-flagship. DeepSeek V4 Pro does math and reasoning at a fraction of flagship cost.
The mental shift is from “which model is best” to “which category is this task.” The first question has no answer. The second has an obvious one almost every time, and answering it is where the money is.
The master formula
- Flagship = the hard problems.
- Mid-tier = the actual work.
- Light = the volume.
- Specialized = the one job.
- Local = your data, your rules.
Frequently asked questions
How do you choose which AI model to use?
Classify the task before choosing the model. Hard, ambiguous, expensive-to-get-wrong reasoning goes to a flagship. Ordinary daily work — roughly 80% of it — goes to the mid-tier. Automations, sub-agents and bulk processing go to light models. Single-purpose jobs like image, video or audio go to specialists. Anything constrained by privacy, latency or cost goes local.
Are flagship models always the best choice?
No. Flagship is a price tier, not a promise. Several mid-tier and open-source models beat flagships on specific jobs — Gemini 3.5 Flash-Lite is the fastest measured model and sits in the light tier, and Kimi K3 is open source at near-flagship capability.
Where do most people overpay?
In the mid-tier, by not using it. If your default is a flagship, every routine draft and code change bills at reasoning rates. The second common leak is sub-agents inheriting the parent’s model, so a one-line chore runs on a heavy-reasoning model.
What are light models actually good for?
Automations, sub-agents and bulk processing — work that is largely mechanical. Haiku 4.5 is built for sub-agent work, GPT-5.6 Luna handles high-volume chat, and DeepSeek V4 Flash and MiMo V2.5 dominate high-volume coding jobs.
When is it worth running a model locally?
When privacy, latency or cost makes the call — any one of the three is enough on its own. Qwen 3 635B fits on a Mac Studio, GPT-OSS is free to run, Hermes 4 is tuned for agent work, and Step 3.7 Flash is fast on Apple silicon given enough RAM.
Build it yourself
Everything written about here gets built in the open — the whole application, on camera, including the parts that did not work first time.
Keep reading
AI agent skillsAI Agent Skills: 9 Free Ones and the Step Most Skip
AI agent skills install from one pasted link — that part is easy. The step that makes them genuinely useful is the one almost nobody runs.
Read article→
Jev decision modelJev Decision Model: The AI That Can't Write, Only Decides
The Jev decision model can't write a word. Vercel and LangChain added it anyway. Why a "System One" classifier is 200x faster than the LLM in your agent.
Read article→
open source SaaS alternativesOpen Source SaaS Alternatives: 5 Repos Worth Selling
Semrush charges $250. This costs ten.
Read article→