PRACTICAL AI OPERATIONS

Stop Wasting AI Tokens: Use the Cheapest Capable Model

Connor T. MacIvor·AI implementation, Santa Clarita Valley·

The strongest AI model is not automatically the right model. That sounds obvious until the usage meter starts moving and every task, from cleaning a list to reviewing a consequential release, is being handed to the same premium system.

I added a routing instruction to my own Codex workflow after watching expensive-model capacity disappear too quickly. Before a task begins, the system should recommend the least expensive available model capable of completing that task reliably. It should consider complexity, privacy, tools, context, and the cost of a quiet mistake. Then it should tell me why it chose that route.

The rule is cheapest capable, not cheapest.

Why is the cheapest model not always the least expensive choice?

A low call price can become expensive when the output is wrong, incomplete, or requires three rounds of repair. Cost is not the price of one request. Cost is the price of useful completed work.

If a smaller model extracts 400 names, normalizes the columns, and flags uncertain rows correctly, it may be the best tool for that job. If it silently drops records and a person spends two hours finding the mistake, the route failed even if the first call cost almost nothing.

The same logic works in the other direction. Paying for premium reasoning to rename files, tag records, or turn a clean table into a predictable format may not improve the result. It just consumes the capacity you may need later.

Which AI tasks should usually move down?

Bounded work is the best candidate. Bounded means the input, output, and checking rule are clear.

Smaller hosted models, approved local models, and simple scripts can all belong in this lane. A script is often better than a model when the job is counting, hashing, joining, or applying an exact rule. Intelligence is not improved by asking a language model to perform arithmetic that a deterministic tool can finish perfectly.

Which tasks should step up to premium reasoning?

Move up when the work is novel, ambiguous, difficult to verify, or expensive to get wrong.

The important variable is not how impressive the task sounds. It is the consequence of a quiet mistake. A short permission change can deserve stronger review than a 3,000-word rough draft.

What instruction can I give my AI workflow?

Use this as a starting point:

Before each task, recommend the least expensive available model capable of completing the work reliably. Consider complexity, privacy, required tools, context, and the cost of a quiet mistake. Route bounded bulk work down. Escalate consequential judgment and final review. Tell me when the route changes and why.

Adapt that instruction to the models and permissions actually available in your environment. Do not assume every ChatGPT, Claude, or Codex account can inspect billing or switch models automatically. Some interfaces expose model choice. Some can only recommend it. Product names, limits, and prices change.

How do I verify that model routing is saving money?

Start with three measurements:

  1. Did the task finish correctly?
  2. How much human correction did it require?
  3. Did moving the task down preserve premium capacity without increasing risk?

Keep a small test set for recurring work. Run the same representative inputs when you change models. Compare missing fields, invented claims, formatting errors, completion time, and correction time. That gives you evidence instead of a feeling.

For consequential work, use a second pass that is independent of the first. The reviewer should check the source, not merely agree with the draft. A cheap first pass plus a capable final review can be an excellent route. Two weak passes repeating the same mistake are not verification.

What does this look like for a Santa Clarita business?

A real-estate office may use a smaller model to classify incoming questions, extract property preferences, and prepare a draft follow-up. A person should still approve claims about a listing, financing, contracts, or representation.

A contractor may use a smaller model to sort job notes and build an estimate checklist. A qualified person still owns scope, pricing, code, and the promise made to the customer.

A local service business may use an agent to identify missed leads and prepare replies. Sending, discounting, refunding, and changing customer records can remain approval-gated.

The model route should follow the work, not ego. The goal is not to boast that every task used the biggest machine. The goal is to complete more useful work per dollar while keeping human responsibility where it belongs.

Watch the short demonstration, then write the routing rule into your own operating instructions. If you want help mapping AI into a real Santa Clarita workflow, book a working session and bring the process you want to improve.

Common questions

Should I use the strongest AI model for every task?

No. Use premium reasoning when complexity, ambiguity, or consequence requires it. Routine cleanup and extraction often do not need the most expensive model.

What does cheapest capable model mean?

It is the lowest-cost available model that meets the task's accuracy, privacy, context, tool, and reliability requirements.

What work can move to a smaller model?

Extraction, cleanup, tagging, classification, formatting, inventories, structured comparisons, and rough drafts are common candidates.

When should I use a premium AI model?

Step up for novel architecture, conflicting evidence, security-sensitive work, irreversible actions, and final review before a consequential release.

Can ChatGPT or Codex automatically switch models?

Only when the product, plan, configuration, permissions, and tools support it. Otherwise the system can recommend a model and the operator can switch manually.

Want this working in your business?

Connor builds the AI systems he writes about, here in Santa Clarita. Book a working session and bring your actual workflow.

Get on Connor's Calendar

Connor T. MacIvor · CalDRE #01238257 · Sync Brokerage, Inc. · DRE #02031490