AI Strategy

AI Just Got Cheaper. Here Is What to Do

Model prices fell sharply between late 2022 and late 2024. The instinct is to wait for cheaper. That instinct is backwards.

Michael Pavlovskyi Michael Pavlovskyi · · 6 min read
Architecture diagram of a basic AI agent by Anthropic, showing the LLM interacting with external retrieval, tools, and memory modules.
Source: Anthropic, "Building Effective Agents" Research Guide

Key Takeaways

  • The price of running an AI model fell by more than 280 times in under two years. That number is real, and the trend is likely to continue.
  • The API bill was never the real cost of an AI project. The cost is mapping the workflow, building it, and getting the team to use it.
  • Cheaper models do not make that work cheaper. Waiting a year just means starting the same project a year later.
  • The right move now is to run AI on the tasks you skipped when it was expensive, then design one real workflow well.

The price of running an AI model has dropped hard. By one widely cited measure from Stanford's 2025 AI Index, the cost of a query at the level of GPT-3.5 fell from about twenty dollars per million tokens in late 2022 to roughly seven cents by late 2024. That is not a discount. That is a different category of cost.

At Bace Agency, we have watched that shift change what is actually possible. Automation that was too expensive to justify for a North Shore agency two years ago is now cheap enough to build and run every day. Heavy document processing, always-on intake, a review of every file before it moves: work that used to price itself out now pays for itself.

So when an owner asks whether to wait for prices to fall further, my answer is no. The model price was never the part that cost you anything.

280x
Drop in the cost to run a GPT-3.5-level query between late 2022 and late 2024, per Stanford's 2025 AI Index.
$0.07
What a million tokens at that level cost by late 2024, down from about $20, per the same report.
30%
Annual decline in the hardware cost behind these models, per the same report. The trend has a real engine under it.

Why Prices Keep Falling

This is not a sale that ends Friday. Prices are falling because the hardware gets cheaper every year, the models get smaller for the same quality, and the vendors are in a fight for your business. Anthropic, for example, has held or cut prices on its Claude Opus and Sonnet lines across several flagship launches, including a steep cut to the Opus tier at the 4.5 release. That is unusual. Most products raise prices on a new top model.

I wrote a fuller explanation of the mechanics in why AI keeps getting cheaper. The short version: this trend is structural, not promotional. So waiting does not catch a bottom. There is no bottom to catch. There is only later.

Why Waiting Does Not Pay

Here is the part owners miss. On a real project, the model bill is the smallest line item. I broke down the full picture in the real cost of an AI project, but it comes down to three things, and none of them are the API.

First, figuring out the workflow. Sitting with your team, watching how intake actually happens, where the renewal list really lives, who re-keys the same client data into three systems. Second, building it and connecting it to the tools you already run. Third, getting people to trust it and use it Monday morning. That last one is the hardest, and a cheaper model does nothing for it.

None of that gets cheaper because tokens got cheaper. An insurance agency that waits a year to automate its intake does not get a discount on the project. It gets a year of staff re-keying data by hand, then starts the same build twelve months behind. One North Shore P&C agency we worked with cut about 22 hours a week of data entry, and the error rate fell from 6.1% to 0.3%. The value was the hours, not the token price.

"The model price is the cheap part. The expensive part is the year you spend doing the work by hand while you wait for a discount that was never the point."

Michael Pavlovskyi, Bace Agency

What to Do With Cheaper AI Now

Cheap tokens do change one real thing: tasks that were too expensive to run on AI a year ago are worth running today.

When a model cost twenty dollars per million tokens, you saved it for high-value work. At seven cents, you can point it at the boring, high-volume tasks you used to skip. Read every inbound email and sort it. Summarize every long PDF before a partner opens it. Draft the first version of every renewal notice. Check every invoice against last month's. The math that said "not worth it" now says "run it."

So do two things. Turn AI loose on the cheap, repetitive tasks where being wrong is low-stakes and a human still reviews the output. And separately, pick one real workflow that costs your firm hours every week, and build that one properly. The first move is almost free to try. The second is where the actual return is.

SAMPLE CLAUDE PROMPT

"Here are the repetitive tasks my team does every week: [list them]. For each one, tell me which are good candidates to hand to AI now, where a person still reviews the output and a mistake is low-stakes. Rank them by hours saved per week and by how easy they are to start. Flag any that touch sensitive client data and should wait for a more careful setup."

How to Get Started

1

List the tasks you skipped because AI felt too expensive

Write down the repetitive, high-volume work your team does by hand: sorting email, summarizing documents, drafting routine notices, basic data checks. A year ago, running a model on all of it felt wasteful. At today's prices it is cheap. This list is where the easy wins are.

2

Pick one workflow that costs real hours and build it well

Choose a single workflow that eats a few hours every week: intake, renewals, document routing. Map how it actually runs today, then build the AI version around that, not around what a vendor assumes you do. One workflow done right beats ten half-finished experiments.

3

Keep a person on anything client-facing

Cheaper does not mean unsupervised. For anything that touches a client, a contract, or money, a person reviews the output before it goes out. The model drafts and sorts. The person decides. That rule is what makes the whole thing safe to run at volume.

What This Does Not Replace

One honest limit. On anything that touches a client, a contract, or money, a person still reviews the output before it goes out. We build that review step in from day one. It is what makes the system safe to run at volume.

Diagram of the Generator-Evaluator workflow pattern for large language models by Anthropic.
The Generator-Evaluator design pattern: while an automated filter corrects errors in drafts, the final decision on complex and sensitive tasks always remains with a human. Source: Anthropic, "Building Effective Agents" Research Guide

Past that, the point is simple. The model is cheap. The architecture is where the value lives. You are not buying tokens. You are buying back the hours your staff loses every week. In two of our North Shore engagements that meant about 22 hours a week reclaimed for a P&C agency and about 32 hours a week for a family office (case studies). That is what we build at Bace Agency.

So the real question is not whether AI will get cheaper. It will. The question is how many hours your team loses every week that you wait. The firms that move first this year will spend the next one running leaner than the ones still reading reports. Instead of reading another report, look at your own numbers with me. Book a 30-minute operational review. I will personally walk through your workflows and tell you exactly which tasks can come off your plate first. In person on the North Shore or by video. Start here, or take the AI Readiness Quiz first.

Frequently Asked Questions

If AI keeps getting cheaper, why not just wait? +

Because the model price is the smallest cost in an AI project. The real cost is mapping your workflow, building the system, and getting your team to use it, and none of that gets cheaper when tokens do. Waiting a year does not buy you a discount on the project. It just means you start the same project a year later, after twelve more months of doing the work by hand.

How much have AI model prices actually fallen? +

By a large multiple. Stanford's 2025 AI Index found that the cost of a query at the level of GPT-3.5 dropped from about twenty dollars per million tokens in late 2022 to roughly seven cents by late 2024, a reduction of more than 280 times. Frontier model prices have also fallen sharply across the major vendors over the same period.

What can I do now that was too expensive a year ago? +

High-volume, low-stakes tasks. Sorting and triaging inbound email, summarizing long documents before a partner reads them, drafting first versions of routine notices, checking invoices against prior months. When models were expensive, you saved them for high-value work. At today's prices, pointing AI at the boring repetitive tasks is cheap enough to be worth it, as long as a person still reviews the output.

Does cheaper AI mean I can skip human review? +

No. A cheap model is still wrong sometimes, and on anything client-facing, legal, or financial, a wrong answer is costly regardless of token price. The safer way to think about it is that cheaper AI lets you run more tasks, not that it removes the person. For anything that touches a client, a contract, or money, a person reviews the output before it goes out.

Related Articles

About the author

Michael Pavlovskyi

Written by

Michael Pavlovskyi

Founder, Bace Agency

Michael builds custom Claude and GPT workflows for insurance agencies, law firms, and PE firms on Chicago's North Shore. Speaker at Northwestern and Lake Forest College on practical AI adoption for professional services.

Connect on LinkedIn

Want to see how AI fits in your firm?

Book a free 30-minute AI audit. No obligation, no pitch deck.

Book a Free AI Audit →