What Actually Drives the Cost of an AI Application
Once a model is available to your application, what actually moves the bill? Most teams answer "the model price." That is only one piece. In Module 1 of the course I boil it down to one formula, and once you see it, optimization stops feeling like guesswork.
The formula
LLM application cost = Work per request × Volume × Attempt multiplier
Three factors. Each one has its own levers.
1. Work per request
How much does a single call cost? Two things decide it:
- Tokens. How many tokens the request needs, including any context you attach, such as documents, conversation history or a long system prompt.
- Model tier. Whether the call goes to a high, medium or low tier model.
2. Volume
How many calls happen? Count both kinds of traffic:
- Interactive: the number of users, times the requests each user makes.
- Scheduled: the data jobs your application runs, and how often it runs them.
3. Attempt multiplier
This is the factor most teams forget. Work per request and volume both get multiplied by the number of tries it takes to get a usable answer:
- Failures and retries, from latency problems or capacity limits.
- Agentic loops. An agent runs the whole request again and again until it reaches the right answer.
- Rework. The workflow checks its own output and redoes the step.
An analogy: running a taxi fleet
Think of the fare on a single ride. It depends on the distance (tokens) and the class of car (model tier). Your monthly bill depends on how many rides your riders take (volume). Then there are the wrong turns, the cancelled pickups and the re-routes. Each of those is a ride you paid for that never got anyone where they were going. That is the attempt multiplier.
A quick worked example
This is an illustration, not a benchmark, and it counts input tokens only. Say a request is about 1,000 tokens on a low-tier model at $1.00 per million input tokens. That is about $0.001 per request.
- 10,000 users × 5 requests a day = 50,000 requests, about $50 a day.
- If the workflow averages 3 attempts per request, the same traffic costs about $150 a day.
Nothing about the model or the users changed. Only the multiplier did.
Why this helps
Each factor points to a family of optimizations: smaller prompts and cheaper models for work per request, caching and batching for volume, and tighter retries and loop limits for the multiplier. Knowing which factor is driving your bill tells you where to look first. We cover the techniques in the modules that follow.
Next: try it on your own workload in the Module 1 token cost calculator lab.