The ceiling conversation happens before the build
Every AI feature has a monthly running cost, and the projects that end badly are the ones where that number was discovered after launch. The discipline here: the ceiling is set before the build, a number the business is comfortable spending per month, chosen with eyes open, with the feature designed to live inside it. This is the budgeting companion to the LLM API costs primer (which explains the token math) and the privacy rules that govern where the money goes.
The cost model, in one formula
Monthly cost ≈ tokens per interaction × interactions per month × price per token (varies by model tier), plus fixed costs: the provider tier, vector storage for RAG, monitoring. A support bot answering 1,000 questions with a mid-tier model costs a different universe than an agent making 10,000 autonomous decisions with the frontier model, the range across "AI features" spans roughly $20/month to five figures, which is exactly why the ceiling conversation is its own step, not a footnote in the build quote.
The levers, in order of impact
- Model routing: the frontier model for hard cases, a smaller model for easy ones, the router pattern routinely cuts spend 60–80% with imperceptible quality loss on easy traffic.
- Caching: identical or similar questions answered from cache, support bots especially (the same questions arrive weekly).
- Prompt efficiency: shorter prompts, trimmed context, the token bill is bidirectional (input tokens cost too, and bloated system prompts tax every call).
- Volume caps: rate limits per user and daily ceilings, the cost can never run away from the budget.
The review habit that keeps it honest
Token pricing and model capabilities shift quarterly, a ceiling set and forgotten drifts. The monthly review: spend by feature, cost per interaction trending, quality spot-checks on the cheap-model traffic. The eval gates include cost as a metric, not just correctness. And the honest ceiling conversation includes the flip side: an AI feature that cannot be run within its ceiling is a feature that needs redesign, smaller scope, different model class, or a human in the loop where the machine was wasteful. The integration builds ship with ceilings designed in; the brief states yours and the build stays inside it.