Why every AI quote needs a monthly number next to it
The build cost of an AI feature is the cheap part. The model usage bill recurs monthly, scales with your success, and, when nobody models it upfront, becomes the reason features get switched off in month three. This page shows the token math in plain terms, the ceilings per feature type, and the levers that keep the bill sane. Vendors publish exact per-token prices; they change quarterly, so this page carries the method and current-ish magnitudes, check the pricing page of your provider before signing anything.
Tokens, translated
Models read and write in tokens (~¾ of a word). A typical customer-support exchange is 1,000–3,000 tokens in (question + instructions + retrieved documents) and 200–500 out. A summarizer on a 10-page document is ~6,000–10,000 in, ~1,000 out. Multiply by your real volume, then by the per-token price of the model tier you chose, frontier models cost roughly 10–30× the small fast models, and most features do not need frontier intelligence for every call.
Monthly ceilings by feature type (real volumes)
| Feature | Assumed volume | Model tier | Typical monthly ceiling |
|---|---|---|---|
| Support chatbot (RAG over docs) | 2,000 conversations | Mid-tier | $50–400 |
| Internal knowledge assistant | 300 queries/day | Mid-tier | $100–600 |
| Document summarizer | 1,000 docs/month | Mid/frontier mix | $50–500 |
| Content drafting assistant | 200 long drafts | Frontier | $100–800 |
| AI agent (multi-step, tool-using) | 500 runs | Frontier | $200–2,000+ |
| Embeddings/vector updates | ongoing | small | $5–50 |
These are editorial ceilings from shipped projects, not quotes, your volume and prompt discipline move them 10× in either direction. The scary rows are agents: multi-step reasoning multiplies tokens per run, and unbounded loops multiply them terrifyingly.
The levers, in order of impact
- Model-tier routing, small model for easy calls, frontier for hard ones. Routinely cuts bills 60–90% with no perceived quality loss.
- Context discipline, send the three relevant document chunks, not the whole handbook. Retrieval quality is cost control.
- Caching, identical questions get identical answers; cache them.
- Ceilings and circuit breakers, hard per-user and per-day caps, so a loop or an abuse spike is a $5 event, not a $5,000 one.
- Batching, non-urgent work (nightly summaries) on batch APIs at steep discounts.
How builds here handle it
Every AI engagement carries a cost ceiling before launch: projected monthly usage at real volumes, a hard cap, and an alert threshold, stated in the quote next to the build price (chatbots, RAG, agents, all the same rule). Post-launch, the bill is a monitored number, not a surprise. If your current AI project cannot tell you its monthly ceiling, that is the first conversation to have: one hour usually settles it.