Every founder asking this expects the answer to be about model pricing. It almost never is. The API bill for a typical feature lands somewhere between a few dollars and a few hundred a month — while the project around it costs tens of thousands. Understanding why is the difference between an AI budget that holds and one that doubles.
What’s Covered
The Three Cost Layers
Most AI budgets fail because they price one layer and discover the other two in month three.
| Layer | What it is | Typical size |
|---|---|---|
| Build | Engineering, data preparation, prompt and retrieval design, evaluation, UI, review workflow | $5,000–40,000 for a focused feature |
| Run | Model API calls, vector storage, hosting, logging | Often under $100/month at small scale |
| Maintain | Monitoring quality, handling model deprecations, prompt and eval updates, new edge cases | $1,000–5,000/month on a live feature |
Read those rows again. The layer everyone researches first — model pricing — is the smallest. The layer nobody budgets for — maintenance — costs more every year than the build did, if the feature matters.
Build Cost by Feature Type
Published 2026 ranges, which match what we see on real briefs:
| What you are building | Build cost | Timeline |
|---|---|---|
| Single LLM feature on clean data — summaries, classification, extraction, semantic search | $5,000–40,000 | 3–8 weeks |
| Customer-facing assistant over your own content (RAG) | $50,000–150,000 | 8–16 weeks |
| AI-first MVP with its own pipeline and evaluation | $50,000–100,000 | 3–5 months |
| Multi-agent system acting across tools | $250,000–400,000+ | 6 months+ |
Those bands are wide for an honest reason. A summariser over tidy, structured records is a three-week job. The same summariser over a decade of inconsistent PDFs, scanned documents and half-filled database fields is a three-month job, and almost all of the difference is data work, not AI work.
A convincing prototype can be built in days. That speed is genuinely new, and it is also the single biggest cause of unrealistic AI budgets — because the remaining 80% of the work (permissions, edge cases, latency, failure handling, evaluation) is invisible in the demo.
The Token Maths, Worked
Let us price a real feature rather than hand-wave. Say you summarise every support ticket: 5,000 tickets a month, each roughly 2,000 input tokens (the ticket thread) and 500 output tokens (the summary). That is 10 million input and 2.5 million output tokens a month.
At 2026 list prices:
| Model | Price per 1M in / out | Your monthly bill |
|---|---|---|
| GPT-4o | $2.50 / $10.00 | $25 + $25 = $50 |
| Claude Sonnet 4 | $3.00 / $15.00 | $30 + $37.50 = $67.50 |
| Gemini 1.5 Flash | $0.075 / $0.30 | $0.75 + $0.75 = $1.50 |
Fifty dollars a month. Against a build cost of twenty or thirty thousand. This is the number that reframes the whole conversation: choosing a cheaper model saves you almost nothing, while choosing a worse model costs you review time on every single output. Pick for quality on your task, measure it, and stop optimising a line item that rounds to zero.
Two caveats that do change the maths. First, scale: at 500,000 tickets instead of 5,000, that $50 becomes $5,000 and model choice starts to matter. Second, agents: a system that loops, calls tools and re-reads context can burn ten to fifty times the tokens of a single call, which is why agentic systems deserve their own budget line.
Build or Buy
Off-the-shelf AI tools run roughly $20–100 per user per month with near-zero upfront cost. The honest rule is simpler than most comparison posts make it:
| Buy when… | Build when… |
|---|---|
| The capability is generic — transcription, meeting notes, generic chat | It touches your proprietary data or workflow |
| It is internal and nobody chooses you because of it | It is part of why customers pick your product |
| You need it this quarter | Per-seat pricing at your headcount exceeds engineering cost |
| Requirements are still changing weekly | Data residency or compliance rules out the vendor |
The predictable bias: teams underestimate build cost and overestimate buy cost, which pushes them to build things they should have rented. If the feature is not in your product’s value story, rent it and spend the engineering on what is.
Where Budgets Actually Blow Up
Four line items, in the order they surprise people:
- Data plumbing. Getting content into a clean, permissioned, refreshable state is routinely half the project. Nobody quotes it because nobody looks at the data before quoting. Ask for a data review before a fixed price.
- Evaluation. Without a test set of real cases with known-correct answers, you cannot tell whether any change helped. Building one takes days of human effort and is the most commonly skipped step — which is why so many teams ship a change and argue about whether it got better.
- Guardrails and review. Anything irreversible needs a human gate, an audit trail and an undo path. That is product and engineering work, not prompt work.
- Maintenance. Models get deprecated on vendor timelines, not yours. Prompts drift as your data changes. Accuracy degrades quietly, without an error in any log. Budget $1,000–5,000 a month for a feature customers depend on.
None of this is exotic. It is the same discipline any production feature needs — it just gets skipped more often with AI because the prototype looked finished.
A Sequence That Does Not Overrun
| Step | What happens | Spend |
|---|---|---|
| 1. Pick one workflow | High volume, cheap to check, measurable definition of done | — |
| 2. Review the data first | Look at the actual records before anyone quotes a price | 1–3 days |
| 3. Build the test set | 50–200 real cases with correct answers, written by someone who knows the domain | 2–5 days |
| 4. Prototype and measure | Score against the test set, not against vibes | 1–2 weeks |
| 5. Ship behind a human gate | Suggestions a person approves, before anything automatic | 2–4 weeks |
| 6. Remove the gate selectively | Only where measured accuracy earns it | ongoing |
A test set turns an endless argument into a number. It is also the thing that lets you switch models later in an afternoon instead of a month — useful, given how often pricing and quality shift.
What We Charge
Since this article is full of other people’s numbers, here are ours. We build AI features with dedicated AI developers from $25/hour, and general web work starts at $15/hour with rates published per market on our hire developers pages. A focused feature of the kind in section 02 is typically 3–6 weeks of one or two developers.
We also ship an AI assistant product of our own, which is where most of the opinions above come from — including the expensive ones we learned by skipping step 3. If you are weighing whether to hire for this at all, we wrote separately about what AI actually changed about hiring developers, and our vetting checklist covers how to test any vendor’s AI claims before signing.
Frequently Asked Questions
How much does it cost to add an AI feature to an existing product?
A focused feature on an existing model API runs roughly $5,000–40,000 over 3–8 weeks. A customer-facing assistant over your own content is $50,000–150,000. A full AI MVP is $50,000–100,000, and multi-agent systems reach $250,000–400,000 or more. The spread is mostly about the state of your data, not the model.
How much do the API calls cost?
Less than you think. Summarising 5,000 support tickets a month costs about $50 on GPT-4o, $67.50 on Claude Sonnet 4, or $1.50 on Gemini 1.5 Flash at 2026 list prices. Inference becomes a real line item only at high volume or with agent loops that re-read context dozens of times.
Should we build or buy?
Buy if the capability is generic and not part of why customers choose you — expect $20–100 per user per month. Build if it touches proprietary data, forms part of your product’s value, or if per-seat pricing at your scale exceeds engineering cost. Most teams underestimate build and overestimate buy, and end up building the wrong things.
What are the hidden costs?
Data plumbing (often half the project), evaluation (the most-skipped step), guardrails and human review for anything irreversible, and ongoing maintenance at $1,000–5,000 a month. Models get deprecated, prompts drift, accuracy degrades without throwing a single error.
How long does it take?
3–8 weeks for a focused feature including evaluation. The prototype is days — which is exactly what sets expectations wrongly. Budget at least as much time after the demo works as before it.
Is it worth doing at all?
Where a wrong answer is cheap to catch and the task is expensive today, yes. Where output must be exact and unverified, or where the reason is that a competitor announced something, no. Reported production failure rates for agentic systems run 70–95%, and the top cause is nobody writing down what “finished” means.
The Short Version
If you read nothing else- Focused AI feature: $5,000–40,000 build, 3–8 weeks.
- API cost for 5,000 summaries a month: $1.50 to $67.50. Not your problem.
- Maintenance is $1,000–5,000 a month and nobody budgets it.
- Data plumbing is routinely half the project — review the data before quoting.
- No test set means no way to know whether a change helped.
- Buy generic capability; build what customers choose you for.
- Ship behind a human gate first, remove it only where measurement earns it.
- The demo takes days; production takes the other 80% of the time.
Want a real number for your feature?
Tell us the workflow and show us the data. We will come back with a scope, a timeline and a fixed range — or tell you to buy something off the shelf instead, if that is the honest answer.
Hire an AI developer Get it scoped