Skip to main content
RoseleapBlog · AI

What an AI assistant actually costs to build and run

A transparent breakdown of the cost of building and operating an AI assistant in 2026, build fees, inference costs, model routing, caching, and where the money really goes.

·14 min read

“How much does an AI chatbot cost?” is the wrong question, and it is why so many businesses end up either overpaying for a build or blindsided by a running bill they never budgeted for. An AI assistant has two costs that behave completely differently: the one-time cost to build it, and the ongoing cost to run it. The build is a known number. The running cost scales with every conversation, and a badly engineered assistant can cost ten times what a well-engineered one costs for identical work.

This is the transparent breakdown: what you pay to build, what you pay per conversation, why those two numbers vary so wildly, and where the money actually goes.

The build cost

Building an assistant properly is not “connect an API and add a chat box.” A production assistant that a business can put in front of customers involves:

  • Discovery: defining what it should and should not do, gathering the content it will answer from, and agreeing what “correct” means.
  • Content preparation and retrieval: turning your documents into something the assistant can search and quote accurately.
  • The assistant itself: prompting, grounding, conversation handling, and the interface.
  • Guardrails: input filtering, output validation, topic limits, and a human escalation path.
  • Evaluation: a test suite of real questions so you can prove it works and catch regressions.
  • Deployment and observability: monitoring, logging, and the ability to see what it is actually saying to customers.

For a focused assistant embedded in an existing site or product, this typically runs ₹3 lakh to ₹8 lakh. A simpler internal-only tool can sit below that; an AI-native product with complex flows scales above it. The range is wide because the content preparation and the evaluation work vary enormously with how messy your source material is.

Beware the ₹40,000 chatbot. It exists, and it is a public model with a chat box and no grounding, no guardrails and no evaluation. It will confidently tell your customers things that are not true, and you will not find out until one of them acts on it.

The running cost, and why it varies tenfold

Every time your assistant answers, it makes one or more calls to a language model, and you pay per unit of text processed (input and output). This per-conversation cost is where engineering separates a sustainable assistant from a money pit. The levers:

Model routing

Language models come in tiers: small, fast and cheap models for simple work, and large, slower, expensive models for hard reasoning. A naive assistant sends everything to the most powerful model because it is easiest. A well-engineered one routes: classifying an enquiry or extracting a field goes to a cheap fast model; only genuinely hard reasoning goes to the expensive one. The price gap between tiers is large, so routing alone can cut the bill by most of its value.

Caching

Assistants repeat a lot of the same context on every call: your instructions, your policy text, your product data. Modern model providers let you cache that repeated context so you are not billed full price for re-sending it every single time. For an assistant with a large stable knowledge base, prompt caching can cut input costs dramatically. It is cheap to design in and awkward to retrofit, which is why it belongs in the build.

Retrieval instead of stuffing

There are two ways to give an assistant your knowledge. Stuff all of it into every request (expensive and it degrades quality), or retrieve only the few relevant passages per question (cheaper and more accurate). Retrieval-augmented generation is both the cheaper and the better-quality path, which is a rare thing in engineering.

Response length discipline

Output text usually costs more per unit than input. An assistant instructed to be concise costs less and, conveniently, serves customers better than one that writes an essay in reply to “what are your hours?”

A concrete illustration, with round numbers. Ten thousand conversations a month at ₹5 each is ₹50,000 a month. The same ten thousand conversations, with routing, caching and retrieval, at ₹0.50 each is ₹5,000 a month. Same assistant, same answers. The difference is entirely in how it was built, and it compounds every month for the life of the system.

What drives your specific number

  • Conversation volume. The obvious one. Cost scales close to linearly with usage.
  • Conversation complexity. A one-turn FAQ lookup is cheap. A multi-step troubleshooting dialogue that reasons over several documents is not.
  • Knowledge base size. More content to search means more retrieval work, though good retrieval keeps this sublinear.
  • Quality bar. Higher accuracy requirements push you toward more capable models and more evaluation, both of which cost more.
  • Latency requirements. If answers must be instant, some cost-saving techniques (like batching) are off the table.

Build in-house or hire it out?

The honest trade-off:

  • Off-the-shelf SaaS chatbot tools (monthly subscription, low setup): fast and cheap to start, limited grounding in your real content, and the per-message pricing often works out expensive at volume. Fine for a simple FAQ deflector, frustrating for anything that needs to be genuinely right.
  • Custom build: higher upfront cost, but you own it, it is grounded in your content, and it is engineered for low cost-per-call, which pays back at volume. Right when the assistant is a real part of your customer experience.
  • Pure DIY: viable only if you have engineers who understand retrieval, evaluation and cost engineering. The API is easy; making it reliable and cheap is the actual work, and it is where DIY projects usually stall.

The costs people forget to budget

  • Content maintenance. Your policies and products change. The assistant's knowledge must be kept current or it starts giving stale answers.
  • Evaluation upkeep. New question types appear; the test suite grows.
  • Monitoring. Someone has to look at what the assistant is actually saying, especially early on.
  • Model changes. Providers release new models and retire old ones. A well-built assistant treats the model as swappable; a brittle one needs rework each time.

These are the same rhythms as ordinary software upkeep, and they fit naturally into a maintenance arrangement rather than being a surprise.

How RoseLeap can help

Our AI Solutions engagements engineer for low cost-per-call from day one: model routing, caching, and retrieval rather than context-stuffing, with an evaluation suite so you can trust the answers. We use Claude as our default and route to the cheapest model that does each job well.

If you are weighing an assistant, tell us your expected volume and the questions it needs to handle on the contact page, and we will come back with a build estimate and a realistic monthly running cost, within one business day. For the wider picture of what AI is worth doing at all, start with our overview of practical AI for small businesses.

RoseLeap.

Rooted in Data · Built to Bloom

Need help with your own?

Tell us about your project. We come back with a clear, honest plan.

Start a Project