AI & Automation

What an AI Chatbot Really Costs to Build — and to Run

Three very different products are sold under the same word. A breakdown of where build cost actually goes, the recurring costs most proposals footnote, and five questions that make two quotes comparable.

Webworx Asia Editorial TeamSeptember 19, 2026 9 min read
What an AI Chatbot Really Costs to Build — and to Run

Chatbot proposals are unusually hard to compare because vendors quote different things under the same word. One quote is a scripted flow on a website widget. Another is a retrieval system over your documentation with human handoff and an audit trail. Both are called "an AI chatbot", and the second costs several times the first for good reasons.

This is how the cost actually breaks down, and which parts of it recur every month after launch — the part most proposals leave to a footnote.

Three things sold as the same product

Scripted flows

A decision tree with buttons. Predictable, cheap, and genuinely useful for a narrow set of tasks: order status, opening hours, routing to the right team. It is not AI and does not need to be. If your top ten queries are all transactional lookups, this is the right build and anything more is overspending.

Retrieval-grounded assistants

A language model answering from your own content — policies, product documentation, past tickets — rather than from its training data. This is the one most enterprise buyers actually want, because it handles questions nobody scripted. The cost is in the content pipeline, not the model.

Agentic assistants

The assistant does not just answer, it acts: creates the ticket, checks the account, issues the refund, books the slot. Every action needs an integration, an authorisation model, and a rollback story. This is application engineering with a conversational front end, and it should be scoped as such.

Where the build cost really goes

On a grounded assistant, the model integration is rarely the largest line. In our experience the weight sits in four places:

  • Content preparation. Your documentation is inconsistent, duplicated, and partly out of date — every organisation's is. Getting it into a state where retrieval returns the right passage is the single most underestimated task in these projects, and it is the difference between an assistant that is trusted and one that is quietly switched off.
  • Guardrails and evaluation. You need a test set of real questions with known-good answers, run on every change. Without it you have no way to tell whether a prompt adjustment improved things or broke them, and you are shipping on vibes.
  • Handoff to humans. Knowing when to stop trying and pass to an agent — with the conversation history intact — is what separates an assistant customers tolerate from one they resent. It also requires integration with whatever desk your team actually uses.
  • Logging and review. Someone has to read what the assistant said last week. Build the review surface into the first release; retrofitting it means the first months of real usage data are lost.

The recurring costs nobody quotes

Ask any vendor to separate build from run. The run side includes:

  • Inference. Per-conversation cost varies enormously with how much context you send. A system that stuffs a large document set into every request costs many times one that retrieves tightly. Ask for a modelled cost per thousand conversations at your expected volume, not a per-token rate.
  • Content freshness. When a policy changes, something has to re-index. Whether that is automatic or a manual job someone forgets is a design decision with a monthly cost attached.
  • Evaluation runs and drift. Model providers update models. Your outputs change even when your code does not. The test set is what tells you.
  • Human review time. Budget real hours, especially in the first quarter.

Regional considerations in the Gulf and Southeast Asia

For deployments in Qatar, the UAE, Saudi Arabia, Malaysia and Singapore, three factors usually shape the architecture more than the model choice does:

  • Data residency. Public-sector and financial buyers frequently require that conversation data and any indexed content stay within a named jurisdiction. This determines which providers and regions are eligible and should be confirmed before architecture, not after.
  • Arabic and mixed-language handling. Arabic performance varies noticeably between models, and real users code-switch mid-sentence. Test with actual transcripts from your own channels rather than clean sample questions.
  • Channel mix. In much of the region the conversation starts in a messaging app, not on your website. That changes session handling, identity, and what "handoff to an agent" means operationally.

How to compare two quotes fairly

  1. Ask both vendors for the same twenty questions — drawn from your real ticket history — and the answers their proposed system would give.
  2. Ask what happens when the assistant does not know. "It says it does not know and offers a human" is the correct answer; confident invention is the failure mode that costs you trust.
  3. Ask who owns the prompts, evaluation set and indexed content. If leaving the vendor means rebuilding all three, you are renting, not buying.
  4. Ask for run cost at three volumes: expected, double, and ten times. The shape of that curve tells you more about the architecture than the diagram will.
  5. Ask what is measured after launch. Containment rate, handoff rate and customer satisfaction on assisted conversations are real metrics. "Number of conversations" is not.

A pragmatic first release

Start with the twenty highest-volume questions, grounded in content you have already cleaned, with an obvious path to a human and full logging from day one. Run it on an internal audience for two weeks. You will learn more from those transcripts than from any amount of upfront design, and you will have an evaluation set built from real usage before you expose it to customers.

See how we approach AI chatbot development and LLM integration, or read the red flags worth checking before you sign with any vendor.