Insight
How much does it cost to build an AI agent in 2026?
Published ranges run from $10,000 to over $400,000. Here is what moves an AI agent up that range, and what it costs to run once it is live.
Ask five agencies what an AI agent costs and you’ll get five ranges, some of them ten times apart. They are rarely wrong. They are pricing different things under the same name.
We don’t publish our own prices, because the honest answer depends on the things below. What we can do is show you what the published figures say, where they come from, and which parts of your idea move the number.
What do published estimates say?
These are the ranges firms published in 2026. They disagree with each other, so read the source column as carefully as the numbers.
| Type of agent | Published build range | Source |
|---|---|---|
| Answers questions from your documents (retrieval, Q&A) | $10,000 to $30,000 | Thinklytics, Sept 2026 (a compilation of other firms’ published ranges) |
| Focused proof of concept or limited internal agent | $25,000 to $50,000 | Neoteric, Aug 2026 (from their own project work) |
| Task agent built on a language model | $50,000 to $120,000+ | Azilen, Aug 2026 (built up from component costs) |
| Knowledge agent over your documents (RAG) | $80,000 to $180,000+ | Azilen, Aug 2026 |
| Production agent with memory, permissions and monitoring | $100,000+ | Neoteric, Aug 2026 |
| Multi-agent platform | $150,000 to over $400,000 | Azilen and Thinklytics, 2026 |
For wider context, Clutch’s AI pricing guide (updated 21 September 2026) puts the average AI development project on its marketplace at about $120,600 over a typical 10 months, based on client reviews. That covers AI work in general, not agents specifically.
The sources: Neoteric, Azilen and Thinklytics. All three sell AI development, so treat them as market context, not independent research.
What moves an agent up the range?
The label on the agent matters less than what it has to do. Three questions decide most of the cost.
How many systems does it act in? An agent that only answers questions is one piece of work. Each system it has to read from or write to, such as your CRM, a booking tool or a payment provider, is its own integration, with its own permissions, error cases and testing.
Does it answer from your own documents? Answering from approved guidance means preparing that guidance, keeping it current, and showing where each answer came from. That retrieval layer is often the biggest single part of a knowledge agent.
How much does a person check? An agent that drafts and waits for approval is cheaper to make safe than one that acts on its own. The more it does unattended, the more testing, logging and monitoring it needs before launch.
What does an AI agent cost to run?
Building it is one cost. Running it is another, and it arrives every month.
Language models are priced per token, roughly a word fragment, for what goes in and what comes out. On the providers’ own pricing pages, read on 29 September 2026:
| Model | Input, per million tokens | Output, per million tokens |
|---|---|---|
| Anthropic Claude Haiku 4.5 (small) | $1 | $5 |
| Anthropic Claude Sonnet 5.5 (mid-tier) | $2 | $10 |
| OpenAI gpt-6-luna (small) | $0.10 | $0.50 |
| OpenAI gpt-6-sol (mid-tier) | $2.00 | $10.00 |
| Google Gemini 3.1 Flash-Lite (small) | $0.25 | $1.50 |
| Google Gemini 3.5 Flash (mid-tier) | $1.50 | $9.00 |
| Amazon Nova Micro on AWS Bedrock (small) | $0.035 | $0.14 |
| Amazon Nova 2.0 Lite on AWS Bedrock (small) | $0.33 | $2.75 |
| Amazon Nova 2.0 Pro on AWS Bedrock (mid-tier) | $1.375 | $11.00 |
Sources: Anthropic, OpenAI, Google, and AWS Bedrock (on-demand rates for US East, N. Virginia, from AWS’s price list published 28 September 2026). OpenAI’s figures are its short-context rates. Bedrock also runs other providers’ models, including Anthropic’s, at their own Bedrock rates, and regions differ. Prices change often, so check the pages before you budget.
For many teams, where the model runs matters as much as its price. Running it through AWS Bedrock keeps requests inside an AWS account you already have, with your existing access controls and billing, which can outweigh a small difference per token.
To make that concrete: a request that sends 3,000 tokens and gets 500 back costs about half a cent on a $1 / $5 model, and about a cent on a $2 / $10 model. Ten thousand of those a month is roughly $55 or $110. Anthropic’s own worked example puts about 10,000 support conversations on Claude Haiku 4.5 at around $37.
Model usage is rarely the whole running bill. Hosting, the database behind retrieval, monitoring and the time to review what the agent does all add to it. Azilen puts total monthly running cost for the agent types above at $3,200 to $13,000. Two published guides give annual maintenance as 15 to 30 percent of the original build, though they word it identically, so treat that as a rule of thumb rather than a finding.
How do you keep the first version small?
Start with one real task
Pick a request that arrived last week, not a category of work. Scope the agent around that one job and measure it.
Keep a person in the loop at first
Let the agent draft and a person approve. It is cheaper to make safe, and the approvals tell you where it can later act alone.
Put rules around the interpretation
Use the agent only for the step that needs reading and judgement, and plain automation for everything that follows a rule.
If you want to see what that looks like for your task, send a brief with one real example. We will tell you which parts need an agent, which don’t, and what would move your cost up or down.