Custom AI Agent Development: Scope, Architecture and Running Cost

11 min read
01 Oct 2026
Custom AI Agent Development: Scope, Architecture and Running Cost

Most pages about custom AI agent development explain what an agent is. If you are reading this, you already know. You want to understand what decides the price, what the thing costs every month once it is live, and what breaks when real people start using it.

So this is the version with the numbers showing. Three architectures with monthly running costs. A scoping table you can score your own project against. The failure modes that only turn up after launch. And a section on when an agent is the wrong answer, because that section is missing from almost every page selling this work.

Every figure below is a planning assumption, written out so you can substitute your own. They are illustrative and not a claim about any particular project.

What custom actually means here

An AI agent is software that takes a goal, decides its own sequence of steps, and calls tools to complete them. That last part is what separates it from a chatbot. A chatbot answers. An agent books the appointment, updates the record and tells you what it did.

Custom means the tools are yours. The agent calls your inventory system, your CRM, your pricing rules, your approval chain. That is the whole reason to build rather than buy, and it is also where the cost lives.

Three things are genuinely custom on most builds.

The tool layer. Every action the agent can take is a function you write and test. Ten tools is a small agent. Forty is a large one. This is usually 50% to 65% of the build.

The policy. What the agent may do alone, what it must ask about, what it may never do. This sounds like configuration and turns out to be the hardest conversation in the project, because it forces a business to write down rules it has never written down.

The evaluation suite. A set of cases with known correct outcomes, run on every change. Without it you cannot tell whether a model upgrade made things better or worse, and you will be asked.

Diagram showing the three custom parts of an AI agent build against the rented model.

Notice what is not on that list. The model. You rent that, and swapping providers should be a configuration change rather than a rewrite. Any proposal that treats the model as the custom part has the cost structure backwards.

What decides the size of the build

Five variables. Score your own project against them before you ask anyone for a number, because this is what any honest estimate is built from.

Variable

Small

Medium

Large

Tools the agent can call

3 to 8

9 to 20

21 or more

Systems it integrates with

1 to 2

3 to 5

6 or more

Actions that change data

read only

1 to 3 writes

4 or more writes

Who it serves

internal team

whole company

external customers

Regulated data involved

none

some

core to the job

Read only is the dividing line. An agent that answers questions from your own data is a retrieval problem with a conversation on top. The moment it writes, four new questions arrive and every one of them has to be built: what happens when the same instruction lands twice, how a wrong action gets undone, who approved it, and what the record shows six months later. That is the whole gap between the small and medium columns above, and it is why a write tool runs a little over twice a read tool in the hours below.

External users move it again. An internal tool can be wrong occasionally and someone shrugs. A customer facing agent needs rate limiting, abuse handling, content policy and a support path for when it gets something wrong in public.

Scoring grid showing the five variables that decide AI agent build size.

A useful sanity check. Count the tools. Multiply by roughly 6 to 10 engineering hours each for a tool that reads, and 14 to 22 for one that writes and must be reversible. That gets you inside the right order of magnitude before anyone writes a proposal.

Three architectures, and what each costs to run

The architecture choice drives running cost far more than the model choice does. These three cover most of what gets built.

Single call with tools. One model call, a handful of tools, no loop. The model picks a tool, gets a result, answers. Cheap, fast, predictable. It cannot handle a task that needs several dependent steps.

Orchestration loop. The model plans, calls a tool, reads the result, decides the next step, repeats until done. This is what most people mean by an agent. It costs several model calls per task rather than one, and the count varies with how hard the task is.

Retrieval first, then act. The agent searches your own content before it plans, which is the pattern behind agentic RAG. Adds a vector store and embedding costs, and cuts the number of wrong turns considerably when the task depends on company specific knowledge.

Monthly running cost, using planning assumptions of 4,000 input and 800 output tokens per model call, and a blended rate of 3 USD per million input and 12 USD per million output tokens. Substitute your provider's current rates.

Tasks a month

Single call

Orchestration loop

Retrieval first

5,000

about 110 USD

about 480 USD

about 560 USD

25,000

about 550 USD

about 2,400 USD

about 2,800 USD

100,000

about 2,200 USD

about 9,600 USD

about 11,200 USD

Work one row through so the shape is obvious. At 25,000 tasks a month on the orchestration loop, 4.4 calls per task is 110,000 model calls. Each carries 4,000 input tokens, so 440 million input tokens at 3 USD per million is 1,320 USD. Each returns 800 output tokens, so 88 million output tokens at 12 USD per million is 1,056 USD. That is 2,376 USD, which rounds to the 2,400 in the table. Change any one of those four inputs and the total moves with it.

The loop assumes an average of 4.4 model calls per task. That average is the single most important number in your cost model and it is the one nobody measures until month three. A task that averages 7 calls instead of 4 costs 60% more to run, and prompt or tool changes move it without anyone noticing.

Three costs sit outside that table and belong in the budget.

Retrieval infrastructure. A managed vector store for a mid sized corpus generally runs in the low hundreds a month. Embedding the corpus once is cheap. Re-embedding it every time the content changes is not, so plan incremental updates rather than full rebuilds.

Observability. Every agent run should be traceable: which tools fired, in what order, with what arguments. You will need this within the first fortnight. Budget for it up front rather than retrofitting it during an incident.

Human review. Somebody reviews what the agent did. At launch that is most runs. After a quarter of evidence it can fall to a sample. It never reaches zero on anything that writes to a system.

What the build itself costs

Build cost follows the scoping table more than anything else. Using the same size bands:

Size

Engineering weeks

What it covers

Small

4 to 7

3 to 8 read tools, one or two systems, internal users

Medium

8 to 16

up to 20 tools, several systems, a few write actions, approval path

Large

17 to 30+

21 or more tools, regulated data, external users, full audit trail

Where the hours actually go, and this surprises people every time:

`` integration and tools 50 to 65 percent evaluation and testing 15 to 22 percent policy and guardrails 10 to 15 percent prompting and model work 5 to 12 percent ``

Prompt work is the smallest line. It is also the only part most demos show. If a proposal is heavy on prompting and light on integration, it has been priced from a prototype.

Put the tool arithmetic through the same treatment. A medium build with 18 tools, of which 4 write and 14 read, comes to 14 read tools at 8 hours and 4 write tools at 18 hours, so 112 plus 72 is 184 hours on the tool layer alone. If that is 58 percent of the build, the whole thing lands near 317 hours, or roughly 8 weeks for one engineer and 4 for two. That is a sanity check, not a quote, but it will tell you within about 20 percent whether a proposal is priced from the real work.

Add discovery before any of it. Two to four weeks to map the process, agree the policy and build the first evaluation set. Projects that skip discovery do not go faster, they just discover the same things in week nine at a worse moment.

The failure modes that only appear in production

None of these show up in a demo. All of them show up within the first quarter.

Silent tool failure. A tool returns an empty result rather than an error, the agent treats empty as "nothing found", and confidently reports that a customer has no orders. Fix: tools return typed errors, and the agent is instructed to distinguish "no data" from "the call failed".

The loop that will not end. The agent retries a failing step forever, or oscillates between two tools. Fix: a hard step limit per task, plus a cost ceiling that stops the run and escalates. Both are one line each and both are missing from most first builds.

Duplicate writes. The agent retries after a timeout, the first call actually succeeded, and the customer gets charged twice. Fix: idempotency keys on every write, generated from the task rather than the attempt. This is the single most expensive failure on the list and the easiest to prevent. Cost it out: at 25,000 tasks a month, a 0.4 percent duplicate rate is 100 incidents, and if each takes 25 minutes of support time plus a refund, that is about 42 hours a month plus the refunds. The fix is perhaps 6 hours of engineering, once.

Context drift on long tasks. As a run grows, earlier steps fall out of the context window and the agent forgets constraints it was given at the start. It is why AI memory design matters more on agents than on chatbots. Fix: carry a compact task state separately rather than relying on conversation history.

Prompt injection through your own data. A supplier writes instructions into an invoice description field. The agent reads it as an instruction. This is not hypothetical and it is the reason securing enterprise AI has to be part of the build rather than a later phase. Fix: treat all retrieved content as data, never as instruction, and keep the privileged instructions outside anything a user or a document can reach.

Quiet accuracy decay. Nothing breaks. The agent just gets slightly worse as formats change and volumes shift. Fix: the evaluation suite, run weekly, with a number someone actually looks at.

Checklist of six AI agent production failure modes paired with their fixes.

When an agent is the wrong tool

An honest section, because the answer is "often".

When the sequence is fixed. If the steps never vary, write a script. It is cheaper, faster, testable, and it does not need a model at all. A surprising number of "agent" projects are a workflow with three branches wearing a costume.

When one model call would do. Classifying a ticket, extracting fields from a document, drafting a reply. These need a model, not an agent. No loop, no tools, a fraction of the cost.

When you cannot describe the policy. If the business cannot state what the agent may do alone, the project is not ready. Build the policy first. That work is valuable whether or not you build anything after it.

When being wrong is expensive and invisible. An agent adjusting credit limits or approving payments needs a person in the loop, at which point most of the speed benefit disappears and a good interface may serve you better.

When the data is not there. An agent cannot answer from knowledge you have not made available to it. If the answer lives in someone's head or a folder of scanned PDFs nobody has indexed, fix that first. AI agents for business fail far more often on missing data than on model capability.

Score your agent idea

How to structure the first build

Pick a task with a measurable before. Current handling time, current error rate, current volume, captured before anything changes. Without it you cannot prove the saving and you will not get funded twice.

Start read only. Ship an agent that drafts and recommends before you ship one that writes. You learn where it is wrong at no risk, and the evaluation set you build in that phase is what makes the write version safe.

Set the step limit and the cost ceiling on day one. Both are trivial to add at the start and awkward to retrofit after an incident.

Instrument before you optimise. You cannot reduce the average calls per task until you can see it. Traces first, tuning second.

Run it alongside the humans for four to six weeks. Compare outputs, collect disagreements, feed them into the evaluation set. This is the phase teams cut when they are behind schedule and it is the phase that determines whether anyone trusts the system.

What belongs in the contract

Short section, but these five clauses save arguments later.

An accuracy number with a definition. Not "high accuracy". A percentage, measured on a named evaluation set, with who owns the set written down.

A cost per task ceiling. With an agreed process for what happens if the average calls per task rises above the assumption.

Ownership of the evaluation set and the prompts. These are your business rules expressed in another form. They should be yours.

Model portability. The right to swap providers without a rebuild, and a statement of what that migration would involve.

A named owner for accuracy after handover. Agents are not projects that finish. If nobody owns the weekly evaluation run, accuracy decays and nobody notices until a customer does.

Summary

Custom AI agent development is priced by tool count, integration count, whether the agent writes to systems, who it serves, and whether regulated data is involved. Read only against internal users is a different project from one that writes on behalf of customers, usually by a factor of two or more.

Running cost is driven by architecture, not by model choice. A single call with tools costs roughly a fifth of an orchestration loop at the same volume, and average calls per task is the number that decides your bill. Measure it early.

The failure modes that matter are silent tool failures, endless loops, duplicate writes, context drift, prompt injection through your own data, and quiet accuracy decay. Every one is preventable and every one is cheaper to prevent than to fix.

And a good proportion of the time the right answer is not an agent at all. If the sequence is fixed, write a script. If one model call would do, make one model call. The agents worth building are the ones where the sequence genuinely varies and the tools are genuinely yours.

Know the monthly bill

FAQs

Build cost follows scope rather than a list price. A small read only agent with three to eight tools against one or two systems is typically 4 to 7 engineering weeks. A medium build with up to 20 tools, several systems and a few write actions runs 8 to 16 weeks. A large build with regulated data and external users is 17 to 30 weeks or more. Add two to four weeks of discovery before any of it.

Architecture decides this more than the model does. On planning assumptions of 4,000 input and 800 output tokens per call and a blended rate of 3 USD per million input and 12 USD per million output, 25,000 tasks a month costs roughly 550 USD as a single call with tools and roughly 2,400 USD as an orchestration loop. Add retrieval infrastructure, observability and human review on top.

A chatbot answers questions. An agent takes a goal, decides its own sequence of steps, and calls tools to complete them. The practical difference is that an agent changes something in a system, so it has to survive the same instruction arriving twice, let a wrong action be reversed, and leave a record of who approved what. A chatbot needs none of that.

Ship read only first and run it alongside the existing process for four to six weeks, collecting disagreements into an evaluation set. Most teams are ready to enable writes after that, starting with one reversible action rather than all of them.

Yes, if the build was structured for it. The model should be a configuration choice, not something the tool layer depends on. Put model portability in the contract, along with a statement of what a migration would actually involve.

Duplicate writes from retries, and silent tool failures where an empty result is read as "nothing found". Both are preventable with idempotency keys and typed errors. The slower problem is accuracy decay, which needs a weekly evaluation run and a named owner.

When the sequence of steps never varies, write a script instead. When a single model call would do the job, make one call. When the business cannot state what the agent may do without asking, build that policy first. All three are cheaper than the agent and two of them remove the need for it.

Vikas Choudhary

Vikas Choudhary

Vikas has around fifteen years of experience building software and now builds generative AI systems at Zyneto. His work covers retrieval augmented generation, agentic AI, knowledge graphs, AI memory, and the evaluation and guardrails that decide whether any of it is safe to put in front of customers. He has shipped enterprise copilots, document AI, chatbots and predictive analytics for e-commerce, fintech and marketing teams, and works day to day in Python, JavaScript and SQL. He follows multimodal models, business process automation and enterprise AI security closely, and mentors engineers moving into AI. He writes about architecture, inference cost and the failure modes that only show up at production scale.

Let's make the next big thing together!

Share your details and we will talk soon.

Phone

We respond to all inquiries within 1 hour.

WhatsApp
Email
Book a Meeting