NLP Build vs Buy: The Deliberation Is About the Cheap Part

11 min read
28 Sep 2026
NLP Build vs Buy: The Deliberation Is About the Cheap Part

An options paper is being written somewhere right now with two columns headed build and buy, and it will be argued over for six weeks.

Most of that argument is about the smallest part of the project. The NLP build vs buy decision moves roughly a quarter of the work. The other three quarters are identical in both columns, they are the part that determines whether the system is any good, and they get one bullet point each.

This piece is mostly about those three quarters. The decision itself takes about ten minutes once the arithmetic is on the table.

The question that no longer needs asking

Build used to mean training a model. Collecting a corpus, choosing an architecture, running experiments, and ending up with something that worked on your data and nothing else.

For almost every enterprise text problem, that is no longer a serious option, and pretending otherwise wastes a quarter. General models handle classification, extraction, summarisation and intent detection well enough out of the box that starting from raw architecture is difficult to justify outside genuine research settings.

So build now means something narrower: you host and control the model, rather than calling somebody else's. That is a deployment and operations decision, not a machine learning one, and it changes who should be in the meeting.

Say that out loud early. It usually shortens the discussion by a fortnight, because the person arguing hardest for build is often arguing for control and residency rather than for training anything.

Three buy shapes, and what actually separates them

Once building from scratch is off the table, three shapes remain.

A general model through an API. Broadest capability, no infrastructure, per call pricing, and the vendor's roadmap becomes yours. Best where the task varies, the volume is moderate, and the text is not sensitive.

A specialised service. Entity extraction, sentiment, translation, document parsing, sold as a narrow API. Often cheaper per call and better at its one job than a general model, and it has been evaluated by people who do only that. Worth checking before assuming a general model is the answer, because teams routinely reimplement a solved commodity.

A small model you host. Open weights, running on your infrastructure. Highest operational burden, lowest marginal cost at volume, and the only shape that satisfies a hard data residency requirement.

Comparison grid of general model API, specialised service and self hosted model across decision properties.

Four properties separate them, and only four.

1. Data residency. If text may not leave your infrastructure or your jurisdiction, that decides it, and nothing else in this article matters. Do not spend three weeks on a comparison whose answer was fixed by a contract clause on day one.

2. Volume economics. Covered below, and it is arithmetic rather than opinion.

3. Latency floor. A hosted small model on a warm instance beats a network round trip to a general model. If you are inside a tight budget, this decides it.

4. Task stability. A narrow, stable task suits a small hosted model. A task that keeps changing shape suits a general model, because you can adjust a prompt in an afternoon and cannot retrain that fast.

Notice what is absent from that list. Accuracy. On ordinary enterprise text tasks the gap between these three shapes is usually smaller than the gap between a good label schema and a bad one, which is the point of a later section.

The crossover is one division

The volume argument gets fought on instinct when it is a division.

Take your per call price from the vendor. Take the all in monthly cost of hosting the alternative, which means instance, storage, monitoring and a realistic share of an engineer's time. Divide the second by the first. That is the monthly call volume at which self hosting starts to win.

A worked example with stated assumptions, so you can substitute your own. If a call costs $0.002 and hosting the equivalent runs $900 a month all in, the crossover sits at 450,000 calls a month. Below that, the API is cheaper. Above it, hosting is, and the gap widens.

Two adjustments make the number honest.

Add the engineering time properly. Hosting is not the instance bill. Patching, model updates, capacity, an on call rotation and the incident when it falls over at 2am are all real. A quarter of an engineer is a conservative floor for anything serving production traffic, and adding it usually moves the crossover substantially.

Check your actual distribution, not your average. A system averaging 200,000 calls a month with a 900,000 call spike every quarter end has a different answer from one that is flat. Hosting handles spikes badly without headroom you pay for continuously. APIs handle them by charging you.

Run that division before the meeting. It ends the volume conversation in under a minute, and quite often the answer is nowhere near the boundary, which means volume was never the real argument and something else is.

Four things you build in every branch

Here is the part that should occupy most of the project plan.

  • A label schema. What the categories are, what each one means, and what happens at the boundaries. Its own section below, because it is where most of these projects fail.
  • An evaluation set. Real examples with agreed correct answers. Without it every change is an opinion, and you will make changes.
  • A fallback path. What happens when the model is unavailable, times out, or returns something unparseable. Every branch needs this. APIs have outages. Hosted models have deployments.
  • Drift monitoring. Also below, and it applies to both branches for opposite reasons.

None of those four depend on which shape you chose. They are 80 per cent of the difference between a system people trust and one that quietly gets bypassed, and in most options papers they appear as a line reading "plus integration and testing".

Diagram showing shared NLP modules across build and buy branches with only hosting differing.

Label schema, where these projects really die

A team builds a support ticket classifier. Sixty categories. Accuracy comes in at 71 per cent and everyone concludes the model is not good enough.

Then somebody has two experienced agents label the same 200 tickets independently and they agree with each other 74 per cent of the time.

The model was not the problem. Two humans who have done the job for years disagree about a quarter of these tickets, which means the categories are underspecified, and no model will exceed the ceiling that ambiguity sets.

That measurement, inter annotator agreement, costs about two days and it is the highest value two days in an NLP project. It tells you three things. The realistic accuracy ceiling. Which specific category pairs are being confused, which is almost always a small number of pairs doing most of the damage. And whether the fix is a better model or a merged category.

The usual outcome is unglamorous and effective: merge the six categories that nobody can separate, define the boundaries of the eight that matter, and watch accuracy rise without touching the model at all.

Forty to 90 hours for schema work, and it is the module to fund first regardless of which branch wins.

Chart showing model accuracy against inter annotator agreement as the achievable ceiling.

Two days that tell you the ceiling

The human review loop is the part that pays for itself

One of the five shared modules deserves more than a table row, because it is the one that turns a 78 per cent accurate system into something a business will actually run on.

No classifier is right every time, and the useful question is what happens to the ones it gets wrong. Two designs exist and only one of them works.

The design that fails. Everything is auto processed. Errors flow downstream, get discovered by a customer or by finance, and someone builds a spreadsheet to catch them. Within six months the spreadsheet is the real system and the model is decoration.

The design that works. Every prediction carries a confidence score, and anything below a threshold routes to a person. The person's correction is captured as a labelled example. Two things follow: nothing wrong reaches a customer without a human seeing it first, and your evaluation set grows from real traffic rather than from a workshop.

Setting the threshold is a business decision rather than a technical one, and it is worth doing explicitly. Pick it from how much review capacity exists, not from a number that sounds good. If your team can handle 200 reviews a day, set the threshold where roughly 200 items fall below it, then move it as accuracy improves and the volume naturally shrinks.

That last part is the compounding bit. As corrections feed back and the schema tightens, fewer items fall under the threshold, so review load falls while coverage stays complete. Teams that measure this typically see the review queue shrink over the first two quarters without anyone touching the model.

Fifty to 110 hours, in every branch, and it is the module that makes the difference between a pilot and a system.

Drift comes for both branches

Both shapes degrade. They degrade in opposite ways and both surprise people.

A bought model changes underneath you. The vendor updates it. Outputs shift, usually improving on average and occasionally changing behaviour on the specific edge cases your business depends on. You were not consulted and often not notified. Version pinning helps where it is offered and it is not always offered.

A hosted model does not change, and the world does. New product names, new slang, a new competitor, a policy change that alters what customers write about. The model is exactly as good as it was and increasingly wrong about the present.

One monitoring design covers both. Run your evaluation set on a schedule, weekly is usually enough, and alert on movement rather than on absolute score. Track the distribution of predicted labels too, because a shift in the mix is often the first signal, arriving before accuracy visibly moves.

Forty to 90 hours, and every team that skipped it discovered drift through a complaint rather than a dashboard.

What this costs, with the hours shown

The rate is $40 to $100 per hour by role. Integration and monitoring sit near the floor, schema and evaluation design near the ceiling, and mixed teams blend to around $65.

Shared, in every branch.

Module

Hours

What it covers

Task definition and label schema

40 to 90

Categories, boundaries, agreement measurement, the merge decisions

Evaluation set construction

50 to 110

Real examples, agreed answers, held out properly

Integration and fallback

50 to 110

Calling path, timeouts, retries, what happens when it is down

Drift monitoring

40 to 90

Scheduled scoring, label distribution, alerting on movement

Human review loop

50 to 110

Confidence thresholds, routing, and the correction feeding back

Worked example, buy branch. Schema 65 plus evaluation 80 plus integration 80 plus monitoring 65 plus review loop 80 equals 370 hours. That is $14,800 at $40, $37,000 at $100, and about $24,050 at a $65 blend. Span 230 to 510 hours.

Build branch adds one module.

Module

Hours

Model hosting, deployment and capacity

80 to 170

Worked example: 370 plus 125 equals 495 hours, about $32,175 at the blend, with a span of 310 to 680.

Excluded from both: per call fees, compute, and any licence. Those are usage rather than build, and they are exactly what the crossover division above is for.

Look at the ratio. The decision everyone argues about moves 125 hours out of 495. The four shared modules are 370, and they are the ones that determine whether anybody uses the thing.

The lock in argument, examined rather than repeated

Every build against buy piece lists lock in as the reason to build. It is a real concern and it is usually stated in a way that cannot be acted on, so here is a more useful version.

Lock in is not one thing. It has three layers and they cost very different amounts to escape.

  • The calling interface. Cheap to abstract. One adapter behind your own interface means swapping providers is a week rather than a quarter. Do this on day one in every branch, because it costs almost nothing when the code is new and is tedious afterwards.
  • The prompt or the tuning. Moderately painful. Work done against one model does not transfer cleanly, and moving means re running the evaluation and adjusting. This is a real cost and it is bounded, typically a few weeks, provided the evaluation set exists. Without the evaluation set it is unbounded, because you cannot tell whether the new provider is worse.
  • The data. The layer that actually traps people, and it is rarely the model at all. If your labelled examples, corrections and evaluation set live inside a vendor's product rather than in your own store, leaving means starting the schema work again. Keep the labels yours. Export them on a schedule. This costs nothing and it is the single most effective anti lock in measure available.

So the honest framing: the evaluation set and the label store are your portability, far more than the hosting decision is. A team that buys an API and owns its labels can move in a month. A team that self hosts but keeps its annotations in a vendor tool cannot, which is the reverse of what the two column table implies.

A decision rule, and what to do first

One. Does the text have to stay in your infrastructure? If yes, host. Stop here.

Two. Run the crossover division. If you are more than double the crossover volume, host. If you are less than half, buy. Between those, it is not the deciding factor and you should choose on something else.

Three. Is there a specialised service for this exact task? Check before assuming a general model. Entity extraction, document parsing and translation are commodities with providers who do nothing else.

Four. How stable is the task? Changing monthly favours a general model, because a prompt change is an afternoon. Stable for years favours a hosted small model.

If the four answers conflict, residency wins, then latency, then volume, then stability.

What to do first, whichever way it goes. Spend the two days on inter annotator agreement before anything else. It will tell you the ceiling, and roughly a third of the time it tells you the project as scoped cannot succeed, which is worth knowing in week one rather than week twenty.

Then build the evaluation set. Then pick a branch. The branch is the easy part and the industry has trained everyone to treat it as the hard one.

The decision moves 125 hours. The rest is 370.

FAQs

Not from scratch. For enterprise text tasks that stopped being justifiable outside research. Build now means hosting and controlling a model rather than calling somebody else's, which is a deployment decision rather than a machine learning one.

Divide the all in monthly hosting cost by the per call price. That is your crossover volume. If a call costs $0.002 and hosting runs $900 a month, the crossover is 450,000 calls a month. Include a realistic share of engineering time in the hosting figure, because a quarter of an engineer is a conservative floor.

Data residency first, which settles it outright when text may not leave your infrastructure. Then latency floor, then how stable the task is. Accuracy rarely separates the options, because the gap between shapes is usually smaller than the gap between a good and a bad label schema.

Most often because the categories are underspecified. If two experienced people labelling the same 200 examples agree only 74 per cent of the time, no model will beat that ceiling. Measuring inter annotator agreement takes about two days and is the highest value work in the project.

About 230 to 510 hours at $40 to $100 per hour by role for the buy branch, near 370 hours in a typical case, roughly $24,050 at a $65 blend. Self hosting adds a hosting and capacity module of 80 to 170 hours. Per call fees, compute and licences sit outside those figures.

Yes, and for a different reason. A bought model changes underneath you when the vendor updates it. A hosted model stays the same while the world changes around it. One design covers both: score your evaluation set weekly and alert on movement rather than on an absolute number.

Enough to cover every category including the rare ones, and drawn from real traffic rather than written by the team. Coverage of the boundary cases matters far more than raw count, because the boundaries are where both models and humans disagree.

Vikas Choudhary

Vikas Choudhary

Vikas has around fifteen years of experience building software and now builds generative AI systems at Zyneto. His work covers retrieval augmented generation, agentic AI, knowledge graphs, AI memory, and the evaluation and guardrails that decide whether any of it is safe to put in front of customers. He has shipped enterprise copilots, document AI, chatbots and predictive analytics for e-commerce, fintech and marketing teams, and works day to day in Python, JavaScript and SQL. He follows multimodal models, business process automation and enterprise AI security closely, and mentors engineers moving into AI. He writes about architecture, inference cost and the failure modes that only show up at production scale.

Let's make the next big thing together!

Share your details and we will talk soon.

Phone

We respond to all inquiries within 1 hour.

WhatsApp
Email
Book a Meeting