Multimodal AI Development: What Images, Audio and Video Actually Cost

11 min read
05 Oct 2026
Multimodal AI Development: What Images, Audio and Video Actually Cost

Most pages about multimodal AI development services list the modalities and stop. Text, image, audio, video. That tells you nothing you did not already know, and it skips the thing that actually decides whether the project happens, which is that the pricing model changes shape once you leave text behind.

Text bills per token and the arithmetic is simple. Images get cut into tiles and billed by resolution. Audio bills per minute of duration. Video bills per frame you choose to look at, and that choice is yours, which means your video bill is mostly a decision rather than a fact.

So this is the version with the arithmetic for each. Three worked pipelines, a comparison where plain optical character recognition beats a vision model, and a section on scoring outputs that have no single right answer. Every figure is a planning assumption written out so you can substitute your own.

What multimodal means once it is built

A multimodal system takes something other than text as input and produces something useful from it. In practice that is one of four jobs, and knowing which one you have is most of the scoping.

Extraction. Document extraction pulls structured fields out of a page, a photograph or a form. Invoice totals, meter readings, licence plates, part numbers. There is a right answer and you can check it, which makes this the easiest job to buy and the easiest to prove.

Description. Say what is in the input. Caption a photograph, summarise a call, describe damage in an inspection image. There is no single right answer, which makes evaluation much harder.

Classification. Put the input in a bucket. Is this receipt or contract, is this call a complaint, does this photograph show a defect. Right answer exists, easy to score, cheapest to run.

Search. Find the input that matches a query. Show me photographs with water damage, find the clause about termination. This needs embeddings across modalities and is a different build again.

Diagram comparing extraction, classification, description and search as multimodal jobs.

Most conversations about multimodal AI for business arrive asking for description and actually need extraction or classification. Extraction and classification are cheaper, more accurate and far easier to prove. Ask which one the business decision depends on, because that is the job you are buying.

Why the cost model is not the one you know

Here is the part that catches teams who have shipped a text project.

Images are tiled. A model does not read an image as one thing. It cuts it into tiles and each tile carries a token cost. Double the resolution and you roughly quadruple the tiles, so a 2048 by 1536 photograph can cost around four times what the same photograph costs at 1024 by 768. Most providers also expose a detail setting that lowers the tile count for a small accuracy trade.

A working assumption: a 1024 by 1024 image lands somewhere between 700 and 1,500 input tokens depending on provider and detail setting. Check your own provider's table rather than trusting that range, because it moves.

Audio bills by duration. Not by how much was said. A six minute call with forty seconds of speech costs the same as a six minute call with six minutes of speech. Silence is billed.

Video bills by frames you sample. This is the one where the number is entirely under your control. Sample one frame a second on a ten minute clip and you have 600 images to pay for. Sample one every five seconds and you have 120. Same clip, one fifth the bill, and for many tasks identical accuracy.

Diagram showing how text, image, audio and video inputs are priced.

Put those together and the useful instinct is this: on text you optimise the prompt, on images you optimise the resolution, on audio you optimise what you send at all, and on video you optimise the sampling rate. Three of those four have nothing to do with the model.

Documents, the image route against the OCR route

The most common multimodal project, and the one where the obvious answer is often wrong.

Assumptions. 10,000 scanned documents a month, averaging 2.4 pages, so 24,000 pages. Blended rates of 3 USD per million input tokens and 12 USD per million output. Substitute your own.

Route A, send the page image to a vision model.

Each page at reasonable resolution costs about 1,100 input tokens. 24,000 pages is 26.4 million input tokens, which at 3 USD per million is about 79 USD. Output of roughly 300 tokens per page adds 7.2 million output tokens, about 86 USD. Call it 165 USD a month, plus no preprocessing step to build or maintain.

Route B, run OCR first and send the text.

OCR the page, send the extracted text. A dense page is roughly 700 text tokens, so 24,000 pages is 16.8 million input tokens, about 50 USD, and the same 86 USD of output. That is 136 USD a month, plus whatever the OCR itself costs, which for a managed service is usually the larger line.

So the cost difference is small and it is not the deciding factor. What decides it is layout.

Route B throws away position. If the answer depends on which column a number sits in, which box a tick is in, or whether a signature is present, OCR has already discarded it by the time the model sees the text. Route A keeps the layout and handles the messy cases: a photographed page at an angle, a stamp over the total, a handwritten amendment in the margin.

Route B wins when pages are clean, uniform and text dense. Route A wins when they are not. A sensible design runs both: OCR first, and route to the vision model only the pages where OCR confidence is low or a required field came back empty. On a typical mixed corpus that is 15 to 30 percent of pages, which keeps most of the cost saving and most of the accuracy.

Audio, and the decision that halves the bill

Assumptions. 5,000 support calls a month averaging 6.2 minutes, so 31,000 minutes. You want a summary, a sentiment label and any commitment the agent made.

The naive build transcribes every minute, then sends every transcript to a model for the three outputs. Transcription bills per minute, so 31,000 minutes is your floor and there is nothing clever to do about it.

The transcript stage is where the choice sits. A 6.2 minute call is roughly 950 words, about 1,250 tokens. 5,000 of those is 6.25 million input tokens, about 19 USD, with output of around 200 tokens each adding 1 million tokens, about 12 USD. So 31 USD a month on top of transcription.

Two decisions change that materially.

Do not send the whole transcript for classification. Sentiment and call type are usually decided in the first ninety seconds and the last ninety. Sending 3 minutes instead of 6.2 cuts that input line roughly in half, to about 10 USD, with no measurable accuracy loss on classification. It does hurt summarisation, so run the two as separate calls against different slices rather than asking one call to do both.

Do not transcribe calls nobody will read. If only escalated calls need a summary, classify from metadata first and transcribe the 12 percent that qualify. That takes 31,000 minutes down to about 3,700 and the transcription bill with it. This is the single largest saving available in audio work and it is a routing decision, not a model decision.

Video, where the frame rate is the budget

Assumptions. 400 inspection clips a month averaging 4 minutes, so 1,600 minutes or 96,000 seconds of footage.

At one frame a second that is 96,000 images. At roughly 1,100 tokens each that is 105.6 million input tokens, about 317 USD a month before any output.

At one frame every five seconds it is 19,200 images, 21.1 million tokens, about 63 USD. One fifth the cost.

At one frame every fifteen seconds it is 6,400 images, about 21 USD.

The right rate is a property of the task, not the footage. Counting items on a shelf needs one good frame. Detecting whether a safety step was performed needs enough frames to catch a movement that takes two seconds. Reading a gauge needs one frame where the gauge is in focus, which argues for sampling more and discarding aggressively rather than sampling less.

A cheap technique that works: sample at a low rate, and only when the model flags something interesting, re-sample that window densely. On inspection footage that typically costs 10 to 20 percent of uniform dense sampling and misses very little, because the interesting seconds are a small fraction of the total.

Comparison of video frame sampling rates against monthly cost.

Price it by modality

Where accuracy falls off, by modality

Vendor pages give one accuracy number. The truth is that accuracy depends far more on the task than on the modality, and it falls off in predictable places.

Printed text in an image is close to solved. Clean documents, screenshots, labels. This is where multimodal is genuinely reliable.

Handwriting is much weaker and varies wildly by hand. Plan for a human review step on anything handwritten that matters, and measure it on your own samples rather than trusting a benchmark.

Counting is a known weakness. Models are poor at counting many similar objects in one image. If the business answer is a count above roughly ten, treat the model output as an estimate and design around that.

Fine spatial relationships degrade. Which of these two wires is on the left, is the gap above or below the line. Models often get the object right and the relationship wrong.

Audio with overlapping speakers loses accuracy sharply, and speaker attribution is harder than transcription. If who said what matters, that is a separate capability and it needs testing separately.

Long video suffers from the same context problem as long documents. Events early in a clip fall out of the window by the end.

Put a number on the handwriting case, because it is the one that gets promised and then quietly fails. If a form has 9 printed fields and 3 handwritten ones, and printed extraction runs at roughly 97 percent per field while handwriting runs nearer 80, then a fully correct form is 0.97 to the ninth times 0.80 cubed, which is about 39 percent. Three in five forms need a human even though every individual field looked acceptable in testing. Measure at the form level, not the field level, or the pilot will read far better than the rollout.

The pattern: extraction and classification hold up, description is adequate, anything requiring precise counting or spatial reasoning needs a human check. Design the workflow around that rather than hoping the next model release fixes it.

Scoring an output with no single right answer

Extraction is easy to score. The invoice total is either 1,284.50 or it is not. Description is not, and this is where multimodal projects stall.

Three approaches that work, in order of effort.

Score against required elements. Instead of grading the whole caption, list what a correct description must mention. For damage assessment: location, type, approximate size. Score each element present or absent. A 200 case set scored this way gives you a number you can track across releases.

Use pairwise comparison. Show a reviewer two outputs and ask which is better. Faster than rating each one, more consistent between reviewers, and enough to tell whether a change helped.

Use a model as a first pass reviewer, and check the reviewer. A second model can score outputs against your required elements at scale. It is only useful if you measure the reviewer against human judgement on a sample first, because a reviewer that agrees with humans 60 percent of the time is a random number generator with a confident tone.

Whichever you choose, build the set before the build finishes. A 200 case evaluation set takes a couple of days to assemble and is the only thing that tells you whether a model upgrade helped or hurt. Teams that skip it end up arguing from anecdotes six months later.

When text only wins

When the input is already text. Obvious, and still worth stating, because projects arrive asking for a vision model to read PDFs that contain a perfectly good text layer.

When pages are clean and uniform. OCR plus a text model is cheaper, faster and more predictable. Only reach for vision when layout carries meaning.

When you need the answer in milliseconds. Image and audio processing adds real latency. If this sits behind an interactive surface, measure before you commit.

When the decision needs counting or precise spatial reasoning. Covered above. A purpose built computer vision model, or a person, will beat a general multimodal model on those tasks for some time yet.

When the volume does not justify the build. Run the arithmetic before the proposal. At 400 documents a month rather than 10,000, the running cost falls to about 7 USD and the saving against a person doing it by hand might be 6 hours a month. Against a build measured in weeks, that is a payback counted in years. Multimodal work rewards volume more than most software does, because the engineering cost is roughly fixed and the running cost is genuinely small.

When you cannot assemble 200 labelled examples. If nobody can produce 200 cases with agreed correct answers, the project has no way to prove it works. That is a reason to pause, not a reason to build and hope.

Summary

Multimodal AI development services are priced differently from text work, and that is the thing to understand before anything else. Images tile by resolution, so a 2048 pixel photograph can cost four times a 1024 pixel one. Audio bills by duration whether or not anyone is speaking. Video bills by the frames you choose, which makes your video bill a decision.

On the worked assumptions above, 24,000 document pages a month through a vision model lands near 165 USD, against 136 USD plus OCR costs for the text route. The gap is small, so choose on layout rather than price, and route only the low confidence pages to vision.

Audio saves most from not transcribing calls nobody reads. Video saves most from sampling at one frame every five seconds rather than every second, which is a five fold difference for often identical accuracy.

Accuracy holds up on printed text, extraction and classification. It falls off on handwriting, counting above about ten, and fine spatial relationships. Build the workflow around those limits.

And build the evaluation set first. Two hundred cases, scored against required elements rather than as a whole, is what tells you whether the next model is better or just newer.

Build the set first

FAQs

Running cost depends on modality far more than on the model. On the planning assumptions above, 24,000 document pages a month is roughly 165 USD through a vision model, 400 four minute video clips sampled every five seconds is about 63 USD, and 5,000 support call transcripts add about 31 USD on top of transcription. Build cost follows the usual pattern: the integration and evaluation work is larger than the model work.

Images are cut into tiles and each tile carries a token cost, so price scales with resolution rather than with content. Doubling the resolution roughly quadruples the tiles. A 1024 by 1024 image typically lands between 700 and 1,500 input tokens depending on provider and detail setting. Check your provider's current table, because these numbers move.

Not always. OCR plus a text model is cheaper and more predictable on clean, uniform, text dense pages. A vision model wins when layout carries meaning, when pages are photographed rather than scanned, or when there are stamps, handwriting or annotations. The sensible design runs OCR first and routes only low confidence pages to the vision model, usually 15 to 30 percent of a mixed corpus.

Mostly it depends on your sampling rate, which you control. 400 clips of 4 minutes is 96,000 seconds. At one frame a second that is about 317 USD a month. At one frame every five seconds it is about 63 USD. Sample low, then re-sample densely only around anything the model flags.

Handwriting, counting more than about ten similar objects, fine spatial relationships such as which item is to the left, audio with overlapping speakers, and long video where early events fall out of the context window. Extraction and classification on printed inputs are the reliable end.

Score against required elements rather than grading the whole output. List what a correct description must mention, then score each element present or absent across about 200 cases. Pairwise comparison also works and is faster. If you use a model as the reviewer, measure it against human judgement on a sample first.

When the input already has a usable text layer, when pages are clean enough that OCR handles them, when the answer is needed in milliseconds, when the decision turns on counting or precise spatial reasoning, or when nobody can assemble 200 labelled examples to prove it works.

Vikas Choudhary

Vikas Choudhary

Vikas has around fifteen years of experience building software and now builds generative AI systems at Zyneto. His work covers retrieval augmented generation, agentic AI, knowledge graphs, AI memory, and the evaluation and guardrails that decide whether any of it is safe to put in front of customers. He has shipped enterprise copilots, document AI, chatbots and predictive analytics for e-commerce, fintech and marketing teams, and works day to day in Python, JavaScript and SQL. He follows multimodal models, business process automation and enterprise AI security closely, and mentors engineers moving into AI. He writes about architecture, inference cost and the failure modes that only show up at production scale.

Let's make the next big thing together!

Share your details and we will talk soon.

Phone

We respond to all inquiries within 1 hour.

WhatsApp
Email
Book a Meeting