Recommendation Systems: Which ML Model to Use

10 min read
08 Oct 2026
Recommendation Systems: Which ML Model to Use

Somebody has asked for recommendations and the first meeting has gone straight to model selection. Two-tower or matrix factorisation. Transformers, probably, because everything is transformers now.

That meeting is premature. The model is the part you can swap in a fortnight. What you cannot swap later is the interaction data you did not log, the evaluation you did not define, and the decision about whether this is a ranking problem or a matching problem at all.

This covers the 5 model families and when each one fits. It also covers the case nobody writes about, which is when the answer is a rule rather than a model.

Ranking and matching are not the same problem

Ranking orders items for 1 user and the item has no opinion. Matching needs both sides to accept, and recommending your 20 most attractive items to everyone makes the system worse, because those 20 saturate and the other 40,000 see nothing.

Two panel comparison of a ranking problem and a matching problem, and how success is measured in each.

Most writing collapses the two into one word. Keep them apart.

Ranking. A feed, a product carousel, a next-episode suggestion. Success is engagement with something in the first 3 or 4 slots.

Matching. Candidates and jobs, drivers and loads, buyers and suppliers. A recommendation counts only when both parties act, so the metric is completed mutual acceptances rather than clicks.

The 5 families below serve both. The objective and the evaluation differ completely, and a matching problem measured with ranking metrics will score well in testing and fail in production. If both sides can decline, say so in week 1.

The five families, and what each one needs

Described by the condition that makes them work, because how they work is in every textbook and what decides your project is the first thing.

Flow diagram of a two stage recommender, from the full catalogue through retrieval and a shortlist to ranking.

Collaborative filtering. Learns from who interacted with what and ignores the content. Matrix factorisation and its implicit-feedback variants, ALS and BPR, are the standard implementations. Needs a median of roughly 20 to 50 interactions per item before the pattern means anything. Useless on day 1 and useless for an item nothing has touched.

Content-based. Scores items by attributes against a profile of what the user engaged with before. Needs metadata somebody maintains. Handles a new item the moment it is catalogued, which is why it is usually the first thing in production. It narrows: 2 articles about Kubernetes and the feed is Kubernetes for a month.

Two-tower embedding retrieval. Learns a vector for users and a vector for items so cosine similarity means relevance, then searches that space with approximate nearest neighbours. HNSW indexes return the top 500 candidates from 10 million vectors in single-digit milliseconds, which is why this is the answer at scale. It needs training data, a vector store and an owner for the refresh cycle.

Learning to rank. Takes 200 to 500 candidates from a cheaper method and orders them with features you choose: recency, price, margin, distance, stock. LambdaMART and gradient-boosted trees still beat deep models on tabular features at this size. This is where business logic belongs and it is the layer most teams bolt on badly 6 months later.

Graph-based. Treats users and items as a network and propagates signal across it. Genuinely better where the relationships carry meaning that a rating does not, such as supply chains and professional networks. Heaviest to operate.

In production, systems are 2 of these stacked: something cheap to retrieve 500 candidates, then learning to rank over the shortlist. 1 model doing both is a prototype.

Cold start is a design problem, not a modelling one

Every recommender has 3 cold starts and the literature addresses the least important of them.

New user. No history. Solved with 2 or 3 onboarding questions, popularity within a segment, or a borrowed signal such as acquisition channel. Not solved by a better model.

New item. Nothing has touched it. This is content-based scoring's reason to exist and it is why a purely collaborative system cannot be your only system.

New system. The one that sinks projects. On day 1 there is no interaction data at all, so every learned approach has nothing to learn from. The honest first version is rules plus popularity plus whatever metadata exists, instrumented properly, running for 3 to 6 months until there is history worth training on.

Teams that skip that ship a model trained on 4 weeks of sparse events, watch it underperform the rules it replaced, and conclude machine learning does not suit their domain.

Your offline metrics are probably flattering you

Offline evaluation has a structural problem. Your logs contain only what the old system showed. An item never surfaced has 0 interactions, so it scores as irrelevant, so the new model learns not to show it either.

Flow diagram of the feedback loop that keeps a recommender trained on its own logs showing the same items.

3 consequences worth planning around.

Precision at 10 flatters ranking and misleads matching. If both sides must accept, measure completed mutual outcomes. NDCG and MAP have the same blind spot when the label is a click rather than an acceptance.

Accuracy and usefulness diverge. A model that recommends what the user would have found anyway scores well and adds nothing. Watch catalogue coverage and the share of results coming from outside the top 100 items.

Only an online test settles it. Budget the experiment infrastructure into phase 1. A recommender with no way to run a controlled comparison can be changed but not improved.

When a rule beats a model

This section costs us work and belongs here.

Under about 2,000 items with clear attributes, filters and a sort order perform comparably, can be explained to a regulator, debugged by an analyst and changed on a Tuesday afternoon.

When the catalogue turns over faster than you can train. Ticketing, classifieds and short-run stock. If the median item sells in 5 days and your retrain cycle is 7, the model spends its life out of date.

When the decision must be explainable. Lending, hiring, insurance pricing. GDPR Article 22 gives a data subject rights around solely automated decisions with legal or similarly significant effects, and "the matrix factorisation said so" is a poor position to defend.

When nobody owns it after launch. A recommender decays. Without someone retraining it on a schedule, rules you can reason about will beat an unmaintained model inside 12 months.

A sensible sequence is rules with instrumentation, then content-based scoring once metadata is reliable, then a learned model once the interaction data justifies it.

Rules first. The model comes later.

What sits underneath all of it

The model is perhaps 20 per cent of the build. 4 things carry the rest.

Event capture. Impressions, position, dwell, clicks, declines and conversions, with 1 identity across sessions and devices. Impressions are the most common gap: without knowing what was shown and ignored you cannot separate an item nobody wanted from an item nobody saw. This is the piece that cannot be retrofitted. Data not logged in January does not exist in June.

A feature and vector store. Somewhere to keep embeddings and features so training and serving read the same values. Skipping it produces a model that behaves differently in production for reasons that take 2 or 3 weeks to find. Those are ordinary engineering choices, covered in our database development work.

Serving inside the budget. Retrieval and ranking need separate allowances. A practical split on a product page is 20 to 30 ms for retrieval and 40 to 60 ms for ranking, measured on production hardware rather than a laptop.

Retraining and monitoring. A schedule, drift alerts and a dashboard somebody opens. Where the signals are text, the feature work overlaps our NLP development practice. Where it runs per tenant, the isolation questions sit in our SaaS platform development work.

Diversity, saturation and the rules nobody models. A recommender optimised purely for predicted engagement converges on a narrow set and keeps showing it. Correct for the objective, wrong for the business.

Saturation. In 2-sided systems the same 20 suppliers get recommended to everyone until they cannot service the demand. The fix is a cap on how often an item appears across users in a 24 hour window, not a better model.

Diversity in the result set. 10 near-identical items is a worse page than 7 good ones and 3 different ones, even though the first scores higher. Apply it after ranking.

Freshness. New items need exposure to earn the interactions that justify exposure. Reserve 5 to 10 per cent of slots for exploration or the catalogue ossifies.

Hard exclusions. Already owned, out of stock, already declined, regionally restricted. Keep this list outside the model, editable without a retrain, because it changes for commercial and legal reasons.

Implicit signals carry the system, and they lie in known ways

Explicit ratings are the cleanest signal and almost nobody produces them. Rating coverage on consumer products typically sits in the low single digits of sessions, so a system built on stars is a system built on 2 per cent of its users, and the 2 per cent who rate are not representative.

So implicit signals carry the load, and each has a known distortion worth correcting for.

Clicks are positional. An item in slot 1 collects several times the clicks of the same item in slot 10, purely from where it sat. Log the position with the impression and the correction is arithmetic. Omit it and the model learns that being shown first causes relevance.

Dwell has a floor and a ceiling. Under about 5 seconds is a bounce rather than interest. Very long dwell on a short item often means the tab was abandoned. Treat both tails as noise rather than as signal.

Purchases are sparse and late. On most catalogues orders run 1 to 3 per cent of product views, so a model trained on orders alone sees a fraction of the behaviour and reacts a week behind.

Declines are the strongest signal in matching and the one nobody logs. A skip, a dismiss, an unmatch or a rejected application tells you more than an accept, because the accept may be the only option shown. Capture it from day 1.

Absence is not a negative. An item a user never saw is not an item they rejected. Treating unseen as disliked is the single most common labelling error in these builds and it bakes the old system's choices into the new one.

What the first 90 days should produce

A phase 1 that ends with a learned model and no instrumentation has built the wrong thing in the right order. A reasonable 90 days looks like this.

Weeks 1 and 2. Event schema agreed and shipped. Impressions, position, dwell, declines, conversions, 1 identity across devices. Nothing else matters if this is wrong.

Weeks 3 to 6. Content-based scoring and a popularity fallback live in 1 slot, not all of them. The hard exclusion list written down and enforced at serve time.

Weeks 7 to 10. The experiment tooling, so a change can be compared rather than asserted. A holdout of 5 to 10 per cent of traffic with the block suppressed, which is the only way to read incremental rather than attributed value.

Weeks 11 to 13. Read the data. Decide whether interaction density now supports a learned model, and in which slot. Often it supports it in 1 slot and not the others, which is a perfectly good answer.

At the end of that you have a working recommender, a measurement loop and the data to justify the next step. Teams that invert this order spend the same 13 weeks and end with a model they cannot evaluate.

What it costs to build

Rates are $40 to $100 per hour by role. Data and front-end work sits near the floor. Retrieval architecture, evaluation design and anything under a tight latency budget sits near the ceiling. A mixed team blends to $60 to $70, and that blend is the number worth planning against.

A content-based scorer over a catalogue with clean metadata is usually 120 to 200 hours. A 2-stage system with event capture built from scratch, a vector store, online experimentation and a retraining pipeline runs 600 to 1,100 hours, and under 150 of those are modelling. The difference is instrumentation.

We do not quote a fixed price before discovery, because a figure given before anyone has looked at what you log today is padded for our uncertainty rather than priced for your project. You will get an hour band on the first call. If the honest answer is to ship rules and instrumentation first and revisit in 2 quarters, you will hear that instead. Where the output is a forecast rather than a match, the work sits with our predictive analytics development team.

Three scenarios, and what we would actually build

2,000 products, 18 months of orders, no recommender today. Content-based scoring over existing attributes, popularity fallback by category, event capture from day 1. No learned model in phase 1. Orders sound like enough and are not, because orders are perhaps 2 per cent of the volume of views and views are what has not been logged. Revisit after 3 months of view data.

A 2-sided marketplace, 40,000 suppliers and 40,000 buyers, both can decline. The decline signal is the most valuable column in the table. Two-tower retrieval for 500 candidates, learning to rank with availability and distance as explicit features, and a saturation cap. Evaluate on completed acceptances. Precision at 10 will look fine and mean nothing.

A media library, 300,000 items, heavy daily turnover. Recency and availability rules first, because by the time a model trains the catalogue has moved. Two-tower retrieval becomes worthwhile only once a durable back catalogue exists that people return to.

In 2 of those 3, the first version contains no machine learning. That is the sequence that reaches a working system fastest and leaves you with the data that makes the next step possible.

The short version. Decide first which problem you have, ranking or matching, because it changes the metric. Expect to stack retrieval and ranking rather than asking 1 model to do both. Treat cold start as a design problem. Assume your offline numbers are optimistic until an online test says otherwise. And be willing to ship rules first, because a rule you maintain beats a model you do not.

$40 to $100 per hour, published.

FAQs

Content-based scoring over the attributes you already hold, plus a popularity fallback, plus event capture from day 1. It works before any interaction data exists, it costs 120 to 200 hours, and it gives you a baseline to judge a learned model against. Starting with collaborative filtering on a new system means training on 4 weeks of sparse events and concluding the approach does not suit your domain.

Look at the median interactions per item, not the totals. Below roughly 20 per item the overlap between users is too thin for the pattern to hold, however healthy the headline looks. Views run perhaps 50 times the volume of orders, so a store with 18 months of order history and no view logging usually has less usable signal than it expects.

A content-based scorer over a clean catalogue is 120 to 200 hours. A 2-stage system with event capture, a vector store, experimentation and retraining runs 600 to 1,100 hours. If a vendor quotes 6 weeks for the second one, ask which of those 4 pieces they are leaving out.

Rates are $40 to $100 per hour by role, blending to $60 to $70 on a mixed team. The spread between a simple recommender and a serious one is instrumentation and evaluation rather than modelling, which is under 150 hours of the larger build. We give an hour band on the first call and a fixed price after discovery.

Often, and it is worth pricing first. A managed service still needs the same event capture, catalogue hygiene and evaluation discipline, so it removes the modelling rather than most of the work. It becomes wrong when your ranking needs business constraints the service cannot express, or when the data cannot leave your estate.

Your logs contain only what the previous system chose to show. An item never surfaced has 0 interactions, scores as irrelevant, and the new model learns to keep hiding it. Offline evaluation rewards agreeing with the past. Only a controlled online test settles whether a change helped.

Only for embedding retrieval at a scale where scanning will not do. Below about 200,000 items an ordinary index and a well-chosen similarity search are usually enough and far less to operate. The question that matters more is whether training and serving read the same feature values.

Vikas Choudhary

Vikas Choudhary

Vikas has around fifteen years of experience building software and now builds generative AI systems at Zyneto. His work covers retrieval augmented generation, agentic AI, knowledge graphs, AI memory, and the evaluation and guardrails that decide whether any of it is safe to put in front of customers. He has shipped enterprise copilots, document AI, chatbots and predictive analytics for e-commerce, fintech and marketing teams, and works day to day in Python, JavaScript and SQL. He follows multimodal models, business process automation and enterprise AI security closely, and mentors engineers moving into AI. He writes about architecture, inference cost and the failure modes that only show up at production scale.

Let's make the next big thing together!

Share your details and we will talk soon.

Phone

We respond to all inquiries within 1 hour.

WhatsApp
Email
Book a Meeting