DeepSeek model integration
DeepSeek is on the shortlist for one reason: careful step-by-step answers at a small fraction of the price, from a model you can download and run yourself. The complication is not the model. It is where the company's own hosted service runs, and that is a purchasing question before it is a technical one.
01Capabilities
Where a cheaper reasoning model earns its keep
DeepSeek is a model family from a Chinese lab, published with downloadable weights and priced well below the frontier providers. This page is judgement about when that trade is worth taking and how to take it safely, which in practice is AI integration work with a governance decision at the front of it.
- Volume
Work that only exists at a low price
Sorting, triage and extraction across whole catalogues and backlogs, the shape behind e-commerce product data. At frontier prices these jobs stay proposals. At a fraction of that they become projects, which is the only reason this model is on the list.
- Boundary
Reasoning that stays inside your account
Downloadable weights mean the model runs in your own cloud account and your own region. That also happens to end the argument about where the vendor's service lives, which is why most of our DeepSeek work is deployment work rather than integration work.
- Tiering
A cheap first pass, an expensive second one
The inexpensive model handles the bulk and hands the hard cases up. With a measured rule for when it hands over, this cuts spending materially. With a rule somebody guessed, it quietly gets things wrong where nobody is looking.
- Overnight
Backlogs nobody is waiting for
Archives, old tickets and document piles worked through while the office is closed. Nothing is urgent, so the only number that matters is how much you get through per pound spent, and that is the number this family is good at.
02Fit
The saving, and what it is allowed to cost you
The technical case is easy to state. On reasoning, mathematics and code these models are competitive, and the price is not close. The complication is deployment: the hosted service is operated from China under Chinese law, the consumer app has been restricted on government devices in several countries, and buyers in the EU, UK and US often rule it out before any technical evaluation. Four of the six rows below end somewhere other than a DeepSeek project with us.
| Your situation | What we recommend |
|---|---|
| Very high volume, nothing sensitive in the data, and price decides feasibility | Use DeepSeekWith a test that proves the cheaper model clears your quality bar rather than nearly clearing it. Nearly is where the cost comes back. |
| You want the reasoning but the data has to stay in your own account | Run it yourselfSame serving conversation as Llama, with a different model in the slot and the same hardware questions underneath. |
| Customer, health or financial data, and the plan was the hosted service | Not that serviceEither run the weights in your own region, or use a provider your reviewers already accept, see Claude or Gemini. |
| Purchasing wants a named supplier with a support contract behind it | Pay the higher priceA low price per call does not survive a security questionnaire. Arguing that it should is how a quarter disappears before anything ships. |
| One decision with four possible answers, and thousands of examples of it | Train something smallA chat model is an expensive way to output one of four labels, see PyTorch. Small, specific and dull beats clever here. |
| A fixed graphics-card budget, and the model has to fit inside it | Different pageQwen publishes a wider spread of genuinely small models, which is the difference between a project and a hardware order. |
Scope
What we own when a cheap model goes into a pipeline is the test that proves it is good enough for that specific job, the rule for when it hands work upwards and the alert when that rate drifts, the deployment location matched to how your data is classified, and a written answer to the question your security reviewer will ask about where the answering happens. unicrew has been building software since 2012, with 100+ senior in-house engineers across six countries, under an ISO 27001:2022 certified information-security management system. Where the saving is not worth the governance conversation, we will say so.
03Delivery
Testing a cheap model without talking yourself into it
Cheap is only cheap if the quality holds. The sequence below keeps price and quality as two separate questions, because merging them is how a team argues itself into a saving that costs it accuracy.
- Define good enough before anyone quotes a priceA test set from your data with the answers you would accept. Without it a cost comparison compares invoices rather than outcomes, and the cheapest option wins on paper every single time.
- Run the identical test across cheap and frontierSame prompts, same inputs, scored the same way. What matters is the size of the gap and where it falls. A model that only loses on the rare hard cases is a routing problem, not a rejection.
- Pick the location from the data classification, not the price listPublic, internal and regulated data get different answers, and the answer is chosen first. Downloadable weights are what make this decision short instead of a three-month negotiation.
- Build the handover rule, then watch itA hard rule for when the cheap model escalates, plus an alert when that rate moves. Rising usually means your inputs changed. Falling usually means the cheap model has started guessing confidently, which is worse. It goes through the same AI QA and evaluation practice as our own work.
04Stack
What a tiered pipeline is made of
The pieces that turn up around a cost-first model, each with its own page if that is the decision you are actually making.
05Questions
Asked by the person who signs the security form
Six that come up before the technical evaluation is allowed to start, answered the way we would answer them live.
Yes, and what is worth having here is judgement about deployment and routing, not familiarity with an interface that looks like every other one. What we will not do is put a cheaper model into your pipeline and leave the quality consequences unmeasured. An architect outside the delivery team reviews the design, and behaviour is covered by our AI QA and evaluation practice. Capacity is managed teams.
It depends entirely which DeepSeek you mean. The hosted service is operated from China under Chinese law, which for regulated data in the EU, UK or US is usually where the discussion ends, and several governments have restricted the consumer app on official devices. The downloadable weights are a different proposition: run them in your own account and region and nothing leaves your boundary. That is a decision for your data protection officer.
On some tasks, and your test set decides which. It is strong on reasoning, mathematics and code, which is exactly where a cheap model is most useful. What the frontier providers still hold is the surrounding package: picture and audio input, tooling breadth, support and the enterprise paperwork. Our usual recommendation is a tier rather than a replacement, with Claude or the OpenAI API taking the escalations.
Yes, with one practical caveat about hardware. The R1 reasoning weights were published under the MIT licence, which is about as permissive as licences get. There are smaller distilled versions built on Llama and Qwen bases, and those carry their base model's terms, so read the exact one you ship. The full-size model wants serious capacity, so most teams run a distilled version and we test at that size.
In the cases nobody looks at. A model that is usually right looks fine on a dashboard while a steady trickle of wrong answers reaches a system that trusts them. That is why the handover rule is measured rather than guessed, and why the escalation rate carries an alert. The saving is real. It becomes real once somebody is watching the cases nobody reads.
Three shapes, and which fits depends on how settled the scope is. Time and materials is billed hourly and quoted per project, which suits work still moving. Fixed price is outcome based, offered once the first read is done, because a fixed number on a system nobody has opened is a guess with a contract around it. Team extension is billed monthly per engineer. The rate depends on the seniority mix, so it is quoted rather than listed.
Is your AI feature too expensive to scale?
Tell us the monthly volume and what kind of data it is. You will get an engineer's read on where a cheaper model is genuinely safe, what the handover rule should be, and where the saving is not worth having.