AI and machine learning engineering
A model is a component. The work is the system around it, and that system decides whether an AI feature survives contact with your customers. What the model is shown before it answers, what it is allowed to touch, what happens when it is confidently wrong, and whether you can prove last week's change helped.
01Capabilities
Four shapes of AI work, and none is a chat window
Four shapes, and the route each one takes. Handing over the outcome is AI and machine learning development. Wiring a model into software that already exists is AI integration. Whether a given idea deserves a model at all is AI consulting.
Document and back-office work people currently read by hand
Contracts, statements, claims and case files turned into fields your systems can act on, checked against a fixed shape with a repair path when they fail it. It is what makes fintech and healthcare back offices expensive. The output has to be readable by a machine, not fluent.
Assistants and agents inside an existing product
A feature that answers from your data and calls your functions over several steps. What it is shown before it answers decides the quality; what it is allowed to touch decides the damage. The model matters less than the blast radius you allow it.
Speech, photos and scans turned into data
Recordings, condition photography and paperwork in whatever state the field produced them, the way logistics and transportation operations generate it daily. Turning that into text is the input stage. The value is in what the next step does with it.
Trained models where a hosted one will not do
Forecasting, ranking, damage detection, and the narrow repeated judgement where a small model beats a paid service on cost and speed. Training commits you to labelled examples and years of maintenance, so we weigh the case for it with you before anyone commits.
02Stack
Ten pages, and where we would actually start
One row per page in this cluster, with the honest reason to reach for it or not. Every row has exactly one destination, so the whole row is the link.
| Page | What that means in practice |
|---|---|
| Anthropic Claude | Long documentsReasoning across long, messy paperwork, and features that call your tools over several steps. On rule-heavy document work it frequently changes the answer.Claude integrationOr book a meeting |
| ChatGPT API | Shortest first buildThe widest tooling around it, worth real delivery time on a first production feature. Build behind an interface so the model stays a setting, not a rewrite.ChatGPT APIOr book a meeting |
| Google Gemini | Input that is not textReads photographs, video and scanned paperwork directly, which removes a whole text-extraction stage and everything that went wrong in it.Gemini integrationOr book a meeting |
| Whisper | SpeechTurns recordings and calls into text. Getting the meaning out of it is a second model's job, and keeping the two separate is what makes either fixable.Whisper engineeringOr book a meeting |
| Python | Where it all landsNearly every system on this page is written in it, including the parts with nothing to do with a model: the queues, the retries, the data plumbing.Python developmentOr book a meeting |
| Llama | Your own hardwareOpen weights, meaning model files you run on machines you control, for data that may not leave your estate. The largest ecosystem of the three here.Llama engineeringOr book a meeting |
| Qwen | Licence clarityAlso files you host, with clearer licence terms and genuinely small sizes, which is what makes an on-site or offline deployment realistic.Qwen integrationOr book a meeting |
| DeepSeek | Self-hosted onlyCheap reasoning, as weights you host. Its own hosted service is operated from China, which usually ends a procurement conversation early.DeepSeek integrationOr book a meeting |
| PyTorch | Only if you trainWhere a model trained on your own data is genuinely justified. Four conditions flip that decision, and they are in the questions below.PyTorch developmentOr book a meeting |
| TensorFlow | Inherited estatesWe would not start a new system on it. Where you already have one, keep or move is decided model by model, never by a rewrite.TensorFlow engineeringOr book a meeting |
03Fit
Model choice is rarely what decides the outcome
It is the decision teams spend the most time on. Evaluation and data quality are the ones that settle the result. So the first job of this table is to send you to the page that owns your case. The second is to name the right first step when it is not a language model, which is where two of the six rows below end up.
| Your situation | What we recommend |
|---|---|
| First production AI feature, an everyday text task, and you want the shortest path to something shipped | Hosted serviceThe OpenAI API, built behind an interface so the model stays a configuration choice rather than a rewrite. Breadth of tooling is worth real delivery time on a first feature. |
| Reasoning across long, messy documents, or a feature that calls your tools over several steps | ClaudeClaude, tested on your own documents rather than on a public ranking. On rule-heavy work it frequently changes the answer, and the comparison takes a day. |
| The input is photographs, video or scanned paperwork | GeminiGemini, which reads them directly and removes a whole text-extraction stage. Audio is a different model again: Whisper for the words, then a language model for the meaning. |
| The data may not leave your infrastructure, or a provider retiring a model would force you to re-approve a regulated process | Weights you hostLlama for the larger ecosystem, Qwen for licence clarity and genuinely small sizes, DeepSeek where cheap reasoning is the priority, as files you host rather than through its hosted service. |
| Rows and columns, a few hundred thousand of them, and a business number to move | Gradient boostingGradient boosting, which is a much older and much cheaper technique. It trains in minutes, it explains itself to whoever signs off the decisions, and on this shape of data it usually wins. |
| Nobody can hand you fifty real examples with the answers you would accept | Examples firstBuild the examples before you pick a model. That gap is the project, and every comparison run without it compares opinions. This is where AI consulting earns its place, before any build starts. |
Scope
Whatever the model, five things decide whether this holds up, and they are what we own. A set of real examples with the answers you would accept, so later changes can be measured. What the model is shown before it answers. A ceiling on spend, and a decision in advance about what happens when it is hit. What gets stripped out of the data before anything leaves your estate. And a second provider behind the same connection point, so one supplier is not a single point of failure. We work to ISO 27001:2022 and ISO 9001:2015, audited by Quay Audit UK.
04Delivery
Two things get settled before anyone names a model
The first two steps happen before any model comparison starts. They are the difference between a project that converges and one that spends a quarter proving nothing.
- Write down what a correct answer looks likeTwenty to fifty real examples from your data, each with the output you would accept, scored blind where the judgement is subjective. It takes an afternoon, and it turns every later argument into a measurement. Most AI features that quietly fail never had one.
- Prove it with a hosted service before training anythingAn API call is an experiment you can abandon on Friday. A trained model is an asset you maintain for years. So we prove the task is solvable at all first, then weigh whether volume, speed, cost per answer or where the data may sit justifies owning one.
- Decide where the data goes, and what the model may touchThis comes before the integration code, because it decides what your legal team has to sign. Credentials scoped to the job, read-only by default, rehearsals for anything destructive, and a human confirmation on writes and payments.
- Ship behind an interface, then watch what arrivesPinned model versions, cost accounted per feature, alerts on slow answers and refusals, and the examples re-run on every prompt or model change. What kills these quietly is drift in the inputs, so we watch what goes in.
05Trust
Two published figures, and where each came from
- 60%Less time on documentationKnowledge management platform, reported by clients (up to 60%)
- 30%Improvement in timely work-time logging complianceAcross the team, on our own internal proof of concept
06Case studies
Read both builds in full
AIRevolutionizing Knowledge Management powered with AISnaplore is unicrew's own product, built and operated in-house. It uses AI to transform how organizations document, structure, and share information, making meetings, training, and project discussions instantly accessible and actionable.Up to 60%Less time on documentation
AIAI-Powered Automation for Project ManagementAn AI bot that enhances operational efficiency by automating a critical, time-consuming, internal administrative task.30%Improvement in timely work-time logging compliance
We choose between Claude, Gemini, the OpenAI API and Llama by testing them on your own data. You bring twenty real examples with the answers you would accept. We run two candidates against them, and most of the choice is settled within a day.
Oleksandr TrofimovChief Technology Officer, unicrew07Questions
What comes up before an AI build gets signed off
There is no single answer, and the comparison is cheap enough that you should not accept one. Long, messy documents point at Claude. Photos, video and scans point at Gemini. A first everyday text feature points at the OpenAI API. Data that cannot leave your estate points at Llama or Qwen. Then run one set of real examples against two of them, on your own data.
In almost every commercial case, call the hosted service first. Four things flip it: sustained high volume of one repetitive judgement, a speed budget an API cannot meet, data that is not allowed to leave your estate, and a task narrow enough to describe precisely. When one genuinely applies, a small trained model in PyTorch is usually cheaper per answer. When none do, training buys the same result plus years of maintenance.
Rarely on model choice. Usually because nobody defined what a correct answer looks like, so there is no way to tell a good change from a confident one. Or because the data underneath is inconsistent, or means two different things in two systems. Third is a model wired into a product with no spend ceiling, no fallback and no evaluation. Measuring model behaviour is AI QA and evaluation.
You can bring our engineers onto your team, and on AI work that often makes sense, because your people know what a correct answer looks like. Our engineers own the evaluation, the data boundary and the cost ceiling, because those three are what turn a demo into production: an architect outside the delivery team reviews the design, and model behaviour goes through our AI QA and evaluation practice. Team extension is managed teams.
A scoping call with an engineer, and the most useful thing to bring is twenty real examples of the task with the answers you would accept. Those answer the questions that decide the design. From them we can usually tell which models are worth testing, and which design fits the problem. You get a written recommendation with the reasoning. Work usually starts two to four weeks after we agree scope.
Bring twenty real examples, and we will tell you what to build
Send the task, a sample of the real inputs, and what you would accept as a correct answer. You will get an engineer's read on which model to evaluate first, what that evaluation should measure, and where a cheaper design does the job better.