Skip to content

Anthropic Claude integration

Claude earns its place on two kinds of work: long documents somebody has to read closely, and rules that have to be followed every time. Two things come with it. You cannot run it on your own hardware, and it will sometimes refuse a request your business thinks is ordinary.

01Capabilities

Four jobs where a long window changes the answer

Anthropic sells a model, not a framework we install, so this page is judgement about model selection rather than a claim about tooling. Turning that judgement into a running feature is AI integration work. Four shapes of work come up again and again.

  • Documents

    Whole files read at once, not in pieces

    Contracts, policies, tender packs and case files, the kind that make fintech and healthcare paperwork slow. A large context window means the whole file goes in together, so nothing is lost at the seam where one chunk ends and the next begins.

  • Rules

    Work where the format is the requirement

    Redaction, rewriting to a house style, sorting against a long policy. The property that matters is obeying twenty instructions without quietly dropping three of them. That is something you test on your own rules, not something you read on a chart.

  • Agents

    Multi-step work that touches your systems

    The model plans, calls one of your functions, reads what came back and carries on. Anthropic originated the Model Context Protocol, an open way to describe a tool once and offer it to more than one assistant. Building one is MCP development.

  • Refusals

    Copy a safety filter keeps blocking

    Clinical wording, arrears letters, insurance decline notices. Models are tuned to be cautious, and cautious sometimes lands on text a regulator expects you to send. Designing the task so the model does the parts it will do is engineering work, not prompt wrestling.

02Fit

When a smaller model wins, and when no model wins

We resell no model capacity and hold no incentive in either direction, which is the only reason this table can say what it says. Claude is strong where the input is long and the rules are strict. It is the wrong answer at very high volume, and the wrong answer when the files may not leave your building. Four of the six rows below end somewhere other than a Claude project with us.

Six situations, and what we would tell you in each
Your situationWhat we recommend
Long documents, read closely, where a missed clause is the failure Use ClaudeIts strongest case. We still score it against one alternative on your own files, because your documents are not the ones any comparison was written about.
A long list of rules, all of which have to hold, every time Use ClaudeWrite the rules down as a scored test before choosing. Instruction following is the property that separates models here, and it is cheap to measure.
Twenty files a month, and somebody already reads them No model at allThe reading is a morning's work. Software to replace a morning costs more than it saves, and the first honest answer is often that there is nothing to build.
Millions of short, near-identical decisions, and the bill decides feasibility Something smallerA cheaper reasoning model, see DeepSeek, or one small model trained for the single job, see PyTorch.
The files may not leave your own network under any terms Different pageAnthropic publishes no downloadable weights, so a genuine air gap means Llama or Qwen on your own machines.
A safety filter keeps blocking copy your regulator expects you to publish Split the taskLet the model do the parts it will do and have a person write the sentence that trips it. Rewriting the prompt for a fortnight is not a plan.

Scope

What we own on Claude work is the comparison that justified the choice, written down and repeatable on your data; the retrieval and context design, which decides most of the quality once the model is fixed; the spending ceiling and what happens when it is reached; what the model is allowed to touch; and a second provider wired behind the same interface. unicrew has been building software since 2012, with 100+ senior in-house engineers across six countries, under an ISO 27001:2022 certified information-security management system. Where the honest answer is a cheaper model or no model, that is the answer you get.

03Delivery

Proving it handles your documents before you depend on it

The order matters. Two of these four steps happen before anybody writes a prompt, and they are what makes the choice defensible six months later when somebody asks why.

  1. Collect twenty real examples with the answers you would acceptTaken from your own files, including the three that always go wrong. This is the artefact every later decision is measured against, and it takes an afternoon to build.
  2. Score two or three models on that set, blindSame prompt, same inputs, scored the same way, with the model name hidden where the judgement is subjective. Public rankings are built on somebody else's task, so they narrow a shortlist and settle nothing.
  3. Design what goes into the window, not just the promptWhat gets retrieved, how it is ordered, what is cached and what is left out. Once the model is chosen, this is where nearly all the remaining quality comes from, and it is cheaper to change than the model.
  4. Pin the version, watch refusals, keep an alternative warmPin the model version, alert on latency and on how often a request is declined, and keep the second provider tested rather than theoretical. Every model is retired eventually, on the provider's schedule and not yours. The whole feature goes through the same QA and test automation practice we use on our own builds.

04Stack

What sits either side of the model call

The pieces that turn up in the same codebase, each with its own page if that is the decision you are actually making.

05Questions

Asked before a regulated team signs it off

Six that come up when legal, risk or compliance is in the room, answered the way we would answer them live.

Yes, and on this work the useful thing we bring is judgement about which model to depend on and how to contain it. What we will not do is drop a model call into your product and leave its behaviour unowned. An architect outside the delivery team reviews the design, and model behaviour is covered by our AI QA and evaluation practice. Capacity is managed teams; an outcome is AI and machine learning development.

You design around it rather than argue with it. Split the task so the model handles the parts it will handle, and route the rest to a person with the context already assembled. Log every refusal, because the pattern is usually narrower than it feels on the day. Where a whole category is blocked, that is a reason to test a second provider, not a reason to keep rewriting the prompt.

It can be consumed through a cloud provider, which keeps the traffic in your own account and lets you choose the region. That satisfies a great many procurement requirements. It does not satisfy an air gap, because the model itself is not something you hold. Where your rule is genuinely that nothing leaves the building, the answer is a downloadable model, see Llama or Qwen.

For a one-off analysis, yes, and it saves a lot of engineering. As a permanent design, usually not. You pay for every word on every call, the wait grows with the input, and material buried in the middle of a very long file gets less attention than people expect. Start with the whole file to get something working, then measure a retrieval design against it before the volume arrives.

A scoping call with an engineer. Bring twenty real examples of the task with the answers you would accept, because those settle in an hour what a workshop will not settle in a day. They also tell us quickly whether the problem is the model at all, since often enough it is the retrieval or the data. You get a written recommendation back. Most engagements start within two to four weeks.

Three shapes, and which fits depends on how settled the scope is. Time and materials is billed hourly and quoted per project, which suits work still moving. Fixed price is outcome based, offered once the first read is done, because a fixed number on a system nobody has opened is a guess with a contract around it. Team extension is billed monthly per engineer. The rate depends on the seniority mix, so it is quoted rather than listed.

Sitting on documents nobody has time to read?

Send us three of them and the rule you need applied. You will get an engineer's read on whether a model helps at all, which one to test first, and what the test should measure.

Book a scoping call

Thank you

Thanks for your message. We will get in touch with you shortly.