Skip to content

Qwen model integration

Qwen matters for two reasons that have nothing to do with leaderboards. The family runs from models small enough for one modest graphics card up to large ones, and much of it ships under a permissive licence. Its language coverage is also wider than most Western models, Chinese especially.

01Capabilities

Where a model that fits and speaks the language wins

Qwen is a family of downloadable models from Alibaba Cloud. What this page offers is judgement about which member of the family to pick, which on this one is genuine work rather than a download, plus what we take on when one goes into production as part of AI and machine learning development.

  • Languages

    Documents that arrive in several languages

    Broad coverage, and notably stronger Chinese than most Western models. Useful for multilingual guest and customer messaging of the kind in hospitality technology, and for supply-chain paperwork that turns up in whatever language the sender uses.

  • Licence

    A model you can resell without a negotiation

    Many members of this family are released under Apache 2.0, which makes putting one inside software you sell a much shorter legal conversation than a community licence does. Check the exact one you ship, because the family is not uniform.

  • Size

    Something that fits the hardware you already have

    A small member of the family, tuned for one narrow job, running on a single card or a box next to the system it serves. On-site deployment beside warehouse and plant systems, as in our warehouse software work, is a common reason to want one.

  • Tuning

    A base worth training further

    Well supported as a starting point for teaching a model your own vocabulary, with coding and picture-reading members in the same family. A smaller base costs less to train and far less to keep running afterwards, which is usually the larger of the two bills.

02Fit

Picking a checkpoint, which is most of the work

Qwen sits in the same decision as Llama, and what separates them is not a score. It is licence terms, the range of available sizes, and which languages arrive in your documents. The thing to be clear-eyed about is sprawl: the family holds a great many versions across generations, sizes and specialisms, and choosing correctly is engineering. Four of the six rows below end somewhere other than a Qwen project with us.

Six situations, and what we would tell you in each
Your situationWhat we recommend
Content arrives in several languages, Chinese among them Use QwenTested against your actual documents. Western models are noticeably weaker here, and it shows up in production rather than in any comparison.
The model has to ship inside software you sell Pick an Apache 2.0 oneWith the licence of that exact version read and filed. The terms are attached to the version you deploy, not to the family's reputation.
English-only, mainstream work, and you want the biggest ecosystem Different pageLlama. More published recipes, more tooling that assumes it, and more engineers who have run it under load.
You want somebody else to host it, with a supplier to hold accountable Not this hosted serviceIt runs on Alibaba Cloud, which for most EU, UK and US buyers is a purchasing conversation before a technical one, see Claude instead.
One decision with four possible answers, and plenty of labelled examples Do not use a chat modelTrain one small model for that single job, see PyTorch. It will be cheaper, faster and easier to explain to whoever signs it off.
Nobody has read a licence and the model ships next month Stop and read itThis is a morning's work, and it is considerably cheaper than the alternative. Terms have varied by generation and by size across this family.

Scope

What we own on Qwen work is the choice of version and licence with the reasoning written down, the testing at the reduced size you will actually serve, the serving stack and how it behaves under a burst, the test set that proves a small model is enough, and an upgrade path across a family that moves fast enough to strand anyone who did not plan for one. unicrew has been building software since 2012, with 100+ senior in-house engineers across six countries, under an ISO 27001:2022 certified information-security management system. Where a rented model is less work for the same result, we will say so.

03Delivery

Choosing, licensing and shrinking a model you will run

This applies to Qwen and to any other downloadable family. Step two is the one people skip, and it is the one that ends up in front of a lawyer eighteen months later.

  1. Write down what rules out renting oneLicence, where the data may sit, how fast an answer must come back, working without a network, or cost per call. If none of those apply, renting is less work and we will say so rather than build you a platform you did not need.
  2. Shortlist versions and read their licencesGeneration, size, tuned or raw, and the specialised members. Terms differ across this family, so the licence that binds you is the one attached to the exact version you deploy.
  3. Test at the reduced size you will actually runQuality claims are made at full precision and production usually is not. We measure the drop on your own test set, because whether that trade is acceptable is a decision about your task rather than a general rule.
  4. Serve it properly and keep a rented baselineRequests batched as they arrive, a queue that pushes back, scaling that accounts for load time, plus the same test run against a frontier service on a schedule so the quality gap stays a number instead of an opinion. It goes through the same QA and test automation practice we use on our own builds.

04Stack

What a small model runs on

The pieces that turn up around a downloadable model, each with its own page if that is the decision you are actually making.

05Questions

Asked by legal before a model ships inside a product

Six that decide which version you are allowed to deploy, answered the way we would answer them live.

Yes, and here the useful contribution is choosing the version and the licence and getting the serving right, not familiarity with an interface. What we will not do is deploy a model into your product and leave the licence question and the quality measurement to somebody else, because both land on you later. An architect outside the delivery team reviews the design, and behaviour is covered by our AI QA and evaluation practice.

Licences, sizes and languages decide it. Qwen gives you Apache 2.0 on many versions, a wider spread of genuinely small ones, and stronger coverage outside English. Llama gives you the larger ecosystem: more published recipes, more tooling built with it in mind, more engineers who have run it. English work on flexible hardware favours Llama. Constrained hardware, resale or mixed languages usually favours Qwen.

For a lot of the family, and not for all of it. Terms have varied by generation and by size, with some versions released under a Qwen-specific licence instead. Since the licence that binds you is attached to the exact version you deploy, we record it in the architecture decision. If you are reselling the model inside your own product, your legal team should read it rather than take our summary.

Technically yes, and for many buyers purchasing will stop it. The hosted service runs on Alibaba Cloud, which for regulated data in the EU, UK or US raises the same questions as any non-domestic supplier, and those questions are usually settled before a technical evaluation starts. Running a downloadable version yourself removes the question entirely, which is much of why this family is interesting at all.

A scoping call with an engineer. Bring the task, the hardware you are stuck with, and whether the model has to ship inside your own product, because those three answers narrow the shortlist immediately. You get back a written recommendation on version, licence, size and serving, plus an honest note if renting one would be less work for the same result. Most engagements start within two to four weeks.

Three shapes, and which fits depends on how settled the scope is. Time and materials is billed hourly and quoted per project, which suits work still moving. Fixed price is outcome based, offered once the first read is done, because a fixed number on a system nobody has opened is a guess with a contract around it. Team extension is billed monthly per engineer. The rate depends on the seniority mix, so it is quoted rather than listed.

Need a model that fits your hardware and your lawyers?

Tell us the languages, the hardware you have, and whether the model ships inside something you sell. You will get an engineer's read on which version fits, and on when renting one is the saner answer.

Book a scoping call

Thank you

Thanks for your message. We will get in touch with you shortly.