Software testing, test automation and the evals AI features need
Quality assurance is knowing whether software is good enough to ship, and there are four ways to buy it: testing run for you, an automated suite you own, a score for AI output that varies, or an outside read.
01Work
Four jobs that all get called testing
These four need different people, land at different points in a release, and come out of different budgets. Bought as one line item they compete for the same money, and it is not obvious from the line which of them is being cut.
Catching defects in the sprint that made them
The cheapest kind. A defect an engineer finds inside the sprint that made it costs a conversation; the same defect found by a customer costs a release, a support thread and some trust. That is what full-cycle testing buys.
Proving a release is safe to ship
The checks you re-run before every release, written down in a form a person can read before they exist as code, and giving the same answer twice. That is what test automation is for. On a seven-country e-commerce platform it meant automation, load and stress testing, and a test plan the client's own team runs now.
Judging an answer that is different every time
AI output breaks pass and fail, so it needs evals instead: real model output scored against criteria you set in advance, plus monitoring for drift, because the model you tested is not necessarily the model running next month.
Telling you the truth about how you test today
An outside read of coverage, process and risk, with every gap ranked by what a failure would cost. A quality audit is the one to buy when the internal argument about QA has stopped being productive, and the findings are yours whoever closes them.
02Services
Six services, and the situation that selects each one
Four of these are testing and two are security. The middle column is what selects them: find your own situation there first, then read what the work actually consists of. Every row is a page that goes deeper.
| Service | What you are actually buying |
|---|---|
| Software testing and QA | The developers are the only testersFull-cycle testing run by people who did not write the code: functional, performance, security and usability, manual and automated, sequenced by what a failure would actually cost you rather than by what is easy to test. Bought as an ongoing team or as one scoped cycle, and it runs on software we did not build as readily as on our own, including a real-estate platform in Japan.Software testing and QAOr book a meeting |
| QA and test automation | Every release is passed by handA suite in your own repository, built from cases written down before they are code and maintained as the product changes: the question this row settles is who owns it a year from now. On one SaaS platform, the time for an update to reach live went from weeks to hours after the refactor and the release automation.QA and test automationOr book a meeting |
| AI QA and LLM evaluation | The output variesEvals for features whose output changes run to run: real model output scored against criteria you set, plus whether the right sources were retrieved, how an agent behaves across a whole task, and whether quality drifts when the model changes under you. The hard part is agreeing what good enough means, and who is allowed to say so.AI QA and LLM evaluationOr book a meeting |
| QA consulting and audit | A second opinionAn outside read of how your team builds quality in, whether or not we wrote the software: what is actually covered, what nobody is watching, and every gap ranked by what a failure would cost. You own the findings and the roadmap, and closing them with us is a separate decision. Buy this before a rebuild, not after it.QA consulting and auditOr book a meeting |
| Security and compliance consulting | Buyers are askingISO 27001, GDPR, HIPAA and PCI DSS: a gap assessment, the controls that close it, and the evidence trail behind both, which is what an auditor or an enterprise buyer eventually asks to see. Because the people doing it are engineers, the remediation can be merged rather than handed over as a report to implement yourself.Security and compliance consultingOr book a meeting |
| Penetration testing | Prove it, offensivelyAn authorized attack on systems you name, run by people trying to get in rather than a checklist held up against the design, and the AI surface is in scope. What comes back is a ranked list of what to fix first. The right purchase when a customer or an auditor wants evidence rather than assurance.Penetration testingOr book a meeting |
03Fit
Where each kind of testing pays off, situation by situation
Confidence comes from testing placed where it gives signal. A suite people rerun until it passes gets fixed before it grows, an interface still changing shape every sprint is tested by hand while the layer underneath is automated, and a verdict that is genuinely a matter of judgement stays with a person. The other three rows each name the service that does the work.
| Your situation | What we recommend |
|---|---|
| A suite that exists, and that nobody believes any more | Fix the suite firstAdding to a suite nobody believes makes it worse, not safer. Fix the causes of random failure first, usually fixed waits, selectors tied to markup that keeps changing, and tests sharing data, then decide what is worth keeping. A smaller suite people act on beats a big one they mute. |
| Screens that change shape every sprint | Automate underneathAutomating an interface that is still moving means rewriting the tests each time it moves. Leave it manual until it settles and automate the layer underneath it, which is stable much earlier. |
| A verdict that is genuinely a matter of judgement | Keep it humanWhether a layout feels right, whether a tone is appropriate, whether a result is plausible to an expert. Automate-everything is a slogan; some checks should stay in human hands permanently. |
| An internal argument about whether QA is worth it | Buy the auditA quality audit is a scoped read rather than a programme, and it settles the argument with evidence instead of seniority. We build as well as audit, so the boundary is written down: the findings and the roadmap are yours, and closing them with us is a separate decision. |
| A stable release path passed by hand every time | Automate thatStable, repeated and expensive by hand is the profile worth automating. Begin with the path a release actually takes and the services under the screen, not with the newest feature. |
| An LLM feature about to go in front of customers | Evals, nowWithout a scored, repeatable measure the launch decision is whether the last demo went well. AI QA and evals, and it is also what makes a later model change a decision rather than a gamble. |
Who does the testing
ISTQB-certified QA sits inside every sprint, on every engagement. ISTQB is the international testing qualification, and it is held by the engineer rather than by the company. The engineers doing it are part of 100+ senior in-house engineers across six countries, at a company that has been building software since 2012, under the Artelogic name until the rebrand. We work as an independent QA partner on software we did not write, not only on our own builds. unicrew holds ISO 27001:2022 and ISO 9001:2015, renewed through a multi-stage audit with Quay Audit UK.
04Trust
What happens between getting in touch and buying anything
Four steps, in the order they actually happen, each one naming what it needs from you. The output of the first two is a written recommendation naming which of the four fits, and in what order.
- Read the evidence that already existsBuild history, the last few releases, and the test cases if anyone wrote them down. Where defects escape, and whether the ones that escaped were ever covered, shows up in that material and almost never in a description of it. From you: read access, and someone who can say what broke last.You getA read of where defects escape
- Say which of the four this is, in writingSometimes it is one of them, sometimes two in sequence, and sometimes the first step is pruning the tests you already have. The reasoning comes with the answer, so your own engineers can argue with it rather than take it on trust. From you: an hour with whoever owns the release.You getA named next step, with the reasoning behind it
- Prove it on one slice before it scalesOne release path, one feature, one dataset. Enough for you to see how we work, and enough for us to find what the build history did not show. Scope grows from what the slice shows, and the slice is priced on its own. From you: an environment, and data that behaves like production.You getA first slice, running
- Hand it over so your team runs itCases written down, the framework in your repository, and somebody on your side who knows how it runs. Quality assurance is an improvement when it keeps working after we leave the room, so the handover is part of the work. From you: a named owner, before the last week.You getWork your team can run without us
Before we add tests, we measure how often your existing tests change their result on an unchanged build. That one number decides whether the plan starts with making the suite reliable or with more coverage.
Ihor PrudyvusDelivery Director, unicrew05Proof
QA on our own builds, and on software we did not write
Real EstateQuality Assurance service for Japanese real estate platformunicrew provided QA consultancy for a leading Japanese real estate platform, which resulted in a tuned QA workflow and test automation.
SaaSWaiverKing: Software Development for Waiver Form CreatorWaiverKing is a document management business serving primarily the health and fitness industries. The platform is partnered with MindBody Online, one of the largest business-systems providers for gyms and yoga studios.Weeks to hoursTime for an update to reach live
eCommerceEcommerce software development for JewelCandleJewelCandle, a mid-size European manufacturer of scented products, sells across seven EU countries (B2B and B2C) via online shops.10,000-plusProducts manufactured
06Clients
Five clients on quality, and what it changed for them
Artelogic’s work had a very positive impact on our team’s morale. As our development quality was improving, our engineers were more confident in what they were doing, allowing them to work faster and with more confidence. As a result, our releases took less time and were less stressful.
Our platform has been refactored to Laravel ensuring the code base is more stable, easier to maintain and easier to add new features. Their professionalism and quality of work have stood out in the partnership. We have had the code audited by a 3rd party who was extremely complimentary of the work.
We were able to get the work completed in the expected time frame. There were little to no defects which was very nice because it allowed us to release and move on to our next project without having to back peddle. They were very accessible and took the time to understand our needs. They truly felt like part of the team.
We have had six out of seven on time and on budget project executions. Each engagement was a minimum of six months of effort. They had a flexibility and willingness to adapt their processes to match my requirements both on communication and the development processes. I never had a problem with any of their code from a code quality perspective.
The work is ongoing, but the impact that Artelogic has had on our development team has been substantial. They can work on developing other projects while still seeking guidance from Artelogic. Artelogic always provides our team with solutions backed up with the right research and analysis.
07Questions
What teams ask before changing how they test
Full-cycle QA: functional, performance, security and usability, both manual and automated, plus test automation wired into CI and AI QA and evals for features whose output changes run to run. It is run by ISTQB certified engineers and sequenced by risk, so coverage matches what a failure would actually cost you rather than what is convenient to test.
Yes, and it is a distinct service rather than an extension of the usual one. Output that is different every time breaks classic pass and fail, so AI QA and evals scores real model output against criteria agreed in advance, and covers whether the right sources were retrieved, how an agent behaves across a whole task, and whether quality drifts as the model changes underneath a live feature. The rest of the feature still needs ordinary testing, and that is the same purchase as anything else here.
Yes, and a good deal of this work is exactly that. On a real-estate platform in Japan, built by the client's own engineers, we set up a QA workflow, QA automation built on a tailored test-case system, and a testing roadmap the client can apply at every stage of development. QA consulting and audits stop earlier on purpose: an outside read of how your team tests today and a prioritised plan, with no build attached.
We scope before we quote, because the size of a testing engagement is set by how much is already covered and how much has to be built from nothing, and that is not knowable from a description. Billing is one of three shapes: time and materials, fixed price, or team extension, billed monthly per team member. There is no minimum engagement period. For test automation, the normal case is that we build the suite inside your repository and integrate it with the pipeline you already have. Most engagements start within two to four weeks.
ISO 27001:2022 for information security and ISO 9001:2015 for quality management, both renewed through a multi-stage audit with Quay Audit UK. Our QA engineers are ISTQB certified, the international testing qualification. When procurement asks for the certificates, they exist and they are current, which is the question behind the question.
What teams usually pair this with
All servicesSend the build history, or the feature you cannot sign off
Whichever you send is quicker to read than a description of it. What comes back is which of the four you need, in what order, and where to start.