Skip to content

Database and data platform engineering

The right data platform is sized to the questions it has to answer, and this page tells the cases apart. Sometimes one well-run database is the whole answer. Sometimes the questions have turned analytical enough to justify a second system built only for reporting. We will say which you are.

01Capabilities

Four shapes of work, from one database to a warehouse

Four jobs account for nearly everything we are asked to do here. The eight technology pages below carry the per-tool detail. This page owns the question that comes before all of them: how much you actually need to build. Delivery sits under data engineering.

  • The data model under a working product

    Ledgers, bookings, orders, clinical records. The rules that stop a half-finished operation recording money twice are set in the database, not in the application. That is why this is the first design conversation on fintech and healthcare systems.

  • Reporting that stops fighting the live system

    A read-only copy of the database, plus pre-computed summary tables, so reporting stops competing with the traffic that pays for the system. It is the cheapest fix here and the one that gets skipped on the way to buying a platform. We will also say when it has run out.

  • Pipelines and warehouse models people trust

    A warehouse is a second copy of the data, built for questions rather than for running the product. These projects fail on definitions and on monthly cost, almost never on the engine. So the effort goes into a model where each business number has exactly one meaning.

  • Estates nobody currently owns

    A database inherited from a team that left, a Hadoop cluster one person understands, a nightly export somebody repairs by hand. We inventory it, prove we can restore it, then say what is worth keeping. Where the answer is automation rather than migration, that is business process automation.

02Stack

Databases that run a product, platforms that report on it

Eight pages sit under this heading and they are two different purchases. The first four are databases that run a product, one record at a time. The last four are platforms that answer questions about that data, in bulk, after the fact. Most teams start with the first group and add the second when reporting outgrows it. One destination per row, so the whole row is the link.

The eight pages in this cluster, in the order we would consider them
Engine or platformWhat that means in practice
PostgreSQL Start hereThe default for a new product's database. It also handles loose, document-shaped data well, which removes much of the reason to add a second database beside it.PostgreSQL engineeringOr book a meeting
MySQL Keep, do not startIf a supported version runs your product today and nothing hurts at the engine level, moving off it buys risk rather than capability. We work on these estates rather than replacing them.MySQL database engineeringOr book a meeting
Microsoft SQL Server Microsoft estatesWhere identity, licensing and the rest of the estate already sit with Microsoft, that settles the choice before any feature comparison. The engine is rarely the problem here. The query plans and the reporting load usually are.SQL Server engineeringOr book a meeting
MongoDB Narrow fitGenuinely right when records vary a lot, are written whole and read whole, and nothing has to report across them. Where the reason given is that the shape of the data is unsettled, it does not solve that.MongoDB engineeringOr book a meeting
Snowflake When questions winThe warehouse we reach for once several teams query years of history at the same time. It is the point at which reporting has genuinely outgrown the product database.Snowflake engineeringOr book a meeting
Databricks Beyond SQLFor machine learning, streaming, and reshaping work that genuinely is not SQL-shaped. Buying it to draw dashboards is the common and expensive mistake.Databricks engineeringOr book a meeting
Apache Spark SpecialistProcessing spread across many machines, worth its complexity when the data truly does not fit on one. It needs a team that already runs clusters well.Apache Spark engineeringOr book a meeting
Apache Hadoop Plan the exitInherited estates only. Where one person understands the cluster and the data turns out to be a few terabytes, the way out is a warehouse, and the inventory that comes first often halves the migration.Apache Hadoop, and moving off itOr book a meeting

03Fit

Six situations, and three that start smaller than a platform

The useful question here is not which database. It is how much data platform your questions need, and we size it to them rather than to the brief. Half the rows below start smaller than a platform.

Six situations, and the answer we give in each
Your situationWhat we recommend
A new product where correctness matters, mixing reads, writes and some reporting One databaseOne relational database, and PostgreSQL unless something forces your hand. Its rules, its data types and the way it plans queries are strong enough that most products never outgrow it.
An existing MySQL or SQL Server database that works, on a supported version, with nothing hurting at the engine level Keep the engineLeave the engine alone. Spend the budget on the schema, the indexes and a restore you have actually performed, because moving a healthy database buys risk rather than capability.
Somebody has asked for a data platform and the whole dataset is tens of gigabytes A read-only copyA read-only copy and a few summary tables answer those questions, and one machine running DuckDB will beat a cluster at that size. Spreading data across machines costs complexity before it returns anything.
Several teams querying years of history, and the pain is everyone querying at once A warehouseSnowflake bills storage and computing power separately, so one team's heavy query stops being everyone else's problem. The record of what actually happened stays where it is.
Machine learning, streaming, or reshaping that is genuinely not SQL-shaped A lakehouseDatabricks, or self-managed Spark if you already run clusters competently. The choice between them is who operates the cluster, not what it can do.
Dozens of scheduled reports generated inside the application, and that is what actually hurts Process automationMove them out of the application. This is business process automation work, and re-platforming to fix it keeps the problem and adds a migration to it.

Scope

What we own is the shape of the data layer and its growth path. Four questions decide it. Do the rules in the database prevent bad states, or merely discourage them? Do the queries still work when the data is ten times bigger? Can a data load be re-run without counting anything twice? And can your team operate the result after we leave? Where one database and three summary tables answer the questions, that is what we build, sized to the problem rather than to a platform programme. We work to ISO 27001:2022 and ISO 9001:2015, audited by Quay Audit UK.

04Delivery

We measure the data before we name a database

Data work goes wrong in the same two places every time, the definitions and the running cost, so the order below front-loads both.

  1. Start from the questions, not the source systemsWe write down the questions the business needs answered and who asks them. Source-first projects produce a platform that contains everything and answers nothing, and the monthly bill arrives regardless. The questions also settle how much history you need, usually the largest cost nobody discussed.
  2. Measure the data before choosing the engineReal row counts, real growth rate, real query patterns. The gap between the data people describe and the data they have is routinely an order of magnitude in either direction. It decides whether spreading the work across machines is justified, or whether one index would have done.
  3. Agree the definitions, then enforce them in one placeActive customer, recognised revenue, delivered order. A platform is where disagreements between systems become visible, and the fix is nearly always a definition rather than an engine. Where those definitions cross a system boundary, that is platforms and integration work.
  4. Make loads re-runnable, and rehearse the restoreEvery load written so a re-run cannot count anything twice, alerts for staleness and failure that reach somebody who cares, and a restore that has been performed and timed rather than documented. A data layer is not delivered until the recovery path has been exercised.

05Trust

Platform growth, measured by the client

  • Around 2,000Companies on the platformWaiverKing, a document platform running on MySQL. Up from a couple hundred, in the CEO's words.

We start with the database you already run, because a well-tuned database often answers the business questions on its own. When reporting starts to slow the product down, we add a warehouse, so the business keeps one version of its data it can trust.

Oleksandr TrofimovChief Technology Officer, unicrew

07Questions

Straight answers before you buy a data platform

Most products can run reporting on the product database for far longer than they expect. A read-only copy of the database and a handful of pre-computed summary tables cover operational reporting, and that is one system to operate rather than two. It stops being true when several teams query years of history at once, which is when Snowflake earns its cost. If nobody trusts the numbers, fix the definitions first.

Considerably more than most people assume. One machine running PostgreSQL will handle tens of gigabytes faster than a cluster will, and it stays debuggable by one person. Spark earns its complexity when the data genuinely does not fit on one machine and the logic is too involved for comfortable SQL. If you already run Hadoop and the data turns out to be a few terabytes, the exit is usually a warehouse rather than another cluster.

Measure which queries actually hurt, then get reporting off the live system before moving any data anywhere. A read-only copy plus a few summary tables resolves a surprising share of these cases with no new platform. If the pain is dozens of scheduled reports generated inside the application, the cheaper answer is to stop generating them there, which is business process automation rather than a database project.

Our engineers can work inside your team, and that fits when your side owns the product decisions. They own the data model as well as the code, because model decisions are the ones that shape the next three years: an architect outside the delivery team reviews the design, and pipelines and migrations go through the same QA practice as our own work. Team extension is managed teams.

A scoping call with an engineer. We mainly want to hear what questions the business needs answered, what the volumes and the growth actually are, and what your team already runs in production. You get back a written recommendation with the reasoning, sized to the questions the data has to answer. Work usually starts two to four weeks after we agree scope.

Not sure how much data platform you actually need?

Tell us what questions the data has to answer and what your team already operates. You will get a straight recommendation with the reasoning, including the case where one database and a read-only copy is the whole answer.

Book a scoping call

Thank you

Thanks for your message. We will get in touch with you shortly.