Skip to content

Snowflake engineering

A warehouse is where questions get answered, not where data goes to be stored. Snowflake does that job well. Its pricing model means a badly designed model shows up on the invoice rather than in a support ticket.

01Capabilities

Four things that decide whether it earns its credits

To place this against its neighbours: Spark is a processing engine, Databricks is a platform built around Spark, and this is the warehouse where the work is SQL. If your requirement is analytics on structured data, this is usually the shortest path. Four things separate a warehouse people trust from one they argue about.

  • Meaning

    Tables where a word means one thing

    Dimensional models, conformed dimensions and business-facing tables with agreed definitions. Loading source tables into a warehouse is a data transfer exercise. Deciding what a customer, an order and a return mean across four source systems is the actual work.

  • Loading

    Ingestion you can re-run without fear

    Batch loads, continuous ingestion, and transformation expressed as SQL in layers with tests between them. In e-commerce that usually means reconciling a storefront, a payment provider and a fulfilment system that disagree about one order, on purpose, for different reasons.

  • Cost

    Compute treated as a design variable

    Warehouse sizing per workload, aggressive auto-suspend, clustering only where it pays for itself, and finding the dashboard refreshing every five minutes for nobody. You are billed by the second while compute runs, which makes this an engineering decision rather than a procurement one. It is core data engineering.

  • Access

    Roles that survive an audit

    A role hierarchy someone can explain, row-level and column-level restrictions, and secure sharing with partners instead of emailed extracts. Worth designing before the third team asks for access. Built under an ISO 27001:2022 certified information-security management system.

02Fit

Buy it now, later, or not at all

Snowflake is strong at one thing: SQL analytics over structured data, with each workload given its own compute so one heavy query stops being everyone's problem. Buy it for that. Several of the reasons people give for buying it are better served by something else, including doing nothing yet. Two of the six rows below send you to no page of ours at all.

Six situations, and what we would tell you in each
Your situationWhat we recommend
SQL analytics over structured data, several teams querying at once, history kept for years Use SnowflakeSeparating storage from compute is the real advantage here, and the model is what decides whether that advantage reaches your invoice.
Operational reporting on one product database, at moderate volumes Not yetA read replica plus materialised views in PostgreSQL covers a great deal, and it is one fewer system to run, load and reconcile.
Model training, streaming, or transformation that is genuinely not SQL-shaped Different pageOur Databricks page, or self-managed Spark. Doing that inside a warehouse is possible and it is not what a warehouse is good at.
Already committed to one cloud, with its native warehouse on the table Compare properlyInclude transfer charges and the skills you already have. Snowflake often wins on ergonomics and the native option sometimes wins the total bill.
Something inside a product feature needs sub-second answers Wrong toolNot a warehouse at any price. That is an application database or a purpose-built serving layer, and a warehouse in that path disappoints everyone.
The real problem is that nobody trusts the current numbers Fix definitions firstA new platform loaded with the same ambiguous metrics reproduces the mistrust faster and at a higher monthly cost. Settle the definitions, then buy.

Scope

What we own here is the model and the economics. Whether the business definitions are agreed and enforced in one place. Whether loads are safe to run twice. What the monthly spend looks like once real query patterns arrive. We work under an ISO 27001:2022 certified information-security management system. unicrew has been building reporting systems since 2012, with 100+ senior in-house engineers across six countries. Where the honest recommendation is a read replica and three views instead of a warehouse programme, that is what you get.

The expensive mistake with a warehouse is loading everything first and deciding what the questions are afterwards. You get a bill that grows with your ambition and a set of tables nobody trusts. We would rather start from the three decisions the business actually makes on a Monday, model those properly, and let the rest arrive when someone can say what it is for.

Oleksandr TrofimovChief Technology Officer, unicrew

03Delivery

Agree the definitions before loading anything

These projects fail on definitions and on cost, almost never on the technology. The sequence exists to put both in front of the build.

  1. Start from the questions, not the source systemsThe decisions the business has to make, and the metrics behind them, written down and agreed. Loading everything first and modelling later is how you end up with a full warehouse and an argument about revenue.
  2. Set the layers and the contract tablesRaw, transformed and business-facing, with explicit contracts on the tables other people build on. Without that boundary every analyst queries raw data and no upstream change is ever safe again.
  3. Size compute per workload and suspend it hardSeparate compute for loading, transformation and reporting, right-sized independently, with short auto-suspend. This is the difference between a predictable bill and a quarterly conversation nobody can explain.
  4. Make loads re-runnable, tested and visibleRow count and freshness assertions, and alerting on staleness rather than only on failure. The dangerous failure is not the pipeline that stops, it is the one that keeps running and quietly stops updating. All of it through the same QA and test automation practice we use on our own builds.

04Stack

Where the numbers come from and go next

The stores that feed it and the platforms it is weighed against. Each has its own page if that is the decision you are actually making.

05Questions

Asked before signing a credit contract

Answered the way we would answer them live. Ask yours on the call and the answer will be specific to your setup.

Yes, and it works well because agreeing what a metric means is your side of the problem and building it reliably is ours. What we will not do is provide capacity while the model goes unowned, because an unowned warehouse becomes several hundred tables and no agreed definition of revenue. An architect outside the delivery team reviews the model. Capacity is managed teams; an outcome is data engineering.

If the question is SQL over structured data for reporting and analysis, Snowflake, because your analysts are productive on day one and there is no platform to learn. If you also need model training, streaming, or transformation that resists being written as SQL, Databricks is the better home. Either can be made to do most jobs, so the practical tiebreaker is whether your team is stronger in SQL or in Python.

A lot of teams buy one a year or two before they need it. If reporting runs against a single product database at moderate volume, a read replica with materialised views and decent indexes takes you a long way, and it is one fewer system to load and reconcile. The line is recognisable: several source systems joined, years of history kept, or analysts starting to affect the application.

Treat compute as the design variable it is. Separate warehouses per workload so dashboard queries do not size the loading compute, short auto-suspend so nothing idles while billing, and query-level cost attribution so you can see what is expensive. Then audit what runs: refreshes nobody reads, full reloads that could be incremental, and dashboards on a five-minute cadence serving a daily decision.

A scoping call with an engineer. For an existing warehouse we ask for the query history and a month of credit consumption, because that pair shows where the model and the spend disagree faster than any workshop. For a new build we want the reporting requirements and the source system list. You get back a written read on the model, the load design and the cost profile. Most engagements start within two to four weeks.

Three shapes, and which one fits depends on how settled the scope is. Time and materials is billed hourly and quoted per project, which suits work where the scope is still moving. Fixed price is outcome based and quoted per project, offered once the first read is done, because a fixed number on a warehouse nobody has opened is a guess with a contract around it. Team extension is billed monthly per engineer.

Standing one up, or explaining the credit bill?

Send the source systems, the reporting requirements and a month of billing. You get an engineer's read on the model and the cost profile, including whether you need a warehouse yet.

Book a scoping call

Thank you

Thanks for your message. We will get in touch with you shortly.