Skip to content

Apache Spark engineering

Spark is an engine for processing data that does not fit on one machine. Whether you need one is a question about your data volume, and we answer it from your own tables and runtimes.

01Capabilities

Four jobs that are worth distributing

Three pages in this section get confused with each other constantly, so here is the split. Spark is the processing engine. Databricks is a platform built around Spark with the operational and governance layer attached. Snowflake is a warehouse where the processing is SQL. This page is the engine, and the four shapes below are what we are hired to build on it.

  • Batch

    Transformation above single-machine scale

    Joining and aggregating across datasets large enough that one machine is genuinely the constraint. Meter and sensor history in energy and utilities is a typical shape: years of readings, reduced into something a warehouse can answer questions about.

  • Streaming

    Near-real-time processing with recovery you can explain

    Structured Streaming pipelines with watermarking, state handling and defined behaviour after a restart. Worth being precise about the limit rather than discovering it: this is micro-batch, so second-level latency is realistic and millisecond latency is not.

  • Enrichment

    The heavy work before the warehouse

    Deduplication, entity resolution, wide joins and quality checks that are awkward or expensive to express in SQL. This is the strongest case for the engine sitting next to a warehouse rather than instead of one, and it is core data engineering work.

  • Rescue

    Slow and expensive jobs, made to behave

    Shuffle volume, partition skew, small-file problems and clusters sized for the worst hour of the month. Most cost complaints here are data layout problems and they are fixable without rewriting the pipeline. All of it delivered under an ISO 27001:2022 certified information-security management system.

02Fit

One machine, or a cluster?

Distribution is a cost you pay in complexity before it returns anything. It is worth paying above a certain volume and a waste below it. The six rows below match each situation to the design that fits it, and the first is where Spark belongs.

Six situations, and what we would tell you in each
Your situationWhat we recommend
Datasets genuinely beyond one machine, with transformation logic too involved for comfortable SQL Use SparkThis is the case it exists for, and nothing in the ecosystem replaces it cleanly. Design around the shuffle and it stays affordable.
Your whole dataset is tens of gigabytes and one job touches a fraction of it PostgreSQL or DuckDBOne machine running PostgreSQL, DuckDB or plain Python will be faster, cheaper and debuggable by one person. Cluster overhead dominates at this size.
The transformation is expressible in SQL, and the data already sits in a warehouse Do it in placeSee Snowflake. Moving data out and back to run something SQL could have done adds a system and a failure mode for no gain.
The engine is right, and nobody on your side wants to run clusters or job orchestration Different pageOur Databricks page. Same engine, and you stop paying for it in operations time. Whether that trade is worth it is the argument there.
Per-event latency in milliseconds, with real event-time ordering requirements A stream processorThe micro-batch model has a floor and no amount of tuning moves it. A dedicated stream processor, or handling it inside the application, is the right design.
Existing MapReduce jobs on a cluster, with a mandate to modernise them Sequence it firstSpark is the usual target and the order matters more. Start from our Hadoop page, because the prior question is whether you need a cluster at all.

Scope

What we own here is the design of the pipeline and its economics. The partitioning and join strategy. Whether a re-run produces the same result. Whether the cluster cost is proportionate to the value of the output. We work under an ISO 27001:2022 certified information-security management system. unicrew has been building systems like these since 2012, with 100+ senior in-house engineers across six countries. When deleting the job and writing SQL is the better design, that is the recommendation you get in writing.

Spark is the right fit when data outgrows one machine, or when batch and streaming need one engine. We build jobs that fit on one machine without a cluster. On a cluster, we design the partitioning and joins first, which keeps it affordable.

Oleksandr TrofimovChief Technology Officer, unicrew

03Delivery

Measure first, then design around the shuffle

Step one is the one people skip, and it sets the scope of the other three.

  1. Measure the volumes before choosing the engineActual sizes, growth rate and the real query patterns. A surprising number of clusters process data that would fit in memory on a laptop, and the fastest performance improvement available is switching the cluster off.
  2. Design for data movement, not for codePartitioning, file sizes, join strategies and where skew will land. Almost every slow job is slow because of how much data crosses the network, so this is design work rather than tuning, and it is far cheaper before the pipeline exists.
  3. Make every job safe to run twiceDeterministic output for a given input window, safe re-runs and no partial writes visible downstream. Pipelines fail. The difference between an annoyance and a data incident is whether you can simply run it again.
  4. Attach cost and quality checks from day onePer-job runtime and spend, quality assertions on the output, and alerting on the failure that matters rather than on every retry. Each pipeline goes through the same QA and test automation practice we use on our own builds.

04Stack

What runs beside the engine

The platforms and languages that turn up in the same conversation. Each has its own page if that is the decision you are actually making.

05Questions

Asked before anyone provisions a cluster

Answered the way we would answer them live. Ask yours on the call and the answer will be specific to your workload.

Yes, and on pipeline work it is often the practical arrangement, because the knowledge of what the data means sits with you. Our engineers own the pipeline design as well as the code, because partitioning decides what the job costs to run every month. An architect outside the delivery team reviews the design, and every pipeline carries data quality assertions. Team extension is managed teams; an outcome is data engineering.

They are not competing answers to one question. Spark is the engine. Databricks is a managed platform with Spark inside it, plus notebooks, scheduling and governance. Snowflake is a warehouse where your transformation is SQL. If the work is SQL-shaped analytics on structured data, start with the warehouse. Most teams that name all three are describing a warehouse requirement.

Ask what the largest table is, and how much of it a job actually touches. Under a few hundred gigabytes, one well-provisioned machine running PostgreSQL or DuckDB usually beats a small cluster, and you debug with a stack trace instead of a job UI. Spark starts winning when the data no longer fits comfortably on one machine, or when you want one engine for both batch and streaming.

Good enough for most business cases, and its limits are worth knowing first. It processes in micro-batches, so realistic latency is seconds rather than milliseconds, and event-time windowing with late arrivals takes more care than the API suggests. If you already run Spark for batch, using it for streaming keeps one engine and one skill set. If you need per-event latency, a purpose-built processor is the better tool.

A scoping call with an engineer. We ask up front for data volumes, growth rate, current job runtimes and the cloud spend on the pipeline, because those four decide whether the conversation is about tuning, redesign or removing Spark entirely. What comes back is a written read on the design and the cost profile. Most engagements start within two to four weeks.

Three shapes, and which one fits depends on how settled the scope is. Time and materials is billed hourly and quoted per project, which suits work where the scope is still moving. Fixed price is outcome based and quoted per project, offered once the first read is done, because a fixed number on a pipeline nobody has opened is a guess with a contract around it. Team extension is billed monthly per engineer.

Sizing a cluster to the data you actually have?

Send the volumes, the transformation logic and the current runtimes. You get an engineer's read on the design, the cost profile, and the scale the job actually needs, from one machine up to a cluster.

Book a scoping call

Thank you

Thanks for your message. We will get in touch with you shortly.