Skip to content

Apache Hadoop and moving off it

Hadoop is a platform in decline, and pretending otherwise would waste your time. If you have a cluster, the useful conversation is what it still earns and what it costs to keep. After that, how the exit is sequenced.

01Capabilities

Four kinds of work on a cluster somebody already has

We are not going to sell you a Hadoop platform. Cloud object storage plus a query engine solved the problem the cluster was built for, at a fraction of the operational weight, and the ecosystem moved. What remains is genuine, and it is legacy modernization: clusters still running something important, on a platform with a shrinking pool of people who understand it.

  • Inventory

    Working out what it still earns

    Every scheduled job, every table, and who or what reads each output. On a cluster that has run for years, a meaningful share of the jobs produce data nobody opens, and finding those is the cheapest part of the whole exercise.

  • Custody

    Keeping it safe while you plan

    Patching, capacity, Kerberos and access review on a platform you intend to leave. An exit takes quarters rather than weeks, and an unpatched cluster holding your operational history is a live exposure in the meantime. We do that under an ISO 27001:2022 certified information-security management system.

  • Rewrite

    Moving the jobs that survived the inventory

    Re-implementing what has real consumers, usually on Spark or in SQL against a warehouse. Hive queries translate more directly than teams expect. Hand-written MapReduce almost never does, and knowing which is which changes the estimate.

  • Landing

    Putting the data somewhere with a future

    Object storage with an open table format underneath, and a query engine chosen for the questions people actually ask. This is where the data engineering design work sits, and getting it right is what makes the switch-off uneventful.

02Fit

Keep the cluster, or leave it?

For new work, essentially never, and we would rather say that here than in month two. The value of this stack today is as a starting point you are leaving. The one honest exception is a hard on-premise constraint, and even then there are better options than the stack you are running. Two of the six rows below send you to no page of ours at all.

Six situations, and what we would tell you in each
Your situationWhat we recommend
Building a new data platform, whatever the volume turns out to be We argue againstObject storage with an open table format, plus Spark or a warehouse for compute. Same capability, none of the cluster.
An existing cluster, jobs with real consumers, and no forcing event yet Keep it, for nowPatched and running while you plan a staged exit. A big-bang migration off a load-bearing cluster is the most reliable way to lose a quarter.
You open the inventory and the data turns out to be a few terabytes Skip the clusterA warehouse, or in some cases a single large machine, will answer your questions faster and be operable by the team you actually have.
A hard on-premise or residency constraint that rules out cloud object storage Still not this stackThe one case for staying self-managed. Look at an on-premise object store with a modern query engine before renewing your commitment to HDFS.
Hand-written MapReduce jobs that are slow, and pressure to speed them up Do not tune themRewrite the ones with consumers on Spark and delete the rest. Optimising MapReduce in 2026 is spending money to stand still.
Your distribution is losing support, or the one person who knows the cluster is leaving Start nowThat is your forcing event and it should be treated as one. Begin the inventory while somebody can still explain what the jobs do.

Scope

What we own here is the exit plan and its risk. Which jobs move, and which get deleted. What the data layer looks like afterwards. The order that keeps the business supplied with its numbers throughout. We work under an ISO 27001:2022 certified information-security management system. unicrew has been modernizing systems like these since 2012, with 100+ senior in-house engineers across six countries. Where a cluster genuinely still fits, we will tell you to keep it and spend the budget elsewhere.

03Delivery

Sequencing an exit nobody can pause

The order matters more than the target platform. Steps one and two routinely cut the scope of the migration in half before anything is rebuilt.

  1. Inventory the jobs, the tables and the readersEvery scheduled job, every table, and who or what consumes the output. Outputs with no reader get deleted rather than migrated, and on a long-lived cluster that is usually a large fraction of the total.
  2. Measure the volumes honestlyReal sizes and real growth, separated from cluster capacity. Clusters get sized for a projection that never happened, and the true figure often changes the target architecture entirely.
  3. Move storage first, leave compute where it isData into object storage with an open table format, while the existing jobs keep running against it wherever they can. Decoupling the two early turns one enormous migration into a series of reversible steps.
  4. Rebuild, run both, then decommissionRe-implement the survivors, run old and new side by side until the outputs match on real data, and only then switch off. That parallel run is not optional, because reconciliation is the only proof anyone will accept, and it goes through the same QA and test automation practice we use on our own builds.

04Stack

What replaces the pieces you switch off

The engines and platforms these workloads move onto. Each has its own page if that is the decision you are actually making.

05Questions

Asked when a platform is on the way out

Answered the way we would answer them live. Ask yours on the call and the answer will be specific to your cluster.

Our engineers can join your team to keep a cluster running safely, and this is the one technology where we push back on the framing. Staffing a platform indefinitely means hiring into a shrinking skills market. Staffing it while you plan an exit is sensible and we will do it. An architect outside the delivery team reviews the plan. Capacity is managed teams; an outcome is legacy modernization.

Not dead, and not where new work goes. HDFS, YARN and MapReduce solved distributed storage and scheduling when cloud object storage was immature and elastic compute was barely available. Both changed, the vendor landscape consolidated, and what survived best were the file formats and query engines rather than the cluster itself. Plenty of clusters still run in production and will for years.

They are not the same category, which is why the comparison confuses people. Hadoop is a stack: storage, resource management, and MapReduce as the original processing model. Spark is a processing engine that runs on several of them and replaced MapReduce on merit. So there are two separate questions. Do you still need a cluster for storage, usually no, and what engine processes the data.

Three things, and none of them need a programme. Inventory the jobs and switch off the ones producing output nobody reads, which cuts cost and load immediately. Get the security position current, meaning patching and an access review, because that is the exposure you cannot defer. Then copy the data you care about into object storage. None of that commits you to a target platform.

A scoping call with an engineer. We ask for the job list, the table inventory and the cluster's size and version beforehand, because the first useful output is normally a shorter list of things that need to move rather than a migration plan. You get back a written read on what to delete, what to move, and what order keeps your reporting alive. Most engagements start within two to four weeks.

Three shapes, and which one fits depends on how settled the scope is. Time and materials is billed hourly and quoted per project, which suits work where the scope is still moving. Fixed price is outcome based and quoted per project, offered once the first read is done, because a fixed number on a cluster nobody has opened is a guess with a contract around it. Team extension is billed monthly per engineer.

Sitting on a cluster you would rather not have?

Send the job inventory and whatever you know about the data volumes. You get an engineer's read on the exit sequence, and an honest view of how much of it you can simply delete.

Book a scoping call

Thank you

Thanks for your message. We will get in touch with you shortly.