A team moves off an expensive proprietary warehouse, excited about open formats and lower storage cost, and a few months later the dashboards that used to load in seconds are taking minutes. Analysts start asking why the new platform feels slower than the one it replaced. This is a common story, and it rarely means the lakehouse architecture itself was the wrong choice.
Adoption of this model keeps climbing. According to Dremio's State of the Data Lakehouse in the AI Era report, which surveyed 563 IT decision-makers in late 2024, 41% of organizations had already migrated from a cloud data warehouse to a lakehouse, more than from any other source system, citing cost efficiency and unified data access as the top drivers.
The architecture works. What usually breaks performance is how the migration itself was executed, which is where data engineering services come in.
Where Lakehouse Performance Actually Breaks Down
Most performance complaints trace back to a small number of decisions made during migration, not a limitation of the platform. Two patterns show up again and again once you look at the query logs:
- Copying the old table structure instead of redesigning it. Teams under time pressure often lift the schema straight out of the old warehouse and load it into the new environment without rethinking partitioning, file layout, or clustering.
A structure that worked well for a proprietary warehouse engine does not automatically perform the same way on an open table format like Iceberg or Delta Lake, because the underlying storage and query engines behave differently.
- The small files problem. Data that streams in continuously, or gets written in small batches across many partitions, ends up scattered across thousands of tiny files instead of a manageable number of larger ones.
The total data volume barely changes, but query engines have to open and scan far more files to answer the same question. Without a regular compaction and clustering process built into the pipeline from day one, this gets worse every week after going live, quietly, until someone finally notices the dashboards have slowed to a crawl.
What Good Migrations Actually Do Differently
The gap between a lakehouse that performs and one that does not usually comes down to whether these problems were addressed before go live or discovered after complaints started rolling in.
Engine selection matched to the actual workload.
Not every query engine performs equally well on every kind of workload. High concurrency, low latency dashboard queries behave very differently from long running batch analytics, and choosing an engine without testing against your real query patterns is one of the fastest ways to end up disappointed. TRM Labs, presented at Iceberg Summit 2025, offers a well documented example.
Working with more than 100 terabytes of data growing 25 to 45% annually, and a strict three second latency requirement for customer facing analytics, the team evaluated several engines before settling on the right fit for their workload. The result was a 50% improvement in P95 query latency and a 54% reduction in query timeout errors compared to their prior setup.
Ongoing file and table maintenance built into the pipeline.
Compaction, clustering, and partition management cannot be a one time cleanup step performed right before launch. They need to run continuously as part of the data pipeline, or the small files problem simply returns within a few months.
This kind of continuous tuning matters most for teams handling high volume, fast moving data, the same challenge that shows up in logistics and supply chain analytics, where shipment, inventory, and tracking data streams in around the clock. We covered this in more depth in Analytics in Logistics and Supply Chain: A Game Changer for Your Business.
A Migration Checklist Worth Running Before You Commit
A short audit before migration usually prevents most of these issues:
- Map expected query patterns and latency requirements before choosing an engine
- Build a realistic plan for file sizing and compaction, not just at launch but ongoing
- Assign a clear owner for table maintenance after go live, not just during the migration project
- Test the chosen engine against real production query patterns, not synthetic benchmarks
Why This Is a Data Engineering Problem, Not a Platform Problem
Most vendors selling lakehouse platforms will tell you the architecture supports both flexible storage and fast queries, and that is generally true. What they will not always tell you upfront is that reaching that performance requires deliberate engineering work throughout the migration, not just at the end. Teams that plan file layout, engine selection, and monitoring into the pipeline from the start tend to catch these issues during planning instead of during a stressful post launch fire drill.
Query performance also depends on what happens after the data lands. Dashboards built on a solid business intelligence layer surface these slowdowns early, before they turn into a pattern of complaints.
Frequently Asked Questions
Does moving to a lakehouse always cause performance problems?
No. A properly planned migration with the right engine choice, file layout strategy, and ongoing maintenance plan can match or beat the performance of a traditional warehouse. The performance problems people run into almost always trace back to migration shortcuts, not a limitation of the lakehouse model itself.
How does Ascend Analytics help teams avoid these migration mistakes?
Ascend Analytics plans migrations around actual query workloads rather than assumptions, building file layout, compaction, and monitoring into the pipeline from day one instead of treating them as an afterthought.
Is a lakehouse always cheaper to run?
Often, but not automatically. Storage costs can drop significantly, but if poor file layout forces excessive computation to compensate, some of those savings disappear into higher query costs. Cost and performance need to be planned together.
What is the biggest warning sign that a migration went wrong?
Query times that were fine at launch and steadily worsened over the following months, without any major change in data volume. That pattern almost always points to the small files problem building up because compaction was not built into the pipeline.
Should every organization consider outside help for this kind of project?
Not every organization needs it, but teams without deep experience running production lakehouse workloads at scale often benefit from working with a partner who has already seen these failure patterns and can plan around them upfront.
Is Your Lakehouse Migration Actually Built to Perform, or Just Built to Launch?
A lakehouse that launches on schedule but slows down every month afterward was not really finished on launch day. It was missing the ongoing engineering work that keeps performance stable as data keeps growing.
Ascend Analytics helps teams plan migrations around real workloads and build maintenance into the pipeline from the start. If your queries are getting slower instead of faster, reach out to us and let's look at what is actually happening under the hood.




