Market Guide for Data Lakehouse Platforms

23 September 2025 - ID G00838518 - 14 min read
By Michael Gonzales, Adam Ronthal,  and 2 more
The evolving role of data and analytics (D&A) in modern organizations has led to dramatic shifts in how data is stored, managed, and leveraged. At the center of this evolution stands the lakehouse — a converged data architecture that combines the strengths of data lakes and data warehouses. D&A leaders should use this research to learn how a lakehouse can serve as the organization’s foundational analytic data store, and to identify the features and capabilities needed to support a unified platform for a wide range of analytic workloads.

Overview


Key Findings

  • Data warehouses are rigid, making them poorly suited for adapting to rapid changes in business needs or handling diverse, unstructured data types like text and multimedia. While data lakes offer flexibility, they can easily become disorganized “data swamps” without proper governance, and they lack the performance and management capabilities found in data warehouses.
  • The lakehouse is now firmly established as the architecture that most organizations will seek to standardize on, either for new systems or when their existing Data Warehouses, Data Lakes, Logical Data Warehouses and other analytical systems are upgraded.
  • There are two broad approaches in the market by vendors and end users:
    • Buy a lakehouse solution from one of the vendors
    • Build a lakehouse by assembling existing query engines, data ingestion capability, AI/ML capability, and leveraging simple object storage
  • Lakehouse architecture still involves some trade-offs — particularly in terms of optimized data warehouse performance, where a dedicated data warehousing platform will almost always deliver superior results.

Recommendations

  • Lakehouse solutions are quickly maturing and are considered transformative. Consider including the lakehouse as a foundational analytic data store in your data strategy to support existing and evolving AI/ML capabilities to remain competitive.
  • When evaluating workloads for the lakehouse, make sure their performance, functionality, and service-level agreements (SLAs) are acceptable — or at least “good enough” — when compared to what competing architectures offer, such as data warehouses for business intelligence (BI) workloads and data lakes for data science workloads.
  • Ensure the solution qualifies as a lakehouse by meeting all mandatory requirements, with no critical elements omitted. For instance, some vendors may market a lakehouse built on proprietary storage instead of an Open Table Format (OTF). While proprietary formats can supplement OTFs, it is essential that the majority of data resides in an open, shareable format that is a fundamental characteristic of a lakehouse.

Market Definition


A lakehouse is a converged infrastructure design environment that combines the semantic flexibility of a data lake with the production optimization and delivery capabilities of a data warehouse. Data lakehouses are considered transformational and can serve as the foundational analytic data store for the organization. They are designed to unify the capabilities of data warehouses and data lakes into a single platform to support comprehensive data management and AI lifecycle.
The market for data lakehouse platforms comprises software vendors that must include the ability to:
  • Execute a broad range of workloads, from BI to AI
  • Persist data in simple object storage and leverage OTFs
  • Decouple storage and compute
  • Unify metadata and governance
The lakehouse reduces redundancy by providing a single infrastructure for processing, governance, and metadata management. Lakehouses are designed to support diverse workloads, such as data science and business intelligence, while ensuring strong data governance and unified management. By using open table formats, separating storage from compute, and managing metadata effectively, lakehouses simplify data architecture, reduce duplication, and improve cost efficiency and performance.
A lakehouse is a converged infrastructure design environment that combines the semantic flexibility of a data lake with the production optimization and delivery capabilities of a data warehouse. It supports the full progression of data from its raw, unrefined state, through the steps of refinement, to ultimately delivering optimized data for consumption.
The data lakehouse market is undergoing a major transformation, fueled by key trends that change the way organizations manage and analyze data. A major development is the integration of data lakes and data warehouses into unified platforms. This convergence simplifies data management, improves accessibility and supports a broad range of analytics within a single system. As a result, organizations benefit from increased agility, flexibility and efficiency when working with various types of data and workloads.
Furthermore, as digital systems generate ever-larger volumes of data, managing big data effectively has become increasingly important. While structured databases, data lakes and data warehouses have each addressed these challenges to some degree, the lakehouse has emerged as a leading solution for extracting insights from unstructured data in distributed environments. By combining the advantages of both data warehouses and data lakes, lakehouse architecture enables rapid data processing and integration, supports fast ingestion of unstructured data and allows transformation and analytics after data is stored. This architecture has attracted considerable attention from the big data research community because of its optimal functionality.

Mandatory Features

  • Converged design architecture: Unifies the architecture and workloads of a data warehouse and data lake on a single platform.
  • Persistent storage: Leverages simple object storage and is expected to be in an open table format (OTF) that may be complemented by other data types.
  • Unified data management: Multimodal data storage, schema flexibility and processing for a broad range of data types.
  • Data sources: Types of information sources that are utilized as inputs, including unstructured, semistructured, structured and streaming data.
  • Data ingestion: Collecting data from sources and transferring to the lakehouse, including batch ingestion, CDC, stream ingestion and file transfer.
  • Data catalog: Ability to identify and discover data objects, data governance, security, lineage and metadata management of information associated with data assets to enhance integration, access and utility across an organization.
  • Workload management: Ability to execute different workloads without conflicts and with acceptable availability and performance.
  • Query engine(s): Execute queries from one or more query engines that share the same metadata and physical assets of the lakehouse.
  • Data management: The lakehouse must ensure features like capacity planning, backup and disaster recovery are performed. This can be executed by the lakehouse or delegated to other services.
  • Data science/machine learning: Involves the application of predictive and prescriptive analytics methods to extract insights and build models.

Optional Features

  • DataOps: Agile and collaborative practices for improving data flow orchestration, integration, automation, observability, configuration versioning and continuous integration and continuous deployment (CI/CD).
  • Governance: Collection and implementation of policies, practices and data ownership to ensure data quality, adherence to regulatory and corporate requirements and efficient use.
  • Data sharing: Making organizational data and metadata accessible to various users and systems essential for AI readiness.
  • Natural language query: Using natural language to access and interact with data will become standard, making data consumption easier and potentially transforming how semantic layers, modeling and analytics tools are used. Users will be able to request reports, recommendations, and predictions in plain language.
  • Feature engineering: Involves creating new features from curated data to enhance model performance, requiring both domain expertise and technical skills; it is a resource-intensive, iterative process.
  • MLOps: Applies DevOps principles to machine learning, streamlining model deployment, monitoring, management and governance with continuous integration and delivery.
  • ModelOps: Manages the full life cycle and governance of all types of AI and analytics models, empowering business experts to interpret results and manage models, not just machine learning (ML) engineers.

Market Description


The optimization goals of the data warehouse and the data lake are different. The former is optimized for production delivery of semantically consistent, well-known data; the latter is optimized for semantic flexibility and rapid access to raw data. Early data lake practitioners frequently tried to deliver the optimization goals of the data warehouse within the architecture of the data lake. Unsurprisingly, more often than not, they failed.
However, the natural landing place for data of indeterminate business value was the data lake which provided an environment for skilled data practitioners to explore, refine, and make sense of new data sets. Where the data lake failed was in optimized delivery - something that the data warehouse excelled at. This required data movement (or promotion) from the data lake into the warehouse where sufficient optimization of data made it broadly consumable while meeting performance requirements. All of the optimization capabilities of the data warehouse — governance, data quality, data integration, schema design and BI reporting — were a means to that end.
However, having two separate platforms led to duplication of data, complex data pipelines and ETL processes, and added expense. The lakehouse addresses this by adding sufficient optimization capabilities to address the last mile delivery requirements of the data warehouse in a single, converged platform.
Mandatory features for this market include:
  • Converged Data Design — that unifies the architecture and workloads of a warehouse and data lake on a single platform
  • Persistent Storage — leverages simple object storage, most frequently supporting an Open Table Format (OTF).
  • Multimodal Data — for data storage, schema flexibility and processing for a broad range of data types, from structured to unstructured.
  • Support for Multiple Data Sources — Types of information sources that are utilized as inputs, including unstructured, semi-structured, structured, and streaming data.
  • Data Ingestion — Collecting data from sources and transferring to the lakehouse, including batch ingestion, CDC, stream ingestion and file transfer.
  • Data Catalog — Ability to identify, discover, govern, and secure data assets, as well as manage their lineage and metadata, to improve integration, access, and usefulness of information across an organization.
  • Data Sharing Allows multiple users, across various use cases, to access the same underlying data. This efficient approach is managed by strong metadata policies that protect data integrity and security. By providing unified access, it eliminates the need for duplicate data copies, reduces manual data management, and simplifies compliance requirements.
  • Workload Management — Ability to execute different workloads without conflicts and with acceptable availability and performance
  • Query Engine(s) — Integration with and support for third-party query acceleration engines (e.g., Trino, Athena, and Spark) that share the same metadata and physical assets of the lakehouse
  • Data Science/Machine Learning — Integrating AI and ML services into a lakehouse offers several strategic advantages for demanding AI and ML workloads. These benefits include scalable computing, a unified data architecture, support for agile development, and enhanced collaboration. Lakehouse solutions should either have built-in AI/ML capabilities or enable integration with third-party services.
Common features include:
  • Data Management — The lakehouse must ensure features like capacity planning, backup, and disaster recovery are performed. This can be executed by the lakehouse or delegated to other services.
  • DataOps Agile and collaborative practices for improving data flow orchestration, integration, automation, observability, data versioning, time travel, config versioning and continuous integration and continuous deployment (CI/CD).
  • Natural Language Query — Using natural language to access and interact with data will become standard, making data consumption easier and potentially transforming how semantic layers, modeling, and analytics tools are used; users will be able to request reports, recommendations, and predictions in plain language.
  • Feature Engineering — Involves creating new features from curated data to enhance model performance, requiring both domain expertise and technical skills; it is a resource-intensive, iterative process.
  • MLOps — Applies DevOps principles to machine learning, streamlining model deployment, monitoring, management, and governance with continuous integration and delivery.
  • ModelOps Manages the full lifecycle and governance of all types of AI and analytics models, empowering business experts to interpret results and manage models, not just ML engineers.

Market Direction


The lakehouse is the latest evolution of an analytic data platform. It has emerged to address significant new demands that remained unfulfilled by traditional data architectures, such as:
  • Data warehouses are effective at managing structured data and running predefined analytical tasks. However, warehouses are less flexible due to their dependency on data modeling, handling only structured/semi-structured data, and limited support for AI/ML workloads.
  • Data lakes emerged as a solution for storing large volumes of unstructured and semi-structured data, enabling new forms of experimentation and exploration. Yet, data lakes lacked essential features such as strong data governance, consistent performance, and reliable transactional capabilities.
  • Logical data warehouses (LDW) were introduced to logically integrate and unify separate physical systems. They often resulted in increased complexity and redundancy instead of simplifying operations
The lakehouse is the next progression that achieves the overall value of each architecture above while resolving its shortcomings within a single platform. The lakehouse is not considered a final solution, but rather a transitional architecture that sets the stage for more advanced systems, such as data fabric and integrated data ecosystems.
The key lakehouse technological components and innovations are focused on integrated infrastructure and modern data management capabilities and the democratization of data. The lakehouse technology stack brings together several key components to enable unified data management:
  • Converged Platform The lakehouse is a converged data platform, integrating the benefits and capabilities of both the data lake and warehouse architecture, achieving a unified data platform for analytics. It separates storage and compute, which is valued for performance and scalability, eliminates the need for multiple storage solutions, and simplifies data processing and governance.
  • Open Table Formats (OTFs) A defining feature of the lakehouse is the use of open table formats such as Apache Hudi, Apache Iceberg, and Delta Lake. These formats serve as an abstraction layer over files stored in object storage, elevate the metadata for all open file formats (Parquet, Avro, ORC) and bring critical data management features like:
    • Performance enhancements like partition management, indexing, and row skipping for queries
    • ACID transactions to track inserts, updates, and deletes
    • Time travel for querying and analyzing historical versions of your data
  • Metadata Store and Governance Layer The metadata layer acts as a data catalog and governance framework, managing metadata at both the table and row level. It organizes data assets and enforces detailed security policies, ensuring data is accessible, secure, and compliant with enterprise standards.
  • Unified Workload Support The lakehouse’s architecture natively supports a wide variety of analytical and operational workloads on a single platform, including data engineering, data science, machine learning, AI, and traditional BI. It accommodates batch, streaming, and interactive processing, removing the need for separate systems for different use cases.
At its core, the lakehouse integrates low-cost, highly scalable object storage with robust data management capabilities derived from open table formats. In doing so, it streamlines the processing of diverse data types and unifies various workloads — including batch, streaming, and interactive analytics — on one platform. This integration not only reduces data redundancy and movement but also provides a seamless environment for rapid data ingestion, processing, governance, transformation, and ultimately, analytics.

Market Analysis


A diverse group of vendors, both established and new, is entering the lakehouse market. These vendors have backgrounds in data warehouses, data lakes, query accelerators, and massively parallel processing (MPP) engines. Data warehouse vendors are enhancing their products by adding data lake capabilities for flexibility, like OTFs, while data lake vendors are integrating data warehouse features like SQL query engine support. Other vendors are broadening their portfolios by introducing lakehouse solutions, often leveraging open-source software alongside their existing offerings.
Most vendors promote their platforms as lakehouses, while others emphasize specific features that support lakehouse functionality, such as native support for open table formats in SQL engines. As a result, it is essential for data and analytics leaders to understand the different vendor approaches, how they compare, and which solutions best meet their organization’s requirements.
  • Vendors with data lake backgrounds: Databricks, Cloudera, HPE
  • Vendors with data warehouse backgrounds: AWS, IBM, Microsoft, Google, Snowflake
  • Vendors with query accelerator, MPP engine, or data virtualization backgrounds: AWS Athena, Denodo, Dremio, Starburst
A key feature shared by all lakehouse vendors is the use of open table formats on object storage.

Representative Vendors


The vendors listed in this Market Guide do not imply an exhaustive list. This section is intended to provide more understanding of the market and its offerings.
Vendor solutions for a lakehouse differ in terms of the scope of services provided as well as the emphasis of the workloads addressed. The vendors listed in the following table are only representative examples and should not be considered comprehensive. Moreover, many lakehouse vendors continue to evolve and enhance their solution stack that may change their scope and workload range.

Vendor Selection

The evolution of data architectures has accelerated trends in the marketplace that signal broad and growing vendor interest in providing lakehouse solutions. Data lakehouses have quickly become a compelling option for both new and legacy vendors.
Vendors are making significant investments to improve their lakehouse platforms, introducing new features and expanding their range of services. The following table highlights a sample of vendors that provide lakehouse solutions; however, it is not a comprehensive list.

Representative Vendors

Vendor
Product
Amazon Web Services (AWS)
SageMaker Lakehouse
CelerData
Celerdata Data Lakehouse
ClickHouse
ClickHouse Cloud Lakehouse
Cloudera
Open Data Lakehouse
Databricks
Databricks Data Intelligence Platform
Dell Technologies
Dell Data Lakehouse
Dremio
Dremio Data Lakehouse
Google Cloud
BigLake
Hitachi
Hitachi EverFlex AI Data Hub as a Service
HPE
HPE Ezmeral Data Fabric Lakehouse
IBM
Watsonx.data
IOMETE
IOMETE Data Lakehouse
Microsoft/Azure
Microsoft Fabric
OneHouse
Universal Data Lakehouse
Oracle
Autonomous Data Warehouse
Qlik
Open Lakehouse
Snowflake
AI Data Cloud
Stackable
Stackable Data Platform
Starburst
Starburst Open Data Lakehouse
Tencent
Tencent Cloud Data Lake Compute
Tensile AI
Tensile Data Lakehouse
Teradata
VantageCloud Lake
Source: Gartner 2025

Market Recommendations


  • Evaluate not only the system’s features but also its service levels. The platform must deliver on performance, capacity, reliability, and maintainability to support every aspect of your overall workload.
  • Confirm that the lakehouse supports the entire data management lifecycle, including archiving, backup and recovery, disaster recovery, throughput, load management, and concurrency for all planned workloads. These capabilities should be clearly defined and built into the lakehouse, not left for the customer to address. While some functions may be delegated to external services, the lakehouse must ensure these data management operations are dependable.
  • Verify that most interfaces comply with open standards, ensuring your data remains shareable.
  • Make certain the lakehouse provides robust support for data governance, metadata catalogs, data lineage, and security, with seamless integration across all components.

Note 1 Gartner’s Initial Market Coverage


This Market Guide provides Gartner’s initial coverage of the market and focuses on the market definition, rationale for the market and market dynamics.