Internal Use Only

Quick Answer: What Is Data Observability?

ARCHIVED
8 June 2022 - ID G00759849 - 5 min read
By Ankush Jain
Data observability has now become essential to support as well as augment existing and modern data management architectures. D&A leaders should use this research to learn what the broader concept of data observability is and what capabilities will be needed to implement it.

Quick Answer

What is data observability?

  • The term “data observability” is an adaptation of the observability concept from the DevOps world, but in the context of data, data pipelines and data platforms. It considers data issues as an engineering problem and strives to empower data engineers in providing accurate and reliable data to consumers and applications within expected time frames. This allows IT and business leaders to have a degree of control over data usage and capacity planning.
  • Data observability uses automation to identify data quality issues, prevent downstream data issues, and augment performance management, capacity planning and production management. The most important underlying features to support data observability are “signal collections” to drive correlation and relationship analysis using data profiling, data monitoring and anomaly detection, active metadata, and data lineage.
  • Data observability provides a mechanism to understand data pipelines that are part of modern data architectures (such as data fabrics/data mesh with multiple proprietary and open-source components like orchestration, transformation, monitoring and metadata tools) with a complex mix of participating applications and multiple involved personas. Traditional data monitoring and data quality tools require manual steps and are limited to monitoring the data landscape against a standard set of expectations (“known known” scenarios).
  • Emerging data observability solutions provide end-to-end monitoring and leverage machine learning to predict unknown data issues and to assess impacts and root causes of failures. The purpose of data observability is to improve reliability of data pipelines by increasing the ability to observe changes and to cover scenarios where the issues and both the underlying causes could be unknown as well.

More Detail


This research is based on several interactions with end-user clients dealing with data outages and complex data engineering issues, as well as with over 30 vendors supporting data observability. Some of these vendors directly market themselves as data observability vendors, whereas many others are from adjacent data quality, data transformation, DataOps and data catalog markets.
Gartner
Data observability is the ability of an organization to have a broad visibility of its data landscape and multilayer data dependencies (like data pipelines, data infrastructure, data applications) at all times with an objective to identify, control, prevent, escalate and remediate data outages rapidly within expectable SLAs. Data observability uses continuous multilayer signal collection, consolidation and analysis to achieve its goals as well as to inform and recommend better design for superior performance and better governance to match business goals.
Data observability can take many forms. Practices and solutions available today primarily look at four things (the broadest we have seen in our conversations). See Figure 1.
Figure 1: Data Observability
The four forms of data observability include observing data, observing data pipelines, observing data infrastructure and observing data users.
  1. Observing data (independent of any other data dependencies) — Data issues could be broader than general data quality issues that look at accuracy and completeness, solutions target testing and monitoring of data for various metrics, anomalies and outliers also employing unsupervised algorithms. Solutions monitor data and associated metadata; changes in any historical patterns can trigger the alerts. This can be done top-down (context-driven) and bottom-up (data patterns/data fingerprinting and inferences from data values).
  2. Observing data pipelines — Data-pipeline-related metrics and metadata are observed to identify issues in transformations, events, applications/code that data interacts with. Any changes in data pipeline metrics and metadata around volume/behavior/frequency, etc. from the expected or predicted behavior can identify anomalies and trigger alerts based on change detections.
  3. Observing data infrastructure — Some solutions are also trying to combine signals and metrics from the data infrastructure layer as one of the dependencies of the broad data life cycle. Solutions can capture logs and metrics around resource consumption like compute, performance, underprovisioning and overprovisioning of resources that can target cost optimization concepts like FinOps, cloud governance, etc. Solutions monitor and analyze processing layer logs and operational metadata from query logs.
  4. Observing data users — Data observability is serving advanced personas like data engineers, data scientists and analytics engineers. The issues mentioned above are related to data delivery (also SLAs) and design of pipelines. Data observability focuses on predicting and preventing before these issues happen by analyzing available additional metadata (analysis and activation). These issues are beyond comprehension of business users, data quality analysts or business stewards; they are more upstream issues and, if they reach production systems, a lot of damage is already done. So far, interactions also suggest that data observability tools are better suited for streaming and real-time data needs where traditional data quality monitoring tools are very limited to even profile and understand data characteristics. This suggests a clear need for solutions that can rapidly observe and help in issue identification for streaming data needs.
See Table 1 for a list of the necessary capabilities for data observability.

Key Capabilities Needed for Data Observability

Supporting Capabilities
Connectivity to various data sources
Integration to data management, data pipeline and orchestration tools
Runtime deployment options and scalability
Data catalogs for search, semantics and metrics views
Usability (low-code/no-code UI and command line as well as support for engineering personas)
Signals Instrumentation and Analysis
Collection of metadata (definitional, query metadata, usage, infrastructure logs, configuration changes, drifts in volume/schema, etc.)
Data profiling and analysis (structure, semantics, anomaly and outlier detection, batch and real-time, automated discovery and analysis)
Correlation (statistical analysis) and relationship analysis of collected signals using AI/ML models for predicting and preventing data outages
Monitoring Capabilities
Data monitoring against standard rules, adaptive rules and automated change detections (nulls, typecastings, min/max, row counts, etc.), monitoring alerts, dashboards, etc.
Data pipeline monitoring for various drifts (schema, code, config) and design issues
Data infrastructure monitoring (compute consumption, config and infra drift)
Rule management (automated metrics, custom rules and test management using UI or support for coding in Python, SQL, etc.)
Root Cause Analysis and Collaboration
Detailed lineage graphs and impact analysis
AI/ML-driven outage analysis over historical patterns
Alerting, incident management, audit and support for SLAs for data
Workflows for team collaboration
Source: Gartner (June 2022)
Key recommendations for end users:
  • Partner with business stakeholders to evaluate and demonstrate the business value of data observability practices by tracking improvement of data issues management within data pipelines to show tangible benefits.
  • Identify the data pipelines that require high standards or SLA in quality, uptime, latency and performance. Pilot data observability principles by building a monitoring mechanism as a starting point to increase visibility over the health of the data.
  • Evaluate data observability tools to enhance your observability based on capabilities to fulfill business priorities, the needs of primary users and how the tools fit into the overall enterprise ecosystem.
  • Include both business and IT perspectives when evaluating data observability tools by engaging with both personas early on in the evaluation process.
Representative (not complete) list of providers:
Acceldata, Bigeye, Collibra, Databand, Datafold, DataKitchen, Kensu, Monte Carlo, Sifflet, Soda, Validio

Recommended by the Authors


Data and Analytics Essentials: DataOps