The Five Use Cases in Data Observability: (#1) Data Quality in New Data Sources

The first of five data observability use cases: evaluating and cleansing new data sources for quality, schema, and anomalies before they reach production.

Written by Chris Bergh on May 10, 2024

DataOpsData ObservabilityDataOps ObservabilityDataOps TestGenOpen Source
The Five Use Cases in Data Observability: (#1) Data Quality in New Data Sources

Key points

  • Data evaluation of new data sources is the first of the five data observability use cases DataKitchen describes, alongside data ingestion, data production, development, and data migration.
  • Adding a new dataset to a production environment without evaluating it first risks data corruption, analytics built on faulty data, and decisions that harm the business.
  • Five challenges define the evaluation stage: profiling the data’s structure, content, and quality, validating the schema, detecting anomalies, confirming data types per column, and assessing freshness and relevance.
  • The anomalies that matter in a new source are specific — delimited data embedded in columns, leading and trailing spaces, pattern inconsistencies within a column, multiple data types under one column name, non-standard blank values, quoted values, and potential duplicates.
  • Profiling every dataset and producing a list of data improvement suggestions turns the patch-or-pushback question into a fact-based one: a data engineer can show the provider, or their manager, exactly what has to change before the data goes into production.

Ensuring their quality and integrity before incorporating new data sources into production is paramount. Data evaluation serves as a safeguard, ensuring that only cleansed and reliable data makes its way into your systems, thus maintaining the overall health of your data ecosystem. When looking at new data, does one patch the data? Or push back on the data provider to improve the data itself? And how can a data engineer give their provider a ‘score’ on the data based on fact?

NOTE

The Five Use Cases in Data Observability

Data Evaluation: This involves evaluating and cleansing new datasets before being added to production. This process is critical as it ensures data quality from the onset.

Data Ingestion: Continuous monitoring of data ingestion ensures that updates to existing data sources are consistent and accurate. Examples include regular loading of CRM data and anomaly detection.

Production: During the production cycle, oversee multi-tool and multi-data set processes, such as dashboard production and warehouse building, ensuring that all components function correctly and the correct data is delivered to your customers.

Development: Observability in development includes conducting regression tests and impact assessments when new code, tools, or configurations are introduced, helping maintain system integrity as new code of data sets are introduced into production.

Data Migration: This use case focuses on verifying data accuracy during migration projects, such as cloud transitions, to ensure that migrated data matches the legacy data regarding output and functionality.

The Critical Need for Data Evaluation

Adding new data sets to production environments without proper evaluation can lead to significant issues, such as data corruption, analytics based on faulty data, and decisions that may harm the business. To avoid these pitfalls, it is crucial to assess new data sources meticulously for hygiene and consistency before they are integrated into live environments.

Common Challenges in Data Evaluation

Data professionals often face several challenges when evaluating new data sources:

Key Data Evaluation Questions:

How DataKitchen Tackles These Challenges

DataKitchen’s approach to solving these challenges revolves around its innovative Open Source Data Observability. Leveraging DataOps TestGen, this platform offers automated profiling of 55 distinct data characteristics and generates 32 hygiene detector suggestions. This automation and detailed scrutiny level allows teams to identify and resolve data issues earlier in the data lifecycle.

The software facilitates a comprehensive review process through its user interface, where data professionals can interactively explore and remediate data quality issues. This not only enhances the accuracy and utility of the data but also significantly reduces the time and effort typically required for data cleansing. DataKitchen’s DataOps Observability stands out by providing:

Conclusion: Getting Facts ‘Patch or Pushback’ on your Data Provider (and Boss!)

By profiling every data set and coming up with a list of data improvement suggestions, DataOps TestGen can give you fact-based evidence to your data provider (Or your boss) on what needs to be done with a new data set before you can put it into production.

The quality of your data can determine the success or failure of your business initiatives. By implementing DataKitchen’s Open Source Data Observability, organizations can ensure that new data sources are ready for production and analysis. This approach saves time, reduces errors, and significantly improves the overall data quality within your production environments.

Next Steps: Download Open Source Data Observability, and Then Take A Free Data Observability and Data Quality Validation Certification Course


FAQ

What are the key points in this blog?

Data evaluation — checking a new data source before it enters production — is the first of five data observability use cases. The work is profiling structure and content, validating schema, detecting anomalies such as embedded delimiters and non-standard blanks, confirming types, and checking freshness. The output is fact-based evidence for the patch-or-pushback decision with the data provider.

What is data evaluation in data observability?

It is the first use case in data observability: assessing and cleansing a new dataset before it is added to a production environment. Evaluation covers the data’s structure, content, schema, types, and freshness, and it exists because a dataset that goes into production unexamined can corrupt downstream tables, feed faulty analytics, and drive decisions that damage the business.

What are the five use cases in data observability?

Data evaluation of new datasets before they enter production, data ingestion monitoring for existing sources, data production across multi-tool pipelines, development work such as regression tests and impact assessment when code or configuration changes, and data migration verification that moved data matches the legacy system.

What should you check before adding a new data source to production?

Profile it first, so you know its structure, content, and quality. Then validate that it matches the schema you expect, confirm the data types are consistent and appropriate per column, look for anomalies in the column values, and assess whether the data is recent and relevant enough to be worth loading at all.

What data anomalies should you look for in a new dataset?

The specific ones that break loads and skew numbers: delimited data embedded inside a column, leading or trailing spaces in values, pattern inconsistencies within a column, more than one data type under a single column name, non-standard blank values, quoted values, potential duplicates, and dates that are out of range or implausible.

Should you patch bad data or push back on the data provider?

Decide it with evidence rather than instinct. Profiling every new dataset and generating a list of data improvement suggestions gives a data engineer facts to take to the provider, or to a manager: here is what is wrong, here is how often, here is what has to change before this source goes into production. Otherwise the patching quietly becomes permanent.

Install Open Source TestGen Free, no vendor lock-in Request a Demo See TestGen Enterprise in action
Chris Bergh

Chris Bergh

CEO and Head Chef at DataKitchen. He is a leader of the DataOps movement and is the co-author of the DataOps Cookbook and the DataOps Manifesto.

LinkedIn →