On-Demand Webinar · 42 min
DATAVERSITY Demo Day: Data Quality with TestGen and Observability
Gil Benghiat's February 2025 DATAVERSITY Demo Day session: where the waste in a data and analytics team actually goes, why DataOps starts with data quality rather than automation, and a walkthrough of the open-source TestGen workflow.
What you'll learn 5 points
- The first problem in data and analytics is waste — wasted time, energy, and trust. 60% of projects fail, 79% have too many errors, 73% of practitioners do not trust their data.
- The second is little boxes everywhere: data stack sprawl, organisational silos, and a dataset count that keeps growing.
- A day-one focus on immediate tasks and individual tools is what produces customer-visible errors. The day-two move is to optimise the whole system of people, tools, data, work process, and deliverables.
- Three places to start, in that order: improve data quality, stop production errors, then automate for faster and safer deployment.
- The open-source TestGen workflow runs in a loop — profile tables, screen for data hygiene gotchas, generate data quality tests, execute them on a schedule, review and refine, produce data quality scores — and shares the issue detail with data owners and engineers.
Slides
Questions from this session
How much time do data teams lose to bad data?
The survey picture is consistent across sources: Gartner finds 60 percent of projects fail, Eckerson finds 79 percent of teams have too many errors, IDC finds 73 percent of data practitioners do not trust their data, and DataKitchen found 78 percent of data teams are stressed enough to want therapy. The common cause is waste, meaning wasted time, wasted energy, and wasted trust, produced by a complex and fragile toolchain, poor data quality, too much work, and errors the customer sees first.
What is the DataOps TestGen workflow?
Seven steps run by one data quality person. Profile the tables, screen for data hygiene problems, generate data quality tests, execute those tests on a schedule, review and refine the testing, generate data quality scores, and share issue details with data owners and data engineers. Source data updates feed back into the loop, so profiling and test generation are repeated rather than one-time.
Where should a team start with DataOps?
In three stages. First improve data quality, which is where DataOps TestGen generates and runs tests. Second stop production errors, which is where DataOps Observability tracks the data journey from source to customer value across the whole estate. Third automate for faster and safer deployment, which is where DataOps Automation orchestrates, manages, and tests the toolchain.
What is the difference between TestGen, Observability, and Automation?
DataOps Data Quality TestGen does generative data quality: it profiles tables and produces and runs data quality tests without hand-writing each one. DataOps Observability is mission control for data journeys, anticipating and tracking production errors from source through to customer value. DataOps Automation orchestrates, manages, and tests a complex data toolchain to cut cycle time. TestGen and Observability are available as open source data observability software.
What is the difference between a Day 1 and a Day 2 focus in data work?
A Day 1 focus is on immediate tasks and on building with individual tools, which is how a team ends up with a fragile toolchain, poor data quality, and errors the customer finds. A Day 2 focus is DataOps: optimizing the whole system of people, tools, data, work process, and deliverables. Gartner's assessment is that teams guided by DataOps practices and tools will be ten times more productive than teams that do not use DataOps.
How does DataOps change how a data team spends its week?
Before DataOps, most of a team's week goes to errors, waste, and operational tasks, leaving a smaller share for new features and data for customers. After, the errors and operational share shrinks, the share going to new features and data grows, and a third block appears that did not exist before: process improvements and technical debt reduction. Less waste, more trust, and higher quality with fewer production errors are the stated results.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.