On-Demand Webinar · 49 min

Rebel With a Data Test: The Solo Data Quality Playbook

You don't need a team, a title, or executive buy-in to improve data quality. One frustrated person, one customer, three fields, and a tool that writes the tests — Chris Bergh's playbook for changing data quality from the bottom up, an hour a week.

Presented by Chris Bergh

What you'll learn 6 points
  • A solo data quality leader is anyone in an organization trying to get value from data. You don't need a big team, a fancy title, or executive buy-in — you need influence, backed by knowledge of the data and a tool.
  • Your insider knowledge is the lever. You already know which data sets cause problems, which stakeholders complain, and which processes break most often.
  • The method is deliberately small: find one customer with a specific complaint, agree on three to five fields rather than the whole customer record, find the one technical person who can actually fix them, then repeat with the next customer.
  • Be concrete, and skip the abstractions. "We have to improve our data governance" moves nobody; naming three fields, a month, and a failure rate gives a busy engineer something to act on.
  • Measure it or it stays wishful thinking. The point of the dashboard is to turn "our data is bad" into three named fields with a tracked score — unknown unknowns become known unknowns.
  • Technology is not the expense; the meetings are. Profiling, 27 hygiene checks, 32 automatic test types, and 53 generated tests for one table arrive on a button click, which is what frees your time for the influencing.

Slides

47 slides

Transcript

Chris Bergh

Show chapters and dialogue 44 chapters · 8,170 words
  1. 0:00 Welcome, and who counts as a solo data quality leader
  2. 3:04 Agenda: a motivation, a method, and a tool
  3. 4:12 Why data quality falls through the cracks
  4. 5:23 You don't need a big team, a title, or executive buy-in
  5. 6:18 Insider knowledge is the lever you already have
  6. 7:23 Pick one customer and one specific complaint
  7. 8:28 The method: one customer, three fields, one person who can fix it
  8. 9:54 Build the tests and the dashboard, then send the issue report
  9. 11:06 Repeat the loop, and keep thinking small
  10. 12:09 Small wins, credibility, and word spreading
  11. 13:16 DataOps principles: small chunks and concrete asks
  12. 14:27 Measure it, or improving data quality is wishful thinking
  13. 15:42 Influence is a service you provide, not a favor you ask
  14. 16:48 Which roles can actually change data quality
  15. 17:53 Finding which part of the assembly line broke
  16. 18:58 The data quality leader and the ingestion team
  17. 20:02 Data engineers, analytic engineers, and DataOps engineers
  18. 21:01 If you are the data quality leader: persuade at the source
  19. 22:00 If you are on ingestion: poll and poke
  20. 22:54 Freshness, volume, and schema checks at the landing zone
  21. 23:52 If you are a solo data engineer: build the andon cord in
  22. 24:57 A virtual andon cord in your DAG
  23. 25:52 Quality gates before production, and where TestGen came from
  24. 26:53 Built for the person with no budget and no permission
  25. 27:57 Two products, and the two roles this session covers
  26. 28:52 What an LLM says it takes to build a data quality dashboard
  27. 29:47 Seven months of writing tests, or two months if you are fast
  28. 30:34 The process TestGen follows, and what it costs
  29. 31:32 Demo setup: the bike manufacturer shipping the wrong bikes
  30. 32:36 Installing TestGen and connecting a database
  31. 33:31 Demo: creating a table group
  32. 34:30 Demo: profiling finds a pattern inconsistency and possible PII
  33. 35:29 Demo: frame size has two different kinds of blank
  34. 36:33 Demo: quoted product names, and other hygiene issues
  35. 37:24 Demo: generating 53 tests with one button
  36. 38:34 Demo: running the suite and reading the results
  37. 39:37 Demo: tagging three columns for the operations stakeholder
  38. 40:38 Demo: building the operations dashboard in the score explorer
  39. 41:46 Demo: scheduling the suite and downloading the issue report
  40. 42:44 The greatest expense is meeting people, not technology
  41. 43:40 What is coming for data ingestion engineers
  42. 44:35 Monitor test suites: freshness, volume, and schema anomalies
  43. 45:33 Conclusion: spend an hour a week on influence
  44. 46:32 Q&A: document databases, stewards per field, and automating tests

00:00:00 Welcome, and who counts as a solo data quality leader

Chris Bergh: Hello. Hello everyone. I'm Chris Bergh from DataKitchen. We'll begin in a minute or two. Hello everyone. My name is Chris Bergh from DataKitchen. Thank you for joining today. So today's webinar is about called Rebel with a data test, the solo data quality leaders playbook. And so I am the head chef and CEO of DataKitchen. I've been in the data field for over 20 years and the software field for another 15 on top of that. and so we come to this with a perspective on how to empower people with data quality tools following DataOps principles. And so before we start we're going to record this webinar and I will share the recording and the slides either today or tomorrow and you'll get that in an email. And so unless you have your AI notetaker in going to this meeting with you in case you'll get that right away. So let me start. So what's today's agenda?

00:03:04 Agenda: a motivation, a method, and a tool

Chris Bergh: So in my mind I I have what what is a solo data quality leader? Well, it could be anyone in an organization that tries to get value from data. at the bottom of this, we're going to talk about different roles that people may have, whether they're they are their role in day job is data quality, or whether they're a data ingestion engineer, or a data engineer or a DataOps engineer. But we're going to go through three parts really, kind of motivation and a method and a tool. So, think of it as like why, how, and what's the lever that's going to get you there. and we'll and it so it should take about 45 minutes. Feel free to ask questions and put them in the chat window. You'll see on the right hand bottom of Google there's a second from the right there's an icon for chat. Feel free to drop questions in there. and I will try to answer those questions. and so yeah, if you have questions, just drop them in the chat window and I will try to endeavor to answer them.

00:04:12 Why data quality falls through the cracks

Chris Bergh: And so I don't see someone raised their hand, but yeah, put put your question in the chat window. So let's talk about motivation. So, you know, we all work in in organizations and organizations are great ways to get stuff done, but sometimes organizations can be frustrating. And so, data quality is especially a frustrating problem in many organizations because it falls through the cracks. It's a tragedy of the commons. And so you may continue to find errors in data or reports or models or hear com complaints from your customers about data reports or models. and you thought to yourself, why isn't anyone doing anything about our data? It's crappy and why isn't anyone helping? And you may have heard from your customers, I don't trust the data or I always have to check your data in Excel. and you may have heard them twice one day and you have to kind of hear excuses or make excuses for your team. It's frustrating. Or you come home at night and your wife or husband or partner or family member or friend is kind of rolling your eyes with you about the 20th time over a beer.

00:05:23 You don't need a big team, a title, or executive buy-in

Chris Bergh: You sort of b**** about crappy data at work for the 20th time. Now, does that sound like you? So, I'm going to give you a different perspective, right? one perspective is to just for the 21st time complain about it and sort of live with it. and and you know we have to do that to certain things in our organization that we have to live with. But I think there are cases where as an individual we can make a difference. And so that transformation can really start with you as an individual and you can yeah you can continue to debug broken pipelines at 2 am or you can kind of defend broken num numbers or you can continue to be frustrated but the point of this is to say that you don't need a big team or a fancy title or executive buyin. It has the the key thing that you need is influence. And influence starts with you sort of backed by knowledge of data and backed by a tool. And that's what we're going to talk about today.

00:06:18 Insider knowledge is the lever you already have

Chris Bergh: And so you don't need permission. and in my career, I've had a bunch of different jobs, right? I've started to work when I was 12, done all sorts of different role, different jobs in my career. And one of the things I've learned is that you can start to influence change just by your ideas and what you say and and what you do. And you have this hidden power. and and you may you know the data sets that are causing problems. You know the stakeholders and you know the or you know the processes that most break. And that insider knowledge is really your superpower for leading kind of a bottoms up data quality transformation. And so that's what this is about. and and we've built some software to help you, but we're but that we can't do the software without wrapping it in a process of influence and how you actually do this. And so what I'm going to do next is talk through a method. It's really a method of of going at this in a very small incremental step-by-step way to deliver small bits of impactful results and build upon it.

00:07:23 Pick one customer and one specific complaint

Chris Bergh: And so this isn't a method where it's going to be a big project or it's going to require lots of investment. It's start small, iterate, show value, and then go on and do it over and over again. And not something that you it's going to be require you to quit your job. It's sort of something that you can do part-time on the side as a way to improve things and and relieve your frustration. And so that's really the the the key idea is like just focus on one customer. and not every customer, not all the data, just one customer and sort of start small. Think from that customers. Focus on a very specific end customer who complains about the data. Like there's a marketer and they complain about email addresses all the time. Oh, focus on email addresses for the marketing customer. and think very specific pain points rather than trying to fix everything once. So pick specific a small number of specific data fields to work on.

00:08:28 The method: one customer, three fields, one person who can fix it

Chris Bergh: Don't work on the entire customer record or the entire customer database. Pick the customer email or the customer name or the customer address. And by fixing small things, very discreet things, it you can create visible successes that that will allow you to advocate for bigger changes. And when you do it, make it concrete and measurable. Don't be abstract like I'm not a big fan of the I've improved accuracy by 14%. Really, it's about h how have you improved a specific field and how have you made a customer successful? And so what is this process? And I'll we'll talk through this in a bunch of slides later, but just to walk through the process that we're advocating. So find a customer with a need, meet them by email saying, "Hey, what is your specific need and talk to them saying, look, if I could fix these three data fields for you, would that improve your life?" And they and often they'll say yes. And you know identify a small set of data elements three five don't try 50 or 100 don't try five tables start very small and and get buy in from them saying hey look if I work on this would you advocate for me to fix this would you go off against the people who actually own the data technically and then you have to get a person who can actually fix the data or improve the data

00:09:54 Build the tests and the dashboard, then send the issue report

Chris Bergh: And there's various roles of those technical people in the organization Sometimes they are called data engineers, sometimes they're software engineers. but if you're going to improve the data, you need someone to help improve the data itself and someone who can crack open a database or do more than that to fix the issue. And then once you've got it, you've got to actually create the rules that monitor the data. You've got to sort of profile and t test and create tests and a dashboard that shows the values. And we're going to talk about that and show how our free open source tool can you can use it and to do this. And I'll give you a demo of like in 10 minutes you can you can create all those things build all the rules and create tests and so and create a dashboard and then as the data changes over time you can spend send specific issue reports to the tech person to fix saying here fix this fix this table this column fix it in this way and then as things change you can monitor and improving as the fixes happen as the data happens and then you can communicate progress to your customers, your tech person and yourself and then you can see improvement.

00:11:06 Repeat the loop, and keep thinking small

Chris Bergh: And so really it's about and then repeat that process. Find another customer with another small set of data elements, meet with them, advocate technical person, build data quality tests, build a dashboard, communicate to the tech person what the fix is, monitor, see progress and go again. So, it's a very discreet, small process, not a big not what we're we're what I'm talking about here is a part-time job for someone to do. not a full-time job. And so, what does it mean to think small, right? don't boil the ocean. Don't get up in multiple tables, multiple things. There's a lot of problems, but start small. And if you choose one customer with a specific pain pain point like the marketing team or the fi finance team and and understand why it's going to make their life better and then creating small wins because you've helped one person with three fields and yeah there is it going to get you promoted? No. But it's going to help you actually do concrete things.

00:12:09 Small wins, credibility, and word spreading

Chris Bergh: So when Sarah starts tr trusting her conversion rates, she starts to become her advocate and then that credibility when the CFO stops questioning these things, word quickly spreads through the organization. And so a lot of organizations, they want data quality, but they haven't at the highest level decided that they're going to do it. they haven't had a CIO or a CEO saying look we're gonna we're going to target these these fields or data quality. So mo a lot of organizations are like that and so one way to affect and influence organizations is through small wins and credibility on your person by making it happen. And you know what? Maybe after you've done this, then finally the CEO is going to realize that this is a an issue that they can fix and is is tractable. and you can have a bigger program. But, the ideal customer profile for this presentation, and honestly, our tool is someone who is frustrated because that doesn't exist. and and that's the majority of organizations that we have today.

00:13:16 DataOps principles: small chunks and concrete asks

Chris Bergh: And so, as you can expect, we're DataKitchen. We're one of the main, proponents of DataOps and one of the founding people on it. So, we we believe that most work should happen in small chunks through highly automated means. So, start small, iterate quickly, and maximize your learning. And try not to be perfect at first. Try to be pretty good. And that means you learn and adapt. kind of perfect one thing. Start, you know, start small, learn, iterate, and improve. And also just make it concrete. Don't, you know, avoid vague things like we've got to improve our data governance. that that doesn't help saying I want to improve these three fields in the next month and be specific and and really provide clear actions because as you know if you you give a person a technical person saying this data field has got a 30% crap rate on it he'll go fine what do I have to do to fix it you've got to be clear because people who are going to do work that under your influence want you want to make it easy for them to do it and and so and then measure everything.

00:14:27 Measure it, or improving data quality is wishful thinking

Chris Bergh: Improve data quality without measurement is kind of wishful thinking. So you want to monitor it continually. You want to show about how you're doing this and you want to provide some visual proof points. Show a nice line graph showing the quality is going up. and that can actually kind of change how people think about you. Instead of you work with data, you're focusing on making business value. And remember that you're you're you're trying to turn unknown unknowns into known unknowns. This data quality is bad to data quality is of these three fields is in this dashboard. That's progress because you've made something that's vague and fuzzy actually concrete. And then influence, you can't do it alone, right? You've got to kind of get allies with you. the right customer a technical person kind of building your army and building people who because people appreciate leadership and if you are nice about it and if you're helpful and say thank you I think people are very happy to help you and I've learned a lot of techniques in my career going from a full-time person who stared at a screen writing code all day to how to get people to to work with you and how to influence influence people.

00:15:42 Influence is a service you provide, not a favor you ask

Chris Bergh: It's it's very satisfying. And and people you're you're not don't think of I'm asking them to do something uncomfortable. Think of you're you're you're helping them. You're you're helping them. And and leadership and influence is a service that you can use to help the organization. And so, you know, what's the bottom line? You can do it. You can get going quickly, right? sort of embrace the imperfection and try to find a small problem, get a specific customer, build some tests in a dashboard, iterate and improve and influence and don't consider this as a destination. and then, you know, make it concrete. That's really what the action plan is here. So, let me check and see if there's any comments, but we're going to keep going from there. I don't see any comments so far. so okay so there's lots of people in organizations right and so maybe you're have a data quality role or maybe you're a business analyst who's interested in data quality or someone who's just frustrated or maybe you actually have a day-to-day role with data.

00:16:48 Which roles can actually change data quality

Chris Bergh: Maybe you're trying to ingest data into a database or maybe you're building data pipelines to deliver facts and dimensions and reports to your customer or maybe you're supporting those those data engineers and data ingestion and data quality people. So I just want to talk through those roles and kind of where you can actually make change. So, one of the problems with quality is it's hard to talk about because if you look at the very end of this, I've got some data coming in on the left and I load it and transform it and predict and report and export. Your customer is always at the right and and the customer says something's wrong. This doesn't look right. And you've got this challenge of like, well, does it look right because the source data that came from the ERP system is crap? Or because the source data is pretty good, but we joined it with some other data and that data was not good or something happened, the transformation didn't work, or all that data is right, but someone misconfigured the report or misconfigured the export from the report.

00:17:53 Finding which part of the assembly line broke

Chris Bergh: And so you've got this customer at the end saying something's wrong. saying and a lot of times they say something's wrong and they phrase it as the data is the data is wrong. And you don't know what part of this whole assembly line where you've gone from raw data to value for the customer that it's broken. And so sometimes it is the actual very source data and that's the best place to fix it. But sometimes it's not. And so finding where those problems are is that along that assembly line is is hard. And so we're going to talk about various roles and where people work in relation to this assembly line. So and this this is in general true. There may be exceptions on how different people organize different roles. This is my view of the world. Different organizations and different titles. Unfortunately, they're not standard in the data and analytics world. So, but this is sort of my thinking. So, on the left here, we see the data quality leader.

00:18:58 The data quality leader and the ingestion team

Chris Bergh: People who in general are interested in data quality, what they really are interested is in improving the data quality of the source systems in the organization, not the extracted data that lands up in a warehouse, although sometimes that is the case. a a lot of times it is I, you know, we've got an ERP system. I want to improve the data in the ERP system. and so we're going to take that perspective here. I know it's not the entire world of data quality, what data quality leaders are, but we're going to take that as kind of our running orders for discussion today. Now, some organizations have a team that's solely focused on I ingest data. So, I've got 50 different data sets. I put it into a bunch of S3 buckets. Maybe I do a little work on it, maybe I don't. I put it in my L1 in my Medallion data warehouse, maybe my L2. They stop there. and so data teams often in bigger organizations, it manages kind of the ingested data and they don't really deal with anything beyond that.

00:20:02 Data engineers, analytic engineers, and DataOps engineers

Chris Bergh: They don't deal with kind of getting value from data. They're trying to get all the data into one place. and you know it's not always the case. you know but we're going to use that as a example for this discussion. And then in the purple here you could say a data engineer who's he or she may actually be involved in taking loaded data and extracting it and transforming it. Sometimes people call them analytic engineers. sometimes called data engineers, but they're sort of building pipelines supporting reports or or running production reports and exports. And then lastly, there's a role that a lot of teams don't have, which we call a DataOps engineer, which is trying to help people move new things into production. So, managing their environments, managing their test data, managing their deployment. And so we're going to talk about those roles as if you're a solo person and in one of those roles and how can you affect change. and so let's start with a data quality leader.

00:21:01 If you are the data quality leader: persuade at the source

Chris Bergh: So what does a data quality leader do? Well, they're trying to make sure that the information and databases and system is is correct, complete, up-to-date, etc. kind of by running audits and quality checks and then trying to get the organization to fix the data problems and trying to set some standards and and create rules for data management. And so how do you actually, you know, if you are this person, how do you actually achieve it? Well, you've got to persuade people at the source. and so that's what we're going to talk about and that looks from from this. It it doesn't mean that you're going to you're going to try to change the source here on the left. you're going to focus on it. you're not really going to worry about too much about, okay, there's a problem in the Power BI report that's sort of out of your domain, right? you're you're focused on finding and proving data at the point of origin and you don't really control much. and you know, this whole downstream set of errors, you're just not considering it.

00:22:00 If you are on ingestion: poll and poke

Chris Bergh: And that's fine, right? you you this isn't a what we're about here is just trying to focus on your world. and yeah, you may have perfect data, but the report will still be wrong. That's not something that you're going to go work on. So now let's let's say I'm in a different role. I'm that data ingestion person. I'm bringing in a whole bunch of different data sets into an L1 or L2 landing zone into a raw data. I've got maybe some pipelines that do that. I've got a lot of different servers and files that show up. what's the right way to do it? Well, I I like to call it polling and poking on ingest. So, you want to pull the data as it ingests and then poke it the people who have problems or will have problems with that data. And so, if you look at it, you're sort of polling and poking the data. You may not you're kind of periodically every hour, every half hour finding it.

00:22:54 Freshness, volume, and schema checks at the landing zone

Chris Bergh: So some reports may get into production wrong. That's you know you don't control the assembly line. You only control the raw material going into the assembly line. but you're looking at things like freshness and volume and schema and other checks to make sure it's right. And there are some gaps in test coverage that's for sure. But you're actually making sure all the adjusted data is good. And that that's a great help to organizations. you know, the problem is, you know, you could find problems after things things get into production before you notify or before it's fixed. and it's fairly simple to do this based on kind of polling a database or or or looking at data lineage for who who the problem is. Now, let's talk about if you're a solo data engineer. Well, what do you do? You build these end-to-end data pipelines. you monitor the performance of those and you work with your customers to change the facts and dimensions to build new aggregates to outs support or maybe run production reports.

00:23:52 If you are a solo data engineer: build the andon cord in

Chris Bergh: How do you actually improve data quality if you're a solo person of that? Well, I think the best way to do it is sort of as part of production. and you've got this assembly line. What you really want to do is sort of build the andon cord in the assembly line and and lean production. You actually in sort of the Toyota production system, everyone on the assembly line's got a cord that pulls down the andon cord that stops production. And you want to do that. You want to be able to go in and check data during the production process. Make sure that that as it goes through every step that every step is correct. and put monitors and steps as part of that production process. And that way you can ensure that you'll never have any reports or things in error because you'll have stopped it before it goes off. and you know it does it's more complicated because it requires some more data and tool integration. but it's the best way for you as a solo data engineer to start your work is to start putting those andon cords in put some checks as part of production check to see if those are right and then then go on to the next step.

00:24:57 A virtual andon cord in your DAG

Chris Bergh: And so that's I think from an ETL engineer's perspective building that virtual andon cord as part of your DAG as part of your workflow is the best way for a solo engineer to start. And then lastly is a a DataOps engineer and and you know like I said a lot of teams may or may not have that but they what do they do? Well they they deploy and manage code releases. So I've got some new ETL code. I've got some new SQL. I've got some new YAML. I've got a new report. I've got a new model. Those are there's they are in charge of putting making that path to production run like a train instead of like backpacking in the wilderness. But they're also and and they're maintaining that infrastructure to do that. The CI/CD, they're also sometimes involved in building test data or managing the systems itself. And then lastly, they're they're also trying to make sure that you have a train that can move your stuff into production easily.

00:25:52 Quality gates before production, and where TestGen came from

Chris Bergh: But they also want to make sure that there's signals on that train, that you have regression tests and quality gates so that you don't push broken code into production and that you can deploy and and that you're running tests in a development environment before things that get into production. And so, how can you help data quality there if you're a solo person? really you want to do it during that push to production. you want to be able to put some tests and regressions into your CI and CD and QA environment. and that way you can run those tests to make sure that they're they stop data poor data or poor code from getting into production. So let's talk about lastly about the tool. So a lot of this has been sort of based on sort of why we built the tool. And so when we started to build TestGen, it was kind of to scratch an itch. One of our engineers, Chip, was seeing the same pattern.

00:26:53 Built for the person with no budget and no permission

Chris Bergh: Other engineers were seeing the same pattern. We built this library of kind of reusable data quality tests that we were using over and over again. And so we realized that like this library of tests that could auto run was a was a good innovation that that we could help. And we started to build a UI on it. and we said who who is the real customer of this? And so in our mind our real customer is this is the person who this webinar is about a solo data quality person who isn't an or who's frustrated over beer and wants to make a change and doesn't have anything. So, it has to be no cost. It has to sort of run on their laptop, right? It's got to connect. It's got to be able to get them 90% of the way there with a few clicks. and there's no large ramp up. It's instead of having them give them YAML or give them lots and lots of forms to fill out, it automatically builds lots and lots and lots of data quality tests and then automatically builds a a dashboard on top of it.

00:27:57 Two products, and the two roles this session covers

Chris Bergh: And so that's the idea is you get influence superpowers by using this free tool. and so we also have more more tools. So we have one that's worried about the data which is our TestGen tool and one that's worried about the tools acting upon data our observability tool. So we have two products that you can start with and so you know these as we go to these roles and where these products align. I'm going to focus today just on two people the data quality engineer and the data ingestion engineer. We've got lots of demonstrations of how you use both products together, but we're going to focus today just on TestGen and and just on the data quality leader and the data ingestion leader. There's lots of information on other ones that we can talk about. So I'm a data quality leader. So the first thing is you know okay this method makes sense. I want to have some tests. I want to identify a few data sets.

00:28:52 What an LLM says it takes to build a data quality dashboard

Chris Bergh: I want to have data set a couple of data fields. you know I want to build some data quality tests and I want to have a dashboard that tracks progress. Well, how do I do that? Well, like if I here I asked ChatGPT or Claude to build me like how do I actually build a data quality dashboard and like it said it gave me this give me a project plan six to eight weeks to build one identify tool and stakeholders and there's a lot of stuff that's involved in here to build the dashboard out especially if you've got other team members who are more technical involved in doing that. Likewise writing tests themselves can be time consuming. So, here's a case of like let's say I've got I want to, do not what I said. I want to boil the ocean. I want to start with, 20 tables, each, you know, having 50 columns. So, I've got a thousand columns to do. I've got to generate 2500 tests.

00:29:47 Seven months of writing tests, or two months if you are fast

Chris Bergh: And so, let's say I hire a data engineer. It would take them and they're working. It takes about 30 minutes to to to write the test for them to figure out what it is. That's like seven months to build all those data quality tests. Okay? And maybe you've got a nice YAML framework or a nice UI. Okay? It doesn't take 30 minutes, it takes 10 minutes. Well, it's still two months of solid work to build all these tests. And so, that's just too long. If you're trying to get influence, no one's got two months to do this work. and even if they're very passionate about it and and, they're doing it on their side. and that I've talked to plenty of people who are building these systems on their own. They're trying to build a dashboard takes months. They're trying to build some data quality tests. It takes time. They're trying to assemble it all into a dashboard that they can track.

00:30:34 The process TestGen follows, and what it costs

Chris Bergh: And it just it's frustrating that they're building all this infrastructure that they don't need to do. And so that's really what we want is and is as is if you're going to do a solo data quality process, you need to start by learning your data, profiling it, screening for gotchas, generating data quality tests. you know, as the source data updates, you execute the tests, you generate a data quality score dashboard, and you review and refine the data quality tests, and you share issue reports. That's really the process that we we follow with TestGen. And so I'm going to kind of talk through that scenario today. The real cool thing about TestGen is it's free. It's has a UI, has a database. our enterprise version is very reasonable. Starts at which is doesn't have that many more features. It just has one great feature. It has a second database connection and a second user login. and so and we're going to talk through how that does today.

00:31:32 Demo setup: the bike manufacturer shipping the wrong bikes

Chris Bergh: And you know what's great about TestGen is it auto builds automatically builds 27 different data hygiene tests, 32 different automatic tests just with a click of button and about 10 custom tests and it runs really fast against your database. So I'm going to do a demo now and kind of talk through this scenario. Okay. So, there's an operations you work at a bike manufacturer, right? And you're a person who cares about data quality. And the operations guys been saying that we've been delivering the wrong bikes to our retail outlets. And he says, "Well, I've been doing exactly what's in the data." And you know, we're delivering the wrong stuff. The data is crap. And like so people we've been shipping wrong color, wrong size, the wheels have been mismatched and how can you improve this? So I'm going to show you in TestGen how you can improve it with by using TestGen. And so I'm going to share my screen.

00:32:36 Installing TestGen and connecting a database

Chris Bergh: I've already in installed TestGen and so it really is quite easy to install it. it takes about 10 minutes, five minutes. you it requires Docker on your machine, Windows Professional, or your Mac or a Linux machine if you're, and so we've got installers and install guides. So, when you install, we actually install a bunch of test data as well that you can play with. And I'm going to show you this test data. And so, but if you're, one of the things that you have to do is you have to make a connection. And so this connection's already here for us because we built this as part of our demo. So I I can show you the connection, but you've got to fill out these parameters to connect to your database. So maybe you have a SQL server database. So that's where the connection is and and we support a bunch of different database types that you can that you can connect to.

00:33:31 Demo: creating a table group

Chris Bergh: So the first thing that you do in it is is you go in and build some table groups. And so if I go in and I view the table group, I've got these tables and I'm looking at every different table in the database. And so in our data catalog in this, I've got four tables. And so from my standpoint, I'm only going to be interested in the product table because that we're shipping wrong products. So what I'm going to do here is create a new table group just to make it easier. I'm going to create a table group. I'm going to name it. all I care about is the ebike customers table. It's all I care about. I say this. And I think the schema's name is is demo. And so I'm going to hit next. I can see it. Can I have access? Yes. I hit next. So now I've got this new table group.

00:34:30 Demo: profiling finds a pattern inconsistency and possible PII

Chris Bergh: I'm going to save it and I'm going to run profiling. So what we're doing is we're executing we've got 51 different characteristics. So we're shooting a bunch of SQL queries at the database to understand the data. And so if I go to the profiling runs, I can actually see I've got this profiling data here. And so I can look at it. And what we've done is I've run hundreds and hundreds of queries. I've got eight potential issues in this one table. And so one case is there's a tax the pattern the tax ID ID there's a pattern inconsistency in the in the columns and so if you look at the source data you can kind of see there's some tax ID that looks different than you than than the other ones. That's interesting. And then it identifies some potential PII data. and so and what I realized is I did ebike customers and I did the wrong table. I want to go and actually create it.

00:35:29 Demo: frame size has two different kinds of blank

Chris Bergh: I want to actually go in and I don't want ebike customers. I want ebike products. So, sorry about that. Ebike products. And so, let me do this again. And so, it's so fast that I can just go in and do it. So, I hit next. it's right. I have access. Next. I'm going to start and run profiling. And so, yeah, this is a small database, but I can go in and do the profiling run. So, I'm going to go into ebike products. And so, I found three issues on that. And so, similarly, I've got frame size as an issue. Hm, that's interesting. So, people are complaining about getting the wrong size. Let me look at the source data. Huh, I've got some non-standard values where it says that there's nothing. You know, sometimes, you know, the frame size is supposed to be of a certain type. Like if I look at my profiling data, I can see the frame size has got, you know, small, medium, large, but it's got both NA and missing as blank.

00:36:33 Demo: quoted product names, and other hygiene issues

Chris Bergh: And that maybe that's the problem, right? There's some cases where these missing values and there shouldn't be missing, right? Those should all be filled out. A frame size can't be missing. It's a physical thing. and then I've also seen some things in product name. So, if I go look at the source data, I've got some products that are quoted, whereas in other cases, they they aren't quoted. So, if I look at the profiling data, yeah, I can see I've got some quoted values. Most of them are not quoted. and so that's a case where maybe that should be fixed. I don't know if that's part of the issue. and then, you know, product type. So, this is all sort of hygiene issues on the data. And so, that's great. I've looked at the data. These are first time I look at it. give it. But I can also go on and now I've got this table group.

00:37:24 Demo: generating 53 tests with one button

Chris Bergh: So I'm going to go into my ebike products and I'm going to add a test suite and I'm going to call it table group ebike products. I'm going to just call this the ebike products suite. So, I'm going to create a test suite, hit add. and so I built this and now I'm going to hit one button and I'm going to from all that profiling information, I'm going to generate a whole set of data quality tests on ebike products. And so what this does is it uses that metadata and builds all these tests. So if I go back and look at my test suites. So I've got this one. And so it's going to generate tests. So if I go look at my data catalog again, I can see that I've got the ebike products here. And I can see all that information in my data my data catalog. Now, in my test suite, I've generated a test suite. Oh, there we go. So, I've got 53 different tests.

00:38:34 Demo: running the suite and reading the results

Chris Bergh: So, I'm going to run the tests now. So, again, it used all that metadata. It used the semantic data model, all the profiling data, and it said these tests apply to these columns. And I didn't do anything. I just clicked a button, right? And so now I'm on what a couple of minutes into this. So I've already got my my test suite on on ebike products. I've got my test results and so I can start to see some some test issues. And so I've got I've got these and I can kind of see all the tests that passed and all the tests that that have failed. And so that's that's great. And so now I've got some tests and I want to identify I want to build a data quality dashboard just for the fields that I'm interested in. So if I go into my data catalog, my customer has really complained about if I go into ebike products, they've complained about things like frame size.

00:39:37 Demo: tagging three columns for the operations stakeholder

Chris Bergh: So, what I'm going to do is add a bit of metadata saying my my my stakeholder is the person who does operations. So, I'm going to they're really concerned about this. So, I'm going to hit save. I'm going to say this field's interested in operations. It's the frame size. It's the color. And so I'm going to actually add some metadata for stakeholder group is here. And then the wheel size. And so I'm going to go and say edit stakeholder group operations. Save. So now I've identified these three fields as the one I'm interested in. So I've got to to repeat, I've connected to a database. I've profiled it. I've generated a whole bunch of data quality tests. Now I want to build a data quality dashboard for them. And so here you can see our what ships with it. We've got a data quality dashboard. It shows the total score. It shows the the the score declining slightly.

00:40:38 Demo: building the operations dashboard in the score explorer

Chris Bergh: This is of our demo data. and you can kind of see the total score going into it. You can actually go in and look at each of the columns in this and what the score is affected. Now I want to build a new dashboard just for this. So, I'm going to go into my score explorer and I'm going to add a filter that says I just want to do my ebike products and I only want to do the stakeholder group for operations and I don't care about CDEs. So, I'm going to call this called my operations dashboard. Add the data quality dashboard. And now I'm done. So, save changes. go back. And what's interesting is I can go in and look at the columns and I can just see the three columns that I've I'm interested in, frame size, color, and here's my score. And so I've got a dashboard on these on where the problems are. So, as new data comes in, every time this runs, every time you hit run test, or you could take this test suite, and you could say here it is, and I want to add a test schedule for it.

00:41:46 Demo: scheduling the suite and downloading the issue report

Chris Bergh: So, I want to add a test schedule for I want to add a schedule for my test suite ebike products. I can say I want to run this every every three days at 2 o'clock and we've got a new UI here that creates the cron job for you. I can add the schedule and and do that and so that way I can have this run every time and I can check it and every time it runs this quality dashboard that I have for my operations is going to get updated. And one of the great things about this is I could look at, oh, there's a problem with the frame size. I can view it and look at it and then I can take this and actually download the issue report and I can send that to my person who can go fix. And so if I look at this issue report and I'll have to share this tab instead. We've got all the information that you need. it talks about what it is.

00:42:44 The greatest expense is meeting people, not technology

Chris Bergh: It talks about where we found it. It says, "Here's the time I did it." And it even gives the queries and the counts. So, as a technical person, I can take this and go, "Oh, yeah. I know how to fix this. I'm just going to make all those missing values blank or I'm going to make them all the same." And here you just email this to the person and you're off and ready to go. And so, you've got a process now where you've picked a you know, the greatest expense here in your time is not technology. It's not hours and days and weeks to build something. The greatest expense is actually in meeting with people and influencing people. And that's right. That's what it should be. You know, influencing people, meeting people, deciding on what tables you want to focus on. That's the hard part in the world. And so this technology that's all packaged up, all ready to go, all free for one user, all the tests come with it enable you to make this happen.

00:43:40 What is coming for data ingestion engineers

Chris Bergh: And so we're really excited about this for you to do it. And that's sort of if you're a solo data quality person. Now to kind of talk a little bit about what a data ingestion engineer does. They could do exactly the same thing as a data quality person. Run tests, schedule it. But we've got some new features here coming out in the fall that I just want to talk about that we're working on now and should come out in the next month or couple of months. and so you know what we've noticed is that a lot of customers want to run a test every hour or every 10 minutes. They want to monitor. They've got you know dozens of data sources that are landing at different times and they're that ingestion person and their goal is to find out if there's some problems in the ingested data. And so they're looking at things like is the data fresh? Is the table fresh? Did it get too much data?

00:44:35 Monitor test suites: freshness, volume, and schema anomalies

Chris Bergh: Did the schema shift? Did the quality drift on these periodic ones? And you could do that before in our product, but it just didn't work out super well. And so what we've done is built some UI around it. So we've got a nice UI that shows up in a test suite built having you able to build these monitor test suites very easily and being able to drill down on this and see it by table and be able to look and see if there's a freshness or volume schema anomaly and then be able to look at the history of these. and the idea here is that test suite that we generate being able to have you automatically be able to have a monitor test suite, add some tests to it, and be able to run this. And we're also working at kind of improving our ML anomaly detection capabilities. So, this is a great addition for that type of user who just wants to kind of set it and forget it and then watch it as the data it ingests.

00:45:33 Conclusion: spend an hour a week on influence

Chris Bergh: And so both these use cases I think for our our types of users I think are really helpful you know for the data ingestion leader and the data quality leader. So finishing up kind of our conclusion look our our view on quality improvement is like it's always going to be tragedy the commons and you know sort of bitching and moaning over your beer is great but it doesn't make a happy life. So how do you get a happy life? do some small things, right? Start with a single empowered influ individual and spend your time influencing people. Don't spend your time on technology. So, that's where the the and it's not a full-time job. You know, spend an hour a week on trying to influence data quality. You know, by focusing on one customer, by focusing on a couple of data fields, work quickly using the DataOps principles and the ideas we laid out here. Give specific concrete remediation actions. continually measure things and use some open source software to do it like ours.

00:46:32 Q&A: document databases, stewards per field, and automating tests

Chris Bergh: We think that that's a great way to do it or if you don't like ours, there's there's other other tools. and a whole bunch of links here where you can learn more. You can download TestGen just by just by going here and and we've got a nice download form for you to go and do it in this. And if you're interested more in the process and the ideas, we've got a great white paper. So that's that's all we have today in this webinar. Thank you. and so I'm going to look at some questions and go to the questions part. So I see one question. so does this work does the tool work for non relational document based databases like OpenSearch or MongoDB? no it does not. It only works on as I said these relational tables and column databases. So sorry about that. and then an attendee has a question. Is there a way to assign sources or elements to stakeholders stewards such as attach a name for a field? Yeah, and that's that there is. And so you know that's in our our data catalog function and then this sort of metadata that you can assign. you could assign. So for instance a you can edit it say this is a CDE or what source what transform level what business domain what stakeholder group etc. and let's see. I don't see anything else. So, yeah, and we're we're super we support our open source users. So, we have a Slack channel. give it a try. See if you like it. thank you for attending today and look forward to hearing you. oh, there is one question. Automating the tests. Can this be done? Yes, I I ask again. That's the whole point. Test generation, automated TestGen. That's what our our tool is about. Thank you much. Have a great day.

Machine-generated transcript, lightly edited: filler words removed, product and speaker names corrected, and audience members anonymised. Speaker attribution is as captured on the call; chapter times are scaled from the meeting clock onto the recording.

Questions from this session

What counts as a solo data quality leader?

Anyone in an organization who tries to get value from data — the session deliberately does not restrict it to people whose job title says data quality. Chris Bergh walks through four roles and what each can change: a data quality leader persuading at the source, a data ingestion engineer polling and poking on ingest, a data engineer building the andon cord into the pipeline, and a DataOps engineer putting regression tests and quality gates in front of production.

What is the actual process being recommended?

Find a customer with a need and email them: if I fixed these three fields, would your life improve? Identify a small set of data elements — three to five, not fifty. Get their agreement to advocate for you. Find the technical person who can fix the data. Profile the data, generate the tests, build a dashboard on those fields. Send specific issue reports to the person who can fix things, monitor as the fixes land, communicate progress, then find the next customer and repeat.

Why not just build the dashboard and the tests yourself?

Because of the arithmetic. Asked how to build a data quality dashboard, ChatGPT and Claude came back with a six-to-eight week project plan. And 20 tables at 50 columns is 2,500 tests: at 30 minutes each that is about seven months of work, and even with a good YAML framework or UI at 10 minutes each it is still two months of solid work. Nobody trying to build influence has two months.

Does the tool work for non-relational, document-based databases like OpenSearch or MongoDB?

That was an audience question at the end, and the answer was no. TestGen works against relational and columnar databases only — it profiles and tests through SQL run inside the database.

Is there a way to assign sources or elements to stakeholders or stewards, such as attaching a name to a field?

Yes, and it is the mechanism behind the scoped dashboard. In the data catalog you edit a column's metadata — whether it is a critical data element, its source, its transform level, its business domain, its stakeholder group — and a dashboard can then filter on that. In the demo three columns were tagged for the operations stakeholder, and the score explorer turned that tag into an operations dashboard.

Can the tests be automated on a schedule?

Yes — another audience question, and Chris Bergh's answer was that it is the whole point of the tool. Tests are generated rather than written, and a test suite can be given a schedule; the demo set one to run every three days at two o'clock through a UI that writes the cron job for you. Every run updates the dashboard for the fields that stakeholder cares about.

Where to go next