On-Demand Webinar · 58 min
Data Quality Power Moves: Scorecards & Data Checks for Organizational Impact
What three dozen data quality leaders told us: they can find the problem but cannot fix it themselves. Chris Bergh and Chip Bloche on the influence-and-action cycle, and the scores that make someone else act.
What you'll learn 6 points
- The session reports on market research, not opinion: between June and August 2024 DataKitchen interviewed over three dozen data quality leaders worldwide, in manufacturing, financial services, consulting, and health care, at companies large and small.
- The question being tested was narrow and practical — can one person who has limited power but real influence use a free tool to cause meaningful data quality improvement in their organization?
- The blocker is organizational, not technical. Data quality leaders can find the problem but often lack the time, skill, access, or role to fix the data, so they have to decide where the change should happen, who should make it, and why.
- The numbers behind the problem: 57% of respondents in a 2024 dbt Labs survey rated data quality among the three most challenging parts of data preparation, up from 41% the year before, and 73% of data practitioners do not trust their data (IDC).
- There is no one perfect scoring model, so pick the scoring data sets that match the goal. A critical-data-element score serves the CFO's compliance case; a DAMA-dimension score serves the CDO's quarterly review; a model-data score serves the ML team. Critical data elements alone are not a panacea.
- Measure first, set standards later. The DataOps position is to start evaluating data quality before standards exist, then use the resulting measurements to establish and improve them — documentation and analysis are a result, not a prerequisite.
Prefer to read it? The written version is in Data Quality Power Moves: Scorecards & Data Checks for Organizational Impact.
Slides
Transcript
Show chapters and dialogue 10,107 words
00:00:00
All right, we'll start in a few minutes here. Just hold on.
There we go. There's Chip.
Hold on a second. Let's see.
Chip, you're on mute and I can't see now. Unmute you.
How's that? How's That? Oh, that's much better. I'm sorry.
So everyone who's joined, uh, give us a few more minutes, uh, and then we'll start.
So are things going today, chip
So far?
Two more minutes.
00:05:00
All right, everyone. So my name is, uh, Chris Bergh, and I'm here with, uh, Chip Bloche. Is it Blocher Block? You know, Chip Bloche, block? How can I like, not know that? Um, so welcome to our webinar. Um, and so just a few things. Um, my name is Chris, I, I'm CEO of DataKitchen.
And my partner in crime today is Chip Lock, who's, uh, everything data engineering and data quality at, at, um, and Chip, do you wanna give a little bit of your background? Yeah, I've, I, I, I've, um, been, I've spent my whole career doing data engineering and, and, uh, uh, I'm, uh, director of data engineering and DataKitchen and, and, uh, uh, vice president for bad kitchen analogies.
No, you, you're president of.
So, uh, slides and recording will be shared. Um, and, uh, put your questions in the chat window. Uh, it's on the lower right, it looks like a little icon there, and we'll answer questions at the end of the session, or we may answer 'em in the middle if we feel like it. Um, and we've targeted 45 minutes, uh, with questions at the end.
We've got, as usual, a lot of slides, um, and so we'd be, uh, we're looking forward to sharing with you. And so what are we talking about today? So, um, as you know, DataKitchen has sort of focused its work on DataOps, which is the idea of bringing agile and lean and, uh, iterative methods to all of data and analytics.
And, and we've been, been lately thinking about data quality, and we realized that we should just talk to some people and ask 'em some questions. And so what we're gonna do is, is sort of go through the market research and what we learned, talking to three dozen people, talk about how we synthesize that, and then talk about what we're doing to, um, change our, our open source data quality testing tool to make that happen.
Um, and so that's our topics today. So, uh, at first I, I guess in during the summer we talked to about three dozen data quality leaders worldwide. And, and they were in sort of big companies and small, um, they were in manufacturing, uh, regulated financial services, consulting, healthcare, you know, some were individual contributors and some had a lot of experience with data quality.
Um, most had some data quality in their title and, and some were in data governance or in data engineering. Um, but we wanted to learn more about the data quality, uh, challenge and get feedback on features that we should add to our open source data quality product. And, and as a background, we kind of had two questions, I guess, you know, could, we could have simply and efficiently empower a data quality leader who kind of has limited power, but influence with the ability to affect meaningful improvement in data quality in their, in their organization, by using our open source TestGen tool.
And, and specifically like, could one user do it? Could we get one user by themselves? Uh, some superpower, some skills using our open source tool. Um, and then the, the second theme is really, is there an sort of an agile or DataOps approach to data quality? That, that's sort of another theme because that's the lens that we look through everything, uh, in, in the world.
And, um, we had short interviews, uh, you know, we asked some generally open-ended questions like, what do you do and are you happy and why? And sort of how do you identify data quality issues? Um, once you find it, how is it fixed and who's responsible? Um, how do you measure progress? And then just some questions about like, how do you put stuff on?
Can you put stuff on your, on your machine, uh, to do work, uh, on your computer or how you actually go about getting software in? Um, and then we showed the slide, which is kind of, uh, the big question and in our mind is, is do scores matter and, and could you score data and use that as a way to effectively drive data quality improvement?
Um, and then, um, we sort of talked about this cycle, uh, in the research, and then we asked some questions about like, what we can call it, um, uh, just from a marketing standpoint. So those, those were our questions. Um, is a short market research really happy, had had great response. Uh, people love to talk, um, and I was very happy.
And, and Chip, do you have any comments on the research you were in? Almost all of them. It, it, it showed me what a hard job data quality people have that it's, it's, it's a, it's a, it's a challenging job at, uh, and it's, uh, really at this kind of nexus between multiple, uh, swim lanes and, and I'm sure we'll get into that later.
00:10:00
Yeah, I think that's, that's really it. Um, and, and so let, let me just define data quality first. Um, and because it means a lot of things. So, so how we're gonna use data quality and discussion, it's, it's kind of a comparison of the current state of your data with its desired state. And, and that desired state is based on what the people who are gonna use the data, the, their usage requirements, quality standards you have had, have.
And this idea of fit for purpose and conformance is the data that you have fit for its intended purpose. Um, and when the quality of your data is poor, obviously you can't use it for a lot of things, right? It's unfit for use and decision making, and, you know, good data results, good decisions come from high quality data.
Um, and so, and high quality information is simply good data evaluated in the context of your business and processes. Um, and, and, uh, as Chip said, there's lots of challenges in data quality. And, and I'll start off with some, some statistics. Uh, DBT labs had sort of a 16% increase in the number of people over one year rated data quality as one of their three most challenging aspects of data prep.
Um, 73% of data practitioners do not trust their data. Um, and then Forrester had sort of millions in loss of potentially billions, uh, with, without, uh, uh, without intervention and data quality. And, and we could go on, um, and, and really kind of improving data quality is not like a light switch. It's not like you go in a room and boom, data quality is better.
Um, kind of improving change is a, is hard. And as one person said is, is like data quality is good enough for the people who are putting the data in the system. It's often good enough for their direct need, and if they wanted to fix it, they would, um, but it's not good enough for other needs.
And so that's kind of the challenge of, of like data quality is like, it's good enough for the person who put it in, but it's not good enough for other uses. Um, and, and really trying to affect change and improve that as hard because, you know, people are, are busy and have a lot of things to do.
Um, and then lastly, we talked different ways to measure with people. You know, some are qualitative, um, some people had different methods, and then there's sort of, I think of 'em as kind of the standard dimensions or the DMA dimensions of the data quality, um, complete list, timeliness, consistency. And we, we'll talk about it.
We got some interesting comments on, on the data, data dimensions. Um, but one of the things really is the challenge of the person who's trying to actually make change. And I think they're optimistic. They believe that data quality matters. They believe that it can help their organization, but, um, the ability to find problems and they need more direct power to change the, in a lot of ways it's kind of data nags as opposed to, um, you know, people who can directly make change and, and so they have sort of influence but not power.
And that's really interesting in organizations. And something, uh, I've learned in, in my career, uh, kind of starting off as a, what in software many years ago, 20 year years ago, product manager, how do you get people to do things for you when you have no power to do it? And so I think influence is actually a whole set of skills.
Um, but influence also calms, I think, from being aligned with their, with organizational goals. So for instance, just because the fax number is blank everywhere and your data quality says that's null, don't, don't make people fix it because who uses faxes, num, nu numbers anymore, I guess if you live in Japan, they use faxes, but like, I don't think any other country uses faxes anymore.
And so how to, how to, how to use that influence, um, as a limited quantity and, and not expend that influence that you have on things that don't matter. That that was a theme that came up with a lot of people, like how, how to get, how to get your limited influence, uh, in there and then get the ball rolling of, of getting more influence over time.
Um, and, and lastly, I think data quality leaders, they're generally not data engineers. So, so they can find the problem, but they may not have the time or skill or access or role to actually fix the data itself. So, um, but they do need to decide sort of where to make the change. Like, is the change best in the source system, or is it best extracted from a source system in some kind of lake or warehouse?
Or does it have to be fixed at use? And then sort of they need to decide where to make the change, who should make the change, kind of why you should make the change and, and sort of what information the person needs to make the change. So this is, um, interesting that just the, almost all the people we talked to, maybe they know some sql, maybe they, maybe they have it, but they were almost all data quality people, but not kind of, um, skilled data engineers or data scientists.
00:15:00
Chip, do you have comments on that, that sort of, that fit with what you, um, what you learned? Uh, yeah, I think so. I think so. Okay. Um, so how, how did data quality leaders change things? Um, and so, you know, not just be frustrated. And so we talked and sort of thought about this process, and so here's my nifty PowerPoint, uh, DI diagram.
And so I think we validated this process with our market research, and it's really about empowering one person who has no power, but only influence and this sort of data quality, influence, and action cycle. And so it sort of starts with desire, right? Um, and then I think for us, our hypothesis is, is can we give them some tool that's easy to use and install and leverage?
And then the process starts with understanding data, finding data quality issues, generating relevant data scores, enabling other people to take action and measuring improvement over time. And I'm gonna go through each one of these blocks here in the next slide, but notice there's a circle here. And, and we'll, we're gonna talk about this sort of circle at the end, um, this sort of iterative process.
'cause that's something else that we, uh, that, that we believe in here. Um, and, and how that sort of DataOps approach to data quality fits in. So, um, mostly these people kind of don't have money and, and sort of don't have budget to buy a data quality tool, but they desire to do it.
We talked to one fellow who, who built his own in its in his spare time, or they use spreadsheets, um, or they have someone else that they've gotten into it. Um, and so this idea of an open source tool, free tool that someone in data quality can use as of interest to them. But there are challenges in installation.
Um, a lot of people they may not have. We talked a bit about sort of docker and getting machines and, and certainly no one, uh, wants to copy their data to somewhere really pushing down queries and, uh, to their databases important. So there is a, a desire and, uh, you know, some people, there's a lot of manual approaches, ad hoc approaches to kind of make this make data quality improvements.
Um, and, and then the second part is really to kind of starting with understanding data. And there's lots of different ways to understand data, you know, looking at it, sampling it, looking at spreadsheets, and of course there's, there's, uh, data profiling and there's lots of tools that do profiling our open source does one, but it, it, it really enables kind of a fact-based discussion of like, what's in the data.
Um, and it's kind of a, a, a, an excellent starting point for you to think about what should be, um, so sort of, uh, measure first test, test second, um, and, and can get you a good sense of what's in your data. Um, and I think that's a, that's an important part of, of what we've, uh, that can help people move forward.
And then the, the third is really about finding data quality issues. And so, um, in our last webinar series that we did kind of, uh, in the spring, we had a sort of data observability and data quality testing series. And, and we talked through sort of places to find quality issues. You know, in one case it's in, in the, the data itself.
And in the other case it's in, in the tool and find basically where to find problems. And so what we're really talking about here is this initial step, the data quality evaluation. Um, you can also do that during, uh, data ingestion. You could also check integrated data per production and development into data migration projects.
There's lots of places that you need to find data quality issues, but one thing is to improve data quality. There is a step sort of on initial, uh, evaluation of data to, to, to load it. Um, and maybe this, this red box should also be extended to data ingestion and change that might be, um, so, uh, in this certification series, I won't talk into it just sort of an ad that, uh, we've had a lot of people sign up.
Um, it really sort of talks through, uh, how to do, um, uh, data observability and data quality testing. So, um, the second need is really how to do data quality validation testing. Well, so starts with understanding the data, and then it sort of starts with sort of making a set of rules or having a set of algorithms make those rules for you.
And that's a challenge, right? Because number one, data teams lack the energy, skill or context to create those data quality rules. And yeah, there are some data quality people who are very adept at sql who can do it, or very adept at Excel, uh, and, and can do those things. Um, but it's a lot of work.
In fact, the, as the data size increases and the number of uses of data increases, it gets harder and harder, right? It's not just, you know, you don't wanna boil every data point
00:20:00
or even, um, data points. There are data points that are related to data science projects, data points related to, um, uh, sort of critical data elements. There's lots of things that you need to, to be able to do, and being able to actually validate those is a lot of work. And so is there, uh, a a way that we, you can automate that.
And, and so we've, uh, automated based on these sort of four groups, hygiene screen testing, anomaly testing, business rule testing, and custom testing. And it's really about being, having a tool as, as your buddy that can kind of get you going, sort of, uh, profiling the data and then generate a baseline of these data quality tests that you can review and refine.
Um, and so this, I think, and we're gonna talk about this idea of sort of upending the processing cycle of data quality. Instead of sort of having a long time to define these rules, it's really about more of a, an agile way to generate these rules. And so I'll, I'll just go through these, these four buckets here, A, B, C, and D, just a, uh, give you a quick highlight of what we mean by data quality validation testing.
So, um, you know, initial hygiene screens are kind of like it, I I think of 'em as a periodic review of your data. And it's really about saying, well, should this be blank? Or should you've got four or five different ways to describe what is essentially the same term. Um, and you can see, uh, in, in this image here, there's the word e-bike, and it's written five or six different ways.
Now, that may be fine, that may not be fine. Um, it's certainly harder for an analyst to pivot on, um, that product type if they're all written in a different string. And so, um, this is really a, a great way to look for inconsistencies, um, and, and kind of confirm what you is, what should be in your data.
And, you know, uh, and what's important here is that these things are dispositional too. That some, there's a lot of things that aren't important. You may have a lot of blank columns that may be fine, or it may not be fine. Um, and trying to understand and kind of react to the data and say, this is important.
This is not as, as important. Um, that's an important part. And so dispositioning this, and you can see in the upper right hand of our ui, uh, is an important part of the process that we think. Um, and then second is really anomaly testing. And there's a lot of ways that people in observe the data observability will talk about anomalies.
Um, I, I think what we mean is that you look at your data, you develop a baseline, and then the next time you see some new data, you're doing a comparison. And there's, uh, and that comparison could be, um, some terms that are common commonly used, like freshness or volume or schema. You know, uh, did we get something new? Did we get enough?
Does it fit? And, and then drift checks, like, does it make sense? Um, and think of it as a lot of red flags and warnings. Um, and so,
um, and, and think of it, and also think of this as like a, like a, like a net or sensors to be able to check on what's happening to see if something changed. And again, this, it's important to disposition these because sometimes things may be an anomaly, they may not. And, and it doesn't replace other testing.
It's just looking for sort of variation from a baseline in a table. Um, and again, disposition's important. Um, and so, uh, here's an example of, of anomaly testing. Um, trying to kind of, um, looking at whether a date fits in a date range that you got before. And maybe you've gotten the, the date and it's all from the sort of past three years, or maybe it's not, maybe you got one from 10 years ago.
Could be fine, could not be. Um, and this is the, the case where interacting with the data as you get it and updating the, the test is important. But it, um, a a lot of, I think what a lot of the metaphor that we're trying to go is go from reactive to proactive, don't react.
When someone looks at this and calls you up and says, the state is wrong, I don't trust it. Try to interact with the data and, and learn these things before your customer sees them or before the people who use them, see them. And then there's a whole bunch of other test examples. Um, chip, you wanna talk about this?
Yeah. So we, we, uh, sat down and, and, uh, brainstormed over multiple, multiple times, uh, the different kinds of testing that we could do on, on data where we could actually, uh,
derive meaningful information automatically, uh, just based on, uh, our, our own experiences and knowledge as, as, uh, uh, longtime professionals. The interesting thing was, I think that it, we started to realize that this, we, we could do more than just that,
00:25:00
uh, traditional static view of data quality. That a lot of these kinds of tests that we, uh, that we do really come down to confirming, uh, uh, uh, you know, upstream accuracy and timeliness, confirming the processes of data integration and ingestion, um, and also confirming downstream standards that, that make the data easier for, uh, data consumers to use more effectively.
Because if you can, if you can save, save them time, if you can give them shortcuts and make sure that their assumptions about the data are true, then you really can supercharge their efforts.
All right, thanks, chip. So, I, I think there's another case here of, of the machine can only be so smart, right? And, and, uh, you need some rules that sort of fit your business. Um, and so, like for instance, um, a list of values like here is the list of our products, and maybe that list is in your data, maybe, you know, you've got a new product coming out, so it has to be added to the list.
Um, there are cases where you may wanna check reference from one table to the other to make sure it's right. And I think every business, every organization has a set of, of unique rules that they need to have and, and, and manage. And, and so, um, being able to enforce these on your data as part of your data quality testing and scoring is, is another important point.
And then lastly, just cus custom tests, like every company's got their own domain. Um, you know that the number of medical practices won't exceed the number of doctors that there's no shipments on Friday, uh, like, uh, the stuff you know in your head, and you need a place to put it in. Um, and that is also, uh, the challenge with that is that it's, uh, you need to actually create it.
And so those are, are harder to maintain and document. Um, they require a little bit of programming or sql a, uh, and then also, you know, you, you need to refactor those. Now in our, our DataOps data quality test in there is a place to put that in. But I also think these sort of business rule and custom tests are, are needed to com to complete, uh, what you're trying to do, you're trying to build, trying to get something fast, get it more configurable, and then go all the way to custom.
And all those levels, I think are important for you to build these data quality tests, it's profile, um, and then build a layer of data quality tests over time. And that's kind of the workflow that's currently in, in our TestGen products, sort of profile tables, do the hygiene screening, generate tests, execute tests, and then review and refine tests, testing.
Um, and so this is our current workflow. The, the, the interesting thing is like, well, what happens if you added something on top of that? You said, okay, I've done all this testing, what if I wanna score my data? And how can that sort of scoring work? And and that's really the, the theme of the research.
And so what the, what the people gave us is a lot of good information, right? They said it's gotta be granular, it's gotta go down to certain data elements. It's gotta be multidimensional, and it's gotta be multi versions. So some people love DMA data categories. Some people said DMA categories are s**t, like people had very opinions on it.
Uh, other people said the only thing that really matters are scores based on CDEs or critical data elements. Uh, other people said, I need a score for one customer who's using one set of data, or I need a score that's aligned with what my company's doing. And so I think of this as you're a person who doesn't have leverage.
Getting a scoring system is a way for you to gather leverage and influence, but just because you can score everything doesn't mean you should, right? And, and focus it and, and find the right way to use your leverage based on a lot of factors. And so having sort of a configurable, repeatable, um, quick to build and implement scoring system, uh, we think is, uh, at least that's what we talked about and, and, and, uh, people gave us some good feedback on.
And so what does that mean? So let, let's look at kind of the world of your data, and here's a bunch of dots. And each dot represents a data set. And maybe this is too many for your company, maybe it's too little, but this idea that there's no one perfect scoring model. And so pick scoring the data sets and the scoring models kind of based on the purpose, the efficiency, how to maximize your influence and leverage and, and really what your organization needs.
And so, yeah, the, the sort of, you may wanna cut across all your data and do the typical data, data quality dimensions. I've got all my data, I'm gonna score all my data. Fantastic. Um,
00:30:00
but you also may want to cut it by some different data sets. Like for instance, there may be a set of data that has one data scientist who has a very critical machine learning model data. There may be another case that you identified critical data elements that you're using for maybe you are in banking and you've got critical data elements that you're reporting to the government that you're managing and monitoring.
Well, that's important. Um, or you maybe have a business goal that your marketing and sales team is focused on that's driven by data. And, and all these things are possible, right? To have these things. It's possible to have them at the same time. Um, and so each one of those could require a, uh, its own scoring model based on their dataset and their use.
And so having the ability to have multiple scoring models and then use those scoring models to try and drive it forward. And, and one thing that we learned is that, you know, critical data elements are kind of, they're really important. Like in this case, your customers, the CFO who's in charge of compliance, they've got reports, so they've gotta give to the government for financial compliance.
And, and maybe there's a person in accounting, but sort of they, they're not a panacea to everything. And, and Chip, do you wanna talk about this? Yeah. Well, I, I, I, I think we're being asked to do so much with so much more data that the CCDE process almost becomes a kind of a classic waterfall process where you evaluate in advance, you identify those, uh, uh, key metrics that are used in KPIs.
It's, it's important, right? It's obviously important and it, and it serve, serves the purpose. Uh, but it very often doesn't keep up with the downstream data need. It doesn't keep up with the kind of innovation that, that people are doing in, in, uh, creating data deliverables with machine learning and, uh, you know, predictive models and, and ai.
So, so you're in a situation where you may not be able to predict in advance the particular data points that may not have been important in a CDE, uh, uh, strategy are crucial for, for, uh, predicting, uh, uh, something in a machine learning model. So, so, so how do you deal with that? You, you can't go back in time.
You, you, you have to, uh, uh, have a wider purview of, uh, of what data you wanna, you know, you need to keep an eye on. Yeah. And I think you need wider purview of data. The data's gotta be based on what your customer needs. Your customer has different needs, um, and you've gotta have multiple scoring models.
And if, you know, um, I, I think sometimes it's the magic overgeneralization is always a problem with human beings, right? You see three things and you think everything's like that. And so, yeah, you may wanna have a data quality dimension scoring model that goes across all your data. And maybe that's something that you give to your boss.
Um, who's the CDO? He is looking across the whole company. And you might wanna have one just for that ML team or just that ML customer for the data that they're interested in, because that feeds into their model chip. And, you know, one, one thing that that strikes me is there's a value in measuring it, even if you are not able to fix it.
So if you, you, you have to make the choice of how you allocate your resources in order to fix different kinds of problems. But you darn well wanna know, you know, you darn well wanna have those issues on your radar, uh, uh, far in advance so you can make that choice proactively and in an educated way, rather than make it based on hope.
Yeah. And so, Frank, Frank raised your hand if, uh, what, what I'd suggest with people is there's, in the lower right hand corner of your ui, there's a little, um, chat with everyone button and, and type your, um, type your questions in there, and we'll, we'll try to get to it, uh, at the end.
So, Frank, I saw you, you, uh, raised your hand twice. We'll put 'em in there and we'll, we'll get to them.
So the other part is you're an influence role. You obviously don't have enough time or energy. It's hard enough just to sort of measure, understand and measure data quality and recommend things. But how do you get other people to, to take action? How do you get other people to fix things? And so, um, uh, you know, the, the, what we think is that there's a way, and I guess I think of it from my software engineering days, like back when I was a software engineer, somebody reported a bug be like, okay, there's a bug.
Where is it? How do you recreate it? You're bugging me by giving me a bug. You know, you gotta write the full ticket out. Um, and, uh, so, um, it's really about creating a package that makes it easy
00:35:00
to fix, um, and easy to recreate. And also there, there's a talk of integrating it into workflow. And so, um, excuse me one moment. Uh, there's construction going on next to my house. I'm just gonna close my window. Hold on one second. Sorry about that.
Okay. I hope that's better. And, and, and lastly, uh, a number of people brought up sort of workflow tool, like creating tickets and tracking tickets. And that's also a, a very good way to get people to say is like, here's your tickets. Let's revisit the tickets. I've created these data quality tickets. Fix them in the source data, fix them in your data lake, but actually get some stuff done.
So, enabling other people to take action by giving them the content they need, and then having a way to track, um, what happens.
And then, uh, you know, kind of developing this package, sort of packaging it up all the pieces, what data set, what time you did it, what rule failed, how do you recreate the rule, what's the documentation of the rule, what's commenting, what profiling, what is the rule, et cetera. And being able to, to put that together.
And it just saves, like imagine you've got 30 data quality issues and you're a data quality person, and you had to go in and spend a half hour writing every one of these tickets out. It just would be painful and, and you wouldn't do it right, because it just takes too much work. Um, can we, uh, can we help automate that?
And, and I think we got some good feedback on that. Um, and then last, the one last thing here is, is really measuring data quality improvement over time. And so I think everyone has a boss. Everyone needs to show, um, that they're improving. And so keeping longitudinally track of your data quality is important and, and really about, um, making sure that you're doing that from day one.
And so in some ways, measure first and then establish STA standards as part of the development process process. And that, that really talks about this, um, sort of, uh, improving the data quality score over time. And, and Chip, you wanna take this one?
Did I ask you to do this one? Is this the right one? I asked, Uh, this is, this is not, but I can, I can do it. Um,
let's see here. What, so what did you have in mind for this? Yeah, I'll, I'll just do it, chip. Sorry about that, folks. So, um, so I think the idea is get, don't get a scoring system working, then tweak and improve it quickly. And as you get new data, your DQ scores will change.
And so hopefully it's moving up into the right, but your scores are affected not only by improvements to the data, but improvements to the scoring system as well. So you're constantly refining your DQ standards, um, improving that and trying to reduce the false, false positive rates. And you're continuing to learn about data. And so trying to get this chart of improving over time, I guess what we're, uh, what we're advocating, and we're gonna talk about this in, in a little bit, is, is don't try to get, don't spend six months getting your data quality standards sort of measure first, check second, and try to improve those standards as you're measuring the data.
Just tweak and improve over time. And, and as you get better data, as your standards change, you're gonna have a system that's much more agile, much more responsive to new data. Um, and, uh, it also allows you to get perhaps someone who's less skilled on the team, uh, to be able to actually adjust these standards.
And so it's about, uh, improving the data and improving the valuation, the evaluation of the data simultaneously over time.
So lastly, what do we think we're gonna do? Um, and so, um, so we've got this, uh, data quality test, Jen, we've talked about it before. Um, so we have this perspective that we've gained from doing the research and, and kind of, if you boil it all down, data quality leaders have two major challenges, right?
How do I deal with data at scale? Like, I've got all this data, right? How do I do this and sort of not kill myself? And then how do I affect change, right? How do, how do I, when I've got no power, how do I get leverage to change things? How do I cause people to do things?
Um, and so our perspective is that we want sort of agile data quality at scale. And so what that means is instead of spending a lot of time upfront, sort of defining as this says, defining business rules, analyzing, assessing,
00:40:00
identifying sort of this multi-month process of trying to set up what you're gonna do before you do it, um, and then do it, uh, which is more a traditional kind of waterfall project management approach. It's about putting in some data, measuring your data, learning quickly, creating scoring. Then, um, as the data changes, remeasure the data, reevaluate the, the scoring techniques and your rules, regenerate the tests and keep improving.
And so start measuring and evaluating data quality before your standards are perfect. Um, it's the same thing that we say in everything. Get it 70% right? Get it 50%, right? But get something and then work on it. And so, uh, and then use these measurements as you asymptotically get better. Maybe your first pass is 50%, right?
Your second pass is 70, your third pass is 75. And as it gets better, you're establishing standards and, and, and really cycle time maximizes your learning as a data quality leader. And that's what it's about because you're interacting with the data and working on it. So we kind of think that this process of applying DataOps principles, which are, uh, shorter cycle times, working on things as opposed to documenting things, iterating quickly, and then learning from the real thing, learning from the data, learning from your customer feedback, we'll get more approach.
And then also this idea of focusing, instead of trying to boil the ocean and, um, scoring everything, try to focus it your quality scoring on, on a person who you can have influence on and data sets that you can have.
And, um, you know, I think gimme this slide. This slide. Okay. There you go, chip. So, so, uh, uh, the, the, the, you know, the thing that we've found really exciting about this is that it can solve a number of different problems, this approach, right? I mean, we're all dealing with enormous amounts of data, uh, more and more data, uh, uh, upstream, more and more data needs downstream.
So thi this, uh, the idea here is by automating as much as possible, we can handle large amounts of data at scale. And as the scales continue to, uh, increase, the, the other thing is we have seen that there are certain data quality principles that you can call them or, or, or, or, or rules of thumb that, that you really can apply automatically to, to, uh, identify issues.
You don't have to have deep business knowledge to get at, uh, this fraction of information about, about data. You can, you, you, you can, uh, uh, start imposing, uh, reasonable standards without, uh, uh, uh, subject matter expertise. And that's really important because, uh, subject matter experts are incredibly valuable to a company. They're, they, they're often hard to get ahold of, you know, larger organ in a larger organization.
They, it, it can be hard to find your way to them to even know who they are. Um, and, and, uh, we think that you can take a lot of significant data quality issues off of the table, uh, through this approach. Um, and, and, and then the other thing is, uh, I think, as I said before, the i, the idea is that, that it's always about trade-offs, right?
Data, data quality, you can never have a hundred percent perfect data quality for all of your data points. But the idea here is that by measuring, you're able to make those trade-offs in an enlightened way. And, and you're able to, if, if there are data quality issues, by surfacing them, even if you can't resolve them, you're educating, uh, uh, business people downstream of you who are doing analytics, uh, uh, and if, if it becomes important enough to allocate those resources, you can go in and, and fix them.
So, uh, uh, um, the, the, you know, and I should say too, it's not only the subject matter experts who are have, whose time is very valuable to be leveraged. It's also the technical experts, right? It's the, it, it's the, uh, data engineers. So as much as possible if, if aspects of this work can be given to people who are, uh, kind of, you know, the, the nice thing about data quality people is you're kind of straddling both worlds.
You can, you can tap experts on both sides as needed to go. And, and, and, uh, you, you know, as, as Chris said, this is, this is a faster way to stand up a data quality program
00:45:00
to get actionable information right away. And we've, we've seen this with our tool, I'm sure this is true, uh, uh, with other tools as well. You know, the, you can get actionable information very quickly and, and start making a difference immediately. And then you continue to iterate and, and, uh, you know, measurement, iteration, refinement, these are the, these are really the superpowers of, uh, of DataOps.
Yeah. Thanks Jeff. That's really great. And so, uh, leverage and, and getting leverage and getting influencing others, it's, it's a, it's a skill I've had to work on in my career, right? Because I started off as a young engineer and I saw something wrong, and it seemed pretty obvious to me that it was wrong.
And like, I told someone it was wrong. Why don't they just fix it? And I, after I was like, wow, I, first part of my career was, yeah, people are stupid. And then I got cynical, and then I started to learn that there are skills to influence others to take action, right? And so, um, it's really, I, I think this fits into that skill.
One of it is, is listening to people and, and finding out what their priorities are. The second is kind of convincing people and, and having data and backing it up and, and rolling people in the problem is really good. And so you can say, I've looked at the data, it's bad. Trust me, I'm an expert.
But it's also so much more powerful if you can show them the data and say, here it is, and here's the report on how to fix it. Just, you know, people, um, I think especially people who are more technical are convinced from data, and especially if you can get them and show them now, it may not always happen, right?
And, and that's where I think other aspects of leverage and scoring can come in. And, and for people who are more mature, getting organizations to set scoring goals on their data or projects to set scoring goals, we're trying to move our data quality from 82 to 85%. And we talk to people who have that goal.
We talk to a bank who has a very good data quality score on their critical data elements that they measure quarterly, that they have quarterly goals that run up to the CFO. Now, that's great. They don't have it for everything, their marketing data, their data science data, but at least they, they're getting impact.
And so, uh, leverage is interpersonal skill. Leverage can come from data, leverage can come from scores, um, and all those things can help you get leverage. And that's really what it's about, because it's too frustrating, honestly, to sit there and, and be a data nag and know it's wrong and not get people to listen to you.
So we're trying to give you a tool to get people to listen to you, but it fits in the, the sort of, I think the broader context of, of how to influence others. And, and, uh, there's books and things that I've read that have been very valuable to me. Um, and so what we're talking about here in, in doing this is really adding a new cycle to our product.
And, and, you know, we do the, what we, you saw before, profiling, hygiene, screen generating data, quality tests. And then, um, the next is you get some updates to your source data. We execute the tests, and then we generate some scores. And then you, as a, a data quality person can review and refine those scores.
And then there's, uh, you can then share the DQ issues, uh, with your data owners or data engineers. And this cycle of getting some more updates, executing the tasks, generating the scores, reviewing the scores, sharing the updates, kind of continually working through this process is what we're going to build. And we've got a bunch of features here that we're going to build.
And this is, um, something that's a bit new to us. We're building the open. So we've got these features on our GitHub. We welcome feedback, uh, we'd like you to partake in this if you're, if you're interested in, um, and we're working, gonna be working the rest of the year to build the, these features into our product.
Um, and so, uh, if you agree with them, if you don't agree with 'em, uh, you have an, uh, opportunity to influence us, and we would love, absolutely. Love that. Um, and, uh, so, uh, uh, I guess in our conclusion here, um, kind of sum it up, DA data quality leaders are challenged to affect change at scale in their organization.
Um, and, and, um, we think kind of influence and an agile and DevOps process is the solution to that, that's done through a scoring mechanism that's backed up by automatically and, and somewhat con automatically created data quality tests. And so we're adding new features to our tool to do this, and we, uh, would love to have you be part of the process, um, in helping us do this over, over the next several months.
Um, but it was, uh, from my standpoint, it was, uh, I really appreciate the people who took time to talk to us.
00:50:00
It's always interesting to talk. It's an area of data that I haven't, I've talked to people on, but it, I haven't, um, spent a whole lot of time on. I've fixed a lot of data quality issues on my own as a data engineer, but never, um, done it.
So I guess there's a question. Um, so we are using brute force very, are we provoking the third, eh, winter? Yeah. Yeah. Um, the, the whole question of the third AI winter as being part of the second AI winter, uh, and, and in school when the first one was going on. I, I, I hear you. Um, you know, I think there's a lot of things that go into making data useful, right?
And data quality's only one. Um, you know, things like ontologies and semantics, catalogs, they're all really, uh, I think important. And so we're not, you know, we're looking at our sort of, we're touching the elephant and, and, and, you know, uh, our part of, we're focused on our part of the elephant. Um, the, the, the challenge with machine learning and any algorithmic process is, um, you know, they're only as good as their data.
They're learning from Patterson the data. And some of these algorithms are good at handling noising data and some aren't. Um, some of them hallucinate a lot, some of them don't, some of their, uh, so, um, and then plus a lot of organizations aren't really, they talk a good AI game, but they're just terrible at their old descriptive statistics, right?
And, and, and so I think what's common of all these is if you've got good, uh, data quality, um, uh, you can help. And so, uh, I think, uh, that's where we, uh, our focus is like, if you can fix it sooner, um, you, and you can fix it quicker, you end up having less problems down the line.
And so think of, uh, error pyramids in manufacturing. If you fix a problem at the beginning of the manufacturing line, it doesn't get to the car ship and have a recall. Recalls are very expensive, and data teams are recalling a lot of their reports. They're re recalling a lot and, and, and stop having recalls find the problem sooner.
And the soonest way you can find it is fixing the data quality. Now there's other places to find it, right? You could, uh, wanna find it when you've integrated the data into your data model for a report, and maybe that's the right place to fix it. Um, but we always have this what level in the stack and the most base level of any data stack is the data itself.
And improving that data, um, in its source system or improving it in, uh, a data lake is, is I think, essential for all these different uses.
Um, so there's a question. What about, um, AI and ML and data quality? Um, so we have a whole bunch of algorithms, um, and that run on data to auto generate these data quality tests. So, so we're a believer, um, now is there a, a believer that everything can be done through ai, that all, all that there's a magic box that'll improve your data.
I'm not a believer of that, but I believe algorithms like a good ui, like having a a good tool are useful for people. Um, and we continue to add more algorithms into the mix to improve data quality. I dunno, chip, do you wanna, do you wanna talk about that? Well, we're, we're focused on, on structured data.
There, there's so many different possible answers to the, to the question, you know, how, how, uh, uh, we can use AI to integrate into workflows. Uh, uh, that's, there's some really exciting stuff that's just starting to happen there. Um, we believe that our metadata that we're creating through test results, uh, can, can feed into, uh, uh, machine learning model to help identify what, uh, what test failures are important and what aren't.
The idea is to is, is, is to use, uh, what we call test disposition. The, the user's decisions that, hey, this was important, this was, this was, uh, a false positive to help refine our, our, uh, uh, prior prioritization and, and, uh, uh, uh, identification of issues in the future. So, so that's something that we're excited about.
Yeah, so the both, uh, Frank and Ann Kush have questions on unstructured data, like text documents or documents, PDFs. So our short answer is, ah, we don't know. 'cause we're not doing any of that. We're not, uh, our, our tool is basically, uh, for data that's resident in rows and columns in the database.
Um, we haven't extended it to looking at completely unstructured data. Um, if it can be turned into, if you can take a, uh, a document and turn it into rows and columns, um, we can, we can then analyze it. But we're not, um, uh, right now our focus is on sort of business level data that's in rows and columns.
O one thing that blows me away about LLMs is the, is their ability to do just that, right?
00:55:00
To take, to take, uh, unstructured data and, and pull structured information out of it, you know, turning emails into spreadsheets, for example. It's, it's, once that data is structured, then we can analyze it in our system. Yeah. Yeah. And so I, I think there's, there's other people who handle sort of text data or hierarchical data or image data.
We're just not, we're just not, um, the group that does that, we're, we're focused on this boring rows and columns that 99% of the people use and, and don't do very well. Um, and so I, I guess, uh, Annie has a question. What are some advice for handling frequent changes to processes and documentations to maintain good data quality?
There are some cases in which a process gets changed frequently due to external factors. Chip, you wanna take that or you want me to take it? Well, j I'll just say that's, that's what I love about automating test generation as well as test, uh, execution. That it is, the idea is to identify changes.
And, and the changes may be perfectly valid because it's part of some business process change upstream, or they may be errors that are introduced into, uh, you know, into your data set. But what you're able to do is to make that decision when you see a test failure, and you can either manually change tests or you can regenerate tests to, to, uh, reflect the new data set.
And, and, and that's what we're trying to do, right? We're we're, we're trying to make this as, as, uh, responsive as possible, rather than waiting for the, the developers to tell the data quality committee, which creates a new set of standards. And, you know, that that stuff can just be so painfully slow. Yeah. Yeah. And I think it also, it, it becomes the question of like, when and where things are fixed, right?
And so, um, like I've spent a lot of time fixing data as a data engineer, not having data quality people involved, but I'm really fixing what is classically data quality problems. And so some new data came in and they forgot a column, and so you gotta patch it back. Um, sometimes you've got some new data and it's so terrible, you put have to push back, right?
And say, look, I, I need a new data set. Um, uh, or sometimes your process change, you're adding new code or new configuration into your process. But all of those, uh, I think are why we, we have pushed for a DataOps approach because change in data, change in processes, acting upon data change in what your customer repo, uh, wants is very, very, very normal.
And so management of change and management of variation is really hard, uh, if you don't adopt these approaches and if you don't have the right tools. And one of the things I do like about our TestGen tool is that it has, um, a really nice UI because management of change doesn't have to be on the data quality person or on the data engineer it, or on the data steward it.
Uh, the UI provides a way for different people to kind of update their data quality standards and their data quality tests. And maybe it, it's one or the other, or all three are collaborating to do this, and maybe you're collaborating just in the, the hour right before you ship. Uh, you, you have to ship this out.
And I think that really, um, having a, a plane to collaborate on not having these things hidden away, uh, I think can, can be helpful. So, um, yeah, I, I just think change is, change is really important. And, and to me, management of change, uh, and, and improvement of cycle time and automation to do that, and lowering of risk of change of changes so that you can be successful is the key to success and analytics.
And that's sort of why we've, uh, founded the company and why we talk about DataOps.
Um, oh, so Andy's got the, a good question. What's your take on the DQ dimensions? Accuracy, consistency, um, chip, you wanna take that? 'cause I, I, I'll have a bad, I'll, I'll, I'll speak, I'll speak badly. Uh, well, so I, I, I, I think it has a role, right? It, it's definitely important. It, it's valuable.
Uh, uh, the thing that I have found that, that it doesn't, the, the question that it doesn't answer for me is particularly where are problems occurring, right? You have, you have processes that happen, you know, uh, uh, upstream, uh, uh, ingestion data integration data, data, uh, conformance. Uh, then, then you have, uh, downstream users who are, who are adding in business rules and assumptions.
It's all across that process that data errors can be introduced. And you might have, uh, missing data that's caused by something that's not, uh, that, you know,
01:00:00
that was never delivered upstream, or it might be caused by an incorrect join that causes, uh, uh, data not to appear in a downstream dataset. So while, while it, it may be helpful on some level to, to, uh, have those categories, it's a very static view rather than an operational view. And, and it makes it harder to, uh, to target change.
Yeah, I think that's true. I, I, I'm, I'm not a fan because I think it's a kind of an officious hierarchy that doesn't really have a lot of use. And like, if I go into a library now, I don't, I, I mean, I used to know the numbers, you know, the Dewey Decibel system numbers and fictions here and NonFiction's there, and science is there, but I, I just like go and I just search the catalog for what I want, and then I go get the number.
And so these hierarchies, I don't think have a lot of, of, uh, importance. And, and one woman really emphasized to me, they don't do much. It has to be connected to action, actions that matter to the person who's gonna make the change. Hmm. So that's what matters. Is it? And so, um, how you organize it and bucket these things, doesn't matter.
You've got a list of changes for people to make. And why, why should they make them? Well, it's because of the, uh, the first half or the second half of 2020 four's, uh, business goals. You have to improve data quality as part of this, because it's used in these, uh, marketing and sales incentives that, that's, that's like real and concrete or, um, that's backed up by a, a predictive model.
And that predictive model is part of this. It's very concrete. Um, and, and getting at it, or we have these data elements that are gonna go into our report to the federal government, they better be, right? And so I like con I like concrete things, and I, I, I, um, not to say that it's, these, uh, groupings are, are important, but I just think that the groupings themselves don't do much.
And so, uh, we've got three minutes left. It's the feedback and questions have been great. Uh, Frank asked the questions. What I will do is take the slides and the recording, um, and I'll post it up on a webpage and I'll send it out to you, uh, uh, probably tomorrow, um, with all this information.
And so you'll get all of it. And I appreciate you all taking the time. Um, go to our, our, our GitHub, uh, site and, and check out the tickets and we'll, um, we'd love to have your feedback and thanks for spending the time with us. And Chip, thanks for helping out. Thanks Everybody. Just remember, you can't be a good chef unless you smell the fish.
Good metaphor, chip. All right. See everybody. Bye.
End call. Okay. I gotta go. Stop sharing.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Questions from this session
What is data quality?
Data quality is a comparison of the current state of your data against the desired state, based on user expectations, usage requirements, and defined quality standards. Good data is fit for its intended purpose; poor data is out of compliance with the standards you set and unfit for operational and decision-making use. It remains a hard problem: 57 percent of respondents to a 2024 dbt Labs survey rated data quality among the three most challenging parts of data preparation, up from 41 percent, and IDC reports that 73 percent of data practitioners do not trust their data.
Why do data quality leaders struggle to get data fixed?
Data quality leaders can usually find the problem but rarely have the time, skill, access, or role to fix the data itself. They have influence rather than direct authority, which is why some get labelled data nags. To move anything they have to decide where the change should be made, who should make it, why it is worth making, and what information that person needs to act.
What are the four types of data quality validation tests?
Hygiene screening tests surface basic problems found during profiling, such as numbers stored in alphabetic columns, inconsistent blank value representation, and string pattern inconsistency. Anomaly tests compare new data against a profiled baseline for freshness, volume, schema, and drift. Business rule tests encode domain expertise, such as verifying that a delivery status is one of the expected values or that a SKU exists in the distributor list. Custom tests handle industry-specific logic that cannot be driven by parameters.
How fast can automatically generated data quality tests run?
Tests are pushed down as database queries, so execution scales with the database rather than a separate engine. In the workflow shown in this session, 1,000 tests run in under three minutes and 15,000 tests run in under 20 minutes. The generation side produces 51 profiling characteristics, 27 profiling hygiene tests, 32 auto-selected test types, two custom test types, and eight fill-in-the-blank multi-column business rule test types.
Should an organization use one data quality score or several?
Several. There is no single perfect scoring model, so scoring data sets should be picked based on goals, purpose, and organizational need. A score built on DAMA quality dimensions, a score for critical data elements, a score for the data behind one machine learning model, and a score for this year's top business priority each have a different customer and a different person who would fix the data: a CDO, a CFO, an ML team, a VP of sales.
What is an Agile or DataOps approach to data quality?
Start measuring and evaluating data quality before standards are established, then use the resulting measurements to set the standards and improve them over time. Documentation and analysis become a result of the process rather than a precondition for starting it. As new data arrives, revisit both the data and the tests behind the scores, cutting false positives and refining standards as you learn more about the data.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.