On-Demand Webinar · 56 min
A Guide to the Six Types of Data Quality Dashboards
Chris Bergh and Gil Benghiat on the six kinds of data quality dashboard — dimension, critical data element, business goal, data source, data consumer, and ticket-driven — and how to pick the one that will actually change someone's behaviour.
What you'll learn 6 points
- Six types, each of which drives a different change: dimension-focused, critical-data-element-focused, business-goal-focused, data-source-focused, data-consumer-focused, and a ticket-driven workflow used as a dashboard.
- Dimension-focused dashboards give the broadest overview and are the easiest to ignore — too abstract to line up with what a stakeholder actually cares about.
- CDE dashboards earn their place in regulated industries, where critical identifiers and financial values carry compliance risk and the definitions are already written down.
- A dashboard is only as good as the actions underneath it. Every score should trace to a data quality test or hygiene result, so a reader knows what to fix.
- The equivalence principle: test results, workflow tickets, and a dashboard are three views of one thing. A dashboard is a collection of potential tickets, each backed by tests.
- Build them iteratively, and expect to run more than one. Pick the type that fits your data landscape and your stakeholders, then refine on their feedback.
Prefer to read it? The written version is in Webinar: A Guide to the Six Types of Data Quality Dashboards.
Slides
Transcript
Show chapters and dialogue 43 chapters · 10,004 words
- 0:00 Welcome, and 20 years in data and analytics
- 3:50 Housekeeping: slides, recording, and the chat window
- 4:50 Data quality improvement starts with one empowered individual
- 5:49 A score is not enough without a concrete fix attached
- 6:48 Why data quality, and the tragedy of the commons
- 8:42 The six dashboard types, introduced
- 9:41 Type 1: the data quality dimensions everybody already knows
- 10:28 Where a dimension dashboard helps, and why some people reject it
- 12:35 Type 2: critical data elements in regulated industries
- 14:34 What is critical to you may not be critical to anyone else
- 15:29 Type 3: tying data quality to a business goal
- 17:11 Motivating a team whose own data is already good enough
- 18:08 Type 4: tracking quality by data source
- 19:03 Holding suppliers accountable when they don't know you exist
- 20:00 A data engineer on arbitrary changes to the feed
- 21:48 Type 5: a dashboard built for one named data consumer
- 23:52 Proliferation risk, and Type 6: tickets as the dashboard
- 25:34 Ticket backlogs, and a framework for choosing among six hammers
- 27:27 Unruly data sources are wild horses that need taming
- 28:26 Q: how do you identify a ticket's root cause as data quality?
- 29:31 Why a mix of dashboards beats one-size-fits-all
- 30:54 Dashboards must be actionable: scores rooted in test results
- 31:47 A score with no action, and what a real bug report contains
- 33:52 Iterate in small assessments instead of running a waterfall
- 35:03 Five data sets assessed, or really only two
- 36:08 Business goal versus CDE, and getting people on your side
- 38:18 Demo: profiling a sale price column and chasing the warning
- 39:17 Demo: the same data read from a CDE standpoint
- 40:15 Demo: dimension scores, and the score trend over time
- 41:12 Fix the data or fix the tests, and managing generated checks
- 42:04 Demo: filtering a dashboard by data source metadata
- 43:11 Data quality denial is cultural, not technical
- 44:09 Workarounds and silence are the dangerous signs
- 45:08 Resources, then Q: what are the terms of the software?
- 45:51 Apache 2, one login and one connection, $100 a month beyond that
- 46:46 Q: how do you link a dashboard to a trust or maturity model?
- 48:01 A score is a point on a maturity model, and Q: where does profiling fit?
- 49:01 Tests get you 70% of the way, then who can you get on your side?
- 50:49 Q: reconciling customer numbers between two systems
- 51:40 Why that one has no technical answer, and Q: data sitting in S3
- 52:36 S3 needs a query execution engine; the tool needs structured SQL data
- 53:27 Where the slides and the recording will be posted
- 54:06 Q: how does the tool generate tests without domain understanding?
00:00:00 Welcome, and 20 years in data and analytics
Chris Bergh: Hello everybody. We're going to wait one or two minutes to start. And welcome to all the AI notetakers who are attending this meeting virtually. I think there's about four or five people showing up who are here. So we'll give it a few more minutes and then I'll start All right. So so why don't we begin the webinar and first let me tell you about myself and and what we're going to do today. So my name is Chris Bergh. I'm CEO and founder of DataKitchen and I've been doing sort of the data and analytic business now almost 20 years, and one of the things that we've gained over those years and kind of doing every role in data and and building this company that focuses on on DataOps and data quality and observability we're going to share today. But before we start, let me take care of some bookkeeping. So I it's interesting there's 53 companies or countries represented today from all over the world in every continent.
00:03:50 Housekeeping: slides, recording, and the chat window
Chris Bergh: I will share the slides and the and the recording and the transcription. So I'll we'll put up a little web page and share that with you. So there's a chat window that's actually if you look in the lower right hand corner you can see it and so that will if you have any questions put it there. I will endeavor to answer them as I can, but I'll also go and and look back at the questions at the end. And so we've targeted about 45 minutes and generally with questions and it should be no more than an hour. So thank you very much for attending today and and let me start. So what are we going to do today? So we're going to talk about the sort of why and the what and the when and the how. And so why a data quality dashboard and and sort of what are the six types of data quality dashboards and when do you use them and and and how do you create them and then we'll do an example demo with our open source software of each of the types of dashboards and then we'll have a have a conclusion.
00:04:50 Data quality improvement starts with one empowered individual
Chris Bergh: So we we have a point of view on data quality improvement that's come you know from experience. It came from some market research we did last year, and let let me go through that because that underlies everything that we're going to say today. And honestly, if you don't agree with it, you should probably leave the call now. So, I think data quality improvement starts with a single empowered individual who who's interested in driving data quality change. And they do that via influence. Most people who want to make change to data quality, they can't tell people to do it. They only have influence on doing it, and second they focus they need to focus on a specific end customer. So it's not general vague data quality. It's like who's going to feel the impact from this. And third, they need to start small and iterate quickly using DataOps principles. And we've been talking about DataOps for better better than a decade. And so we've got books and training that we can share later.
00:05:49 A score is not enough without a concrete fix attached
Chris Bergh: And then giving specific and concrete remediation actions, and that's something that's going to be a theme through this. A score is not enough. It has to be a score backed by here, you should fix this. This column, this table, this is exactly what's wrong. And then the whole point of what we're here today is quality needs to be measured in some way. Why? Because that is good in of itself. Like we're all data and analytic people. But it's also good because as a single empowered individual, I have to prove that I have value. And that line that goes up and to the right is really important, and so showing that your quality is improving actually helps extend your influence, helps show that you're making a difference. And so that's really what we're going to talk today is is kind of sort of high, you know, how do you set up a measurement system? And it's enabled by free open-source technology. So we've written a lot about this over the past six months year on data quality improvement.
00:06:48 Why data quality, and the tragedy of the commons
Chris Bergh: We've got webinars, we've got white papers. I'll share this at the end. We've got a certification. We've even got a video series with Uncle Chip that's really fun to watch. And so those will be shared, but they're also linkable from this. And so every presentation about data quality has this slide. Why data quality? And I guess because you're here, you believe in it. But as you know where data quality can derail operations and screw stuff up and erode trust erode trust with your customers and our view is that data quality is is has this ecological term the tra tragedy of the commons that is everyone wants everyone's grazing on the data quality field but it's been cut down to nothing because you know no one's investing in making that that shared common field good grazing grass for us to get data off of. And then data quality leadership itself as I said before it's an influence challenge. It's not about power. It's not about telling people what to do. And how do you get someone to take action? And so that's really our perspective. And so why a dashboard? Well they're essential, right? They provide a they they take something that's invisible like the health of your data and make it visible. And well done ones enable targeted action improvements. And sometimes dashboards are called reports or scorecards or assessments. They have different terms and really not all just because you have a dashboard they're not it's not great. They're not all created equal. It's got to be aligned with what people care about. What the data landscape is, what their challenge is. And then you know why you're here. There's just a bunch of distinct types of data quality dashboards each serving a unique role. And so you want to be able to sort of maximize your impact, streamline your efforts, and and really drive that change in data quality. So that's all about why you should do a dashboard. So what are the six types of data quality dashboards?
00:08:42 The six dashboard types, introduced
Chris Bergh: And let's go through them. And so the first one is called dimension or data quality dimension focused dashboards. The next one are critical data elements. And so I'll use the term CDEs. It's but that really means specific elements that the corporation has found valuable. The last one focuses not on a person but on a goal. And maybe it's you know there's always a person responsible for a goal in the organization. The fourth one doesn't again it for it doesn't focus necessarily on a person but but perhaps it does. Somebody owns a data source and how good are your sources?, you know, we've worked with companies who have a hundred different data sources that come into their warehouse and infrastructure. And then how do you actually focus on the person who receives the value at data consumer? And then lastly, one that's kind of not a dashboard in in in any sense, but it's ticket- driven workflow as a way to to judge change in data quality.
00:09:41 Type 1: the data quality dimensions everybody already knows
Chris Bergh: And so what I'm going to do in this I have sort of four slides on each. What I'm going to do is have one slide that's an overview, one slide that talks about benefits. I'll have a quote and then I will actually talk about some challenges in this. So data quality dimensions everyone knows what they are. They have there are ways to categorize what is data quality and so people have written a ton about it. And there's the sort of fundamental things of of data quality. Is it complete? Is it accurate? Is it time timeliness? And they're kind of philosophical in some ways like what does complete mean? What does unique mean? I'm not going to talk very much about that at all today. I I think of them as just a grouping. They're grouping of the areas of concerns and and they do support targeted interventions based on those sort of dimension specific insights. So what is it?
00:10:28 Where a dimension dashboard helps, and why some people reject it
Chris Bergh: Well well sometimes you want to have a very broad brush to cover lots of data elements like how's all our data doing? And it's a relatively standard grouping of data quality issues and categories. It does facilitate sort of consistent and repeatable evaluation across data sets and it's kind of I think like high level data governance reporting I think it's good like if your co wants to report on all the data across all the data quality dimensions and so however some people don't believe in them so we talked to a senior data quality person 20 years experience and basically she said data the data quality dimensions are crap they don't they do nothing to motivate people or make improvements to their data and and she also on that the sort of philosophical things about what is you know what is consistency what is accuracy kind of derail the process and so you know that that's kind of goes in challenges they may be too abstract or general it's really about influence and not science here that's another theme it's not about saying exactly what dimensions are not scientists we're just trying to get people to do stuff and and the sort of context specific definitions like well why should I improve accuracy what does it mean to me and they end up kind of being check in the box and they're not aligned to stakeholder needs and they don't really resonate with business people like 70% consistency. You know, WTF, I'm trying to make my number. I'm trying to, you know, improve quality of manufacturing. What does consistency matter to me? And and they can't really connect to what people care about. And and you know I and the reason I say that is is remember that data quality is almost always good enough for the people who own the data. So if I'm putting in data in a transactional database and I run a call center, I'm putting it enough to meet my goals. Maybe the data in that call center is not good enough to do marketing insights. Or maybe that data in the call center is not good enough to to help decide how we're going to hire or fire people, but it's good enough for me to run the call center. So why do I care about 70% consistency? So that perspective of data owners in some ways the data is always sort of good enough for their needs.
00:12:35 Type 2: critical data elements in regulated industries
Chris Bergh: But the tragedy of the commons is not good enough for the corporation's needs, for the whole group needs. So now the next one so so the critical data elements. So they're kind of essential I think for business who operate in more regulatory environments. So a bank has got to have reporting and that reporting has got to be based on critical data elements, right? So their financial values, they could be key identifiers like customer ids, account numbers. And so this could be used to look at credit scores or regulatory compliance or risk modeling. And so they are really effective in regulated industries where CDEs are clearly defined by regulatory bodies or by industry standards. And they can be very powerful tools for, you know, risk mitigation or compliance and and they really make make it clear about what matters. These external mandates, hey, we've got the risk report to get out or else we're going to go out of business or the regulators are going to stop us acquiring another bank. That that really can help a company focus. And then everyone understands, well we've got a compliance report and a compliance report's made of data and the data is made of these CDEs, so our CDEs have to be good. It kind of directly connects to people. And so as an example we talked to a sort of a large US bank and they had sort of amazing results in aligning people on CDEs because they had corporate goals like they got leadership bought in and and they had to do them right focusing on data quality and reporting affected their line bottom line and it wasn't just fear it was also that was like, oh, we're going to the regulatory bodies are going to be be on us. They also said there's sort of a benefit to it. If we get our reg regulations in right, it'll enable us to, I think they said, buy another bank or grow up a banking level or something like that, however, that does have some challenges, right? How do you identify what a CD is when you don't have a regulator breathing down your neck?
00:14:34 What is critical to you may not be critical to anyone else
Chris Bergh: What's critical to you may not be critical to someone else, and that sort of alignment is tough sometimes, and then CDEs may change, right? And then in cases where you're doing things like AI or predictive modeling, how do you predefine CDEs where you're actually trying to figure out what the features are in the model? And that's part of the modeling process. And so this iterative approach are is is sort of necessary to kind of actually define what your CDEs are. And so kind of going on to the next one which is so so CDEs are great. I think however there's another type here that's business goal focused and this is a this is different. This is sort of what is the goal of your company. Well maybe you have Q3 sales metrics and you're trying to have a a a goal to get three new customers and to get three new customers, you're going to have to do a cross-sell campaign or 30 customers or 300 customers.
00:15:29 Type 3: tying data quality to a business goal
Chris Bergh: And to do that, you need good data to be able to actually do that cross-sell campaign to customers. And so aligning data to a goal that a company has to get more customers, to cut costs, however it is, can actually drive change because then you're saying, well the data has to be good to do this activity that we're going to do to meet the goal. And so it links it directly to specific organizational objectives. And it actually really highlights how data quality influences those objectives, and and you know, it's it's a challenge, right? The call center may be very good at collecting phone numbers, but in order to but you may need email addresses because you're going to do an email campaign to get those cross-sell customers. And so leadership can see these operational and financial implications, right? Because if you're going to do a cross-sell campaign, well you're going to have to do it to someone and do it by email. So and then it makes data quality more strategic, I think, by tying it to business goals. And so that connection makes data quality relevant to people. It's not just a distraction. It's like, well we need these data elements to be good because we're trying to make some money, and it shows that data quality h can have business impact. It gets executives bought in because they it's their, you know, it's their bottom line, and then it enables this impact or some way to judge the ROI of actually doing the work, and then it kind of comes from like this data quality is sort of a technical nerd thing. Oh, it's these people in the background. They're data nags, you know, pat them on the head and go away to the people who really run the business, but it takes that from from like a technical nerd concern to like putting it in the business part of the organization. And I think that's actually good, and so here's a here's another quote. So how do you get business to take action?
00:17:11 Motivating a team whose own data is already good enough
Chris Bergh: Well by definition, they're inputting data that meets their data quality needs with all their data problems. Why should they care about data quality when when providing data other teams? Well the really is is that why you got to mo motivate them by directly affecting their quarterly b business goals. They care about that, and I think that's a really important perspective. And so you know the the challenge is also how do you define and maintain alignment between this goal this cross-sell campaign these data elements that's sometimes hard to do. And then business objectives change frequently right and so you you may get halfway through and then well the quarter's over, and then you know this there's some complexity in maintaining this alignment between goals and data. But I I I like the business goal challenge because it it provides a way to motivate people to make change. Now this is also one of my favorite ones, right? Is is because data comes from a lot of places.
00:18:08 Type 4: tracking quality by data source
Chris Bergh: It may come from your internal Salesforce system or your internal ERP system or you may get data from an external supplier or you may purchase data. And so these are data sources, right? And and some of them are good and some of them are bad, and so you can look at it and track metrics on the quality of the data that they're giving you, the amount of error rates, and it helps hold suppliers accountable because your data people are often taking it on the chin when they decide that suppliers decide, whoops, we're just going to transpose columns or we're going to change the meaning of a columns or we're going to send you twice as much data that we did before. We're just going to duplicate it twice. All of which I've had the experience of happening and and it kind of fosters accountability. And this graph on the right is a something that we call a tornado report and it actually shows the number of errors kind of balance with the amount of work it took the the data team to go kind of broken out by source here.
00:19:03 Holding suppliers accountable when they don't know you exist
Chris Bergh: And so looking at it saying I've got five or six sources and this one source is giving me really crappy data. That that's very helpful, right? And this accountability is important because in my years of doing data, your data suppliers kind of don't care that you exist, right? They're like, "Oh, let's there's some data people. We just changed the schema on Thursday." Well they'll figure it out and work the weekend to make it all happen. And so the by having this accountability, by having a dashboard that's backed up with specific recommendations, you can then they know that they're being watched, and I've learned to not shame and blame the providers, but say, "Hey, look, we found these four problems. We fixed them, but here, could you fix it next time?", and they actually the whole thing gets better and suppliers get better once they know that you're sort of watching them and and if you're kind about it, instead of saying, you know, you you jerk, it they tend to get better.
00:20:00 A data engineer on arbitrary changes to the feed
Chris Bergh: And so it also helps your data team not struggle. It helps improve the quality of the the reports etc. And just saves a lot of errors. And so here's a case you know from a data engineer. You know our suppliers don't know that we exist. They don't care about quality. They'll make arbitrary changes to our feed at any time. And it's sort of up to us to deal with it. And so they have no control over their data quality. And I think that's true, right? It's it's unfortunate that you know the contractual relationship doesn't really matter but it is it is the world that we live in and so testing it the data and making a data source focused dashboard can help you gain leverage and can also help on the other hand if you're having a team that's having to fix a lot of things it can also show sort of why your team is struggling look we had 27 issues from our supplier last week it took us 10 hours to fix them that's why we didn't get X done. And it also helps you sort of getting at the root cause of where things are. And it I think a lot of it is you actually read nearly you need really good data source tests. You need to check the data as soon as it arrives. And we've talked a lot about testing and where to test. Testing data as soon as it arrives is important, but also testing data as it's integrated with other data sets, as you clean it and and put it into a an analyst ready format is also just as important, but the data source one is should really be focused just when the data arrives, and and lastly, the or second to lastly, there's this data consumer one, and and so think of this as in your mind, think of a I've got a data scientist, right? And he or she's got this really cool model and maybe they've worked out all the features that go with it and they care about this model and this model may be tied to a business goal but but they're invested in its success and how can you make sure that they're successful.
00:21:48 Type 5: a dashboard built for one named data consumer
Chris Bergh: Well you could build a dashboard saying there's these 16 data elements that are very focused and very important for this predictive model. And so let's check the quality of those and let's monitor that quality over time and see that it's improving and be able to give recommendations to the data providers to fix them. And so it helps the data consumers then use their influence to be able to make change. And again, this is an influence problem. How do you get other people to to drive influence for you? Is is is the key and and not be the nag, but say, "Okay, data scientists, you're you know, you're the new cool kids in school. We have AI as a goal. Your model is is really important to the CEO. We all know that garbage in and garbage out's a problem. So maybe you can help. I I'll set up the model for you, but maybe you can help get these get, you know, corral the cats to actually improve their their data quality." And like I said, it's it it if you've got an analyst or a data scientist, they get accurate, correct data, they're very happy. And it gives them leverage to do it because they can see it. It's concrete. It's measurable. You can have more than one, right? A a highlevel CEO report, BI report that goes to the exec team that tracks 27 data elements, but maybe you've got a data scientist that tracks a different 13. Those are two good dashboards each for one. And it really by being able to have these sort of person focused or consumer focused dashboards really helps extend your influence leverage. And you know here's this quote. Our data scientists are forever complaining about the quality of the data. You know they have these highv value predictions in production. They're motivated to do it. So if we could focus on what they need then we get they give us help in some ways and that's a a consultant told us that and you know the the challenges are always you know no technology can help trying to define exactly what data elements they're needed and what's more important and then trying to work with a lot of people sometimes to get the data elements that are important to them.
00:23:52 Proliferation risk, and Type 6: tickets as the dashboard
Chris Bergh: And then resource allocations. What what if you've got three critical reports and three critical models and only certain amount of time in the day? And so there could be a risk of proliferation here when you've got seven data consumer focused dashboards all that have a lifetime. So you've got to it it's you've got to sort of manage the manage that. And so lastly there's a dashboard here that's not a dashboard at all. And so this is just using tickets. And so fix these four things of data quality and that's what matters the number of tickets that are created and the number of tickets that are closed. And in this case each ticket represents a well- defin fined issue. So in this case we're actually not in some ways measuring data quality. It makes but we're making it actionable trackable. And it really shifts from traditional quality metrics to sort of operational progress. How many tickets were done this week? How many tickets will you do? And that is a very actionable way to think about it. And so by focusing just on here, do this and saying you've got to, you know, your organization has got to fix 10 data quality tickets this quarter, I think that's that's a reasonable way to drive behavior change, right? And and so and then it helps you also look for bottlenecks and saying, "Okay, you only fixed seven tickets. Why is that? And you agreed to do 10 tickets. Why haven't you done done them? So there is a lot of benefits of thinking in and tickets and workflow. And here's one one customer that they don't use a dashboard. They just use workflow. There is no data quality dashboard. They just use the Jira ticket UI, right? And that and it's all about the number of tickets created and fixed during that particular period. And people have specific goals. Data owners have specific goals around tickets. And so I think that's pretty cool, right?
00:25:34 Ticket backlogs, and a framework for choosing among six hammers
Chris Bergh: You know the challenges here you need you need very direct tests to create the tickets right because if someone is going to have a ticket they you say oh the quality of this column sucks well what does that mean you need to give very well-defined good tickets and I think that is backed up here with every one of these and I'll talk about that in a bit and then you also risk creating backlogs that are overwhelming you know you have 3,000 tickets in backlog how do I how do I groom you need a ticketing system and governance. You need some consistency and in some ways it measures work done and not data quality and and maybe that's enough for your team but but it it it isn't one way that people have been successful. And so those are the six really and that I've talked about, and so before I jump into kind of questions, is there any comments or questions I can answer? So, An attendee had a comment and I'll talk about that in a bit. It's hard to see beyond all the AI notetakers who are putting comments in.
Gil Benghiat: Yeah, that's the only that's the only one we got. An attendee's comment
Chris Bergh: All right. Thanks, Gil, so how do you think about this? Right? There's these six dashboard types and and what's the framework to be able to say which one should I use?, and so you've got six hammers. Which one do I use to pound the nail in? And so let let me talk a little bit about that. So you know, dimension focused dashboards I think are good when you've got a broad kind of undifferentiated data landscape and you're trying to, you know, look broadly. I think CDE ones are great when you can define CDEs and people will sign up to actually push CDEs. I think business goal ones are great when you can have when you have business goals and and you can directly connect those to data elements and actions to meet those business goals.
00:27:27 Unruly data sources are wild horses that need taming
Chris Bergh: I think data source ones are important especially when you have unruly data sources. Maybe you don't need them for, you know, in my experience, data sources are kind of like, you know, crappy data sources are kind of like wild horses. They need to be tamed. And so you can tame them through tests, you can tame them through test bed, feeding on dashboards, but once they're tamed, they're okay. So you don't you you know maybe you need them in your back pocket but you should always so when you've got definitely when you've got an unruly data source that that could be helpful data consumer ones or when you have you know everyone's got data consumers of data. So that's that's kind of a if you've got a vocal data consumer and ticket- driven workflows. I'm going to talk a little bit about this because I think in some ways a data quality dashboard and a ticketing system are two sides of the same thing, and so I'll talk a little bit about that.
00:28:26 Q: how do you identify a ticket's root cause as data quality?
Chris Bergh: So let's go on to the next one.
Gil Benghiat: Oh, Chris, we got a question from an attendee on tickets.
Chris Bergh: Oh yeah. So the problem with tickets is how to identify a ticket as a root cause as data quality. So so let let me I'm going to answer that in a second. There's sort of the next slide talks about it. And so but in general the way that we think of a data quality dashboard is it's rooted in actions and the scores that make up a data quality are rooted in data quality tests. And so that's a but let me let me finish this thought and then I'll jump to I'll jump to the attendee's question. So you know the key insight here is every dashboard type kind of touches the elephant in a different way. It addresses a different stakeholder needs and organizational perspectives. And so you really need to ask who is the customer for my data, right? Because let me find the customer that's going to be able to drive influence.
00:29:31 Why a mix of dashboards beats one-size-fits-all
Chris Bergh: And you know, you could even say that for source dashboards because your customer is the data engineer who's going to suffer the consequences. And so I'm a believer in multi-day dashboard strategies. You you should have a mix of data consumer or CDE or data source. It's not there's not a one-size-fits-all here. And that's what I I I I believe is like these these types of dashboards are different techniques that you can use to drive influence to show that you're doing well by tracking the improvement over time and to drive improvement at data quality and so so what are some heruristics right you know I think use dimensions ones for foundational monitoring across all data think of it as your your your base layer if your CDO wants to say how great am I doing overall and use CDEs when you can have when it's pretty clear to the organization what's what CDEs are when it's regulatory or high impact data use business goals when it's pretty clear that you can translate them into ROI or translate the business goals into data sets and use the source or consumer or ticket driven when when you need when it's targeted to specific cases And so we're 12:30 actually.
00:30:54 Dashboards must be actionable: scores rooted in test results
Chris Bergh: There's a bunch of great questions. So I'm I'm just going to go through in order to have the next 15 minutes I'm there's some good questions I'll answer at the end. So I want to get to this because that answers the attendee's question. So data quality dashboards must be actionable. So the basis of a dashboard has to be concrete actions. It has to be test results or in our another way data hygiene test results. So it has to say these actions make up a dashboard. So they have to be built from discrete concrete actions. And in some ways data quality tests are the foundation of your data quality score. Not some fuzzy map based on profiling, but like you should be able to another way to say it is you should be able to click through your dashboard to say here's a score and here's all the actions from it. And so that way is the only way that you can get people to do things.
00:31:47 A score with no action, and what a real bug report contains
Chris Bergh: I think is because if you're giving them a score and there's no action, who cares, right? What it's like what do I have to do? And that action itself has to be clear. I've worked in the software industry for quite a while and every time I have to find a bug in our software I have to write a ticket because the software engineer says well how can I recreate this and so you have to not only make it concrete but you have to make it actionable that is you have to describe exactly what the problem is and you know since the foundation are test results there's kind of this equivalence principle right that data quality test results or workflow tickets and a dash dashboard are kind of the same. Whe you could express a dashboard as a series of tickets in workflow or you could express express it as you know a numeric sum and so you must be able to con you know a score must be could be able to convert into workflow tickets and workflow tickets should be able to be converted into a score. And so that means that every data quality ticket must be supported by a test, and so this is our perspective on it because it's an influence challenge. It's tragedy of the commons. The only way you're going to get influence is to say do this discreetly and fix this. So if you base your data quality in tests and results and have the scores calculated from the test results, you're going to get that action. And so I think that sort of answers hopefully it answers the attendee's question. Because that equivalence principle is like is there that's how you identify where the source is because the ticket is on a particular table in a particular column that's the root cause. So let's go on to the next slide, and I guess we kind of believe that assessing data quality or building dashboards because of our background in agile methods and DataOps and lean that don't spend months building a data quality dashboard. Don't analyze and plan and design and build, and so if you get an assessment or a dashboard here at the end, it could take months.
00:33:52 Iterate in small assessments instead of running a waterfall
Chris Bergh: A traditional sort of waterfall way to do it. We actually believe it's much better to iterate quickly, do small assessments with less data elements over time. And what that means is that this sort of we've applied our background in agile and lean methods to data quality. And so we've written some papers and talked about this is sort of get a data quality process working right away. Sort of generate the rules, data quality rules automatically, generate the tests automatically. Start measuring and evaluating data qualities before you're sure that all the standards are right and then kind of use the measures itself to help establish this the standards where there aren't any. And by cycling quickly by saying I can and this is the principle that we built our open source software on that in an hour I can connect the data sources profile the data build a bunch of data quality tests and then get a score and then start tweaking is so much better than I think the exercise in word documents and powerpoints because people will react to things that are real and so get something real quickly and what that means is that this scope here of learning and value.
00:35:03 Five data sets assessed, or really only two
Chris Bergh: What happens in a waterfall thing? All your learning and value is is sort of pushed to the very end. And all of this is really about individual empowerment, individuals giving value quickly, and individuals learning. And if you can learn more, you're likely to do things that really matter. And so here, notice I've had some different data sets assessed on the bottom. Data set one, data set two, three, four, five. I've assessed five data sets in that time, but in a wonderful way, I've only assessed two. And I think that's or the worst is I could have assessed a bunch of data sets that don't matter at all because I've taken so long to do it. And I think that waste is is another way that you avoided by this iterative cycle, and so let's see. So some comments. So An attendee says the first three types feel quite similar. Business goals tend to drive reporting which is usually how CDI are identified. DQ dimensions help us define measurability for data quality but still given by the business scores.
00:36:08 Business goal versus CDE, and getting people on your side
Chris Bergh: Yeah, I think that's c certainly true. Like there is kind of a philosophical aspect of how you say something's a business goal versus a CDE because your CDE is a business goal. But in in cases where companies don't haven't defined CDEs attaching something to a business goal and and these words source focused, business goal focused, customer focused, they're just I think themes to help you identify how you should get your dashboard. I it's really you are in trying to get people to help you influence data quality change. You're a a lone person out there and how do you get a band of people together to help you? And that's really I think the the goal here in each type of dashboard is is a method to help get a team together to to to do data quality and and by focusing on what that team needs, a specific customer, a specific business goal, a specific regulatory report. So let's let's keep going. So what I'm going to do is just demonstrate our open source software here. And to do that, I'm going to switch and share this screen. And so this is something that our open source software, it's Apache 2. You can download this full application here. Again our philosophy is it's a fully functioning software contains everything that you see today. Allows you to score and build as many dashboards as you want against as many database connections. And so let's talk about this. So what I've done in in the tool is actually built these dashboards. I have a business goal dashboard here in the upper right. I have a CDE dashboard. I have a data consumer dashboard. I have a data quality dimension dashboard and a data source dashboard. And so let me let me talk a little bit about that by by going in. And so let's first look at this business goal focus dashboard. And so I drill into it and I can see I've got some tables, but I've also got some columns. And so here's what I'm I'm our belief is that you need data quality needs to be made up of discrete issues and these things are driving the score.
00:38:18 Demo: profiling a sale price column and chasing the warning
Chris Bergh: So like if I look at this the the price here, the sale price, and I drill into this, I can see, okay, what's the sales price? I can kind of look at my profiling for this and see the distribution. And then, well what's happened here? Hm. So, I've got a warning on this. So let me drill into that to see what happens. And so here I've got the sale price and the distribution has changed. So it's actually kind of gone up. So people are giving sort of more discounts. I can look at the source data here to see what happens. And so something's changed here that I don't really know. And this could be an indicator that some people are giving more discounts or an indication that there's a statistically significant shift in the percentage of unique values versus baseline. And so that something's funny in the data. And so that means we should investigate. Perhaps it is right and we can disposition it saying it's it is an issue or it's not an issue up here in the upper right.
00:39:17 Demo: the same data read from a CDE standpoint
Chris Bergh: But it's a way for us to kind of drill into the data and look at it. And this kind of let me go back and share this this this tab. Being able to kind of go in identify specific issues and then see what's what happens here. And so I can look at this from a CDE standpoint. And so the way that this also works is we if I drill into it, if we identify where things are specific columns that are CDEs, each one of these like customer ID, first name, gender, income level, last name, they've been identified as CDEs. And so here we actually they're all pretty good. They don't have any sort of hygiene issues or data quality test failures. And let's go back and look at some of the other dashboards. And here we have sort of a typical data damma data quality one if I drill into that. And so again each one of these are have specific columns. We can look at these for the columns.
00:40:15 Demo: dimension scores, and the score trend over time
Chris Bergh: We can actually look at them by the data quality dimensions and see how they contribute and then we can kind of drill into each one as before. And here we have a score trend. It's not very illuminating in this example but this is you know the score trend increasing is kind of I'm it shows the effect of a data quality person because as what you want is that upper and to the right improvement because that's the I'm as a data quality person I'm awesome I've made an effect and it's improved over time and what that means is either someone's fixed the source data or you've gone in and looked at these issues and said, "Well there's a whole bunch of issues here that go into data quality dimensions. Should I disposition these in a different way? Does the product ID or the total amount or the price is that really relevant?" And so I think that's an answer here that you have to kind of go through and do this. But that's there's sort of two ways to fix the score.
00:41:12 Fix the data or fix the tests, and managing generated checks
Chris Bergh: One is to fix your data and the other one is to fix your tests acting upon the data or your checks. And that's where this iteration can come through. You can actually work on both at the same time, and I think that's having data quality checks means managing them and our system automatically builds a whole bunch of data quality checks. There's several dozen. It just scans and creates automatically, but it also you can build custom tests and and change them yourself. And so kind of going back to this quality dashboard, we can see, we've got one by data source. And so what we've done here is we also have a data catalog. And in this just sort of where this comes from, we have metadata that's defined by different data sources. And so we can kind of go in and say, okay, here's the bike suppliers. We can look at this and we can say, okay, this data product is supplier data. Everything in here is comes from supplier data.
00:42:04 Demo: filtering a dashboard by data source metadata
Chris Bergh: So this has all been set with some metadata that describes this the data in this column as coming from supplier data and then we use that as a filter here into the in into the quality dashboard. So what we go in and and our data source dashboard if I want to go in this and edit it I can go in and say okay I've got it's for this table group which is our way of identifying what database you talk to and what source system here it's this is pro happens to be product master and I can change this configure it how however however much I want and I can you know show it in different ways I could actually say I want to you know I want to group this by source system etc. So it's a very flexible way to create and make a make a dashboard. So, I'm going to stop sharing this and go back to the presentation. And then we'll answer some questions. So so just to conclude, you know, I so what are we trying to do here?
00:43:11 Data quality denial is cultural, not technical
Chris Bergh: Basically, we're trying to fight data quality denial like a lot of people or they kind of, you know, data quality issues have have for a long time they've just gone unaddressed and I think partly it's because people don't have tools to do this and that's what we've tried to do with our our tool is given individual us some superpowers through some open source software. But on the other hand, it's there's these cultural issues, right? And denial is cultural, right? Ah we don't have everything works you know I'm I'm not getting yelled at so things are great and but that's a cultural thing not necessarily careless and there's a lot of superficial stuff that that aren't grounded in action and a lot of tools out there that just give you a score and like so what I got a score what exactly should I be doing here and so we believe that that's really important to be grounded in specific actions and and share those actions in the details of those actions and so we have what's called initi report and and ways to get at the details in the UI.
00:44:09 Workarounds and silence are the dangerous signs
Chris Bergh: And I think these and and you know when you see workarounds and and and sometimes when you don't see any complaints, they're actually dangerous signs that people have just given up or they've put a data quality system on top of your data and they're not telling you. And so someone else is fixing your problem. And so I think this idea of dashboards and observability can break the cycle, right? By making problems visible and traceable and actionable. And I think again that's only as if they link the metrics to real actions that people should take. And to kind of reiterate our final our final belief here is that basically, you know, if you're a data quality leader, you know, how do I deal with all this data data and how do I influence change? That's really the the core problem. And our belief is that you get going quickly, you focus on a specific customer, you pick the right type of dashboard, you iterate, you improve, you get influence, you give concrete action, and you start with, you know, maybe you don't use our open source, maybe you do, but it's a great place to start.
00:45:08 Resources, then Q: what are the terms of the software?
Chris Bergh: And we've got a whole bunch of resources for you. There's a link here to install TestGen. If you move on from TestGen, we have and you want to start not looking at data but also looking at your ETL and BI tools. We have an observability tool that does that. And we have a a data quality 7-hour data quality certification and another three-hour certification on DataOps and and of course several books. So I want to thank you for that. I hope this was valuable for you. We had a lot of people here today. And let me go back and I'm going to stop sharing here and I'm going to look at the questions. And Gil, is there any particular questions I should answer?
Gil Benghiat: You know, there there's a lot of good a lot of good questions in there. You can maybe just go up from the bottom might be the easiest way
Chris Bergh: Okay, so what are the what are the terms of the software?
00:45:51 Apache 2, one login and one connection, $100 a month beyond that
Chris Bergh: So the the software itself is what's called Apache 2 open source. That means basically you can do whatever you want with it, there's some terms, and so how we make money is we focus on giving everyone a full featured an individual a full featured software. So the UI, all the data quality rules, the AI that generates all those rules, the the profiling engine is all there, and but we limit it by one user login and by one database connection and if you want more users and more database connections, it's $100 US per month, so that's our our business model, so an attendee's got a question. Each dashboard type are different in terms of which data points they have. So the main difference in the inventory of the attributes, not the technique or the dashboard. I I I think so. I mean, I think I think of the world as you've got a bunch of tables with columns and then on top of that, you've got a bunch of data quality rules that have found problems in in the columns.
00:46:46 Q: how do you link a dashboard to a trust or maturity model?
Chris Bergh: And so what we're in the process of is aggregating those those those rules or specifically the results of those rules running into a dashboard. And so the the question is not so much what's the cool graphic design of the graph dashboard, but how you're taking those columns and their tests and putting it together to make action in the organization. And so so I guess an attendee's question, how could you link a data quality dashboard to a trust or maturity model that gives us easier context to business for understanding and assurance?, you know, I think the in some ways isn't isn't a data quality dashboard a trust score? In my mind, isn't the result of a score on a dashboard and and the aggregate score on that dashboard this a proxy for trust or and isn't the improvement of those a so and so I guess in my mind a maturity model on data quality has got a lot more things than just a score right there's there's the programs and the process and the understanding.
00:48:01 A score is a point on a maturity model, and Q: where does profiling fit?
Chris Bergh: And so a score is kind of a point on a maturity model. And so there's a whole program that has to be put in place to to do data quality. And and you know our technology can help that and be a sort of a point on a maturity model or a trust score that fits into a maturity model, but it doesn't replace a maturity model. There's other aspects that and processes that people need to use and Okay, so there's a lot big question here. How can dashboards be totally actionable unless someone performs data quality assessment or root cause analysis based on profiling results to figure out various actions taken up for improving data quality? Do you think that statistical column profiling insights, counts, range, distribution, doesn't have a place in these six types? No. Oh, I I guess I just I I think those I think every dashboard, like I said, has to be backed up by a series of data quality tests and their test results when applied to the columns.
00:49:01 Tests get you 70% of the way, then who can you get on your side?
Chris Bergh: And our tool profiles the data, learns the data, and automatically generates data quality tests in addition to giving you the option to to do manual ones. Our whole goal in tools sort of get you 70% of the way there so you have a broad range of data quality tests and then make it easy to to tweak them and easy to add custom tests. And so what makes action is the test results and and getting people to fix them. And there's two sets of of things to do that with the test results. One is you can tweak the tests, say this doesn't matter or improve it or you can deliver the test results in an actionable format to another person to fix. And so that's our our sort of view on it is is action comes and that's why we built the product along the ideas that you suggested. So An attendee has got a question. What kind of dashboard will be the best suited when they are implementing data quality for the first time and have done only statistical profiling based on some ha based on same created business rules. Well, I I I I think I put it in another way. How are you going to get people to listen to you? And who can you put on your side in the organization who's going to help get stuff done? Because I I don't think, you know, as an engineer, I've spent a lot of time in my career kind of focused on the technical and built these big technical artifacts and then they haven't done anything in the world. In order to do stuff in the world, you need to solve things extrovertedly. You need to get people on your side. So who are you going to get on your side? And what's going to cause them to be on your side? Is it business goals? Is it their report that they're generating?, is it some CDEs?, is it that they run a data engineering team and they want their data source? How are you going to get people lined up with you? And what's the leverage to do it?
00:50:49 Q: reconciling customer numbers between two systems
Chris Bergh: And then how are you going to make that actionable through through tests? That's the way I think about it. And an attendee has a question I'm not sure I can answer. He says, "In case where you have a lot of reconciliation between two systems, which one will you recommend? Verifying customer numbers from two systems. It's such a context dependent case when you've got two systems. I mean one is it is it the case where you're trying to move data from one system to another? You're doing data migration or you've got the same data in two systems and then how you reconcile it. I think it sort of depends on on which case, and then like if you've got two systems with the same numbers, which one wins?, that's a really good question and I think the the the best cases answers that is to have a have a test that's that you'll run against both systems that sort of contains what should the what should the numbers be?
00:51:40 Why that one has no technical answer, and Q: data sitting in S3
Chris Bergh: But I can't answer it because it has to also do if you've got two systems that have customer records and which name wins. That's a sort of a a question on business context that is sort of hard to answer from a from a technical standpoint. It has to do more with the the data quality program that you put in place, so I think that's it. I don't see any more. Gil, do you see any more questions?
Gil Benghiat: Yeah, there was there's one coming at at the end, what do you do if your data is sitting in S3 and it's not in a warehouse?
Chris Bergh: Well don't use our tool because we need a database to run our stuff. So I think you know we our tool doesn't copy data, it injects data into a database. So there are ways that you can a lot of databases now have the ability to read directly from S3 and create sort of ephemeral tables and so I know Redshift, I know Snowflake, I know a bunch of databases has that.
00:52:36 S3 needs a query execution engine; the tool needs structured SQL data
Chris Bergh: So for our tool you need to have a query execution engine to run it. And so a lot of databases you you basically reflect the S3 bucket into a table in a database and then you profile it and run it. And that's the way I I would work. We're not a believer that you should have a def a definite separate data warehouse that you copy all the data into just to do data quality. That tends to be more of an expense, but I know other solutions do it. So it could be it could be debatable.
Gil Benghiat: Yeah. And and just just questions about our tool. It's it it works on you know structur structured SQL data. So so if it's in a file or one drive etc. It it it's not accessible. So it needs to be in one of the supporting databases or supported databases or loaded into a supported database and run it from there. And I think there there's a lot of concern about getting copies of the slides.
00:53:27 Where the slides and the recording will be posted
Chris Bergh: Yeah. So, I what I'll do is I'll put up a little web page. It'll have the slides. It'll have the recording, and here I I'll if you want them, I'll just put the I'll put the Google Drive link here in the in the chat so you can go. I'll have to give you access, but I by the end of the day or by tomorrow, I'll have the website up that'll have you you'll get you'll get an email with it.
Gil Benghiat: Yeah, we usually, you know, anybody who wants the slides who signed up, anybody who signed up for the webinar, we send an email with a link to the slides and a recording, so I shared a moment ago the I I'll put it in again. Put make sure we have your email and we'll just send you the we'll send you the slides.
00:54:06 Q: how does the tool generate tests without domain understanding?
Gil Benghiat: There's the There's how you can send us your email.
Chris Bergh: All right. Well thank you so much everyone. I really appreciate the questions and the activities. I hope this was hope this was valuable for you. And so there is one more question. Gil should I okay I I'll try to answer it here while people are leaving. So so when you say data quality test how exactly the tool generates tasks to improve data quality without domain understanding and context is it totally rule-based or procedural etc. So I guess kind of answering it it is a way it uses some algorithms but not a large language model to to build a whole series of data quality tests based on profiling. Our goal here is that we want to get the basics. We want to give you the syn we look at the syntax of the data and not the semantics. So if we generate a whole bunch of tests, we're getting you 50 60 70 80% of the way there and we give you a UI to when you have the sort of domain understanding it gives you the time to do that. And so you know we in our data work we have some domain specific tests that we use. So we got a whole bunch of tests that we use for pharma companies to check to check ID numbers that we that we use but we don't have a feature yet to kind of share those among other people, but that's our goal is like get the syntax of the test done and then give you time then to focus on the domain specific rules. We want to get you 70% of the way there so you can do the 30% that really matter and that's the domain specific rules, and so that's our view. All right. Thanks everybody. Hope you have a great day and we'll follow up with all the information. Yeah. Thank you Gil. Thanks. Bye.
Machine-generated transcript, lightly edited: filler words removed, product and speaker names corrected, and audience members anonymised. Two things were dropped — the recorder bot's announcement, and a stretch where an attendee's microphone bled into the presenter's audio and the words interleaved unintelligibly. The audience questions were asked in the chat window and read aloud by the presenters, so they appear here in the asker's words. Timestamps are scaled to the published runtime: the source transcript's clock ran about three and a half minutes longer than the recording.
Questions from this session
How do you identify that a ticket's root cause is a data quality problem?
By building the dashboard out of actions rather than scores. The scores that make up a data quality dashboard are rooted in data quality test results, so you should be able to click through from a score to every action behind it. That is also why every data quality ticket has to be supported by a test: a ticket saying the quality of a column is bad gives the person receiving it nothing to act on.
What are the licence terms for the software, and what does it cost?
It is Apache 2 open source, so you can do what you want with it. An individual gets the full-featured product — the UI, all the data quality rules, the AI that generates them, and the profiling engine — limited to one user login and one database connection. More users or more database connections is $100 US per month.
How do you link a data quality dashboard to a trust score or a maturity model?
An aggregate dashboard score is already a proxy for trust, and it is a point on a maturity model. It does not replace one: a data quality maturity model covers the programs, the process, and the understanding as well as the score. The technology can supply the score and its trend over time; the rest of the program still has to be put in place around it.
Do statistical column profiling insights — counts, ranges, distributions — have a place in these six dashboard types?
Profiling is the input, not the dashboard. Every one of the six types has to be backed by a series of data quality tests and their results applied to specific columns. TestGen profiles the data, learns it, and generates tests from that profile, and it is the test results — tweaked, or handed to someone in an actionable format — that produce action.
Which dashboard type should you start with if you are implementing data quality for the first time and have only done statistical profiling?
Chris Bergh reframed the question on the call: it is not which chart, it is who will listen to you. Ask who you can get on your side in the organisation and what would give them a reason to be — a business goal, a report they own, a set of critical data elements, or a data source their engineering team depends on. Pick the dashboard that matches that person, then make it actionable through tests.
What if your data is sitting in S3 rather than in a warehouse?
TestGen needs a query execution engine, so S3 on its own does not work. The route is to reflect the S3 bucket into a table in a database — Redshift, Snowflake, and others can read S3 directly and create ephemeral tables — then profile it there. The tool works on structured SQL data in a supported database, so files in S3 or OneDrive are not reachable as they stand. DataKitchen's position is that you should not stand up a separate data warehouse and copy everything into it purely to do data quality.
How does the tool generate data quality tests without understanding the domain?
With algorithms rather than a large language model, working from the profiling results. It reads the syntax of the data instead of its semantics, which gets you somewhere between 50% and 80% of the way, and the UI is there for the domain-specific rules that need a person. DataKitchen keeps its own domain tests — ID-number checks for pharma customers among them — but there is no feature yet for sharing those between users.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.