On-Demand Webinar · 1 hr 12 min
A Masterclass in the Six Types of Data Quality Dashboards
Six kinds of data quality dashboard, what each one is for, and which person it is supposed to move. A data quality leader has influence rather than authority, so the dashboard that works is the one aimed at somebody who can act.
What you'll learn 6 points
- There are six dashboard types, not one: dimension-focused, critical data element, business goal, data source or provider, data consumer, and ticket-driven workflow. The right answer is usually several of them at once.
- A data quality leader has influence, not authority. The type of dashboard you build should follow from which person you need to move — a VP of sales, a CFO facing compliance reporting, an ML team, or a chief data officer.
- One senior practitioner DataKitchen interviewed said dimension dashboards "do nothing to motivate people to make improvements to their data". A score of 70% consistency leaves a data engineer with nothing specific to fix.
- Every dashboard has to be built out of test results against a named table and column, so a reader can drill from the score to the failing check. Anything less arrives on someone's desk as "fix the accuracy" and gets skipped.
- The tornado report ranks each data provider by issue severity and the hours your team spent patching their data, which turns an invisible cost into something you can put in front of the supplier.
- Don't spend months on a dashboard. Profile the data, generate tests automatically, publish something small, get one column fixed, then rescore — the loop is where the learning and the influence come from.
Prefer to read it? The written version is in Webinar: A Masterclass In The Six Types Of Data Quality Dashboards.
Slides
Transcript
Show chapters and dialogue 43 chapters · 8,200 words
- 0:00 Welcome, and what makes a data quality dashboard useful
- 4:24 Agenda, and why improvement starts with one empowered person
- 5:33 Influence, a specific end customer, and small iterations
- 6:49 Why you need a chart that shows quality improving
- 7:49 Data quality as a tragedy of the commons
- 8:57 Not all dashboards are equal: align one to a person
- 10:04 The six types, and what a data quality dimension is
- 11:22 “Dimension dashboards are crap”, and why they don’t motivate
- 12:22 Type 2: critical data elements, and where they come from
- 13:49 The bank that got fined, and why CDEs became a senior goal
- 15:09 Defining “critical” outside a regulated industry
- 16:19 Type 3: business goal dashboards
- 17:29 Thirty columns out of ten thousand, in the language of the business
- 18:58 Tying quality to bonuses, and type 4: data source dashboards
- 20:27 The tornado report: severity and hours spent per provider
- 21:31 Fifty to a hundred suppliers, and how to get leverage
- 22:36 Providers are often glad you found it, and type 5: consumers
- 23:43 Influence by proxy through a data scientist’s 40 model inputs
- 25:03 Human limits, and type 6: a ticket list rather than a dashboard
- 26:00 Counting tickets the way a software manager counts bugs
- 27:01 Dashboarding ticket flow, and what a closable ticket needs
- 27:57 Tickets measure work done, not data quality
- 29:02 Why the answer is several dashboards, not one
- 30:05 Matching the dashboard to the customer: sales, finance, ML, CDO
- 31:20 Rules of thumb for when to use each type
- 32:25 Every dashboard has to be built out of test results
- 33:32 Why “fix the accuracy” is not an actionable request
- 34:33 Iterative delivery versus a waterfall assessment
- 35:52 Doing just enough, and the iterative playbook
- 37:15 Profile, generate tests, dashboard, fix, rescore
- 38:26 Demo: five dashboards in the open source tool
- 39:22 Drilling into a shift in the sale price distribution
- 40:16 Dispositioning an issue, and the CDE dashboard
- 41:14 The DAMA dimension view, and reading a score trend
- 42:14 Two ways to move a score: fix the data or fix the tests
- 43:04 Data source dashboards built from catalog metadata
- 44:26 Why data quality stays broken, and what breaks the cycle
- 45:38 The two challenges: data at scale, and influencing change
- 47:01 Q: can I group critical data elements by dimension?
- 48:14 Q: adapting the six types to LLM fine-tuning data sets
- 49:28 Q: fine-tuning on raw data that hasn’t been validated
- 50:37 Q: which dashboard types are most popular with clients?
- 51:45 Q: 20 million records — should I test all of them?
00:00:00 Welcome, and what makes a data quality dashboard useful
Chris Bergh: Hello everyone. I'm Chris Bergh. We will start in a second just to give us about 30 more seconds. See if more people join. All right, let's start. My my name is Chris Bergh. I'm one of the founders of DataKitchen and I'd like to talk to you today about data quality dashboards and what makes them great, what makes them useful, what makes them impactful. I have a long history in data engineering and a technical person and I run DataKitchen and I'll and so from a kind of logistics we will share slides and recordings of this probably tomorrow you'll get an email and I want to welcome everyone today. We've got 39 countries who who have signed up for this. And on the right hand side of your Google Meet, there's a the button that's kind of third from the lower right hand corner. That's the chat button. If you click on that, you can put your questions in there and I promise to either answer them at the end or answer them if I have time or an appropriate place during the discussion.
00:04:24 Agenda, and why improvement starts with one empowered person
Chris Bergh: And today's presentation should take about 45 minutes. I may go a little long. Sometimes I like to talk. So the agenda for today is we're going to talk about why a data quality dashboard and what are the six types of data quality dashboards and when do you use each type of dashboard and we'll give an example of how to create a data quality dashboard or each type. We'll go through a demo and then we'll have a conclusion. So let me talk through kind of I have a cold today so hopefully I won't be too hoarse. But our view on data quality improvement and this is our perspective that we've gotten you know having worked in the industry for quite a while and and talking to a lot of different data quality people. So first is that improving your data quality starts with a single empowered individual driving data quality change. It's not about large projects. It's not about teams. It can be but in a lot of organizations it really has to start with someone.
00:05:33 Influence, a specific end customer, and small iterations
Chris Bergh: And so our view is it starts with a single empowered individual who's trying to make change kind of via their personal influence. They care about the company's data quality, their organization's data quality. They're trying to make a change. And to do that, they need to get the stamp of approval from a specific end customer a VP, a data scientist, a person who has a need for that data. And their influence in that specific end customer. That has to be done not in a long waterfall project way, but by starting small, iterating quickly using DataOps principles. And what we think is that it really isn't a philosophical descriptive activity. It's a specific concrete actionable activity like go fix this column this way. That's the sort of remediation that you need to have success. And then to do that, you need to continually measure your improvement. Why? Because you have a boss and your personal influence and your end customer can sing your praise, but you need to actually have a chart that says my quality is improving.
00:06:49 Why you need a chart that shows quality improving
Chris Bergh: It go it's going up to the right. And that's sort of why we're here today to talk about how to measure data quality. And then lastly, we think it's enabled by free open- source technology of which we have we have one that will demonstrate. And so we've talked a lot over this last year or two about data quality improvement. We've got one thing I want to do is we've got a certification series on data observability and quality. We've also got a lot of different great content for you to look at and just check out our blog and our our previous webinars. And so, in any data quality webinar, we've got to talk about why data quality is a problem. And probably everyone here believes what I'm going to say. But I just have to repeat it because we're in a data quality webinar. And so, of course, data quality can make things operations poor, h lead people to make poor decisions and erode your trust.
00:07:49 Data quality as a tragedy of the commons
Chris Bergh: And our view is kind of data quality is a tragedy of the commons. It's an ecological term meaning you know a shared field that everyone grazes their cattle on and so no one owns it but everyone uses it. And so like that person who says hey we should take turns grazing cattle in our shared field it's a challenge of influencing others to and data quality leaders typically they don't have power they only have influence they can't make people do things and and the difficulty here is how do you get influence to have someone to take action that's really what we're trying to help with here that that influence influence challenge. And so a data quality dashboard is part of that influence and it's a way to kind of make concrete the effect that you're trying to have in a specific domain and they provide a visibility into health of your data. Sometimes they're called reports or scorecards or assessments. We're using dashboards today. You know, you can pick your favorite word.
00:08:57 Not all dashboards are equal: align one to a person
Chris Bergh: But we don't think all dashboards are created equal. Because we think the effectiveness of that dashboard, scorecard, etc. Is really about how it's aligned with a person who needs the data to do something of value to your organization. And that person and their role and the context of the company you're in determines kind of the type of dashboard that you want to do. And and picking the right dashboard also will help you kind of maximize your impact. It'll streamline your efforts and and really help you will achieve sustained improvements in data quality. And so we are you know what we're going to talk about here is is a strategy of of specific dashboards focused on specific customers that help you iterate to drive your influence forward in an organization. So what you won't hear from me today is saying you know have one dashboard to end them all that fixes your whole organization. So to do that we're going to talk about these six types. And so what are the six types of d the dashboards?
00:10:04 The six types, and what a data quality dimension is
Chris Bergh: And so we're going to go through each one sequentially. And so we're going to talk about dimension focused dashboards, critical data elements or CDE dashboards, business goal focused dashboards, data source or data provider focused dashboards, data consumer focused dashboards, and ticketdriven workflow dashboards. So I'm going to walk through each and following a similar pattern in each one. So what is a data quality dimensions? Well, if you know about DAMA or other organizations, it's a way to say there's different attributes of data. Is it accurate? Is it timely? Is it consistent? Is it unique? And kind of grouping the attributes of those data together to highlight specific areas of concern. And what it does, it says, okay, we've got a problem with, you know, timeliness of data or consistency. And that's an area of concern and can help sort of thematically organize your thoughts and it does enable sort of a broad brush to cover lots of ground. And a lot of people do know about these standard groupings of the the data quality dimension and it does allow you to sort of look broadly across different areas.
00:11:22 “Dimension dashboards are crap”, and why they don’t motivate
Chris Bergh: But, you know, we we've talked to a lot of people and one of the people that we did some research on last last year or 18 months ago is said really that data quality dimension dashboards are crap. She said they they don't motivate people to make improvements to their data. And so why is that? Well, they're kind of abstract, right? It it this isn't about science. This isn't about an academic exercise. It's about influence. It's about getting someone to do something, getting someone who has their hands on their keyboard to actually be motivated to make a change. And so the sort of context a lot of the dashboards are are really not context specific. They don't really have an actionable insight. They become kind of like a check the box exercise. And sometimes the scores like 70% consistency don't resonate with business. And and and execs sometimes have a problem connecting these sort of philosophical metrics to what it means with business.
00:12:22 Type 2: critical data elements, and where they come from
Chris Bergh: And so what is the next step when you have 70% consistency? How do you actually make a change? And so one way some organizations have addressed that challenge by identifying what are called critical data elements. So if you look at the thousand tables that matter in your organization, maybe there's 20 columns in each one of those tables that really matter. And and sometimes organizations that have regulatory compliance like a like a bank or people who are doing pharmaceutical clinical trials they they know their critical data elements and they have to actually from a regulatory compliance report on those and so those are high impact fields. Other organizations they don't have those but it's kind of obvious to them what what are critical right sort of customers customer names credit cards things like that costs profits products things like that that are really critical to their organization and so you know these kind of CDE focus are really effective in regulated industries where you kind of get in trouble and where you don't actually report on the quality of these CDEs because you have regulatory compliance reports that are based on these critical data elements.
00:13:49 The bank that got fined, and why CDEs became a senior goal
Chris Bergh: So if you screw up the quality of those, you're going to screw up your reporting and then you're going to get fined. And so these kind of compliance focused things are actually focus important. And when you have someone at the higher level of the organization who sees it important or you've actually gotten a fine in your organization, it makes a it makes this a much bigger deal. And saying and also it by focusing on CDEs, it allows you to have allows you to not boil the ocean and say okay these are the critical data elements. Now, there's a whole process that may happen in your organization of defining what a CDE is. And maybe that comes from regulatory or maybe that's a meeting that you have to have with your team. But it does allow you to get more of a direct connection to your business and and compliance object objectives. And so when we talked to people we talked to a bank and they they had a really great CDE focused plan and because they sort of did get punched in the nose with some fines on their regulatory compliance and so they it became a very senior goal in the organization to get their data quality gathered to focus it on CDEs because they didn't want to get fined again and that fine
00:15:09 Defining “critical” outside a regulated industry
Chris Bergh: Was also going to stop them from expanding. And so there's for them it was not only that they had been fined, but also that there was a potential for them to to grow their banking, the size of their bank by doing by making sure their regulatory compliance because I think they said they wanted to buy another bank or something like that. So it was really connected to corporate goals. And you know, I think there there are challenges, right? If you're in a non-regulated industry, what constitutes critical? How do you get alignment on on what's critical? And critical may change over time. And you know, when you've got cases where what's critical to one person isn't critical to another person. And so you may actually have to your your CDEs may not be fixed. You may have to work from other people and through other teams and figure out what it is. So defining CDEs is is sort of non-trivial in non-regulated industries. But that's you know part of the part of the challenge in trying to to extend your influence.
00:16:19 Type 3: business goal dashboards
Chris Bergh: So we've talked about sort of the DAMA dimension type of dashboards. We've talked about CDE dashboards. And now we're going to talk about business goal dashboards. And this one's I think interesting because it really is about linking data quality to a specific organizational objective. And so and how can specific data elements actually affect the goals of a company? And so and it really helps you position data as strategic. So, for instance, if you talk about how to make this a if you've got a specific sales goal for product, a type of product, you've got a specific marketing goal, you've got a specific efficiency goal, that may all the way go up to the top of your corporate goals for the year. And what that means is that your data then becomes a way to measure the effectiveness of your goal. And then you can't actually tell if your goal is right if you don't have the effective the the necessary data. And so that can help you with buyin.
00:17:29 Thirty columns out of ten thousand, in the language of the business
Chris Bergh: It can help you with influence by saying okay out of those thousand tables and those 10,000 columns these 30 are actually really mostly focused to our corporate goals. And so and that helps you speak the language of of of business. And so because it this is after all an influence game, right? And how do you stop becoming a data quality nag and well you make data quality align with what people need to get their bonus to you know to to make sure that their end-of-year reviews go well. And so I think that's a way to motivate people and so people are motivated and sometimes companies will tie bonuses to the fact that you made the corporate goal and people can say look your bonus is 30% of your bonus has to do with meeting corporate goals. These are directly related to you having your bonus so maybe you should care about that. And I think that's a great way for you to kind of connect and influence people to make change because you know of all I've talked to so many data quality people who really care really want to improve data quality but this influence is hard and so what we're trying to do here with these data quality dashboard types is connect them to influencing modes that you can use and connecting it to business goals is a great
00:18:58 Tying quality to bonuses, and type 4: data source dashboards
Chris Bergh: Way because there's this alignment between data quality metrics and your business goals and that is a great part of doing it. But you know the the challenge with this is business goals change right and sometimes you've got to change what's important in data quality with the business. These aren't sort of fixed forever because businesses change often. So we've talked about now we're going on to our fourth type and this is a different one. So we focused on the dimensions, we focus on CDEs, business goals. Now we're talking about data source and this is a different one. This is not this is about measuring who is giving you the data, what system, what partner, what tool is giving you the data or even what t table. And if you've ever done data engineering, you know that some of your data providers are great and some of them are just problematic. And so how can you track source level quality by saying okay I've you know my ETL process is running fine but I'm getting short records or I'm missing columns or they're changing columns and so this chart on the left here is one that we've used in our consulting organization is called a tornado report and it it sort of tracks it you can see your sort of source by the provider and
00:20:27 The tornado report: severity and hours spent per provider
Chris Bergh: Looks at both the severity and the number of hours it took to actually patch the issue or deal with the issue. And this way you can kind of say look you're here is if you can see this green one supplier E or whatever is problematic and it's a way to sort of foster accountability on your providers because sometimes the data providers they forget that you exist and they're going to give you crappy data and this is a way for you to say look you've you've given me you know five se one errors which has cost my team 20 hours of time of work and it's a way that's one way to sort of get leverage that they've done problems to you and giving that accountability to your providers and saying, "Look, you're giving me poor quality data or you're giving me great data, but you're changing things." allows you to take something that's invisible and make it visible. And that's an important part of dealing with your data suppliers. And and oftentimes people who are on data teams are dealing with not one but dozens of different data providers.
00:21:31 Fifty to a hundred suppliers, and how to get leverage
Chris Bergh: And having dashboards focused on specific providers can help you get leverage to do them or explain to your boss why this one provider is so problematic and you know maybe it's a a a CYA thing. And so you know we've you know we've talked to a lot of different people and you know in in the pharma world the commercial pharma world they're dealing with 50 to 70 to 100 different data suppliers and and and oftentimes they just don't care. They'll make arbitrary changes and and it ends up that the data teams I have a feeling just have no control over the data quality. And so here's a dashboard that's about trying to get motivate not the organization to make change but your suppliers to make change sometimes internal or external to the organization. And but you know the the challenge is again with influence you don't these people may not work for you. They may not even be employed at your company. They may not have motivation. And you know you just may have to take it on the chin.
00:22:36 Providers are often glad you found it, and type 5: consumers
Chris Bergh: But that really means is that you need good tests. You need to tell if there's problems before it get gets into your into your environment. But oftentimes on the human beings are are are good not bad idea. I've often found that when you actually say here are the problems with your data, they're they're excited. They're like, "Oh, you found some problems. I better fix it." and they're happy with you. And so which is a surprise because oftentimes it means that you found it first and they have other people they're delivering the data to and and you're saving them problems. And so working with your providers sometimes can take a noisy problematic data provider and by having tests and having dashboards you can influence their you can influence them to being sort of a better data citizen. So going on to one more there's a data consumer focused dashboards and this is not necessarily like a business but let's say you have a person a data scientist who's really interested in their model and their model has 40 inputs 40 columns of inputs in well they care about the quality of those don't they?
00:23:43 Influence by proxy through a data scientist’s 40 model inputs
Chris Bergh: And so that specific model that specific person may care more about it and may help you influence the people who are actually caring and trying to take care of the data. And so it's really about saying I'm using it's influenced by proxy. I've got a a consumer, a a data scientist, a data engineer, a marketer who's using this data for specific work. And you want to make sure the data is good for them. And you know, it becomes a personal task. And and why? Well, because those person they're using that data for their role. They're using it to build a model or to analyze data or to export to whatever project they happen to be working on. And so by having these reports or views or models that are using the data, it gives you leverage by proxy to be able to go in and help the people who are making changes and that you know we found sort of doing research is that you know data scientists are maybe they're forever complaining about the quality of data right and and by saying I've got a dashboard you've got these 30 to 40 fields that are going to your model I'm measuring those and by focusing data quality quality reporting on those inputs to their models.
00:25:03 Human limits, and type 6: a ticket list rather than a dashboard
Chris Bergh: It can be very helpful for you getting again gaining that influence. And so you know the the challenges apply here, right? The the human part of data quality always remains. Well, what are those fields? Where do they come from? Are they actually does the data scientist or the person actually know? And if they've got a lot of them, can you do a Pareto analysis and find the ones that are most one? And then, you know, how many customers can you handle? So, it's it's not a perfect world here, but there are the sort of human challenges in any data quality project that I that are are important. And lastly, here's so we've gone through five types. This last one I think is really interesting because it's not a dashboard actually. It's a list of tickets. And really it's sort of using tickets as change requests and and sort of counting the number of tickets that are fixed or created in a certain time.
00:26:00 Counting tickets the way a software manager counts bugs
Chris Bergh: And it's a it's a way saying, okay, we don't have to say data quality is 70% good or 90% good. It's like, well, how many data quality tickets have we had? And that actually is is interesting, right? Because a ticket has to be based on something in fact, right? It has to be based on there's an issue. And so I I I I like this one, right? Because it's it's it it is a way to prioritize and assign and resolve. It's saying how many tickets have created, how many tickets have been closed. And it kind of is a reminiscent to me as a software engineer is bugs. It's how many bugs you're having, how many bugs are coming in. And way back in my day, 20 years ago as a software engineering manager and a waterfall team, I used to look at bug creation and close rates to be able to see if I could predict the end of a end of a release. And so and this is allows you also sort of to identify bottlenecks because what tickets are being created, how many tickets do people have?
00:27:01 Dashboarding ticket flow, and what a closable ticket needs
Chris Bergh: And and so we did talk to some people who aren't using really dashboards per se. They use workflow or tickets and they look about the number of tickets created and fixed during a particular period and they have goals on those tickets saying okay how many data quality tickets have you have you created and so this is you know you you can dashboardize your ticket flow like looking at number of tickets created and closed that's a a bar chart that you can that you can have and and any good ticketing tool like like Jira has that and so you know one of The challenges here is with tickets is you need good tests. You need to actually have a ticket that's closable. So when you look at the ticket, someone can actually open it up and say, "Oh, here's what I have to do to fix." another thing is you could create backlog that doesn't have you could have a lot of tickets. And then you've got to manage the ticket process, right?
00:27:57 Tickets measure work done, not data quality
Chris Bergh: On on how who gets assigned, it's done. And it's a measurement of work done, not of data quality. So that's a different perspective. It's it's saying, "Okay, I don't care about measuring data quality per se. I'm just trying to measure the amount of work done." and none of these are participle ones. You could have a ticket workflow and a data quality dashboard, but that's just a different perspective. So, kind of going on here, when do you use each one of these dashboards? So, what are some of the rules of thumb? So you know we've gone through each one of these the dimension focused ones the CDE ones the business goal focused dashboards data source focused dashboards data consumer and ticket driven workflow and so I guess there isn't one I think I I I really think a multi-dashboard a multi-ype dashboard is the best approach. It sort of depends on the context sort of who's your customer and and what's the problem you're trying to solve.
00:29:02 Why the answer is several dashboards, not one
Chris Bergh: And so in this case I'm I don't think there's one size. I I think having multiple dashboard types is really important. And having dashboards that are small and discreet and iterative and focused I think are are the trick to maximize your influence. And so if you're building dashboards that have 10,000 data items in, it's kind of cool, but you're really not going to get anything done and not have any change. And so, this is like trying to, you know, having one dashboard type is kind of like building a house with a hammer, right? You need saws, you need nails, you need drills, you need all the tools to go in. And this, this is kind of focusing on your customer is key. And so like let's just talk through this. So let's say I've got a VP of sales here in the upper upper leftand corner. And you know you your person who's going to fix it is a data engineer. Well the VP of sales of course carries cares about sales and so what things have to go with their sales program this quarter.
00:30:05 Matching the dashboard to the customer: sales, finance, ML, CDO
Chris Bergh: And so what are the actionable tasks? What are the things that have to be done to fix this? And so the the goal may measure this business goal of of you know the first half 2026 plan goals for sales but you've got to identify the specific data elements and the specific actionable tasks to fix. Likewise if you've got a CFO and they're looking at sort of financial compliance right there may be an accounting person who has to go into SAP to fix these things. And here you may actually look at critical data elements because they may be used for compliance or reporting. And then you may have someone on the ML team who's got, you know, some data that they're using for AI and maybe they're have their, you know, greatest LLM labelled integration. And so you do need a data c customer because that those fields that are be being put into the model or being put into the LLM are are really cared about by that ML that specific ML customer. And then lastly, you may you may have a chief data officer and he's just trying to look overall at how's our data quality and he may care about the thousand or 10,000 element data quality dashboard.
00:31:20 Rules of thumb for when to use each type
Chris Bergh: So that may be sort of looking in general how's our data quality do doing and that you may want to use a DQ dimension. So who your customer is what their need is having more than one is is certainly fine. And so the challenge now is like how do you actually when do you actually do it and so I think of use these dimensions ones when you're looking sort of foundational monitoring across large swaths of data. Use CDEs when you can identify CDEs, right? When you've got these CDEs that are either from a compliance or a legal reason or you can corral people into saying that these are your CDEs. Use business goals when you have business goals and use source or ticket driven when you're looking at specific systems. And if you can use a ticket driven workflow and you can agree that your organization works that way, I think that's a I I like tickets as well. So how do you create a data quality dashboard? So what let's walk through a couple examples and then we'll get to a short demo.
00:32:25 Every dashboard has to be built out of test results
Chris Bergh: So all dashboards have to be based in here fix this right they have to be based in a reality or a check or a test that is based on a specific table and a specific column of data. So they have to be discreet and concrete actions. They can't be sort of based on oh this column looks funny, right? Or I something's weird in this data. So data quality tests I think are the foundation of what happens in data quality dashboards and they have to be built from that. So it's not a measurement. And so we we have this sort of equivalence principle here that you've got a bunch of data quality checks. Well, you can use those to calculate dashboards and then those can actually be put into workflow tickets and all these things can kind of work together, but mostly that your dashboards are really built out of test results. That's the most important part here that it's not a that you can drill into a dashboard and actually see the specific test result.
00:33:32 Why “fix the accuracy” is not an actionable request
Chris Bergh: And so every data quality ticket should be grounded in a specific test result. And that makes it clear, it makes it actionable and you don't waste your influence saying giving it to a data engineer and they say, "Well, it's 70% it's 70% accurate. You've got to fix the accuracy." And the data engineer scratches their head and like, "I don't know how to do this." And they'll go on to the next thing, right? And so you've got to be discreet and clear if you expect people to actually change the source data to fix it. And and lastly, I think or another thing I think is don't wait for months to deliver a dashboard. So, a lot of typical dashboards are I've got to build a database and a warehouse and I'm doing it my part-time and I analyze it, I plan it, I design it, I build it, and you've got months. And so, for us, we actually think starting small, pick just a few data elements, do an assessment, add a few more.
00:34:33 Iterative delivery versus a waterfall assessment
Chris Bergh: Spoke and you know we think that this idea of iterative value really increases the amount that you can do and so this kind of funky diagram here if you imagine in a waterfall way I'm doing this over a period of months and at the end I have two sets of data D1 and D2 I've assessed and at the end I really get my learning and value like I've deployed my dashboard and I've started to get value from it I've started to learn what works and what doesn't work. And so the fundamental idea of doing things in an iterative and agile way is that you have much more opportunities for learning and value delivery. And so you're continually learning and improving and then you're contri contributing or continually improving the amount of data sets that you're trying to analyze or the amount of people you're trying to fix. So here you end up with D1 and T2 but in this sort of DataOps world you end up with five sets of of data that you've done and so smaller cycles more focused promotes your own personal learning promotes own value and a lot of times in projects the the tyranny isn't the the tyranny or the the success is not doing too much.
00:35:52 Doing just enough, and the iterative playbook
Chris Bergh: It's trying to do just enough. And what happens in large waterfall projects is you do too much and you end up with lots of waste. You have a an assessment of 30 data elements when only five really mattered. And so finding out what really matters and finding out where you can actually make changes, this iterative approach we we think really works for for teams. And so so what what's the playbook here? What's this iterative playbook? Well, we think that learning your data and profiling your data is important and and and trying to do a lot of things automatically. And so kind of start measuring and learning your data before all the standards are established like you can profile your data, you can understand your data with with our open source tool. You can generate data quality tests automatically and you can start learning kind of before all everything's done and what that means is you can kind of very quickly build a data quality dashboard and and and show it to someone and start getting that learning and value quickly and start trying to get problems fixed and seeing if you can track the changes and then this idea of I want to generate data quality tests I I generate a dashboard.
00:37:15 Profile, generate tests, dashboard, fix, rescore
Chris Bergh: I share those changes that I want to make. The people who own it, I'm trying to influence them. I update the data. I reprofile the data. I get another score. This cycle is what we're trying to engender is like I learn the data. I test the data. I see what's in the dashboard. I take concrete actions to fix it. They're fixed. The data is updated. And I see my chart move up to the right in terms of improving the score. And this cyclic iterative process is what we call the DataOps way to data quality. And so I just like to go and share a sample share a quick demo. And so to do that, I'm going to I'm going to share a different screen and then we'll go And so this is something that our open source software it's Apache 2.0 you can download this whole application here. Again our philosophy is it's a fully functioning software contains everything that you see today.
00:38:26 Demo: five dashboards in the open source tool
Chris Bergh: Allows you to score and build as many dashboards as you want against as many database connections. And so let's talk about this. So what I've done in in the tool is actually built these dashboards. I have a business goal dashboards here in the upper right. I have a CDE dashboard. I have a data consumer dashboard of a data quality dimension dashboard and a data source dashboard. And so let me let me talk a little bit about that by by going in. And so let's first look at this business goal focus dashboard. And so I drill into it and I can see I've got some tables, but I've also got some columns. And so here's what I'm I'm our belief is that you need data quality needs to be made up of discrete issues and these things are driving the score. So like if I look at this the the price here or the sale price and I drill into this I can see okay what's the sales price?
00:39:22 Drilling into a shift in the sale price distribution
Chris Bergh: I can kind of look at my profiling for this and see the distribution and then well what's happened here? Hm. So, I've got a warning on this. So, let me drill into that to see what happens. And so here I've got the sale price and the distribution has changed. So, it's actually kind of gone up. So, people are giving sort of more discounts. I can look at the source data here to see what happens. And so, something's changed here that I don't really know. And this could be an indicator that some people are giving more discounts or an indication that there's a statistically significant shift in the percentage of unique values versus baseline. And so that something's funny in the data. And so that means we should investigate. Perhaps it is right and we can disposition it saying it's it is an issue or it's not an issue up here in the upper right.
00:40:16 Dispositioning an issue, and the CDE dashboard
Chris Bergh: But it's a way for us to kind of drill into the data and look at it. And this kind of let me go back and share this this this tab being able to kind of go in identify specific issues and then see what's what happens here. And so I can look at this from a CDE standpoint. And so the way that this also works is we if I drill into it, if we identify where things are specific columns that are CDEs, each one of these like customer ID, first name, gender, income level, last name, they've been identified as CDEs. And so here we actually they're all pretty good. They don't have any sort of hygiene issues or data quality test failures. And let's go back and look at some of the other dashboards. And here we have sort of a typical data DAMA data quality one if I drill into that. And so again each one of these are have specific columns.
00:41:14 The DAMA dimension view, and reading a score trend
Chris Bergh: We can look at these for the columns. We can actually look at them by the data quality dimensions and see how they contribute and then we can kind of drill into each one as before. And here we have a score trend. It's not very illuminative illuminating in this example but this is you know the score trend increasing is kind of I'm it shows the effect of a data quality person because as what you want is that upper and to the right improvement because that's the I'm as a data quality person I'm awesome I've made an effect and it's improved over time and what that means is either someone's fixed the source data or you've gone in and looked at these issues and said, "Well, there's a whole bunch of issues here that go into data quality dimensions. Should I disposition these in a different way? Does the product ID or the total amount or the price is that really relevant?" And so I think that's an answer here that you have to kind of go through and do this.
00:42:14 Two ways to move a score: fix the data or fix the tests
Chris Bergh: But that's there's sort of two ways to fix the score. One is to fix your data and the other one is to fix your tests acting upon the data or your checks. And that's where this iteration can come through. You can actually work on both at the same time. And I think that's having data quality checks means managing them. And our system automatically builds a whole bunch of data quality checks. There's several dozen. It just scans and creates automatically, but it also you can build custom tests and and change them yourself. And so, kind of going back to this quality dashboard, we can see we've got one by data source. And so what we've done here is we also have a data catalog. And in this just sort of where this comes from, we have metadata that's defined by different data sources. And so we can kind of go in and say, okay, here's the by suppliers.
00:43:04 Data source dashboards built from catalog metadata
Chris Bergh: We can look at this and we can say, okay, this data product is supplier data. Everything in here is comes from supplier data. So this has all been set with some metadata that describes this the data in this column as coming from supplier data and then we use that as a filter here into the in into the quality dashboard. So what we go in and and our data source dashboard if I want to go in this and edit it I can go in and say okay I've got it's for this table group which is our way of identifying what database you talk to and what source system here it's this is pro happens to be product master and I can change this configure it how however however much I want and I can you know show it in different ways I could actually say I want to you know I want to group this by source system etc. So, it's a very flexible way to create and make a make a dashboard. So, I'm going to stop sharing this and go back to the presentation.
00:44:26 Why data quality stays broken, and what breaks the cycle
Chris Bergh: So thanks we went through that demo. So now we're going to go on to the conclusion. So but you know data quality is a challenging problem right because of this idea of that teams sort of deny that it exists or don't want to pay attention to it or tragedy of the of the commons. It's kind of a cultural issue. And and a lot of teams, they're not careless. It's just sort of they don't lack the tools to get it done. Or it seems too big, right? And and one way to sort of do it or is to say these superficial overly broad dashboards that aren't grounded in action that somehow it people it's going to magically be fixed if I've got this sort of broad and I don't think you know the fact that when you're doing sort of workarounds or lack of complainings aren't aren't really a good sign if you can measure data quality you should be able to improve it and so you know I think this sort of denial or tragedy of the commons can really be broken where you can make it visible and traceable and actionable and by linking it to directly to data quality tests.
00:45:38 The two challenges: data at scale, and influencing change
Chris Bergh: And so you know our our belief here is that data quality leaders have these two challenges right how do I deal with all this data at scale and how do I influence change to improve data quality it really comes down to that and and for us it's about get going quickly focus on a specific customer pick the right kind of dashboard start small iterate improve influence give concrete actions and and you know use some open source or or or do it yourself. And you know we think that these six type of dashboards are really are really great and and you know picking the right one is is part of part of the journey. And so lastly we got some links here at the end of this on where to install TestGen or our open source observability. We have a data quality certification and a DataOps certification if you're interested. And so why don't I go and jump in on some of the questions that I've seen and so the first question is was how about dimension focused for critical data elements or attributes and so I I I think what that means is can I just take my CDEs and organize the organize them by the data quality dimensions and and I think that's fine.
00:47:01 Q: can I group critical data elements by dimension?
Chris Bergh: And I think however you want to organize the data elements where there are problems and sometimes being able to say I want to look at it from the CDE standpoint. I want to look at it from the table standpoint. I want to look at it from some other attribute. However you group your dashboards together I think is is important, right? And so that's one of the reasons we've built our our tool with such flexibility. You can kind of pull the tests and the columns and say well I want to group them by DAMA dimensions or I want to group them by table or I want to group them by some other bit of metadata about the data and so I I think grouping is fine it's it's if that motivates change then that's fantastic so to me I think the answer is is yes with whatever motivates and makes your influence and and action happen and So some some two other questions and so how would you adapt the six dashboard types to monitor quality specifically for LLM fine-tuning data sets where issues like hallucination bias and inconsistent labeling are as important as values and schema drift.
00:48:14 Q: adapting the six types to LLM fine-tuning data sets
Chris Bergh: Well I think there's sort of two questions here, right? One is that the output of an LLM is is always going to be fundamentally wrong, right? And and so it's probably going to be 80% right or 85% right. And so the input to that is is is important and like if you improve the quality of your input, the expected value of the quality of your output is going to increase. And so being able to number one make sure that you have high quality data going into your LLM and then having high quality metadata about the data. So things like data catalogs or profiling data or test results can also improve the output. I think there's a separate question that I'm not sure I'm the right person to answer is like how do you tell a hallucination from a real answer or bias? That's a that's sort of beyond the scope of of our talk today. But I do know that that LLM's that internal your internal data plus an LLM if you look at the combination of those improving the quality of your data and improving quality of the metadata will improve the LLM output.
00:49:28 Q: fine-tuning on raw data that hasn’t been validated
Chris Bergh: I'm not sure where how to tell bias or hallucination in the LLM output. That's a that's a whole separate almost research area that people are looking at. And so how do you deal I think the other question is sort of sort of in LLM fine-tuning workflows we often deal with large volumes of raw unclean data. From your perspective how should teams manage fine-tuning when the data has been not fully validated? How significant the impact of poor quality on downstream LLM performance? And that's a really good question. You know, I I I think it comes down to my sort of expected value calculation, right? The the LLM is going to be wrong 10 20% of the time. Your data quality, if it's wrong, is going to the less quality of your data, the less likely your predictions are going to be accurate. And so from a DataOps perspective, the ability to iterate and improve and deploy something to your customer, have them say this is right or this is wrong and then check the data, refine the data.
00:50:37 Q: which dashboard types are most popular with clients?
Chris Bergh: That process overall of the system that you've embedded your internal data in with an LLM can be helpful. And then an attendee has a question, which dashboards are the most popular with your clients, aka which are the top two? And so probably the first one the data quality based dashboards because like if you go into ChatGPT and they ask data quality dashboards they say organize them by dimension. I think unactionable sort of DAMA data quality are probably the most popular. I I don't think it should be to be honest but I think it is the most popular. And I I I find the other one really tends to be where people have focused on a specific customer sort of business or customer focused. Not to say that that's popularity implies usefulness. I I I don't think you know I think data quality is a challenging field and saying looking at what other people do is not often the best case because there's just a lot of poor companies with poor data quality out there.
00:51:45 Q: 20 million records — should I test all of them?
Chris Bergh: Excuse me. And an attendee says, "I have a table with 20 million records. Should I apply data quality to all of them? And how expensive is that process?" Well, so let's say you've got a table of 20 million records. That's not that many these days, honestly. The question is not about the size of the data. It's your ability to get someone to actually change the source data. And do you have someone who you can have time to actually if you focus on one column to improve? Can you get them to improve it? And like what do you have to give them? And so and I think starting it from that perspective. I apologize for my cold and coughing. I'm losing my voice. But that's the perspective. It's not about data size. It's about how you can get someone to make the change to improve data quality and then how you can show that you're awesome that you've actually made that change happen in the organization. I think that's it for today. So appreciate it while I choke to death here from my cold, but I appreciate you taking the time. I again I will share the slides and the recording and in an email here tomorrow. So I appreciate it and Oh, okay. I I think that's it. Have a great day.
Machine-generated transcript, lightly edited: filler words removed, product and speaker names corrected, and audience members anonymised. Speaker attribution and timings are as captured on the call. The transcript covers only part of the recording, so the chapters stop before the end of the video.
Questions from this session
How do you choose which data quality dashboard to build?
Six types of data quality dashboard and when to use each: dimension-focused, critical data element, business goal, data source or provider, data consumer, and ticket-driven workflow. For each one Chris Bergh gives the case for it, the failure mode, and a real example, then demonstrates five of them built in DataKitchen's open source tool and closes on the iterative loop of profile, test, dashboard, fix, rescore.
What are the six types of data quality dashboard?
Dimension-focused, which groups checks by accuracy, timeliness, consistency and the like. Critical data element, which reports only on the columns that matter for compliance or the business. Business goal, which ties quality to a named corporate objective. Data source or provider, which scores the systems and partners sending you data. Data consumer, which scores the specific columns feeding one person's model or report. And ticket-driven workflow, which is not really a dashboard but a count of data quality tickets opened and closed.
Can I take my critical data elements and organise them by data quality dimension?
Yes. Chris's answer was that whatever grouping motivates change is the right grouping — by CDE, by table, by dimension, or by any other piece of metadata about the data. It is one of the reasons the tool lets you pull the same tests and columns into different groupings rather than fixing one hierarchy.
How would you adapt the six dashboard types for LLM fine-tuning data sets, where hallucination, bias and inconsistent labelling matter as much as schema drift?
Chris split the question in two. On the input side the six types apply directly: improve the quality of the data going into the model and the quality of the metadata about it — catalogs, profiling results, test results — and the expected quality of the output goes up. On detecting hallucination or bias in the output he said plainly that he is not the right person to answer and that it is close to a separate research area, so the session does not claim to cover it.
In fine-tuning workflows we deal with large volumes of raw, unvalidated data. How should teams manage that?
As an expected-value calculation. The model will be wrong some of the time whatever you do, and lower-quality input makes it wronger. The DataOps answer is to iterate: deploy something to your customer, have them tell you which answers are right and wrong, check and refine the data behind it, and repeat, rather than trying to validate everything up front.
Which dashboard types are most popular with clients?
Dimension-focused ones, by a distance — partly because if you ask ChatGPT for a data quality dashboard it will organise one by dimension. Chris was clear that popular is not the same as useful, and said he does not think it should be the most popular. The type he sees working is the one focused on a specific customer, business goal or consumer.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.