On-Demand Webinar · 38 min

The Seven Deadly Sins of Data Quality

Shame, denial, avoidance, passivity, laziness, gluttony, and ignorance: seven habits that keep data quality broken, told through real stories from real data teams. Plus the five organisational patterns underneath them, because the sins are usually the system's fault rather than the person's.

Presented by Chris Bergh

What you'll learn 6 points
  • The seven personal sins are shame, denial, avoidance, passivity, laziness, gluttony, and ignorance. Each one comes with a story from a real data team rather than a definition.
  • The five organisational sins are low relational coordination, blame culture, data blinders, project rather than product focus, and lack of process curation.
  • The personal sins are almost always a reflection of the system. Chris cites Deming's split — 97% of problems are the process and 3% are specific — and says a leader's first job is recognising that the disarray is theirs to fix.
  • Avoidance has a measurable price: teams that keep putting data quality off spend up to 40% of their time firefighting, and their stakeholders find the issues first.
  • Gluttony looks like ambition. One insurance team spent seven years building a data warehouse that never got a customer, and a pharma team expected a platform migration to fix data quality by itself.
  • The counter to all twelve is to get something done rather than something perfect: pick one field, add a check, and iterate to influence.

Prefer to read it? The written version is in On-Demand Webinar: Seven Sins Of Data Quality.

Slides

42 slides

Transcript

Chris Bergh

Show chapters and dialogue 34 chapters · 6,326 words
  1. 0:00 Welcome, and why the sins are a metaphor
  2. 5:25 Seven personal sins and five organisational sins
  3. 6:31 Why data quality behaviour is worth naming
  4. 7:30 Sin 1, shame: the 24-year-old who wept in my office
  5. 8:22 What shame costs: burnout and lost safety
  6. 9:25 Sin 2, denial: the call about a blank report
  7. 10:20 The war room, and the question nobody could answer
  8. 11:11 “I don’t want to know”, and the rationalisations that follow
  9. 12:17 Sin 3, avoidance: the data quality debt of a busy team
  10. 13:39 Sin 4, passivity: waiting for the problems to compound
  11. 14:58 Sin 5, laziness: “nobody called me, so the data is fine”
  12. 16:04 Broken windows, and whether laziness is really incentives
  13. 16:59 Sin 6, gluttony: boiling the ocean on a new platform
  14. 18:05 Seven years, no customers, and sin 7: ignorance
  15. 19:06 Why “I can’t test what I don’t understand” is wrong
  16. 20:05 Deming: 97% of problems are the process
  17. 21:11 Process cause versus specific cause
  18. 22:11 Organisational sin 1: low relational coordination
  19. 23:08 Bronze, silver, gold, platinum, and four teams that don’t talk
  20. 24:15 Warring tribes, and what a simple request costs
  21. 25:08 Organisational sin 2: blame culture in a multinational
  22. 26:04 Defensive engineering and the shield of deflection
  23. 27:10 Safety culture, and sin 3: data blinders
  24. 28:10 From controlling tables to orchestrating workflows
  25. 29:21 Organisational sin 4: project focus, not product focus
  26. 30:16 “Are people using it?” “I have no idea.”
  27. 31:13 Organisational sin 5: no process curation
  28. 32:12 Tribal knowledge, and putting the process in one place
  29. 33:12 Takeaway 1: something done beats something perfect
  30. 34:10 Takeaway 2: improve how you lead the team
  31. 35:05 Why code and process beat data and metadata
  32. 36:07 Where to read about each sin, and the survey
  33. 36:59 Q: the data quality tool is priced too high to start slow
  34. 37:50 Q: how do you shift from reactive to proactive data quality?

00:00:00 Welcome, and why the sins are a metaphor

Chris Bergh: Hello everyone. I'm Chris Bergh. We'll start in just just about two minutes. Again. We'll start in about one minute. So, hello everyone. My name is Chris Bergh. I'm CEO of DataKitchen and we've got a fun webinar for you today called the seven deadly sins of data quality. And so, I'm going to go and share, of course, the slides and the recording and the transcript transcription of this. We'll set up a small website and we'll email you, today. And put your questions in the the chat window. I'll answer at the end or perhaps during during the webinar. And it and it's fun. And so, I hope you enjoy my use of Google NotebookLM and and ChatGPT. It's just meant to be a fun webinar where I use se seven sins as a metaphor. Of course, they aren't literal sins, but I think hopefully you can can enjoy some of the stories.

00:05:25 Seven personal sins and five organisational sins

Chris Bergh: And so, I'm going to talk about kind of two groups. So, there are seven sins. And I realized in preparing for this, it's it's awfully personal. Like you take it as my fault. Things like shame and denial and avoidance and passivity and laziness and gluttony and ignorance, they're they're awful personal. And I was thinking about this and and realized that like people live in organizations. And so there's kind of another set of sins that are really have to do with the team that you work in. And so if you're a manager, you may want to think about these, right? Things like low relational coordination, we'll talk about that. Blame culture, sort of data blinders, project not product focus, and lack of process curation. And we we'll talk through those. So, that's really our agenda is to go through. And what I've done here is you know, you've got the seven personal sins and the five organizational sins. And I put a survey if you want to check off which ones that you've seen or which ones that you think are you know or you think you will see.

00:06:31 Why data quality behaviour is worth naming

Chris Bergh: So we'll we'll look at the end and I'll I'll the link is shared for that Google form in both the chat. So let's start. So we're going to talk about the seven sins of data quality. And since I used AI to create these it came out kind of amazingly good but also a little bit verbose. So, I ask you to bear with me on the the wordiness of the slides, but they look cool. So, you know, it's got that going for it. Well, we all know why data is important, right? And we all know if you're here, we all know that data quality is important both in source systems as well as being perhaps patched or improved in your in your data and analytic warehouses. And so there's a lot of problems in in data and so a lot of problems with that people have and behaviors that they have around data and data quality that I think can be improved. And so we're calling these the seven deadly sins.

00:07:30 Sin 1, shame: the 24-year-old who wept in my office

Chris Bergh: And I'm going to go through each one. We've actually blogged about each one over the past year. And so I'll give links to those at the end if you're interested in in learning more. So the the first one is shame. And so it's that feeling that something's wrong and it's my fault. And so I've got a story for you. When I turned 40, I had a data engineer work for me. He was actually turning 24 and we shared the same birthday, October 18th. And I was at a company that did data and analytics for the healthcare industry. I was in charge of kind of the COO, in charge of all the people doing the stuff. And so, I tried to, you know, get to know him in in a one-on-one. I think it was my first or second week, and he just broke down crying. He said, "I can't do anything right. Things are late. I can't go fast enough. Things are broken." And he just wept in my office.

00:08:22 What shame costs: burnout and lost safety

Chris Bergh: And I felt so bad for him that you here he was 24 year old, and he just had this deep shame that he wasn't doing anything right. And so, you know, that has a a lot of things, right? Is that you know he did have his customers you know literally yelling at him and you know the culture in the company that I had joined was very much a blame culture. The CEO was a a former doctor who was had sort of a medical school view of errors like oh you just killed the patient and and people you know felt it was their fault. They sort of it's it's my bag it's it's it's alone and so these things sort of fester in the dark right and it leads to burnout. It leads to in some cases crying and you've even done surveys of data engineer burnout and data engineer unhappiness and it's it's incredible how many want to leave the industry or just in general frustrated. So shame just does take its toll right both in performance as well as in just success of your team.

00:09:25 Sin 2, denial: the call about a blank report

Chris Bergh: So, I don't know if anyone has seen shame or has any questions on that, but I'll I'll keep going. And, and, you know, it's hard like we've all performed, with shame. We've all tried to do well, and when you don't do well, you know, shame fers and it it it makes you loses safety and you become protective. And then you also become less willing to take risks, less willing to try, more wanting definition, and that just slows everything down. So the the second one is is denial. And it's kind of a culture of rationalization where signs of trouble are you know ignored or sort of they hope it go away. So a few years ago I got a call from the head of a data team. It was the 12th or 13th biggest company in the United States. And so, what he told me the story is that he had just got off of this sort of war room of about 12 people, because there was a report that was blank.

00:10:20 The war room, and the question nobody could answer

Chris Bergh: And how he learned about it was the CEO of this 12th Vegas Fortune 500 called him up and said the report was blank. And how he learned about it is another CEO of a Fortune 100 company called him up and yell at him for it because it was their partner and had to do with those companies working together. Successfully. So what happened? Well, he took 12 of his best people and said, "I'm gonna fix it and they worked all night and they found, you know, some one line of code somewhere in this sort of complex system that led to the report being blank." And so I talked to him and I said, "Wow, that's hard, right?" And how are you going to, you know, the simple question, how are you going to avoid this in the future? And he didn't know because he they didn't test data, they didn't check data, they sort of hoped that both the data was right and all their systems worked. And I asked him if this was prevalent.

00:11:11 “I don’t want to know”, and the rationalisations that follow

Chris Bergh: He goes, "I don't want to know. I just don't want to know the problems." And so, because he's afraid that he'll find even more. And so, denial is it's it's a real problem, right? Is that people look and say, "Well, that they take things they take signs and overgeneralize them." Like, "Oh, my pipelines are green." Well, that doesn't mean the data is right. It just means they they ran. Well, or yeah, the data is wrong, but the users have a workaround. Or nobody's complaining or hey, we got the data exactly as we received it. And so, what happens in denial is sort of this institutionalization of sort of failure is that bad data is not my fault and it just sort of goes through. And or they get into the war room cycle where they're always in reaction. Nights and weekends are bad. I've I've told stories of one guy I met at a conference who spent Saturday afternoon while his kid's birthday was party going on fixing a data error and that was pretty standard for him.

00:12:17 Sin 3, avoidance: the data quality debt of a busy team

Chris Bergh: So when you deny that there's problems, you deny that there's a way to fix the problems, you end up with this this toll. So the third one is avoidance. And we you know we do some consulting with sort of management consulting. We worked with another company that did kind of analytics consulting for the media business and and they were super busy, a small team, really smart, kind of working all the time and they were kind of always crushed, right? Because they were looking for new customers, they were kind of selling new things. And the team was just busy and so they ended up with this debt of data quality. Both the data was unknown whether it was good quality, the pipelines were in production, there were no data quality tests and and so they just were kind of avoiding it. Because there is always a new crisis and so you know that sort of ignorance of kind of you know I'm I'm avoid I'm avoiding it and trying to make it work is you know it takes their tolls right because you end up kind of spending up to 40% of your time firefighting and and a lot of your issues end up being found by stakeholders and that just is is cost right and and you feel like I've got this big task list that I've got to get done and I don't have time for this Kov

00:13:39 Sin 4, passivity: waiting for the problems to compound

Chris Bergh: Quadrant 2 stuff. And so the next one is passivity. And so I've seen this quite a bit. So like here's a large organization a pharma company and they have a giant data team in in India or and maybe it's in India and some other countries, I don't remember. And so they're just not into it. They're like, I don't know what to do, so I'm just going to wait for the problems. The sort of being proactive is kind of it's not on a project. I don't really need to do it. And so what that means is they just wait, right? The you know, the the debt of the data, the problems in the data compound. No one really looks at the data. The people who use it end up having to check the data, right? So and then it becomes sort of a joke right on on that this idea of passivity ends up meaning that people wait and wait more and so the fifth sin now laziness so I I've got a I got a story here again back in my data consulting company days 20 years ago and so I just sort of started getting the data testing religion and and you know being new at the company maybe nine months

00:14:58 Sin 5, laziness: “nobody called me, so the data is fine”

Chris Bergh: In I talked to an engineer he had gotten some new feature in for a customer he's very proud and I asked him you know how did you know it worked and he's like well no one's called me no one's yelled at me so it's great the data must be fine no one called me up about it and so like that sort of I think of it as laziness that they they didn't actually take the time to build a repeatable data quality check. And they just sort of were hoping it worked, right? And and that sort of laziness on testing ends up kind of biting you, right? Because what happens is like some of the other ones, you get firefighting or people don't trust the the the process or this you know you're constantly trying to make a quick change to patch something and so you have this sort of you know there's this sort of broken windows theory in policing, right? Where if you if you fix the small things, you start making sure that people don't litter or that small crimes are are prosecuted, it ends up improving the overall crime rate.

00:16:04 Broken windows, and whether laziness is really incentives

Chris Bergh: And that sort of applies to data quality as well, right? Is that you even for your smallest change, you need to make sure that you put in data quality tests that that also run in the development process but also run during production. And so it is sort of a thankless task, right? And and maybe it's not that they're lazy, maybe they just aren't incented properly to do this work. Maybe their boss doesn't care. They only care that they get their check marks done. But these data quality becomes this thing that no one does. And and you know, a lot of these sins have this characteristic. They they they seem when you look at it and we're going to talk more about the the the culture that organizations set up and and here in the personal part we're being very personal and and and laying it on people but it is in a sense laziness and we'll talk about maybe why that laziness happens in an organizational context in in a little bit.

00:16:59 Sin 6, gluttony: boiling the ocean on a new platform

Chris Bergh: Now the sixth one is gluttony and and this one is one of my personal favorites the boil the ocean. I got to do everything. And so, we just talked to, a big pharma company, and guess what? They're going to go from Redshift to Databricks. And guess what? Databricks is going to do everything. Once the data is in Databricks, everything's going to be great. And I go, I've seen this so many times. You pick the platform. Once we go to Hadoop, everything will be great. Once we go to Teradata once we go to Oracle it's this new platform thing where I'm going to boil the ocean u another case in data quality is you do have a data quality improvement project but it's so broad there's you know 62 different things you have to do right and and this idea that we're going to do a lot we're going to do a lot at once instead of doing things quicker and is a telltale sign that you're not going to have success right because the grand initi the new data warehouse, the you know list of a thousand data fields that you're going to fix ends up in kind of analysis paralysis.

00:18:05 Seven years, no customers, and sin 7: ignorance

Chris Bergh: It ends up on being hard and unmeasurable and they sort of collapse under their own weight and then when business users have to wait a year or a year and a half they're like well what's going on? And like I talked to an insurance company 3 years ago. Their team had spent seven years building a data warehouse that had no customers on. And so I'm sure it's beautiful, but like seven years is a long time to deliver any value. And so I think that team doesn't exist anymore. And I don't think the actual warehouse did end up getting in production. So this idea of gluttony the the big piece and then the last sin here is really ignorance and I've seen this in larger companies u multicontinental IT organizations and people are very separate from the business right and they don't understand the data or don't even attempt to understand the data and so data becomes kind of held at a distance by by people who do it they don't someone's got to tell them the meaning and I I'm I I don't agree with that.

00:19:06 Why “I can’t test what I don’t understand” is wrong

Chris Bergh: I think data is a representation of your business and is fairly straightforward to understand products and customers and manufacturing lines and sort of understand the theory of the business and understand the context that it's in. And so a lot of cases they don't test data because they're like I don't understand the data and don't even want to conceptualize why they don't understand the data. And so this sort of you know the fact that you sort of can't test what you don't understand is a problem right because first of all there is a lot you can test even without understanding the data but also just trying to collaborate with people to do can help you learn more and you know we've we have a consulting organization that's had a lot of success by assigning people to work in pro data product domains and having them learn the data really well and then be able to add a whole bunch of work very quickly both from small instructions, you know, a few sentences as well as being able to to build very custom data quality tests.

00:20:05 Deming: 97% of problems are the process

Chris Bergh: And so having your data teams understand the data is a really good thing and and not teaching not taking every data engineer as a completely fungible asset. And so this this goes from and so you know it's harder in organizations with thousands of data sets and that's where Stuarts can come in to help but trying to get your data team to learn a little bit more. And so that's really the seven sins of data quality. And so, you know, from my perspective, you know, I I've I've grown I I see these things in people that, you know, that they're in denial or that they're lazy or they're passive. But as a leader, I've of teams, I've grown to not look for the specific cause. I'm more of a person who believes in Deming and that it's the system that the person works in. And so if someone is lazy or someone is ignorant, it's because of how I manage them as a team. And it's just like in manufacturing, 97% of the problems have to do with the process and only 3% are specific, i. E. The people.

00:21:11 Process cause versus specific cause

Chris Bergh: And so yes, sometimes you do have passive people who aren't doing anything and and but almost always it's the system that people work in and so it's organizational that shows up at these problems. And so that's my that served me well in my career. To look for the what Deming calls the process cause as opposed to the specific cause. And so a lot of these things can reflect in organizational sins and data teams. And so this perspective is kind of based on our books on DataOps and our thoughts on DataOps and and it goes a little bit it goes beyond data quality to the entire work stream that data and analytic teams do. So we're going to talk about these these pieces now. And so we're going to go through five of them. And so that's so the first one is like yeah I I guess this kind of to reiterate what I said it's really about the system that people work in and not their particulars. And so finding the root cause is is is almost never that person is an idiot and we have to fire them.

00:22:11 Organisational sin 1: low relational coordination

Chris Bergh: Right? It's it's almost always the manager isn't managing their team in the right way and the incentives aren't aligned and they're not thinking about the problem. And this just comes from my experience both as a leader of data teams myself and as consulting and and working with our customers and trying to it's a hard nut to to swallow for people but like improving yourself as a leader of your team starts with sort of recognizing that your kingdom is is in disarray and and it's your job to fix it. So so the first thing is sort of low relational coordination and and this is sort of a fancy word for that people don't work well together. Now in organizations they could have great relations. They could like each other go out to drink but they also could like each other and not work well together. So sometimes we mix up the fact that people get along the fact that they actually work well together. And so this shows up in a lot of ways, right?

00:23:08 Bronze, silver, gold, platinum, and four teams that don’t talk

Chris Bergh: You see this where the work is very divided. Here's one team and then I hand it off to the next team. I don't know what they do. They pick it up and hand it off to the next team. There's not a lot of co and getting anything done takes a while because of this barriers of relational coordination where people are like in this diagram sort of sending tickets back and forth and not communicating. People don't understand each other's perspectives. The data engineers don't know the data. The business users don't know the challenges of the data team and and it's this ends up sometimes there's one wall, sometimes there's several walls that that happen. And so you know an example of that is we worked with a a company they had a medallion architecture which has layers in it and sort of bronze you know bronze, silver, gold. They had another layer, a platinum layer and all the layers basically had different teams on and those teams didn't talk to each other and so you had kind of four layers of data movement and then of course you had the utilization of data the people who were building charts or data science and then of course you had the customer.

00:24:15 Warring tribes, and what a simple request costs

Chris Bergh: So no one really knew what was going on. And so yeah, you can organize teams in that sort of balkanized way, but they had extremely low relational coordination. They were kind of each level was sort of a warring tribe and sort of fingerpointing at at the other one. Everyone worked separately. And then what does that mean? It means you see that when a customer makes a request for a simple new data set and it's got to crawl through each team, right? And and so even if they like each other, it can drink with each other and have a great time together, just actually getting these teams to work and think about how to do it. And then what happens is you you hire people and they live in that they're like, "This is dumb and I, you know, I want to understand what's on the other side of the wall." And they can't. And so this isn't really done with like I just want to have a better workflow tool or a better Slack.

00:25:08 Organisational sin 2: blame culture in a multinational

Chris Bergh: It has to do with actually rethinking the process that people work in. And so it has a lot to do with seeing the data and analytic production process as a manufacturing line where data comes in and insight comes out. And so I saw Attendee, you have a question. I'll I'll I'll I'll answer that at the end. So the next sin is sort of blame culture. And so I'll talk about this and get give an example of kind of a large Japanese multinational. And they have this they've got kind of this layer of like the business finds a problems. They call up the analysts and say you're an idiot. The analysts call up the data team and they say you're an idiot. The data team calls it who runs the system and then sometimes loads at the data and say they're an idiot. Then it blames the data provider. So it's just this fingerpointing, right? It's your fault. It's your fault.

00:26:04 Defensive engineering and the shield of deflection

Chris Bergh: It's your fault. And so no one wants to be left holding the bag, right? Because it looks bad and everyone wants to avoid that blame and in the organization and those that's a terrible honestly a terrible place to work. Because what you're trying to do instead of make your customer successful you're trying to be defensive. And that sort of fear makes you hide problems. Hit fears, makes you do defensive engineering, makes you do a lot of like bureaucratic things like emails and documents. And then you know what happens is the energy goes into kind of like this the science says the shield of deflection. You're spending all your time kind of deflecting blame as opposed to creating new things. And so blame is a real problem in organizations and some organizations just live on blame and it starts almost with the top. They're just you know you didn't make revenue I'm going to yell and so that that's a tough one to change in organizations.

00:27:10 Safety culture, and sin 3: data blinders

Chris Bergh: But I know that this has happened in auto manufacturing. It's happened in software. The sort of safety culture, the fact that you can bring up problems, the fact that you see problems as an opportunity to improve as as opposed to an opportunity to blame can really help. And this was back in the 20 years ago when I first took the data team trying to get people onto that m mantra away from blame to opportunities for improvement and and let's not just let's the problem isn't the problem. Let's fix the problem and and and learn from it so it doesn't happen again. And I think that attitude can go a great way to to help defer blame culture in your team. And then the the third one here is I think of as sort of data blinders or data blindness. And so some people may may this I've seen this at many companies, right? Teams focus on data, right? They're sort of centric on data and metadata. And I get they're sort of blind of the processes that are acting upon data.

00:28:10 From controlling tables to orchestrating workflows

Chris Bergh: And so their work is seen as like proportional to the number of tables. I've got more data. I have more work. And so they're sort of blind to the sort of processes acting on data, the processes and the code that act across those tables and those deliverables, that manufacturing line, that pipeline that's taking the data. And so it's it's a focus kind of from controlling tables to orchestrating workflows. And it's seeing those workflows as kind of cattles that that are cattle, not sort of pets that need to be done. And so the factory is really it doesn't say that data is not important, but it says when you have problems with data, look to the factory first and see if you can put checks or or patches or data quality checks. And so that sort of focus away from data and onto the processes, the systems, the pipelines that are acting upon data has a whole lot of benefits in terms of making your team more efficient. And so you also can see here in the AI generated image the weirdness of the water flowing backwards into town.

00:29:21 Organisational sin 4: project focus, not product focus

Chris Bergh: So and so I saw Attendee, you asked a question. I'll that's great. I'll I'll I'll do that at in a little bit. So, this is one of my favorites, one that I believe in quite a bit and and focusing on the manufacturing line, the process and having it it really means that doesn't mean that data is not important or metadata is not important. It just means the code and the processes acting on on the data are just as important too. And so, the fourth thing is really sort of project and not product focus. And so I've heard this from a company, right? They so I talked to the data team. They're like really proud. They said, "Our project was a success. We followed our software development life cycle to the letter and you know it was on time. We got all the project artifacts done and I was like, "Oh, great. I was happy for them." Then I asked him the question,

00:30:16 “Are people using it?” “I have no idea.”

Chris Bergh: "are people using it?" And he goes, "I have no idea." you know, and then the next thing he said, we got an award. And he was very proud, right, that he was focused on the project, the rituals and artifacts that go into the project and the timeline was the goal. And I guess I think the goal of being product focused and there's a lot of terms of data products out there that I think frankly are are kind of BS, but it really means focusing on value delivery to your customer and kind of is like it's not that it's on time or on budget or I followed my project. It's like is it valuable and and value is defined by does your customer use it? And so what happens is teams are often assembled for projects. They follow it. They deliver it and then they go away. And so the product focus is really about building improving products that happen time and time and again and you're constantly never done and improving.

00:31:13 Organisational sin 5: no process curation

Chris Bergh: And so I think that's a a sin that I I've seen and like it's not great just because you've followed your process. And then our last one here is a similar one. And so here's another large organization. They have data ingest teams. They have data analytic engineers. They have people who do charts and graphs. They have data scientists. And then they have specialized consultants who come in to do data science or other work. They sort of five groups, right? And and so at the end of the work of those five groups, maybe all of them, they have some deliverable to their customer. Maybe it's a report. And so what happens is people are doing little bits of work on that. They're writing some SQL, maybe there's some Python. You know, maybe it's in a database, maybe it's on their laptop, maybe someone put it in Git. It's that there's no there there on all the process artifacts.

00:32:12 Tribal knowledge, and putting the process in one place

Chris Bergh: And so there's no 360-degree view of their process. And so what happens is the next time they want to fix it, they've got to go finding where the work is. And so it gets this tribal knowledge. It's maybe in in you know, maybe it's in Arthy's laptop or maybe it's over here on on on Tom's Python code. Nobody knows how it works. And so changing it, improving it becomes very difficult because it's all about emails. And so the idea here is that both the process needs to be put in one place i. E. The code. So and then second is that a lot of times when you see different metrics like I have a common metric for calculated sales you want to be able to find and refactor that you and that has to be with well what code was used to do it what's the logic and so having iterating and improving upon your process and trying to curate it centralize it refactor it is is is is is very important.

00:33:12 Takeaway 1: something done beats something perfect

Chris Bergh: And so those are the kind of challenges that I see with process curation. And so I think that you know like I said that that sort of leads to reinventing the wheel. I've got to every time I do something I have to find it and then start over again. What that means is more work, more slowness, less quality. And so we've gone through kind of the seven deadly sins. Passivity, laziness, gluttony. We've gone through the five organizational sins. And so I want to kind of sum up with just sort of two sort of big thoughts, right? The sort of the main idea on these personal sins is kind of get something done. Like something done is better than something perfect, right? And in data quality, it's really about iterating to influence. Sort of start small. Develop some data quality checks, try to influence or improve the data. We've talked quite about the this in our ideas of the DataOps way to data quality.

00:34:10 Takeaway 2: improve how you lead the team

Chris Bergh: We have an open source tool that can help you do that very quickly. And so the main idea that can help data quality and fix the data the seven data sins is just get something done, right? And and that will counter all these seven sins and sort of take it upon yourself. Maybe it's just one field or two fields that you want to improve. And that's and maybe you don't have to improve them perfectly. Maybe just a clearing up a zip code field or clearing up an e email address would help the organization. And then the second idea is and and this I mentioned this before is that individuals are great but actually to make improvements you have to improve how you lead as a team. And so often the problems in teams have to do with how they work together. And so in data and analytic systems that there really two ways to look at the work. There's sort of data and metadata and then there's code and process.

00:35:05 Why code and process beat data and metadata

Chris Bergh: And I'm of the view that code and process is actually more important than data and and metadata because you can use the code and process to improve the data and the metadata. And so getting your people to start thinking about the end to end process start thinking about how they can deploy that better. The ideas that we've written about in the DataOps manifesto and cookbook and our our our training are actually really important. As well as you know thinking about data quality and testing and so so that's where I think it's it's really is like you know this has got this overly religious theme and I hope I didn't offend anybody but it's not meant to be that they're sins these are typical problems that organizations have and dysfunctions that I'm sure all of you have seen and so I think one of the things that we have done is built this open-source tool that creates data quality tests for you and develops data quality scorecards for you and does data observability for you and it's free.

00:36:07 Where to read about each sin, and the survey

Chris Bergh: It can use as many tables as you want. And so love you to to download it and start using it today because that'll help you work on some of these problems by yourself. And lastly we do have a whole bunch of writing about the ideas of DataOps and how to manage a team. We've written two books. We've got a certification that would be great. And so lastly, you know, we have kind of articles about every sin. And so but there's a blog about the organizational sins and a link to everyone. And then there's sort of a blog on the seven sins and a link to the articles that talk about them. So, if you're interested, want to learn more, check out these links. And then lastly, if you, go to the survey, I'd love to have you click and see which one of these are relevant to you. All it takes two seconds. All you got to do is have some check check marks and hits hit a button.

00:36:59 Q: the data quality tool is priced too high to start slow

Chris Bergh: I'd love to see you do that. And so, the link again is in the notes. If you don't want to go to this insanely crazy URL. So, let me, let me go back to the questions and kind of walk through those. And so, so Attendee, question on boiling the ocean. We understand to not, but we understand to not implement it slowly, but the DQ tool is priced too high to go slow. Value realization. They want a proven main player tool that can be implemented below 10K and ramps up to need it. Our VPs would be all years. Yeah, I honestly, Attendee, you're talking to the right company. We you know our our price point is $100 a user a month. So a small team with three users and three databases that's about 10k a year. So Attendee have I got the data quality tool for you. You can start free unlimited tables prove the value. So TestGen is something that you should look at.

00:37:50 Q: how do you shift from reactive to proactive data quality?

Chris Bergh: And so Attendee how can you shift from being reactive on data quality and proactive in use profiling? Well, I think partly one of the reasons that we do this is is that your perspective and your mindset matters and so you know the the structure of the organization you work in and and sort of your own behaviors are something to reflect on and improve and you're not stuck. And so I think one of the reasons that we've, you know, spent millions of dollars and given away an open source tool is that you can overcome these things yourself by just downloading something and picking a data field or tool and trying. You can profile them. You can press a button and get data quality scores. You can press a button and get dashboards. You can press a button and monitor those continually. In our release coming out here in three weeks. And so all these things are are enable you to be very proactive in data quality by yourself. And so and Attendee, thank you. I I was just a little worried at the beginning whether sort of the the medieval theme would would bother people. So give me your comments on on Google forms or in the goo on using NotebookLM and AI to generate these. I will share the deck. I will share the recording with you probably today or tomorrow. Thank you for attending and thank you for your questions and have a great rest of your day.

Machine-generated transcript, lightly edited: filler words removed, product and speaker names corrected, and audience members anonymised. Chapter times are scaled from the meeting clock onto the recording, which is shorter than the meeting. Speaker attribution is as captured on the call.

Questions from this session

What behaviours keep data quality broken?

Twelve behaviours that keep data quality broken. Chris Bergh walks through seven personal sins — shame, denial, avoidance, passivity, laziness, gluttony, and ignorance — each illustrated with a story from a real data team, then five organisational sins: low relational coordination, blame culture, data blinders, project rather than product focus, and lack of process curation. The session closes on why the personal sins are usually a symptom of the system.

What are the seven deadly sins of data quality?

Shame, the feeling that something is wrong and it is your fault. Denial, a culture of rationalisation where “the pipelines ran green” stands in for “the data is right”. Avoidance, letting data quality debt build because there is always a new crisis. Passivity, waiting for problems instead of looking for them. Laziness, not taking the time to build a repeatable check. Gluttony, boiling the ocean with a grand initiative. Ignorance, staying far enough from the business that you cannot test what you do not understand.

What are the five organisational sins?

Low relational coordination, where teams hand work over a wall and nobody sees the whole. Blame culture, where the energy goes into deflecting fault rather than making customers successful. Data blinders, focusing on tables and metadata while ignoring the processes acting on them. Project rather than product focus, where finishing on time counts as success even if nobody uses the result. And lack of process curation, where the SQL, the Python, and the logic live in five places and only tribal knowledge can find them.

Are these sins really the individual's fault?

Chris's answer is no, and he is explicit about it. He follows Deming: 97% of problems come from the process and 3% are specific to a person, so if someone on your team looks lazy or passive it is usually about how they are managed, what they are incentivised to do, and whether anyone is thinking about the system. Improving as a leader starts with accepting that the disarray is yours to fix.

We know not to boil the ocean, but data quality tools are priced too high to start slow. What can be implemented under $10K?

TestGen. The open-source version is free with unlimited tables, so you can prove value before spending anything, and the enterprise price is $100 per user per month — a small team with three users and three databases lands around $10,000 a year. That is the point of the pricing: it lets you start narrow instead of forcing a big-bang rollout to justify the licence.

How can you shift from being reactive on data quality to proactive?

Chris put the first part on mindset: the structure you work in and your own habits are things you can reflect on and change, and you are not stuck. The second part is mechanical. Download the open-source tool, pick one data field, profile it, press a button for a data quality score, press another for a dashboard, and set it to monitor continually. Doing that on a narrow scope is what turns reaction into prevention.

Where to go next