On-Demand Webinar · 1 hr 2 min
Redefining Data Governance with DataGovOps
Laura Madsen, author of Disrupting Data Governance, and Chris Bergh on why command-and-control data governance fails the way modern data teams work, and what DataGovOps puts in its place. Recorded July 2020; updated August 2026.
What you'll learn 6 points
- Data governance historically claimed data usage and quality, compliance and risk, and security and protection. This session argues governance actually lives in a narrower place, and proposes a weighting: 40 percent increasing usage of data, 25 percent data quality, 25 percent data management including lineage, and 10 percent data protection.
- DataGovOps is data governance done with DataOps practice: business glossary and data catalog as code, process lineage, automated data testing, self-service sandboxes with test data management, and agility in defined roles and responsibilities.
- Deploy data catalog and glossary changes simultaneously with the code, model, and report changes that caused them, rather than updating the catalog after the fact.
- Process lineage means keeping every version of the end-to-end process in a git repository and storing the as-run version — source code, test results, timing data — so the exact path the data took to value can be reconstructed.
- Deming's finding that 94 percent of causes are common cause is the argument for governing the process rather than looking for a person to blame.
- Automated production testing for governance spans five test types: traditional data quality, statistical process control, location balance, historic balance, and business-based tests.
Prefer to read it? The written version is in Continuous Governance with DataGovOps.
Slides
Transcript
Show chapters and dialogue 10,583 words
00:00:00
Good afternoon, everyone. Thanks for joining us today. My name is Beth Beverly. I'm the VP of marketing at DataKitchen, and I will be the host for the webinar today. Our topic is Redefining Data Governance with DataGovOps. We're very excited to have a very special guest, Laura Matson, join us to share her experiences and insight on data governance.
Laura is the co-founder of the Minneapolis-based consulting firm, Via Gurus. She also has over 20 years of healthcare data governance and analytics experience. She's written three business books, including "Disrupting Data Governance," "Data-Driven Healthcare," and "Healthcare Business Intelligence," and she's a frequent conference speaker in the US and internationally. She'll be joined today by Chris Bergh, who's the DataKitchen's founder and CEO.
Chris is a leader in the DataOps movement. He has more than 25 years of research, software engineering, data analytics, and executive management experience. At various points in his career, he's been a COO, CTO, VP, and director of engineering. Through these experiences, Chris realized that there had to be a better way to quickly deliver innovative analytics, which led to the founding of DataKitchen. He's the co-author of "The DataOps Cookbook" and "The DataOps Manifesto." So Chris will kick off the webinar today and provide a high-level background on DataOps. Then he'll hand it over to Laura, who will talk about disrupting data governance, and then we'll have an open discussion during the last 20 minutes or so about DataGovOps.
We hope to have lots of audience participation, so please enter your questions in the question and answer box on the webinar control panel during the course of the webinar, and we'll make sure that Chris and Laura are able to address all of these at the end. One last thing I wanted to mention before we jump into it, as we mentioned on our registration page, we'll be giving away a copy of Laura's book, "Disrupting Data Governance," to 10 lucky webinar attendees.
So we'll choose those winners after the webinar and notify you via email if you're a winner, and at that time, we'll coordinate with you the best way to get you a copy of the book. So with that, I'm going to hand it over to Chris to take it away.
Chris, your sound is off. Okay. Hi, everybody. Thanks for attending today. So my name's Chris Bergh, and I'll be setting it up here and then giving it over to Laura, who'll be doing most of the talking. So what are topics today? So first, I'm going to do a short introduction to sort of DataOps in a big picture, and then Laura's going to talk about sort of redefining data governance, and then we're going to actually do the intersection.
We're going to sort of smash DataOps and data governance together and talk about what DataGovOps is and some use cases. And at the end, we're open to having a whole bunch of questions and discussions. And so Laura is a Midwesterner like ourselves, so we may have some Wisconsin-Minnesota humor in there, so watch out. But I'll start off with what is DataOps? And so, the concept of DataOps is a little bit different in the sense that we're trying to pull you away from data, actually, and focus less on what you do with data and the data itself, and more on how you do things with the data. And so that comes from a couple of ideas.
The first is if you think about building things like a car or an assembly line, the factory is really important. In fact, the machine that makes the machine is actually in some ways much more important than the result. And so thinking of your analytics like a factory and an assembly line, I think, is a lot of good ideas and maps my experience in managing data and analytic teams when they just had a lot of challenges with poor data quality, errors, complaining customers.
And then the second part is in managing people who did data science or data engineering, when I first started off, I often thought when things went wrong, I should blame somebody, and that it was their fault, that they were being dumb. And in reading this guy, Deming, I realized that what he said about industrial plants actually fit very well to data science and analytics teams, that it's very rarely a person's fault. It's not a specific cause.
94% of problems are a common cause. That means you as a leader haven't built a system to make people successful. And so in a lot of ways, that's what DataOps is. What kind of system makes people successful? What kind of an assembly line? How do you actually put that together? And a lot of times, we spend time talking about the model and the algorithm, the pipeline, and the visualization, even the data itself, but we spend a lot less time talking about the process of development, the process of deploying those or monitoring those or iterating or collaborating and measuring.
And so these are the perspectives of DataOps and, in fact, the process and the people, and in fact, the operations in some ways, I think, are more important than
00:05:00
the tools and the technology and here's what's frightening the data itself. And so it has kind of this contrarian perspective. Not to say that they're not important. Everyone loves a good algorithm or good metadata, but the process at which those things are created and managed and the operations are really important. And so there's a mindset shift that we talk about.
And it's really taking away from it, not to say it's not great to have well-defined data or a good transformation, but the cycle time at which you can deploy those and get them into production, if you can get an idea from an individual contributor, data scientist, or data engineer's head into production quickly and successfully, and to do that with very low errors.
And what we mean by errors, it could be poor data quality, it could be just that the data quality is perfect, but something happened on its journey to value. And then a lot of data and analytics has challenges with how people work together. Data scientists versus data engineers and data governance has a lot of challenges working with different people and trying to get well-governed data.
And then also just how do you measure your process of how the people work and how it goes. And so that's the short introduction to DataOps, and we've of course written a book and I've talked a lot about what it is. But for those, now I'm going to switch to Laura and she's going to talk a lot about what is in her book, "Redefining Data Governance."
Well, hello everyone.
Assuming everyone can hear me okay, thank you for that, Chris. I'm very excited to be here. So I wrote the book, "Disrupting Data Governance" after a couple of things happened. So I have been in the analytics space for a really long time, a couple of decades, and a year and a half ago or so, I quit my corporate job and took a little bit of a step back. And as I was sort of preparing to decide what I was going to be when I grew up, one of the things I did is I took a state of the state in terms of data and analytics and programmatically what we were doing and maybe where we were
not doing as well as we could. And one of the things that kept popping up for me was data governance. And I have run a lot of data governance functions. I wrote a lot about data governance in my first book because I have always believed that in order to do data analytics very well, you need to have some semblance of governance.
So it was a little disheartening to me that after all this time in the space, I recognized that data governance was the Achilles heel of most analytic functions. It was an area that we just didn't do particularly well. It actually tended to become a little bit of the other bucket for our data programs. So one of the things I did, which you normally don't do when you do this work, right, I googled data governance.
And I've asked a number of live audiences this question, but it's like, has anybody ever googled data governance? And I was shocked to find out that the standard definition, the one that we broadly all agree is how we define data governance, it's like a paragraph long, and it doesn't really get to what is the actual function of data governance. Why do we all feel like it's important, right?
And it was just really
disheartening to find out that our function is to provide and support analytics for our organizations to get to the metrics. But when we define data governance, it's impossible to create a metrics associated with that because it's a paragraph long, and it literally sort of included the kitchen sink. So here's what, from a visual perspective, I think a lot of data governance functions think about.
It's security and protection, it's compliance and risk, it's data usage and quality. It's kind of all of those things historically all mixed up. That's the frame, right? It became training a lot of the times, right? If it didn't fit somewhere nicely in another data management function, it ended up being thrown into data governance, and it really did become our other bucket from a data program perspective. But as I was doing this research and interviewing people like Chris, what became very obvious to me is data governance actually lives where all of these three things come together.
It's not all of it all the time. It's not one of these things all the time, but it is a conglomeration of each of these things some of the time. And these bubbles will move around. Sometimes you have to focus more on security and protection and less on data usage and quality. But the reality is it's where all those things come together is what organizations should be focusing on from a data governance perspective.
00:10:00
So we can flip slides here. So here's what I talk about in the book about the proposed scope of data governance. One of the key things about this as we think about transitioning to a more agile model of data governance is this idea that when we talk about data governance as being usage of the data, which quite frankly, if we're not focusing on people using the data, what the heck are we doing?
And I can't tell you how many organizations I've worked with that it almost seems like the usage of the data is so much an afterthought from a governance perspective because we all have always been historically worried about protecting that data, and we have to be worried about protecting the data. But it's a little bit like the fox guarding the henhouse.
If you're so worried about protecting the data, it becomes very difficult to support usage of that data. And so we have to really think about flipping that on its head a little bit and saying that increasing the usage of data for our organizations is the only reason that most data governance functions should exist.
That doesn't mean that we have to give up all of the neat things that we have to be focusing on in terms of compliance and risk and protection. It just means that it has to be coming from a different function. And we'll talk about that in a minute. Obviously, data quality is a big aspect of data governance.
And one of the things I talk a lot about in the book is you don't have data quality without data governance, and you don't have good data governance without data quality. And I got a fair amount of pushback on this one when I was talking to people. But the reality is that good data quality is contextual, and in order to have that context, you need to understand how people are using the data, where it comes from, all the things that are sort of the data governance bread and butter. And so data quality is actually the best indication of how good or not good your data governance is.
It's almost like those sets of metrics that drive the data governance function is really whether or not you have "good data quality". Now, the challenge, of course, with good data quality is it's all relative. But that's a topic for a different webinar. Then there's data management, which is lineage. And I get really excited when I think about data lineage because for a very long time as we were running our data governance functions, data lineage tools didn't even exist. Data catalog tools didn't exist, and you had to kind of do it the hard way, giant Excel sheets. And so in order to understand where all of those data pieces come from, your sources, and where they land, the targets, and everything that happens to them in between is really what data governance can operate in terms of those data gap catalogs and live and breathe in that space.
And then we can really talk about how do we define it? What transitions in that data happen over time? How do we ensure data quality as we make those steps? So data lineage becomes very important. And then data protection, because we all recognize that we have a role in data protection. The distinction here that I think we all have to make in redefining data governance is for a very long time, data governance was the team that would end up being almost like the default data protection people.
They would do a lot of the asset management because they had the environment for themselves. They ran the environment where all the data existed. But what's happened over the last decade is that compliance and risk and privacy and information security, or InfoSec, became a career in and of itself. And those of us that are data professionals, data analysts, even data architects, are not professionals in information security.
But for many organizations, we were still the default people that were supposed to be doing that in the environment in whatever you have, your data lake, your data warehouse. And so what really has to happen there is we have got to create a happy alliance with our InfoSec teams, and they are accountable to make sure that we understand what those requirements are, but we still have a responsibility to ensure that we're appropriately protecting our data.
And so in that context, that's what the data protection aspect of our proposed scope is. One of the other key attributes to this as we think about how we define data governance, and these will be in the four attributes of that definition of data governance, is that these things should change over time. So in that bubble graph earlier, we talked about the idea that some of those bubbles will move and become more or less important as we make our transition through data governance. And in the timeframe that is appropriate for your organization, once a year, twice a year,
00:15:00
you should look at these four and allocate the percentage of importance for any given entity here. So for example, I worked with a company last year, and they were really struggling with data quality. So they made data quality and usage of the data both the same amount of importance. And how that sort of works from a functional perspective is in your backlog, anytime you pull things forward into a sprint, you have to make sure that you're allocating the work that you're doing that aligns well with each of these four entities here. So if you're focusing a lot on increasing usage of data and data quality, then that should be reflected in every one of your sprints.
If you're focusing more on data lineage and protection, then that also should be reflected. But most importantly, as you're setting the percentage of importance for these at a more higher level in the organization, like an executive level, that is revisited with an appropriate amount of frequency, and then you're always shifting what's important to the organization. And it allows you to be agile in that sense.
You're no longer just sitting there trying to figure out how to define something, talking to people forever, and figuring out where the data comes from, and then 18 months later, maybe you finally have a definition, but your executives have moved on probably 17 months before then. So that's the idea here of you always modifying your scope and always making those shifts and aligning that to your backlog and your sprints.
We can go to the next slide.
So the one other thing, and there's lots of things that we can cover, but we only have an hour and the main thing I think we really want to get to is this concept of DataGovOps. But one of the things that I think is incredibly important for all of us to recognize is that data governance is really a change management function.
I am a data analyst by training. It's how I started my career. It's one of my favorite things to do when I'm bored. I even took my wedding and turned it into an Excel sheet. If you let me, I would sit quietly at my desk with earphones in and bang away at data all day long. But the reality is that data governance is so much more about change management than it is about analyzing data.
It is about making sure that everybody understands, and this is the awareness piece of it, right? Understanding what it means to enter data into a system and what it means to get data out of the back-end systems. So an example is most people, the vast majority of the people that you will interact with that aren't analysts, understand data in the context of how they enter it into a system, not how they get it out of the back of the system.
So if they get a report out of the back of the system that doesn't align with how they think about it from the front end of the system, that dissonance is going to be a challenge for those folks. So you have to create an awareness and an appreciation for what happens to data when you enter it into the system and the impact that it has when you get it out of the end of that system, whether that be a dashboard or a report or something like that.
Most people don't appreciate that lineage, that storyline that happens with data. And you have to help them understand that because their role in entering that data in the front end drives so much of our data quality and data quality challenges. So that's an awareness. And first and foremost, when you're creating awareness, we have to provide facts and not fear. For a long time when we talked about data governance and getting support for data governance, we used that heavy hand of compliance and these threats of giant fines coming down if you didn't follow these governance types rules. And those are all very true, but if we can create this happy alliance with our compliance and privacy and risk folks, and they give us those requirements, and we execute those requirements, then what we really have to figure out is how do we make sure that we tie the value of the data to a effort that the business people are trying to execute. And providing the facts associated with that and not fear about, "If you don't do this right, we're going to have fines" because that's a very short-sighted way of getting motivation.
And then you've got to kind of acknowledge the discomfort, right? But create some desire associated with making sure that people understand what it means to use data.
And that leads directly into providing some training. So it goes from providing facts, not fear, to acknowledging some of that discomfort immediately into providing that training for people so that they understand what they're doing and why they're doing it. And I don't mean training in like a software piece, right? I'm talking about how you use data and what's important there, and helping them understand what they get from the front end and
00:20:00
what they get from the back end, and all the things that happen in between. And then giving them some action, like how can they help you, and then reinforcing that and doing that over and over and over again. And I can tell you from personal experience that this is not a linear path. You will stop at any one of these multiple times, and you will do that reinforcement over and over and over and over and over again.
But one of the things that I know for sure is if you don't have a communication and marketing plan associated with your data governance function, regardless to how you approach it, whether you use an agile method, this DataGovOps method, or some other way, you're not going to be successful. You have to communicate and you have to communicate more.
And when you get tired of communicating, you have to communicate more, because this is where the work really resides. And I think that is the end of my slides, Chris. So when we talk about the definition of data governance, how does that align to your world? Well, I think that's a great introduction to the next section.
So, and yeah, let me talk a little bit about this idea of DataGovOps. So before that, I want to talk a little bit about why are we hearing all these ops suffixes out there in the world. There's DevOps and GitOps, and we're using DataGovOps, and some people may have heard of ModelOps or MLOps, and it's like there's ops sort of jumping all over the place and so what's that about?
And so the cynical engineer in me just thinks that's all sort of BS marketing, but there is something behind it. And so we wrote a blog post to try and disentangle that. So I think at the very highest level of this slide, there's a business management concept that goes to a lean or learning organization or Deming or that... And there's a bunch of books that started in manufacturing and sort of started to make their way into software and things like Agile and Kanban and Scrum and Six Sigma and total quality management.
And a lot of the idea is to find that balance between sort of top-down control and chaos, and find the right balance between the two. Because top-down control has lots of order and control, but it also can be slow, and chaos can be fast, but it also has problems. And we're trying to get the balance between those two and agility.
And if you look at that organizational method, whether it's Agile or Kanban or Scrum, and you apply it to teams that build software like websites or transactional systems, generally that is called DevOps. It's the technical process, the environment, the factory that does it. And there's all sorts of variations, DevSecOps, GitOps, AIOps, CloudOps. And that's starting to happen to the application of these principles to data science and engineering analytics teams, and that's called DataOps.
And there's some subsidiaries of that, how you apply it to data governance or ETL or data science, and that's kind of what these op suffixes mean. In general, how do you run an organization with a more iterative way? And then how do you see the system? How do you build a system to enable those iterations to be successful where you are able to balance between control and chaos? And so if we look in the data and analytics realm, and you think of on the left-hand side here, think of broadly speaking, and there's different words to use for each one of these. There's data governance, data engineering, data science, data visualization, and maybe these are, you could argue that there's more than four groups of everything that you do in data analytics, but let's just suppose those are the four groups. And each one of them has an ops side to it.
What we're going to talk about today is DataGovOps, and some people are starting to talk about DataSecOps. In data engineering there's a version, in data science, ModelOps and MLOps and BI, and there's an operational side to it. And so that's kind of what we're going to talk about with regard to data governance. And I'm going to go through a couple of examples.
And a theme here in this ops side is a couple high-order thinking and that the first high-order thinking is if you're doing something manually, like you're having a meeting or you're doing something more than once, you should try to make it into code and script it and automate it, and making things repeatable and testable is an important concept.
And so that is one of the ideas about how the focus area of data governance and DataGovOps. And let me go through a couple of these, and I'm going to go into detail into the first three. So in data governance, as Laura said, having a business glossary and a data catalog is really important because what the terms mean, what the definitions are, are incredibly important because many
00:25:00
organizations have multiple definitions for the same term. It creates lack of trust in data, it creates slowness, and really trying to understand that. And so the idea here is can you treat the business glossary and data catalog as code or configuration? And I'll talk a little bit what do I mean by that, and when we change something, can we deploy a piece of that change along with the work?
And I'm going to show you an example. And so it doesn't mean that it's a replacement for a business glossary, it's really talking about the change, the process of adding to a glossary or changing a glossary or a catalog. And then the second data lineage, trying to find out sort of what's the story in human-readable form as your data traverses systems and sources, and where did it come from and what's happened.
And I think those are also data catalog systems provide that, but DataGovOps is a little different focus. It's less on the data and more on the processes that act on the data and trying to get ahold of those, because if you think of the data's journey from a source system to a database, to an ETL tool, to a data science tool, to a biz tool, in multiple versions, it has a whole bunch of process steps.
And trying to get ahold of what happened on those. And then the third one I'm going to talk about is really... And I think it is important to sort of profile your data, understand its quality. DAMA has a whole set of characteristics to tell you whether data is good or not, and I think those are accurate, but they're kind of like a balance sheet in a business. They give you an assessment at a point in time.
But there's another thing when you are running a business, you look at its cash flow, and that's over time how things are going in and out. And I think automated data testing is like that, because as data goes from its journey from source to value, it gets transformed and stored and artifacts are added. And so how can you tell that that journey is correct and that you're having a low-error environment? And so I'm going to talk a lot about those three things.
And then the last two I'm going to come back to after we talk about those three things, data security and defined roles and responsibilities, because I think it'll be a little bit more helpful. So the first part ... a data catalog. Now, if you look at my screen here, it has this thing that looks like a graph.
And if you're doing some work, let's say you have an existing system that does some ETL work, some visualization work, some data science work, and you're going to put some new data into it. That may mean you have a new table and perhaps you're joining that table to an existing factor dimension in a database.
Perhaps that new data that you're doing means that you're going to have an update cost to your predictive model, and you're going to have a report that changes. But the fact that you have a new table means that you have new information, right? And you have something that should go in your data catalog to say, "This table is new, and here's what this table means, its definitions, what each column means." And those are actually really important.
And so our view in DataOps is that all those things need to be deployed together. The I'm going to introduce new data into my database, the new schema that goes with it, changes to the model, changes to the visualization as well as to the catalog. All those things are a deployable unit that should be deployed at the same time from your development environment into production.
And the only way you can do that is to think of your data catalog as code or configuration. And the delta, the change to the data catalog, is what you move from your development environment into production. And
what helps about this is a lot of times data catalogs will lag production systems. And so if you can make it a deployable unit, it becomes easier for more people to get involved in the data catalog process instead of just, "Okay, it's in production now, you've got a ticket, and you've got to go fix it, and maybe it'll get done, maybe it won't." It becomes part of your change management process.
And seeing it as code, seeing it as part of a deployable unit, I think is sort of the DataGovOps perspective on a data catalog and updating a data catalog.
And then another idea is process lineage. And so in this graph where you can see the steps, where it goes, gets some data from SFTP, it builds some facts and dimensions, it does some sales forecasting, puts some stuff in Tableau, and updates a data catalog. Now, there's a whole bunch of process steps here, and in some organizations, that may mean they all run in one tool. But many organizations have multiple tools and multiple locations, cloud, on-prem, where this data processing happens. Sometimes there's centralized teams, sometimes there's decentralized teams that do self-service tools on top of it, and just trying to find all the steps are really important, especially when the VP says something's wrong with your chart, and which team owns it, what process happened,
00:30:00
and trying to keep track of that process. And so we think an important part of metadata is the code that acts on data. And so whether that could be a script, an ETL code, the Python model, the Tableau workbook, all that stuff needs to be stored in Git in a version control, so you can see what process acted on this to get it there.
Likewise, we also think that the as-run process, I did this at this time with this test results, and here's what I got to produce this data set. That's also an important thing. It's not a replacement for data lineage. I kind of think of it as a brother or sister to it. It doesn't, in a human readable way, say where the data came from, but it tells the processes that act on it. And I think that's an important part, both from an understanding standpoint as well as from a policy or a compliance standpoint, saying, "Here's what's happened.
I've recorded the journey of the metadata of the journey as the data's been transformed to value," and that ends up being stored in a process store. And I think that's a very important part of trying to understand what happens. And then the third part is of automated testing. So in my career in being a software engineer and managing data and analytic teams, I spent 15 years in software, now I spent 15 years in data and analytics, and by far, managing data and analytic teams was harder.
And mainly because I just never liked coming into the office in the morning and finding problems, where the VP didn't like the report or didn't understand the report, or something changed, or the data was wrong, and I just wanted to focus on lowering error rates. And so this is where I started to look at factory and Deming and trying to do it, and I think trying to lower your error rates.
And so to do that, you need
to not only profile your data, but you need to automatically test and monitor and observe your data as it's working its way through production. And do that on top of all the tools that you currently have. And if something's wrong in the data, maybe you've got a bad raw data file, maybe somebody committed some code that is wrong in the data transformation or modeling process. Get a notification as soon as possible before your customer sees it.
And that can really help you improve confidence, because it's not about only well-governed data. What well-governed data means is that it's trusted, and if people trust it because they know that you, as a data and analytic organization, have sweated the details. And manual testing and checking is fine, but again, can you automate this? Can you make it easy and repeatable?
And can you take these automated tests and data tests and put them in production so they can run over and over and over again? And that focus on-- I see error reduction as kind of a superset of data quality, because poor data quality is a problem, but it produces an error. Perfect data quality and a bad transformation also produces an error, and your business customer doesn't know, care which. They just know it's an error.
And a lot of teams are suffering, like I did, where they come in the morning with dread, and trying to get rid of that dread gives you more time for innovation and improves customer data trust. And so those three aspects are, I think, an important perspective on data governance that is about how to change a business glossary, how to have process lineage, how to do automated testing instead of just data quality definitions.
And then the two that I don't really have too much time to talk about, because we want to get to questions, are data security. And so in the DataOps perspective is the first part is that a lot of organizations have self-service teams doing data analytics, and that gives a lot of data governance professionals heartaches because they're like, "Well, what are they doing with the data? Is it right?
Are they actually documenting it?" And we have some technology that allows you to sort of take self-service environments and kind of put it in a nice basket and basically put a speaker on it or put a line in it to see what's happening and seeing if they're working properly. And that's one aspect of data governance.
And then the second is to do development, you need test data. And in a lot of organizations, this seems like a really nitty, boring thing, but actually having really accurate test data that also is privacy aware, that passes all the security things is a pain point in a lot of organizations. And so having good test data enables you to
deploy faster with lower errors, which means you can iterate faster and therefore learn more. And then finally, the defined roles and responsibilities. And what do I mean by agility?
I guess what I mean by this is, I think the process of data governance is not just for data governance people. I think the people who do data science and data
00:35:00
engineering and data visualization should all be part of it. And so if you are making that change, like I suggested in this example, you're doing the ETL, you're doing the modeling or the visualization. Perhaps those people should update the data catalog because they know the data best. And that could be your first version, and your data catalog could improve over time. But always having the data catalog work and the data description done by the data governance person, it seems suboptimal to me, because I think that one of the spirit of Agile is that everyone should own the end results, and the team who owns it should treat data governance as part of their process. And so testing and automating and documenting is something that the development team should own.
It shouldn't be just thrown over the wall to your data governance team. And one of the spirits and generators of DevOps was that the developers were throwing their code over to the wall of operations. And I think that is one of the things that the greater Ops movement is trying to do, is to not have things that you create thrown over the wall to other people to sort of pick up your trash. And so I think that's another perspective.
And so that's it in terms of my discussion of DataGovOps and what I'd like to do is stop now, given that we've got about 20 minutes left, and see if we've got some questions that have come in that we can answer. Yes. Thank you both, and I encourage everyone to enter your questions into the webinar control panel in your screen, and we'll try to get through as many as we can in the next 20 minutes. So this question has come in, and this is for Laura.
It's about definition of data governance. "I agree with the idea that we need to focus on usage, but I find analytics too restrictive when talking about usage. For me, governance is also about ensuring quality, availability, compliance in the operational space of the company, not just in analytics." What's your view on that? Sure. Yeah. Yeah, I would agree. Usage shouldn't be limited to just your analytics department. And most analytics departments couldn't keep up even if they tried.
I don't know very many analytics departments that are much bigger than
2% of your total employees, and it's just not conceivable. So the thing to me is really about what I call in the book radical democratization of data, and the only way to do that is to get more people to look at it. The more people that look at it, the better your engine from data is going to be because you're going to have people looking at it and challenging it.
And we really have to make that transition to ensure that when we have people look at data and challenge it, that those challenges are met with, "Thank you. That's awesome," instead of, "Oh, crap," right? And that's probably one of the most powerful things that can happen as you make that transition. So I would agree with you. Usage isn't just being determined by the analytics group. I think that's very important.
Great. Thanks. Wow. Yes, we have a ton of questions, so I'm trying to pick some good ones here. Okay. So is DataGovOps a mirror of data as a service?
Is that for me? I think that's for you, Chris.
I don't think so because the idea of ops is the process that you use. It's a people and process as opposed to focus on data and tools. So whether you do data as a service or data modernization or whether it's a data warehouse or a data lake or self-service data versus centralized data, they all have an operational part to develop things and deploy them and test them and monitor them. And the DataGovOps idea things equally apply because it really is about the process that you use to build things and the people, and less about the actual artifacts and methodologies. So it supports data as a service, but it's not really about data as a service. It's about how you get there.
How do you get to data as a service or get to a data lake or a data warehouse or self-service data? Okay, great. Here's another great question. How do we assess the DataGovOps maturity of the organization to define a path to formally implement and improve?
Should I take that one, Beth? Yeah, you can start. We've actually been working on a DataOps maturity model, and we've got a survey out, and we've got a model that we've been using in some consulting engagements, and that we're writing a white paper on. We actually don't have a DataGovOps maturity scale, mainly because it's pretty new, and these sort of ideas, at
00:40:00
least at conferences that I talk to about deploying, about the sort of five points I made are still pretty new. I don't know, Laura, would you have a better answer than me? No. No, I do not. And I'm excited that you guys are developing a maturity model because I think a lot of people would appreciate that.
What I can tell you for sure, and Chris, I'm sure you feel the same way, but to me, it's really about progress over perfection, right? We just have to continue to make progress, and even if you feel like you've totally failed from a data governance perspective or a DataGovOps perspective, or even DataOps perspective, you learn an enormous amount in your attempt, right?
And so you take those lessons learned, and then you continue to improve. That to me is really the main goal. Yeah. Because I think that sort of growth mindset or love your errors or when you have a problem, it's not an opportunity to shame, but an opportunity to improve. Those psychological states that teams have I think is important, and I've seen having to work with teams to help them grow from that.
Because if it's 2% or 3% of the people or 1% who are doing data analytics in our organization, they do feel put upon and when something goes wrong they take it personally. And so having more of a growth mindset and, "Okay, something did go wrong. We understand why, and we could put in a test so it doesn't happen again," I think is a much better way than saying, "Oh, it wasn't our fault.
It was our supplier's fault." And your business customer doesn't care if it's your fault or your supplier's fault or the server went down. They just want it fixed. Yeah. And they don't at all care about your world. You just need to be able to build systems that enable robust delivery and enable them to understand and trust the data.
And so I think all those things have to go into a mindset about a continuous improvement mindset. Yeah. Agreed. Here's another great question. Our organization is currently implementing an enterprise data warehouse and shifting the focus to data governance as we implement. Any tips or focus points other than these that we could take into consideration?
Oh boy.
Well, let's see here. I'm sorry. You can probably hear the pounding upstairs. It's lunchtime, and my son is being super loud. So you're currently implementing an enterprise data warehouse and shifting focus to data governance as we implement. So incredibly important that as you are making strides from a data warehousing perspective, you are also making strides from a data governance perspective.
Just the fact that you're doing those two things together is very, very important. I would encourage you first and foremost to not boil the ocean. The faster you can get to high-value items that need to be governed, the better off you're going to be, because not all data is important to data. And I know that's sacrilege, but it's true, and so I would encourage you to really try to focus on high-value data and find some
I feel like I'm stepping into a little bit of a pit here, but find some data stewards, and I use that term lightly, to focus on those high-value data pieces because I think that's very important. Chris, how about you? Oh, yeah, 100% agree. It's focus on value. There's that Kevin Costner movie from 25 years ago called "Field of Dreams".
So don't build it and they'll come. Right? Don't spend a year building something that's the perfect data warehouse. Incrementally focus on value for your customer, and as part of that, do everything that you need to do. Do data governance, do data security, build your data warehouse, build the visualizations, and work in a more iterative way that's focused on customer value.
And I'm always saddened to see the people who follow the "Field of Dreams" approach, and they have teams of 10 or 20 or 30 people, and they've spent millions of dollars on a data warehouse or a data lake or whatever happens to be the fancy term at the time, and then they don't succeed, and their projects a year or two end up getting cut. And I'm concerned that actually that may be happening now with a lot of companies. They've invested in the cool, shiny new thing, and that is usually a technological fix on the assumption that if I get the technology in, I've built it and they'll come. And so I think that is, in my experience, that just is a setup for failure. So you need to find a way to focus on delivering increments of value to your customer that they can use.
And as you do that, you can build a data warehouse. Okay, great. We have several similar questions related to tools and products
00:45:00
that support these DataGovOps processes. Laura, do you want to take that one? Kick us off on that one? Well, I will happily kick that one off, and I will tell you that my best advice to you is do not focus on the tool. And while I love a good data lineage, data catalog, and I genuinely do, any tool that you buy in which you don't really know how you want to use, is going to be shelf wear And I think this is one of the true things about data governance, probably more than anything right now is I really want to encourage you to think about what your program objectives are.
Think through those four items that we talked about early on, which is usage, quality, lineage, and protection. Make sure you understand what each of those are. Make sure you get some great executive buy-in. Ensure that you understand the process and how you're going to execute that process. Make sure you know what the heck that tool is going to do for you and why long before you buy it. And here's the primary reason why.
Data governance has always been really pressed with return on investment. It is an incredibly difficult thing to be able to tell you, "I got X dollars because I spent X dollars." And one of the challenges that happens in most corporate environments is when you buy software, the second you buy that software, you are on the hook for proving a return on investment for that investment that you just made in software. And if you don't really have a great understanding of what you're doing, how you're doing it, and how you can drive the value of that data, you're going to be back on your heels almost immediately.
So I would encourage you almost to, and sacrilege as this sounds, almost I'd prefer that you use Excel and really get yourself a base understanding of what you're doing long before you invest in a data catalog tool. And when you're ready to invest, then please do that with no regrets. But make sure that you are understanding what your organization is going to get out of that tool, so that you can be as successful as possible.
So that's from a data lineage perspective, but I think there are other tools that will help you solve the DataOps perspective, right? Oh, yeah. So that's me and- Yeah ... I always hate when I say this, but I'm going to say it anyway and drive them nuts, and that I agree with you. It's like if you're going to do DataOps, start small. And my DataOps journey started in 2006 where I read a Japanese management book, and I said, "Do a quality circle." And my quality circles technology was Excel, and all I did was get people from data engineering and science and visualization together, and make a list of all the problems we had in the last few weeks. And then I kept doing that, and pretty soon we started to see patterns after a couple of months. Hey, these three problems are related to this root cause, and you know what? We fixed that root cause.
And then every time we keep doing it week in and week out, or month in and month out, and you know what? The amount of errors went down. And did that involve any technology beyond Excel and meetings? No. But it also involved a change in mindset that errors are a learning experience. It's not about an opportunity to shame someone. It's an opportunity to improve.
And I think that is a very low-tech way of getting into the fact that you need to build your organization for high change. And analytics isn't about building a house and walking away. It's about building a system that can continuously change and grow based on the needs of your customers. And so, we do have some software that we think is awesome to do DataOps. And we certainly want to talk about that, but we think DataOps or DataGovOps, it's better to start with increments of value and get some value first. And then you'll actually start to see the need for more robust software.
Yep. Totally. Great. Thank you both. So do you have a point of view on how data access should be managed in a way that balances utility and data privacy and protection? Who should own this, given that the scalable access control is technology-heavy and the business risks are distributed across functions? Yeah. Good question. So I think about it like a little bit of a trifecta. So what you need is a relationship with your compliance, and every group is different, but it's compliance, privacy, risk.
Sometimes it's three different departments. But they are responsible for creating the policies and procedures associated with this. They're the ones that are responsible for keeping up with legislation, understanding how that impacts your organization. Then there's a group in the middle that hopefully is in the middle, called information security, that is responsible from a technology perspective to ensure
00:50:00
that all of the things associated with the policies and procedures actually are played out in the technology in some way, shape, or form. There's an aspect of that InfoSec that ends up being a relationship back to your data people, because that's where we live and breathe, and oftentimes from an access management perspective. I really encourage you to ensure that access management, the literal providing access to data, is the responsibility of your InfoSec department, not of your data department.
And I have been a manager or a leader of those teams that have to do that because InfoSec didn't. But it is not the way to do it, and it particularly challenges it from a compliance perspective for a lot of legislation that occurs. Those things should be distinct and separate. So your access management in a perfect world should come out of InfoSec.
And from a data perspective, you make sure that InfoSec understands how that data is protected, and you follow standard data classification matrices to ensure that data is truly protected. If you do not have an InfoSec department that can do that for you, then you have to do some asset management protection. And while that's not optimal, it is feasible.
So, I think that's kind of in a best case scenario, that's how that works. Your compliance, risk, and privacy departments, policies and procedures. Your InfoSec is responsible for the technology, and your data group has some responsibility to ensure that that data is protected and partners with that InfoSec team. But from a reality perspective, that is one of the reasons why I think it's very important that you separate out the data governance from the compliance and risk, to make sure that we are really defaulting to usage where it's appropriate.
There are some data sets that should never be exposed, and that is where that partnership with your compliance and your InfoSec team will become very important. Yeah. And I think my perspective is very operationally focused. If you're looking at who can access data and trying to get the ACLs and roles and permissions of people who can access data in your data warehouse or in your data visualization tool, that data about who can access what and what tables at what time and what location, that's code. That shouldn't be manually put into a system, and you shouldn't have to register a ticket for someone to do it manually because these sort of keyboard-based manual processes end up being the lag in time and end up having errors. So number one, if you're going to deploy
that new table and the new model and the new visualization and the new description of the table, you should also deploy the security for that table at the same time in code. And again, it's about having a unit of deployment where it's complete, and being able to handle that as code and not as, "Oh, I've done it in the last mile. I got to call someone up and put a ticket in." Because that's where a lot of things happen. And then the second is monitoring whether the data is being used or protecting it. And so we've got a partnership with a company called X8 that is trying to promote this idea of DataSecOps and allows you to sort of put some limits on raw data access or data access that people have.
But I think it is an incredibly important one because it's these sort of data defense, and we all hear about Twitter yesterday getting hacked or the company X that lost 10,000 records. And compliance and security is incredibly important, and InfoSec is incredibly important because a lot of times these errors don't happen because they're malicious.
They just happen because somebody forgot something, or somebody forgot to key something in, or something wasn't tested. And the degree that you can automate that and take the human out of it through scripts, through toolage, I think is a great boon to making things work. Great. Thank you both. This question is somewhat related.
How does DataGovOps apply to more ad hoc use cases such as self-service analytic and BI teams? These folks can sometimes not have practical knowledge about version control or practical coding. Should we be upskilling these groups or changing their tool sets? Well, yeah, I think one of the things that we didn't have time to talk about, and we actually have another webinar about it, is we are calling them self-service sandboxes. And it'll help teams get a sandbox environment where they can do their work, and then providing them a control that if they go out of the bounds of that sandbox, someone knows about it. And then third, it gives them a path to production. If they want to be able to take their work and take it from a one-off and turn it into something real, what's that like?
And so I think there is an opportunity in the DataOps framework to focus on maybe... We're still trying to get the name right, self-service sandbox, governed
00:55:00
self-service. But trying to take all the great tools like Alteryx and Tableau and all the power that comes with those tools, but put them in a nice little nest and have a line on it that the central IT can make sure things are going right and take it away if things are going wrong.
Any thoughts on that, Laura?
No, I think that's it. Great. Okay, this one is related to banking, so let's see if we get a good answer here. Have you tried using DataGovOps in BCBS 239 compliance data governance framework environment? Any advice from both speakers is much appreciated.
So I don't know if you're familiar with that standard, but if you have any kind of general comments, that would help.
Yeah, unfortunately, I'm not familiar with that standard, so I can't guess. Okay, great. Okay. Who defines who has access to data?
Who are the data owners in the organization? Ah, yes, the longstanding data owner conversation. So it is a mix of the compliance risk, privacy, right? They get sort of a determination as to what data sets have always should be protected. But longstanding sort of best practice in this space has been that you identify a data owner that is generally closest tied to how the data is initially created.
So if it's finance data, it's a finance person that owns that data. And they get to determine whether or not somebody has access to it. That often becomes problematic when you have what happens often in healthcare, or in healthcare, in analytics, the Freudian slip. In analytics, where it's sort of like people become very protective of their data, right? And so it's like, "It's my data, you can't have it.
You'll never understand it." But really, there is no easy, fancy way around this because you do need to protect your data. So you identify a data owner who's closest to how you source that data, and then your best approach, if and when those scenarios happen where people are trying to protect it too much, is to have a conversation about improving the transparency of data throughout your organization.
Chris, thoughts? Yeah. I think I brought up test data management, and it's, again, it's a small part, but if you're going to do new things with data, you want to do it on test data and not real data. And so you've got to de-anonymize it or pseudo-randomize it so you can actually have test data that's representative of your production data, but isn't. Because if you've got healthcare information or Social Security numbers or credit card information, you certainly don't want copies of that floating around and ending up in an S3 bucket that someone forgot to apply a lock to, and then it was scanned, and suddenly it's in the news.
So protecting your data in development, I think is a small thing, but it's a part of being able to do agility as well as protecting your data once it's in production. And of course, that's an even more serious thing because that has the real data on it. And also goes to how people design their development environments versus their production. And in some organizations, the development team, the data scientists, people who do data engineering, they don't even have access to production. And perhaps that's good, but they still need test data to do their work. So it's an effort that is sometimes seen as building test data sets is an effort that's sometimes seen as an afterthought, but actually is an important driver of helping teams deliver things faster with higher quality.
Okay, great. Thank you. So we have one minute left. I think this is probably our final question, and it's, I think, a good question to close with. I apologize to you if we didn't get to your question, but we will follow up with you directly afterwards with some feedback. So this question is, would Chris or Laura have any recommendations on how to prioritize data governance projects and proposals once a data governance framework and policy is established in an organization?
Laura, do you want to take that?
I would say go with value, and it helps to identify any objective method of identifying that value, whether it be people impacted, money associated with it. It's not very exciting. But though if you can start with the objective and then shift to the subjective, but it really has to be about value and the impact to your organization overall.
01:00:00
That's the best, cleanest way to prioritize those things. And I think I do talk a little bit about that in the book. Shameless plug.
Yeah. Laura, why don't you do some more plug of your book?
Well, you can buy it on Amazon. Please buy several copies. But it's called "Disrupting Data Governance," and it really does talk a lot about sort of upsetting the apple cart when it comes to how we define data governance, and then aligns well to this concept of agile and introduces the, I call it DG Ops in the book, because honestly, I just got tired of typing DataGovOps.
But it aligns well to the conversation we had today.
Any thoughts on that, Chris, or people would want to- No, yeah. Buy Laura's book, and we've got some resources, a manifesto in a book as well. So if you're looking for some summer reading, follow these links and you can find it. Great. Well, thank you so much, both of you, for taking the time today to do this webinar. Thanks to all the attendees for joining us.
I hope you all found this as useful as I did. Just a reminder, we'll be sending out the recording and the slides in the next 24 hours, so be on the lookout for that in your email. We'll also notify those of you who want a copy of the book, so we'll let you know about that via email as well. If you have any additional questions, please don't hesitate to reach out to us at DataKitchen or to Laura directly, and I hope everyone has a great afternoon. Thanks again.
Thank you. Bye.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Questions from this session
What is DataGovOps?
DataGovOps is what data governance looks like when DataOps practice is applied to it. Where a governance program produces a business glossary, data catalog, lineage, data quality definitions, security, and defined roles, DataGovOps produces the same things as running machinery: glossary and catalog as code, process lineage captured from real runs, automated data testing, self-service sandboxes with test data management, and roles that can change without a committee.
How is DataGovOps different from traditional data governance?
Traditional governance documents intent; DataGovOps automates and verifies it. A catalog maintained by hand drifts from the pipelines it describes. A catalog deployed as code alongside the schema, model, and report changes cannot drift, because the same deployment that changes the data changes its description.
Where should a data governance program spend its effort?
This session proposes weighting governance toward use rather than control: 40 percent on increasing usage of data, 25 percent on data quality, 25 percent on data management such as lineage, and 10 percent on data protection. The point of the rebalance is that governance justified only by risk reduction tends to be experienced as a tax, while governance that measurably increases data usage pays for itself.
What kinds of automated tests support governance?
Five types are named: traditional data quality checks, statistical process control, location balance tests, historic balance tests, and business-based tests. They are meant to run automatically in production, across the whole tool chain rather than one tool, sending alerts and keeping history so trends are visible.
Why does DataOps claim how you work matters more than what you build?
Because most failure is systemic rather than individual. Deming found 94 percent of causes were common cause — properties of the process, not of the person who happened to be on shift. A team that responds to each error by finding someone to blame keeps the process that produced the error, and so keeps the errors.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.