On-Demand Webinar · 59 min

DataOps For Beginners: The Art of the Possible With DataOps

Chris Bergh and Beth Pfefferle on making the internal case for DataOps: how it supercharges rather than replaces the staff and technology you already have, how to build the business case, and how to collect the small wins that show its impact. Recorded February 2022; updated August 2026.

Presented by Chris Bergh

What you'll learn 7 points
  • The 2021 DataKitchen and data.world survey of data engineers found 52 percent hope and pray that things do not break, 78 percent are stressed enough to want a therapist, 70 percent expect to change jobs within a year, and 79 percent have considered leaving the career entirely.
  • Gartner put innovation at 22 percent of a data team's time in 2020. The other 78 percent goes to errors and manual execution.
  • At Celgene, later Bristol Myers Squibb, seven data and DataOps engineers ran hundreds of integrated data sets with more than 10,000 automated tests and more than 100 schema and data changes a week, with very few errors or missed SLAs.
  • Production error rate is not only a data quality number. Late data that misses an SLA, a data processing failure, a code change that broke something, and a broken report or model all land on the customer the same way.
  • Production testing covers five test types: traditional data quality, statistical process control, location balance, historic balance, and business-based tests.
  • A DataOps engineer works on the pipelines rather than in them, taking nuggets of ETL, SQL, Python, and XML from other roles and putting them into pipelines, tests, deployments, and process measurement.
  • Bringing DataOps to an organization follows six steps: educate, find a first project, establish a community of interest, demonstrate value in a month or two, iterate onto more use cases, and expand with a staffed center of excellence.

Slides

65 slides

Transcript

Show chapters and dialogue 9,632 words

00:00:00

So hello everyone. My name is Chris Bergh. I'm the head chef of DataKitchen, and I'll be your host and presenter today for this DataKitchen webinar on DataOps for beginners. But before I start, I'll have a few housekeeping items to share. The first is that this, like all our sessions, will be recorded and posted up on our website in a few days, along with the slides that go with it, so you'll be able to share it.

And then second is that as you have questions, feel free to drop them in the chat, and we'll target to answer those questions at the end of the meeting. And so, with that being said, why don't I start? My name is Chris Bergh. As I said, I'm an old data nerd. So I spent 15 years doing software and then the last 15 years kind of doing data science and data engineering and managing teams.

And so my background is really the genesis for this. And so what I would like to do is kind of start off with a few slides to sort of talk a little bit about me and where this idea of DataOps came from, and what it is about it that has

kind of captured, I think, myself and the problems that I've had. And so if I go to my next slide. So I spent, as I said, I spent sort of 15 years writing software at companies like NASA, laboratory at MIT, internet startups. One got acquired by Microsoft. So I thought I was a pretty good software engineer, software engineer leader.

And I thought in 2005, when my kids were small, I'd do this data and analytics thing because I'm a big software guy and it's easier. And you know what? I was very wrong. It was not easier at all. And so oftentimes I had small kids. We'd bring them out to the soccer field. I'd be watching a game, and then I'd get a buzz on my BlackBerry saying something happened.

This weekly data build was wrong, and I'd have to sort of skulk off the dirty looks from my wife saying, "Oh, I know there's a problem," and either fix it myself or find someone to fix it for me. And the company I worked for did analytics for kind of sales and marketing people. And so there's nothing worse than delivering the wrong data to 5,000 sales and marketing people, and then have the VP call you up and basically dress you down and say to them, "I'm sorry that we did that. It's wrong.

It won't happen again," and have them sort of yell and scream at you, or even worse, passively aggressively talk to you and talk down to you. So that's just not fun. Another experience that I had as a data team leader is hire some engineers, and they were very proud that they put last-minute changes in because customers always have requests.

And so I talked to different data engineers, and they'd be very proud that they put in a last-minute change. It got into the build. It worked. And when I asked the question, "How did you know it was right?" They say, "Well, I haven't got any calls, so I know it's right." And so,

I think that sort of hope that things right is also a characteristic of the team I inherited. And then also, I just sort of hated going into work. I got the Monday morning blues. I'd just wait.

00:05:00

I wouldn't look at my BlackBerry while I was driving in, mainly because I was just afraid that I'd get that email saying, "Your data's wrong and your team sucks." And I don't know if you've had that sort of feeling of dread when you go into work. And then in this application, we just had a lot of different data providers, sort of 50 to 75 per application, and most of them didn't really care that we existed.

They would give us late data or wrong data, or change the format of their data, change the meaning of their data. And so, when I first got there, we didn't notice it. The data providers would change their data. We wouldn't notice it, and sometimes our data and our dashboards would have been wrong for a month or two, and we didn't even know. And that's another email I didn't like writing, saying, "Yes, we noticed that your data's been wrong for the last six weeks." And then the company I joined was run by a physician who went to Harvard Medical School and he was very good in analytics, didn't know a lot about technology, and he would go off and talk to someone in a healthcare company, come back with a great idea.

I'd sit down, look at the data, work with a data engineer and a data scientist and maybe someone who knew visualization. We'd sort of pound on it and I'd come back to him and say, "You know, David, this is going to take us two weeks to do," with a lot of excitement. And he'd look at me like I had killed a patient on a table and said, "Chris, I'm sorry.

I thought that wouldn't take two weeks. I thought that wouldn't take two hours." And so my team never could go fast enough. And we hired a lot of people, and one of the engineers I remember is sort of a few weeks in, it was my 42nd birthday and he was 24. So we had the same birthday, except it was reversed.

And we were doing a one-on-one, and he was so upset. He actually cried in front of me because he just felt like he couldn't go fast enough. Things were breaking. He just felt like he couldn't be successful. And then that was a tough conversation because he was a great guy, and he was very smart.

And then we just hired a lot of smart people, right? And they like their favorite tool. Some people like R, some people like Python, some people like Tableau, some people like another tool. Some people like to do the work in SQL, some people like drag and drop tools. And so people love their tools.

And so finally, after sort of a year, year and a half of this, my wife sort of got tired of me, and I've been married to her now for 30 years, and said, "Could you either stop complaining or just fix the problem?" And so, being a good husband, I actually listened to my wife, and this is sort of my story. My co-founders and I have kind of spent 15 years trying to fix these sort of emotional problems, and realize that it's not our fault, but we haven't taken the time to build a system to fix these problems.

And so, unfortunately, my life story is still playing out today. And so we did a survey about six months ago with a company called Data.World, and we sent it out to sort of 600 data and analytics professionals, I think most of them were data engineers. And we got these results back that's saying sort of 52% hope and pray that things don't break.

78% of data engineers are so stressed that they admit they need to see a therapist. 70% expect to change jobs, and 79% have considered switching careers entirely. And so if you've listened to what I said first, when I got these results back, I was not surprised. I was sad, but in the last 15 years, there's been a lot of technology change. There's been a lot of new things.

We've gone from on-prem to cloud, we've gone from small data to big data. We've gone from almost no predictive data science to it being everywhere. And these things are great, and they're great, but the life of people who are delivering data and analytics has almost gotten worse. And so how do you solve this? How do you actually stop the suffering that data teams are having? And to us, it's-- I want to talk about that today, and that's really about the idea of DataOps for beginners, because the solution to that suffering is not sort of a new-- It's in the people and the process you use to produce insight.

And so I'm going to talk about that today. So I'm going to go through what you do is much less important than how you do it, sort of talk about data and analytics in general, go through some possibilities, and then sort of talk about, have a conclusion. And it should take about 45 minutes, and as I said, we'll do questions at the end.

So there's a lot of work that we got to do, right? There's sort of a bucket of requests that we get from our customers.

00:10:00

Could you change the dashboard? Could you integrate a new data set? The model looks off. Oh, I don't understand this data. Is our data catalog right?

How could I segment my customers? There's a lot of data tasks that we have to do and a lot of data tasks that we want to do, right, because we're interested in the power of data and how it can improve the organizations and the customer experience for the people that we influence. And so I think a lot of us are kind of just really focused on getting the next task done. I've got a list, and a good day is I've got a lot of tasks done. And so I think at one sense, that next data task focus is making us blind to a greater systemic problem. There's a bunch of problems upstream from where we're currently working.

And that lack of focus in the problems upstream is causing not only the emotional pain, but also general project failures. Like 60% of all data and analytic projects fail in some ways. 87% of models never get into production. Many pipelines have too many errors. And so this idea that you need to look upstream at the source of the problem, and it's about how you do things, not what you do, not the task, it's about how you do those tasks.

How do you develop? How do you iterate and change things? How do you deploy things? How do you monitor things? How do you test? How do you collaborate? How do you automate? That's what we're going to talk about today. So if you're looking about DataOps as the thing about how do I write some better SQL? How do I actually develop a dashboard and change my dashboard? It's not quite there.

And so

as I said, just one quick question. The slides and the video will be shared after the meeting. We set up a webpage to do it. So, let me keep going here. So the idea here and the concept that I'm going to talk about, and this is really a concept discussion, is this idea of DataOps.

And think of it as a set of technical practices and cultural norms and architecture patterns that try to drive sort of four outcomes. And one is that can you actually rapidly cycle change, iterate quickly, get new things into production in the hands of your customers with minimal problems? And then second is can you run the things that you have in production with really low errors?

And we're going to talk a little bit, a lot about that. And then the third is how do you not spend so much time in meetings? How do you collaborate when you've got lots of people and lots of tools and lots of environments? And then finally, how do you get some measurement of what you are doing and these processes that your teams are using, how much work they're doing, how many errors they're creating?

And so from an intellectual idea standpoint, DataOps is just sort of stealing some ideas that have worked well in lots of other cases. And in some cases, one way to think of DataOps is that it takes the same sort of principles that people use to make Toyota and automotive factories better to help software ship quicker and applies them to the world of building data science models, data engineering, delivering insight from data.

So it owes its heritage to Agile and DevOps and lean manufacturing. And so I'm using the term DataOps in a very broad way here. There are other terms that people have. There's sort of Model Ops and MLOps and DevOps, and I can talk about those in the questions if you want. But right now, I'm really speaking about the entirety of the analytic development process. And so what is that?

What is the way that people develop data? So let me talk a little bit about that and how the ideas of cycle time and error rates and collaboration and measurement for DataOps apply to it. So we've all heard, one thing that's happened in the past 15 years, that's for sure, is that data and analytics is a team sport.

It's not done with just one person. And so there's different types of people on teams, and here's a gross simplification, but there are people who put data together, data engineers, analytic engineers, ETL engineers, sometimes they're in IT. Their job is to source data, load data, clean data, integrate data. And then there are people who take that data and actually apply mathematical models to it that group data together, that separate data into buckets.

And then there's a third team that's kind of, you could think of them as

00:15:00

analysts or BI teams. They're-- or a self-service team. They take the data and visualize it and sometimes mix other data in. And then there's another group that tries to explain to the customers what it is. They develop catalogs and models, and that's called data governance, and also assesses the risk and security in data. And man, there are a lot of tools to actually operate on data.

And so over the past three or four years, every cloud vendor has got their whole stack, but they kind of fall into some patterns. There are tools to visualize data, develop dashboards. There's sort of 50 of those. Tableau is one you may have heard. There are data science tools like SAS or Python. There are tools to do data work, ETL, transformation.

There are tools to store the data, different database types, different data catalog tools. And there's different patterns to put all these things together, data lake, data lakehouse, data warehouse. And so there's a big, massive set of tools out there. It's almost a $100 billion industry of tools. And the tools are great. And so let's just talk about how those tools are used and talk a very sort of simple scenario. So let's say a data engineer goes to a CRM system or an ERP system and grabs some sales data, and they load it into a table in a database.

It doesn't really matter what it is, it doesn't really matter how they load it, what their favorite ETL tool is, or the database. And then the data science team runs a model on it, and they determine that Kelly is high value and Joe is low value. And again, it doesn't matter. It could be Python or SAS, however they do it.

And then there's another group that actually takes it and visualizes it. So they develop some nice charts, and then they actually mix in some data that no one knows about, and they use a, perhaps a tool like Alteryx to do it. Again, it doesn't matter. So there's a dashboard, there's some data loaded in, some data prep, and then there's finally the poor data governance team who's trying to say, "Okay, where did this data come from? What's its lineage?

What's its meaning?" And assessing the risk of security on it. So this all happens together. And the processes that people work and develop are almost largely separate. There's a data engineering team, data science team. Sometimes there's a hub and spoke model with multiple self-service teams. And so these teams work-- There are teams in themselves, and then they have to work together. And so this value chain actually spans teams and systems. So as an example, from a technical standpoint, we bring a lot of data together to do data and analytics, right?

ERP, CRM, supply chain, financial website, syndicated data, APIs, et cetera. But the process of putting data together, and that doesn't matter whether it's big data or small data, whether data is structured or unstructured, whether it is updated every month or in sub-second. You're sort of running a factory to produce insight on data. And so the data is loaded and transformed.

It is modeled and segmented. It is visualized and governed. And so the metaphor here is literally a factory, where each one of the workstations on the factory is that tool, all the different tools that you use and all the different scripts that you run. And that's what makes this so interesting. And when you look at people's architecture diagrams of data architectures, it's sort of a sea of boxes of different tools, and really because they are building a factory.

And so this factory, the question comes to you, do you want to run a factory that produces 1970s AMC Pacer cars that last 50,000 miles and have a lot of defects, or do you want to produce Toyota Corollas that have very low defects and last 200,000 miles? And so I think the difference between those is not only the fact that one's newer and one's not, but also the processes that have been used to do that. And that's where we're going to talk a little bit about.

But number one metaphor is think of it as a factory of insight. And then this factory itself is really interesting because in a factory you can imagine sort of, hey, there's one building and we all work together. That's not quite true. In a lot of organizations, there are different parts to the organization. There's a centralized IT team or data platform team that puts data together.

There are sometimes more than one data science team. There are multiple lines of business, spoke teams on there. They sometimes don't even have the same boss or the same boss's boss's boss. They might actually organizationally meet up at the CEO, and so they're just completely separate. And there's this term that's called Conway's Law, which means that when we build things that are technical systems, the parts of the technical system often reflect the organizational structure, which is, I think, true in data and analytics, but from an engineer standpoint,

00:20:00

that may not be the optimal way to do it. And so that actually can be problematic, and that's why thinking about this from an organizational and teams really matters. And then the idea here of DataOps, I think one of the key sort of tenants here is that be focused on delivering value to your customer, and then know who your customer is.

And so DataOps is, think of it as in broad speaking, an agile methodology. That means that you're delivering small bits of work, and maybe it's a new data set or a tweak to a data pipeline, maybe it's a new visualization, but instead of spending four months, five months building something, you're spending a week, a day, two weeks, and getting it into the hands of your customer and iterating and learning. So, I guess I don't ever want to write a large spec, spend three or four months on something, put it into the hands of my customer, and then they go, "Eh, that's not quite right." And so my feeling is, in everything in data and analytics, working in small iterations of time, delivering small bits of value, and getting feedback from your customer, whoever your customer is, learning from that enables you to do more work with less effort.

You're maximizing the amount of work you don't have to do, because by delivering in small batches, you're learning more accurately what your customer wants. And that's hard, right? To be able to iterate quickly, to be able to change whatever you've done quickly. And so that's where this idea of sometimes from software, from agile development also comes in.

So, let me keep going.

And

so the challenge in any data and analytics system is we've got lots of data rapidly changing on the left. We've got these data customers, and sometimes those customers aren't even human beings. Sometimes they're other systems. But they've got requests and changes that they want to make. And if you've ever worked in data and analytics, you know a mark of success is delivering something to your customer and getting 10 follow-up questions. And those 10 follow-up questions often have more follow-up questions, and it's sort of a river of questions that you get.

And so how do you handle that, where you want to run a really good factory of insight, data passing through. Man, you want to know that it's right and that it's perfect. But then your customer's going to ask follow-up questions and going to ask you to change it and then share it with people. So how do you do that?

How do you run fast and sort of not break things? And so

that's an interesting challenge. And then second is you've also got this collaboration aspect, where the factory itself isn't owned by one VP, it's owned by a bunch of people around the organization. And then lastly, if we're data and analytic teams, how do you measure all this stuff? Are you measuring that you're having low error rates?

Are you measuring how fast you deploy or how much work you get done, how much time you're spending in meetings? And so we're data people, perhaps we should have data about how we work. And so all these things make it quite hard to actually implement DataOps and work in this agile way, deliver bits of value in small bites, and focus on quality.

And so I'm going to give some examples here about how companies have done this, and as a way to sort of take these concepts. Our first part was kind of talking about DataOps in an emotional sense, what the problems are. The second part was more conceptual, what's the definition? How do we actually get this done? And the third here is a little bit more concrete, and I'm going to try to walk through some scenarios and mention some specific tools. And now DataOps isn't limited to these, but it's just these are concrete examples.

So let me keep going. The first case here is,

think of it as this. When your customers say to you, or have you ever heard these things from your customers, like, "Your data's wrong," or, "Your team is too slow," or, "What's your team working on? Can't you measure them?" I know I've heard them in my career, and it's not my favorite things to be said by my customers, especially, "Your data's wrong." And so let's talk a little bit about this in terms of some people.

And so, the first thing is imagine when your data's wrong, what does that mean? Well, that first goes to someone who's in charge of your production process the day-to-day, and maybe that's a whole team of people, maybe it's one person, maybe it's a shared responsibility. But think of Eric, who's our production person.

00:25:00

He's running the factory. And then you've got another person who's actually in charge of getting things into the factory, right? The person who does deployment. And then you've got a third person who sort of manages these processes, Stephanie. And so let's talk about each one of these in turn. So everyone's got a production perfectionist on their team, right?

And as I said, maybe it's a separate team, maybe it's an individual, maybe it's a shared responsibility among your team. But when it's your turn to run production, you want it to be perfect. And so, finding problems late in on Friday when the data drops and trying to figure them out is a grind and painful.

And so people who are in charge of production want to protect and perfect the daily grind of delivering data. They want to minimize errors and chaos, and they want to know where the problem is, and know what's wrong as soon as possible, and know it before their customer sees it so they could possibly correct it.

And so this production process, if you've done this and run and had to do the weekly build or, and have your time and support is hard. And so let me give a concrete example of this just to show it. So here's an example. So a customer has what's called telematics, and it's an insurance company.

And so what it means is that they put this box in your car that reports where your car is, how fast it's going, and what time it is, and sends that information constantly to a vendor, and then the vendor aggregates it and then sends it to an insurance company. It goes through a process, and then if you drive too fast or drive too crazy, your insurance goes up.

And if you drive reasonably, your insurance goes down. And so you're kind of trading this information for discounts on your insurance. And so this company had a bunch of different telematics data providers that were going, and it set up this beautiful system that was done in the last few years. And the data would come in from these telematics providers all the time. And it was called streaming data.

That means it's happening quickly. Every minute, new data is coming in. And it gets into a data lake. It has an ETL process. It has a predictive model and integrations to their policy system because, like, hey, if something's wrong, it goes in databases, it has a reporting application. It uses a streaming technology called Kafka.

And they built it, and it was great. And you know what? The day they got it in, there was some wrong data loaded in. They started to look at it and said, "Wow, my car can't go a thousand miles in one second." Or they started the trip, and then it suddenly stopped, and then they picked it up.

And the data wasn't looking right. And it wasn't the data team who was having to answer this problem. The customers were finding it. And then when they found it, this whole group of people would have to get together and say, "Well, where's the problem? Is the problem from the raw data, or did a container stop, or did somebody muck it up in the data warehouse team or the integration team or the web team?" Like where's-- sort of finger-pointing.

And so people would be having phone calls and typing on terminals and, again, it's Friday night, pulling your hair out, trying to find where the problem is. And this is sort of the latest and greatest up-to-date technology stack you could want. So what's the issue here? And again, this goes back to the emotional challenge. You build something, you spend all your time on the task of building something, but you don't look upstream to where the problem is. And the upstream is that there is no way to look across all these systems to be able to see if there is a problem during the production process and then alert people on this. And so, we're a software company, so some of our examples are from the use of our software.

And so, by putting something-- But the principles apply here. Put something on top of what you're doing. Monitor not just that the servers are up, but actually dig into the data itself and see if it's right. So in this case, go in and look at a car and make sure it's not driving a thousand miles an hour or doesn't move from the West Coast to the East Coast inside of five seconds. Just look at the data, look at what's happening with the data, and be able to make sure it makes sense.

And then if the data is weird, perhaps stop the process or in this case, quarantine the process so life can go on, that someone can check. And don't rely on your customers to find data problems. And I think that's a really important step because then it reduces the amount of pain your team has.

And inevitably, things are going to go wrong. You're going to find problems in data. You're going to have a condition that you haven't checked or something else is going to happen. These systems are complicated. They're combinatorially complex. But the best case is like you say, "Yeah, we found it. We found it quickly, and we're going to put in a test or a QC to make sure it doesn't happen again." And so what are the principles here?

00:30:00

So focus on error rates and embarrassment, is really what we're saying here. And to do that, automatically check your data In production, while things are running, simultaneously of the building and working of the data, and kind of on top of your entire tool chain. And send alerts and notifications if things are going wrong, because you want to know immediately if there's a problem.

And there's a lot of ways to check data, and we don't have time to talk about it. We have written some books on it. But the idea here is by focusing on errors, you actually get more time to innovate because you're not spending Saturday and Sunday trying to debug something in production and pulling your hair out. And you actually get more customer data trust.

And think of when your system has an error, it's down, it's in failure mode, and it's just less stress and embarrassment. And conceptually, what we mean by that is that this idea of focusing on error rates means that you-- and it's not only us who talked about this. If you look at Gartner says that there's only 22% of the time that people in data and analytics are actually doing what you think they should be doing.

And in our survey, 52% said errors are a major source of burnout. And conceptually, the idea here is that we've talked for many years about data quality and the dimensions of data quality, which is fine, which is a nice static measurement of data. But that's only one source of errors, having bad data. You could have perfect data but still be late.

You could have perfect data, be on time, but there's some processing issue that happened, or someone put some code into production that creates a regression or has an impact on the system. And so all these things impact the errors that are happening in your system. So drive errors down and of course, focus on data quality because bad data quality, it all starts. But it's not just about data quality.

Data quality is really one component of a bigger idea, which is driving error rates down in your system. And that's really an important thing. And know what your errors are. And as some of our team says, "Run towards your errors, don't run away from your errors." And so the second part here is really talking about Ahmed.

And so think of Ahmed as a person who's in charge of trying to get things into production. And there are many people who make changes to production, data scientists, data engineers, people who do data visualization, a data doer. And did their change have an impact on what's already working? Did it have what software engineers call a regression?

And Ahmed is like, "Is Eric going to be mad at me if I let something get into production and it breaks?" And so if we think about that, companies will have a lot of different ways of doing work. And so on the left, there's different systems. Sometimes there's two, sometimes there's five. There's a development system, a system test and QA system, a production system. They'll have different copies, and they'll have a process to move things from development into production. And that could be using software CI and CD tools, continuous integration deployment.

So they could have built a railroad to do it, to move things, or it could be manual. They're sort of walking along a path. But in this case, there's a picture of an Airflow, what's called a DAG. And so how do you tell in the development or QA that there's been an impact, that there's been a change, that this change had a problem? And to do, we think that what it really means is that you need to have a test here and build tasks across all your tools so that you can find where the problem is. And this term, from a principle standpoint, we like to think of automated testing, and I like to think of the idea of pulling the pain forward.

And what that means is you have lots of smart people on your team. They're really good at doing SQL or doing Python or doing visualizations. They know their world. And what they need to do is say, "I have changed a little bit in my world, and what's the impact of that on the rest of the system?" And pull that, make sure they can see that as fast as possible, and make it judge the impact across all their tools in the value stream and sort of pull the pain forward in front of them.

And the idea here is that it's test automatically. And to test anything in data and analytics, you are testing data. And from our way of thinking, the best way to do that is to reuse those tests that you've done in production, perhaps add some more, and treat them as regression or functional impact analysis tests. And so why do all this work?

00:35:00

Well, less regression errors, the more problems you find in development, the less rework you have to do. It means that you can deploy faster. It means that you can have a railroad with a good signaling infrastructure. And if you've removed those individual blinders, you've also removed the need for meetings and steps and time.

And lowering the risk of change means that you can actually satisfy your customers better, and you get less eye rolls. And so this sort of velocity of change, I think, is really important, but velocity with controls and automation to make it happen. So the third principle on this is really about measurement. And so how do you measure, as a person who runs a data and analytic team, how do you show that your team's awesome?

Right? Because you're only as good as your last insight. And then, God forbid, if you've had bad data, that's going to follow you around. And I don't know if you've had the experience. I've had the experience of people telling me, "I can't go into the cafeteria today because we had a data error and everyone's going to look at me and give me dirty looks." And that's not what you want to be as a team leader.

So you want to know if there's errors in production. You want to know if you're making your customers a success. And I think there's a way that you can build data across the work that your data team does. And here's an example of a metric that we do, and I'm going to explain a little bit of what this is.

And start in the middle here, because there's a time on one axis in the middle, and this is actually the number of automated tests that are running both in development and production. And one of the things that people in general push back on DataOps is we're saying a very simple thing. You can change production fast with low errors. So you can do two things simultaneously. And if anyone's been in the field for a while, that sends chills up and down their spine.

You're saying you can go fast and not break things? And most of the time, they're saying, "Choose one." You can either run things and not break them or change them, and it'll break. And we're saying the way to do that here is through automation. If you increase the number of automated tasks, you actually can increase the deployment frequency. You can increase the amount of the times that you're late. That's another metric here in the corner.

And then there's associated sort of velocity of work that's being done. And so these metrics can help both prove that your team is awesome, but also help your team move to a new kind of world. And then second, just being able to look at what's happening with your production, your factory. You've got a lot of things building. How often are they going?

Is it coming? And your data providers, are they good? Are they bad? Getting some metrics to drive them. And so what we've seen is when these kind of things happen, the sort of DataOps idea happens, these metrics happen, you're actually able to have an incredibly productive team because they're spending much more time on the work that they're doing. So by building a lot of automated tests, by automating deployment, they can actually make hundreds of changes per week with an incredibly low error rate in a very small team.

And so this actually means that the productivity enhancement is on the order of 10X because you've lowered your collaboration time. And why does this happen? And I think what it really means is that there's a set of tasks that we have to do that we're going to put under the bucket of DataOps engineering.

And I'm going to talk about that and what that means, right? Because somebody has to own the assembly line in your organization. And maybe it's a data engineer who's inspired to do it. Maybe it's a separate role. I don't want to get into whether it is or not, but someone's got to run the factory, right?

Someone's got to enable the people who do the work to do it better. Enable your data scientists. So the customer of a DataOps engineer is the teams who are doing the actual insight generation, data scientists, data engineer people are doing this, and they own the pipeline or the factory. And what DataOps engineers, they work on the pipeline.

They don't work at a workstation in the pipeline. They own the factory itself. And so the role of the DataOps engineer has highly collaborative across all these different types of people, data engineers, data scientists, data governance, and it's really a very collaborative role. And the challenge is that when you've got this collaboration and you don't focus on DataOps engineering, is that a lot of people are tempted to have these sort of bad behaviors, where I do some work as a data engineer, and then I just throw it over to production, and weeks later, it comes back with a problem.

And what does it mean to be done? What's the definition of done for my team?

00:40:00

And being able to focus on just tasks and not systems, hoping that things work.

Kind of willful blindness to the errors that are going on, focusing, "It's someone else's problem, not my problem. I only focus on my little part." And so what the DataOps engineer does is sort of collaboration through this idea of a shared abstraction, the idea of the factory. And one of the ways that we think of it is take the nuggets of work that they do.

They've built a new model. They've changed some ETL code. There's a new visualization, new data prep, new governance code, new security code. The process is to put it into pipelines, test it, running the factory, automating deploys, working across people, measuring success, enabling self-service. The idea of DataOps engineering is to take a lot of these things that we think are invisible and make them visible. Take those nuggets and put them into a factory and use that factory to kind of, that shared abstraction to drive collaboration. And that's really about automation, about automating things and intentionally automating things.

And one of the challenges with automation is it's kind of seen as not important. It doesn't have an owner. People aren't spending enough time on it. No one cares. It's always someone else's problem. There's a perception that automation is sort of for lesser beings and that-- Or data's different. We don't have to automate, and manual work's always the way we've done it.

And DataOps is an automation discipline. And so the kind of tasks that we do, we think is automation is really important. Automated testing, automated deployment, automation running your factory. And these tasks about production orchestration, production data monitoring and testing, building and self-abstracting the development and test environments, regression testing, test data automation, deployment automation, shared components, process measurement.

This stuff hasn't been important on our teams, honestly, and it's sort of not done very often. It's not seen as an important thing, and that's what we're trying to change. And whether you call it a DataOps engineer or not, think about the percentage of effort that you do. A lot of teams spend very little effort on those tasks, 3%, 1%.

And What we're saying is, if you look at software teams, high-performing software teams, 23% of the staff is on these automation tasks. They're not directly doing work. They're not working at the task at hand. They're working upstream. And in fact, the role of a DevOps engineer over the past 15 years has gone from being lesser paid than software engineers to being equal or higher paid, and because building a good factory enables teams to work together.

And so I'm just saying, don't go all the way to 23% with DevOps, just do 15. Try that. Do 10. Do more than 3%. Put some time into automation, and perhaps what that means, if you've got a 100-person team, that would mean a 15-person DataOps team, or 15% of everyone's time. And in a lot of ways, we think DataOps is everyone's responsibility, not just a certain team.

And so this idea of automation is worthy of investment, and that this worthy investment actually pays to solve the emotional pain, improves productivity, improves cycle times, and lowers error rates is the discussion on DataOps. And so I'm sort of running a bit out of time here, but I just want to talk a little bit about how organizations make this change. And it's a tough change for organizations because there's a lot of steps to go through.

And in smaller companies with a 10-person team, it has one set of challenges, but bigger companies that work across lots of teams, you take a big multi-lines of insurance or financial service companies, or they have brokerage and insurance and high net worth. There's a lot of different data teams across, and all these different data teams are working on different things.

And how does someone who's a chief data officer say, "This DataOps idea is right. How do I actually influence everyone in the organization to start making this change?" And whether it's that your team is 10 people or you've got 1,000 people to influence across an organization, there's a set of steps that we've identified on how you actually bring that change to your organization.

And so we actually wrote a book on how to help called "Recipes for Enterprise DataOps Transformation." But how do you actually kind of get people to make this? Because we've been at it quite a while, and we've seen all sorts of pushback on how to make this happen in the organization, because people are struggling, and it is hard, and they're frustrated.

And it's just another thing, and they've got a whole lot of tasks to do, and so why

00:45:00

am I going to do this DataOps when I've got a new model to get out and things are breaking left and right? Why should I focus upstream when what's in front of me is such a pain? And so that problem of how you actually make that sort of mental and emotional change, walking through these steps of educating, finding a project, establishing a community, demonstrating real value, enrolling more people, iterating on the rollout, and expanding is really about trying to bring these ideas to your organization in a very systematic way. And this is not the first time companies have done this kind of cultural transformation on the IT and software development side.

The DevOps transformation is still ongoing. And in some cases, companies, this idea of a DataOps transformation is part of or secondary thing that happens when the IT team has gone through a DevOps transformation, and then the data team, data analytics team go, "Well, I don't know how to do this." And so this is sort of a step.

And one way to start is actually, we have a benchmarking thing. It's on our website. You can actually go through and do a sort of a benchmark on where people are in terms of these sort of dimensions of maturity on DataOps. And it has to do with, as we said, the cycle time that you can work, how fast you can deploy, how many errors you have in production, do you measure, and sort of customer satisfaction and culture. And you can feel free to take it, and we've actually--

It's a great way to kind of understand where you are as a team, and if you have more than one person at the company do it, we can kind of help organize the data for you. But it's free, and you can try it. And so what are the next steps? So if we look at the way the world is today, most companies are spending weeks or months deploying new insights into production.

Their cycle time is slow. Most companies constantly have errors, and maybe they don't even know. But they see their customers don't trust the data or things are breaking left and right. Their teams are not very productive. 60% of their time is spent on things that aren't value add, or 80% if you count. And then they have no idea. They don't measure their process at all.

They don't know about their error rates or cycle times or productivity. And so what that means is the cost is very high. And so the data and analytic teams, we're hiring more people. They're not as productive as we'd like, and we have unhappy customers. And my concern is that when the market turns and we've gone through this great boom in data and analytics, it's been fantastic, that when eventually we fall into recession again and the CFO says, "Are we really getting value out of our data and analytic teams?" And we're not able to show it.

And so I think the idea here is that you think of these as a graphic equalizer. You can push up all of these simultaneously. You can go fast and not break things. You cannot live by having a lot of meetings. And you can take this graphic equalizer, push it all up, and if you do The actual productivity of your team increases and the costs go down, and your unhappy customers go down.

And this idea that you can simultaneously improve all these things, it isn't a trade-off between making a lot of changes and errors, having less meetings and more errors. That you can work on all these things together and see the benefit is

a really intriguing idea. And also it comes directly from software as well, and from manufacturing, that you can not just improve on one metric, but on multiple metrics. And so how do you do that? So let me just conclude here and get to questions. So, here's the advertisement. We have software that helps you do this, right?

That is specifically built to help you focus on decreasing cycle time and lowering error rates, increasing collaboration and measurement. And so we've got a website. Please check out our software. The second is really an idea, a set of ideas. And so there's ways that you can learn more about this DataOps idea. So if you watch this video, let's look at the slides, it's been 45 minutes. So if you want something quick to send to your customers, we have an 18-point manifesto, one page, that you can sign and you can believe.

We also have 160-page book. We have a hundred and some odd page DataOps transformation book, and what's not on here is we actually have a three-hour DataOps certification program, and you can find that on our website. And I've been surprised we let this thing out in the wild in

00:50:00

December, and we've already had, I don't know, 1,200 people sign up and some percentage of those go through it. So there's a lot of ways for you to learn about the concepts of DataOps, what it is, and the book is relatively easy chapters, lots of pictures. Both books are. I've gotten comments that they're easy to scan.

And so, there's lots of places for you to learn more about DataOps, from manifestos to books to even a certification program where you can watch more videos. So I'm going to stop there, and try to answer some of the questions. So I guess a two-one is people would like a copy of the video and the deck, and again, they will be posted.

And so let me talk about, there's another person named Chris who's asking questions. So the first question is, "Is DevOps a prerequisite for DataOps?" And so it depends on what you mean by the term DevOps, and that's one of the problems with DataOps. If you Google it, there's all sorts of people who are using it in different ways, and then there's shade terms of Model Ops or Analytic Ops or MLOps.

Sometimes there's the term AI Ops. And so I'm not particular believer in the terms that matters because I think it's the concepts that matter. And both DevOps and DataOps conceptually mean focusing on automating. And I think DevOps more focuses on automation of the IT infrastructure, being able to do infrastructure as code. It also focuses more on the software teams.

And so people who are building HTML database back websites, whereas the idea of DataOps also has infrastructure as code, but focuses more on data and analytic teams and the unique problems that they have. But basically they all fall into the same sort of conceptual category. So is DevOps a prerequisite for DataOps? I think they're both part of the operational problem and the automation problem that people have to do. And then the second question that Chris had is, "How do you tackle ownership, data ownership?" And so who owns what? And I think that's a really interesting question, right, about how you own data, how you own changes to data. And so, I'm more a fan of

also thinking about who owns the processes that act on data, the code, the configuration, the pipelines, in addition to the data itself. And I'm also a fan of the concept of a data mesh, where there's a small team that really gets to know the data that they're working with, and can iterate quickly and this idea of a domain that they work in, having multiple domains in an organization, and I'm a big believer that small teams working very closely with their customers, and automating a lot can get a huge amount done.

And so to me, I think it's who owns the data is less important than who's trying to make the customer happy and who owns the things acting upon data. And the sort of data mesh idea, data domain-driven design, kind of decomposing your teams into smaller groups and working on the relationship, I think is important.

And so,

a question is, "Where do the data management engineers fit into this in your eye?" And so the question is like the name of who does data has changed. I used to call them ETL engineers, and then we used to call them people who do data management. We sometimes just call them data engineers. Now we have data engineers. There's a term analytic engineers.

And to me, I think, in general, the way the market's talking about it is data engineers tend to be more back end, and analytic engineers tend to work more closely with the business, but somebody is transforming data in some form using a visual tool or ETL tool or Python code. And so I think to me, they're part of the value chain and part of the factory. And honestly, from the way the tools work, there's a bunch of high-code tools where people are writing SQL and doing Python. There's a bunch of low-code tools with nice UIs, but they're all acting upon data, and they're all part of the factory that's delivering value.

And so, one last question is: how can I start with DataOps? And I've given some pointers on education. I think that there's one way to start, I think that I've

seen is really start small.

00:55:00

And the first thing that I did sort of in 2006 was put together a quality circle and just have a spreadsheet of all the errors that we had on, and just kind of take the shame out of errors in your production process, put them in a spreadsheet, look at the tickets, and then once every three, four weeks, get a team of people to say, "Is there a common problem across all these errors?

And can we write a script to automate them away?" And take the shame and hiding of the errors and run towards them. And I think that can be just a first way of the cultural change of looking at your error square in the face. I think that's one small step that I-- And also just looking at your errors and trying to find out where they are.

Just getting a list is actually a good thing. So I think that could be a very good first step.

And so have you seen an example of successful DataOps teams that wear multiple hats? A data scientist that handles the role of data engineer, a data scientist that does data and analytics, and yeah, I think there's a lot of cases where people are wearing multiple hats. And so, everyone sort of wants the magical person who is full stack, who can understand data, integrate data, do a data science model, visualize it, and then have a good talk with the CEO about it. Those people are so rare and get so bored that it's hard to keep them, honestly.

And I love full-stack, amazing people, but they also create a lot of chaos in their wake. They build things really fast, really quickly. They show business value, and then they don't want to run it. They don't want to automate it. They don't know, and they hand it off to someone, and no one knows how it works. And so I'm not saying that innovation isn't there, I'm just saying that when you build something new, the process of putting it in the factory and automating it is also a very technically complicated fact.

And so someone that sort of-- I call that a DataOps role, and sometimes the DataOps role is fulfilled by people who are called data engineers or data scientists, sometimes it's not. But a lot of times, organizations are rushing towards insight, and then they get the insight, and the customer loves it, and then the operationalization, the industrialization, the productization is a second thought.

And I'm saying that when there are cases where you're going to do things quickly and throw them away, but also, it's not that much more work to think about the operational part of this from the get-go. And let's see. I think that's it. So we've gone through 60 slides in an hour. I appreciate everyone kind of sticking with us here.

And as I said, what we'll do here is set up the PowerPoint and the actual copy of the deck, and we'll have a webpage. And then from that webpage, you'll also be able to get to our books and our certification program. So thank you very much for attending, and I hope that you all found this valuable. So have a great day. Thank you very much.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What is DataOps?

DataOps is the set of technical practices, cultural norms, and architecture that enable rapid cycles of experimentation and delivery of new insight, low error rates, collaboration across complex sets of people, technology, and environments, and clear measurement of results. Gartner's Sumit Pal wrote in November 2018 that organizations adopting a DevOps and DataOps approach are more successful at implementing end-to-end, reliable, scalable, repeatable solutions.

How much of a data team's time goes to errors instead of new work?

Gartner put it at 78 percent in 2020, leaving 22 percent for innovation. The 2021 DataKitchen and data.world survey found 52 percent of data engineers named errors as a major source of burnout, and 78 percent said they were stressed enough to want a therapist.

What does a DataOps engineer do?

A DataOps engineer owns the assembly line rather than the product moving along it, working on the pipelines rather than in them. The job is to take nuggets of code from data engineers, scientists, analysts, and governance staff and put them into pipelines, create tests, run the factory, automate deploys, and measure success. The goal is making invisible process visible.

What should a data team automate first?

Eight things, in the order a team usually needs them: production orchestration, production data monitoring and testing, self-service environments, development regression and functional tests, test data, deployment, shared components, and process measurement. The common failure is not technical. Nobody owns automation, so everyone stays on the task in front of them.

What tests belong in a production data pipeline?

Five types: traditional data quality checks, statistical process control, location balance tests, historic balance tests, and business-based tests. They should run automatically in production, across the whole toolchain rather than one tool, send alerts, keep history, and be easy enough to create that the team keeps adding them.

How do you get a large organization to adopt DataOps?

Six steps. Educate on the ideas. Find a first project by talking with individual teams about where the pain is. Establish a community of interest and a shared resource center. Demonstrate real value in a project of a month or two. Iterate onto more use cases. Expand by staffing a full-time center of excellence with common tools and metrics.

Where to go next