On-Demand Webinar · 57 min
Why Do DataOps?
Chris Bergh on what DataOps actually is and how a data team uses it to cut development time, reduce errors to near zero, and improve collaboration within and across teams. Recorded March 2020; updated August 2026.
What you'll learn 6 points
- The failure numbers are the argument: 87 percent of data science projects never reach production, 60 percent of all data analytic projects fail, and 79 percent of data projects have too many errors, while the share of organizations calling themselves data-driven fell from 37 percent to 31 percent.
- The diagnosis is that this is a people and process failure, not a technology one. Data analytics in 2020 is compared to the US auto industry in the 1970s: high production errors and deployment latency measured in weeks and months.
- Data teams sit between three competing forces — unaware data providers sending late and error-prone data, demanding consumers expecting Amazon-speed delivery, and critical production operators needing flawless output. The result is a work environment where teams cannot innovate and cultures run on heroism or fear.
- Conway's Law applies to data: the structure of how teams are organized for engineering, science, analytics, and production shows up directly in the shape of the data pipelines, and a common platform is what lets a team escape that constraint.
- The DevOps precedent is quantified. High-performing IT organizations deploy 200 times more frequently, recover 24 times faster, have three times lower change failure rates, and spend 22 percent less time on unplanned work and rework.
- The prescription is seven steps plus three: orchestrate two journeys, add tests and monitoring, use version control, branch and merge, use multiple environments, reuse and containerize, parameterize processing, then add architecture, metrics, and inter- and intra-team collaboration.
Prefer to read it? The written version is in Why Do DataOps.
Slides
Transcript
Show chapters and dialogue 10,000 words
00:00:00
Hello, everyone. My name is Chris Bergh. I'm CEO of a company called DataKitchen, and today's webinar topic is something called "Why do DataOps?" And just as a little bit of housekeeping, please keep your questions and put them in the chat window. We'll gather them up at the end and leave about 10 minutes left at the end to talk through those questions. So we have a number of topics today in the webinar. And the first thing is we're going to talk kind of at the 10,000-foot level, like why does DataOps exist or why is this a thing at all? And it's really a more of a people and process thing than a tech thing. And so we're going to focus on the people and process part.
And then we're going to dive into sort of why DataOps fixes this problem and how DataKitchen helps, and then we'll have questions and answers at the end. So,
during these trying times of COVID-19 and everyone working at home, our hearts go out to everyone who's being impacted either directly or indirectly by the current situation. So that being said, I'd like to kind of go into why DataOps exists as a thing. And I think, first of all, you'd have to be living under a rock to not know that data analytics is a hot thing. You can't walk through an airport or watch a TV ad during a football game and not hear something about data and analytics and machine learning and AI.
And there's just a lot of buzzwords out there, originating from big data or data science or data lakes or MLAI, and there's just a lot of some, I think, warranted hype about it. Yet, if you look at every one of the projects that people do, most of them actually fail, and they fail in different ways.
So for instance, 87% of models don't get into production. Companies calling themselves data-driven are actually going down. 60% of all data and analytics projects fail. And so why is that, given that it's such a cool area with so much buzz and the number of master's degrees in data science or machine learning have gone up, why is it that the success in this field is not so great?
And so I think, if you look at it, we did a survey with Eckerson group, who is a Boston area data and analytics consultancy, and we asked a couple of questions. The first question was, how many errors do you have? And we categorize errors, and you'll hear about it in this presentation in a very broad way. It could be that your data quality that you get from your source systems is poor, it could be something done when you get the data makes it wrong, could be that you're just late.
And most companies just have way too many errors per month. And then we ask another question, how fast can you get something from the brain of your data scientist or your data engineer into the hands of your customer? And most companies are just way too slow to be able to make that happen. And so if we look at, at a high level, what the data and analytics industry is like, I grew up in Wisconsin in the '70s and '80s, and my dad was a telephone repairman, and he drove a Toyota Corolla, which wasn't a great thing to do in the '70s and '80s in Wisconsin because Milwaukee had lost about 50,000 jobs mainly to the time of Japanese imports.
And my dad loved the Toyota because he said it was cheaper, it lasted longer, and it was a better car. And so I think in a lot of ways, the data and analytics industry is producing sort of Pacers. They're cars that they build the assembly line and they can't change the assembly line for years, and they're full of errors, and it just takes a while to change or improve it. And so what we see in a lot of organizations, taking this analogy to a model or an ETL process or a new data set or a new visualization, is that it takes weeks or months to deploy something into production.
And when it's in production, it has a high error rate, and the teams themselves are actually, I think, suffering quite a bit. And if you look at it in a different way, looking at the percentage of time your team spends every week,
we think that there's just a lot of challenges. First of all, there's just a lot of errors in operational tasks that the teams are working on. Instead of doing what we want in the green on top, they're actually doing things that are just not time well spent. And I think that's really because of complex organizations and complex tools and complex data and complex collaboration.
There's just a lot of complexity in data and analytics. And I think as a result, data teams are suffering. And so when I started, I come from a software engineering and a data analytics background, and when I started my career, I was a software engineer, then a software manager, and then about 2005, I started
00:05:00
to work in data science, data engineering, data analytics. And running teams, I was sort of caught between these three forces. One is that the people who provided data kind of didn't care that my team existed. They would often just sort of give me crappy data, and it would break. And then on the other side, the consumers of data just wanted data that was of extremely high quality, delivered on time with low errors so they could trust it. And a lot of times delivering something into production, the production team wanted things that were flawless.
And so, how do you go fast given that your data providers don't give you? How do you innovate and give them trusted data sets for your data consumers? And I think, in fact, my experience was for many years, I always sort of felt beaten down and distraught and disempowered. And part of the genesis of DataOps is, I think, trying to make for a team that can actually not have these problems, that can not create and innovate, that can not live in both one end of just being a hero and working nights and weekends or burying themselves in paper process so that their fear causes them not to want to change anything.
And so in a lot of ways, if we look at DataOps, it's kind of an empowerment idea. That instead of being sort of crushed by these forces, you can take these set of ideas and apply them to your dissatisfied data consumers and your unaware data providers and your unstable production operations. And it gives you a set of ideas and methodologies to empower your team to sort of retake the initiative. And so it's more about taking back control and empowering your team, given that we all sort of live in this tough environment with all these forces against us.
And so, let me start off first talking about why does this problem exist, and we're going to talk a little bit about the people, then the process, and then go back to the people. And so first, there's just a lot of people who work in data. Some of them are called data analysts who work with tools like Tableau.
Some of them are data engineers who work with Informatica. Some of them are data architects or database administrators. There's a whole chain and type of different people who work in different roles, and they've all got great different skills. And if you look at it, let's just take a look at how they work in a team.
So first of all, here's an example of a team, and there's a data engineer and a data science team who work together, and then there's a self-service team and a data governance team. And if you look at how they work, they all have their different tools. There's a sort of $100 billion industry in data and analytic tools. And it's just massively fragmented.
So there's different tools to actually visualize data. Some of which, like Tableau or Qlik or Cognos are very large, and there's a bunch of niche vendors out there and open-source vendors. There's tools to do data science like Python, and there's data science platforms like SAS. There's data catalog tools and governance tools like Alation.
There's ETL tools like Informatica and Talend or just writing SQL. There's a whole variety of databases like SQL and Redshift and Postgres, and there's all sorts of other sort of tools out there to build data lakes. And so there's a massive set of tools. And if you're listening to this, you probably are coming from a larger company, and most larger companies probably have two of each already.
And from my standpoint, I think people love their tools, and I never want to get into a R versus Python or Qlik versus Tableau or writing SQL versus a visual UI discussion in my life because people love their tools. And to a certain extent, companies want to standardize on tools, but how people express their ideas into tools, I think they should have a little bit of freedom with, in my opinion. And so let's look at that.
So let's look at a scenario of how these teams work together. So there's these two data engineers, and they source some data, and they create a really simple, silly database table with a name and some sales in. So they load the data from the source, put it in a database table. And then the data science team They have to create a model, so they actually take that data and they cluster it, and they
create two segments, a high-value segment and a low-value segment, and it's a batch model. And then there's a self-service team, or you could sometimes call them a business intelligence team or a data analytics team, who take that. And then they've got their own favorite tools, and so they're going to visualize it in a tool like Tableau. But they're also going to use what's called a data prep tool, and that's where they can add or mix in some more data into it, and they can add an owner. The west team and east team. This is very common.
They may be embedded in a line of business and who owns it may be subject to a line of business, and that data may not be stored in a data warehouse. And so the business user actually sees the summary of all three of these things. They see a chart that has those four columns in.
00:10:00
And then, there comes the data governance team who just wants to know where the data comes from and what it means and what its description is. And so all these people have to work in teams together to get their job done. And so they also source data from a lot of internal and external systems.
So,
in some ways, every database, website, internal IT tool, is a source of data that you can use to understand your business, understand what your customers do, understand your products and your operations. And there's a lot of, they call them CRM or ERP, your supply chain, your website, your finance, your HR. There's data that can come in from inside your company.
There's data that can come outside your company, like census data or open data. It comes in a lot of different forms. And, one of the things that we believe here is that as that data comes into your company, and it's invariant of whether it's big or small or whether it's batch or real-time or whether it's structured or unstructured, it goes through a series of steps, kind of like a manufacturing line here, where it's usually put in some space, a place. It's transformed.
It may or may not have a predictive model, may or may not be visualized, or may or may not be cataloged. And all these steps happen. And from our perspective, it's really a factory, a factory of insight, and each one of these steps in the factory is kind of a manufacturing line. And so this is an important metaphor that we're going to talk about later in the discussion, in that this factory view of analytic production we think is important because every system that's in production has this sort of, think of as an assembly line, or we're going to call it a pipeline later.
And some companies have one pipeline, but a lot of companies have hundreds or thousands of these pipelines running in various parts of the organization. And how do you know that these pipelines are producing? How do you give them a virtual hand on cord to stop the assembly line? And so there's a whole set of metaphors on that we're going to bring up that have been learned in the Toyota production system, lean manufacturing, statistical process control. These ideas, I think, actually are very important in order for us to understand how to run our own factories that are producing data given all our tools.
And then the second part of this is that, if you look at whether it's a data engineer or scientist and whether they're sitting down at R or Python or Informatica, they're creating. And the process that they create things, they innovate, goes from the idea in their head into a development environment into production. And so how does that happen?
How do you deploy quickly from, how do you get the ideas from your individual contributor's head into production? And a basic idea of DataOps is that if you can get feedback quicker, if you can work in a short cycle time of delivery, you can learn more about what your customers want because there's a set of humbleness about we don't really know what our customers want. And then you can maximize the amount of work that you don't have to do because you're constantly learning what your customer wants.
And you may think they want A through F, but giving them feedback, they quickly could learn they only need A through C. And so you maximize the work and you don't have to do. And so both of these things actually have to happen at the same time. You've got to run a factory like Toyota that produces really high-quality cars, with a very low error rate, and understand it with tools like statistical process control.
But you also have to be able to pick up a piece of that factory and change it, like in software. You've got to be able to do continuous deployment or another term is continuous integration. You've got to pick up a piece of the assembly line, give it to an individual contributor, have them be able to tweak it or add to it.
And maybe it's not one piece, it's two piece or all of the assembly line, and then tweak it and push it back into production. And so these things are actually pretty opposite. One is about fear. I don't want to break production. And the other is about change. And so fear and change are kind of opposites, and you've got to reconcile these opposites in order to do data analytics well.
And so let's look at the second part. Why does the people problem exist? And so if you look at this hypothetical organization of data engineers and data scientists and people in self-service or data governance teams, they may work for the same boss, like a chief data officer or a chief analytics officer. Maybe sometimes they have a different name, a VP of data or something like that.
But they could work for the same boss. But sometimes they don't. You may have your data engineer or data warehouse team or data enablement team working for the CIO. The self-service team may be embedded in various lines of business across the company. The data science team may be new and working for the CEO.
And the organizational structure of how these teams sort of lay into companies, is
00:15:00
quite varied. In general, there's not a set pattern yet. I think the number of chief data officers has rapidly increased over the past five years, and that the former organizational pattern is winning out, but there's plenty of companies who are organized like this. And so there's a different type of organization, too, where, if you look at the people who create things, the data engineers and data scientists, and let's call them developers, and then the people who run things once they've created it, and call them ops or operations. In a software development team, there's kind of a development team and then an operations team, sort of a one-to-one relationship.
But in data and analytics, because of this organizational differences, there's actually a many-to-many relationship between dev and ops. There's different ops teams supporting different groups. And that actually goes to an example that I want to show you here. And so the first is that you may have a home office team. So your company may be organized with a home office team, and in this picture, there's SQL Server and there's some Python in there. The home office team works for the CIO.
It's got data engineers and data scientists, and they live in Boston and maybe they update their data warehouse once a week, which is pretty good. But then you may have local office teams around the country. They may be working with Alteryx, which is a data prep tool, and Tableau. In this example, it's in New Jersey.
They're able to respond very quickly to a request from a VP of marketing. And so you've got these two teams and two tools and two sets of deployment operations going on. And so that can actually create a lot of problems. So for instance, just if you're looking at it from the perspective of a home office team, if I've got a schema in production and I want to change it, how do I know if I'm not going to break anything? There could be all these reports.
I don't know how many reports are using the system. Or if I'm in the local office, I can use a data prep tool and add in a new dataset. Do I have to manage it? How do I make that available for everyone? And a lot of cases, one of the symptoms of a problematic data and analytics organization is that there's not a trust in data.
And sometimes the trust in data, it fails because your data providers give you crappy data, which we're going to talk about. But it also could be because you've got reports that have the same idea, but expressed in different calculations. And a lot of times people will copy and paste these reports and put them in different locations. And so, these challenges with coordination are quite important. And so actually, if you look at this, and this is a bit of an eye chart, but I want to take a second to walk through it.
And so if you go to the left-hand side and look at that D in the middle, and let's suppose that's a development team, and they've got BI people and data science people and data engineering people working together, and maybe they work for the CIO. But then if you go to the second column and look at the development team and the dashed line to that P team, the production team.
Well, the development team's got to push to production. Maybe they've got a dev database and a QA database in production, but they've got some way that they're moving things into production. And so that's a challenge of coordination. But then you go to the third column and you look at this decentralized development, and that's the case where people are using Tableau and Alteryx and all these great self-service tools. And they may be completely decentralized.
And then going to the fourth case, where those tools, like Tableau, they may be pushing to something like production just by pressing a button in Tableau, saying you're pushed to Tableau Online. And so you've got this sort of network of cornucopia of complexity between different development teams and different locations, different production. And so that creates some challenges.
And one way to look at these challenges is to actually look at them in a different perspective. And there's something called Conway's Law, which states that any organization that designs a system will inevitably produce a design of that system that is reflecting of the company's communication structure. And so,
when I was a software engineer, if you brought up Conway's Law, most people would say, "Oh, that's designed poorly." If it's designed not the way the technology wants, but it's designed the way your managers have sat in some meeting and put the org structure, that means it's inferior. And I think that's to a certain extent true in data and analytics that DataOps follows Conway's Law. And let me give you an example.
So if you look at a typical team, and let's say there's a production team and a data engineering and a data science team. And one way to look at it is that all the work that everyone does in data and analytics, you could express it as a pipeline. And here I'm showing these pipelines, and maybe the pipeline does some data
00:20:00
engineering, does some data modeling, some data flows through that pipeline. There's different tools. But each data engineering team and data science team is kind of pushing to a production team independently, their own separate pipelines. But, and if you look at how they develop, where there's data engineers and they work together on different projects and they've got to roll their sub-projects up into major projects, you end up with a more complicated view, right?
Where the data engineering team has got two dev tasks that there's five different people working on, and some people are working on both, and they have their own individual tasks. And each one of these have these pipelines, and they're working on these sub-pipelines that are going to be rolled up into a task pipeline, that's rolled up into a QA pipeline and a production pipeline.
And the data science team has got their own organizational structure. And so what happens is that this complexity of pipelines is the complexities of the organization and their development process and their team boxes is really reflected in their pipelines. And that creates a lot of challenges. And so one way to think about this is that that organizational complexity ends up reflecting your pipeline complexity.
And if you've got that pipeline complexity, how can you solve both of those together? And so we believe that a DataOps platform or the ideas of DataOps can help you. So, I'm going to keep going here, and we're going to ask questions at the end. So,
how does the idea of DataOps fix this problem? And so, first of all, the term DataOps, it's been around for a few years. Actually, there's a company in Arizona who owns the term DataOps.com, and it's two women who've been working in data since, I don't know, the '80s or '90s. And so the term's been out there. But just the past three or four years, it started to gain traction. I think Gartner put it in its Hype Cycle starting in July of 2019, to actually look at Google Trends over the past two years as a proxy of market growth. And, you can kind of see, if you look at the sort of 12-week moving average, there's sort of a 500% growth in search volume over time, which is great.
So that's one indication that it's good. There's a lot more media articles, a lot of companies branding themselves as DataOps. My company got to be to have a cool vendor. There's just a lot more buzz around the term DataOps, and as a result, there's also just a lot more confusion about what it is and why you should do it.
And like most tech terms, they sort of inflate and big data became everything in data and analytics, and unfortunately, I think due to DataOps' popularity, probably will get to that point. However, we've strived to kind of have a very clear definition of why do DataOps and what DataOps is. And so, that's what we're going to talk about really here, what is DataOps?
And so I think if you look at mostly all definitions of DataOps, say it's sort of like doing what those DevOps people do in software, those agile DevOps people do, and apply it to data science and analytics, like in general. And one of the interesting things, having bridged both worlds, is that a lot of people who spend their career building software, and I spoke at the DevOps Enterprise Summit about DataOps this last year, don't know that much about data science or data engineering. They sort of know the words and principles, but they don't really know or think about it. It's sort of that other group in another room.
And then likewise, there's all of us who do data and analytics day in and day out. We sort of have a glancing idea of DevOps and what software people do. It's as if we graduated from college and took one door left and one door right, and you either took the software door or an IT door, or you took the data door.
And we all have the same set of talents. But what's happened in the software industry, and starting kind of around 2009, 2010, is this idea that you need a technical platform to be agile, that you need to-- There's a set of ways to help you release your software in an automated way with fast time. And so, what that means is that teams now, and in the '90s, I was a software development manager.
I was a director of engineering. I could ship software every three months, and that was actually pretty good in 1999. Now, if I were to go get a job 20 years later saying I could ship software every three months, no one would hire me because the expectation is I should be able to do it in three days or three hours or three seconds. And so the entire industry's perspective of how fast you should get software out the door has changed.
And that idea is really bounded in the same set of ideas that drive DataOps, is that cycle
00:25:00
time, getting ideas into customers' hands, automating it, and handling the complexity of systems and coordination through technology produces a better outcome. And in almost every way, people can release faster, the quality of softwares goes up, and the teams are spending more time focused on what matters. And then, if you actually look at lean and the idea of building cars, lean manufacturing or the Toyota production system, it does improve efficiency, reduces waste, increases productivity, and makes cars that are better.
And so, there's some books, a guy, W. Edwards Deming, and "The Machine That Changed the World" I think are really interesting in this regard, and also the ideas of statistical process control. And so if you take these ideas together, agile, which is a way of organizing software teams, the Agile Manifesto was written in 2001, and then DevOps, which is sort of a technical way to do software development. And then lean manufacturing, you apply those ideas or principles to the more complicated world of data and analytics, you get DataOps.
So if we think of it from a definitional standpoint, so DataOps is a set of technical practices and cultural norms and architecture patterns that enable really four things: fast cycle times, low error rates, high collaboration, and measurement. And so what does cycle time mean? It means the ability to rapidly experiment and innovate. So you can get your hands on a keyboard and create a model and deploy it, change a database and deploy it, change something in your data and analytics value chain and get it in the hands of your customer to get feedback quicker.
And what that really means is that you're maximizing the amount of learning your team has, because cycle time and feedback forces your teams to learn. But you've got to be able to do that in a way that has low errors. And in this discussion, I use errors in a very broad sense. I don't mean just data quality, because if you have poor data quality, that could result in an error. But you could also have perfect data quality, and you could have an error, meaning something happened in your processing or you were just late.
And then, we've already talked about the idea of collaboration across the sort of complex set of the tools that people have, who they work for, the environments that they work in. And then finally, how do you measure? And one of the things that I think the success in DevOps and the success in manufacturing is they see the systems they work on as a source of data, and then they can use that systems that their teams work on and understand the data that that system throws off, and then be able to analyze it for improvement.
And so the systems that we work on are these pipelines. And, I think they actually produce a whole bunch of interesting data that can help us improve as teams.
So, why do companies implement DataOps? And so if we forget the right-hand column and look at the left-hand column, they come at it in different ways. So like you, they start off hearing this term DataOps and start wondering what the heck it is. They do some reading, and then they say, "Well, what..." They'll put their business hat on and say, "Well, what can I do? What difference does it make?
Sounds cool." Well, the first part is they're going to try to focus-- One, companies focus on errors. They're getting yelled at by their business customer because the data's wrong, or something broke, or they're late, and they feel beaten up because their VPs are calling up. And I had this very similar experience in my life of having VPs call me up and yell at me because the data's late, and I'm an introvert, and I don't particularly like to be yelled at. And so, one way, too, a lot of organizations have taken that fear of errors and instantiated it in this sort of Word document, meetings, handoff, technical review board processes.
And that is a way to review the low errors. It has the result of slowing down deployment of new things. So instead of being able to change your data warehouse every week, you start to have data warehouse changes and new data sets come every three months or six months. And so that's another way that people come at DataOps, is we just want to change things faster and our warehouse, our models aren't getting to production, or our warehouse takes four months to deploy 20 lines of SQL.
Other organizations come at it from the sort of centralization versus freedom. They have a lot of self-service teams out there, and perhaps the IT team is worried that they're going to do stuff that's out of compliance. Or there's poor collaboration between their development teams and their analytic production teams. And, finally, I think some leaders want to be more data-driven about the work that their teams do, and can you measure that?
And so these sort of business effects around lowering error rates, increasing the deployment velocity and cycle time, about resolving these collaboration and
00:30:00
coordination problems, and just management measurement are ways that people come at doing DataOps. And you notice what's not here. There's nothing about buy a new database, buy a new ETL tool, use the latest, do this type of predictive model versus that type of predictive model. This is actually nothing to do with databases or models or visualization tools.
You can implement DataOps on a 15-year-old SQL Server. You can implement DataOps on the latest and greatest cloud-based streaming event-driven architecture. It's really invariant. This is about a process in which you develop analytics, and the substrate, the substance that you develop analytics on is up to you. And there's reasons, and I'm a technologist.
I love what kind of model you have or what kind of database or how you use. All those things are really cool, but it's not about that. And fundamentally, I think you as a leader, you own the stuff, you own the process in which your team works. And the technology itself will change. So how do you do DataOps? And so if we think about this,
we have another presentation where we break it down into seven steps plus three, and we've hinted at some of them here. So the first step is what we call orchestrate two journeys, and in there it's about doing, running a good factory, being able to deploy. The second step is about adding tests and monitoring.
And so what enables all this stuff to work?
A good DataOps team should be able to take someone who just graduated from college, have them be able to change, tweak a model, tweak an ETL process, and deploy it to production within their first week and not have any problems. And that's done because you've built a system of tasks that happen in development, regression, functional performance tasks, as well as monitors that are in production that tell you if something goes wrong.
And we're a big believer in automated testing and being able to have that as a way to allow you to actually do DataOps. And there's a whole set of how do you test data? How do you test the code that acts on the data? How do you test your pipelines? What's a good test? What's a bad test? What are top types?
And we've got a whole series of discussions and blog posts about that. And then the other idea is that in DataOps, we're actually much more interested in the code that you write. I think DataOps is more interested in the Tableau workbook itself, the SQL code itself, the Python model, the Informatica job, all that stuff, and all that work that you're doing should be put in a version control system.
It should be put in one place, and then you should be able to branch and merge it because it's about the process that acts on data. And that process itself is code, and that's very much in the software engineer sense of code. And then we talk a little bit about using multiple environments, how to abstract dev and production, how to parameterize your processing and reuse.
And then we're going to talk a little bit about here in the next one about sort of architecture and metrics and some comparison to sort of DevOps and DataOps and other ones, just to whet your appetite. So the first one is, being a technical guy, I've seen so many different data architectures, and over time there's been different data architecture patterns.
For a while, it was put everything in Hadoop, or then Hadoop next to another database, and now there's sort of the S3, Redshift, Snowflake pattern. And there was sort of Kimball's data warehouse, and then there's different ETL tools. But they all sort of have this, whether you're talking about a 1995 architecture, a 2015 architecture, or 2020, they all have the same sort of patterns of I get some data in somewhere, I do some processing on it, I get some refined data, I do some data science and visualization, some governance.
And data architectures are very focused on what happens in production. They're designing the production system, but it doesn't reflect the operations, the ability to change, and that collaboration angle. And so we think that in DataOps, think of it as a right to repair. You should think of as a first-order item that you should be able to change your architecture and be able to focus on automatic deployment from dev and test to production, environment creation and management, orchestration and monitoring of testing and production. That all these features here are first-order things that you should think about when you're designing your data architecture.
And if you stick to just the production-only architecture, you're missing half the picture. And we're not the only one who agrees. Gartner just came out in their 2020 Planning Guide for Data Management, a very similar idea about being able to think through how you deploy from QA to dev
00:35:00
to prod and be able to think about your data architecture with that. And so, the second idea here is more of my experience as a manager, and there's the old managerial adage of you can't improve what you can't measure. And so I think there are things that we need to measure, and analytic teams, I think, are not very analytic about measuring what the work that their team does. We're trying to get our customers to be data-driven, but we're not sort of data-driven about the day-to-day work of our teams.
And I think the things that we need to measure are how productive are your team as a team or individuals? What is the error rates in production? Are we late? What are your data providers doing? Which data provider is giving you erroneous data sets and how often?
Am I meeting my SLAs? How fast can I deploy from dev to prod? How many release environments do I have? What's my coverage in terms of testing? And I think these sort of metrics can be derived from tools like my company's or other tools. But if you can start to measure things, you can start to make it visible, and then you can start to help people improve.
And one simple thing I did in 2006 at this job I talked about was I made a spreadsheet, and I wrote down every error we had every week. And then every three weeks, I'd sit with people and have a quality circle, and we'd look at every error we had in that past three weeks.
And then over time, we just started to look at that report of here's all our errors, and then we started to see patterns. This error's happened every four weeks for the last three months. And it doesn't have to be nice charts and graphs, just being able to put this in a place and reflect and look on it with your team, and then be able to make changes based on that. So you can start simple and move forward.
And then there's a lot of question of what is this ops stuff? I hear about DevOps and DataOps and ModelOps and AIOps and DevSecOps, and there's all these words. And actually, I think it's interesting because there are differences between DevOps and DataOps, and actually the most popular blog article, it's got over 40,000 views that they wrote, is a comparison between DevOps and DataOps.
So I think they are different. But I have a little bit more subtle view in that if we look at DevOps and DataOps, in fact, all these ops words that you hear, they're all based in, at the highest level, a kind of similar business management concept and approach to, I've got these people who are working on this technically complicated thing. It's an assembly line, it's software, it's a data analytic process.
And whether you call it a learning organization or the Deming principles,
lean, there's a bunch of ideas in business management there, but they're all based on, I've got a bunch of people working on a technically complicated thing. And in particular, there's ways that teams can manage themselves. There's Agile or Kanban or Scrum. There's Safe Agile. There's the Spotify method. There's ways that you can organize your teams.
And likewise, there's, in manufacturing, there's sort of Six Sigma and total quality management. There's certifications here. There's experts. And these are all really good ways. But if you kind of break it down into the organization that that applies to, I think more Six Sigma and total quality management applies to industrial manufacturing teams. You can apply Agile or Kanban or Scrum to software teams.
And then the technical environment that those teams work in, that enables them to be agile, has these words. Sometimes it's called DevOps or DevSecOps. It also has the term AIOps, which is applying algorithms to the streams of monitoring data. There's GitOps, which is making your technical environment centered in Git. And I think the similar idea here is instead of taking these business management concepts applied with an organizational management, whether you do Agile or Kanban or Scrum or Safe, whatever, and you apply that to data science, engineering, or analytics team, you end up with DataOps.
And I think, the idea of ModelOps or MLOps certainly applies because it's really doing what I think of DataOps for data science teams. There's other people who've written about governance ops, data governance ops. You could call data engineering ops. But for brevity, I just end up calling it DataOps. So in general, whatever term it is, they're all based on the same idea, that this is how you manage your teams.
And so, I hope that clears it up for you. So the last part, what does my company do and what does DataKitchen help? So, we've been at this for about six or seven years. My co-founders and I really started the company because we suffered that emotional problem that we had at the beginning. We were getting beat up by data providers who didn't care about us, by customers
00:40:00
who said we were too slow, and production problems left and right. And at the end, we just didn't believe that there was a new magic tool that would save it for us, and we tried. In a previous company, we tried to build one. And so what we've done is build a software product that solves these problems. It does these seven steps and these three steps, and we use a lot of food metaphors and kitchens and recipes and tests and ingredients, and we'd love to talk to you about it, because we think it's pretty cool.
But more so, the purpose of our company is in two ways. One, we want to make money by selling software, but also we want to promote these ideas. And we think that these ideas can be instantiated in our software, but you don't have to use our software to do it. And so, one of the reasons that we spend time in improving these discussions and writing a book and talking about it is I think these ideas are important and really need to go throughout the data and analytic industry.
And so finally, what does our software do? And we talked about the percentage of time people spend each week, and if you think about what happened, if you all remember the Iowa Caucus and the data problem with the Iowa Caucus. They had a system that recorded the votes. The recording of the votes was right, but then they were extracting the data from the system that recorded their votes to analyze it. They had a big ETL problem.
And when that happened, sort of the hair stood up on the back of my neck and I really felt sorry for the team, because I've had problems where my ETL process broke and I've had a VP yell at me. I've never had a problem where the entire country was yelling at me because I had a data error.
And so, the idea is using our software and applying DataOps principles, you can reduce the amount of errors, you can reduce the Iowa Caucus problems down, and then that gives you more room to actually develop more features. And it actually gives you more room to focus on process improvements or a term from software engineering, improving your technical debt or fixing things that you've already done. And likewise, what we've seen is that the deployment latency, how fast you can get something from dev to prod, instead of going from weeks to months, goes to hours or minutes. Your production errors goes down, and I think your team actually ends up being sort of happier and more productive.
And so, the last idea I want to share before is this perspective that DataOps is something else. This something else idea is really about the people and process, and by focusing on the people and the process and the operations, you gain a benefit. And so if you look at Tesla and Elon Musk, he talks a lot about the machine that makes the machine, and I think that's true.
We focus a lot on the model and the transformation, but we don't focus on the system that builds that. And I think that's an important perspective. Or if you look at Satya Nadella, who's the CEO of Microsoft, who's done a great job turning around the company, he has a quote that says, "If an engineer who's working on software has to choose between working on a feature or working on development productivity, he should choose productivity." He also has another quote that says he wants the smartest people in the organization not to work on features, but to work on the environment to build features faster. And if you look at the software engineering, in '99, I had a release engineer who worked for me, and he actually made less than all the other cool software engineers who were writing Java.
Now, that role is called a DevOps engineer, and they're actually paid more than software engineers because having a team of people focused on the productivity, the deployment efficiency of your software teams is so valuable that their starting salaries are higher. And in fact, if you look at Google, they have over 2,000 engineers who focus on engineering productivity.
So this, as a leader saying, we really need to invest in operations. It's not going to help us tomorrow, but it's going to help us build tomorrow much better. And so that ends my discussion. We've got about 12 minutes left, and I'm going to turn it over to Beth, and she has been hopefully grabbing all the questions and
putting them there. Yes. Thanks very much, Chris. That was a great overview, and I encourage anyone who has a question to please enter it in the question box, and we'll try to get through as many as we can. So I have a few that came in. Here's the first one, Chris, is, what is the best first step if you've decided to subscribe to DataOps principles?
Well, that's a really good question. And so I think what we try to do is say, is not think of DataOps as this all-encompassing thing. We think it's try to find a first step.
00:45:00
And a first step is to say, find a problem to solve. And it could be that problem on you're having too many errors in production. Well, just focus on that. The second is, well, it just takes me forever to deploy from dev to prod. Just focus on that. And if you can find a project or a team that's willing to focus on this one thing and improve it. So don't boil the ocean.
Find one part of the value with one team and then improve that, and then go from there. And so, actually our last webinar was about, that you can record, was about sort of three or four cases of our customers who started with us and what they focused on first and why they did and what kind of results they have.
And so DataOps may seem like a big thing, but you can break it down into smaller components and then start working on each. And there's also part of it, too, I think, in some organizations is the sharing of ideas and trying to get people sort of bought in that you can sort of reclaim control. You don't have to be sort of paralyzed with fear that someone's going to yell at you when the data's wrong or running around like a hero on the weekends, fixing things when things break, that there's a better way.
And we've written a book on it, we have talks on it, and it's helpful to sort of attack the emotional part of change as well as just finding good projects to do it.
Great. Thanks, Chris. Okay, here's a long one for you. Okay, so at times production teams can see what DataOps engineers and data scientists in dev to be constantly creating technical debt, especially when they are delivering iteratively to users to provide value and obtain feedback. How can these two teams complement each other and avoid the conflict or the perception of conflict mentioned above?
Yeah, and I think that is a challenge, right? Because what is DataOps? It's saying work iteratively. That means be able to get something from the hands of your data scientists or data engineers into production quickly. And so some technical things actually are better probably with a long cycle that I think about it, and that's true, but it doesn't mean that you have to go back to a waterfall method.
And so what I found is that if you can get some... We have a term in our company, get some value in the bank, bank a little bit of it, and then bank a little bit more. And perhaps, take some time in one of your periods not to deliver value, but to refactor or improve what you have.
And so it becomes a question of, what are you doing that week? And try to allocate with your business customer. I've got some of the stuff I'm doing new things for you. I've got some of the steps I'm trying to do to improve my process. I'm adding more tests. I'm upgrading the server. And then I've got another part where I'm trying to refactor or improve or document what I currently have. And so it's a balance with your customers week in and week out what you're doing. And if you can gain more trust to them by at least delivering some of the things that they want every week, what we found is they will work with you to say, "Okay, this week I'm going to refactor a bunch of code and improve it." Or, "This week I'm going to add a bunch more tests." Or, "I'm going to document what I previously didn't document." And so seeing it as, it's not an all or nothing, it's this period, this week or two weeks, we're going to do this. Maybe this week I'm doing all new features.
Maybe next week I'm doing a whole bunch of refactoring. And the balance between how you allocate resources every week and the dialogue you have with your customer is important, and that's all built on trust. And so, a lot of organizations end up having trust, distrust between their business customers and their development teams, and distrust between their development teams and their production teams. And so having everyone have a seat at the table, and the production team saying, "Look, you've given me crappy code the last three months. Can you find a way to make sure that you catch all these regressions?" Well, maybe that, a whole bunch of regression testing should go into your sprints for the next few weeks to fix that.
Okay, great. Here's another question. Can DataOps be driven in an organization from a top-down approach?
I think it helps a lot to have senior leadership sort of bought into the idea that of DevOps or DataOps or Agile as a way to work. It isn't a necessary requirement, but it makes it a lot easier. And so, some organizations have made the change over to DevOps and Agile principles a fundamental part of what their CIO's goals are.
And then the data and analytic teams are kind of left saying, "Well, how do I do this DevOps stuff with data?" And then they put Jenkins in, and they put a couple
00:50:00
of unit tests on their ETL code, and then they're still having a bunch of errors, and they still have the same sort of problems. And I think there's not
just doing Agile, and I've been in some data organizations where they have rooms set up for Agile, and they have all the posters, but they're still basically doing waterfall. They just do a dev sprint, dev sprint, dev sprint, and then a QA sprint, QA sprint, QA sprint. They do Wagile or Waterfall Agile or Potemkin Agile.
There's a bunch of terms. And so I think you do need executive sponsorship is helpful, but you also need to sort of think deeply about exactly how you do it. That doesn't say that two or three people can't start using these principles on a small team right away. I'm a technical person. I don't ever want to deploy a bunch of data code to production without running a whole series of tests on it first.
I don't ever want my Saturdays interrupted. I don't ever want to sort of walk off the soccer field again with a bunch of data errors and my wife scowling at me, that I have to go in the car over my cellphone to fix it. And so I think it's better to start with executive sponsorship, but you don't necessarily have to.
Okay. So we had a couple questions about the best organizational structure for DataOps. So one is, what size of IT environment is suitable for applying DataOps principles? Say we have a BI team of five guys consisting of one DBA, two ETL developers, and two report developers. Is this environment suitable or should it be bigger?
So it depends on how far you want to go with heroism, right? You can actually, on a smaller team, you can heroically divide, and people who are willing to work smart and work weekends, heroism gets you actually quite a long way. The problem is heroes burn out and heroes quit. And so if you've got that small team and they're comfortable fixing problems on the weekends and comfortable going the extra mile, and you're not afraid that any of them are going to leave, then well, maybe DataOps isn't your highest priority.
But what I found is that people who end up doing heroism week in and week out get bored and get frustrated, and then they quit. And then you're left with a hunk of ETL code, a hunk of data code, a hunk of visualization code that nobody has any idea how it works. And then you throw it off to another person, and now they got twice as much crap to deal with, and then they leave. And so partly what you're trying to do at DataOps is make it so that people can easily work in a way that they create something, and they know it's going to work, and they know if they hand it off to someone, it's still going to be work because it's automated.
It has a whole bunch of tests about it. You built a system that can handle the loss of any one individual. So
I guess that's the answer is like small teams, I think should follow DataOps principles. In fact, individual data engineers should follow DataOps as individuals, but it's hard for some people to overcome the bias of heroism, because it's nice when your business customers give you praise, saying, "Oh, you worked the weekends. Isn't that awesome?" And so it really does take a little bit of leadership to say, and often I've said, you sort of praise people in public for working weekends, but then privately you talk to their manager and say, "Why is that guy working on weekends? Why have you not set up a system that they don't have to?" So that's it.
Okay, great. So we have time for about one more question. I think this one segues nicely from the last. How mature are the customers you work with in their DataOps practice? What's the biggest challenge they face in terms of growing immaturity?
I think they all say they start with different levels of understanding of sort of the idea of an iterative process and what it means. And so companies that tend to be bigger and more mature have a lot of sort of Word doc processes and meetings set up to handle collaboration and trying to unlearn some of that is harder for them.
And then the journey that they go on is trying to find a project and trying to start doing an iterative development path and then seeing that success and then getting more momentum behind it. And so it starts small, show value, communicate that value, and get more people on board, because I don't think once... My experience is it's very rare that once people understand these principles, understand how the systems work, that they don't want to work in that way.
Okay. Well, great, Chris. Thanks so much. We are at the top of the hour now, so there were a few questions we didn't get to. So if we didn't get to your question, we will just follow up with you directly afterwards and with an answer. And then also just wanted to let everyone know that
00:55:00
we'll be sending out a recording and the slides for the webinar shortly afterwards. So be alert looking in your email for that. And then I'd just like to thank Chris for a great overview. That was a fantastic presentation. And thank you all for joining us today. Yeah. Thank you, and just if you can see my screen, we've got sort of four webinars, scheduled to come up later.
And then, we also have a lot of ways that you can learn more about DataOps. You can actually go read that DataOps manifesto I talked about. We've got a whole almost 200-page book that you can download for free that talks about these ideas in long form. Gene Kim, who's sort of famous in software's DevOps world, has an excerpt of his book that you can go. We have a blog and even you can see this rotating GIF.
We have a bunch of aphorisms that you can see. So we've a lot of resources for you to go and that are free that you can just learn about DataOps. So, thank you very much, and just finally, given our environment, I hope everyone stays safe and stays warm and stays sane and, of course, stays healthier.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Questions from this session
Why do DataOps at all?
Because the base rates are bad and they are not improving on their own. 87 percent of data science projects never get to production, 60 percent of data analytic projects fail, and 79 percent have too many errors — while investment rises and the share of self-described data-driven organizations falls from 37 percent to 31 percent. Spending more on the same way of working has not moved those numbers.
Is the problem technology or people and process?
People and process. Teams already have capable tools — ETL, warehouses, notebooks, catalogs, BI. What they lack is a way to deploy safely, catch errors before customers do, and coordinate across teams and locations. DataOps is framed here as technical practices, cultural norms, and architecture rather than a product category.
What does Conway's Law have to do with data pipelines?
The structure of the teams doing data engineering, data science, analytics, and production is reflected in the pipelines they build. Three teams with separate handoffs produce three disconnected pipeline segments with brittle joins. A shared platform with common pipeline and environment constructs is what allows collaboration to cross those team boundaries instead of being shaped by them.
What evidence is there that this approach works elsewhere?
DevOps in software and Lean in manufacturing. The State of DevOps Report finds high performers deploy 200 times more frequently, recover 24 times faster, have three times lower change failure rates, and spend 22 percent less time on unplanned work. Lean manufacturing produced higher quality, less rework, better employee satisfaction, and higher profit.
What are the seven steps to DataOps?
Orchestrate two journeys, add tests and monitoring, use a version control system, branch and merge, use multiple environments, reuse and containerize, and parameterize your processing — plus three more covering architecture, metrics, and inter- and intra-team collaboration.
Why focus on operations rather than the next feature?
The session quotes Musk on building the machine that makes the machine, and Nadella's rule that an engineer choosing between a feature and developer productivity should always choose productivity. Google has over 2,000 engineers contributing to engineering productivity. The claim is that capacity to deliver compounds, while any single feature does not.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.