On-Demand Webinar · 1 hr 2 min

Real-Life Examples to Inspire Your DataOps Initiative

Chris Bergh on where to start a DataOps initiative: a Theory of Constraints framework for choosing a first project by finding the bottleneck, and three companies' examples of how they kicked theirs off. Recorded January 2020; updated August 2026.

Presented by Chris Bergh

What you'll learn 7 points
  • The starting-point framework is the Theory of Constraints: improvements not at bottlenecks are illusions. Find the constraint in your analytic process, pick one, and iterate — rather than launching a broad DataOps programme.
  • Three constraints cover most teams: too many errors ('I don't want to learn about data quality issues from my customers'), slow deployment ('I don't want to break production when I deploy'), and poor team coordination.
  • An American transportation company attacked errors first, orchestrating a toolchain spanning Oracle, SQL Server, Hadoop, S3 and Redshift, with Informatica, PySpark, Nifi and Kafka, across on-prem and AWS, and six separate owning teams. Tests validated streaming data was fit for purpose and alerted to Jira, email, and Slack.
  • A European telecom took four months to move an idea into production through a four-stage manual deployment, with almost all testing done by hand. Dozens of automated tests embedded in the pipeline brought deployment down to minutes.
  • A consumer goods company used orchestration to hand ML models from data scientists to IT, so updates to a model stopped being a crisis and expensive data science time went back to creating rather than operating.
  • Automated tests serve a dual purpose: data tests and monitoring in production, and regression, functional and performance tests in development. Quality the customer receives is a function of data, code, and environment.
  • Measure the process, not just the output: production error rates, data provider error rates, SLAs, deployment rates, release environments, and test coverage. Analytic teams are typically not analytic about their own work.

Slides

58 slides

Transcript

Show chapters and dialogue 10,813 words

00:00:00

Good afternoon, everyone. Thanks for joining us today for our webinar. Our topic is Real-Life Examples to Inspire Your DataOps Initiative. My name is Beth Beckley, the VP of marketing at DataKitchen, and I will be your host today. We're thrilled that we've had such a positive response to this webinar with nearly 700 registrants. This shows that DataOps is really more than a passing trend, and as many companies get their DataOps programs off the ground, there's a real need for practical solutions.

Before we get started today, I just wanted to cover a few housekeeping items. The webinar is being recorded, and we will send a recording to all participants after the webinar, so be on the lookout in your emails for that. You're also all on mute, so we will use the last 15 minutes today to answer questions. If you have a question, please enter it in the question box on the webinar control panel. We'll collect all of those questions during the webinar, and we'll answer them during the Q&A session at the end.

Next, I'd like to introduce our speaker. Chris Bergh is the founder, CEO, and head chef at DataKitchen. Chris is a real leader of the DataOps movement. He has more than 25 years of research, software engineering, data analytics, and executive management experience. At various points in his career, he's been a COO, CTO, VP, and director of engineering. Through these experiences, Chris realized that there has to be a better way to quickly deliver innovative analytics without errors, which led to the founding of DataKitchen.

He's also the co-author of "The DataOps Cookbook" and "The DataOps Manifesto," and a regular speaker on DataOps at many industry conferences. So with that, I will hand it over to Chris to take it away. Thank you, Beth. My name is Chris Bergh, and our topic today is Real-Life Examples to Inspire Your DataOps Initiative.

And so our company's name is DataKitchen, so I'm going to give you a warning that there will be an egregious use of food and cooking metaphors in this presentation. So if that bothers you, I think you should log out now. So let's talk about what we're going to do. So the first part of this presentation, we're going to talk about what DataOps is, and I think DataOps is really about reclaiming control of the data pipelines, the process that you work with, and by reclaiming that control, you can deliver more significant business value.

And we're going to do that in two ways. We're first going to talk about the journey of some of our customers and prospects and how they start off with DataOps. And so we're going to present a theory called the theory of constraints, which comes from manufacturing, about how they picked that particular constraint or bottleneck to follow. And then we're going to show in three examples what they did, how it worked, and the success that they had. And we're also going to end up with a little bit at each section, talking about best practices.

Good. So, let's go into the agenda. So the first slide in my deck is sort of what is DataOps? And so, to set the background for DataOps, I think we have to kind of zoom back to the whole analytics industry as a whole. And in this presentation, I'm going to use data analytics to mean anything in the value chain of getting data into the hands of a customer. It could be data preparation, or ETL, or data science, or AI, machine learning, data visualization, data governance.

And, I think everyone here knows that data analytics is hot. You can't walk through an airport or watch a football game and not have something about data analytics. And that's been a real change over the past 10 or 15 years, in my experience. And it's got a lot of buzz, all these terms are, and companies are forming.

But my perspective is, not to be Debbie Downer, but there's actually a lot of failure in analytics, and there's statistics to back that up. There's a lot of data science projects, creation of models that don't get into production. The number of organizations self-reporting that they're data-driven is going down from 37% to 31%. 60%, some people say 80% of all data and analytic projects fail.

And in a survey that we did with Eckerson Group, 79% of all data projects have too many errors. And so, there's great potential. People believe in it.

00:05:00

There's lots of new technologies, new acronyms coming out, yet the situation is that these projects and teams are not being successful. And so why is that? Why are people not being successful? Well, I grew up in the Midwest in the '70s and '80s, and the US auto industry went through a big trial, mainly from Japanese imports and also just from unhappy customers.

And I think in a lot of ways, the data and analytic industry is like the US auto industry in the '70s. The artifacts that we're creating are really full of errors. You would get an AMC car, and it would have 5 or 10 defects, and people would leave wrenches in the cars. They only lasted 50,000 miles. They would rust through.

And the amount of time it took a company to tool up for a new model to build a new car was five or seven years. And along came the Japanese, using some techniques in manufacturing and lean and Deming and were able to produce a better car that was faster and lasted longer. And so I think that's similar to what we're facing in the data and analytics world, because in a survey that we did with Eckerson, it's taking customers months to deploy some new code from their data scientists or data engineers' fingertips and get it to production.

They have abnormally high errors. And I think if you actually talk to your data and analytics team and sit down and have them talk with a beer, you're going to find a lot of frustration. And as a result, some people are actually leaving the data and analytics field, which is, I think, tragic. And I think the other core idea is that your team's time is not really well spent. And first of all, there's just a complexity in data and analytics. There's complicated organizations and roles between data scientists and data visualization and engineers.

There's a complex tool chain. And what I mean by tool chain is all the tools and engines that people use to derive data. There's obviously complex data, big and small, unshaped and shaped, and there's a complex collaboration across people and locations and tool chains. And as a result, the innovation that the customer desires can't be delivered, and people are spending too much time on not the right thing.

And so if we look at that, and we've been talking about this for about six years, the problem is data and analytic project failure. And the given is that there's lots of data, lots of tools, lots of people, and lots of frustration. And so the solution that we think is something called DataOps, which combines ideas from Agile and DevOps and lean manufacturing and applies those sets of ideas to the more complicated world of data science and analytics.

And so it's not us who's, even though we've been talking about it a while, it's not us who believes this. Actually, the industry is starting to buy in on DataOps. So for instance, two years ago, Gartner put it on its hype cycle, which if you look at that graph, it's at the bottom, but it's still on the hype cycle.

Looking at sort of Google Trends, there's been 500% search growth on the term DataOps in the last two years. If you ever just do a Google search on DataOps, every day there's some new blog post or new article talking about it, and that's a big change from six months or two years ago. And Gartner was kind enough to make my company a cool vendor, and there's just a lot of people sort of writing and talking about DataOps, and I think because of this realization that there is a sort of a general process and people problem with data and analytics.

And so, we've got lots of talks where we describe DataOps in detail. And in this talk, I'm just doing a sort of a front or giving you a quick view of what DataOps is in general, or why DataOps or how DataOps. And so if we look at this from a definitional standpoint, what is DataOps?

And this definition is sort of modified from Gene Kim's version of DevOps, a description of DevOps. And so, first of all, DataOps is a set of technical practices and cultural norms and architecture patterns that enable the ability of teams with all their tools to rapidly experiment and innovate for the fastest delivery of insights to customers. So that means rapid experimentation and innovation. One way to think about that is the cycle at which you can do that.

How fast can you go from a keyboard to a customer and back again to learn? So the cycle time is important. And then DataOps also enables low error rates. And for us in the DataOps community, we think of errors as not only involving poor data quality, your source provider gives you poor data, but have you transformed it in the wrong way?

Has the server gone down? Is it late in production? And focusing on errors, which we think is a superset of

00:10:00

data quality problems, I think is an important part of thinking about DataOps. And then this collaboration across all your tools and technology and environments is an important part of DataOps. And then finally, just measuring. We are trying to help our organizations become data-driven, and so measuring these results is important. So cycle time, error rates, collaboration, measurement, these are the sort of things that DataOps is focused on.

And so these are primarily process things. And so, in a lot of ways, DataOps is not about data. It's sort of misnamed. It's about the processes that act on data and the teams that act on data. And so I've got a couple charts that look at what DataOps is. And the first thing to do is think of the first idea behind DataOps is that what you do in analytics is a factory, is a manufacturing line.

And in this metaphor, data comes in one side and a series of, think of them as workstations or manufacturing stations happen and value comes out on the end. And that value may be a set of charts and graphs, it may be a system, it may be a file. But as data goes through each one of these manufacturing stations, code acts upon it. So Python code, ETL code, R code. In this manufacturing process, data is being transformed, artifacts are being added, and if you think of it like that, then you start to think, well, I want to produce really high-quality cars, really high-quality analytics.

And so, this metaphor of a factory I think is important because whether your systems are batch or streaming, whether there's little data or big data, we all have a production process that it goes through. And from my perspective,

having lived in this role of having these production processes run every day, having time delivery, having unhappy customers who would yell when the data was wrong or late, really focusing on lowering error rates in this was a much better, I think, way to work. And then the second part, metaphor with DataOps is how do you take ideas, how do you take innovation from your team and get it into the hands of your customers?

Because if you boil down the entire world of Agile and Lean manufacturing and DataOps, DevOps, they're all about trying to create a learning organization. And so the way that you learn is to do things and get it into the hands of people who can give you feedback. And the faster you can do it and the smaller chunks you can get into the hands of your customer, the more likely you are to learn and the more likely you are to

have to be successful. And so the metaphor here is it's kind of like software development and the software teams through DevOps and Agile techniques and a bunch of other things, infrastructure of code, have really worked on this over the past 10 or 15 years. How do you get something from a developer into production?

And they have a whole bunch of metrics on it and a bunch of ways of working that I think are really a good inspiration. And so you've got to do these two things together, though. You've got to both have this value pipeline and this innovation pipeline, the factory and the software development shop work together. And you got to do it in this really conflicting way. First is you want to run your factory that has low errors, and you don't want to learn about problems from your customers.

You don't want to get the phone call with the yelling that something's wrong. And then second is, if you've got something working, you don't want to live in fear. You want to be able to change it, because analytics at its highest level is kind of a river of questions. We're not in the house-building business.

We're in a service business, trying to help our business customers understand what's happening with the data. And oftentimes when you deliver some charts, a model to a customer, it's just the beginning. It's just the start of the dialogue of follow-up questions. And so how do you do this? How do you run something that is really high quality, runs like a factory, but then be able to change it on a whim?

And sometimes those changes happen at various places and so along the process. And so doing both of these things together, we think is important for DataOps. And so that's my quick summary of what DataOps is and why you should care. We've actually written a book about DataOps, a manifesto. We've got a bunch of videos, and we're actually going to do more discussions about the ideas of DataOps in the future. But we're here to learn about how other companies have started DataOps.

So we're going to focus on that now and go on the assumption that at least you have a passing understanding of what DataOps is from my quick introduction.

So if you think about DataOps as a thing, you draw a circle,

00:15:00

put the word DataOps in, and there's just some stuff you have to do. So why do people do it? What are they trying? What effect are they trying to have on the world? And so the first effect is a lot of companies are producing data, and they don't know if it's right. They're hoping it's right.

Their customers find out, "Hey, this Tableau workbook's empty," or, "Hey, you were going to give me a file, and I don't know if it's there," or, "Hey, this data looks weird." And so they end up in this high error rate environment, and that means that their business customers don't trust the data. And they also don't have visibility into whether all the pipelines that are running, that are sourcing data, transforming data, visualizing data, if they're actually working.

And so this idea of errors and trying to reduce errors and producing quality products, quality insight is one area where people try to focus on making a change with DataOps. And the second is this idea that it takes a while to innovate. And a lot of analytic teams have a backlog of requests from their customers, whether it's in Jira tickets or ServiceNow or a spreadsheet.

There's much more that they possibly can do,

much more their customers want them to do than they could possibly do. And so how could they rapidly be able to get things into production? And so there's a lot of great tools out there that help an individual contributor build a model easier, build a workflow easier, build a new visualization easier. And that's part of it, but also, we think the bigger part is that the bigger problem is not just in building a data transformation, but it's in how fast you can get that data transformation from your developer's fingertips into the hands of your customer to learn.

And then another business constraint that people have, and this is pretty rife across a lot of organizations, is poor collaboration. And this may be between a data engineer and a data scientist sitting at the table across from each other or even across different teams and different organizations. And so how do you work in a world where there are centralized teams and decentralized teams?

And then finally, some managers come in and have seen our product and have seen the reports and go, "I need that. I need to prove my worth. I need to understand what my team is doing in terms of productivity and error rates and time." And so these four things, errors, deployment speed, collaboration, measurement, these are some of the business value propositions that people are trying to affect by doing DataOps.

And so if we look at those four things and we think about, well, where should I start? And, yeah, I talk to customers every day and prospects, and they're like, "Hmm, this..." Sometimes they get a little daunted because it does seem like a big thing to change. And so what we try to do is help them through a process of focusing on where is the biggest problem.

And there's a book that I'm going to refer to in a bit by Goldratt called "The Goal." And in there, there's an idea from lean manufacturing called the theory of constraints. And The idea in this book, and it's sort of well-written, a novelized book, is that if you're going to change something, change it at the bottleneck where the biggest constraint is.

Because if you're not fixing that, then the improvements that you're doing are an illusion. And you can kind of see this in a manufacturing line. If you've got all these steps in a manufacturing line and one is slowing down production, it doesn't matter if you make changes before or after, it's that bottleneck that's going to matter.

And so the idea is find that bottleneck, find the biggest constraint in your system, and then figure out what impedes you from fixing it. And then change it, pick it, improve it, and then go and find another bottleneck. And so it's a very agile, very incremental way of making changes. And I think this theory of constraints, that the idea that improvements not at a bottleneck are an illusion, I think really does apply to the processes in data and analytics.

And so if we take that idea of what kind of bottlenecks or constraints and where to start, and we could kind of bucket them into three areas. And the first is, this is what I've experienced in managing data science and engineering teams. I never liked hearing about data quality issues from my customers. Maybe it's because I'm an introvert or maybe because I never liked getting the shame and blame from people when things were later wrong, but I found errors to be really challenging.

Another area that people start in is it will take them sometimes three months to deploy 20 lines of SQL, and they hear about how fast digital teams do it, how fast software teams can do it. Maybe they've hired a CIO who's saying, "Oh, we've got to do DevOps. We've got to do agile." And then they're in their data and analytic team and scratching their head because it takes three months to deploy 10 lines of SQL.

And then other organizations are very cognizant of that analytics is not

00:20:00

a centralized function anymore. There are pockets and lines of business, and sometimes these fiefdoms, Hatfields and McCoys, are at war with each other, or even just how do they get a data engineer and a data scientist to collaborate better? So bottlenecks, there can be different bottlenecks in an organization that they can focus on. And so the idea here is that you pick a bottleneck, pick errors, pick deployment, and find that bottleneck and try to solve it, and then measure the effects of that solution, and then go ahead and iterate.

And so these ideas that I'm presenting aren't new. They're in Goldratt's book, they're in books about manufacturing, they're in "The Unicorn Project," and other books by DevOps thinkers. But the theory of constraints, I think, is helpful because from a manager, it helps sort of cut through the noise about everyone saying you should do this, but let's find the biggest problem and work on it. And bottlenecks are often the biggest problem that people want to solve.

And so I'm going to go through three examples and three constraints or bottlenecks that people solve first. But I do want to give a little background, and these bottlenecks aren't always completely compartmentalized. Sometimes people want to lower errors, but also improve deployment. I think from a thematic standpoint, it makes sense to talk about what these people have done. So I'm going to start with the first one, and this is an American transportation company. And transportation company has a bunch of vehicles, and these vehicles actually give a lot of data.

And so also their internal systems give a lot of data. And so they ended up, like a lot of organizations, have a very complicated infrastructure where some of it's batch, some of it's streaming, there's different tools and technologies of different generations running, and they work in real-time and in batch, and they've got different teams working together.

And so they want to be a data-driven organization, yet all this infrastructure is there, and all this coordination makes it hard. And so they looked at it and were like, well, what is it that they really wanted to address? And the biggest thing is as data's flowing through these systems, as it goes from one system to other, gets transformed in data science and data visualization and governance, their business users weren't sure it was right.

And that's an incredible constraint to having any organization be data-driven because if you've been in the analytics field a while, you know that sometimes your business customer doesn't want to be data-driven, and the best thing that they can do when they hear something that they don't want to hear is attack the quality of the data.

And so we all face this in the greater data and analytics field. How do we get people to be data-driven when, in fact, they aren't data-driven? And so having good data quality and having evidence of why it's good can help take that one barrier away. And also, a lot of times when you have errors, it creates these sort of chaotic, stressful sessions of fixing and patching and digging in the data to see if it's right.

And that increases risk and also just increases the sort of hassle factor that you live with. And I think that's partly another motivation for DataOps, is that we don't have to live with a high hassle, high problem environment. There is a way to make it more sane and more predictable. And so they chose to focus on errors, and it's really errors plus orchestration of the entire tool chain because they go hand in hand.

When you've got a complex tool chain, you want to be able to understand all these tools and whether each one is working correctly, and whether the data that's an artifact that are created are in fact right and won't affect things down the line. It's an assembly line. It's a complicated assembly line with lots of branches and merges and lots of different types. But they've got multiple databases, Oracle, SQL Server, Hadoop, Redshift. They've got multiple data tools, Spark and Informatica, Enterprise Service Bus and SQL. They've got streaming, NiFi and Kafka. They're both cloud and on-prem, and they've got multiple tools, Python, Jupyter Notebooks, Tableau, for people to get value. And so it's a very diverse environment. And so another tenet of DataOps is that there is not the one system to rule all, that this diversity in tools is not a problem, but a feature, that because the amount of tools that are creating is exploding, and this is actually a good thing.

And so if we look at it from an architecture standpoint, so you've got sources on the left and customers on the right, and then you've got this mix of ingestion types, streaming, fast batch, streaming in Kafka, batch in Informatica, predictive models, and that data itself gets rested in a big Oracle database, and they're trying to work out with S3 and Redshift. And then there's notebooks and other ways

00:25:00

to visualize that data by the data consumers. So the factory's in the middle, right? And then there's sort of different paths to the factory. But the people who work in the factory are kind of organized in different teams. And those different teams, maybe there's a warehouse team and a data science team, an analytics team. They're trying out the cloud.

There's an IT team working on that. And so you've got this complex sort of team ownership and orchestration. And as you know, if something happens in the block on the left, it may be the block on the right that notices it. And so how do you make sure that you find the proper owner of errors and catch it before it gets downstream?

So what we did is that our DataKitchen software has a thing called a recipe, and a recipe is kind of an abstraction of all your tool chains and all the processes that happen. So we created a DataKitchen recipe, and this is a visualization of what it looks like, and you can kind of see how it reflects that architecture.

There's some processing that happens in NiFi. There's some processing that happens in Informatica. It's sort of being marshaled into a big database. Then it's being put onto Tableau and Spark jobs. And so the idea is that by orchestrating all these technologies and then decorating each one of these steps with tasks, that you can prove that your system is working before your customer does, and that the data is validated and fit for purpose at the very end, while all that transformation and actions. And this is, again, why that metaphor of a factory is true, because this is running all the time. And

as you know, if you've got a part that is going into your car, if you've got the wrong engine in your car, your car's not going to work very well. And so we need to test and make sure things are working as the system's being assembled. And so what happens when things aren't working?

Well, it could be that they're not working because the server's down, or it could be that somehow we were expecting 10,000 rows of data and we got two, or we got 10,000 rows of data, but they dropped three columns, or we got 10,000 rows of data, but one of the columns has a 50% difference in the sum of the data than last time. Something anomalous has happened.

So you need to alert your team before it gets to your customer. And so there's a variety of ways that our software does that, through Jira and email and Slack and others. Because partly in these data systems, you need to design the system that orchestrates all these tools, but you also design the system that can patch or fix or recover from these, or understand these data errors.

And alerting and finding out the problem soon is better than finding it out three weeks later when the report's been on your CEO's desk for that time. And then you've got to go back as a team and say, "Oops, sorry, we didn't notice that this column wasn't there anymore." And so alerts and monitors, I think, are an important part of that.

And so what's the result? Well, if you focus on errors and focus orchestrating the system, you can ensure high quality, and that increases the overall trust of your analytics. And if you can automate the orchestration and orchestrate these complex pipelines, you can then have a chance to do much lower errors and have better team coordination.

And so,

ending example one, I just want to talk about what is the best practice here. So what are the ideas behind this? Well, the first is that you need to test in production. You need to monitor the servers, you need to monitor your systems, you need to test the data, test the artifacts across the whole assembly line to make sure that it's right, that it's working.

And so how do you do that? Well, you do it automatically in production every time the data's flowing. You don't do it manually, you do it automatically and across your entire tool chain. And when something happens, you send an alert and a notification as soon as possible. And partly, since there's lots of ways, and if you look at that middle column, there's lots of ways to test, but you should keep track of your test history because that's an important part of what's happened and actually helps you prove your worth.

And finally, make it easy to create tests. And our view on this is that people shouldn't have to replace, whether they like to write SQL like I do, or like to do visual UIs, or like to write Python or Scala, they should have a way to use their native tools and techniques to do that.

And there's lots of ways to think about testing data, testing systems to make sure that you don't have embarrassing production errors. And there's a whole set of industry around traditional data quality and profiling tools, and those profiling tools can give you a bunch of rules that you can implement. There's something called statistical process control, where you look at-- and that comes directly from manufacturing, where you look at the upper and lower bounds of data and fields and sizes and see if a trend breaks there. And we're going to talk about location balance and

00:30:00

other types of tests, but the biggest test is, do you have rules that reflect the heuristics that your customer has? Because I know that many of you have done work in data and put it in front of your business customer, and they within five seconds go, "It's wrong." And we're all smart people, we've all done lots of calculus, and this businessperson instantaneously knows the data's wrong.

And that's because they have a model of the world in their head. They have a set of heuristics. And those heuristics should be boiled into tests. And so why focus on this? Why focus on test automation, and why focus on error rates? Well, I think less errors mean more innovation, and I think it means more customer data trust.

And what it also means is that you have less stress on your team and less embarrassment. And I think this idea of doing errors is good. And the first way I did it was actually to have a quality circle where everyone sat around and just counted how many errors we had that last two weeks and then said, "Is there a pattern in the errors? What's the error that we can fix?" And there's a social part of getting the shame and blame away from errors that I think is important.

So now let's go on to the second point, which is: How do you handle slow deployment? And what kind of customer started with DataOps on this path?

So this was a European telecom company, and they'd like to increase the rate that new data features are delivered into their enterprise data warehouse. And for them, it takes four months to move through that process from-- And this could be a schema change, it could be a new dataset, it could be an alteration of some facts and dimensions or an aggregate, but that cycle takes them four months, and they'd like that to be faster. And so why is that? Well, they've got a four-stage process to go from development to system test, to a pre-production, to production.

And moving from each step is pretty manual. It's kind of open the Word doc, copy from this file. And there's lots of different tools at each one, each step in this process. And they chose to just focus on the data warehouse changes, not all their other tools, the visualization tools. So it's very different than the customer I saw that we talked about before, with all sorts of tools. In this case, they just focused on their database and making changes to their database as really the first bottleneck.

And they thought that by speeding up this development process, instead of taking four months, they could get it down to four weeks or four days. That could actually have a significant short-term business benefit.

And so what makes this speed? What makes it so you can deploy something from a dev to a production? Well, the first idea is automation, that replace manual processes with things that are automated. And the second is try to make sure that you've got a place to keep what you've done, a repository, a central place to keep all your SQL, all the artifacts, all the ETL jobs.

And finally, the biggest one is automated tests again. In development, you need to run a series of tests to prove that what you have done still works, or that what the new thing that you've added works. And they had unit tests, but they didn't have this more wider system regressional function tests to be able to make sure in development that you could deploy it into production.

And again, like we talked about before, embed those tests into the recipes. And so they built dozens of tests in the DataKitchen platform. And here's an example of a recipe, and these numbers actually indicate the number of tests that we're running. And think of it this way, when you're in development, you want to make sure if you move to the next level, that you're not going to be-- as an individual data engineer or ETL engineer, if you're going to move it to the next level, you want to know it works. And there's all these relationships upstream, downstream in data and analytics that are captured in a DAG, in our recipe, but you want to run these tests on top of it.

And so when you change something, let's say you change something here in this create views node in our recipe, you want to add tests, but also tell if you've broken anything and run packages or run QA tests at the end. And so they wrote these in SQL because they have ETL tools, but they're also a SQL shop.

And so what makes the speed? Well, you want to build automated tests and embed them in recipes. And we're a believer that you should, if you like to write SQL, well, use SQL to write your tests. Because I just don't ever want to get involved in a discussion between R and Python, between Looker and Tableau, between ETL tool versus SQL. And they're just-- people love their tools,

00:35:00

and so let them use their tool of choice, but take those frameworks and have that part of what their work do embedded in a framework. And the tests actually help reduce the time at which they find problems and help them locate the problems. And so one specific type of test that we're showing here in our UI is called the location balance test. And it's a little bit hard to see at the bottom right.

It says, "Final table row count equals expected final table row count." And so a type of test step that we do that can be both in production and development is that if you look at, for instance, as data's flowing through a system, it comes from a place, it's stored somewhere initially. Maybe that's a lake, maybe that's a staging area.

It goes through a series of processing steps and ends up in a set of facts and dimensions or other aggregations. Then perhaps that data is used as features that drive a predictive model. It ends up being put in a report, and maybe that report has a caching system like Tableau has. So you've got this sort of following the bouncing ball as the data.

So if your source data has a million rows, as it's moving through the system, you want to make sure that you haven't lost a million rows, and maybe it ends up with 300,000 facts and 700,000 dimensions in your database and then in the report. And so you want to be able to balance and follow the bouncing ball as it goes through your systems to make sure you don't have any loss along the way. And while you ask, could it be that this is very common?

It's not uncommon, and when it happens, it's such an incredible pain. It also helps you prove in development, because a lot of times you have organizations, these pipelines are broken up by who does the work, and you may have a team working on your ETL code and a separate team working on your R code.

And you want to make sure that as you hand off that data in the processing, that it's good enough for the next team to pick up.

And so another point I want to make is that this deployment speed can be scary. And

in data and analytics, a lot of teams start out with the idea that they're going to be fast, and they end up being heroes and working nights and weekends and fixing things, and they own their process. And then people start to leave, things break down, and they add more and more human process in, checklists, process meetings.

And when you get comfortable in that, the thought that you can actually move on from that, the thought, I think, in a lot of people's minds is you go back to that place of heroism, that you go back to working nights and weekends, and that can be scary or concerning for people. Because what if you make a mistake and will you get yelled at, and will you make a wrong decision?

So how do you balance those two opposites between, okay, I'm living in a world where I've taken care of my fear of change by putting a bunch of process around, or I'm kind of being a crazy hero. And I think DataOps offers a middle way between those that you can kind of do both. And as part of what we do, we have much like this webinar, we talk about the best practices and the why of doing DataOps and help organizations go through that psychology change, where focusing on some organizations, they're scared to make a change, some organizations, there's a lot of shame around making errors.

And how do you help the team make that transition? And I think that's part of the value that we provide and our team provides. So what's the result of this focusing on deployment? Well, if you can focus on fixing slow deployment, then you can speed it up, right? And you can deploy safely from dev to production, which means at the end of the day, you can learn more.

And what's the best practice here to go on? So the best practice is that tests themselves have a dual nature to them. Some of them are, you can run a bunch of tests in production, and they're really helping you know that in that case, your code that's acting upon the data is fixed, but the data varies. But in a development process, in that innovation pipeline, your data is fixed from test data, but your code varies. And so the same tests actually are useful in both, maybe not all of them, but a good majority can be useful.

And so I think of automatic tests and monitoring in production, as well as regression, functional performance, unit tests, and development as kind of two sides of the same coin. So let's go to part three, a consumer goods company. So in this consumer goods company, they had a data science team who had created a lot of good things.

And they're expensive, they're hard to hire, as everyone knows. And they had sort of developed some useful machine learning models that were ending up in a dashboard, and this team wanted to move on to the next thing. And so how do you do that? How do you hand off work from a small data science team to a greater IT team?

And so another way to look at that is how the IT teams sort of free up

00:40:00

data science resources. What's the path to production from a data science team who's incented to create new ideas and a production team who's incented to run things with low errors? And so, as we do in DataKitchen, we created a kitchen, which is a place that people work. And then we created some recipes, and these recipes can, in our UI, you can look at them in different ways, kind of the components in the tool chain.

You can look at them in a test view. But again, these tests and errors help prove that the system works. And it proves it in production, and again, as we spoke about, it proves it in development. And our kitchens provide this abstraction because if you're going to be able to help hand off something from one team to another, the team needs to be able to fix the pipeline or develop the pipeline in a place that's not production.

And so for us, a kitchen is where you work and enables that development of that new pipeline capacity. And so the results is that a data science team is able to do that work. And I think there's patterns in a lot of cases in analytics where there are people who are incented to create new ideas, and that's what they want to do, and they're not that interested in operationalizing things. And so how do you hand that off? How do you operationalize between one team and the other?

And I think DataKitchen, through our kitchens and testing and recipes, can help that. And so that's just one example of how teams can collaborate. And so there's actually a lot of collaboration complexity in data and analytics that I want to talk about. And the next few slides are sort of a best practice slide.

So the first is, how do you get a data engineer, DE, or data scientist, DS, or BI to work together without stepping on each other? Because in a lot of cases where there's this D and that circle in the middle is a development team, how do you hand off work? Because sometimes a data engineer does some work and then a data scientist does some work and then the BI person has some work.

And how do you make sure they don't step on each other? How do you make sure that they-- And this sort of collaboration is handled a lot by version control. And our product has embedded version control in, there's other things like Git, and there's features that we have, like recipes and kitchens. But that's not quite it, right? The data person does the data work, the viz person, the data scientists do their work.

They're all coupled together, well, great. But you've got to also take that and push it to production because in a lot of organizations, there's a separate P team, a production team that runs these things day in, day out. And so how do you move what you have in development to a production team with a button push and not a lot of problems? And also, since we all live in the data world, sometimes you get poor data. How do you patch that quick in production so that it can be used?

And this sort of centralized development team, centralized development production team is where some organizations are. And this is a type of collaboration between a team, between two teams that I think is important. However, a lot of organizations have chosen to go the decentralized development part, where maybe they have a centralized data enablement team or data warehouse team or data lake team, and then they have teams all over the organization who are using self-service tools, like Trifacta or Paxata or Tableau or Looker. There's even self-service data science tools out there, and they're doing some good work, trying to build upon that data. And so how do you sort of balance this sort of centralized control with the self-service freedom?

And I think this decentralized development model is actually, I think, becoming the norm in a lot of organizations. And then finally, the last part is the delivery to production. You may have a centralized team who's doing a data warehouse or data lake, and when they make changes, it goes off to the production team, who watches the data and catches the alerts.

But these decentralized teams may be using Tableau and hitting that publish button on Tableau Desktop, and it publishes it to Tableau Online, and that's how they get it to production. And so you may have multiple dev to production groups, and this sort of way of working, this sort of tree or relationship, I think, is important. And then actually, how do you handle this decentralized dev and production? How do you incorporate other people?

Because the trick here is that in these decentralized models, your customers are seeing the end result of this whole network working together. And they don't know that one part of the network had the error and the other-- They just know it doesn't work, and they want it fixed. And so this kind of cornucopia of collaboration complexity is something that I think is what we try to solve with DataKitchen and is an aspect of doing DataOps. Because I think

00:45:00

everybody wants to have low errors, everybody wants to deploy to production, but the complexity of how we've organized ourselves is really quite interesting. And there's a reason for that, and it's actually called Conway's Law, which is we're actually having a blog post. So for you nerds can Google Conway's Law and think about how it applies to data pipelines, and we'll be publishing a blog post later on it.

So, we're getting near to 44 minutes. I want to leave a little bit for questions. The last point is, we talked about these constraints, focusing on errors or deployment or coordination, but you need to measure success, and so you need to be analytic about your analytic processing. And so one thing that I think is important in all of this is that you focus on process analytics or DataOps process analytics.

And so, I find it very ironic that many teams who say that their goal in life is to help their business customers be data-driven, are in fact not data-driven at all about the work that they do. They don't track things like productivity of their team and individuals. What are the error rates in production?

What are the error rates of their data providers? Are they delivering things timed? How fast are they deploying? How many environments are they using? What's their test coverage? And so all these metrics, I think, are an important way to help prove and understand how your team's working. And so I believe that you should be very analytic about your analytic processing, because I think that actually can help you get better.

And reports can help. There's the old adage, you can't improve what you can't measure. And so one way that we do this in our product is through the data that our system throws off. And here's an example of a dashboard, and there's a lot here, but I'll bring you to the attention of the middle part of this dashboard.

And if you look at it sort of over time, and what it shows is that error rates on this red box are declining. Like in production, you've had these errors. And conversely, one of the reasons errors are declining is the number of automated tests are increasing. And we also have metrics on productivity, like for instance, the test-to-node ratio is an important measure of productivity.

We can help you measure your SLA, look at your productivity of individual teams, your deploy rate, your collaboration rate. And so being able to get, and this is a project-based report, not an overall system report, get sort of analytic about what your team is doing, I think is important. It also helps from a manager perspective, sort of prove your value.

And so, we are hosting this webinar because we have a software product that we love to sell. And, so our software product is what we call a DataOps platform. It's a process tool, not a data tool. And it enables you to do DataOps, we think, better than anyone in the world. Fast delivery of analytics, high quality, low error rates, and allows you to use your tools and data stores. And what it does is, like we showed in this, it allows you to orchestrate these and monitor these complex data pipelines. And you saw that in the first example.

Allows you to automate tests and monitor quality. We saw that in every example. Generate sandbox rapidly. We saw that in the third example. Deploy new ideas to production. We saw that throughout. And then also increase the kind of collaboration and communication. And so we think that the benefit of applying DataKitchen is that, if you remember the slide from the beginning, the percent of time that your team spends, I think is too much in non-value add activities, too much time in errors and operational tasks. And maybe this blue is smaller than you think or bigger than you think, but we're trying to reduce that.

And we think if you reduce that, you get more time to add more features and actually improve your technical debt. And then the other benefit is that we think that deployment latency Instead of spending weeks or months to get something from a dev environment to a production environment, you can do it in hours and minutes.

Instead of having a lot of errors and a lot of fear, you can have a low error rate environment. And I think as a result, you can have teams that are more productive and happier, frankly. And so the closing thought I want to leave with you here is some kind of social proof or inspiration.

So if you look at Elon Musk, and he has a company called Tesla that makes cars, and he said his focus is not on the machine, the car, but on the machine that makes the machine, the factory. And that factory is where the sexiness is for him. And here's another one, Satya Nadella, who's then sort of turned Microsoft around, and it's almost a trillion-dollar company now.

And he has a quote that says, "If any engineer has to choose between working on a feature or working on developer productivity, always choose developer productivity." Or if you look in the software industry, DevOps engineers, and we're trying to create a standard role for a DataOps engineer. DevOps engineers used to be called release engineers, and when I managed software teams in '99, they made less than every other software engineer.

00:50:00

Now they actually make more. And finally, Google's got over 2,000 engineers just devoted to productivity. And so if you focus on operations and productivity, not the next feature, you actually enable more features to go out. And so it's a mind change that I think people like Tesla and Microsoft and Google have made that switch.

And I think that's an important way of thinking that we need to change in data and analytics because we're all very harried and worried, and we're trying to get the next feature out the door. But we're always going to be on that treadmill, and what gets us off this treadmill is a focus on operations and productivity and DataOps.

And so to conclude, we have a whole series of webinars planned, some around DataOps best practices, sort of why DataOps, how DataOps, some on our product. But in March, actually, James Royster from Celgene and BMS is going to talk about his experience in bringing a team up on the DataOps curve, and we're really excited about that.

And lastly, if you want to learn more about DataKitchen, we've got a website. We wrote the DataOps Manifesto. Please sign and read. We have a cookbook that's free. There's actually a Kindle edition, too. And Gene Kim, who I referred to previously in the presentation, was kind enough to let us use a section of his new book that has a chapter on kind of DataOps and data-related things.

And so now there's a bunch of questions, and so I'm going to switch over to Beth, and you've been great asking questions, and she's going to present her screen.

And I'm going to answer some questions.

Well, thanks, Chris, for that terrific overview. I do want to remind everyone to please enter your questions in the control panel box, and we'll try to get to all of them in the few minutes we have left. And before Chris jumps into these questions, I'll just answer the most popular question, and the answer is yes, this session has been recorded, and we will send out the recording and the slides after the webinar, so please monitor your emails. We should send those out in the next day or so.

Okay, some questions. So I'm using Airflow. How does DK complement? Well, I love Airflow because I'm a Python programmer as well. And I find the way you write DAGs in Airflow very compact. And so I think our feeling is that Airflow is a tool to do a data job. And if you like to write your work in Python, great, and use Airflow. If you like to write it in Informatica or StreamSets or any other of those sort of data tools out there, go ahead. We're not trying to replace those.

But you generally don't use Airflow for doing visualization. You don't really use it for data science a lot. And so I think it's a multi-tool world, and I think you should have the freedom to pick the best tool for the job. And I'm a fan of Airflow, and I think a lot of ways Airflow is a cool tool.

And so how does DK help with building data pipelines like Striim or StreamSets or SDF? Well, I think our perspective is that we're not a data pipeline tool, is that there's these pipelines out there, tools that already act upon the data, like StreamSets, that are great. It's a great streaming ETL tool. And they use DataOps in their marketing literature. But I think DataOps is about having multiple tools.

It's about orchestrating those multiple tools, monitoring those tools in production, and helping iteratively deploy those tools. So if you're using StreamSets or Striim, go ahead, keep using it, or Airflow. We'll help monitor it. We'll help deploy it. We'll help orchestrate all your different tools and wrap that in creating kitchens and all the other features that you saw today.

And for DataOps, do I need to change any of the tools that I'm currently using? I think of a different way, is you shouldn't be wedded to your tools. And so if you love your tools, great. You should have a process that you can swap in and out different tools in a way that doesn't cause complete chaos.

So you don't need to change any of the tools or the tool chain, but you should think about having a right to repair or a right to change architecture, and I think that's an incredibly important idea. And

then the fourth bullet, how are companies organizing their teams for DataOps? Is there a best practice?

So there's a couple answers to that. So first is there's a general body of literature on teams organize themselves to be agile. There's agile and kanban, there's Spotify method, there's scaled agile. There's different frameworks to do it, and I think DataOps can exist in all those And so I think one of the tenets is that there's a role for a DataOps engineer in this, and having a DataOps engineer as part of the team to help increase the flow of

00:55:00

new features from development to production, help extract environments, I think is important. But the actual process of how teams organize, in general, we think if you're going to do one of these broad agile methods of Scaled Agile or Spotify or Kanban or Scrum, that's fantastic, and DataOps can work with any of those. And do you need to hire a DataOps engineer?

It depends on the size of your team. Our view is that you need to first buy into the principles of DataOps and start focusing on errors and writing some automated tests, start doing some orchestration and continuous deployment. And then you'll see that I think as your team grows, you'll end up with a DataOps engineer or maybe a data engineer who has part-time doing DataOps. And then finally, does DataKitchen manage version control? We're kind of a front end to Git.

So Git is our engine to manage version control. And so our software works in conjunction with Git and other systems. And our recipe abstraction is actually stored in Git, so you can use a lot of Git features if you want, or you can use Git directly. Todd, a bunch more questions come in. So we do have a few more minutes.

We can address a few of these. How can this help teams with testing in BI tools such as Power BI? Oh, that's a great question. So the question is, how can it help with Power BI and testing? So, if you saw in my slides, I had pictures of tools like Tableau or Power BI. And so we have a perspective on those tools, in that those tools they don't just draw charts and graphs. They are low-code development environments. So you can build an if/then/else in Power BI or Tableau.

You can build a calculated field. And that is business logic. And you have a choice in designing your analytic systems where that calculation or that if/then/else is. And there's good reasons why it should be in Power BI, or there's good reasons why it should be in your attribute of a dimension in your database, or it should be created from a dataset in a model, or it should be in the access code. And those are design decisions that you should be free to make.

And however, there's a lot of benefit to sometimes having this business logic in Power BI. And so we think that Power BI should be part of the DataOps ecosystem, that you should deploy it and test it and monitor it so that when some data is running that breaks the bounds of that if/then/else, or in the case of where you take that if/then/else and you mail your Power BI file to another file, another person, that you don't end up with any regressions.

And so, DataOps is scoped not just to data or data science or data engineering, but also data visualization and also things like data governance and security. They're all deployable things that need to be monitored and automated.

Another question is, does this include data governance practices? Oh, that's great. That's a great segue. So yeah, it depends on what you mean by data governance, right? And so I think when you're deploying a pipeline, you should deploy the security that happens on it, the database security, the user security. You should do security as code, right?

And if you're deploying a new change, well, why aren't you deploying who can have access to that change? If you're creating a new table in a database, you should have the security deployed with it. Likewise, you should have metadata about that new table. So for instance, how do you know what's in that table?

How do you know what those columns mean? And there's a lot of tools like Alation or a wiki where people create data descriptions. And I think those things are part of what you should do in DataOps. And so if we boil it all down, you should work in short incremental releases. And what you do in those releases is up to your dialogue with your business customer.

Maybe it's a new table, maybe it's a new model, maybe it's a new visualization. Maybe it's some documentation of the tables that are already in your database. Maybe you need to improve your security. All these things are deployable pieces that you should orchestrate and organize. And so I think data governance is part and parcel of DataOps, if you define data governance in the restricted way I did.

And so there's a lot more to data governance than just data catalogs and data security. We probably have time for about two more questions. Here's one we've already touched on, but what competencies should a DataOps engineer have?

Well, for us, we hire DataOps engineers, and we look for, number one is that they have some knowledge of data and analytics. Maybe they've done a little SQL, maybe they've done some reporting, maybe some data science. And then second, because a DataOps engineer is part and parcel of dealing

01:00:00

with servers and hardware and versions, they should have kind of a modicum of DevOps skills, and maybe your environment is AWS or maybe it's on-prem. And third is that they should have a service attitude because what you're trying to do is help team members deliver features faster, and you're building infrastructure and libraries so that they can do their work.

And I think that best is accomplished by having kind of a servant or a service attitude. Okay, I think I have time for one last question, and this is a good closing question is, any quick tips on how to frame the benefits of DataOps to an executive audience?

So, the way I frame it is similar to what I did, is that in general, the promise of data analytics isn't being realized. And you're not going to get success by buying another database or buying another tool. It's about a people and process problem, so we need to affect how our people work. And

how people work in these technically complicated environments, these ideas from Lean and Agile and DevOps apply. So we need to start doing that type of process. And so unless we adopt DataOps, you're just sort of throwing money away, spending $2 million on another database. Great. Well, thanks, Chris. Thank you everyone for attending. If we didn't get to your question, we will follow up with you directly.

So thanks again everyone for coming today. I hope this was helpful, and we look forward to seeing you in a future webinar.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

Where should a team start with DataOps?

At its bottleneck, using the Theory of Constraints — improvements not at bottlenecks are illusions. Ask where the constraints are in your analytic process and what stops you creating new insight, then pick exactly one and iterate. Starting everywhere at once is how DataOps initiatives stall.

What are the most common constraints?

Three: too many errors, slow deployment, and poor coordination between teams. Each has a recognisable complaint attached — not wanting to hear about data quality problems from customers, not wanting to break production on deploy, and the standing feud between data science and analytics teams.

How did a company reduce errors in a complex toolchain?

An American transportation company running both streaming and batch — Nifi, Kafka, Informatica, an ESB, across Oracle, SQL Server, Hadoop, S3 and Redshift, on-prem and AWS, owned by six different teams — built end-to-end orchestration with tests that validated data was fit for purpose, so the assumptions made downstream stayed true. Failures alerted into Jira, email, and Slack rather than being discovered by customers.

How much can automated testing speed up deployment?

For a European telecom, from four months to minutes. Their existing process was a four-stage manual deployment with almost all testing done by hand. Dozens of automated functional, regression, unit and performance tests embedded in the pipeline removed the manual verification that made each stage slow.

Isn't deploying faster riskier?

Speed without tests is riskier. The session is explicit that moving code quickly into production is scary — the fear is making a mistake and having the business act on wrong data. What makes speed safe is that the tests proving correctness run automatically on every move between environments, so faster deployment also means more checks per change, not fewer.

How do you measure whether DataOps is working?

By instrumenting the process itself: production error rates, data provider error rates, on-time delivery within SLA, build time, deployment rates between environments, test coverage, and per-project productivity. The pattern to look for is error rates declining while deploys and automated test counts rise.

Where to go next