On-Demand Webinar · 57 min

Building the Business Case for DataOps

Every data team knows DataOps delivers better analytics faster. Getting the rest of the organization to fund it is a different problem. DataKitchen CEO Chris Bergh and VP of Marketing Beth Pfefferle cover how DataOps builds on the staff and technology you already have rather than replacing them, how to construct the internal case, and how to collect the small wins that show its impact. Recorded January 2022; updated August 2026.

Presented by Chris Bergh

What you'll learn 7 points
  • The numbers behind the business case: Gartner puts 60 percent of data analytic projects as outright failures and says 87 percent of data science projects never reach production, DataKitchen finds 30 percent of companies have more than 11 errors a month, and 79 percent of teams want weekly therapy.
  • Gartner's 2020 figure is that only 22 percent of a data team's time goes to innovation and 78 percent to errors and manual execution. The 2022 DataKitchen and data.world survey adds that 52 percent of data engineers call errors a major source of burnout.
  • The ROI case has three separate calculations. Production errors cost team size times average salary times number of errors times percent of time on rework and triage, plus compliance risk and business opportunity cost. Downtime costs cost per outage times mean time to restore times change fail rate times deployment frequency. Lost productivity costs the time in unneeded meetings, context switching, waiting, and rework.
  • Published ratios of data engineers needed per analyst or data scientist range from 1:1 to 12:1. The claim made here is that DataKitchen flips it, with one data engineer supporting 12 analysts. Even a conservative 20 percent improvement across 10 people saves the equivalent of two people, which is more than the cost of the system.
  • One customer's before-and-after: cycle time to publish a new visualization went from weeks or months to next day, schema changes per data engineer per week from 1 to 12, analysts supported by one data engineer from 0.5 to 12, automated tests per build from zero to 1,000, and errors per build from frequent to none.
  • Rajesh Gill, former Celgene Senior Manager of Business Analytics, Forecasting and DataOps Strategy, reported errors dropping to about one per quarter when DataOps was first implemented, and several years without a major glitch after the team kept adding tests.
  • The advice on selling internally is to skip the DataOps explanation and show results instead: fix low-hanging fruit, publish metrics so everyone is accountable, establish a baseline before you claim improvement, and attach the work to an initiative already funded, such as a cloud migration, a data mesh, a data fabric, or a Snowflake purchase.

Prefer to read it? The written version is in The Business Case for DataOps.

Slides

56 slides

Transcript

Show chapters and dialogue 9,484 words

00:00:00

Welcome, everyone. Thanks for joining our webinar today. My name is Beth Befferly, and I'm the VP of marketing at DataKitchen, and I'll be the host for today's webinar. So our topic today is how to build the business case for DataOps. We'll hear from Chris Bergh, who's the founder, head chef, and CEO at DataKitchen.

For those of you who don't know him, he's a leader of the DataOps movement and co-author of "The DataOps Cookbook" and "The DataOps Manifesto." So Chris will start with some background and cover how to think about the ROI for DataOps, and then I'll jump in and share some tips for selling the business case that we've heard from our DataOps community.

So before we get started, just a few quick housekeeping items. We are recording the webinar. We'll send the video and slides out to everyone as soon as they're ready, likely within the next twenty-four to forty-eight hours. We've also reserved the last fifteen minutes of the webinar for questions, so please enter your questions into the Q&A box on the control panel during the webinar, and we'll make sure we have time to answer those and get through as many as we can at the end.

So with that, I'm just going to hand it over to Chris, who's going to kick things off.

Hello, everyone. Thanks for the introduction, Beth, and I'd like to talk a little about building the business case for DataOps today. And so I hope everyone is warm. We're here in Boston today with five degrees and, of course, being a global company, we have people on the other side of the planet who are having ninety-degree days and summers and people like Jessica from Atlanta who are up here in the frigid north who are very, very cold.

So I hope everyone's doing good today. So what we're going to talk about today is not specifically what is DataOps, but how do you actually get DataOps to happen in a business sense? So how do you actually make a business argument that increasing agility and cycle time has an ROI? And then how do you actually go about convincing people?

And so to do that, we're going to have a few slides just to talk about what we mean by DataOps, and then we're actually going to go into three buckets of ROI. And the first is in production, lowering error rates. The second is in development, increasing cycle time. The third is actually in team productivity.

And then the fourth, we're going to give some customer examples, and then Beth's going to take over and talk about how we get some stakeholder involvement. And so, I think that the background of data and analytics is in some ways failure, in some ways the opposite of these things, that we aren't delivering new and original insight fast.

It's taking a while, that there are lots of problems and failures and a lot of expense is going into analytics for not ROI. And so for us, I think the most important idea is adopted analytics, that you're making your business customer successful, that you're finding value, and to do that quickly with high quality and low cost.

And from some of the justification of DataOps is that it's a very high failure rate environment. Lots of projects that I've seen in my career have failed. Data lakes, data warehouses, data marts, business intelligence, predictive analytics, data science, they've failed. And a lot of companies are putting out data and analytics that are problematic, that require rework, that have errors.

And a lot of data science models actually never see the light of day. They never leave the desktop of the data scientists. And finally, last year, we did a survey that a lot of people, especially data engineers, just want a therapist because that rock and the hard place that they live in has made their lives very stressful.

And so in a lot of ways, the sort of field that we work in is struggling. And for us, the solution to that struggling is not to sort of buy new tools because there's just so many great tools out there. It's actually really think about the people and process that you work in and to apply some ideas from Agile, from software DevOps, from lean manufacturing, do it.

And focus on a few things. Lowering the amount of errors, wrong data, misconfiguration of your reports in production, being able to take the ideas that get in your data scientists' or data engineers' head and get them live in production and get feedback from your customer quickly and change. And then also think about how much your team can do, how much time they're actually spending on task as opposed to not on task.

And then measurement. And so this is what we think of as DataOps. And if you look at grossly what we think of as the ROI of DataOps, imagine you've got a graphic equalizer and each one of these bars on the left are where we are. Companies are taking instead of

00:05:00

hours to deploy, they're spending months to get a new twenty lines of SQL into production. And every day they're running into data errors or report errors. And teams are spending a lot of time in meetings, a lot of time on rework, a lot of times on costly coordination challenges. And we don't really know for being data and analytic teams, we don't really know how well we're doing on these things.

And so that means that we actually have very unproductive teams, which is high cost, and we have a lot of unhappy customers. And one very simple way to think about the ROI of DataOps is that it's possible to actually shift all these levers up, and it's not a trade-off. It's not a trade-off between going faster and breaking more things. It's not a trade-off between productivity and scary changes. That best-in-class teams can do all these things together.

And if you actually work on them together, you actually end up driving costs lower by increasing productivity and your customer happiness. And so this is an incredibly important idea that these choices, these opposites that we may have learned growing up in data and analytics, that if I go fast, I'm going to break things aren't true, and that best-in-class are able to break this really through the ideas of automation and DataOps And so hopefully you believe that because you're listening.

And so how do you actually explain that to people and your CFO? How do you put a hard-headed business look on this? So what we're going to talk about is how to judge the return on investment or how to calculate it. And we're really going to break it into three sub-components.

Sort of production, deployment, and then productivity. So let's first talk about production error and the ROI on production errors. And so what do we mean by that? So everyone's got a data and analytics system, and data comes in and gets in a bucket store and a database. You do some transformations on it in your favorite ETL tool or ELT tool.

You have an orchestrator like Airflow or Control-M, you run models on it, you have visualizations, you have governance and security, and you know what? There's problems. And the error rates, and that means wrong data, late data, misconfigured reports, are just staggering in companies. And in a lot of ways it's staggering, not only because they have way too many errors, but because they've leaned away from looking at it.

And what it means is that lots of reports from business customers are that they don't trust the data or there's two versions of everything. And the idea here is that error-free days are all too uncommon in our data and analytic teams. And so what does that mean? That means that we're reacting to errors.

And so if we look at it from a Gartner report, that only 22% of the time is spent on innovation, 78% is being reactive to errors and doing things manually. And it also means that in our last report, 52% of data engineers said that errors are a major source of burnout. And so because we've built these systems that have a lot of errors, we're ending up with spending the wrong amount of time, and that wrong amount of time actually has a cost. And so what would be the right way to do it?

Well, I think first of all, what we talk about is to be able to go in and automatically check for errors in production, observe them, test them on top of your tool chain. Whether you've got

Control-M in production, whether you're using SQL or ETL tools, whatever your tool chain is, make sure that you're grabbing bits of data and checking it and send alerts if you find something wrong. And be able to actually do a lot of different types of tests, from profiling to statistical process control to location balance tests. And the business value is really that if you focus on automation and testing in production, you're going to actually drive more time on task, more innovation.

You're going to drive higher customer data trust. So how do you actually calculate that? How do you calculate these principles so that you can actually get time to do this? And so here's a very simple equation. So think of production data errors as having three components. The first is team time, the second is compliance, and the third is opportunity cost. And so if we look at this equation on the left, you take your team size and your team salary and how many errors you have, and what's the time that you spent on those errors?

And just do the math. And so if you've got a team where the fully loaded cost of five or six members, let's say a five-person team at 150,000 person team, and then they're spending 20, 30% of their time on reacting to errors and fixing things that should have been fixed, that should have had problems in the first place.

00:10:00

You've got a multi-hundred thousand dollar productivity problem there. And then that's one part, the team time. Then you actually look at the risk of having data errors from a compliance standpoint. Some of you may be in data domains where compliance is really important. For instance, financial services or healthcare. Some may not. And so if you get the wrong data from a report to the government, that may be very problematic.

It may be less problematic if you're just giving your sales VP a report. And then lastly is, this is actually really fairly hard to quantify, is what's the business cost here? What's the business opportunity of the lack of trust in data from your business customer? Because we're all a team to really have the effect of having our business customer or our end consumer of the data trust us and trust the data. And when they stop trusting us, that means they're actually not making data-driven decisions. And there's a huge opportunity cost to not affect that change. So if you think about it, having wrong data or wrong reports and models are very costly.

And that cost could come from business customers' lack of trust,

lack of compliance, and then just your team spending so much time on find and fix and triage and repairs. And I don't know how many times I've sat with groups of people where, yeah, we have 100-person data and analytics teams, and the 20 of their best people are on the phone trying to figure out what the problem is, looking at the logs, trying to fix it, and then everything else stops because there's a major data outage. And I think these sort of factors, team time, compliance risk, and business opportunity costs, are a way to calculate the end effect of production data errors.

And our hope here is that by taking and reducing these things, you can actually shrink the amount of errors in operational tasks, and you can shrink the amount of time your team spends doing these things. And in some cases, we've seen customers get it to almost zero because they've spent time upfront doing automation.

And I'll show you an example from Celgene about where these sort of automations and the time and the ROI on doing it. And so developing these ways to do DataOps, I think has a good ROI. And so how can you get to these error-free days? Well, I think thinking about it in that team productivity suffers, customer data trust suffers, compliance suffers. And if you think about not changing your existing production process, but decorating it with tests and automations. Test and monitor, check the data, and really don't worry about whether you're going to have to decide what the tests are. Put in basic tests and work from there.

And when something happens, give context on the change. And then just try to do this to prevent issues from happening in the first place, which actually leads to our second point. But this is really a justification, this sort of equation and this way is a justification of how much time you're spending in this reaction break-fix mode and putting things into production that could be caused by poor data, but also could be caused by the fact that you've put something in, you've deployed something into production during your development process that's problematic. And that's our second section to talk about.

And so, the challenge here is how to get something like a change to an ETL or a schema or a visualization or your data governance or data security into production from a development environment. And there's two aspects to it. There's sort of velocity and risk. And so one aspect, some organizations, it's very slow because there's manual checks and manual tests. Some organizations have automated the process, but it's still slow.

And so either you're going to move things from one place to another, one environment, development, or do it with CI and CD. And so how do you actually do this? How do you do fast and low risk at the same time? And if you look at it from a complexity standpoint, just think about the top columns here. An individual development environment, maybe a team development environment, a test or a UADT development environment in production, and all the different versions of code and tools and test data that happen across.

And so how do you actually manage this life cycle? Because the most important thing is that in these cases, if I make a change in my specific tool, my Python or my ETL tool, how do I know the impact of that change on the whole system? How do I do impact analysis or from what a software engineers do, regression testing? And so there's complexity in the environment, and then there's complexity in just figuring out what the problem is.

And so likewise,

00:15:00

what's done today is there's errors and differences in environments. There's sort of a patchwork of manual operations. There's sort of hope and heroism. I throw it over the fence and see if it works. And that causes a lot of wasted time, wasted energy, and wasted rework. And as an individual contributor, you build something, you want to know that you built it and not come back to it weeks later and say, "Oh, someone found a problem. I've got to switch my context and fix it again." And so this slow movement or this risky movement through this life cycle, just ends up being a cost that you've got to bear.

And so the solution here is really to automatically test in lower environments, in your development environment and QA environment and pull the pain forward. Have your data engineer or data scientist, when they make a change, see the impact of that change on everything downstream from it. And do that in an automated way through testing, through being able to standardize on your test data, being able to abstract your environments. And from practices here is that code-based unit tests are just not enough.

You need to not only check your piece, but your piece in relation to every other pieces. And the principle here is that really, whenever you're testing data in analytic systems, you're testing data. And you're testing data as it lands, as it flows through your particular tool, particular code that's running. And you can reuse a lot of those tests that we talk about in production to help do regression or functional impact analysis tests.

And if you can lower regression errors, if you can find problems in a QA environment as opposed into production, if you can find it in a development environment as opposed to QA, that means you've got faster deployment. That means there's less context switching, and that means you've empowered your people to make a change. And so what we want here is to think about it from a math standpoint.

Just think about, I've put things into production that cause a problem, right? So assume your data is perfect, and I've started to deploy things that are incorrect, right? And what that means is I have an outage, right? I've misconfigured my ETL process, so my data's wrong. The raw data's perfect, but my ETL process is wrong.

And so while that system is down, you have a time and energy to restore it for the person who made the regression or didn't see the impact to go back and change it. And then how many times has that happened? What's your failure rate? And then how often are you deploying? And if you look at that, these deploy failures of, I've put in code into production that's incorrect, that doesn't work, or that has a problem.

And that could mean that if it's not caught by your business customer, that just means it's going to take your team's time. But if it fails with some downtime, it fails where your customers are not trusting the data, well, that ends up being another case of another factor that goes into it. So if you look at it from the overall cost, the total cost is really how much downtime you have because of deployment failures.

And so these things are something that you can go ahead and calculate. And so our vision here is to be able to go in and, say, take your deployment latency from weeks or months to hours or minutes, and to do that with low risk. And the net effect really comes in happiness and productivity.

And another way to look at this, to sum up, is what we want is low errors in production and fast, low risk deployments. And what that means is that if you deploy quicker, another way to think about it is our whole point in data and analytics, which is often hard to judge, is insight generated. And the effect of an insight is often hard to calculate.

Is it the insight, or is it the actions from the insight that were taken? And I've sat in meetings with executives debating both. But the more often you can actually get insight to your business customer, the time to new revenue is faster. And so if you look at this chart, if you can release faster, that's what that product release cycle is, you end up with potentially more revenue.

And so better insight, better cycle time, lower risk means potentially more groundbreaking insight. And so if that's very hard to calculate the value of a great insight to an organization. And so I didn't include it here as part of the ROI, and because it's notoriously hard to calculate. But it ends up showing up in the reputation of your team.

If you're delivering a constant stream of new, important insight that your business customers find value, if they trust the data, if they trust you to actually put things into production with low errors,

00:20:00

then you're able to actually drive ROI, but it's notoriously hard to calculate.

And so the last is productivity, and this is the sort of the production of your team. And honestly, data and analytic teams are not very effective or efficient. And that's because there's a lot of rework, a lot of disease, a lot of excess meetings, a lot of context switching. And that's why we get these reports of people leaving the field, people wanting therapy, people being upset.

It really makes for a slow and frustrated workforce. And in a lot of ways, team members almost have been set up to kind of work with their head in the sand, work on your model and throw it over to someone else, and they sort of throw it over the wall and hope it works.

And they have no context in their work. Or they've worked really hard to build something, and then they're afraid to change it. And so what we want to do is get our teams to spend kind of more time on keyboard, more time on task, and really have a system that allows them to see the impact of their changes, allows them to understand what's happening.

And so building automation around their work, not manual systems, drives productivity. And so let's look at one case that we got from one of our customers. As part of their justification for DataOps, they just sent a survey off, like, "Where are you actually spending your time?" To their whole data and analytic team. And it turned out that 60% of their time in this survey was in just completely non-value added work, meetings, responding to bugs, unplanned work.

And so that's incredible, that 60% of this team's work is just in overhead. And so how do you actually get that down to be a small amount? And that's where your team is judged. If your team's doing a lot of work but not getting a lot of done, it's because you have these huge sources of wasted time in your organization.

And let's look at it from a graphical standpoint. And so what we're trying to do is look at your team, right? At your whole staff, how much does that cost? They put out a certain amount of insight, dashboards, models, reports, et cetera. And the key question is can you get out more useful insight with the same team and not have to every time people want more, you've got to add more staff, because that's untenable, right?

And we're reaching a maturity in the data and analytics industry. We've gone through a great boom time, and just adding more staff isn't the answer. How do you improve the productivity of the staff you have? And so what we want to do is say, "Look, can you invest in DataOps?" What we call the process hub. Can you invest in this automation?

And that automation, by putting time on task to do that, results in productivity increase of your team, and the result is you get more insight at the same cost. So what we're talking about is kind of an efficiency ratio. By not just doing day-to-day work, but investing in automation and testing and deployment in these DataOps processes, you actually get more done.

And so by building the system and building the factory and not just focusing on the individual work in the workstation, you actually get more done. And so here's an example. Look at it this way from an industry standpoint. What's the number of ratio of data engineers to support a data analyst? In some, it ranges from one engineer to support one data analyst to 12 engineers to support one analyst, kind of depending. And in our experience, that ratio is actually flipped. We've got sort of one data engineer can support six or 12 analysts. And you'll see what that means is we've actually seen literally 12 to 144 times increase in efficiency by being able to use DataKitchen, by developing the system in the factory.

And even if you're conservative on it, you're starting to get a 10X productivity increase because people are working on automations, and you've invested that time. And this very much parallels the kind of productivity increases you've seen when people go through a DevOps transformation. And so how do you calculate that? Well, let's look at it, right? Again, if you've got your data team size and your average salary, how much time are they spent on context switching and meetings, and how much time are they spent on rework?

So if you take that last example, 60% is doing non-value-added work, and maybe 20% of their time is reacting to big errors. So you've got 80% of your time is non-value-added work. Well, just do the math. If 80% of your team's time is non-value-added, that is a huge opportunity cost. And if you can actually make a dent in that productivity in a significant

00:25:00

way, you've got a huge ROI from any investment. And also, you just got an ROI because your team is doing better. And so we're just making this long, extended argument that investing in what we call DataOps engineering, being able to help support your data governance, your data team, your data scientist, and by trying to fix these kind of problems.

Are they throwing it over the wall in the prod? Does done mean done, or done mean it works for me? I've got my little part working. It's not my problem. A lot of times data scientists and data engineers are really willfully blind. It works on my machine. And a lot of times we find organizations, they're task and not value-focused. I've done my work.

I've closed my Jira ticket. I don't really care if the customer's getting value from it. Or they're focused on their project and not the product. They've got a whole list of a Gantt chart, and as long as they're working on the Gantt chart, that's what matters. And there's a lot of hope and a lot of heroism and lack of willingness to pull the pain forward. And so what we believe is that by taking the work that your data engineers or DevOps engineers, taking those nuggets of code and putting it into pipelines, creating automated tests, automating the factory, automating deploys, being able to share and collaborate on work, being able to measure success, and kind of taking what's invisible and making it visible is what drives high team productivity. So it is this productivity is a direct result of investment in these kind of DataOps activities.

So let's talk about a few examples, and then I'll finish up. So, one example is a customer of ours where there's sort of a multilayer team, where there's a team that lands the data, there's a team that sort of prepares the data in usable formats, and then there's a team that actually are business analysts who use the data.

And one of the challenges that the business analytic team, the people who use the data to get insight from their business customers, really feel the pain of when the data is wrong or when the data's late and really have their own sort of production process to do this. And these three groups of people sort of landing the data, the group of preparing the data, and the group of getting value out of the data,

have a constant challenge in working together. In a lot of ways, the final team, the business analytics team, is kind of constantly firefighting to get things done. And any time they want new insight, there's just a lot of rework and a lot of manual processes. And so the way they looked at it is, either we could kind of continually throw more people, more business analysts to get done, or we could kind of invest in automating.

And so one of the ways that they look at it is think of three ways that these automations could help and three ways to think about it. The first is kind of the wedge. Could you actually build integrated data sets that are business useful for your data analyst team? Think of them as marts or sometimes people are calling them, a better way to look at it is a data mesh, where it's under their control and they can be changed quickly. And the benefit here is that business analysts can find, instead of having to go through a thousand tables, they can go to just 10 tables and get their answer quickly.

So it kind of lowers the cost of business analysts to be productive. And then the second is, well, can you actually fully automate the production of analysts from landing of new data all the way to the report? And there's a lot of sort of manual steps that happen in report production at this organization, and can that be automated?

And then the third is kind of think of it as a store. A lot of when you have dozens and dozens of business analysts in different parts of the world, there's a lot of rework, and no one knows if someone else has done something else. And there's a lot of redundancies and productivity because there's no sort of sharing of best practices. And so these three business benefits.

So one is that give your data and analytic team, your business analysts, a data set that they can trust, a data set that they can change, and a data set that makes sense to them. Because they're actually doing a lot of different things. Some of them are using tools like Tableau. Some of them have people that are voice-enabled to be able to do it. Some people are writing SQL.

But if instead of having them go against hundreds of tables, can you give them the 10 that matter that's integrated? And then be able to actually answer specific ad hoc questions by giving them sort of aligned data engineers who can help answer their questions. So the first benefit of the wedge is what I think people in data and analytics have been talking about, all the way from data mesh to data marts, saying

00:30:00

a great representation of data that fits what your business customers need, that can be rapidly changed, that can be trusted, that has an up-to-date data dictionary, has up-to-date automated tasks, can really help your data and analytic teams be productive. And then the second, and here's a case, and this is a bit of an eye chart.

So imagine this is a pharma brand, and they've got a very small team, sort of one and a half DataOps engineers, kind of two analytic engineers, and then sort of one analyst. And this is a launch brand, a very complicated new drug that requires a very interesting supply chain And so how do you integrate 70 different data sets and be able to do that updated every 30 minutes, and to be able to automate the QC of the data, automate it so that these different teams working for different companies in different locations don't run into each other?

And so from the pharma world, this sort of

example usually requires not a team of four people, but a team of 40 people, from what I've seen. A lot of times there's overhead and meetings and slowness, and it doesn't even allow the sort of rapid new insight. And so this sort of small team being able to deliver insights at Amazon speed where the data can trust it, where the data drops, everything runs through, and all automated testing and report exception happens at every level is a remarkable fact. And this sort of hammer to automate the actual production of analytics frees a team from doing the mundane day-to-day work of managing data and exceptions and problems. And so that drives productivity.

And the third case is really the store. And here's an example. So imagine you've got some SQL that maybe a data analyst has done, and is it right? Is it not right? Where did it come from? Who wrote it? Is it tested? How do I find it? Does it live on someone's disk drive?

Is there the resulting data? Are the codes in it? And there's tribal knowledge that happens in any business analytic team. It happens in any data engineering team. And so how can you make sure that tribal knowledge lives in one place under version control, that it's shared? And it seems like a small problem, but there's a lot of repetitive work that happens in organizations where they don't know what happens. They can't find it. They can't search it. They can't get at it.

And so is there a way that you can take out that tribal knowledge, the hidden code, and put the processes that act on the data in one place. And we spend a lot of time as an industry talking about getting your data in one place, but I think we need to spend time focusing on getting your processes in one place. And so the benefits here are really kind of extraordinary, right? Being able to make a schema change sort of once a week to 12 weeks, to being able to publish new visualizations every day.

The ratio of support of data analysts by data engineers is remarkable, and just the amount of automated tests that happen and the number of errors. And so a lot of organizations in marketing or in retail, they live with a very costly, very unautomated, very error-prone production. And so by developing DataOps automation around it, you can see these kind of benefits.

And here's another case, another customer, Celgene. So, I think one of the things that is remarkable of this case and is that they had almost no errors in production, yet they were making

... And here's a quote from Rajesh saying they've been several years with no major glitches. And what that means is that instead of having major glitches,

that means that they have more time to do things. And what that means is you've got some automation to do it. And so if we look at the math here on this, and so imagine you've got $130,000 FT salary, and then you look at their fully burdened cost, and then you look at the team size of 10. So let's say 50% of the time is spent on operational execution, and then 20% of your time is actually spent on automation and testing and deployment, some of the things. You end up with a 36% net savings, or about a third of your team costs can be redeployed either on other things or on delivering more value. And so by doing automation, you've kind of net effect increasing the productivity of your team by a third, according to this calculation. And if you want more, there's a blog post that our customer wrote a year or two ago.

And so here's a great example of ROI and a very clear calculation.

00:35:00

And so ROI is good, right? And these kind of calculations I hope you find useful. But that also doesn't only make the business case. Perhaps it should by showing good ROI, but it's also there has to be convincing of different shareholders. And our community of customers and people who've been involved in DataOps transformation, we've been talking to, and I'd like to turn it over to Beth to kind of talk about what their feedback is on how to sort of sell DataOps.

Yeah. Thanks, Chris. So, as Chris said, as you know, we've been engaging with the DataOps community for many years, and we've certainly learned a lot along the way. We also recently started a DataOps community of practice where DataOps practitioners can get together and share their ideas and tips for success, which has been really fruitful.

So I'll share some of the great tips we've heard and seen work on how others have sold DataOps within their organizations. So you can go to the first slide, Chris. Next slide. So probably one of the most important tips we've heard is just to get started. There are likely a lot of low-hanging fruit within your organization where you can make improvements and make a difference.

So you could certainly spend a lot of time explaining DataOps, but the biggest impact will come when you can just show some results. In fact, in most cases, you may find that the stakeholders don't even really care about the technical details of DataOps. They just want to see the results and the business outcomes.

So, we often use a factory analogy to describe DataOps. So using that analogy, most senior leaders, they don't care why the factory's efficient. They only care that the output right, is faster or cheaper. We have one customer who created an operational excellence team, and the sole charge of that team was to find and fix low-hanging fruit.

So they created a Tableau dashboard where everyone could understand what the problems were and track how they were doing. So this made everyone accountable, and it also made very clear the benefits of doing DataOps. So consider that if you're in a position to make small changes without senior leadership approval, this can be a really important and powerful way to gain larger buy-in for a larger DataOps program.

You can go to the next slide, Chris.

So where else can you find low-hanging fruit? You can certainly look for bottlenecks, and one way to help identify them is with a DataOps maturity model assessment. We have an online maturity model on our website, but we can also come and do it with you, that we created with the Eckerson group. And this tool helps you to identify key areas of weakness along six different DataOps dimensions.

So errors, cycle time, collaboration, team culture, process measurement, and customer happiness. So by taking a tool like the assessment here, you'll know where to focus and be able to put in place a roadmap to improve going forward. Yeah. This also, having done this with a few customers, allows you to sort of benchmark the growth, right? To say, "Here's where we are.

Here's sort of the North Star where we're going." And as you implement aspects of DataOps, you can start seeing the changes that people have. And I think that's a very powerful way to show how well you're moving forward. Yep. Great. So next slide. So in order to prove the value of DataOps to your stakeholders, one of the most important things you can do is to establish a baseline.

Otherwise, you're not going to be able to show results in any of the projects that you tackle. So you need to know where you're at today and what improvements you're seeing. And to do this, you need to really start measuring everything your team is doing. Chris shared an example of one of our customers earlier who did some time tracking, and that's just a great way just to categorize how much time your team is spending on certain activities, like finding and fixing errors, how much time are they spending in meetings, and hopefully how much time they're spending in value-add activities.

You also want to track things like what's your cycle time, what kind of error rates are you seeing, and the number of incidents that you're seeing. This should go beyond just ticket reporting, which is often about just servicing issues and moving on. These issues really need to be tracked systematically with a spreadsheet or software so you can also track improvements.

So we found that those who have done this have really found the results to be very illuminating and sometimes even very shocking, as Chris said earlier. In some cases, the teams find they have less capacity than you'd imagine because there's a ton of time spent on waste, 60% in their earlier example, right? So this can really have positive effects once you uncover these types of issues.

In one case, it led to a complete mindset change, right? People went from actually thinking that their job was to fix things and find workarounds, and instead, it changed to focusing on eliminating problems

00:40:00

entirely. We've also found that it can be scary when you do this, because at first, you're probably likely to see an increase in incidents as you identify and log them for the first time. So it may look like things are going the wrong way. Stabilization could take a few quarters, so you just need to be aware of that. It also may feel like bringing skeletons out of the closet, but we've seen that kind of honesty is really required if you want to improve.

So all the data you've collected here gives you the ammunition you need to set goals and track progress. It really shows the impact DataOps has on your operations. Without this information, it's almost like you're shooting in the dark, and it'll be much more difficult to prove the value of DataOps.

Okay, you can go to the next slide, Chris.

So there may be some stakeholders who do care about the factory. In that case, be prepared to give them a tour. But instead of pounding them with the technical details of DataOps, you really need to demystify the factory. Many business stakeholders think the data is magic, and they really have no idea what goes into generating these reports and dashboards that you give them. So this complexity is really invisible to most people.

So if they've given you an opening, this is a great opportunity to educate leaders about everything that can and will go wrong. Just make sure to stress the benefits of having accurate, trusted data for decision-making. Tell them there's a solution, but it will require their support. And we've found that most stakeholders are supportive, once they understand the problem, the solution, and the benefits.

Next slide. Yeah, but it's a minority who actually care how the factory runs, unfortunately. Right. So, another thing that comes up often is when educating stakeholders, DevOps can be a useful analogy, especially when working with IT. So if you have a strong DevOps culture at your organization, this analogy can help many come to that aha moment.

I will say, though, we've seen that this is something to approach with caution, right? I've heard a few interesting debates around this suggestion. Some think it really helps, others not so much, so you really need to understand your audience. In some cases, DataOps, as some people feel as an analogy, can oversimplify the concept and dumb down the power of DataOps, because it's not just CI/CD for data. And you don't want people to miss the point that DataOps involves complex systems and tools and teams.

The last thing you need is your IT team to come away thinking that they've got this, because they can just do it with their DevOps tools.

Okay, next slide

So another way we've heard some organizations have successfully sold DataOps is to piggyback on other DataOps, other data initiatives. So DataOps can be the foundation that enables many of these initiatives to be successful. So we've done whole webinars, and we have white papers on how DataOps enables your data mesh or a data fabric or a cloud migration. Without DataOps, many companies implement these initiatives, and they don't get their expected results, especially when it comes to agility. So DataOps can be the secret sauce that helps you get to success. Right now, we're working with a large pharma company that's implementing a data mesh, and DataOps has been a very critical part of that process and implementation.

The data mesh was really a differentiator in accelerating the adoption of DataOps at that organization. So a good success story there. Can we go to the next slide? So lastly, one of the biggest pushbacks we hear is that we're too busy fighting fires to focus on making any long-term change. We certainly get that the pressure to do things the wrong way is very strong.

Data organizations, as you know, don't always have the budget, schedule, or buy-in required for DataOps, especially when it's thought about as an enterprise-wide change. It can certainly feel overwhelming. So however, as we've said before, just start small. Focus on reducing errors in production pipelines. We call this production DataOps. Another name for it is data observability. This can be implemented by a small team or even one person with no process changes, so you really can get started on a small scale.

That said, we've seen that the best organizations are very strict and don't allow workarounds. So at one of our customers, a high-level data leader operates as the cop, not allowing workarounds by the team. And at the same time, he also gives visibility to senior leadership, telling

00:45:00

them, "You're not going to get this report, and here's why." So once senior leadership understands the problem, they're usually supportive of the solution. Over time, they'll see that things get better, and these incidents happen less frequently.

So those are a few things we've heard and seen success with when others have tried to sell DataOps at their organization. If you have anything that's worked, we'd certainly love to hear and let us know. Chris, do you have anything to add on this, or? No, it's been interesting to hear from our customers and their day-to-day differences. And it is a

change that they have to convince people to make, right? Because there's still a lot of techno-determinism in data and analytics. I buy a new tool, and magic will happen, or I go to the cloud, and magic will happen. And so if anything, DataOps is a set of principles that doesn't believe in magic, and what I've seen in the software industry is we've all stopped believing in magic, and we put a lot of work on the factory and the processes and the automations that our team work in.

And that belief is becoming more and more prevalent in data and analytic teams. But I think the assumption is that it's not there yet, and that's why you need to do ROI arguments and to convince people with us. And so it's been very helpful. And at the end, it really does come down to this future vision that you can work quickly, get things into production, run production with incredibly low errors, have a highly productive team where they're spending a lot of time on task, and show your worth by measuring. And if you do all these things, you actually end up with much happier customers. You end up getting promoted, and your career moves ahead, and your team is not so costly. And the principles that we've talked about here are actually in two books. Our manifesto, which is 18 points, our first cookbook, and then our second sort of recipes for DataOps transformation. And so if you want to learn more, come to our website or check out those books.

And that's it for our discussion today. And so in the remaining sort of 10 minutes here, we are open to answering some questions. And I don't know, Beth or Jessica, if you've been able to peruse any questions that we can answer. Yeah. So as Chris said, definitely now's the time to put your questions in the question box on the control panel, and we'll get through as many as we can.

So here is one to kick it off, Chris. Can DataOps be driven from a top-down approach?

Well, I think it is essential to get buy-in because we're talking about the way people work. And so I think having your chief data officer or your VP of analytics sort of bought in that we're going to work on agility and deployment cycle time, and in essence, DataOps is a big help. However, in a small team, if you are a leader and just your five people, you can start doing it yourself and getting a small team working on it.

And so I think it does help to have leadership buy-in to be able to make this change because it does have to go with kind of the operating model your team runs under. And so I think it's really important, but not absolutely essential. Okay. Along those lines, who are the people who tend to champion DataOps from within the organization from a leadership perspective?

So we categorize them in different ways. So there tends to be people who are Are senior, but not super senior, and they're true believers. And so maybe they're a data scientist or a data engineer or a director who've been around the block and have just hit their head against the wall so many times, like I did, and just want to find a better way to work. And they tend to be practitioners or managers who do it. And they tend to be change agents.

Maybe sometimes they're on the architecture committee, maybe sometimes they're just senior people, and they've sort of had enough and want to make a change. Other times it does come from senior management and the CIO talks to the CDO, and the CIO says, "Hey, we've, for the last two years, been doing this agile DevOps transformation.

What are you doing?" And the data person, the CDO goes, "We're doing AI." And they're like, "Yeah, it still takes you three months to deploy any new sequel into production." And so I think it can come from senior leadership championing at the C-suite level of saying, "We need to make changes." And so for us, what we look at where people are the case where they're having a lot of errors in production that are embarrassment, it takes them forever to change things, they're having productivity challenges, and they have this inkling that it's a people and process problem, and that they need to work on it.

00:50:00

Okay, great. Chris, what are some of the pros and cons of having a separate DataOps team that does the plumbing, or having the same team do both?

Well, I think it's essential to think about developing of testing and automation as a specialty that people can work on. And it also depends on the size of your team. But let's say you're a sufficiently decent size team. Having the role of a DataOps engineer to help automate deployment, develop automated testing, build environment and test data management, help the team of people who are doing the work, the data engineers and the data scientists.

In some ways, DataOps engineers, their customer is the data engineer or the data scientist. They want to make them successful by freeing their burden, by giving them the ability to make a change and judge its impact. The ability to make a change and see it to go into production quickly. And helping them with the tools and techniques of version control and all the other automation tools that go with it.

So I do think it's worth it to invest in DataOps engineering when the team is significant size. And I actually think it's a great career choice for people, and I think we're starting to see it. Certainly, I see the role of DataOps engineering being more and more customers having it. And I think perhaps not now, but I think if I look at the software industry, when it went from release engineers to DevOps engineers, and now DevOps engineers are very hard to hire and actually very expensive.

And actually very important. And the average ratio of DevOps engineers to software engineers, front-end, back-end, is 28% DevOps to the rest being actual people doing hands-on work. And I think what we want to get to in the data field is something like 10 or 20% of the people are involved in DataOps. So having those by title or having those by desire, I think is great.

Okay, great. Thanks. Here's a good question. What about the cost of implementing DataOps? Do you have pointers, ideas on what that might be?

So one way to think of it is just I like what Celgene did. Assume 20% of your time is going to be put on DataOps, and assume that by doing that, you're going to actually end up making your team more productive. So if you look at it from a time ratio, 10%, start off with 5% of your time, start off with 10% of your time focusing on the factory, the automation, the testing, the system that your team works in, and devote time to that.

And what you're going to find is the actual net effect of people being able to do more work is going to go up. And so by investing in the system, you're actually going to be able to do more work. And so from a cost standpoint, I look at it that way. I like what James Royster did at Celgene. Saves 20% of your time, enables a 36% increase in overall productivity.

Great. How would a DataOps approach complement a data and analytics team who are using Agile Scrum already? I think they go hand in hand. I think Agile, how the teams run and work, Agile and Scrum is great. I think Agile without the sort of automation doesn't work. And I think that's happened in software.

Agile without DevOps is not successful. And so you need the system, the automation to be able to make the Scrum and the Kanban and the team meetings work. And so I've seen the opposite, where people just do Agile, but they don't do DataOps, and they end up kind of falling back into Agile waterfall or Wagile, where it's sort of dev sprint, dev sprint, dev sprint, QA sprint, QA sprint, QA sprint.

And they use the words, but they're not actually getting the point, because the point is I've got to get value to my customer quickly and not kill my team, and develop an automated system to make all that work. So I think they go absolutely, essentially hand in hand. Great. A great answer. So how have you supported DataOps with ERP systems like SAP in place? What are the challenges?

Well, what we think of as DataOps, it's really about the process of getting insight from data. And so that insight could start with data in SAP and ERP systems, or it could start with CRM systems, or could start with supply chain systems, or could start with data outside the company. And every data system, whether it's SAP or not, has challenges with the quality of its data and the ability to turn that data into useful insight. And so that's what we focus on in DataOps.

It's the

00:55:00

integration of data across different data sources, ERP, CRM, into building useful artifacts to help people make better business decisions, and to drive efficiency or be able to drive revenue or compliance. And so I think from that standpoint, whether it's data that we're analyzing and the team needs to work in an agile sort of DataOps way.

Okay, great. So looks like we're up against the hour. Those are all of our questions, unless there's any last-minute questions anyone wants to add. But in the meantime, I want to--

Nope, I think we're good. I want to thank everyone for taking the time to join us today. Thank you, Chris, for that great presentation. I hope everyone got a lot out of it. So we will be sending out the recording of this webinar and the slides in the next 24 to 48 hours, so please be on the lookout for that in your email.

If you have any additional questions about the presentation, please don't hesitate to reach out to Chris or I directly. We'll send our emails in the email to you. And again, thanks all for attending and have a great afternoon and evening. Thank you, Beth.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

How do you calculate the cost of production data errors?

Multiply data team size by average data team salary by the number of errors by the percentage of time spent on rework and triage. That gives the team time cost. Two more costs sit on top of it and are usually larger: compliance risk from wrong data, and the opportunity cost to business users who lose trust in the numbers and stop acting on them.

How do you calculate the cost of failed deployments?

Cost per outage, mean time to restore, change fail rate, and deployment frequency together give a cost of downtime per year. A deploy that fails without an outage still costs the team's time. One that fails with a data error or downtime costs the value of the data system not being available for as long as the restore takes.

How much of a data team's time goes to errors rather than new work?

Gartner's 2020 figure is 22 percent of time on innovation and 78 percent on errors and manual execution. A customer survey of its own data team found over 60 percent of team time was not value-added work at all, once meetings, context switching, waiting, and rework were counted. That overhead is the pool DataOps is meant to reclaim.

How do you sell DataOps to internal stakeholders?

Do not sell the details of DataOps, because most stakeholders do not care how the sausage is made. Fix low-hanging fruit for fast wins, publish metrics so everyone is accountable, and give leadership visibility into what is broken and why. Establish a baseline of cycle time, error rates, and where the team's hours actually go before claiming an improvement, and give the reporting time to stabilize.

Is DataOps just DevOps for data?

No. DevOps is a useful way to frame the conversation with IT, but DataOps is not only a CI/CD problem. It spans complex toolchains and multiple teams, and the outcomes it targets are error rates in production, cycle time of change, and team productivity rather than build and release mechanics alone. Oversimplifying it to DevOps loses most of what it does.

What is a DataOps maturity model assessment used for?

It assesses organizational readiness and produces a roadmap, which makes it a way to find the first project worth doing. It identifies weaknesses across error rates, cycle time, collaboration, team culture, measurement, and customer happiness. Those weak spots are where a short pilot has the best chance of showing a measurable result.

Where to go next