On-Demand Webinar · 1 hr 6 min

Elevate Your Data Strategy with DataOps

A data strategy is about more than the next new tool. James Lupton, CTO and founder of Cynozure, joins Chris Bergh to work through where DataOps fits in a strategy, what its key components are, and practical examples of it in use. Recorded December 2020; updated August 2026.

Presented by Chris Bergh

What you'll learn 7 points
  • Cynozure defines a data strategy as a framework that lets an organization generate business value from data and analytics with pace and agility, and builds it on six pillars: vision and value, operating model, people and culture, technology and architecture, data governance, and roadmap. DataOps belongs inside that strategy rather than beside it.
  • The obstacles named are almost all organizational rather than technical: silos of people and silos of systems, poor collaboration inside the data team and between the data team and its customers, slow processes, inflexible architectures, technical debt, bureaucracy and politics, and access to data.
  • The team-effort slide sets a data team without DataOps at 97 percent of effort on data and analytics development against 3 percent on data operations, and a team with DataOps at 85 percent against 15 percent. The argument is that the smaller development share produces more, and that software engineering already runs a higher operations ratio than data does.
  • DataOps engineering builds five things: meta orchestration across tools from one view, automated testing in production and development, environment management, better sharing across roles and teams, and measurement of quality, deploys, tests, errors, and SLAs.
  • DevOps and workflow tools alone do not add up to DataOps. Beyond continuous integration and deployment, the list adds continuous self-service environments, continuous meta orchestration, and continuous testing and monitoring.
  • Adoption runs in six steps rather than a big bang: educate on the ideas, find a first project, establish a community of interest, demonstrate real value in a month or two, iterate onto more use cases, and expand into a staffed center of excellence with common tooling and metrics.
  • Success is measured in two areas: production metrics covering build and data provider errors, test result history, timings, SLAs, and model metrics; and team and project productivity metrics covering collaboration, deployment frequency, and test coverage.

Slides

32 slides

Transcript

Show chapters and dialogue 11,130 words

00:00:00

Good morning, good afternoon, and even good evening to some of you. Thanks for joining our webinar today. My name is Beth Beverly. I'm the VP of marketing at DataKitchen, and I'll be the host today. Our topic today is elevating your data strategy with DataOps. We're excited to have a special guest, James Lupton from Synojure, is here today to share his experiences and insight. Synojure is a data and analytics consultancy based out of London, and James is the co-founder and CTO.

He has over 10 years of experience in the delivery of data initiatives across a variety of different industries, both as a consultant and in-house. He's also been a software engineer, senior solution architect, and big data team lead. He'll be joined by Chris Bergh, who is DataKitchen's founder and CEO. Chris is the leader of the DataOps movement.

He has more than 25 years of research, software engineering, data analytics, and executive management experience. At various points in his career, he's been a COO, CTO, VP, and director of engineering. He's also the co-author of "The DataOps Cookbook" and "The DataOps Manifesto." So James will kick off the webinar today and talk us through why and how to include DataOps in your strategy, and then he'll hand it over to Chris to give us some tips on how to get started.

First, though, before I hand it over to James, just a few housekeeping items. You're all on mute. We'll save the last 15 minutes at the end of the webinar for Q&A. So if you have any questions, please enter them in the Q&A box on the webinar control panel, and we'll make sure we address all of those at the end.

And then lastly, the webinar is being recorded, so we'll send out a recording in the slides to all attendees shortly after the webinar. So with that, I think you can take it away, James.

Brilliant. Thank you, Beth. That was a great introduction. I always hate writing my own bio, so I'm going to have to get you to send that to me afterwards so I can steal that for future use. So hi, everybody. As Beth said, I'm James. I'm the CTO at Synojure, and I'm really looking forward to talking to you today about embedding DataOps in your data strategy.

I'm not going to talk at length about Synojure. Please check us out afterwards or link in with me if you'd like to talk further. But just to frame a little bit about what we do in the context of what we're talking about today, we help organizations go on this journey with data, helping them understand what they're going to do with data, what they're trying to achieve, and how they're actually going to build the teams, build the platforms, and get those all stood up to achieve those goals and visions. And that's the bit we want to talk to you about today, setting that data strategy. So I'm just going to take myself off webcam for the rest of this so you can see the screen properly, and then we'll jump in.

So we're here to talk about embedding DataOps in your data strategy. And really, the first thing we need to do, and the first thing we need to talk about is why do I even need to think about DataOps, and why do I even need to think about having a data strategy? What good are these going to do to me?

What kind of problem are they going to solve? And to do that, we first need to really zoom out and look at what's going on in the marketplace as a whole. And it's obviously going to come as no surprise to anyone that data and technology have been massively disruptive influences in the market for a long time now.

When we look at the kind of numbers on screen here, we can see that the biggest companies in the world are all putting data, are putting technology at the heart of what they do. And that disruption has continued in so many different industries in so many different ways. What that tells us, and what we'll look at a little bit more, is that in order to keep up, we need to be on top of this.

We need to be leveraging our data to the best of our ability, making sure that we are innovating, that we're adapting, and that we are meeting the needs of the market. And really, there's so many examples of this. You can't stand still. And I think a classic one from recent years, Blockbuster and Netflix. This is a story I'm sure we're all familiar with, where we've seen a massive market leading brand go into bankruptcy because of its failure to adapt to the needs going on.

And Netflix is one of those companies that's really demonstrated how it can use data to its advantage. If you really think back to where Netflix algorithm really started, it was about maximizing their small portfolios of

00:05:00

movies so they actually had the licenses to stream. They didn't have masses and masses of big hits and of big well-known names, certainly not compared to a competitor like Blockbuster. So the algorithm was about using the data they had on your viewing behaviors, combined with the viewing behaviors of everyone else, to surface to you the titles from there to make sure that you were engaging with their brand.

And we all know how that story has ended for them. They were really able to leverage data to adapt to their circumstances and drive new, innovative ways of getting these services out to people. More broadly, we are all obviously operating in a very difficult climate at the moment. Whether the country you're joining us from is in recession, whether the industry that you are specifically working in has been heavily hit.

We happen to do quite a bit with the arts industry, for example. And one of my clients, a big theater group in the UK, they've just, in response to our increasing tightening of regulations in the last couple of days, have had to cancel the pantomime And a number of more shows that they were launching.

And it's a very difficult climate for a lot of industries out there. But on the flip side, I've talked to other clients. I was talking to the CEO of a global recruitment company recently. She was talking to me about how both themselves and some of her peers that she's talking to are using the current climate, the current situation, as an opportunity to invest and innovate. And time and time again, in the past, whenever we've been in these kind of situations, it has been the companies that are able to innovate and to do something new and to adapt to the circumstances that have come out the back of this doing better than they ever have before.

And really, you need to be making sure that you are making use of all of your assets, data being one of the really key ones, to position yourself ready to bounce back and to take advantage of the opportunities and to adapt to your changing customer base and their changing behaviors. Similar to that Blockbuster example from earlier, how you respond in these kind of situations is crucial.

And Kodak and Fujifilm, another great example where one company focused down on their core business, decided that analog photos, that's where we are, that's what we do, that's what we need to focus on. Well, Fujifilm diversified and started looking at other investment opportunities, other areas, and obviously, take a look at the data on the right, that drop off and that speed of drop off with analog photos versus digital photos is so extreme and fast.

And the right data with the right mindset was something that people could have seen coming, could have adapted to, and could have been prepared for, to change their business around.

Beyond that, more recently, COVID has massively accelerated the digitization of businesses out there, with many of them having to really quite fundamentally change some of their core propositions in order to still reach their customers and to adapt to the changing needs that we've seen these customers want and have. Pret's shift to digital subscriptions for a company that was almost anti-digital isn't quite the right word, but very much focused on that bricks and mortar approach, intentionally not having a sort of website offering.

They've had to shift. And we've seen that across a huge amount of our clients. There's some great statistics out there that show the rise and the speed of digitization of new products and services. And fundamentally, data is at the heart of a lot of what these companies are doing. There's going to be so many businesses out there that we haven't seen adapt.

Part of the reason they've not been able to do that is because they've had fundamental underlying issues that they haven't chose to tackle that far. Maybe that's a data quality that means they just haven't got the insight. They can't digitize their products because their data's not in a shape, a form that they can put in front of customers that have meant they've really suffered, and they've been unable to tackle the changes in their industry.

There's a lot of lessons as well that we can learn from data and digital native companies that have been doing this for some time. All those brands that you can see on screen there are obviously made data a huge part of what they do and what they offer to their customers. Whether that's Spotify with 232 million monthly active users that they are selling

00:10:00

subscriptions to, that they're customizing and creating curated playlists for, whether that's the, what, $55 billion in 2018 that Facebook did on advertising revenue based on what it knew about you. These are companies that have really harnessed that core asset and made data something that is just at the heart of the business, whether that's about how they make decisions, whether that's about the kind of services they offer.

They've truly placed that asset at the heart of everything that they do. So in order to be like these companies, and a lot of our customers talk about, "I want to be the Amazon of X," construction or distribution of whatever it is. You have to be able to move quickly, and you have to be adaptable.

Now, those are easy things to say and easy things to

think that you can do, but they're fundamentally very difficult regardless of the size of your organization. And there's a whole load of things that can cause that challenge in there. And I imagine a lot of the things on the screen here is familiar to people and the problems that you're all facing in your day-to-day. So whether that is silos of people who aren't cooperating and collaborating, that might be within a data team because you have a lot of people that work in the data team, or it might be across data teams and business teams that are trying to get something done.

That might be silos of systems that mean it's difficult to integrate the data that you need together in order to have the kind of understanding of what's going on that you want. Technical debt is a major issue. Bureaucracy and politics that can get in the way. Access to data, whether that's the source data to build a platform and a data science algorithm on, or whether that's as an end user just being able to come in and write a simple query. There's so many different things that can get in the way and can slow you down, and eventually, while you might be able to make progress You're going to hit some walls, and those are going to be very difficult to move through until you actually sort out some of the underlying basics that sit behind all of these things.

So great. We need to be adaptable. We need to be able to move quickly. We need to be able to respond to our customers' changing needs and to the marketplace's changes. And there's lots of things that might get in the way. Well, how can DataOps help? And DataOps is very much focused on trying to move past those issues and trying to get teams to be more efficient and more effective.

So when we talk about DataOps, what we mean is this. It's an approach to shipping high-quality data products at pace. I think there's three things in that sentence that we really need to break down and to understand. So firstly, there's treating data as a product. And when we first kind of introduce that concept to people, they often ask, "What do you mean by a product?

What counts as a product in there? Is the data platform a product? Is it the data set? Is it the algorithm?" And what we tend to mean by a product is the data itself and the things that are created from it. So a data set, say, a sales transactional data set, that can be a product. A sales dashboard I create that lets my management team monitor our performance.

That's a data product. A data science algorithm that is recommending next best product to customers based on historic sales data. That's a data product. And those are all things that have a specific outcome, drive some value, are measurable, have features. There are characteristics that make them a product. The platform, the things I use to create those products are enablers for me to achieve those things. At the heart of that is value. If those products aren't adding value, they should be dropped.

The second piece of this then is about high quality and what we mean by a high-quality data product. Now, quality can come in a range of different forms, and we can mean different things by it. We could mean that with the things that we build, with the products that we create, they don't have many bugs.

They don't have many issues, or hopefully no bugs and no issues. The data quality in them is high, so data's accurate, it's timely. All of our six elements of high-quality data. But it also can mean that they drive value in the organization. That they've been driven by iterative development, that they've been refined to best serve the purpose that

00:15:00

they've been created to serve. So high quality is absolutely at the core of it, and it's really about quality in all aspects of what we're creating. The last bit of that then for us to take note of is pace. So DataOps is definitely there to help us move faster, to shorten our time to market, to help us need less resources to achieve the same results, to reduce our cycle time so that we can innovate products faster, that we can ship products more quickly, and generally just do everything that we need to as an organization, as a data team, when it comes to creating those high-quality data products as quickly as possible.

That concept is directly into response to that last slide that we were looking at with all those different kind of issues and many more that weren't listed on there. And in order to do that, in order to help us as a team ship these high-quality data products at pace, there's three elements of it that we think of when it comes to DataOps.

So at the top, modern software engineering approaches. So this is the technology side of it, because that's a crucial part of what we need to do when it comes to DataOps. And it borrows heavily from that DevOps world that we might be more familiar with and that's certainly more mature and more established. So environment management, infrastructure as code, source control, CI/CD pipelines, automated testing, production monitoring, all of these kind of tasks.

The idea being that we want to automate everything that we possibly can so that it's quick, so it's effective, and so that it helps drive that high quality nature by automating our test and check, by removing the chance for manual error, and so on. It's all about speeding up delivery and removing that error rate.

We use the term synergy modern software engineering approaches for this element of it, because while it borrows heavily from DevOps, there are some unique challenges that come with managing data that are not really common and not typically faced in the software engineering world that necessitate some different tools, some different approaches to how those kind of tools are deployed.

So we've got that technology stem of DataOps, how we automate everything to make it faster, to remove that need for manual intervention. The second part of it then is our process side of things, our agile delivery methodologies. So this is all about bringing over many of the behaviors that Agile

really champions. So constant innovation, releasing little and often,

the ways of prioritizing work or breaking down work packages, of tracking work as we go through. Now One of the key aspects of this and one of the key aspects of DataOps is a sort of statistical process control and lean manufacturing history that forms the intellectual heritage of the topic. Now, I tend to not use that description because many people might be familiar with those terms at a very top level or not at all.

So I try not to introduce too many new concepts in short sessions like this. But I think one of the key bits of that that we put under that agile delivery methodology piece is about measurement. So it's about measurement of all of your pipelines, so how they are performing, measuring the quality of them, that there's no issues, measuring how long things are taking to process, all that kind of stuff, but also measuring how we're performing as a team.

So are we trying enough POCs? How accurate are our estimations? How many bugs and errors have we had over the last 90 days? How many new products are we delivering? What's our customer satisfaction? All of these kind of things, but basically constantly measuring everything that we do from a build task through to a how we run the team task, so that we can focus that constant iteration and constant improvement on the areas that are causing us the biggest problem.

The last area then, value-focused cross-functional teams. This is about rethinking how we actually put teams together from capability-aligned verticals. So in a larger organization, that might be a data engineering team, a testing team, data science team, a BI team, a governance team, and so on, to teams that are pulled together with all of the skill sets needed to deliver one of those valuable data products where everybody is focused on actually the end result and not just their part of it.

00:20:00

We see so many teams get lost in, well, I've put some data in the data lake, and that might be great. It might be something that was needed. It might have been delivered really well, but in and of itself, that is not something that's added value to the organization. Not until that's then consumed and used in a dashboard to make a decision or put into some algorithm for a purpose.

Not until it's reached its end goal has that bit of work actually added value. And teams often get bogged down in thinking about those units of technical delivery rather than the ultimate goal they're trying to achieve. So this area is about rethinking the skills we need, focusing them around our customer, and thinking about what our customer wants, thinking about the value, thinking about the final output that we want.

And it's also one of those things that is, again, different in that data world from a DevOps world, that we have a vast range of skills involved, and not all of them are hardcore software engineers who are comfortable with command line tools and so on in order to deliver their work. So we've got these three elements that make up DataOps, the people that we have in the team, how we run the team, and the tools they use to achieve their goals.

All things that can help solve a range of problems. But when we try and explain the value to it, we need to think about it in a few different ways. And at a top level, we think anything that you do with data can be put into one of these three categories. Any of your data initiatives or use cases that you're delivering fit neatly under these headings, and some will obviously do more than one.

So when we think of DataOps specifically and what value is that going to drive, how is that going to help my organization? When you're trying to build your business case for investing in the tooling you need or the team you need and making the changes, what are you going to be describing? How are you going to be approaching that?

So risk reduction, great for DataOps to solve, and that driving high quality is really where that risk reduction piece comes in. Making sure that we don't have bugs in products that are released because we've got a proper automated testing process that ensures that errors are just not introduced, that the data quality that we're putting out is being constantly monitored and remaining high.

And there's some huge stats around this for your organizations, especially the bigger you are. Gartner have a couple of studies out in this that show on the low end, the average impact to an organization of poor data quality is $9.7 million a year. And they've got later research that shows it as high as $14 million a year. So having the right tooling in place to make sure that you catch data quality problems early, that you've got a real handle on how you do that, has a significant financial contribution to make to your organization.

Now, efficiency and cost saving is definitely heartland for

DataOps because of that pace bit and because of that automation bit that makes it up. We are moving faster, and we're doing it with less people. That just naturally drives efficiency and cost. And this is the easiest place to explain to people what benefit that this is going to drive to us.

Trying to tackle poor teams, poor silos, poor process. If you get that right when you're implementing DataOps, it can really help drive those things forward. And again, there's some great stats that can help bring that to life very practically as well. Because it's one thing to say we're going to be more efficient, that it's going to take us less time to do things.

But actually, when you try and explain to someone, well, what's that worth to our organization? Why is moving faster going to help? You've got studies like a big landmark study from the National Institute of Standards and Technologies in the US that showed it takes five times longer and thus costs five times more to fix a bug in production than it does One found early on in development.

The time it takes basically increases the closer you get to production through the various release stages. So when you think about the large-scale organizations that might be releasing a lot of features, a lot of new products, and so on, that starts to add up very quickly. I think Microsoft released something like 30,000 bugs every month.

So you can imagine the cost to an organization that scale if they catch those in production versus in dev is huge. The last bit, though, is about value driving. And I think this is

really where you want to explain why DataOps is important to people. So while at its core,

00:25:00

the things it contributes to most directly are risk reduction and efficiency and cost saving and how you do things. The point of doing those is so that you free up more time to do more value driving things. So by being more efficient, I can deliver 20 of my use cases this year instead of 10. And those 20, those extra 10 use cases that I'm realizing the value of earlier are worth however much they're worth to your organization.

And we're going to come back to that topic of use cases when we talk about data strategy in a moment. So it's really important in order to be able to sell the value of DataOps and understand what it's going to be worth to your organization to understand that wider picture of, well, what am I working on? What is the point of doing data in my organization?

What am I building and how is it going to help? And that's where we start to talk about data strategy. So data strategy is a framework that enables you to generate business value from data analytics with pace and agility.

It solves a number of different problems, and we're going to break some of the things it looks at down in a second. But it's there to give you a framework that drives business value. And it's very easy for data conversations to get lost in technicalities. But ultimately, this is all about how you align the work that you're doing in data to the overarching goals of your organization, whatever those things are about at the current time.

It's a narrative to help you explain how you're going to contribute to those things and the things that you need to do that. So it shows alignment to your business strategy, to your strategic goals, and outcomes and objectives so that everybody is clear that this is the company's goals, this is how data is going to contribute, and this is the things we're going to do to enable that contribution.

It helps you target your investment. It drives priorities and decisions around people, tools, technology, all of these kind of things. It's important to treat these things as a FUC. It should be a living tool that adapts and changes with your organization. And adaptability is, again, a core part of what a data strategy should be enabling in your organization.

So we break down a data strategy into six pillars. And you'll see this if you follow up or are familiar with any of our content, you would have seen us talk about these before. So six main things that make up a data strategy. Firstly, vision and value. So what is the vision for data in this organization? And specifically, what are the use cases that we are going to deliver, and what are they going to be worth to our business as a whole?

Crucial to getting the buy-in, the backing, the understanding, and for making decisions on all of the remaining pillars, because if you don't understand that value piece, if you don't understand what you're going after, what it's worth, when you might need to do it by, you're not going to be able to make the right decisions about skills you bring in, about technology you bring in, and so on.

The rest then are relatively self-explanatory. So people and culture is all about the skills that I have in my organization and how I organize them. Am I in a big multi-region, multi-department organization, and I need to create a federated model with teams embedded in different business lines? Or am I trying to centralize my data into a single central data team?

What kind of skills do I want to build in-house? Do I need to hire data scientists, or actually, do we just aspire to create some reporting? Do I want to own data engineering? All these kind of questions. Operating model is then about how you

gather the requirements that go with those use cases, prioritize them, track them, and ultimately deliver them through. It also is going to cover how are you going to manage and support what you're creating. How are you going to effectively manage the governance elements of that to make sure that relevant regulations, rules, and so on are applied, that you're looking after your data to a standard that is acceptable to your business.

Technology and architecture then is about making the broad directional decisions when it comes to your platform and the various bits of technology that will surround that, all the way from your core data store, reporting tools, ingestion tools, governance tools, and so on, so that you've got a clear strategy and approach to what you want to build,

00:30:00

what tools you're going to need, and the order that you're going to go about putting those together.

Data governance then very much a risk-based approach to saying, right, what do we need to do to control our data? And where do we want to sit from a heavy command and control kind of point of view to more of a handholding, self-governing kind of approach?

We like to think about data governance as how do you use data governance as an enabler rather than a horrible bureaucratic step that gets in the way, or just something that we have to do because, well, there's GDPR and we're required to do it, or whatever regulation that you might be operating in. And then lastly, the roadmap, bringing together the plan.

How do we go about doing all of these things? Because we can't do it all at once. And we'll come back to in a second a little bit more about that journey that you need to go on. I think the easiest way to think about how a data strategy works alongside and embeds with DataOps is to think about it like this really simple picture.

Once we've understood what makes DataOps, we've got our technology side, our modern software engineering approaches, we've got our agile delivery approaches to how we deliver our work, and we've got our

cross-functional teams organized around value, the people and how we organize them. It's clear to see the parallels that there are with our data strategy, and what a data strategy is covering. But it's about thinking not just in isolation of DataOps and the topics in DataOps, but thinking about how they fit more broadly with the decisions we need to make.

So, for example, technology. Well, DataOps has got a very specific focus in technology about how we automate our workloads, and enable our environments. But we also need to think about, well, fundamentally, am I going in the cloud? And if so, which vendor? And if so, what tools within that vendor should I be deploying?

And do I need different types of databases? Do I need real-time data ingestion? What tool set do I fundamentally need to enable what I want to do? Same with people, the skills I'm going to bring and how I organize them. Heavy overlap with the kind of topic of DataOps, and DataOps gives us a great approach to how we should be thinking about structuring those teams.

It's going to be highly effective at building a data team that can actually deliver. Again, same with process and operating model. We can, from the start, build in those concepts of agile, of iteration, and so on into our data strategy. So data strategy is also giving us that value and use case piece as well.

So why are we doing this? It's giving us the why, and it's helping us understand the value that DataOps could add, when it comes to delivering more of that why at pace. So there's such a tight synergy between these two things, and DataOps, just the concepts behind it neatly fit in and can help inform a lot of what you need to think about and do as part of a data strategy.

Now, a lot of what is in DataOps can feel quite mature. So it can be quite a scary topic to approach at the start, particularly that technology and automation side of things. And there's a huge amount that you have to tackle when not only implementing DataOps, but also looking at implementing a data strategy.

There's so many things, especially if you're starting, relatively new to this journey. It's a greenfield site, or indeed, you're throwing away a lot of legacy and trying to start again. There's so many things to do that can't be done overnight. So we have to think about how we prioritize and how we tackle that stuff.

So we talk about that in terms of this framework on screen, the level up framework. Through a lot of experience, both in consulting and in-house in places, and through our own research, we've seen everybody move through a very similar journey to this, where you are at the start having to establish the agenda, get the buy-in, set the direction, understand the value data's going to deliver.

Prove value to start to iterate out some MVPs, start to build initial skills, start to try and demonstrate some proof of concepts that are worth, that are delivering on the value that you described. Maybe building out MVP data platforms. Scale, taking that out more widely to the business, delivering more use cases for more people, building out your technical capabilities, building out your teams and so on.

Accelerate, where you're really refining what you're doing. You're getting faster at it, and you're just getting

00:35:00

quicker at bringing products to market. And finally, optimize where you've really made data at the heart of your business, and it's all about the small tweaks. It's all about how the 1% improvement in your data products and your outputs can have a big impact on your bottom line and your performance as an organization.

So as we go through this journey, this takes a long time to go through. We need to think about how we prioritize and what we do. DataOps and a lot of the concepts, particularly on the technology side in DataOps, really come to shine in the scale and accelerate phases here. That doesn't mean we don't want to be thinking about them early on, knowing that we're going to get there. So while it might not be something that is our priority to invest right away in, by building it into our data strategy, by understanding it in the concept of a roadmap, and by recognizing that eventually we are going to need it, and when we might need it, we can much better plan and bring that stuff in.

All of the prioritization should be done around prioritizing value, and prioritizing quality and pace. And what are the things that are going to deliver value? What are the things that are going to help you move faster to that in the long run? Now, some of the bits of DataOps are very easy to get going with.

Others, not so much. So technology can be a bigger barrier to entry, but definitely you can start thinking on day one about data as a product, about agile delivery approaches, about building those cross-functional ... teams. So start to think about which bits of those could I bring in and make a core of my data strategy, and which bits might wait till later.

Just to wrap up then from me before I hand over to Chris. If you're new to data strategy, if it's the first time you've been thinking about it, we've recently put together a data strategy scorecard, really quick survey that you can go and take. Only takes about four minutes to put in some straightforward yes or no answers to questions, and you can get a score on how you're performing across those various pillars that we just talked about and see where you've got kind of development needs and where you might need to be focusing.

But thank you very much for listening so far. I hope you've got some value out of that, and I will pass over to Chris. No, thank you, James. That was great. Thank you.

Okay. So now that we've all figured out you could have a jarring change in slide design, so, yeah. Thank you, James. And so I guess from my standpoint, I think of data strategy kind of first off from what you want as a leader. And in my experience in leading data and analytic teams, I guess, the first thing I really learned was

that I owned very little as a leader of a data and analytic team, and the thing I owned was the process people worked in. And so when things were going wrong, I was very tempted to find a data engineer or a data scientist I could blame and maybe fire and say it's all their fault. And I started to read Deming and about industrial process control, and I realized that most of the problems are not just individual problems.

I mean, sometimes they are, but mostly it's the system that people work in. And so as a leader, you own that system. You own the processes that people work in. And so, some of the key processes that DataOps works in is, is your production process, the journey that data takes from source to value. And for that, you really want to focus on lowering error rates, because errors are problematic. And then the second process that you own, really as a leader of a team, is your deployment process.

And as James said, iterative development, agility, cycle time really matters in terms of sort of maximizing the amount of value you create for your customers. And then you've got all these people in different parts of the organization, and some, as a leader, may work for you, some may not. You may have self-service analytics or a data science team.

And how do you actually get people to collaborate in development? And so you as a person who's running the data strategy, you own these processes, and that's your leadership lever to be able to make your team work better. You own the factory. You don't really own the assembly of an individual piece in the factory.

That's what your team does. And so to think about that, the next slide I want to share is, really think about the work that your team does and the amount of effort in terms of kind of two big buckets. And the first bucket is the stuff that they do to give value to your customers. They're developing some data transformation, some ETL. They're building a model.

They're doing a visualization. They're governing data. And so that really-- a lot of organizations feel pressure because, oh, I've got to get my stuff done. I've got people waiting. I've got a list of backlog of 100 things, and I want

00:40:00

all my effort to go to that. And this other stuff, these processes that I work in, is kind of not worthy. Maybe 3% of the time, maybe people do it on nights and weekends or someone who's got some passion, but as a leader, you don't really focus on it. And so I think this is very similar to my experience in leading software teams back twenty years ago in that I really wanted to get my software team to push features out, and I didn't spend much time on the system to help them push features out. And I think what the software industry has learned, and that's this graphic in the lower left, is that the ratio of DevOps to development, the ratio of stuff to help people build end features to the customers and the people who help own the processes and improve the processes is almost 20, sort of 23% and higher in some organizations.

There's a whole group of people who are talented, well-paid, who are just there to try and make the people who are doing the real work better. And so that's one of the things I want people to think about in their data strategy, is the ratio of the work that you do to make a system to help people deploy their work in order that they can get things done faster and better and happier.

And so that group is, in general, I think of as DataOps engineering. And what they're doing is kind of creating a superstructure for your team to work in, to create these development and deployment and collaboration processes. And that involves a couple of things. And the first is really about automated testing, both in development and in production. Observing your system, running the tests, making sure that you don't have problems before your customer sees them.

And the second is because you've got all these tools, right? As James talked about, you've got lots of people and lots of tools. You need some kind of meta-orchestration across all those tools, so you can kind of see what's happening on a single pane of glass. And then you need to manage the environments that people work in if they're going to have a process to go from dev to production, making sure that they do that quickly. And then it's really about how you get your teams to be productive.

And there's communication, there's tools like Jira. How do you get productivity and collaboration across teams and do that in a way that minimizes meetings? It's not just like, let's have better meetings. Can you build a superstructure that removes the needs for meetings, and have that superstructure handle the collaboration? And then finally, as a leader, you own these processes, and you want your company to be data-driven, so you should be data-driven about your own processes.

Development, deployment, collaboration, testing, errors, SLA. These things should be metrics that you measure as a leader. And it's part of that sort of 15% of the effort And so for a lot of companies, and if you think about these tasks that you do in DataOps engineering, that 15% I'm arguing that you should put into your data strategy. Most companies are kind of not doing very well at it. Their ability to deploy is slow.

They have a lot of production errors that cause a lot of chaos. And when you have your CEO report, we were just talking this morning with a transportation company, and they have a major CEO report, the morning report, and it had a problem, and everyone was running around. Was it the source data? Was it this part? Was it that part? And that just makes life kind of suck, actually, when you're chasing errors.

And how many error-free days have you had in producing your analytics? You go into a factory and they say how many cases where people haven't gotten hurt. Well, people do get hurt psychologically when you have production problems. And so these attributes, this sort of graphic equalizer, is low in a lot of organizations, and that's a challenge.

And the end effect of that is you actually have a much higher cost in productivity and much all-too-frequent unhappy customers. And so to me, this feels like an important part of your strategy. How can you push up these levers? How can you increase the cycle time, decrease the amount of errors, improve your collaboration without meetings, and measure your process? And the end result of that is that you can actually end up lowering cost and delivering more value to your customer. So it's a very upstream idea here in saying, stop focusing on, "Okay, I got to get the next feature out for my customer." Start building a system that makes it easier to deliver features in the first place.

And the end result is you're going to be able to do this faster and kind of with less problems. And so, the other part of this, and that's what we're going to talk about in our next part, is pushing up this graphic equalizer, cycle time and error rates. A lot of people, when you're doing your data strategy, who've been in the industry for a while, are going to not believe it. They're going to say, "If I do something fast, I'm going to have errors." And you can slide one up, but the other one goes down.

00:45:00

And they really, in their heart of hearts, are going to believe that that's impossible. And so, those sort of best-in-class companies that James talked about, the big companies, that they're able to actually push all these levers up all at the same time. And being able to go fast and not break things and not have fights between people and being able to measure all of it is something that you can do.

And in fact, working on them actually together at the same time helps you get that effect. And so, for me, I think one of the things that James started with that I found interesting was that a lot of people see that it's the big FANG companies that have all the valuation are good because they have a lot of data.

And I think that's true, but they're also good because they've learned to iterate quickly. They've learned to try things, measure, iterate, and improve. And their agility, in my mind, is almost as more important than their ability to have a lot of data. And it goes from the two-pizza teams at companies like Amazon to sort of Netflix's ongoing journey in developing their own data platform.

And so, I'm going to skip this next slide. And so that leads me into, okay, this sounds good. So let's say you do believe in your heart of hearts that this is possible and you believe that the analogy from DevOps and Agile is real, and you, as part of your data strategy, have to build a system to make this happen.

And so you have to do some DataOps engineering. And so what is that like? How do you take your team from sort of the basic part to how do you help them transform their organization? And I think

both James and my views are very similar, right? In that it's a series of steps. And the first part really, I think, begins with education on what is DataOps, the possibility of DataOps,

and talking with people because you will find people who will push back on this because, A, it's number one, in their heart of hearts, they think it's impossible. And so, but you're also going to find people who are very interested in it because it's a better way to work and they see the value.

And can you find a project and establish a community and start demonstrating projects iteratively, and then kind of expand on how that works. And so, if you're not already working in an agile way or you don't follow things like scaled agile or Spotify, that process of thinking, "Okay, our team is going to work in a more iterative way and more agile way.

How do I do that, and how do I get that percentage of time in my organization to do the DataOps engineering required to support it?" And then another part, again, from data strategy and the leadership perspective, is you have to measure how your team's doing. If the only thing that you really own as a leader is the processes people work in, it's a bit hypocritical that you aren't actually measuring your own processes.

And so you've got production processes. You've got hundreds of pipelines that are building data and analytics in batch or streaming every day. You've got data sources that are problematic. You've got pipelines that could be late. You need to measure that production process and instrument your pipeline for success. And are things on time? Are things on weight? Are the trains running on time?

It's very basic things, but I'm surprised at the level of organizations who don't have this visibility because you're running an assembly line or hundreds of assembly lines, and you want to be able to make sure that they write, that they all work. And in that assembly line, you've got lots of tools and lots of models and visualizations, and you want to make sure they're right before you get to the customer. And then the second part is you are running a team or a set of teams who are delivering insight, and that's very much like a software development process. How well are those teams collaborating?

How often are they updating? How fast are they deploying features? What's their burndown chart in the sprint plan? And there's a set of productivity metrics of your team that you want to look at because you want to say, "Okay, at the end of the day, if we do this work and believe that DataOps is a thing that we should do, and we invest in DataOps engineering, you should see that in increased feature velocity that you're delivering to the customer." And at the end of the day, this sort of report shows that you're awesome.

It shows that your data strategy that's been in fuse with DataOps is working. And so I think this measurement part, and again, my perspective just comes from my experience leading data and analytics teams. I think data strategy is an expression of leadership, and I think if you focus on the things that we talked about, the sort of cycle time and error rates and developing out DataOps engineering capability in your team, I think you're going to be successful.

And of course, I would be remiss, and my head of sales would also be remiss if it didn't say, "We've got some software to help you." And so, to do these, to instantiate these processes in

00:50:00

your organization for lowering error rates and observability, for decreasing cycle time, and improving on non-meeting collaboration. And lastly, if you're more interested just in DataOps in general, we have a very cool, newly revamped website that has a lot of great content on. We also have a book if you'd like to grab a PDF, or mail us and if you're interested, we can send you a physical book, and a manifesto. So there's a lot of resources from our website out there to understand what DataOps is.

So that's it in terms of my presentation. Beth, do we have any questions from the audience that James or I could answer? Yep. So we have a few minutes for questions now. So if you have any, please enter them into the question box here, and we'll get through as many as we can in the next few minutes.

So here's one to get us started. So who are the people within the companies that you work with that tend to champion DataOps initiatives from a leadership perspective? I'm sure you both have some thoughts on that. James, do you want to go first?

Yeah, sure. Let me come on to webcam so you can see me. Yeah, so who champions DataOps? I find that that tends to come from the more technically minded people, so often solution architects, enterprise architects, people who are in a technical leadership capacity in a team, or who are potentially something like a data product owner and are quite switched on when it comes to technology, so they recognize the opportunity to bring these things in. And it's becoming more and more common now for us to have these conversations.

I think as adoption of these kind of things increases in industry, its visibility becomes bigger. So we're starting to see that championing bleed out a little into roles that maybe weren't originally doing it, where that might be heads of data and CDOs who are saying, "Well, I've seen my friend over at so-and-so has said his team have put all this stuff in place, and we think we need to be doing that. Can you come and talk to me about what it means and what it looks like?" So definitely we get a lot of it from technical champions in teams, but that's bleeding out increasingly to senior data leadership who want to help their teams go faster.

Yeah, I'd agree. Chris, if you want to say something different, yeah. Yeah. It's best if you have leadership buy-in, and your chief data officer, chief analytic officer, your VP saying, "Okay, we're going to try to change this organization from a waterfall to an agile, and we're going to follow DataOps principles." And I think that sort of air cover from leadership really helps.

Now, you don't always need to have it, right? Because you could have a smaller team, a director saying, "We're going to do it ourselves," and that certainly works. And even individual contributors start following the principles themselves. But at least we've had the most success when it has bubbled its way up to the senior part of the leadership of the organization and saying, "Look, we're going to start. We need to change how we deliver." And that mimics what's happened in software where the sort of DevOps and agile mindset has penetrated the CIO suite, and bigger companies who've had success have been really kind of focused from the leadership down.

Great. Thank you both. Here's a great question. In your experience, what is the most common obstacle to establishing DataOps? Culture, values, other? James, do you want to start on that one? Yeah. I would say it's a combination. Well, at top level, it's investment, but that can be both investment from a money point of view, but it can also be an investment from a time point of view, because it would be remiss to say that day one, you immediately start seeing the speed increase.

It takes time to put these technologies in place. It takes time to change people's ways of working and bring in the relevant approaches. And that investment, both from a time perspective and a team, but also potentially in purchasing software, bringing new skills and so on, can be the biggest blocker to it all. The people are bought into, quite often, the outcomes and say, "That sounds brilliant." As Chris was saying when he was talking earlier, a lot of people think it's impossible or very difficult to get to that And part of that impossibility can be, well, the time it's going to take me or the money it's going to take me. So that's why we start to try and talk about, well, there's definite tangible value to this, and you can start seeing that value quite quickly. And it's significant if you do this

00:55:00

right and you put these things in place. But certainly, and I don't know whether this is just symptomatic of the type of people that are reaching out and thinking about this, have the right culture in place already, have the right mindset of thinking in these right terms. But yeah, I think that's the one that I tend to come across as the biggest barrier for people.

Yeah, I guess I could follow up on that. And so there's certain things I've seen in people's data strategy that makes me want to take a two-by-four and whack myself over the head. And I think in some ways, people focus their data strategy all on the wrong things. And so number one, there's a sort of a techno-fetishism that happens in data and analytic organizations.

And maybe if I get to the data lake and Hadoop, or maybe I get to the cloud, or maybe I buy Teradata or I buy tool X, everything will happen. And data strategy becomes an ancillary of the tool strategy where I sort of hope that in implementing this tool, all this magic is going to happen. And I find that incredibly frustrating because I just don't think that ever works. And then the second I think is there's a sort of trap that people get into where they focus on sort of deferred value instead of immediate value. And what does that mean?

It's like, okay, if I build this whole system, then magic's going to happen. And so a case in point of that is things like trying to do data valuation, where part of my data strategy is to say, I've got to go and do a data-- count the value of my data and say, "This data is worth a lot of money." And I just find that, again, intolerable because it's sort of deferring actually delivering value. And you're just sort of-- if your restaurant is producing crappy meals, you don't do an inventory in your storeroom to find out how to fix it.

You fix how you're making crappy meals. And so that's another thing that makes me want to get a two-by-four and whack myself in the head. And then the third part, just because I'm an engineer, I find a lot of data strategies have a data architecture component, and a lot of the data architectures have nice little boxes in and here's we're going to do streaming and here's our new tools.

But they're solely devoid of the process of changing that architecture. How do you deploy? How do you monitor? And so they build these sort of end state static production-centric architectures and they don't really build for change. And so I think data strategy, honestly, should have one line in, which is make your customer successful. And then the second line is know who your customer is.

And it's all about how do you get your team to focus on delivery of value. And these-- pardon my rant, but these sort of three things actually really-- I don't see them as signs of success. And that's why I think focus on value delivery and have your data's strategy all aligned behind that is where you need to be.

Okay, great. Thank you. That technology one rings so true for me, Chris, just to add on there, that the number of conversations I have where we're going to buy database X or visualization tool Y, and that's going to solve all our problems is-- and it's definitely not. It never has for anyone before. So you're definitely focusing on the wrong thing if you think you could just put in a new tool and that's going to solve everything as part of your data strategy. It's a very common mistake we come across all the time.

Yeah. I run a software company, right? And so we want people to buy our software. But on the other hand, the fundamental problem isn't that we don't have a great tool. Right? There are tools that are closed source, open source, that are expensive, cheap, that are self-serve and code driven. You can do everything that you need with the tools now.

And so the tools aren't going to provide you leverage. And it always reminds me back in-- I grew up in the Midwest in the '90s. And American Motors, which a car company, went out of business. And I remember reading an article about how they're going to put in new robots in their assembly line, and the robots are going to have American Motors make better cars. And you know what? That didn't help.

They got beat by Toyota, which eventually put in robots, but really because they focused on the process. And that's what their leadership was focused on. How can I make the process of my team better to make better cars? And so take the lesson of American Motors. Don't buy robots. Don't believe that the technology's going to change everything.

You've got to own these processes and you've got to focus on these things like cycle time and error rates and collaboration, and that's your leverage and your strategy.

Thank you both. So we are up against the hour, but maybe we can do one more question if everyone can stay for one more minute. We had two questions on the role of data analysts and insight teams. So maybe James, you can address these kind of together. The first was how do data product teams work alongside data insight teams? Do the first replace the second and insights is treated as a

01:00:00

product? And then someone else asking if analysts working in data teams is a big shift in the way they traditionally work compared to engineers, and on any tips for doing this successfully.

Lots to take in there. I think it's quite a telling question, I guess, about how the organization, whoever's asked that, is set up that those are-- they sound like they're different teams and insights is not sitting with data products. So, I'd start to ask, how has data products been defined there? What do we mean by data product in that organization?

There's a need in organizations to have capabilities grouped together still for career development, for the development of best practices and standards. You want people who work on the same kind of topic and the same kind of subject to come together. But where insights fits in alongside data engineering, alongside the rest of it, is they're all pieces of a puzzle. One of them alone doesn't deliver the value, doesn't create the product that's needed.

They all have a role to play. They all need to come together, and it doesn't necessarily mean that they need to be merged into the same team, for example, from an organizational structure point of view. They don't need to be the same department. They can be handled as virtual teams, for example, that says, "We're building a customer recommendation toolset, and we need some data for that.

We need some infrastructure, and we need some testing. We need some data science. We want some dashboards. Let's bring those skills together who are all working towards building this customer recommendation tool." And day to day, their workloads, their priorities, what they're focusing on is going to be driven by that product, by that product owner who owns that product.

I'm still reporting maybe to a line manager in an entirely different department. I think you can separate out the actual organizational structure piece from how you actually run day to day, how you collaborate, and how you prioritize your tasks and work together.

I don't know if I got both bits of the question covered in that answer, but that's my initial thought on that. Any thoughts to add on that, Chris, or... Oh, of course. There's a book called "Project to Product" by Mick Kirksten that's actually really good. And the way I think of it is,

if I work at a company, why am I successful? How do I feel good about going home at night? And I think the transformation is you don't want to feel good because you've got your three tickets and you've done your piece of ETL code or viz code, and then you're done. And okay, you sort of throw it over the fence and like, "Okay, I've done my task." And people feel successful. What you want to do is people not to feel successful when they've done their task, but to feel successful when their customer is successful. And so those are two different things, right?

That I've done my piece of the pie versus I'm trying to really focus on how to help my customer be successful, and I'm giving them. And that's what the product attitude is. My job is to make the customer successful, and that's a very different mindset. And what's come up is that there's a bunch of self-service tools that people can do data visualization or data science or data prep and I just got an email from Tableau saying, "Okay, you can do everything in Tableau.

You can do data science. You can do data prep. You can do data visualization." And self-service tools are actually really good, and people are very focused on trying to answer the next business question. And so that has its own challenges, right? How does that fit into a system that makes things that are repeatable and low error rates, et cetera? And so people are getting there, but it all comes down to this attitude of what's my job? My job is not just to take a ticket and close it.

My job is to know who my customer is and to work with a team of people to make sure that those people are successful every day and day in and day out. So it's that sort of service and servant mentality I think is very important.

Great. Well, thank you both. We are over time now, so I just want to thank everyone for taking the time to join us today. Thank you, Chris, and an extra thanks to James for joining us and sharing his insight. I'm sure everyone found this really useful. So to all attendees, we'll be sending out the recording of this presentation and the slides in the next 24 hours, so be on the lookout for that in your email. We'll also send out registration information for our next webinar, which is How to Build a Successful Cloud DataOps Program, and that'll take place in January, so we hope to see some of you there.

We didn't get through all the questions today. There are one or two we missed, so we'll follow up with you directly if we missed your question. If you have any additional questions, please don't hesitate to

01:05:00

reach out to us at DataKitchen or to James directly. James, how can folks reach you best? James@synertor.com? Yeah. James@synertor.co.uk. Okay. Or you can find me on LinkedIn as well and reach out to me on there. Okay, awesome. Well, thanks everyone again. I hope everyone has a great afternoon and also a great holiday.

Great. Thanks for having me. Thank you. Bye. Bye.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What is a data strategy?

Cynozure's definition is a framework that enables an organization to generate business value from data and analytics with pace and agility. Their model rests on six pillars: vision and value, operating model, people and culture, technology and architecture, data governance, and roadmap. The strategy is the thing that turns those pillars into a sequence of moves rather than a wish list.

How does DataOps fit into a data strategy?

DataOps sits inside the strategy as its delivery method, combining modern software engineering practice, agile delivery, and value-focused cross-functional teams. Its value shows up in three places a strategy already cares about: risk reduction, efficiency and cost saving, and driving value. Adding a DataOps engineering function is the concrete step.

How do you start DataOps without a big bang?

Six steps. Educate the organization on the ideas, find a first project that can demonstrate value, establish a community of interest and align with data and agile leaders, demonstrate real value in a project of a month or two, iterate onto more use cases, then expand into a staffed center of excellence with common infrastructure and metrics. Each step is meant to earn the next one.

Is DataOps just DevOps plus a workflow tool?

No. Continuous integration and deployment are only part of it. The full list adds continuous self-service environments, continuous meta orchestration across a diverse toolchain, and continuous testing and monitoring of data as well as code. DevOps and workflow tools on their own will not produce DataOps outcomes.

How do you measure a DataOps transformation?

Measure two areas. Production metrics cover the pulse of the current build, data provider error and success rates, test result history, timings, SLAs, and machine learning model metrics. Team and project productivity metrics cover collaboration, deployment frequency, and test coverage. The four focus areas those roll up into are cycle time, error-free days, team collaboration, and process measurement.

Where to go next