On-Demand Webinar · 1 hr 1 min

Practical DataOps: Delivering Agile Data Science at Scale

Harvinder Atwal, author of Practical DataOps and Group Data Director at MoneySuperMarket, on aligning the people, processes, and technology of an analytics organization with the rest of the company's goals, and the steps his team took at MoneySuperMarket. Recorded May 2020; updated August 2026.

What you'll learn 8 points
  • Harvinder Atwal of MoneySuperMarket opens with the state of the field: NewVantage Partners' 2020 survey found only 7.3 percent of organizations rate the state of their data and analytics as excellent, only 22 percent of companies see significant return from data science spending, and Gartner predicted that through 2022 only 20 percent of analytic insights would deliver business outcomes.
  • Technology matters less than teams assume. Research across 23,000 survey responses from more than 2,000 organizations by Jez Humble, Gene Kim and Nicole Forsgren found no significant correlation between system type and delivery performance, and the share of firms naming technology as the principal challenge to becoming data-driven fell from 19.1 percent in 2018 to 9.1 percent in 2020.
  • Lean process mapping of a data science delivery cycle exposed 57 days of value-adding work against 234 days of waiting, an efficiency of 20 percent. The waits were IT resource provisioning, software installation, data access and model recoding, not modeling.
  • Fixing those waits moved real numbers at MoneySuperMarket: lead time to make a new data item available for analytics went from 2.3 years in 2017 to three weeks in 2020, and the time to test a change to a machine learning model pipeline went from 11 hours to five minutes.
  • Data is a product, not an application by-product. DJ Patil's definition is used: a product that facilitates an end goal through the use of data. Work back from the impact and outcome you want, not forward from the data you have.
  • Conway's Law is not academic. Microsoft research found organizational structure predicted code quality better than code churn, code complexity, dependencies, test coverage or pre-release bugs, and nearly 60 percent of breakaway organizations use cross-functional teams against less than a third of everyone else.
  • Functional teams organized by expertise optimize for utilization of scarce talent; domain-oriented cross-functional teams optimize for speed. Cross-functional teams still form silos inside themselves unless members cross-skill, moving from I-shaped specialists toward T-shaped, Pi-shaped and M-shaped people.
  • Fitting a model is the easy part. Around it sit data governance, data quality, data security, test data management, version control, access control, team organization, stakeholder buy-in and outcome measurement, and DataOps is what makes those repeatable rather than heroic.

Slides

77 slides

Transcript

Show chapters and dialogue 9,033 words

00:00:00

Good afternoon, and good morning to those of you who are in later time zones. I know we have people joining us today from all over the world. Thanks for joining us for our webinar. My name's Beth Beverly, I'm the VP of marketing at DataKitchen, and I will be the host today. Today, our topic is Practical DataOps: Delivering Agile Data Science at Scale, and we have a special guest, Harvinder Atwal.

Harvinder is the author of a new book on DataOps, which was released late last year, called "Practical DataOps." He's also the group data director at MoneySuperMarket, a $2 billion UK-based consumer finance company. At MoneySuperMarket, he's responsible for the entire data life cycle, including everything from data acquisition to business intelligence. Previously, he led analytics teams at Dunhumby, Lloyds Banking Group, and British Airways. So before I hand it over to Harvinder today, I'd also like to cover a few housekeeping items. You're all on mute, so we'll use the last 15 minutes of the webinar to answer questions. So please enter your questions in the question box on the control panel, and we'll collect those and answer those at the end.

The webinar is also being recorded. We will email a recording to all participants, so be on the lookout for that in your email in the next day or so. And we'll also send a link to the slides in that same email. Lastly, as you may be aware, we will be giving away a copy of the book, "Practical DataOps," to 10 lucky webinar attendees.

So we will randomly choose the winners after the webinar, and we'll notify you via email if you're a winner. At that time, we'll coordinate with you the best way to get you your copy of the book. So with that, I will hand it over to Harvinder.

Thanks, Beth. Thank you, DataKitchen, for hosting me, and welcome, everyone. Good morning, good afternoon, good evening, depending where you are in the world. I'm going to talk about DataOps. I'm going to potentially be a little bit controversial in places, but don't worry, I've got data on my side. So

let's get going. So yeah, just a quick recap on who I am and what I do. So I've actually been working in analytics for about 25 years now. I studied in operational research a long time ago, which I guess was the data science of its day. I've done insight roles. I've done sort of data strategy roles.

I'm currently group data director at MoneySuperMarket. So as Beth mentioned, we're a sort of integrated data function. So we own the whole data supply chain all the way through from the point we capture data through to the users of data, whether that's data science, BI, product analytics, you name it.

So a little bit about MoneySuperMarket. We're actually a group. So we're the UK's most popular price comparison website. So our main brand is MoneySuperMarket itself. So customers come to us when they're looking to compare prices across insurance, credit cards, loans, energy providers, broadband, mobile phones. There's a vast array of products which people can find the cheapest deal on, or the best products on our site. We also have a similar brand called TravelSupermarket, which, surprise surprise, is all about travel products.

So on that site, you can find deals on hotels, flights, car hire, petrol stations. We also own a consumer affairs website, MoneySavingExpert, which is the UK's biggest consumer affairs website. And people go there for advice. And it also has a large forum, so it's also a reasonably sized social media platform as well. And then lastly, we have a B2B business, which is Decision Tech, and they power price comparison for other companies.

So we're actually very well trafficked. So we estimate that 80% of the UK online adult population at some point will visit one of our websites during the course of a year, which makes us more popular than Facebook in the UK. And what I really like about working at MoneySuperMarket is our mission. So our mission is to help customers save money and get a better deal.

So last year, we helped UK consumers save £2 billion on products on our website. So as part of what we do, we capture a lot of data. We use data in many different ways. So whether that's in product creation, so working with our product teams to test new features, new products. We also use data in our customer experience, so using a fair amount of machine learning to deliver personalized experiences to our customers.

00:05:00

And also helping the business make better decisions, so helping develop and deliver reporting.

So you guys will know that applications for data are in short supply. So there's many, many use cases. I've just got some up here. Predictive maintenance, all the way through to image detection, anomaly detection, and so on. So like many organizations, we're not short of use cases, we're not short of data. So you'd think we live in a world where data is making a massive difference to consumers, to organizations. But if you actually look at data about data, it's actually quite worrying. So it's actually saying something else.

It's actually saying that there's a big problem.

And this is just one survey. So this is a recent survey by NewVantage Partners, and they found that only 7.3% of organizations said that the state of their data and analytics was excellent.

And it's not the only survey that comes up with the same conclusion. So here's another one from one of the consultancies, which said that only 22% of companies are seeing significant return from their data science expenditures. So I don't know about you, but I find this worrying because there's massive value in using the data correctly, many different ways, from making people's healthcare better, making them financially better off, making the environment better off.

There's many ways we could be using data very effectively to make the world a better place, but it's just not happening. Now, if you ask the people at the other end, the users of data, so this is survey data from Kaggle, you also get a fairly depressing picture as well. So many people are having challenges within their organizations.

So you can see some of them here. A lot of people have complaints about dirty data, a lack of talent, a lack of support. So you'd think with all these problems, people would be talking about them, trying to solve them. And there's a whole industry out there devoted to data. A lot of conferences out there.

But when you actually look at what people are talking about, so I did this just sort of word count based on the synopsis of talks at London Strata, you find that they're talking about something completely different. They're really talking about technology. So let that sink in. So people are complaining that they don't have clean data, that they don't have support, and people are saying, "Let's talk about Spark and Kafka." So there's a major disconnect.

And the truth is, although technology gets all the focus, it's actually not as important as people think. So this is another massive survey, this time by Jess Humble and Gene Kim, who were founders of the DevOps movement. And they found no correlation whatsoever between system type and delivery performance. So they were expecting companies with greenfield sites, with the latest technologies to outperform those that were on legacy systems, and they did not find that at all. Again, this isn't a freak survey result.

Again, here's NewVantage Partners. So they asked a question around the barriers and challenges within organizations that stop them becoming data-driven. And actually, the number of people mentioning technology is declining. So fewer than 10% of people say that actually their biggest challenge is technology. The technology's actually keeping up. It's actually people and process that is the problem.

And then there's another unhealthy obsession out there apart from technology, and that's an obsession with algorithms. So many people coming into this industry get excited by machine learning, deep learning, and especially about spending a lot of time tuning those algorithms to make them as accurate as possible. But that's a very bizarre measure of success, and it's pretty counterproductive.

And actually, what's happening is that auto ML is improving all the time. So this is a chart showing the performance of Google's auto ML on various Kaggle competitions, and you can see it scores pretty well. So apart, I appreciate there are areas where extreme accuracy on machine learning problems is important, so I'd say fraud detection, healthcare, and so on. But in the vast majority of cases, 10 good enough models are better than

00:10:00

one perfect model.

And actually, machine learning was never really the problem. So this is a famous-- You might have seen this picture before. It's from a famous Google paper called "The Hidden Technical Debt of Machine Learning." And they said, actually, machine learning code is a very small part of what you need to worry about. There's lots of other things.

But actually, I think they missed out a lot more. There's a lot more around this in terms of governance, data quality, stakeholders that really need to be considered.

So there's a little bit of their secret out there, which is that people want to talk about algorithms, they want to talk about technology. But no one really wants to talk about the problems, the real problems, which are all to do with people, process, and culture.

So this challenge in large part arises because not many people in the organization actually understand data and how to use it properly. So if you look at people in the organization who understand what you do as a data person, there's you, there's your teammates, there's hopefully your manager. Then there's people who don't understand what you do, and that's pretty much everyone else.

And the challenge often starts right at the very top of organizations. So businesses, leaders that get caught up in the hype around AI. So they think that they need to hold data because large tech companies have a lot of data. They need to hire some PhD data scientists, and then some magic s**t happens, and lots of money will flow.

Money does flow, but it flows in the opposite direction because the technologies aren't cheap and data scientists aren't cheap.

Then there's lots of legacy thinking out there. So for many, many years, people have been told that they need to deliver actionable insight. But the problem is, this might've been relevant when I started my career when there was nothing much else you could do with data. You analyze it, you put it into a document, you put it in front of someone as a recommendation, and they decided either to act on the recommendation or not.

But the world has completely changed now. Data can directly action

via customer touchpoints, consumer experience. So Google, Amazon, they can make millions, billions of recommendations using data and make far more money doing that than they ever could using the same data for insight.

So what happens is there's a whole lot of insight that goes on that is actually never used. So Gartner puts only 20% of analytical insights actually deliver business outcomes. It's just not a scalable way of using data.

So the first step really before we get into the technicalities is really to make sure you're focusing on the right thing. So most organizations, they think in terms of you get resources, people, money, technology. You give them activity, so projects, and they produce an output. But really, this is the program logic model. That's not where success arises.

Just because you're producing output does not mean you're actually creating any sort of beneficial impact or outcome. So what we do is actually flip things on its head. So our starting point at Money2Market is actually we want to understand the impact we want to make. So our starting point is to make sure our objectives are aligned with the rest of the organization and strategy.

We have OKRs, so we make sure our OKRs, objective key results, are very much aligned with the rest of the organization.

So there's a lot of legacy thinking to overcome. And I just want to use a few analogies here. So this is Edison's first electricity generator. He installed it in Lower Manhattan in the 1880s. And for the next 30, 40 years, nothing happened. The world didn't all switch to electricity. Factories still ran on steam.

Streetlamps were still lit by gas. And that's because the move from steam to electricity required a complete change in architectural factories. So rather than a single steam engine running belts and rods throughout a multi-story factory, it required a completely different architecture with local motors. And no one was willing to throw away what they had unless they had to.

And that change didn't come until the 1920s. But when it did, it had a massive economic impact. So we're really at that stage where we're switching from the steam age of data analytics to the electricity age. So you might not remember that electricity dynamo, but you might, if you're old enough like me, remember one of these.

There's only one phone in the house, and it wasn't mobile.

00:15:00

So actually, one of the first uses of computing was actually in telephone companies for billing. So companies had to send out bills. That required a lot of calculations, so they installed computers. And because computers were expensive, there was a lot of rigorous systems architecture, a lot of upfront development requirements, gathering, testing to make sure that what you built was actually going to be of use because you could not waste any of that computing resource.

There was some reporting capability added, but it was pretty basic.

And once that data was used, well, actually there wasn't much value in it. So once that bill had been sent out and a report created, what else were you going to do with that data? So you just archived it.

And then later on, people discovered actually, maybe there's some value in keeping some of that data. So data warehouses were created, but they still had that same development methodology applied, which was a lot of upfront design, a lot of upfront requirements gathering, very slow development processes, a lot of rigor and testing, which meant that it was really slow to basically capture new data and make it available to people.

But the world has completely changed where we're not in that steam age anymore. So if you look at a telco today, it's completely different. So it'll have multiple sources of data. It'll be coming in from various platforms, from its mobile network, from its landline network, from its call centers, from social media, from its website. And it'll be coming in multiple formats, so there'll be structured data, semi-structured data, there'll be free text, unstructured data, video, you name it. And they'll end up in multiple data silos.

But thankfully, these days, storage can be achieved, so it's not as much of a challenge as it was in the early days. And that data needs to be shared and combined. So there's many analytical use cases, but that data may reside in lots of different places. So for instance, if you're trying to find marketing effectiveness, you may need to get your product revenue lines from one database, your marketing costs from another.

And data has now become pretty critical for any organization in any decision-making. So whether you're making decisions around your pricing, your tariff, your marketing, your manpower planning, you need to use data. And it becomes a source of competitive advantage. So what you have now is data is no longer that IT application byproduct that is just thrown away and archived. It's actually something which is incredibly useful in its own right, and it requires a completely different way of using data to create value than was the case 20, 25 years ago.

So critically, data is no longer that byproduct. It is a product in its own right, and it has to be treated as a product. So number two is you have to think in terms of products, not project, when it comes to data. So data products, they need maintenance, they need iteration, and they need a team who maintain them

So if data is a product, then data analytics is now manufacturing. It's complex manufacturing, incredibly complex manufacturing. So if you think of what you need to essentially create data analytics and create data products, you need data storage layer, databases, you need compute infrastructure, the ability to query that data. You need development tools on top.

You need orchestration monitoring, reproducibility deployment tools on top of that. You need very strong data management capability, so metadata management, data cataloging, data lineage tracking. You need data analytics, specialist software, and you need tools for data integration, data processing, so ETL tools, stream processing, master data management.

And there are actually two processes at work here. So the first one is there's a production system which takes the raw data and puts it into a data product by processing. And there's a product development process, which is creating those data products, but also iterating and improving them over time.

Now, data isn't the only domain that has this problem of dealing with complexity. There are two other areas where we can learn from. So one is manufacturing, where lean thinking comes from. And the other is software engineering, where agile and DevOps originated. So we can borrow a lot of the practices and concepts from these areas and apply them to data analytics, and that's where DataOps comes in.

So number three is you have to apply lean thinking to what you do.

00:20:00

And lean thinking comes from Toyota's just-in-time philosophy, and it's a philosophy for the absolute elimination of waste. And in analytics and when dealing with data, waste can come from many sources. So, from rework in particular, from errors, and so on. Now, the challenge here is that data people use data to improve everything except their own data cycle. So we're always looking to see how data can affect other people in their decision-making.

But very rarely do people actually apply data analytics to their own processes. So one of the first things that I recommend is just do some processes mapping, just to identify where some of the waste is in some of your process and data product development. So if you're not familiar with lean process mapping, it's fairly straightforward. So you just draw a horizontal line, which is time, and then you plot your activities, in terms of time, on that graph. Anything above the horizontal line is actual value-adding work.

Anything below is waste. So waiting around is waste. Any sort of rework is waste. Any correcting of error is waste. So here we can see there's a process whereby someone's trying to develop a model from initial design through to deployment. And actually, they're only working on that 57 days. The other 234 days are actually waiting around for other seams.

So that's not an efficient process.

It can be enlightening and eye-opening the first time you do this. So don't underestimate just very simple data visualization of your processes. So this is a real-life example of making data analytics available for analytics at MoneySuperMarket. So three years ago, it's not a typo, the lead time is 3.3 years. Admittedly, a lot of that is just people waiting, or requests waiting in backlogs.

But nevertheless, it's still waiting. We've now got that down to three weeks just by focusing on the process. The team has done a fantastic job of working through the bottlenecks, improving all their development processes, and testing and creating a hell of a lot more automation to make data available much more quickly for analytics teams.

So that production line requires deployment and orchestration of data pipelines, and they are not simple. So you may have hundreds of pipelines. So you may have data coming from many different sources in many different formats. You may be using many tools in your pipelines. So Airflow DAGs can have many nodes, and there may be multiple data products. So there's a lot of complexity here.

And the challenge with working with data is the data, basically. So, unlike software development, where if you code business logic up correctly, you can rely on it. So for instance, if you code up a code for a calculator and you test it and it works, it'll work correctly in 100 years' time. The challenge with data is that it's always changing.

So you may create a data pipeline, it may work, and the very next day it will break because of some change that you did not anticipate upstream. So it could be missing data, it could be data in a completely different format to what you expected. It could be unexpected schema change. There's many reasons.

So that's a major challenge. So what's required is, so this is number four, is to put a lot of monitoring and testing in place. So you need to trust that your pipelines are healthy. There's many different types of testing you can do. So there's a bit of an endless challenge actually, because there's always something new coming up that you didn't anticipate.

We actually have a specialist team whose job it is to do a lot of the automation of our monitoring and sort of level one and level two support, and automate some of the incident reporting so we can have as much data as possible to basically reduce the time to recovery.

One of the other sort of points of conflict in organizations is data. So as a data person, I want access to all the data. But there are good reasons why you shouldn't be allowed access to all the data. So for instance, we capture what personal data is. It's not very good practice for people to have access unless they absolutely need to.

So a lot of information security, data security teams will want to lock everything down, which does create a major challenge. As a data consumer, you want to access data, you want to combine it, you want to understand how you can use it to create value. But there are ways to get around this. There are lots of things that you can do.

There's identity and access management policies, there's encryption, there's data

00:25:00

masking. So, there's a lot that as a data consumer, as a data analyst, a data scientist, you can ask to be put in place or you can put in place yourself so that you can work securely with data. It's not a binary choice between having data and not having it.

So all this still has to be delivered. So all products and projects have a life cycle to be managed. So it starts with a concept, an idea. You assess it for feasibility. If it's feasible, and it has an investment case, then there's an inception phase where you engage people, you set things up. Then you start developing, and that process is usually iterative.

And then you have a deployment process. Your product goes into production. It needs monitoring, fixes have to happen, and then eventually, the product may be retired. So there's lots of ways that you can manage that life cycle. So for me, number six is you want to deliver outcomes in a way that allows you to be adaptable. So no one knows upfront whether their great idea is going to work with their consumers or customers, whether that's internal stakeholders for, let's say, a business report, or a recommendation, or external consumers, who are going to be served with a recommendation for our machine learning models.

So we have to be adaptable. So many organizations choose Agile as their frameworks for managing product delivery. And Agile is extremely badly misunderstood, especially given how long it's been around. Many people just assume that Scrum is Agile. However, that's not the case. There are actually many Agile frameworks. We don't typically specify what teams should use.

But we've kind of settled on Kanban in most places. And the reason is that Scrum works well if you're, say, a software developer, and you have reasonable certainty over what you can deliver in a sprint. But with data, there's always a lot of uncertainty, around what you're actually going to find in the data or what you might actually be able to create at the end of a two-week sprint. So we find that Kanban is actually a better approach for the type of work that we do. And Scrum also requires larger cross-functional teams, which typically don't exist in a lot of analytical functions.

But the main thing is that you, rather than adopting many of these practices superficially, that you actually pay attention to the Agile Manifesto and what it says in terms of values. So it really is about people over processes and tools. It's about make sure that you get your work in front of your stakeholders very quickly and get their feedback.

It's being prepared for change, expecting things to go wrong, and it's about collaboration with your stakeholders rather than working in a corner and presenting something that they didn't really want. Now, Agile would be nothing if you couldn't actually put those iterative changes into production and get that feedback. So there's no point working in two-week sprints if it takes you two months to put something in production. This is where DevOps comes in.

So number seven is you have to embrace development best practices, especially DevOps best practices and data analytics. So DevOps came about because of the conflict between developers who wanted to make changes, wanted to make improvements, and ops teams who had to look after that development in production, who really wanted things to be stable. The minimum of changes. They didn't want to be woken up at 3:00 a.m.

in the morning to resolve a problem.

So one of the first steps you can do is put in automated reproducibility. So this is all about version control. So version control of code has many, many benefits. It creates an archive. It allows you to go back to known good versions of code. Separation of your dev from your test and production code.

Allows your people to review each other's code very easily. And so there's many benefits.

We use Git big bucket, but obviously there's other version control software out there. But it's not just code that needs to be reproducible, but it's entire environments. So there's many ways you could create reproducible environments. What you want is to know that the code you're going to run in your development environment will run the same way in your test and production environments.

So you could recreate entire operating systems with VMs. That's less popular these days. There's package managers or things like Conda, use those more on data science. Docker for reproducible containers.

00:30:00

And then there's configuration orchestration management software like Terraform platform, which makes infrastructure reproducible. So typically on Google Cloud, we set along using Terraform for the infrastructure and Docker for reproducible software environments.

So the idea is that you want to automate as much of your deployment pipeline as you can. So from your dev through to production. So the best place starts with continuous integration. So that's you take your development code in. Once you commit it to your repository, it should trigger some basic unit tests. If it passes those, it should trigger integration tests in a test environment.

But you can go further than that. So you can also automate the acceptance stage. So, that's your acceptance testing. And then typically, your deployment to production is still a manual step. But you could even automate that and go all the way to continuous deployment. But we're not there yet. And you can have some massive

benefits from doing this. So this is just an example of some pipelines for machine learning. So we've managed to reduce the time it takes to develop and test, and run these pipelines quite dramatically. So one of the first things to do is to break up that monolithic code into sequential steps. That allows each to be developed independently and in parallel, and also to be run in a parallel way, in Docker containers in our case.

And we've been able to reduce the time it takes to make and test changes in our pipeline from 11 hours to five minutes. So you can make significant gains by doing this.

But the thing you have to watch out for is, as I mentioned earlier, is not just your code that's changing, but data is changing all the time. So test data management, number eight, is really important. So what you want to do is change one thing at a time. So if you're changing your code, you want to be using the same test data so that you know any changes you see are a result of changes in code and not test data.

So a lot of cases for us, it's quite straightforward because our data is event-based, we can create, reproduce a snapshot of state of that data relatively easily. In some other environments a little bit harder because we're using sample data, and we're using obfuscated personal data, so it's a little bit harder to test. But this is something that you have to take very special care of.

I did want to spend some time talking about organization as well, and people. So some of you may be familiar with Conway's Law, which says that an organization's output is basically a reflection of its internal communication lines. And Conway's Law is an academic. This is some research from Microsoft. So what they did was they looked at why some teams were better at deploying code successfully, more than others. And there was so many things they looked at, organization structure, code complexity, bugs, test coverage, code coverage, and so on.

What they found was actually the biggest predictor was the organizational structure of the teams. So organizational structure is one of the most important things. Probably more important than technology used, to be honest. So historically, teams have been organized quite functionally by skill. So you'd have a team of data scientists, team of data engineers, potentially reporting into different functional reporting lines.

So this structure kind of comes about because it's there to minimize cost effectively, to maximize resource utilization. So if you can imagine, you don't want an architect in every single team because that will be quite expensive. And what you want is a central pool of architects, and to be assigning them projects. So as soon as they become free, there's a queue of projects for them to work on.

So the good thing is, yeah, it helps reduce costs, reduces duplication, keeps everyone busy. But the downside is that it makes everything much, much slower. So what happens is that your data has to flow, and your development has to flow amongst many functional teams, who are working on many different projects at the same time.

And what happens is you may finish your work, but your output ends up in the

00:35:00

backlog of another team who may not get around to it for weeks, months. There's all kinds of battles around escalations, around prioritization. My work's more important than this piece of work. And everything just becomes bogged down and incredibly slow.

Here's a little bit of research this time from McKinsey, and they looked at what makes some analytics teams excel, versus the vast majority. One of the things they found was that actually cross-functional collaborative agile teams were more common in successful teams than others. So 60% of those that they call breakaway organizations had cross-functional teams as opposed to functional teams.

So, this route we're having to go down. But even if you try to create cross-functional teams, what you'll find is that there are certain special skills or personas I call them, because every organization has different job titles, slightly different skilled people, who just won't fit in a cross-functional team. So they will always form part of a central resource pool.

So even if you create cross-functional teams, you'll still have another problem,

which is that you'll still end up with silos in a team. So you may break down silos between teams, but you will have a different kind of silo within teams,

and that is a skill silo So you'll still end up with people working in a sequential way. So you may have a data engineer who's helping build data pipelines, which bring the data for data scientists to work on, but you then may have a specialist QA, who then has to test the work, and then you may have, let's say, an MLOps engineer who's then responsible for deploying that model.

What you want to do is basically cross-skill people so that there's no bottleneck within the team.

So historically, within teams, you typically have two kinds of people. You've had people who are what are called dash-shaped, so they're generalists. So no deep knowledge of anything but a broad range of knowledge over a lot of things. So typically, they'd be your managers. And then you have I-shaped people. So they are very specialist, and they have deep expertise in one thing.

And what you want to do is you want to create what's in Agile terminology called generalizing specialists or T-shaped people. So they're capable in a lot of things. So not necessarily specialists in a lot of things, but they understand a bit of data engineering, a bit of data analytics, a bit of data science, a bit of testing, and so on.

But they have deep knowledge in one area. And you want to increase their skills. So you want to get them to gain knowledge in other areas. So this is typically a lot easier with some people. So getting people who have data analytics or data engineering backgrounds interested in data science is usually a lot easier than vice versa.

And ultimately, what you want is poly-skilled people. So then there's another question around orientation of teams. So what do you want their remit to be effectively? So what we have settled on is cross-functional domain teams. So for us, a domain might be one of our brands. It might be customer, so anything sort of horizontal that touches the customer, like marketing or a touchpoint, like our mobile app would be a domain. Some of our vertical products would be another domain.

And then to tie them together, you have to have a center of excellence. Otherwise, you end up with teams who, again, are siloed, and you then end up with inconsistent patterns, inconsistent hiring. So you need something to tie them together. And then we have a central data management platform team, whose job it is to make sure that those domain teams, they have the tools, the platforms, and the data to be able to do their job effectively.

And this arrangement is optimized for speed, so we want to be able to work as quickly as possible.

And we do that through trying to gain enough, or allowing our self-service access so that no one is a blocker, and that people can deploy their own work into a production as easy as possible. So that was people. Then there's the aspect, which is you have to measure and act on feedback. So I spoke a little bit about monitoring, but that's not the only measurement that you need to do.

So I've stolen this from Matt Phillips, so I can't

00:40:00

claim credit for this. So he says that there are two dimensions of measurement and feedback. So, the first is the viewpoint. So one is the internal viewpoint of a team, and the other is the viewpoint of the customer. And then they have separate concerns. So one is the product itself, and the other is the service around that product.

So if you imagine a restaurant, a customer is interested in how tasty the food is, whereas the restaurant team will be interested in how fresh the ingredients are for that product. So same product, but they're interested in two slightly different aspects. So in terms of a team, they will be interested in monitoring that product to make sure that it's healthy, that nothing is broken.

In terms of the service, they will be interested in their own internal processes. So we do regular retrospectives. We look back at the work we have done, to see what went well, what we could improve on. In terms of customer, we need to measure the benefit. So I'm really surprised in the 21st century that people just stop with the output, and they don't measure the impacts that they're having.

Because without that, you don't know whether you're working on the right thing. You don't know what you could do to improve. And then finally, there's a service measurement from the customer's viewpoint, which is, are your team delivering? Are they working on the right things? Are they delivering for the organization? So there are the four aspects of measurement and feedback that you should be doing.

The final thing is really talking about tools and technology. I deliberately left this to last because I said it is important. The right tools will make a difference, but it's actually a lot less important than people think. So just as DevOps is more than using Chef, Puppet, Ansible, configuration management software, DataOps is more than tools itself.

So long gone are the days when you could go up to one or two vendors and say, "Build our analytics capability for us." So go to Oracle or IBM or Teradata and SAS and say, "We want to build a brand new analytics platform. What can we buy from you?" Those days are long gone.

So if you want to be best in class, you have to buy best-in-breed products, and you have to buy products that fit a niche and bring them together. That brings some complexity. And then in the DataOps world, there is specialist software. So DataKitchen, obviously, have a great tool, great software in this space. There are others.

And one of the challenges here is that you can get these tools to work together, but as yet, there's not a huge amount of interoperability when it comes to passing metadata between them. So there are a few challenges here around, for instance, if you're doing a lot of transformation of data, being able to pass metadata around that from the source all the way through to the final end data product to understand what's actually happened.

But hopefully, someone will recognize those challenges and come up with a solution for them. But I do want to say that it is complex. There's no simple solution. That you are going to have to bring together a lot of technology, which is not bad necessarily. You get to use what's best for the job as opposed to what's best on average.

And that complexity can be quite difficult to keep on top of. So one of the things that we do is we have a tech radar. So again, I've stolen this from ThoughtWorks, the IT consultancy. They published one of these, but I've got one which is specifically the sort of DataOps area. This is just an example of one.

This kind of helps you manage what technologies are there, what you might want to assess and adopt. So I recommend, if you're starting down this path of DataOps and looking at technology, that this is one of your starting points, which is to maybe draw up one of these and see what you might be interested in assessing.

So finally, before we go on to the questions, apologies a bit of shameless self-promotion. I can only really scratch the surface in 45 minutes. There's a book I've written, as Beth mentioned. Hopefully, some of you will be lucky enough to win a copy. If not, I hope you found this talk useful enough that you want to find out more.

There's a lot more detail in my book, if you are interested enough to go out and buy it. So that was it for me. I guess, over to you and your

00:45:00

questions. I'll be happy to answer them. Great. Thank you so much, Harvinder. That was a lot of great practical advice. Just to let everyone know, I did put a link to Harvinder's book on Amazon in the chat, so you can follow that and get a copy of the book yourself as well. So I'm just going to dig into the questions here, and I encourage everyone to please continue to enter your questions into the question box, and we'll try to get through as many of those as we can in the next few minutes.

So here's the first one for you, Harvinder. Any tips on documenting data assets to enable fast access and easy reuse? I find we spend a lot of time rewriting existing code and duplicating existing models because we don't know what we have.

Yeah. So that's quite a common challenge. So we actually use something called Domino Data Labs, which is a data science platform. And it has a lot of collaboration tools in, which help you easily find work that other people have done. So for instance, if you're going to be creating a piece of analysis or a model rather than start from scratch, your first point will be to go search inside their repository to see what other people have done.

We do a lot of documentation and things like, so we use Confluence, for instance,

MS Teams as well. That's not probably as useful as the formal documentation in Confluence. Although Confluence is not great for search either. So there are solutions.

But I think there is some... Is it Tamar?

No, I'm thinking of something else that's for helping find data, as opposed to other people's work.

Okay. How do you measure the KPIs end-to-end on the DataOps AI pipeline data lineage?

KPIs to data lineage. So, we track the transformations we do, the data lineage we do.

I'm not sure we have KPI. We have different KPIs for the end-to-end. So, for instance, we're interested in things like data completeness. So of the data we start with, how much do we end up with at the end, to make sure we don't drop records, for instance. We're interested in KPIs around latency. So how long does it take to run a process?

Freshness of data. So when was, let's say, a record or prediction last updated, those kind of things.

But I'm not sure whether we have KPIs around the data lineage itself. Okay. I guess we should try and have one, actually. I think a good idea would be a coverage of data lineage across all our jobs. Yeah. Okay, great. How do you see the importance of DAGs in the DataOps space?

DAGs are pretty critical to a lot of data pipelines. So we

kind of have two approaches. So we have a little bit of a traditional ETL approach. So we're multi-cloud. So one of our challenges is that actually some of our DAGs actually not just span tools and platforms, but entire clouds, which is a bit of an issue, because it creates a bit of a challenge around using the same tool across all clouds.

But, so DAGs are pretty important. So we have a sort of more traditional approach on AWS, which is ELT, ETL approaches, and we use more traditional ETL tools there, things like Talend. And then on Google Cloud Platform, we typically use what is managed Airflow on Google Cloud Platform, Cloud Composer, I believe it's called. And that's critical to basically taking data from the storage buckets where it typically lands, through to where it needs to be processed. We also use things like Dataflow as well.

So some of our DAGs can be reasonably complex, and they require a fair amount of monitoring. But we try not to make them too complex. We typically then try to look for a different tool. So what the cloud providers are typically doing is offering software as a service

00:50:00

for many of the tasks, so for loading data where we traditionally might have used something like Airflow. Or there's software as a service, ELT, ETL tools like Segment. So yeah, DAGs are important, but we try not to make them too complex. If they are too complex, we'll look for another solution. Okay, great. This is a great question.

Where would you suggest a group start in order to start implementing DataOps? We're a team that has our flow created already, but the automation part is lacking, so it seems like we're putting the cart before the horse.

That's a good question. So,

when it comes to automation, you really need access to the right tools and platform first to automate your goals. So that's kind of prerequisite. So if you don't have those, that would be where I'd start building a case to have access to those tools and platforms, and explain the benefit. So,

when you're only in control of a whole end-to-end process, it can be quite challenging to kind of create any value, because you can do some local optimization in your space, but the benefit may not be so great in the end-to-end chain. So what I would typically ask or kind of recommend that people do when they're starting off is not try to focus too much on optimization in their local area, but try and find an end-to-end use case.

So find an area where you have some friendly stakeholders either side. So either it's the people who are helping you in terms of data capture, and providing data, or the consumers of your data at the other end. And try and look at that entire end-to-end process, map it out, highlight where the problems are, highlight where you might be able to get the most benefit, and essentially create a benefits case, and attack that specific bottleneck first, the one which is most promising. So for us, to give you a real life example, we started down this journey of DataOps with personalization in marketing.

So we started working with data engineers who were supplying the data science team with data. And the output of that was used with our marketing teams to use machine learning to personalize the communications our customers were receiving. And we took that entire sort of problem space, and basically we looked at that and we said, "Right, what can we do to create a pipeline which allows us to iteratively test and learn as quickly as possible?" So we came together as a group. We did a lot of test and learn.

As we were going along, we were building the automation. So we were building the automation of the machine learning data set builds, the scoring, the testing, and also the measurement at the other end. So we were starting to build measurements that automatically measure the results of the A/B testing we were doing with the models.

And so we started with the process, and then we built up the automation steadily. Because we could see from doing it in a string until a takeaway, sorry, that's a UK colloquialism, but very basic way, the benefit, and that allowed us to invest in the automation. Rather than starting with the investment in automation, only to find that actually there wasn't much value in automating what we were doing.

Sorry, that was a really long answer, wasn't it? No, that was a great answer. We've got actually a ton of questions, so I'm going to pick a couple more, and if we don't get to answer your question today, we'll follow up via email. So, here's one. Any advice on how to go about replacing legacy reporting and analytics systems using DataOps methodology and avoid having to reverse engineer existing complex systems built in technologies like SAS?

That's a good question, because we went through that same process. So we were using SAS a few years ago.

There was some things which we basically just had to throw away and start from scratch. So there was a lot of parallel development going on, on new platforms. And we kind of had to incrementally Basically swap out outputs from legacy systems. So started with

00:55:00

marketing. So marketing solution was the first we migrated, then some of the analytics solutions.

So we took a piecemeal approach rather than burning with fire. It meant that it did take a long time. It took longer than we were anticipating, but it was the least disruptive way of doing it.

So for us, what we did was, it was make sure that we had the right foundation, so the right common data layer in the cloud for us to be able to build those applications that we needed to rebuild. So when we're ready to go off SaaS, then we could basically have the data and basically build an application, a data pipeline to replace what we had.

There's always a challenge when you're switching data sources, the two things will never produce exactly the same output. So, we had to put tolerances around that. So we were willing to accept a certain amount of difference. Beyond that, then we would investigate by exception.

Okay, great. Harvinder, how does one manage technical conflict between DataOps/data teams and software engineers, especially in environments where the product is software-driven? So the interface between teams is always the most challenging area. So we have to interact with our product engineering team. So we are relying on them in two ways. One is for data capture.

So we're a website, most of our data comes from event-based capture.

So we have to work quite hard at that type of relationship, basically to make sure that data is at the forefront of their mind, that they understand why capturing data is an important way. We have to work in a collaborative way.

So one of the things which will guarantee to break our pipelines is a front-end change. So we need to be notified of these kind of things well in advance, so that we're prepared for them. And then we're reliant in another way, which is, basically some of our outputs need to be incorporated back into our customer touchpoints, so the webinar, for instance.

And typically the challenge we have there is more technical. So, the software developers, they will be used to rigorous software development life cycles, a lot of rigorous testing, especially around non-functional requirements. And so they want to see that we are going through some of the same rigorous methodology. So there, the way we get around that is to basically sort of learn as much as possible about their domain as possible, so that we are prepared for their questions and their challenges and have done as much as we can at our end to show that we are as professional as possible, even though we're a data function, and that we are sensible and we can be trusted.

Okay, this may be a good follow-up question and probably our last question. Any tips or strategy to grow T-shaped team members?

Yeah, so, one of the things was just with the move to cross-functional teams. So by putting data analysts, data engineers, and data scientists together in a single team, they were able to see how each other worked. They were able to express their learn-- they wanted to learn from each other. And so teams do spend time cross-training each other.

We also make sure that we have backup for other people, which means that people have to basically learn each other's jobs to some extent. And then also, training as well. So, we encourage people to primarily online training. There's a lot of help. But we also sponsor people. So we sponsor a few people for master's courses, in data science. We sponsor people for data engineering certifications and so on.

So

it's a mix. So it's a mix of informal learning from each other and formal training. Okay, thanks. Well, we are up against the hour now. So I just want to say thank you so much to Harvinder for joining us today and coming and sharing your experiences and wisdom on DataOps. It was great practical advice, and I hope everyone appreciated it.

And also thanks to the audience for joining us and attending the webinar. If we didn't get to your question, we'll follow up directly with you. We will also be sending out the slides and a recording of the webinar via email, likely in the next day, so be on the lookout for that. And I also encourage everyone to check out "Practical DataOps," the book, and

01:00:00

let us know what you think. Thank you again.

No, thank you. Thank you, everyone who attended. If you do ever want to reach out to me, you can find me on LinkedIn or Twitter. So I'm always happy to talk data. Okay, thanks. All right. Have a good afternoon, everyone.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What is DataOps in the context of data science?

DataOps applies three proven methodologies to data analytics: Agile, DevOps and Lean thinking. The goal is quality and speed in delivering data products, not better algorithms. Because data analytics behaves like complex manufacturing, from ingestion through transformation to data products, the practices that made software delivery reliable transfer directly to the analytics production system.

Why do data science projects fail to deliver business outcomes?

Because success is treated as starting with data, data scientists, models and technology, when it ends with them. The Program Logic Model runs resources, activities, outputs, outcomes and impact, and data teams stop at outputs. Working right to left, from the impact you want back to the activities that produce it, is what connects an analytic insight to an organizational objective.

What is a data product?

DJ Patil, the former US Chief Data Scientist, defined a data product as a product that facilitates an end goal through the use of data. The shift is from thinking about projects, which end, to products, which have a lifecycle: concept, inception, development, transition, production and retirement. Data is no longer an application by-product, so it needs a strategy and rigor of its own.

How do you find waste in a data science delivery cycle?

Lean process mapping. Map every step from initial design through searching for data, data access, cleaning, proof of concept, model build, IT resource provisioning, software installation, recoding, testing and deployment, and mark each as work or wait. One such map showed 57 days of value-adding work against 234 days of waiting, or 20 percent efficiency, with the waits sitting in provisioning and access.

Should data teams be organized by skill or by domain?

Functional teams grouped by expertise, such as data scientists or DBAs, optimize for utilization of scarce talent and only work well when every functional team shares goals or delivers genuine self-service. Domain-oriented cross-functional teams optimize for speed and are what nearly 60 percent of breakaway organizations use. Microsoft's research found organizational structure predicts code quality better than code churn, complexity or test coverage.

What is a T-shaped team member?

A T-shaped person is a generalizing specialist: expert in one thing and capable in many. An I-shaped specialist is expert at one thing only, and a dash-shaped generalist is capable in a lot but expert in nothing. Cross-functional teams still form silos inside themselves when they are staffed only with I-shaped people, so cross-skilling toward T, Pi and M shapes is what makes the team work.

Where to go next