On-Demand Webinar · 1 hr 1 min
Practical DataOps: Delivering Agile Data Science at Scale
Harvinder Atwal, author of Practical DataOps and Group Data Director at MoneySuperMarket, on aligning the people, processes, and technology of an analytics organization with the rest of the company's goals, and the steps his team took at MoneySuperMarket. Recorded May 2020; updated August 2026.
What you'll learn 8 points
- Harvinder Atwal of MoneySuperMarket opens with the state of the field: NewVantage Partners' 2020 survey found only 7.3 percent of organizations rate the state of their data and analytics as excellent, only 22 percent of companies see significant return from data science spending, and Gartner predicted that through 2022 only 20 percent of analytic insights would deliver business outcomes.
- Technology matters less than teams assume. Research across 23,000 survey responses from more than 2,000 organizations by Jez Humble, Gene Kim and Nicole Forsgren found no significant correlation between system type and delivery performance, and the share of firms naming technology as the principal challenge to becoming data-driven fell from 19.1 percent in 2018 to 9.1 percent in 2020.
- Lean process mapping of a data science delivery cycle exposed 57 days of value-adding work against 234 days of waiting, an efficiency of 20 percent. The waits were IT resource provisioning, software installation, data access and model recoding, not modeling.
- Fixing those waits moved real numbers at MoneySuperMarket: lead time to make a new data item available for analytics went from 2.3 years in 2017 to three weeks in 2020, and the time to test a change to a machine learning model pipeline went from 11 hours to five minutes.
- Data is a product, not an application by-product. DJ Patil's definition is used: a product that facilitates an end goal through the use of data. Work back from the impact and outcome you want, not forward from the data you have.
- Conway's Law is not academic. Microsoft research found organizational structure predicted code quality better than code churn, code complexity, dependencies, test coverage or pre-release bugs, and nearly 60 percent of breakaway organizations use cross-functional teams against less than a third of everyone else.
- Functional teams organized by expertise optimize for utilization of scarce talent; domain-oriented cross-functional teams optimize for speed. Cross-functional teams still form silos inside themselves unless members cross-skill, moving from I-shaped specialists toward T-shaped, Pi-shaped and M-shaped people.
- Fitting a model is the easy part. Around it sit data governance, data quality, data security, test data management, version control, access control, team organization, stakeholder buy-in and outcome measurement, and DataOps is what makes those repeatable rather than heroic.
Slides
Questions from this session
What is DataOps in the context of data science?
DataOps applies three proven methodologies to data analytics: Agile, DevOps and Lean thinking. The goal is quality and speed in delivering data products, not better algorithms. Because data analytics behaves like complex manufacturing, from ingestion through transformation to data products, the practices that made software delivery reliable transfer directly to the analytics production system.
Why do data science projects fail to deliver business outcomes?
Because success is treated as starting with data, data scientists, models and technology, when it ends with them. The Program Logic Model runs resources, activities, outputs, outcomes and impact, and data teams stop at outputs. Working right to left, from the impact you want back to the activities that produce it, is what connects an analytic insight to an organizational objective.
What is a data product?
DJ Patil, the former US Chief Data Scientist, defined a data product as a product that facilitates an end goal through the use of data. The shift is from thinking about projects, which end, to products, which have a lifecycle: concept, inception, development, transition, production and retirement. Data is no longer an application by-product, so it needs a strategy and rigor of its own.
How do you find waste in a data science delivery cycle?
Lean process mapping. Map every step from initial design through searching for data, data access, cleaning, proof of concept, model build, IT resource provisioning, software installation, recoding, testing and deployment, and mark each as work or wait. One such map showed 57 days of value-adding work against 234 days of waiting, or 20 percent efficiency, with the waits sitting in provisioning and access.
Should data teams be organized by skill or by domain?
Functional teams grouped by expertise, such as data scientists or DBAs, optimize for utilization of scarce talent and only work well when every functional team shares goals or delivers genuine self-service. Domain-oriented cross-functional teams optimize for speed and are what nearly 60 percent of breakaway organizations use. Microsoft's research found organizational structure predicts code quality better than code churn, complexity or test coverage.
What is a T-shaped team member?
A T-shaped person is a generalizing specialist: expert in one thing and capable in many. An I-shaped specialist is expert at one thing only, and a dash-shaped generalist is capable in a lot but expert in nothing. Cross-functional teams still form silos inside themselves when they are staffed only with I-shaped people, so cross-skilling toward T, Pi and M shapes is what makes the team work.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.