Failing Your Way to Success with DataOps

Borrowing 'fail faster' from lean manufacturing and design, DataOps lowers the cost of data errors so teams can take risks and innovate with confidence.

Written by DataKitchen Marketing Team on January 21, 2020

DataOps Principles
Failing Your Way to Success with DataOps

Key points

  • Good people underperform inside a bad system: when a data team cannot deliver quickly enough or its quality is poor, the methodology is the likelier culprit than the people.
  • “Fail faster,” attributed to David Kelley of IDEO, is less about speed than about minimizing the cost and consequences of failure, which is what frees a team to take bigger risks.
  • Lean manufacturing makes the same argument in money: cost-of-goods-sold rises as a product moves through the process, so screening out a faulty component before assembly costs far less than mid- or post-production rework.
  • DataKitchen’s 2019 DataOps survey found that 30% of respondents reported more than 11 data errors a month, so errors reaching user analytics is the normal case rather than the exception.
  • DataOps attacks four classes of error — data pipeline, development, integration and product management — and lowering the cost of all four is what gives analysts room to experiment alongside business users.

Have you ever failed at work? Most people have done something cringeworthy at some point. I have plenty of funny stories to tell. There’s the time I set off the fire alarm on my first day of a new management job. Perhaps even more notable are the epic fails. I once worked at a company where it seems, on a near-daily basis, I would find myself on the phone with an angry customer. There was another data error. If it wasn’t fixed immediately, the customer would find a different vendor. My reputation and career were constantly on the line. Not fun. At first, I worried that the problem was me. I always thought of myself as a dependable person who gets the job done. I wasn’t used to coming up short week after week.

Fail Faster

Over the decades, I have come to understand that good people underperform when operating within a bad system. If the data team can’t deliver quickly enough or suffers from abominable quality, the culprit is likely the methodologies that they are using. Some companies become overly dependent on a star employee who they drive to work long hours. When people work in teams, task coordination and the elimination of process inefficiencies is far more consistent and impactful than individual ability or heroism.

In the industrial design domain, attitudes towards failure have matured. The adage “fail faster,” attributed to David Kelley of IDEO, expresses how failure fuels innovation. Failing fast implies that it is about speed, but it is more generally a call to minimize the consequences and the cost of failure. Failing fast enables designers to feel freer to take big risks, leading to creativity and innovation that powers progress and growth.

Operations managers widely apply the axiom “fail fast” to lean manufacturing. As a product progresses through a manufacturing process, its cost-of-goods-sold increases. Every factory manager knows it is much less costly to screen out a faulty component prior to assembly than to invest in mid- or post-production rework.

Failing Faster in Data Analytics

When data professionals apply these same principles to data analytics, they can reap the rewards of lower costs and higher creativity that we see in industrial manufacturing. In the data industry, the application of these methodologies across the data lifecycle is called DataOps. In summary, DataOps reduces the cost and consequences of data errors and data-analytics bugs. You could say that “failing faster” is the common theme that unifies all aspects of DataOps.

Data Pipeline Errors

Data operations consist of a set of data sources that progress through a series of processing steps, for example, integration, cleaning, processing, transformation and publication (as charts, graphics and reports). 30% of respondents to our recent DataOps survey reported more than 11 errors per month. In a significant number of enterprises, data errors are regularly flowing into user analytics with potentially catastrophic results.

DataOps places tests at each stage of the data-operations pipeline. It checks and monitors data at its source before it enters the pipeline. Does it conform to business logic? Does it fall within statistical norms? DataOps also tests inputs and outputs at each stage of transformation. If a fork or join fails, DataOps will catch it before it corrupts analytics.

DataOps implements automated statistical and process controls on data operations, much like a manufacturing plant. If data flowing through a multi-stage data-analytics pipeline violates business logic or statistical norms, then tests alert the data team or, in an extreme case, stop the flow of data.

Development Errors

In many organizations, new analytics are developed directly on operational systems. DataOps allows data analysts to create development sandboxes. With virtualization technology, sandboxes closely match the target production environment minimizing unexpected regressions. Sandboxes inherit analytics components and automated orchestrations along with associated tests so data scientists can better leverage each other’s work. If a data scientist takes a development risk, they can abandon a sandbox and revert to the baseline analytics code and configuration.

Integration Errors

DataOps employs continuous integration methods like those that enable leading software organizations to deploy millions of code releases per year. Development sandboxes are isolated from production unless they progress through an automated release workflow that includes integration, functional, unit and other tests. Tests catch issues before analytics migrate into data operations.

Product Management Errors

Developing analytics that no one wants or needs is a costly product management failure. When DataOps teams implement Agile Development, they create short-term value by iterating rapidly. It’s much easier to be correct about what feature you need this week as opposed to 24 months from now. Also, Agile teams receive immediate feedback on what they have produced so they can course correct. Often, users don’t know what they want until they see it. Teams can be much more innovative when they create a rough, approximate solution and keep iterating on it.

Failing Your Way to Success

DataOps focuses on the identification and elimination of data pipeline, development, integration and product management errors. When DataOps minimizes the cost and consequences of errors, data analysts are free to work more closely with business users. Together they can play with ideas, try new things, and act on hunches. When this process plays out, it unlocks tremendous creativity. With failures minimized and identified early, DataOps enables enterprises to deliver on the promise of leveraging data for competitive advantage. We have seen many companies use DataOps methods to vault forward, taking a leadership place in the market. Fail faster using DataOps.


FAQ

What are the key points in this blog?

DataOps borrows “fail faster” from industrial design and lean manufacturing: the point is not failing more often but lowering the cost and consequences of each failure. Applied to analytics that means tests at every pipeline stage, isolated development sandboxes, an automated release workflow, and short Agile iterations. It targets four classes of error — data pipeline, development, integration and product management.

What does “fail faster” mean in DataOps?

It means minimizing the cost and consequences of failure rather than failing more often. The adage is attributed to David Kelley of IDEO, where cheap early failure fuels creativity and frees designers to take big risks. In data analytics the same logic holds: when a broken transformation is caught in minutes and costs little to repair, a team can afford to try things.

Why do good data teams still miss their commitments?

Usually because the system around them is bad, not because the people are weak. A team that cannot deliver quickly enough, or whose quality is poor, is normally running a methodology that produces those results. Task coordination and removing process inefficiencies turn out to be more consistent and impactful than individual ability or heroism.

How many data errors does a typical data team hit?

More than most people would guess. DataKitchen’s 2019 DataOps survey found that 30% of respondents reported more than 11 data errors a month, which means errors regularly flow into user analytics instead of being caught first. That rate is why lowering the cost of each failure matters more than promising failures will stop.

What kinds of errors does DataOps try to catch?

Four kinds. Data pipeline errors are caught by tests at each processing stage that check business logic and statistical norms. Development errors are contained by sandboxes that closely match production. Integration errors are caught by an automated release workflow with functional and unit tests. Product management errors — building analytics nobody needs — are caught by short Agile iterations.

How does lowering the cost of errors help a data team innovate?

It removes the reason to say no. When failures are small, caught early and cheap to fix, analysts can work closely with business users, play with ideas and act on hunches instead of spending the week repairing yesterday’s report. Creativity follows from a low cost of failure rather than from a bigger tools budget.

Install Open Source TestGen Free, no vendor lock-in Request a Demo See TestGen Enterprise in action

DataKitchen Marketing Team

The DataKitchen marketing team curates industry news, resources, and thought leadership on DataOps, data quality, and data observability.