DataOps Mission Control And Managing Your Data Infrastructure Risk

The Head of data got a call from the CEO of the entire company about a compliance report that was empty, with no data. So, he had to rally 26 different people across his team all-day And what was the problem? A field passed through the pipeline that was blank. Can you imagine how embarrassed he is at the error? How frustrated all those 26 people -- most likely the best he has on his team -- at having to chase a crappy error? And he has 1000 other pipelines in the same 'hope it works' position, just waiting for some customer to find a problem. High risk, indeed.

Written by DataKitchen Marketing Team on June 1, 2022

DataOps Mission Control And Managing Your Data Infrastructure Risk

Key points

  • Most data teams cannot answer basic operational questions about their own pipelines: whether source data arrived on time, whether every scheduled job actually ran, or whether the numbers in a report are fresh.
  • The same gap covers the development process: what code is in which environment, how many deploys failed, how many tests ran in QA, and which tickets shipped.
  • Five things keep it unsolved: teams are already overloaded, they fear changing an architecture that is running, no single view spans their tools, they do not know what to check, and problems surface as blame without shared context.
  • The mission control idea is borrowed from spaceflight: one interface covering every aspect of the flight, used for decisions and communication, stored for later analysis, and set up to alert automatically.
  • The goal is visibility of every journey data takes from source to customer value, across every tool, environment, data team, and customer, so problems are detected, localized, and raised immediately.

DataOps Mission Control

Data Teams can’t answer very basic questions about the many, many pipelines they have in production and in development. For example:

Data

Jobs

Quality/Tests/Trust

Tools/Models/Dashboards

Root Cause

They also can’t answer a similar set of questions about their development process:

Deploys

Environments

Testing/Impact/Regressions

Productivity/Team/Projects

Why does this happen? Why is this problem not solved today?

  1. Team Are Very Busy: teams are already busy and stressed – and know they are not meeting their customer’s expectations
  2. They have a Low Change Appetite: Teams have complicated in place data architectures and tools. They fear change in what already running
  3. There is no single pane of glass: no ability to see across all tools, pipelines, data sets, and teams in one place. They have hundreds or thousands of existing pipelines, jobs, and processes running already.
  4. They don’t know what, where, and how to check. They need to make sure their customers are happy with the resulting analytics. But they often don’t know the salient points to check or test.
  5. They live with lots of blame and shame without shared context. Problems are raised after the customer has found them, with panicked teams running around trying to find who and what is responsible.

How do other organizations solve this risk problem? The biggest risk of all is space flight. How do SpaceX and NASA manage risk? Build a Mission Control

A New Concept

DataOps Mission Control’s goal is to provide visibility of every journey that data takes from source to customer value, across every tool, environment, data analytic team, and customer so that problems are detected, localized, and raised immediately. How? By being able to test and monitor every data analytics pipeline in an organization, from source to value, in development and production, teams can deliver insight to their customer with no errors and a high rate of pipeline change.

However, if you solve this problem, you will see:

TIP

Learn more – watch our on-demand webinar


FAQ

What are the key points in this blog?

Data teams usually cannot answer basic questions about their own pipelines — whether source files arrived, whether every scheduled job ran, how many tests pass in production, or which code sits in which environment. DataOps Mission Control applies the spaceflight answer to that gap: one view of every journey data takes from source to customer value, across every tool, environment, team, and customer, so problems are detected, localized, and raised immediately.

What is DataOps Mission Control?

A concept borrowed from spaceflight rather than a product component: one place that shows every journey data takes from source to customer value, across every tool, environment, data team, and customer. DataKitchen builds that view in DataOps Observability, which tests and monitors pipelines in development and production so a failure is detected, localized, and raised before a customer finds it.

What questions can data teams usually not answer about their pipelines?

Operational basics. Whether source files arrived on time, whether the data in a report is fresh, which supplier keeps causing trouble, whether one job ran only after an earlier group finished, whether every job that was supposed to run actually ran, how many tests exist in production and how pass and fail rates move over time, whether a model is still accurate, and where a problem started.

Why do most data teams have no single view of their pipelines?

Five reasons recur. Teams are already busy and know they are falling short of what their customers expect. They have a low appetite for changing an architecture that is already running. Nothing spans every tool, pipeline, data set, and team in one place. They do not know which points to check or test. And problems surface as blame and shame without shared context.

What does a mission control approach borrow from NASA and SpaceX?

Four practices. Build one interface carrying information about every aspect of the flight. Use that information as the basis for decisions and for communicating with everyone who has an interest. Store it so a failure can be analyzed after the fact. And alert and flag automatically instead of waiting for someone to notice. Applied to data operations, those four make pipeline risk visible before a customer finds it.

What changes when a data team can see every pipeline end to end?

Four things. Less embarrassment, because breakage is found before a customer reports it. Less hassle, because production status becomes self-service instead of a queue of people asking. More space to create, because time goes to new work rather than chasing errors. And a step toward transformation, which stays out of reach while customers do not trust the data or the team.

Install Open Source TestGen Free, no vendor lock-in Request a Demo See TestGen Enterprise in action

DataKitchen Marketing Team

The DataKitchen marketing team curates industry news, resources, and thought leadership on DataOps, data quality, and data observability.

Don't want to give us your email address? Go directly to the webinar recording here.