On-Demand Webinar · 51 min

DataOps Risk Insurance & Mission Control

Chris Bergh on treating DataOps as risk insurance for a data investment: how you know the data is right, where the risk hides in models and reports, and what a Mission Control for the data organization does. Recorded June 2022; updated August 2026.

Presented by Chris Bergh

What you'll learn 7 points
  • At a top-20 US company by revenue, the head of data took a call from the CEO about a compliance report that came back empty. The cause was one blank field passing through the pipeline, 26 people spent a day chasing it, the deploy cycle was six weeks, and a thousand other pipelines sat in the same hope-it-works position.
  • The 2022 DataKitchen and data.world data engineer survey found 52 percent hope and pray things don't break, 78 percent are stressed enough to want a therapist, 70 percent expect to change jobs within a year, and 79 percent have considered leaving the career entirely.
  • DataOps Mission Control takes its model from how NASA and SpaceX handle the risk of space flight: one interface carrying information about every aspect of the flight, decisions and communication driven from that information, history stored for after-the-fact analysis, and automatic alerts.
  • The unit Mission Control watches is the observational meta-pipeline, a layer above the DAGs, jobs, and schedulers already running. It records only the major steps a team cares about, treating the dozens or hundreds of sub-steps inside each tool as noise, and it holds fan-in, fan-out, causal, temporal, manual, periodic, evented, and inferred relationships that most organizations keep as tribal knowledge.
  • IT hardware and application performance monitoring are lagging indicators of data problems. APM tools do not check the data, the integrated data, or the reports and models built from it, and they give no context for the pipelines, jobs, and tools acting on that data.
  • A data catalog and data lineage answer what the data is and where it came from. Two governance questions remain: can I trust the data, answered by test results attached to each artifact, and is the data fresh, answered by process lineage recording when it was last updated and at what version.
  • The stated adoption sequence is Mission Control first as the pain pill that reduces risk, then DataOps Automation as the vitamin that drives customer value. The closing framing is that it is not about data quality, it is about a low rate of errors.

Prefer to read it? The written version is in DataOps Mission Control And Managing Your Data Infrastructure Risk.

Slides

58 slides

Transcript

Show chapters and dialogue 8,627 words

00:00:00

of DataKitchen, and I'd like to welcome you to today's webinar about DataOps as risk insurance for your data infrastructure. But before I start, I'll do a few housekeeping. The first thing is that the slides and the recording of this webinar will be shared after the event. The second is I will take time for questions at the end, but if you have questions, feel free to type them in the question part of your GoToWebinar meeting, on the right or the left or wherever it is on your screen. And I'll have it open during the talk, and if I get a chance, I'll try to bring it up, but I will have time for questions at the end.

And so I should talk for about 45 minutes, and then we'll take some questions. So again, my name is Chris Bergh. I'm the CEO and head chef of DataKitchen, and welcome to our webinar on DataOps as risk insurance for your data infrastructure. And so let me go ahead and walk through some slides here.

So the first thing we're going to talk about is when risk goes bad, and I'm going to share some horror stories from our prospects, from our customers, from my personal life, and then we're going to talk about a way to control risk through something called mission control or DataOps mission control. And we're going to talk about how that resolves risks and give the characteristics of it. And then we're actually going to talk about some software that we're developing around the mission control and give you a preview of what it looks like.

So let's go through. I have a little theme in this of TV and music. So I don't know if you know "Breaking Bad," but it was a little bit slightly different, but I heard this story just a few months ago at a company, sort of top 20 company by revenue. And so there was the head of data, so he had, I don't know, 100, 200, 300 people working for him.

Got a call from the CEO of the entire company about a compliance report that was empty, that had no data. And how did the CEO find it? Well, the CEO of the other company who it was supposed to be given to called him up to yell at him. And so imagine the CEO of a Fortune 20 company calling you up and saying that you're a moron. And so he immediately got 26 people on his team to spend all day on the problem. And of course, you're a data person.

What do you think the problem was? Well, it was a field that had gone blank that passed through that caused the report not to render. It's something dumb, as a lot of these things are. And his time to fix, usually, of these things is about six weeks. So can you just imagine how embarrassed he is as a leader of that error, how embarrassing it is to work on that team, and how frustrated all those 26 people are to be interrupted from their daily job, interrupted from their workflow.

Most likely it's some of the best people on his team to dig into logs and find it out, and then really having to chase this kind of crappy error. And the problem is he's got 1,000 of these pipelines running in this organization that move data from one place to the other, that give insight, that have sort of the similar problem. You just sort of hope it works.

And he's waiting for some customer to find a problem. And that's very similar to my experience. So I spent 15 years kind of building software and managing software teams at companies like NASA and MIT and Microsoft and some internet startups. And then about 2005, I got the bright idea that I'd join the data and analytics industry. And so I managed a team that did what we now call data engineering and data science and visualization, and we did it as sort of an outsource basis for large healthcare companies. So we had thousands of users of our reports and dashboards and models and datasets, and

I'd have to go out and fix things. So for instance, I'd be sitting at my kid's soccer match and I'd get a call, and I think had a BlackBerry at the time, and go, "Uh-oh. I've got to go fix something. I've got to rally the troops. Something's wrong with this week's data build. Yeah, I know it's the second or third or fourth time this season." And my wife would give me a dirty look, and I'd sort of slink away and try to figure out how to fix it or get the right person to fix it.

And then, I had smart people who were really trying to help customers, and they'd put last-minute changes in. They'd say, "Oh, my customer wants me to do this," and our warehouse builds on Friday, so here it's Friday morning. "I'm putting in this last-minute fix, and I kind of hope it works." And then I'd talk to them weeks later and they'd say, "Oh, yeah, it worked because I haven't gotten any calls." And so, then I'd always sort of have this pit of my stomach kind of going into work that something was wrong, that there would be a delay, that the data would be wrong, that some data provider would have given us bad data, that a server would've gone down, that somebody had put some bad code in, and so I started bracing myself every time. I started dreading going into work after a few years, that I would get this

00:05:00

nasty email, and then I'd have to placate the people who were complaining, the heads of marketing and sales of our sponsoring customers. And oftentimes we integrated data from dozens, maybe 50 or 100 different data sources for every customer. And in that case, as we all know, your data providers don't care that you exist.

And they may give you bad data or change the meaning of data or forget to give you data, and all those things end up sort of producing problems. And so

one of the things that we did last year was actually do a survey of data engineers. And oftentimes data engineers are the ones doing the data processing, but they're also doing the production in some cases. And we just found these statistics that I think are kind of shocking, but fit my view of what I've experienced, is that 52% hope and pray that things don't break, and 78% are so stressed that they need a therapist, and 70% want to quit, and 79%, because of the stress, have considered switching careers entirely. And now there's just a shortage of people who do data work, not only data scientists, but people who do data engineering.

And I think one of the shortages is a supply problem, but it's also a problem that people are leaving the field because it's so bad. And so, what I want to talk about now is the solution to that sort of examples of risk gone bad.

And so why haven't people really solved this problem? If it is so bad and people are stressed and things are breaking left and right, why haven't they solved it? Well, and by and large, most companies haven't solved these problems. I think the first reason that the teams are just very busy. They have a backpack full of customer requests and a lot of things to do, and they know they're not meeting their customer expectations, so they go into work saying, "I've got to get some stuff done or else my customer's unhappy." And then they build things and they're afraid to change them.

And they've got complicated in-place data architectures and tools and code, and they know it's working. They don't want to change it. They've sort of built these kind of very precious tools and systems, and they're just afraid to change them. And they're also afraid because they have no single pane of glass. They have no ability to see across all their tools and pipelines and datasets and environments and people and organizations.

Some companies have literally hundreds or thousands of existing pipelines or jobs or processes, or as that journey, as data comes in to go to value, they just don't really know where any of it is. And likewise, a lot of teams don't know what and where to check. They want to make sure their customers are happy, right?

And their customers are not going to yell at them with, "This data looks wrong" or, "This data looks weird." But oftentimes they don't know what are the salient points to check. What should I look at? How do I find things that are important? This seems hard and complicated. And then there's lots of blame and shame in organizations where people say, "It's not my fault.

It's your fault," because there are lots of teams working on the data part or the ingest part or the transformation part or the biz part or the science part, and when things go wrong, they sort of blame each other. And one of the big reasons is there's really no kind of shared context within teams about where it is. It's just sort of people kind of saying they know their little part, but they don't know the bigger world.

And so I think these things of busy teams and they build things that they don't want to change, and they don't have any visibility, shame and blame culture, and then sort of a lack of general knowledge of where to check things. And so, I think there's a set of basic questions that people should be able to answer about anything that's running to produce analytic insight in their organization.

Like, is my data going to arrive on time? Is the data good? Is the data fresh? When I process the data and integrate it with other data sets and build schemas and storage, is that right? And the process needs to do that, whether they're batch or streaming. Do they actually run? Do they run on time? Are they late?

How many are running today? How many ran yesterday? And then, to produce good,

trusted data and reports and models for your customers, you have to check it or test it. And so, is the data actually right? Did you actually look through the telescope and see it's not blurry?

00:10:00

And are my models still predicting? Are my dashboards still showing the right data? And if there is a problem, how can I get to the root cause, and can I fix it before someone notices? And so there's a lot of very basic production questions that are hard, or they're known by certain people. They're known as pockets, so they're not known generally.

Maybe it's the one person who happens to know the very details of your Informatica jobs, and some other person knows the details of your data science thing. And if they do know, it's certainly not in one place. And then there's a bunch of development process questions about getting things into production that people don't know a lot.

Like how many deploys are happening, and how often are we deploying, and what parts are affected by the deploy if something goes wrong? And what code is in what environment? And how many things are actually in production that we're changing? And how often have I done changes, and have they passed? And how productive is my team? How many tickets did we release?

How many deploys and tickets are happening for a particular project? And so there's very basic questions about what's happening in development and what's happening in production that I think are very hard for teams to answer, or those answers are only in certain parts of an organization with certain tools and certain teams. They're not widespread at all.

And so this lack of being able to answer the questions, these frictions are really causing risks. And so how do organizations handle serious risk, right? And I'm not saying that data and analytic processing is anywhere near as risky as going into space, but there are organizations who have managed to do high-risk situations. And think of SpaceX or NASA.

They create something called a mission control. And in a mission control, there are a lot of screens. There's a UI that has a lot of information about all aspects of what's happening with the flight and the launch and every phase of the flight. And that shared information is kind of the basis for making decisions and changes, and actually communicating to interested party.

It becomes the context at which people can understand what's going on. And then that information actually can be stored and used for after-the-fact improvement. And then a lot of it is having people sitting looking at it or automated alerts to understand what goes on, and that's sort of the movies and the red flashing lights, et cetera. And so organizations that have serious risk, like space flight, have decided that they need something called a mission control to manage that risk. And so I'm going to introduce this concept today called DataOps Mission Control, or DMC. And its goal is to provide visibility of every journey that data takes from source to customer value, across every tool, environment, data and analytic team, and customer, so that problems are detected, localized, and raised immediately.

And the how on that is really being able to test and monitor every data and analytics pipeline in the organization, in development, in production, so that when teams can deliver things with no errors and a high rate of pipeline change. And so I'm using the term pipeline in a very broad sense here. I just don't mean data pipeline. I mean model pipelines and visualization pipelines, the whole end-to-end value chain of where data comes in, and then at the finally some business customer or website's actually interacting with it.

And we've talked a lot about DataOps in these webinars, and really, I think DataOps' job is to deliver customer value is really number one, making sure that your customers have good insight, that they can ask as many questions as they want of the data, and you and your team and tools can satisfy that.

And so it is really about rapid cycles of experimentation, low error rates, collaboration, and clear measurement. But what we're saying is that there's-- think of it as kind of two steps to DataOps. So the first is sort of take a pain pill, reduce the risk and your hassles first, and then focus on the vitamins, on iterative development and customer value.

Because we think that if you're really stressed and really unsure of what's going on, you're likely to want to sort of close in and not make a lot of changes. And so sort of reduce the pain first. And that's what we're going to talk about, is what is this idea of mission control that can reduce that pain and reduce this risk? And what are the problems that need to be solved?

And so I'm going to kind of walk through this sort of step by step.

00:15:00

And so let's just start off with the emotions and what sort of feelings. So let's say you have this sort of mission control for all your data pipelines across all your organizations, both in production and development. Well, you're going to be embarrassed less, right? Your customers are not going to find things that are broken because you've checked them, and they're not going to find reports that are blank.

And if they do, you're going to be able to say, "Oh." You're going to be able to look quickly and say, "Part three of a 20-part process didn't work. Let us go investigate it and put a check in and make sure it never happened again." And you're going to have less hassle because you're building a system that allows you to get people off their back and sort of self-service the checking of the problems. And so we're going to talk about sort of views for your customers, and whether they're data analysts, if you're building a hub-and-spoke solution or business customer to kind of self-service what's happening in your organizations and sort of get them off your back, and so you can focus on creating and building.

And that's really the third benefit, right? Instead of chasing problems, context switching, answering simple questions. Like if you think about those 26 people when that report was blank, well, they were doing something, right? And they had to stop doing it. And I'm a technical person. I like long periods of uninterrupted technical five, six hours because it takes me a long time to sort of get up to speed and work in that zone of creativity and then come back down, and context switching and meetings and all get in the way of actually doing the real work. And so that more space to be create, you can build a system that can kind of notify you if there's a problem and self-service people working in that way. And I think it's a big step to working on DataOps transformation. At least a lot of people realize now that there are problems in data and analytics, and they didn't solve by magic tools, that you need to work on the people and process. And you can't focus on delivering iterative customer value if your customers don't trust you or don't trust the data or don't trust your team.

And so think of four constituents to this. So there's always a team in your organization who manages the day-to-day production. Maybe it's a part-time person who's a data engineer, maybe it's a full-time group in a different country. And then there's people who are developing things acting upon data. Maybe you call them a data scientist or a data engineer or a BI developer.

They have lots of different titles. And then there are people who manage those groups. They run data teams. And then there's, of course, a customer of data, and maybe that customer of data might be a business user looking at a dashboard saying, "Oh, this is blank," or it may be a business analyst who's looking at data in a data lake and has a whole bunch of Tableau and Alteryx work that they're doing.

But they're really customers of data. And so the idea of mission control is could you give them all a shared context about what's going on, about what pipelines are running or will run? And I trust them what development's going on, so that everyone can kind of see the same thing. And much like mission control, they have all these different roles and different stations, but they're all sort of looking at the same data, so you see the same reality. And I think that becomes a good way of creating that shared context of everything in the organization that's going on, every journey, is a great way to kind of start reducing your hassles and focusing more time on creating. And let's just talk about what happens.

So in every organization that we talk to, there is a process to get data from ingest to transformation, to integration, to modeling, to visualizing, to governing, and they have lots of tools. There's lots of databases out there. You may have heard of companies like Snowflake or Oracle. There's 50 or 100 of them, and there's new ones that are happening every day.

There are lots of great tools out there to transform data. Informatica, Talend, Airflow, DBT. There are lots and lots of data science platforms out there. There are lots of tools to do visualization. There are also data catalog tools. And so you have a lot of tools in your organization. And so in some ways, what we've said in previous discussions is think of this as a factory, and each one of these factories are going on, sort of these factories or pipelines, and each part of the pipeline is a tool, and then tools have code that's actually acting upon the data and doing something.

The challenge is you get a lot of these all over the organization, and they could be batched, they could be streaming, they could be done scheduled, they could be done from an event, they could be manual, and they're kind of all over the place.

00:20:00

And kind of order of magnitude of an organization is like you can have several dozen for every thousand in your company. So if you've got a 10,000-person company, you may have several thousand pipelines going on of some sort in your organization. And that's complicated, right? And

the problem also is that those sometimes are organized in the context of how your company is organized. So let's say you're a big financial company, and you've got banking and high net worth and insurance and brokerage. Well, you may have different teams, and those teams may have different production pipelines. And that makes it-- You've got a lot of owners on who this is. And likewise, you have a process to put things into production, right?

And one of the challenges, it does take forever to get things in. And there's a lot of different methods to deploy to production. Some organizations have development and QA and UAT and prod. They may have a cloud and an on-prem. They may have multiple deployment techniques. Some cases are manual and sort of have it documented, move things.

Some cases, they use tools from software, CI and CD, continuous integration and deployment. Some cases, they may not use either of those. They may have to have a more functional or a continuous variation method, and that's my favorite. But they're all just ways to get something from a development place into a production place. And then multiple teams may have different methods, so with their own rules. So some organizations may have Tableau, and they just press a button, and that goes into production.

Other organizations may have rules and UAT and manual tests and checks and methods. And so there are many paths to production in an organization. And if you look at it from a tool standpoint, there are, generally speaking, two classes of data teams. There are sort of big technical IT teams, and they may use tools that are sort of harder, like Informatica or Cognos or Databricks or Python.

And then there's self-service teams who may use Alteryx or Tableau or Power BI. They may be doing things themselves. And then bigger technical teams may have their own set of automation tools. They may like to use GitHub, or they may have Control-M or use Cron. They open up the Azure or AWS console. They use Jenkins for CI and CD.

And a lot of people don't have the ability or very spotty sort of data testing and observability, and sort of nobody's got this sort of mission control idea. And so why-- This sort of thing, a lot of people say, "Well, aren't you just talking about monitoring? And isn't that done? That's like monitoring your database, right?

And put one of the database monitors on, and everything will be fine." And that class of software is called application performance monitoring, and it's a big sector. And so there's tools like New Relic out there or Splunk that people use, and those are great tools. And they're very focused on kind of, I think of them as IT infrastructure monitoring, Kubernetes clusters, disk, CPU, network.

And from my experience in having worked in this field in lots of companies is those are lagging indicators, almost always. And there are some cases where they're predictive, but a lot of times is if you're getting your disk, but it doesn't matter if the data on that disk is incorrect or being integrated incorrectly, or the dashboard's gotten an error. All those things happen at such higher frequency than you're running out of disk space. And I'm not saying that you shouldn't monitor disk space or CPU. Those are important things, because you could have high CPU, which means things could take longer.

But I actually think there's lots of reasons things could take longer. You should actually check if things are taking too long, and that should be your test. And then looking at disk and CPU are often diagnostic tools as opposed to it. And these tools don't actually check the data, the integrated data, or the things that are created from the data. They don't provide this context to see the pipelines, jobs, and tools. And they don't synthesize the data and kind of production and development tasks or data analytics in a coherent context that people can do. They're about servers and CPUs, and I mean, that's all good, but that's very far from, "My report looks-- Something's wrong with my report," or, "This report is blank." And so I'm going to present kind of a new idea, and pardon my language here.

00:25:00

It's called an observational meta pipeline, and I'm going to talk in the next slides about what this means. And so in organizations, there's lots of pipelines or DAGs or jobs or workflows or schedulers. Then they also coordinate lots of data tools and data stores, and orchestration tools, reporting tools, data science tools, and that all sits on sort of granular data tables and files.

And all these pipelines and DAGs and jobs and workflows and schedules do not stand alone. They're also part of a larger conflict and have antecedent and downstream dependencies. And so I'm going to use this term observational meta pipeline because the word meta means above, and it's sort of above your existing work, and it's a representation of it. And so let's just take a really concrete example.

And so when we were starting the company, I acted as a data engineer, and I've seen this with other data engineering tasks, because you sort of have some FTP file sources. You put it in S3 buckets. You kind of fill a bunch of data tables. In this case, there's sort of 30 S3 buckets going to 120 database tables.

There's a Python model that runs. You make some Tableau extracts that publishes Tableau reports. So you've got all these pieces, right? And then you've got tools. Maybe you're using the modern data stack, and you've got Fivetran and DBT and some SQL. Maybe you've got a Jupyter Notebook, and you're using Tableau, and you've got this tool level. And so all these things run at some point, right?

They have a job wrapper on it, and maybe you're using Airflow as your job wrapper with its schedule. Maybe you have a Cron job, or maybe you just press a button. And so the context of all these things at various levels is what I'm talking about to put into an observational meta pipeline to say, "Did it run? Is it on time?

Will it be late? Is it right?" Happens to be related to this whole context of the work that's being done. And so the challenge with this is that it has multiple levels, right? It's the tool level, the test level, the code level, the actual data level. And that leveling and being able to go up and down is an important part of this, but also its relationship to other observational meta pipelines or other contexts in an organization. And so you may have a data hub Where people are coming in and you're loading a bunch of raw data, but then you may have other organizations taking that raw data and building things from it, which could have other people using it. And so you have these relationships between the work and the organization, and that also fits into how the organization and projects are done.

Or if you want to help, if you find a problem, you want to know sort of who owns the problem and what project it's part of. And so you have this level issue that goes on in these that's really important to represent in some way. And oftentimes, these relationships between meta pipelines are kind of tribal. They don't know, they're not anywhere in an organization. What do I need to get done before mine?

And it's sort of, you have these fan-in and fan-out relationships, causal and non-causal relationships. You have to move up and down the abstraction level, and all this stuff is sort of tribal knowledge. It's kind of in a document or somebody's mind. It's not really active or actionable. And there's actually a nice visualization study that Airbnb wrote that talked about one aspect of trying to capture this.

And so let's give some relationship between these meta pipelines. And so here's a fan-in case. So let's say you have a data hub and you've got three streaming pipelines that are driven by events that build at the end of the day, fill up some S3 buckets, or maybe they fill up some S3 buckets and build some tables in Snowflake. And then you can have some other work that depends on those, that takes those tables in Snowflake and builds a schema on top of it, a star schema and some reports. That's sort of a fan-in relationship.

You could have a fan-out relationship where I've got a team that actually

works on a data mesh project, and that data mesh project has downstream customers from it. I'm focused on the customer domain, but then they have a product domain, or they may have a sales analysis domain. And you have these, I build something and you've got customer users. And then you may have this complicated sort of directed acyclic graph or DAG relationship. And then you may have a sub-component relationship where I've got one pipeline, but I've also got other sub-components within it.

And so, that also is very important to represent because if these workflows are complicated and being able to work through their relationships, I think is important. Likewise, how they connect to each other is interesting. And the first one in the upper left is pretty straightforward.

00:30:00

One is done, and then the other starts after that one is done. It's a causal relationship. And think of it as a step in an Airflow DAG. Others could be temporal, and this happens a lot in organizations. I've got my 2:00 a.m. batch job and then my 5:00 a.m. job that runs afterwards. And they don't really have any relationship except they have to happen one after the other. And of course, there's problems when the 2:00 a.m.

job takes four hours and the 5:00 a.m. job starts, but that's a different matter. And then there's manual ones. Like you have something that runs overnight, but then somebody's got to push a button to make it happen. And then there's event-driven, right? There's a periodic group, or you may have a streaming infrastructure where I'm getting data based on events, and I'm filling some data lake storage, and maybe I've got some Synapse tables, but then you've got a daily batch cycle that runs.

Or you may have ones that are purely event-driven, where I have a streaming pipeline, you have some events, and then another case happens completely based on those events. So it's sort of a one-to-one relationship. And then there's sort of inferred relationships that may not be there, that are actually not known in an organization.

So if you've got a pipeline that creates table one and then another observational meta pipeline that uses table one, that was a relationship, maybe it's not known, but it sure would be good to know. And so all these types of relationships between our tribal knowledge, they're not known in any organization, or they're only known by select people. And so building this sort of representation of it can be a great service to an organization. And why is that, right?

Because, the first is

the idea of a pipeline, a meta pipeline, it has expectations on schedules and durations and start times and upstream and downstream dependencies and parameters that go into it. This should happen. And then you can judge variations between those expectations. Is it late? Is it on time? Can I trust the results? Sort of what is happening now or what will happen later today?

And will I be able to trust the data or the reports? And so this idea of what is happening in reality, but versus what my set of expectations are, are important and why is that? Well, your customers care about it, right? We talked about sort of having a shared view. Well, we've done work at a number of customers to build sort of dashboards.

We call them a pulse report that actually lists, here's all the datasets, here's when they arrive or when they have arrived, what's the latest, and here's all the workflow that's happening, all the builds. And we use our software called DataKitchen Recipes, but you don't have to. But is it on time? Is it scheduled? Is it late? Can I trust it?

And looking at it from it's scheduled and when it started actual versus reality. And it's actually really important, right? Because let's say I'm an analyst on a data lake and I'm looking at 15 datasets and I've got one that just arrived and I'm waiting for it. I don't want to do my analysis and then, oh man, I'm going to get another one in.

I missed the latest version. And of course, they ask you from an IT perspective questions, when's it arriving? And so all these things, it can provide a, if you build this sort of set of observational meta pipelines, these observable meta pipelines, your customers can take advantage of it and help you see, you can use this as a way to sort of lessen the hassle of your life and provide a shared context to communicate with your customers.

And also these meta pipelines, the information needs to be stored over time. And so, you can look at variations from a historic mean. And so, for instance, here's some graphs that look at a particular pipeline and saying, "Look, is the data right? What happened last time? What happened last month?" And that can help you get to the root cause to analyze, learn, and predict. And you want to collect data longitudinally about these pipelines and build a database on it, because it's not just about mission control, it's about what happens after the mission and the data there about what happened during it, so you can analyze and improve.

And that context is useful not only for doing things that we've talked about in the past, like statistical process control on a particular pipeline, but also from the greater organizational context to figure out where the root cause of the problem is, and which group or which team or which data provider is being particularly wonky and giving you problems.

And then lastly, you need some kind of an event engine, right? Because all these things are coming into this context.

00:35:00

This pipeline started, it's ended. Is it late? It's predicted end time, the data test results. And a lot of times that becomes a lot of data. And not everyone wants to sit and look at a UI and see. They want to manage by exception. So they want alerts to go off and notifications. And those notifications need to go to different places, Slack or Jira or Teams.

And those notifications, you need to watch the signals-to-noise ratio on them. You can't have too many notifications, right? And has a particular meta pipeline has an error? Did it run? Did it take too long? Did a test fail? These kind of notifications are very important. Did a deploy not happen? And so having this idea of a meta pipeline, capturing all the events on that meta pipeline, putting it in a database, but you also need an engine that can react to what's happened and react to that difference between what you expected and what the reality is, and build rules on it. And why is that? Well, we talked about data and analytics as a factory.

And in factories, there's this thing called the Andon Cord, which is the ability to sort of stop production and then fix the problem before production goes on. And I think that empowerment to stop production is very important and empowering your employees to do it. And so there are cases when the data is so bad that you want to stop, and maybe you don't want to do what goes on. And so how do you coordinate this processing when you have no way to understand the bigger context of an organization? And so

The idea of a meta pipeline is something that sits on top of all the work, your Airflow pipeline, your model prediction, your reports. And the only way to prove that it's right is to actually check it. And that doesn't mean did it run or did it run on time. Those are also good checks.

But is the stuff in it, the data, the model, the integrated data, the reports, is that right? Because that's what your customers see. And so you need what we call tests. And the industry hasn't really settled on a term. Maybe some people call them quality control checks, or

production QA, data observability checks. There's a lot of different terms, and the industry really hasn't settled on, but I'm going to use the word tests here. And so tests check your data or things created from your data, like models. And what's interesting is that these tests themselves can work in production, but they also can work in development. They have a dual nature.

And so in production, your code, and that's acting upon data, is fixed, but the data itself varies. But in development, you have a fixed set of data because that's your test data, but your code acting upon it varies. And so a lot of these tests can perform double duty. And really you're testing, again, not just data, but the integrated data and the things that are created from data.

And so we've talked before about the idea of tests, and one type of test is this, you're running a factory, so do statistical process control. And I think of it as variation over time. If you've gotten a million rows consistently from a data provider, and suddenly they give you 10 million, that could be interesting, and that's something that you should know about, right? And perhaps that's an alert that should go out, and that you should also decide whether that's a and encore to stop the assembly line or a warning or an info or some categorization of that error so people can know what to expect. But likewise, I think testing is not a bag on the side. Having it as a bag on the side is better than nothing, but it should be integrated as part of the work that you do. And why is that?

Well, the sooner you find a problem, the more time you will have to fix it. And so having it being all at the end or having it being done in a manual way is just not an acceptable way to run when you have 1,000 of these pipelines running or 10,000 in production. And then there's lots of ways to check for consistency. And we've talked about this, and I'm going to skip through this.

There's location balance and historic balance. There's ways to check consistency over space. There's consistency over time in terms of data sets. And then find these errors and send alerts. And so there's lots of ways to create tests, and I think some people--

There's some tests that can be done based on kind of the syntax of the data. Did the file arrive on time? Are there column-level and source data consistencies?

00:40:00

And then there's ones that really are more about kind of conforming or custom tests that based on your business logic. And really, they're specific to the domain. You can't expect a drug discovery data test to be the same as one that runs in a bank or a brokerage. They're just really sort of business or use case specific. And that's really think of it as tests on the syntax of data and then tests on the semantics of how the data is used.

And so there's just many sources of tests in an organization. So you may have, again, this example, you may have-- DBT itself has a test framework. There's an open-source tool called Great Expectations. Your data engineers may have wrote some SQL tests. Your data scientists may have wrote Python tests. You may be using DataKitchen as a test engine.

There's all these different test engines on top of it. So you have many sources of tests in an organization. You don't have to replace them all. They're all good information that needs to be bubbled up into the observational pipeline. And the last thing is to see how this idea of mission control enhances data governance.

Because looking at data governance provides the case of what data it is and what does it mean and where did it come from. But I think this idea of mission control allows you to tell, can I trust the data? And is it fresh or will it be fresh? And I think having and building these observational meta pipelines and putting them in a mission control structure will help you do that.

And so in the final few minutes, I just want to give a preview of a product that we're building that's happening now and that we're working on. And so

we're building this data observability or DataOps mission control module, and our goal is to actually get something out this summer and later this year. And so let's just talk about the idea of mission control. So what we're doing is building a database to store these events from all across the organization. And this data could be test, the process's start or the process's end. Here's the schedule of the process.

And being able to look across the entire organization and drill down on it and then lay an organizational model, so a company-level overview of all the processes in an organization. And then being able to drill down between what's happening, the expectations versus the reality. So we're building this mission control. And one of the reasons that we're building it is we've actually done this in a consulting role already. Our DataOps software throws up some data that we build databases on it. But we've realized there's a lot of companies, they've got a lot of different tools and a lot of different test frameworks, and they just want to have a sort of a neutral third-party place where they can put all this mission control information. And so we've learned enough from doing a one-off that we can kind of organize it into a product. And so here's this idea of building a meta pipeline. And so here you can see the case where there's multiple tools, like you may have Airflow and you may have Databricks and you may have Tableau, and they're all these sub-components and they're sub-tests that go with it. And so we can pull in that data from all these different tools and configure expected schedules and dependencies and trigger alerts and actions based on rules and sort of visualize the estate and allow you to sort of drill up and down on this. And really that's the idea is you can have a common database to store all this stuff and be able to build events from it.

And it provides that sort of historic view of what's happened, of what's happened last week, what's going to happen today, that business UI to be able to have people look at it. And a lot of the things that I talked about in the requirements section are really requirements that we've used to build the product.

And so going on to the next part is that we realize that companies have different ways to do data tests, which is fantastic. And so they may have tests that they've written running in Snowflake and SQL. Fantastic. Let's not throw those away or re-implement them. But we also are going to build our own test suite because we've done this already at a bunch of customers. We've got our own sort of auto-generated test engine and our own way to build custom tests that you can use. And so I think all these ways to sort of fit into the idea that having mission control should be an open platform to bring in all this data.

And so what's the conclusion here in the final minutes? What are our key learnings?

00:45:00

Not narnings. Our key learnings is that you really run a factory, right? And I'm going to change that because I can't stand the error.

So what are the key learnings? So you run a factory, right? And you've got a lot of these assembly lines running in your organization. Monitor, check them, look to see if the results are coming out make any sense across every tool in your dataset in your organization. And Demi was right. The problem's rarely in a person and often in the process, so you need to build a meta pipeline that represents all these processes. And why?

Because errors are just a huge drain on your productivity. And the sort of work harder, work overtime where you're suffering as a badge of honor is not working given the surveys that we've done. And so make sure your data and the integrated data and the artifacts that you're creating from the data are right before your customers see them.

And it's really not about data quality, it's about low rate of errors. And so the two things that we think that organizations need is they need this idea of a meta pipeline, observational meta pipeline. They need to build these layers in the organization because the world's not changing. You've got 200 Airflow jobs and 50, 1,000 Control-M jobs, and you've got ELT pipelines and ETL pipelines, and that stuff all exists. What you want to do is get a view of it in one place, so then you can see if it's actually working. And so build these meta pipelines and develop a mission control is what we've set out to do.

And so if you're interested more in DataOps and DataKitchen, we've got a manifesto that talks about our ideas of DataOps, our DataOps cookbook, and our DataOps transformation book. And so in the last few minutes here, I want to see if there's any questions that people have asked.

And so there's a question: Is this module you're releasing this summer an extension to the existing product or a separate offering that can be installed on its own? And so we're seeing it as a separate offering that can be enhanced by our existing product, but not required from it. And why are we doing that?

Well, we think

that people need this sort of-- There's this concept that Starbucks has called a third place that's not home and not work, and I kind of think it's the same thing. It's sort of not development, not operations, it's mission control. It's the third place where you can keep all of what's happening in your development and in your operations together in one place.

And that ability to build a database of all those actions, the schedules, the events, the tests, the production, the development processes, and having that all in one place, I think is good, and that's why we've released it and we'll be releasing it kind of as a separate product with its own separate work. And so, yeah, and then I'm very excited about the ideas that we presented today, because I think it's something that the industry really hasn't solved, honestly, which is too bad.

It's because in this world, with a very complicated multilevel world, and so technology is a great place to do that. And a lot of organizations just struggle with, they don't know what the there is there. And we've run into organizations who are trying to build this on their own. They're trying to build database tables where they fill it up, or they have other organizations that they've done it through spreadsheets and checklists.

And so we think that by kind of building a mission control, filling the mission control with these observational meta pipelines, having people construct them, being able to infer them, would be a great value to a lot of organizations. And so it's not something that-- We just see there's a gap in the industry. We see there's a gap in different data architectures, and so I just think this is great.

And so I'm very excited about it, and

so you hit the spot where it really hurts, how to convince the CEO and CTO to move to DataOps. Well, I think where it hurts is the hassle and the pain, right? I think where it hurts for a lot of people who do the day in and day out work is this, is that they want to, but they're embarrassed by the fact that they don't know what is going on. They're hassled by the fact that people keep asking them things. And, "Is it done? What about this? I don't trust this.

Can you look into it?" And so they're spending a lot of time kind of reducing embarrassment and hassle, and that impinges on their ability to create great

00:50:00

things. Because I think people who got into data are creative people who want to build models and data transformations and get insight out of the data, and they're sort of suffering. And the root cause of that suffering is that mismatch between the day-to-day job, which is embarrassment and hassle, and the vision of creation.

And I think building this kind of a mission control is a good sort of first pain pill to take that done. And so we're very excited about it, and we think it sort of extends our mission to bring DataOps into every organization. And so,

yeah, and I think if there's no more questions, I appreciate you taking the time. Like I said, I will put the slides and the video up on our website. Next week, I'll send out a follow-up email with it, and if you have any more questions, feel free to reach out to me. I'm Chris C.

Bergh, C-B-E-R-G-H, @DataKitchen.io. And if anyone has any questions or follow-up, I'd love to hear from you. So thank you much, and have a great rest of your day.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What is DataOps Mission Control?

DataOps Mission Control gives visibility of every journey data takes from source to customer value, across every tool, environment, team, and customer, so problems are detected, localized, and raised immediately. It works by testing and monitoring every data analytics pipeline in an organization, in development and in production. The model is borrowed from NASA and SpaceX: one interface with the whole flight on it, history stored, and automatic alerts.

What is an observational meta-pipeline?

An observational meta-pipeline sits above the operational pipelines, DAGs, jobs, and schedulers a company already runs, and represents the process at the level a team actually cares about. A warehouse build might reduce to five steps: ingest data, create dimensions, build the fact table, run predictions, extract and report. The sub-steps inside each tool are noise until you need to drill down and diagnose.

Why is IT infrastructure monitoring not enough for data pipelines?

Disk and CPU metrics are lagging indicators of a data problem, and by the time they move the report is already wrong. Application performance monitoring tools do not check the data, the integrated data, or the reports and models created from it. They also give no context that ties pipelines, jobs, and tools into a coherent picture of production and development work.

What production questions should a data team be able to answer?

Whether source files arrived on time, whether the source data is the right quality, whether the report being read is fresh, and whether a particular supplier or pipeline is a repeat offender. On the job side: did every job that was supposed to run actually run, did job X run after group Y finished, and how long did yesterday's jobs take. Most teams cannot answer these without a manual hunt.

What is the Andon cord, and how does it apply to data?

The Andon cord came out of the Toyota Production System: a cord or button any worker could pull to stop production. The related idea is jidoka, which empowers operators to detect an abnormal condition and stop work immediately. Applied to data, it is the mechanism for halting a pipeline when the data or the processing is bad enough that continuing would push the error to customers.

How does mission control relate to data governance?

A data catalog answers what data exists and what it means, and data lineage answers where it came from and where it moved. Two questions are left over: can I trust this data, which test results attached to each artifact answer, and is it fresh, which process lineage answers by recording when the data and the reports built from it were last updated. The combination is called DataGovOps.

Where to go next