On-Demand Webinar · 1 hr 4 min

Orchestrate Your Production Pipelines for Low Errors

Part one of Orchestrating the Three Pipelines of DataOps. Chris Bergh covers why and how to orchestrate a multi-environment, multi-tool production pipeline along the whole journey from data access to value delivery, and how to build testing and monitoring into it. Recorded April 2020; updated August 2026.

Presented by Chris Bergh

What you'll learn 7 points
  • DataOps has three pipeline orchestrations: the value pipeline that runs production, the innovation pipeline that moves changes to production, and the environment pipeline that both of the others stand on. The value pipeline is the one that runs day in and day out, and it is where the production errors live.
  • Gartner's March 2020 survey, Data Management Struggles to Balance Innovation and Control, put only 22 percent of a data team's week on new initiatives and 56 percent on operational execution.
  • Deming found that 94 percent of causes were common cause, which is the argument for fixing the process rather than finding a person to blame. Elon Musk's version on the same slide is that the real difficulty and the greatest potential is building the machine that makes the machine.
  • Every tool in a data toolchain has its own workflow or DAG, and each category has 50 or more tools to choose from, so a production value pipeline is a meta workflow: a DAG of DAGs sitting over Informatica, Airflow, Redshift, Python, Tableau, Alation, and the rest.
  • Testing data is not just pass or fail. Three severities carry different responses: error stops the line, warning gets investigated later, and info is a list of changes. Test history is kept so statistical process control can spot a trend break.
  • A location balance test compares the same quantity at each point on the pipeline: a million rows at source, a million in the database, 300,000 facts and 700,000 dimensions in the report. A historical balance test compares this run's aggregates against the last run's to catch a shift no single-step check would see.
  • The practice to start with is a production quality circle: keep track of every error and failure, meet periodically to review them, find the patterns and root causes, and write a new test or procedure for each one. The framing throughout is no shame, no blame, and it is not about data quality, it is about low rates of error.

Slides

55 slides

Transcript

Show chapters and dialogue 10,959 words

00:00:00

Good afternoon, everyone. Thanks for joining us today. My name's Beth Befferly, I'm the VP of marketing at DataKitchen, and I will be the host today. We're very excited to kick off this webinar series on orchestrating the three pipelines of DataOps. Today's webinar is the first in a series of three, and we'll cover how to orchestrate your production pipelines for low errors.

But before we get started, a few housekeeping items. This webinar is being recorded. We'll email a recording to all participants, so please be on the lookout for that in your email in the next day or so. We'll also send you a link to the slides in that same email. Also, you're all on mute. We'll use the last 15 minutes of the webinar to answer questions. Please enter your questions in the questions box on the webinar control panel during the course of the webinar, and we'll collect all of these and answer them during the Q&A session at the end.

Finally, I'd like to introduce our speaker. Chris Bird is no stranger to many of you who closely follow DataOps. He's the founder, CEO, and head chef at DataKitchen. He's the leader of the DataOps movement. He has more than 25 years of research, software engineering, data analytics, and executive management experience. At various points in his career, he has been a COO, a CTO, a VP, and a director of engineering. Through these experiences, Chris realized that there had to be a better way to quickly deliver innovative analytics without errors, which led to the founding of DataKitchen. He's the co-author of "The DataOps Cookbook" and "The DataOps Manifesto," and a regular speaker on DataOps at many industry conferences. So with that, I will hand it over to Chris.

Hi. Good afternoon or morning, or evening, or wherever you are. Thanks for taking the time to talk with us today. So yeah, this is the first of a three-part series on orchestration, and of three different pipelines in DataOps. So it's a bit of a different cut. And so what are we going to talk about today?

For those of you who are new and haven't heard of the term DataOps, we're going to have a few quick slides to define it. And then we're going to actually talk about what we mean by these three pipelines. And then we're going to go into one of them, the value or the production pipeline, and then we're going to talk about the benefits of automating it.

And at the end, we'll talk about our next webinars, and we're going to have some slides that intermingle some what we call DataOps aphorisms, short phrases that capture kind of what we talk about. And the first aphorism that I'm going to bring up is, "What you do is much less important than how you do it." And so how does that apply to DataOps?

And what does this aphorism mean? Well, if you look at two areas where people are working on a technically complicated thing, for instance, manufacturing. Elon Musk has got this quote that he's more interested in the machine that makes the machine or the factory than the actual result. And I think that's a very DataOps perspective, and we're going to bring that metaphor through the discussion today.

And then, a man behind one of the biggest revolutions in manufacturing, Dr. Deming, basically says that when you have a problem, it's mostly often the process. It's very rarely the person or a specific case. So instead of looking for a person to blame, try to fix the process. And again, that's a very

Data Opsy idea. And so in data and analytics, we spend a lot of time on the model or the algorithm, or the data feature, or how we transform data, how we visualize data, how we govern data, even the data itself. And so there's a very opposite view here that how you do these works, not what you do, not what model or what data or what visualization.

How you develop it, how you deploy it, how you monitor it, how fast you iterate, how you collaborate upon it, how you measure is actually much more important than the things that you do. And it's a very contrarian perspective because a lot of people talk about tools and technology and data, and that's all we talk about.

And I think at DataOps, we're much more interested in people in the process and at least in this, the pipelines that make all that process possible. And so that's what we're going to talk about today. And so one of the backgrounds in DataOps is so why, and this fits my experience and why bother, why are we talking about DataOps at all, is that a lot of times teams don't spend enough time on the things that they want to do.

People are creative individuals. They want to create new ideas, create new insight, but we're caught with this blue area of a lot of errors in operational tasks and meetings and coordination, and it has a lot to do with the complexity of everything that we see, organizations, tool chains, data, collaborations. And we're not the only one who've talked about this.

00:05:00

Actually, Gartner did a recent survey, and about data management that struggles to balance innovation and control. It talked about that very point that only 22% of the time is spent on new initiatives and 56% spent on operational execution, which on the whole fits pretty closely to this estimate chart that we made. And so the emotional core in my experience managing data and analytic teams, data science teams, is that the teams themselves are sort of caught in this sort of trio of pain. One is that the data providers who are providing the raw materials are unaware that you exist as a data team.

They often send crappy, late, or error-prone data sets that you've got to turn into good insight. And then your data consumers feel like they live in an Amazon world, and so they can ask you for new cool insight every day, and you should be able to deliver. And then you've got all these different people in the organization, from production or other teams that you've got to collaborate with who are actually quite critical. And so I think a lot of data teams are suffering.

They're sort of beaten down or distraught or disempowered. And a lot of teams cannot create and innovate, or they're just defensive. And so this fits my experience in running analytic teams. I felt all these pains. And so I think in some ways, DataOps is the solution to that suffering. And instead of focusing on the new tool or the new data set, you really focus on a set of technical practices and cultural norms and architecture patterns that enable really rapid experimentation or cycle time, low error rates, collaboration across this complex environment, and clear measurement and monitoring of results.

And as a leader, it's a different focus. And that focus can then, I think, empower you as a leader to actually make your organization deliver a lot more. So let's talk about these three pipelines. So the first

metaphor is that analytic processes are like manufacturing. If you look at this value pipeline, I'm going to use the term value and production pipeline kind of interchangeably in this because, really they're about taking data on the left-hand side and manufacturing artifacts from it, charts, graphs, refined data sets, all the way out to the right-hand side to the customer.

And it goes through these, think of them as manufacturing stations where you could access it or transform it or model it, visualize, and think of the tools that are working on those manufacturing stations. They could be data viz tools or science or catalog or pipeline. There's tons of tools in each category. But the fundamental metaphor is that we are really manufacturing, and whether it's streaming or batch or big or small, you are creation.

You are running this factory of analytics. And that metaphor's going to come through this discussion quite a bit today. And the funny thing about this manufacturing line is the ownership of it is often not by one person. Part of it is owned by IT, part of it may be a data science team, maybe line of business teams.

This value pipeline, who owns it, and if something breaks, who has to fix it? Sometimes who has to run it is differently, and so the organizational structure laid across this is hard. And then, the third idea here, or the second idea actually, is that you have another pipeline in addition to what's in production, is how you get things into production, the innovation pipeline.

And that's sort of like software development. How do you deploy, how you continuously integrate. And the funny thing is, we have to do both things. We've got to run this nice Toyota assembly line, and then we've got to pick things up from that assembly line and change it and put it back into production without breaking anything. And so those are sort of two pipelines, the value pipeline in production, the innovation deployment, which is deployment.

But we've actually got a third one that underlies them both, which is kind of the environment pipeline And so this is really the... Because pipelines are processes that run in things. And what I mean by environment is the hardware, software, data, security environment that they run in. And they're kind of the foundation of each one of these pipelines. And so for today, we're actually just going to focus on the left and the value pipeline, but the other webinars are going to be about all three, the other two, how we deploy, and then the environments that go under every pipeline.

And so for today, we're really going to focus on this production pipeline. We're not going to talk about how to get development in, we're not going to talk about the environments that run in. Those are for another day. But we're going to talk a lot about how they have a lot of errors, about how we kind of construct and monitor these pipelines, how they map into organizations, and of course, how we can improve them.

So, here's another aphorism, if you don't find this too cheesy. It's really not about data silos, it's about people and tool silos. So let's take that as a metaphor. So let's look at this production or value

00:10:00

pipeline. Well, there's a lot of tools in data and analytics and a lot of tools for individual contributors to do their work, and here's a market map of all of them. And it just keeps growing, and they're open source and closed source. There's tools that are in the cloud or on-prem. It's just a big marketplace, and people love their tools.

And so, in my experience, I don't ever want to get between a data scientist who loves R versus Python, a data visualization who likes Tableau versus Qlik, someone who does data engineering versus someone who likes to write SQL or do ELT or ETL, someone who does visual or metadata-driven. There's all these great tools out there, and so there's tons of tools.

And

the second thing is that we did a survey last year with Eckerson, a great analytic company or analyst company, and we just asked this question: how many errors, incorrect data, broken reports, late delivery, customer complaints do you have each month? And most companies have just way too many. And I think that's a problem. If you're going to run a good Toyota production line, you don't want to produce cars that have defects.

And so we as a industry are just producing a lot of defective analytics, which leads to a bunch of problems. And so that's where another metaphor that we talk about in DataOps, it's really not about data quality, it's really about low rates of error. Because you could have poor data quality, which actually could lead to an error, or you could have perfect data quality and something happened in the processing.

The user doesn't know at the end, or customer, it's just wrong or it's just late, or it's problematic, or it's not right. So really focusing on error rates as opposed to data quality is another sort of aphorism that we talk about or perspective in DataOps.

So, you've got a lot of tools. We've obviously got a lot of data, got a lot of errors in production. But these teams who actually create things that go into these pipelines kind of work and think of it as an integrated value chain for their customers. And maybe it's someone in the IT, in the data center providing data, some data engineers transforming data, data scientists modeling it, putting algorithms on, people doing data visualization and governance.

And if you look at these people, this value chain, there's lots of tools that go on. Here's kind of an example of one where if you look at it, sort of data comes in from a service or CSV files. It gets in

some kind of lake. Maybe there's some MDM applied. It gets into, does some transformation in Talend or SSIS, does some data science, ends up in a database, actually ends up in being visualized. It gets in a wiki. There's just a lot of different tools that go on. You may have some of these tools, you may have a completely different set, but each, that value chain has an associated tool chain that goes with it.

And this is a term that we sort of borrowed from DevOps. And if you look at what those tools do, like when you're getting data or accessing data, there's servers and storage and databases and software and FTP. There's a lot of great tools to do data engineering: Airflow tools, ETL tools, Informatica, tools to do it on streaming data, et cetera.

There's data science tools, Python, R. There's data visualization tools and self-service data prep tools. There's data governance tools, and if you go and think about, I just want to put a new table and have a report on it for my customers, I've got to access it through IT. I've got to put it in a schema and a database and data engineering.

Perhaps I've got to, in data science, I've got to segment that table. I've got to visualize it and get it to my customer, and then I've got to actually keep track of what those columns mean and where they came from, and I do that in a catalog. And it's just putting a simple table and to making it visible to your customer, you've got a lot of tools and perhaps a lot of teams that are involved in that, and so no wonder things are slow.

And here's a perspective, that since everyone's got their own tool, right, to do that, they're working independently in their tool, yet they're dependent on what someone else does. As a data engineer, I'm dependent upon IT to get that data somewhere. And if I'm a data scientist, I'm dependent upon someone in data engineering. And we each own our part, and they have a specific task that they're working on. And what I want to do is say, each think of what every tool does.

Every tool has some format of the file that it creates. It's a Tableau workbook or a workflow. And think of those tools as having their own series of steps that they have to happen, their own workflow, their own DAG. And a DAG is a directed acyclic graph.

00:15:00

It's a graph of steps. So every tool is doing some things in succession. And maybe it's just like a Jupyter Notebook, where it's one step after the other. Maybe it actually is a directed graph, and Airflow is like that. And so because you've got a lot of tools and you have a lot of workflows or DAGs, you need a meta workflow, a workflow on top of the workflow to organize the production system.

And so, the production value pipeline is really a meta workflow or a DAG of DAGs because we've got all these tools, and because each one of these tools has their sub-workflow. And so that's an interesting idea, right? And so,

but I think this is true because most organizations have two or three different ways that they want to do their data engineering, a couple of different biz tools, a couple different data science platforms, to make this work. And so another DataOps perspective here is that there's not sort of too much data in the world. It's that there are too many people and too many tools and too slow processes to take advantage of that data.

So think about that. So let's talk about given the situation with our production pipelines. Let's talk about the benefits of automation or how can we improve them. So the first idea is think of this as...

Think of your production as a manufacturing line. And so when I managed back in 2005, 2006, we had thousands of users at a different company using our analytics, and I just would get phone calls as a COO when things would go wrong, and I really disliked having defective merchandise. And I sort of read a little bit about this book, "Machine That Changed the World" in college, and I'd heard a little bit about Deming, and I started to actually research into it. And this idea that if you think of that we all work on this technically complicated thing together, that's an assembly line.

And if you focus on the processes that act on it, you can have increased product quality. You reduce the amount of rework that you do. You end up with higher employee satisfaction and therefore higher profits. And this sort of lean manufacturing idea, total quality management, a bunch of different words, has really been transformative and it's gone from kind of something that just the Japanese did in the '80s to being standard throughout the world. And it is at its face, sort of a governing philosophy.

And so in production pipelines, what we want you to do is think of them like a factory. And so this diagram at top is illustrative. So you've got your production data going in on one side. It goes through these series of steps that we talk about, the manufacturing line, and that's on the other side, you've got your analytic customers And think of a big stoplight at the top. When this is running, you want to have a green light that says everything's working, and my customers are not going to find any problems with the data that's happening, going through the processing.

I know that it's going to be right, and I know my customers are going to be happy. And maybe you have a yellow light saying something's weird. So think of these tubes as pipelines. And everyone's working and everything is working in a pipeline, whether they believe it or not. And there's lots of tools and lots of data go in, and lots of customers on the other side. And if you think about it, companies have hundreds of these pipelines running, all over the organization on a day in and day out basis.

And one of the most important things here is that you need that light at the top saying, "Hey, it's working or not." Because if you don't know, if you just hope it works, then you end up in this cycle of rework. And it's just like in a factory, if you're producing cars with defects, it's so much more expensive to fix the defect than rather fix it at the first place.

And if you think of all these pipelines themselves, they're all interacting with your tools and data. You've got your database tools and science tools and catalog tools and ETL tools and databases and data lakes, and they're all trying to work with these tools because it's a DAG of DAGs. They're each doing their part of their work.

And so how do you then, if you want to have this sort of stoplight on top that's going to tell you whether your customers are happy or not when all these pipelines are running, what does that mean? How do you actually do that? Well, the first thing is you need to add automated tests and monitoring in production. So you need to test every step and every tool in your value pipeline, because basically hope is not a strategy.

Don't trust your data suppliers, don't trust your servers to run. It's sort of trust but verify. And I think you need to put on top of your

00:20:00

entire process a set of tests or logic to see that, am I getting bad data in? Is my business logic still correct based on that data? And are my outputs consistent for my users, or is there some big variation between what a business user saw, current versus previously? And I actually think that these sort of tests are actually a really interesting way to look at and get analytical about your production, and we'll talk a little bit about that in a bit. And so, here's an example of what, in our product, we call a recipe, which is a DAG, a workflow, a series of steps. We'll use all those names kind of interchangeably.

And here's an example of a customer that was trying to build basically a sort of a data warehouse. And these numbers represent the numbers of tasks. And you can see sort of create dimensions, create facts, create source tables, and all these things are happening. But one of the ideas is that if you have tasks that happen earlier in your production process, and let's say something happens in the creation of a dimension, well, if the dimension's empty, everything else doesn't make any sense. You can do all your work, but why not stop what's production, learn about that sooner, and maybe you can correct the problem before it gets out and you're not late. And again, think of an assembly line.

If you're going to put a wheel on a car and you notice that it's flat, well, don't put that wheel on the car. And maybe even stop the assembly line so you can get another wheel to go on it. And so I think that's an important part of the perspective here is that don't trust that your system's just going to work. Hope is not your friend.

And when you think about trying to monitor what's happening, it's not just things like server and CPU monitoring, and those are important, but it really is getting into the data and the artifacts that are created from that data and then classifying them. If something's wrong, if you've got a file that you expect is going to always have a million rows and you've got 10 rows in it, well, stop the assembly line, pull the virtual Andon Cord, and raise a red alert and get people, someone to look at it because that's going to affect everything else.

And don't wait for your customers to find out the problem, and don't live with the idea that that's just acceptable. And be able to find that there's these errors and warnings before your customer does. And then another feature in our software is that, once you find these problems, you can actually alert and maybe create a Jira ticket, send an email, send a Slack message, because everyone has an SLA.

Every customer, once they start getting a report or an analytic, gets used to it coming at a certain date, certain time. And some of our work, and sometimes it's hard more weeks than others, that you've got to patch work or kind of run around. And the more time, the earlier you have it, the more time you have to patch it.

And perhaps, it's as simple as asking the data provider to send the data again, or maybe it's more complicated, where you have to write a patch on the data. Or it's maybe so complicated that you have to stop and wait for something to be fixed. But it's important that you know about it before your customer does.

And there's just a lot of tests that can happen on data. And there's tests that look at the data itself, so the whole class of data quality profiling tools that'll profile the data and look for things like, does this column have three values in it? And you can test to see if there's three values.

There's business logic tests, and you can write those tests in lots of different ways. You can write them in SQL, you can write them in your favorite tool, and there are test engines out there. And from our belief in our software, we think that you should be able to write tests using whatever tool you want.

And so if you're a SQL person or an ETL person or you like to write MapReduce jobs, you should be able to write your tests there. And so one example of a test that is one of my favorites is called the location balance test. And I was talking with a customer last night who's working in Azure, and they get data from internal systems. It gets into a sort of blob storage.

It gets into three levels within their Snowflake infrastructure, and then it gets cached in some views in Power BI. So you've got basically the same data item in four or five places. And so how do you know that something hasn't gone wrong in each one of those steps? And so we have something called a location balance, where you can basically say there's a million rows in, and you're going to get the equivalent of a million rows out. And what that means is you've got to keep track of variables on this. You need something, number one, that goes across the whole system. You need something that can keep counts and variables and sub datasets that you can do comparisons across. And so our tool provides that.

And if you look at it, here's a way that we wrote one, and here's the case. You can see it's stop on error if final table row count not equal to the expected

00:25:00

row count. And you can see the logic here that the final table row count should be equal to the expected final table row count. And so this is a case where you just want to make sure that if you're moving data from a blob store to a database table or from one level in a database to another level, that the counts are the same.

And so it's a very simple way to make sure that you don't have any losses on these transforms.

And another one of my favorite tests is, and this is another aspect of running a good production pipeline, is you want to make sure that you produce items that your business customers know is right. And business people and consumers, they're heuristic. They don't know all the details in the data, but they do know certain things very well.

Like they often are business people and have 80/20 rules, and so 80% of the value comes from 20% of their products or product groups or regions or customer types. And so I've had the problem of putting data in my past in front of customers and having them within three seconds say it's wrong, and that's very deflating.

And sometimes it's wrong not because the source data is wrong. Sometimes there's cases where there's just small files with hundreds of rows that affect it. And here's a case. Here's one type of test that tries to capture that business heuristic. So if you look on the left, there's what's in production that's live, and then on the right, what's in pre-production.

And so if you look at this table, and let's say it's a hypothetical table, there's SKUs and products, and there's a grouping of products. And there's volume on each, and the volume's 575 for this group, and they're grouped up in these two groups. Now, let's go to the next day or the next week.

I've still got these products and the same SKUs, but notice the product grouping has changed. However, the total volume is still pretty much the same. It's 575 or it's 587, so if you compare just the total volume, it's probably within the lines of what you're going to see every week. But look at comparing the grouping here.

So as a business person, I know that group one's got about 225, group two got 350, but suddenly group one's got 358. That means that that's your competitor's product, and suddenly your competitor's product went up 150% in one week. That's how a business person would know what's wrong, because they memorize what their competitors do. So one example of a heuristic test that you should run is just compare these main items between what a customer has said.

And I don't know how many times these small data files that product groupings or even the data that comes from suppliers have by mistake, have given an error. And so we need to be able to create tests that represent this. And so again, what happens if this is a problem and you get it?

Well, I think there's a philosophy here that it's not having problems. You're always going to have problems, and it's better to see it as instead of as a failure, that you did something wrong as your team, seeing it as an opportunity for improvement, that you can get you and your team to say, "Okay, what did we learn?

How can we implement a test or a change in the process to do it?" And I think that's actually a really important thing, because the data world is always going to have failures because data providers are always going to give things that are broke. And so you can help notice those before, or even if you don't, you can have a way to improve.

So let me look and give you just an example of what I mean in our software here for a few minutes. It's 1:30, so I'll try to do this in maybe 5 or 10 minutes. So if I go to our software and I've got a simple example, and you know, we're DataKitchen, so we've got this overuse of food metaphors.

So there's this thing in these four steps. They're called a recipe. This is our workflow or DAG. And this is a really simple one. It's got four steps. And most of these have hundreds of nodes. They're in a graph, but it kind of represents the way we think of the world. And this one, it's not doing everything in analytics. It's just loading a database.

So again, a super simple example. But even in the simple example, it's complicated. We've got some scripts that run in Windows, we've got some Python code, we've got SSIS, which is Microsoft's ETL tool, and SQL servers. You got four different tools here running to do this work of loading things in the database. And when these run, we want to be able to monitor what's happening.

Not only did these run correctly, but we want to sort of dig in and find out. And at this end, this number here at the end is a three that represents that there's three tests. And so when the system runs, it actually provides a good set of data. And we actually keep track of all the processing that has happened and the counts and even the code that was acting upon it.

And that creates something called an order in our system. And the set of orders or order runs in our system, we group together, and it

00:30:00

allows you to start looking at trends across what happens with your order runs. And if I go to this one on the right, I can see some time on the bottom. So the time it ran in volume. And one of that sensors is just a simple row count. And all we're doing is counting. And then it suddenly drops, and then suddenly it goes up again. And hey, that may be fine for your data or hey, that mean that this day that you start getting that call from the CEO, that his dashboard's wrong and suddenly you're scrambling on Saturday to fix it.

And that's another reason why we want to have this company. I'm pretty sick of scrambling on Saturdays, and you should be able to find this stuff out before it goes to your CEO and make sure that it's right. And so the whole point is notice these things and be able to set warnings or failures or logs and have alerts that go off, because if something's wrong, you want to know about it. And this idea of statistical process control, breaking bounds, is a very lean manufacturing concept.

Likewise, even looking at timings, seeing if something's taking longer. And I can't tell you the number of times I've had people say, "Yeah, it ran in two seconds. That was unusual." That means the data didn't show up. And well, that's probably a good way to alert and find out if something's wrong. And so, we have these recipes, which is our abstraction.

They can run from a schedule, or they can run from event, or you can just sort of hit a button and say, "Run it now." And I did that, and when I ran it now, it creates a simple order. And this order run actually got an error in it. And this is also part of our aphorism is that you want to love your errors and find out your errors before your customer has done it.

Here, we've got a system that's ran, but the very last step didn't work. And, what does that mean that it didn't work? Well, if I look at my test results, this test, this count raw orders row failed. And so the system was running, maybe every tool gave a correct finishing code, but the data itself was wrong.

And so we're digging into the data, or we can dig into the artifacts that create that data to be able to tell. And likewise, when you find it's wrong, maybe it's not the test. Maybe there's something that it did go wrong with the system. So we bring all your logs together. So just to finish out this short demo, what does it mean to actually dig into the data and look at it?

And so there's lots of ways that people like to do their data work. As I said, some people like to do MapReduce jobs, some people like to write Python, some people like to write use tools like Airflow or tools like Informatica. I like to write SQL because I'm kind of a nerdy guy. And so I interacted with my data by writing SQL. And so this is the source.

There's lots of different tools and ways that we integrate with tools, but this one is just simple. It's a select count. And that was actually pushed into our test framework. And here's how that went. That came in. It was a stop on error test. We took it from that key, and then we wrote this very simple logic.

It's a count that's greater than a hard number. And this number doesn't have to be a hard number. It could be a variable, or it could come from the history of a variable. And so, why does this all matter? Well, let me go back to the PowerPoint to explain. So

you want your production pipelines to run automatically. You want them to tell you if something's wrong before your customer sees it, because you want to lower your error rates and embarrassment. So first of all, test and monitor automatically on top of production, and even behind that, meta orchestrate all your tools so they can run automatically.

And then do it on top of your entire tool chain.

And so I know there's some questions here, but I think I'm going to keep going, and we'll try to hit those questions at the end because Beth is going to collate them. So, we want to be able to send alerts and notifications and, first of all, keep track of the history of this running system, because that history is actually a good source of analytics and make it easy.

And then we talked about these different test types. But again, why do you want to do this? Focusing on low errors actually has this ability to give you more innovation, because in some ways, it freezes out the fear of people. If people have less fear, they're more willing to try things, and then your customers will give you more trust because they'll stop having so many errors, and it leads to less stress and less embarrassment of your team.

So there's this almost to us, it's counterintuitive. If you focus on errors, you actually end up getting more value for your customers.

So let's keep going. And again, another aphorism, "No shame, no blame, love your errors." And this is actually a phrase I've used quite a bit, because in a lot of organizations, when errors happen, it's seen that you screwed up and something's

00:35:00

wrong with you. And so I think one of the ideas of Deming is that, we're all sort of touching the elephant in these big, technically complicated things, an assembly line, a big analytic production process. And so no one can have the whole thing in their head. And so when there are errors, errors are going to happen. And don't blame someone.

Try to find the root cause and try to find a way to fix them.

So let me go on to the next slide. So one of the things is, if you imagine that you've got these production pipelines and you've instrumented them, they're running automatically. A lot of analytics teams aren't very analytic about their production process, and so let's keep track of some things. What are your error rates in production? What part of the pipeline?

Which pipeline are they running on? Which data provider is giving you crappy data? How are you meeting your SLAs? What's your coverage in the tests? And these are very important metrics. And if we're trying to help people be data-driven, and it's kind of hypocritical if we, as data and analytics teams, aren't analytic about what we do.

And this one is a report that actually comes out. It's called Tornado Report. And it shows on one side, here's the weeks and here's the different data providers. And, we built this in conjunction with Jira, where saying- On one hand, what's the severity of the error and where did it come from? And then the other, how many hours of time that took to fix it.

And this gives you leverage on your data provider saying, "Look, for the last three weeks, we had a 71 error because of your data, and you cost me three days of time of my team to fix your data problem." And I think that kind of quantitative discussion, as a leader, can help you provide leverage on your data providers.

And then second, looking on a project basis, or even on a pipeline basis, you should try to look at over time, and here's the middle bar, look at how many errors you have in that pipeline or how many tests you have in that pipeline. And try to look at, is that pipeline on time?

And very simple things that I think can help you look, and if you look across all these, you can start to find some patterns. And another part of metrics is that by measuring, you provide evidence to change the behavior of your individual team. You say, "Look, we've had these errors. Here it is in the report.

Do you believe this or not? Let's try to focus on reducing these errors." And it becomes less of your opinion and more becomes here's the facts, and we can all rally around to fix these facts. And here's another aphorism. So do you know what "don't be a hero" really means? Well, it doesn't mean be heroic. It says don't solve problems.

Figure out how to never create them in the first place. Again, a different philosophy. A lot of organizations in data and analytics, they praise the individual contributor who worked the weekend to fix the problem. And yeah, I think that's good. People are good. But figure out how not to have that problem in the first place.

And so one of my ways of doing it is, yes, if you've got an individual contributor work the weekend to fix it, that's great. But then as you talk to that manager and tell that manager, "What the hell are you doing? Why did that happen? Why is that person working? How are you not going to make sure this doesn't happen again?" Because the manager owns that problem. The individual contributor being a hero is wonderful, but the manager owns the process. And if there's a process problem, that's a problem for the manager, and they need to figure out in conjunction with their people how to make sure it doesn't happen again.

So what are the technical characteristics of good production pipelines? Well, the first is this idea of meta-orchestration. You've got all your tools, you love them, plug in your existing tools into the pipeline. The second part is that these pipelines themselves need to be monitored, tested, observed. There's a term from DevOps called observability. You should automate those tests. You should keep track of the test data and artifacts. You should send alerts.

You should do things across the entire pipeline, statistical process control, end-to-end, and other test types. And then themselves, store the pipeline definitions as code, and keep your tool code and pipeline code together. And then once each one of these

pipelines run, keep a database of what the run history of the operational metrics, because you can use that to improve.

And so that's the technical characteristics. So let's look at what are the organizational characteristics of good production pipelines. Well, first, everyone's got a development team, data scientist, data engineer, and a production and operations team. And so having a good relationship with the team and being able to share kind of common artifacts like here's the pipeline, here's the run results.

00:40:00

We can both look at it. We both can see what the problems are. We both can fix them quickly. Because your production team always hates errors. They want things that run automatically, easy, with low errors. And that relationship, don't see it as I'm throwing the things as a data scientist or a data engineer, I'm just throwing it over the wall to your production operations team. You need to have a set of shared objects and a shared relationship. And part of that is a reflection and improvement mindset.

And I'm going to talk a little bit about something called a quality circle that you can do right away. And then kind of be able to have fast reaction to the inevitable issues. And I'm going to talk about something called mission control. And there's another idea called safety culture, and that means if there is a problem, you should feel safe to bring it up and not shamed or blamed or bring it down. And then finally, I'm going to talk a little bit about trust in team members, or the sort of deference to vendor control.

So first of all, we've got a development team at the bottom. There's Eric and Priya, they're kind of data doers. Maybe they're a data scientist, data engineer, data visualization. They want to create insight out of data, make their customers successful. They've got their favorite tool, their favorite language. But then in production, you've got Eric, who's our production perfectionist.

He wants to protect and perfect the daily grind of delivering data. He wants to minimize errors, and he's a bit of a taskmaster, right? And this relationship between these people, the developers and production, they're a team, and they need to work together. And partly, I think one way to do that is to have a quality circle in production.

And you don't need-- our software can help you do that, but you can also just do this with a spreadsheet. Keep track of every error that you had in production. Put it in a spreadsheet. And then on a periodic basis, meet with a team to review those errors and sort of find patterns and root causes.

And then as part of that meeting, just pick one or two things and create a new test or a new procedure to mitigate those problems. And then focus on error reduction and also focus on your team's psychological safety, that if they bring up a problem that they're not going to get stepped on or blamed. And have both dev and ops teams as part of that.

And have everyone try to look at it because we all have different perspectives on it. And maybe you do this every three weeks, and it's a spreadsheet, and it's a half-hour meeting, and try to assign someone to fix that. And then you'll see actually gradually over time that by focusing on the psychological safety, by taking those errors and just making them visible on a spreadsheet, you'll actually start to have it.

And this is just a super Japanese management technique that is, I think you can do with zero cost, and you can do today. And it's all focused on lowering errors in production.

So a second thing, and this is a bit more philosophical. So there's this book by Dostoevsky called "The Brothers Karamazov," and there's a part of it that's the grand inquisitor. This guy basically says people can't be free. You've got to put a boot on their neck. If they're free, people's lives suck, and people need to be controlled because they can't be trusted with freedom. And so, I was talking with a vendor last night, and they want to have the Uber tool that does everything.

And they've gotten a lot of startup funding. They're really quite enamored with their tool, and perhaps for good reason. But the promise of a single package offering that can do everything in data and analytics is years away. And a lot of vendors don't trust you with the freedom to use their own tools. And so I don't believe in that control.

I think, like the grand inquisitor, you should be able to have the freedom and the choice to use your own tools. And so don't believe that voices of control, like Dostoevsky's grand inquisitor, that there is one tool, and if you just standardize on one tool, magic will happen. And so I guess beware of the unified platform or beware the tool vendor that's selling you a promise that everything happens. And there's a lot of tool vendors have adopted the DataOps terminology, and I think that's great, but are they really doing DataOps or just putting the DataOps on their existing tools?

And then another one is set up a mission control. So if this diagram's got a lot of pipelines running, and you see all the pipelines running in a different part of your organization, well, you should monitor those and have, maybe it's your production organization who's monitoring them. Maybe it's you, maybe it's your team, but try to look at what's happening.

And I think actually making that visible, having reports, having views, I think is an important way to rally whether things are working or not. And again, our view is it's not that they're just running, it's that the data and the artifacts that are creating on them are correct so that your business customer knows that that's going to be right.

00:45:00

And so the stoplight has a little bit more meaning on it.

And so

that's it. It's 1:45, and we've got two webinars coming up next. One about the deployment pipeline, how to deploy from dev to production, and how to adapt these ideas of CI and CD from software to DevOps. And then actually the next one is how to actually set up your environments and what that means, how to create reusable environments.

And I'm going to stop there and then answer, see if there's any questions.

Thanks, Chris. That was great. I hope everyone found this informative. I encourage everyone to enter their questions into the question box on the control panel, and then we can get to those. I did see some questions. And so where are they?

So I saw one question about Kubeflow, which is the Kubernetes way. And so there was, I don't know, somebody asked a question about Kubeflow docs, and I think you just have to go and find out where they are. And so,

the question about, is

it possible to get the slides? And sure. We're going to be able to deliver the slides and the audio of this. And

there was a question around do we support Kubeflow, and the answer is yes. And also we have an example of SSIS.

And there was a question around the actual testing in our software. And so the question is, I think was,

how many tests should you have? And should those tests be across every node in the pipeline? And I actually think that that's true. And this demo, it's meant to be simple, but

the first one I showed, the first demo, it had some zeros here. And the idea is that we think that having zero tests on an individual node is not great. You should try to have several tests on each node. And so that way the tests kind of span everything. And why is that? Well, the sooner you find a problem in the processing of your analytics, the better. And that way, and if it's such a bad problem that it doesn't make any sense to go on, then you save all that work.

And perhaps you can then stop it, patch it, and fix and restart it. And so we think that having tests across all your nodes is a good standard practice. And perhaps our demo doesn't show that. And that's also why in one of our-- We also try to help people keep track of, and if you look at this graph on the right, trying to look at how many tests you have per node, trying to get at least one test per node or the total number of tests in a project.

And if you think of I'm a manager, how do I know I've got enough coverage on testing? And in software industries, there's actually coverage tools that do that. And, I think it's something that we're working on, but we've got a first-order approximation of it. We're just looking at how many tests you have and how many nodes and how many keys you have to get some gross metrics on it.

So, there are questions here, but I'm just having a really hard time.

So the question is, so are the tests written within the DK code and then DK consumes them, or are those tests written outside of it? And so I think the answer is, from our perspective, and is out there,

if we look at the way our tool is written, and if I go into the... Think of the test as kind of having a data-gathering part and then an execution of the test part. And the data-gathering part is written in whatever tool you want. So for instance, here's a Python Anaconda node. And so we're writing something in a Python notebook, but at the very end of that notebook, we're creating a file that's actually pushed into a container that we're

00:50:00

actually entering in and creating a test from. So every

tool itself, we allow you to use whatever tool that you want. And this goes with my philosophy that people love their tools, and so if a data scientist does all their work in Python and Jupyter Notebook, they're not going to want to have yet another tool to work with. And so if you can build the tests themselves, or at least the logic of those tests or the data-gathering part of those tests in their tool framework, they end up actually being more productive, and more willing to create the tools.

And so there's a question of,

so, what is the scope of influence I have?

So what if my scope of influence is only to lead an ETL team, but data scientists are in a different department, do I implement tests within my area of influence?

Yeah, I think first of all, you should test, and test within your area of influence. And the challenge is how do you get other organizations to adopt a testing framework? Because one of the problems in data and analytics is the

local global problem, right? If you are a data scientist and/or data engineer, and you only own part of the process, right, you only own maybe the data part, but then you've got your consumers and they're not testing, well, then you're just hoping that things work. And so I think it's better if everyone works under DataOps principles. But if you can't control them, I think the first thing you want to do is sort of make your bed in the morning. Make sure that you are running a good production pipeline that has low errors, that has automated testing, and then try to influence those data scientists to adopt those same principles.

And the best way that you can do that is that they're going to complain sometimes that you've made a change and their code broke, that their model broke. And so you could say, "Well, look, if you added some automated tests that I could have run, I could have known that ahead of time, and then I wouldn't have broken your feature code and then therefore your model." And so I think that's a good way to help those teams. And also using the kind of metrics that we talked about. And I think there's just having a discussion with people about these principles, because oftentimes people aren't being, sort of Hanlon's razor, people aren't being malicious, they're just not paying attention, and they're really focused on other tasks.

So I'm going to keep going here.

The question, are the slides available for download? We're also doing a recording. Where do I get started? What are the first few baby steps? Well, I'm a fan of the quality circle, honestly, because I think errors lead to a lot of poor behavior on people. And if you start people on the path of loving their errors, you actually set it off, and that was my personal experience back in 2005 and 2006.

Obviously, we are a software company, and we want to sell software, and we think talk to us and buy our software, but we also think that you can start with doing even a simple level of automated tests, quality circle, and just writing some tests. If you're an ETL engineer, just write some tests and throw the results in a test table and say, "Is it right?" And I think those are the few baby steps you can do today.

But it'll get you on the path of making sure that things work.

So, is the database to track operational testers from something that's part of the DK tool or something that has to be built? It's part of the DK tool, so we have what we call an order run database that keeps all the orders. But we also, given that we know about analytics and that typically people are going to want to look at not only the results of DataKitchen, but also maybe mix in their Jira tickets. We have an integration to Jira, and you can build databases from the combination of Jira and ServiceNow and DataKitchen and whatever other systems you want to do to build your own reports.

But we do have our own database for keeping track of all those operational metrics. But, we also try to practice what we preach and be a good citizen of the data ecosystem and know that people to really understand what's happening with their process, DataKitchen's just going to be a part of it

00:55:00

How do we go about beginning the transformation to DataOps? And so, to me, I think what I answered before is true, is just do something, right? Add some tasks, create a quality circle. Start focusing on stop trying to hide errors under the rug, and just list them and get people-- so there's a mental part of getting your team psychologically prepared. There's a technical part of just trying to do some of the small sort of automated testing, across the project that you can do.

We've talked to organizations who've even done things like a book club on our book. But I do think you should start small because I think

it also helps build momentum for DataOps. And we've tried to be good idea mavens here. We've got a book and presentations and discussions that we can do. But I'm a believer in just start small. Don't buy some software, just start doing it. And I think you'll start to see results even with the small things.

So does the source code in the end-to-end DAG of DAGs have to be in a single repository?

Yes and no. I'd like it to be, but if it's in two, does it really matter if it's having it in three? Having it in source code, I think, is the most important thing. Having it all in a single source code, I think is better, but not essential. And so you do have to have some more manual steps, but keeping track of all your DAGs in source code, and being able to branch and merge across those is important.

And so having them all in one is, I believe it has effects because having a unified view of the pipeline, across your entire organization and having to know where it is and instead of having to know who to talk to and know which version, I think taking complexity out of systems actually makes teams more efficient.

But having things in source code in the first place when they're not at all, I think is even a good first step.

So how do you think about how granular a node could be? For example, could you have an ETL node that could contain dozens of transformations? Better to have granular tests or at a high level? That's a really good question. So, a lot of people have built thousands of hours of work in their ETL process.

And you can treat that all as one node, call all my Informatica jobs. And that's fine. And in fact, if you call all my Informatica jobs and then you put a bunch of tests at the end of your Informatica jobs, great. As long as you're testing before your customer sees it. Now, the problem is, if those Informatica jobs take four hours and you may want to break them into two chunks and put some tests in the middle because you want to find out where something is before that four-hour process happens.

And so decomposing them into smaller pieces, into more nodes is not required, but it does allow you to dig into the data more, find out if there's problems sooner, and then localize where the problem is, in a quicker way. And so, because remember, we're going to talk about in our next time that these tests are not only useful in production, they're also useful in development and trying to find where problems are in development, it can also help you center in that fast. So you don't have to, but I think it's better to decompose it. And, I think first of all, if you have no tests, putting tests as a bag on the end is automatically much better than none.

And do that first and then second, start breaking it apart into pieces so that the tests themselves can be intermingled in with the production process, so you get more visibility in development and production. And

could you talk more about the internal state management of DataKitchen? Is it primarily for programmatic retry capabilities on a task level, or can it be thought of as a form of data versioning where the state is persisted for longer periods based on configuration?

I'm not entirely sure I get that question, but so we have a thing called a recipe. The recipe itself is the word in GitHub along with all of that. And then when it executes, it has the state of its executing, is kept in real time in another database, in Mongo, called an order run.

And so if it stops, we can restart it. It keeps track of the process lineage, all the code and the versions happen, but it doesn't particularly data versions. Because we're a process tool and not a data tool. So one of the ways that people will create

01:00:00

repeatable systems is that they will take, and have pointers to where the data resides as variables in our system. So for instance, the data's in 20 of these S3 buckets, and these S3 buckets are based on certain time variables. And the combination of those time variables, plus the process pulled out of Git, gets you a reproducible way to run.

And so it provides a version of data versioning, but it's not sort of a Git for data. It's versioning by using the pointers to the data in the various bucket stores or the various places that it is. And, so that state is persisted,

in Git, for each run, and then the as-run state is kept in a database. So that's our order run database. So hopefully I answered that. It's a very good technical question.

How can you introduce human-in-the-loop feedback testing into the process, particularly at the end of the value chain? When the client customer is a real person, the test can be more ambiguous.

Well, first of all, fantastic that you're actually having a real in-person test. And so, it's better to start with a manual test than an automated test. And so ways that some people have done it is taking our graph, and at the very end when it runs, we can create some Jira tickets, and that person can go say, "Go run your manual tests." Or they do it by time.

They say, "Okay, the building of the system's got to be done by Friday at noon, and then the manual testers come in from 12:00 to 3:00 and check things." And so you can do it either by events and Jira tickets, or you can do it by time. And I think the idea here is that even with ambiguous tasks, that you can start to see patterns. And a lot of times the way people who are doing manual testing are doing it is they're extracting data, they build spreadsheets, and they're kind of building automated tests there. And I think it's that you can start to encapsulate some of those ambiguous things into tests. And therefore, you can start to think about trying to, if I've got a three-hour block of manual tests I've got to do, start thinking about how I can get that down to 90 minutes, to an hour, and replacing some of that work with automation that's run at the end. And I think that can be a help.

But I think

starting with the ambiguity and then trying to remove the ambiguity with automated tests is a good way to get on the DataOps path.

So I think we've run out of time. So, Beth, I've answered all the questions that are listed here. Thank you, Chris. Thanks for answering all those questions. I'd like to say thanks to everyone for joining us. As I mentioned, we'll be emailing out the recording and the slides within the next 24 hours, so check your email for that.

And please don't hesitate to reach out to us or to Chris if you have any additional questions. Have a great day, everyone, and stay healthy and safe.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What are the three pipelines of DataOps?

The value pipeline is production, the work that runs every day to turn source data into reports and models. The innovation pipeline is deployment, moving new code and configuration into production. The environment pipeline creates and manages the environments the other two run in, which is why it is described as the foundation for both.

What is a location balance test?

A location balance test checks that the same quantity holds at every point along the pipeline. If a million rows leave the source, a million should land in the database, and the report built from them should account for all of them, for example 300,000 facts and 700,000 dimensions. It catches silent losses that a check on any single step would miss.

What is a historical balance test?

A historical balance test compares an aggregate from the current run against the same aggregate from a prior run. If product group G1 totalled 225 last run and 358 this run while G2 moved the opposite way, the numbers have swapped groups or the mapping has changed. The check compares pre-production data against production data rather than validating a single value in isolation.

What kinds of tests belong in a production data pipeline?

Five types are named: traditional data quality checks, statistical process control, location balance, historic balance, and business-based tests. They should run automatically on top of the whole toolchain, send alerts to Jira, email, or Slack, keep their history, and be easy enough to create that people actually add them.

What is a production quality circle?

A production quality circle is a recurring meeting where a team reviews every error and failure it has logged, looks for patterns and root causes, and writes a new test or procedure to stop each one recurring. It includes both development and operations people, and it is run for error reduction and psychological safety rather than for accountability.

Why is a production data pipeline called a DAG of DAGs?

Because each tool along the way already has its own workflow or DAG, and each group of people owns one part of the value chain. Data engineering runs Airflow or Informatica, data science runs Python, visualization runs Tableau, and none of them see the others. A meta workflow above all of them is what organizes those separate DAGs into one production system.

Where to go next