On-Demand Webinar · 47 min
Data Journey Is the Missing Piece
The modern data stack is complex, the toolchains are fragmented, and the data changes constantly, so you cannot tell whether what is happening right now is what should be happening. Chris Bergh makes the case that the missing piece is the Data Journey: the thing that connects a data system's expectations to its reality.
What you'll learn 7 points
- The Data Journey Manifesto opens on a single test: at any time, in a data analytic system, know what should be, what is, and the exact difference between the two. Hope it works is not a strategy, and a customer finding a problem in your data analytics is not an acceptable outcome.
- A Data Journey is the expectation layer over the paths data takes from source to customer value. It observes rather than runs: it tracks components across the toolchain and down the technology stack, taking in logs, messages, status, and data test results, and it does not execute anything itself.
- Two manifesto principles cut against the usual assumptions. Your data providers will send you bad data, so protect against it rather than trusting it. And perfect data quality at ingestion is not a cure-all, because many things still go wrong after the data is in.
- Data preparation itself has become a source of chaos: 5 percent of dbt's customer base has more than 5,000 models or tables, and the same data is used many times across many of them.
- Data Journeys follow Conway's Law, so the organization's design shows up in the technical design. The relationships between journeys, whether hub and spoke, data mesh, producer and consumer, or streaming and batch, usually live as tribal knowledge rather than as anything active or actionable.
- Tests have a dual nature, and the majority perform double duty: the same test that monitors production catches a regression in development. That is what makes a Data Journey usable for judging the impact of a change before it ships, instead of relying on manual or static analysis.
- Paraphrasing Anna Karenina, the manifesto's line is that all happy, error-free Data Journeys are alike, and each unhappy Data Journey is broken in its own unique way.
Slides
Transcript
Show chapters and dialogue 8,442 words
00:00:00
Hello everyone. My name is, uh, Chris Bergh. I am c e o and head chef of DataKitchen, and I'll be your host for, uh, this webinar for the next 45 minutes to an hour. Um, and so just a a bit of housekeeping. So, uh, we will share the slides and the recording of this. Uh, after, after we finish, uh, please feel free to drop some questions in the question box. I'll try to, uh, if I get time, I'll try to see those during the webinar. If not, I'll, I'll get them afterwards. Your questions are welcome and, uh, thank you for taking the time to, uh, listen in today.
So I'm gonna turn off my webcam is who wants to see me, and you can focus on the slides. And so, uh, what's this webinar about? Well, it's about the idea of a data journey and why it's important and sort of why you should care. Um, and, um, as I said, I, I'm a c e O of a company, uh, called DataKitchen. Uh, we help people do something called DataOps, which is helping them deliver insight with lower errors and, and higher rates of change and, and higher productivity. Um, and so we've been in business about, uh, 10 years, been profitable, uh, and I have lots of experience both in software and, and data and analytics systems. I'm a, I'm a technical guy. And so what, let's, we're really gonna talk about three things. So this idea of, um, the missing piece. And so, uh, we had DataKitchen wrote something called the Data Journey Manifesto, and we're gonna walk through that and then we're gonna talk some principles behind it with some pictures. And then I'm gonna do, um, a short demo of our, of our software that instantiates those, those principles and the manifesto in some software.
To give you a concrete example. So, um, I guess why another manifesto in the world, and, you know, maybe that's all, um, junk. And, and, and so one of the interesting things is about six or seven years ago, we wrote the DataOps manifesto. And, and what's interesting is, is kind of showing you the term DataOps is actually gone up. And, uh, this manifesto, when we wrote it several years ago, we kind of like, we're talking about DataOps, but no one quite got it. Um, this was, uh, for a conference. So on the flight Black Back, I sort of wrote this and sent it.
Some people got some comments and we put up this webpage and now, I dunno, 10, 20,000 people have actually signed this. Um, and it's interesting that like, um, even people just yesterday, um, uh, Microsoft had an announcement on data fabric and okay, some of the terms from, uh, the DataOps manifesto end up showing up in people's blogs.
And so these things can have, uh, influence. And so I guess from my, uh, perspective, the idea of a data journey is, is it's fundamental. Um, and so we need some way to think about it. And so, so, uh, from my standpoint, sort of reducing errors and defects, it's a very lean idea in the production of data insight. Um, charts and models and data sets and integrated data sets and visualizations is the key to success. So errors and defects are a killer,
00:05:00
and that's a very lean manufacturing idea. And, and from my, my experience and our customers experienced sort of being shamed and blamed with problems that you didn't start right, problems with the data, um, being trapped by having sort of existing data processes that you don't understand and then sometimes fail. And sort of the, the morning dread sort of sitting. And, and I experienced this, um, in 2006 and seven, sitting in my car every morning not wanting to look at my Blackberry, waiting for something to go wrong with, you know, the, the th the thousands of people who were using our data. And, and I think we've talked a lot here about the stress and, and wasted productivity, um, from sort of chasing errors, um, and, and not having things perfect and, and also having your customers find problems.
And so this manifesto is, is really about a method to watch over our data's complicated path, kind of to avoid problems and errors and customer frustration and also, you know, increase your productivity and increase your happiness. So I'm gonna have, uh, just a few slides. There's 21 points here, and I'm, I'm gonna go through each, and I, I, I think the, the first big idea is, is the difference between what should be and what is.
And so when we build a data and analytic system, right, you've got a database and tools to transform data and languages and visualization tools, um, and you can kind of, and the difference between what is, is it running, um, is it late? Is the data right? And what should be is often missed in an organization.
And so you should know what should be, um, and then you should not hope that it's perfect. Um, hope is not really a strategy. Hoping things work is kind of a recipe for failure. And, and honestly, it's surprisingly, a large majority of people run their data and analytics systems on hope or, or it's converse. They, they sort of wait for your customers to find problems and, and that's just not okay.
It's not okay to deliver something to your customer and have 'em call up and say, it's look weird. And, and, and you go, oh, yeah, it's, and I'll, I'll fix it quick or I'll fix it in a month. Um, and, uh, I'm gonna talk a a bit about this, but we've got really complicated data architectures.
You look at any data architecture and it just dozens of little boxes everywhere with servers and software and tools and lengths. And, you know, maybe this is pejorative, but it's kind of built on almost performative complexity. How many boxes you have are almost your cooler. Um, and so when you have a lot of little boxes that are all working together, you need a mission control.
I need to see if things are working through those boxes. Um, and as all of us who've done data and analytics, don't trust your data providers, um, they're, they're gonna, you know, whether they mean mean, whether they wanna do it or not, whether they know they're doing it or not, they're gonna give you bad data and kind of get used to it and protect against it. Mm-hmm. Um, and don't assume what work last week will work today.
Um, you know, you're, if everything ran great last week, well the data could have changed. Some code could have changed, some server could have had a problem. Um, and then when you find problems, hopefully before your customers do, find the exact source in those little boxes, find the instance of that little box and what workbook, what model, what, um, job is running it.
And I think a lot of people focus on sort of perfect data quality is the raw data that we're getting. And that's really important, but it's not a cure-all. Uh, even with perfect data quality, uh, you still could have problems, you could still could have errors, um, because the data's being joined and put together and visualized, et cetera.
And a lot of us try to test, which is check to see if things are, are right in a manual way. I'll, I'll eye it up and see if it's right. Um, and sort of, uh, the idea here is to avoid manual testing of, of anything data server. C P U is the report showing the right numbers?
Just avoid it like the plague because, um, it's, it's causes, it, it's sort of akin to hope. Um, if, if it sort of works in development, it, it's guaranteed to break in production. And we've talked a lot about this in our books on DataOps, but really you're running a factory. Um, and the lessons of Toyota and Lean and Deming, that every tool along the way, the ingestion tool, the transformation tool, the database, the model, the visualization, data prep tools, data governance tools are kind of work stations on that assembly line.
And so you don't wanna buy a car that is produced on a crappy assembly line that has a lot of errors, nor do you want insight produced on a, on a crappy assembly line. And, you know, we've written a lot about the, a a more general concept and, and wrote a, a manifesto years ago that I talked about. So, um,
excuse me. So I, I think we need a new idea and, and that's called a data journey.
00:10:00
And I think the idea here is that you need a layer kind of on top of what already is running. You need an expectation layer, and that expectation says, um, it tells you, um, of all the myriad paths that data takes from where it gets into your organization to where it goes out to your customer. Um, is it right, can I trust it? Is it on time? Does it make sense? Um, and it's a, a layer that really observes things and doesn't run things run anything.
Like we've got lots and lots of orchestration tools and, and there's just a, you know, another open source orchestration tool starts every week, right? And there's like 56 orchestration tools, there's 50 vis tools, uh, every, uh, data stacks there. And so really you've gotta observe kind of across those tools and down the stack that they run. And so that across and down problem, uh, for every instance is hard. And that's really, you're not running those things. You're observing them. And, and when things go wrong in them, you want to know before your customers.
So getting realtime status, um, and then kind of creates, what's missing here is the production context. And this is, um, often sits with sort of a person on your operations team and they know, maybe they've got a spreadsheet or a Word document that talks about, okay, this is run and after this ha hasn't run. And um, they have, okay, we've got 50, a hundred, a thousand jobs in their organization and maybe they got an Excel spreadsheet cuz they use it, but there's no sort of context of what's running. Um, and that context also requires that there's a lot of components, right?
Your servers and databases and jobs, um, how do they all fit together? And here's sort of a literary reference of paraphrasing Anna Corrina. Now sort of all happy error-free data journeys are alike and each unhappy data journey is broken in its unique way. Um, so if you've ever read any Russian novels, you'll get that quote. Um, and then here's another one that's also a very nerdy quote from the eighties.
So Ronald Reagan said, of the Soviet Union Trust, but verify, and I think that's right. So trust comes from monitoring and checking the data in every tool on your data journey, um, and test and look for anomalies at every step.
And then, uh, finally, this idea of your data journey and having a place to put it actually makes it really useful, right? One is that you can share it, right? Everyone wants to know the production schedule and, um, between the engineers who built it, the people who use it, um, and all the managers who get yelled at like that, that schedule, the what's running, what will be running is really important.
And then keeping track of what happened in the past, sort of learning from errors, trying to look at over the last two weeks, have we met our SLAs? Um, have we had any serious data errors? Where was the problem? What can we do to fix it? This idea of root cause analysis and being able to go in and look and learn. I think these are all great ideas. And so what, why would you wanna do it? Well, I think there's two parts.
One is that of course it reduces errors and by reducing errors, it drives team productivity. That's the bottom bullet. The second is that it also can actually, by having a data journey that's sort of tested, you can use it in development. Um, and so, um, you know, and, and I think that's great, right?
And I think there's also the idea of a data journey to help regression test or do impact analysis is a good way to think about how data journeys exist from a similar concept called data lineage, which is a way, sort of a static analysis of your data system, but doesn't tell you if things are are late or if the the data's wrong. Um, and so this idea of automatic, like any software developer knows, you could have a bunch of static analysis of your code and it's good, it'll tell you, okay, the, the function has, uh, ex is expecting a string and you sent an inte integer that's really great. Um, but nobody ships software just with static analysis. Um, just like no one should ship, um, uh, data and analytics just looking at lineage. You actually need to, to run the system and test it to see if it's right.
And that's where this idea of regression or impact analysis testing comes in. So let, let's kind of go to, um, that, that's the manifesto and, and, uh, tell me what you think. We actually have it up on the internet now, so you could go in and look at it and see all these points.
And if you're interested, you can sign the, the manifesto. Um, and so, uh, thank you for that and let me see, open up the thing and see if there's any questions so far. Great, there's no questions. So let's go back to, um, our slideshow here and let's sort of give some concrete examples of sort of why you need this now. Now this is from, um, and Horowitz, they're a, you know, of a top-notch VC firm, and they're also top-notch content creators.
00:15:00
And, uh, here, here's all the little boxes, right? And so like, there's all these little tools between, uh, you know, reverse E T L and your data lake and your metrics layer and your dashboards and your embedded analytics and your sources and e TL tools, visualization tools. It's, it's pretty crazy, right? Lots of boxes and arrows.
And so somehow you're supposed to tell if things are working on that. And so I, I just think that's a crazy, um, it, it, it's, it's a crazy piece right there and, and trying to see across all of it. Um, and I think that's one part. And then then second is like every cloud now has the full stack of data tools, right? They have databases, they have e TL tools, they have VIS tools, they have data science tools, they have catalogs, um, they have different types of, of, of database engines.
And so just pick your pattern here. Again, the law of little boxes everywhere. Um, and so, and, and there's different patterns and you know, this is normal, right? It's a very, we're doing very complicated things. Um, there's never gonna be one tool that does everything. Um, little boxes mean there's lots of innovation. And so people love their tools, right? They love their little boxes. And, and a lot of times your individual contributors are on their team.
That little box is their world, right? That's, that's the thing, their, their little box of working on. Um, you know, the, their little box of working on the E TL layer is gonna, that's what they do. And so these little boxes actually, um, are important to people. And so I don't mean to diminish the, uh, the innovation and importance of those tools. I just, you know, happen to think that the whole system is more important than the individual part. And, and we take these parts and we actually arrange them in really complicated design patterns. Now, a lot of bigger companies have what's called a hub and spoke team.
They have a central team that ingests data, maybe governs it, and then spoke teams who take data and do, do with it, transform and integrate it. Um, you've got sort of, um, data producers, data consumers, which is a very common pattern. Uh, you have the new modern data mesh, right? Where there's sort of individual teams opening domains and the interconnection between domains and then streaming, um, and in addition to batch, right?
And all these things are great patent, right? And, and you should try them. And they're sort of, and the reason you have them is because either your data has a high velocity and it has business value, or you're trying to organize these large teams to make, um, to make them more efficient and more agile. But again, it's just more boxes and more links between boxes. Um, and then interesting D V T published this case of, of their models. Um, and so, you know, it's, it's sort of like places where data goes in a database and, and 5% of the D B T models have, over 5% of the customers have over 5,000 models. And this is, you know, whether, whether this is valuable or not, or they should refactor it and get it down to a thousand models.
This has been common, not just with d this isn't DB t's fault, right? The sort of view of data happened. It still happens with Cognos, right? You have a, uh, 400 Cognos of reports e each of which have 20 embedded lines of sequel, and then they actually do data prep. Or, you know, there's companies who have use Alteryx and have lots and lots of alters jobs running to do, kind of take the data from its raw format or semi formatted and put it into a position to get value. And so again, that's that follow the path through 5,000 things that that, and find out where the problem is, let alone how those 5,000 things end up in models or visualization. Um, and so, so what's a picture?
So data journeys are sort of across the little boxes, right? Um, where you take raw data, it's batch or streamed, you place it in the database and you ETL it or e l t it, you transform or clean it and you use it right? In visualization or data science or sense to somewhere else. And, and the characteristics are, there's multiple tools, multiple data sets, multiple paths through those data sets. Um, multiple jobs, um, multiple architectures, multiple customers, and multiple people.
So that's a lot of word multiples everywhere in addition to boxes, right? And, and then the other little thing that's problem about this is those boxes are not just, it's, the boxes are deep. So the problem could be in, um, all the way to the bottom could be in, you got something in the lower right here off your ftp, wrong, or that could be perfect, and maybe your server is, is spinning outta control or maybe your tool isn't running, um, or maybe the code that's existing in that tool, a yaml, a sequel, but workbook has a defect, or maybe the thing that's running it, your orchestrator, um, uh, fell over. Um, or, uh, and this is great if you've got some tests, maybe your tests are returning errors. And so all these, you have this sort of a cross and down problem in finding where things are.
And that, that's the challenge here because, uh, things will always go wrong.
00:20:00
And so where to find the problem and fix it before your customer, uh, is important. And so that's sort of the idea here after 19 minutes. And I'm gonna sort of talk about, um, some of the principles here quickly. And so, um, feel free to jump in questions. Um, and so what are the principles behind a journey? Well, first is that you have lots of them in your organization already, right?
Maybe you call them jobs or pipelines. Um, maybe you call them workflow and, and you know, they could be sort of embedded in different teams, and they're batch and streaming and manual, and they're everywhere. And if you look at larger organizations, they could have hundreds, even a team of, you know, three to five people, they could have 20 to 50 of these things running. Um, and, and the whole point here is you just want to know that they're running and know that they're right and just think about other things. And, um, you know, we all have lots of tools, as I've talked about before.
And so people have different tool chains, lots of boxes. There's lots of tools in every category, and I don't think, uh, I I need to express this anymore. And then the second is, there's lots of pieces that things run on. It's the across and down problem. And then data journeys. These, these pipelines actually fit in in an organizational context, right?
So you may have, um, a central hub team who ingest data, they've got their own journeys, and then you've got each team in a, in a sort of line of business running separately. This could be mapped into a data mesh context, or each are domains in the inter, inter domain, uh, relationships. But like, there's lots of hub and spoke dependencies here.
And so data journeys themselves have a development process, right? You, you change them. Um, you wanna get new code or new data or new changes to configuration into production. And so the, they have, and, and oftentimes that development process is different depending upon your organization. Maybe your hub team is a very formal multi-month s sl d c, whereas your spoke teams are kind of just pushing a button and eyeing it up. Um, and so it th this is also, uh, interesting in, in, in, in a source of error.
And so, um, you know, because we've got these organization designs, hub and spoke data match, your data journeys end up following Conway's Law, and that they all have to be, you know, people own the data journeys based on how the organizational structure not based on the, uh, you know, the, the, perhaps the best technical way to do it.
Um, and this happens in software, this happens in manufacturing. Conway's law is a very common thing. And, and so we talked about this, right? They need correlation. So if something goes wrong, right? You want to know it, was it the raw data? Was it something happened when I put the raw data together with other data integr integrated it? Uh, was it because the actual predictive model went outta whack?
Or the visualization? Is it making sense or it's just late and the server's getting pinned? Um, and so all these things, and then what will happen, uh, what's related to this that will break it. And so this idea of correlation across and down and between journeys is actually really fundamental and actually makes it complicated, um, because of this, uh, complexity and coupling in, in these systems.
And so, um, oftentimes, as I said, that's not really anywhere in an organization. You know, like for instance, a very common case, if you look at the, the, how does one data journey relate to another? Well, one is causal, it completes, and then something, an orchestrator or a signal causes something else to happen, right?
That's in the, uh, upper left ear. Sometimes it's temporal, like your ET TL process has gotta finished by 5:00 AM and then at 7:00 AM your, your dashboard build works. And they're, they're uncoupled, but they're obviously related to each author. Um, and then there's manual cases. You know, your, your, uh, Airflow schedule are finishes at 1:00 AM and then your, uh, India team's gotta press a button to do something at 3:00 AM. Um, and then there's ones that are more event driven.
So something new data drops added in, uh, update the dashboards and, and then these relationships are complicated between journeys. They sort of fan in, uh, where I've got a bunch of new data sets and I'm building the warehouse, or I built the warehouse, uh, fan out, and now I've got a bunch of dashboards and data journeys that are related. Um, and so this tribal knowledge, unfortunately really isn't, you know, it's not anywhere in an organization, and that's why it's crazy hard when things go wrong. And you're having to, uh, take all your smart people and, and putting 'em together in a room, and that they're wasting their time trying to find out what the problem is, uh, rather than having a journey just tell you. And so I think that the, you know, the cornerstone of this is that you need an expectation layer on what has happening with your, uh,
00:25:00
with the production of analytics and expectations. Could be on logs, could be on errors, could be on metrics. They could be on data tests. And so you need to judge what is versus what should be. And so, um,
and so let me, uh, if I'm ask a question, let me, let me finish this, uh, expectations and I'll answer the question here. And so, um, and, and so journeys react to events and notify. And so what's happening in your system, the sort of living, breathing data, moving from place to place, something happening with data results being pushed to customers.
This is good sort of living and breathing and going on. And so what's the variance between what should be and what's not? Like obviously schedules and durations and dependencies, this should have run before the other thing. Um, and then trust, like is it trust? And so that variation between what you think should happen and what is, um, is the source of a great way to notify. And so, um, you know, because you've got this database of what happens, you can actually do some historic analysis and look at it and say, well, what did happen? Um, and what happened last week? What happened last month?
And where are we? And sort of, uh, looking at it from the boss's context, um, how many tests did you have? How often were you late? Um, and then this idea of context, right? Between someone who does the work, right? Uh, maybe I'm a data developer, a data scientist or engineer, um, maybe someone who runs the work, a production engineer, maybe someone who manages people, and then maybe someone who uses it, like they're all interested is, is this working or not? Um, and so this shared context is actually really important for people, and oftentimes it's missed. And some people try to do it, right?
They'll have a run table and they'll manually create them things. And that, that's, that's very helpful, but it's just, it's, um, um, it, it's saves a lot of time if you can give the pulse of what's going on with everyone else in the organization. And so I guess the question is, um, we saw a lot on analytical data, but how can we approach transactional systems to apply some, uh, data journey on it? And I think that's interesting, right?
Because I'm primarily talking about analytics systems. And what that means is there's a transactional system like Salesforce, c r m or your company website or an e r P system where people are trying to put data in, in a very highly normalized form and, and data and analytics. We take data from that and try to answer insight questions.
And so how do you actually sort of push back on your transaction systems to find out if it's right? And I think that's actually a really good point, is like, um, and something I believed is that you just don't trust your data providers. And so the journey starts with them and how do you actually improve their data, uh, both the quality and the errors that they have.
And I think that comes from having and the ability to find out problems in data very quickly. And so, um, you know, uh, and the sort of software systems that build sort of three tier architectures where you've got a front end and a back end and, and they test their transactions and they have, they don't, they don't call it a data journey that they just call it a, um, a regression or an end-to-end test where someone puts something in the ui, it goes through the UI layer to the database, and they see that the data's all right. And so they have kind of a concept of, of the, um, front end backend end to end test that that is sort of equivalent to a journey.
But really, I think journeys, um, one of the benefits of, of making sure that your data is right, as soon as you get it, um, uh, a data, a data journey is making the first step of your data journeys, right? I think that's a very helpful thing to do. And so, um, this is an interesting diagram showing like, there's just lots of problems and potential problems in data right now. Uh, if you look at all these cases where you're watching a file and putting it in a warehouse and putting in a dashboard and a model, and where is the problem, right? Was there a file count mismatch?
Was there an external refresh that failed? Um, was there some legitimate or illegitimate shift in the model prediction? Um, is there some business rule that doesn't make sense? Like, do you suddenly have, um, twice as many sales as you had last, last week? Um, and so is the EMA different? Is the file format different?
There's just a lot of places that you could find these problems, right? And so the point of a journey is to say, okay, which one of these red boxes is it? Um, and where can I find it quickly and hopefully fix it before my customers do? Um, and then the last idea, I think, is that these journeys themselves, um, have have a use in development in being able to do, um, development regression testing in addition to production. And so, um, uh, so that, that's sort of the kind of summary here after a half an hour.
00:30:00
And so the idea is that, you know, we are missing a piece. That piece is the concept of a data journey. It's something that is your expectation layer on top of all your systems. It tracks as an instance. Each one of these has that data takes through your organization, um, and allows you to store that data, get notified of that data, set expectations on that data. And so, um, you know, being a software company, we've instantiated these ideas, um, into a piece of software.
And so what I'd like to do is, is spend about 10, 15 minutes showing you that software and, and discussing kind of how we did it and why we did it. And so if I kind of go into that and show our software, so let, let's go in and I'm just gonna show this data journey. And so here, here's one.
And um, what's interesting is there's sort of a bunch of tools. There's a Azure Data factory that's putting data in. It happens to use Databricks as the database that run it. There's a Databricks notebook that runs, there's a Python file that runs, and then there's a Tableau dashboard. And so your customers sit at the right hand side here. And, and now remember, a data journey doesn't run anything, right? Um, it, it's a, it's sort of a collection of expectations. And the first expectation as well, a Azure data factory's gotta finish before Databricks and the models run and then they have to finish before Tableau runs. That's one level of expectation.
And the next is that they ran that they started and stopped, and there wasn't errors in any logs. Uh, and then the, the third level is like, what's going on inside of these things? Um, is the data correct? Um, has it been loaded, right? Is the, is the Tableau dashboard showing those things? And so the important thing is to collect all that information and start to understand the timing of them, um, start to understand all the events that happen on it.
And so in this, we're collecting a whole bunch of event types, like for instance, just messages and logs. Those are actually really interesting things like, hey, it ran, or Hey, there was an error. Um, there's run status. I started and I stopped. Um, outcomes of metrics, like, here you go. I've gone in and if I want to look at different, like a metric event, I can actually go in and see the number of files written, the amount of data written, I can look at, um, various messages in here, just, um, what's running, what's not. And then I can look at things, what are called, uh, test results, which are ways to actually check test.
And here we've got some test outcomes. And so the, what we do is bring all that together and then allow you to say, well, what happened? Here's the, uh, I've got my azu data factory job running here, and what job is that? And links us into whatever tool you happen to be using so you can dig more, because the tool's doing all the work, right? You're just observing it.
So being able to go in and get at it is, is important. And so, um, you know, this sort of data journey collects all this information and collects all these events, all these tests that are running against it. And we're gonna talk a lot, a little bit about tests and puts them all together, um, into a way to look at the world.
And then there's ways for you to actually go in and judge what happened when the world doesn't meet that. So for instance, when one of the components in a journey fails, well, I get an email when Journey has a late start or a late end. Well, we, uh, call for instance a specific Slack channel.
And then the third is like when the server capacity is greater than save 75%, uh, you get an email and that may not be an error, that may be just a notification that, that you go on. And so being able to look at, um, schedules and timing and order and uh, metrics across the tools and down the stack and be able to make, uh,
uh, alerts off of it, I think is important. And so all these things kind of go into a data journey, but seldom does one company have one data journey. You actually have a bunch. And so let's look at this case here, and I'll show you a slide to kind of talk about what I'm about here.
So in this case, we've got sort of, uh, in the example, there's a nightly exports. They sort of have a data factory pipeline, and then some Azure Logic apps that do the export. Um, we've got an Airflow DAG that does the daily data load. And then after it you've got sort of what I just showed you, the dashboard and, and, and the model and any other, there's some databases and, you know, uh, using the typical bucket stores and what happens when you don't do this, right?
You can get data errors or tool failures. You can get downtime on one of these that they're not running, and that's the red box I showed you. It's difficult to judge the impact when something goes wrong and your customers find problems and you have much longer resolution time and you end up kind of
00:35:00
doing the same thing over and over and over again. So going back to the product, well, we found the nightly exports, well, it had a problem. So let's kind of go into that and say, well, what was the problem here? Hmm. Well, Azure Data Factory had an issue. Okay, well let's go into there.
And I can actually see the instance of Azure Data Factory and oh, the third step in that Azure Data Factory, that data transform had an error. And that shows up in my list of events, um, that there's a specific error. And that actually came from one test failing. And so that means that everything ran, everything's great, but the data in that was not right. And so that's again, one of the reasons why you want to pull this together is that things are running, uh, logs look good, there's the right amount of disc space, but it ends up not being right. Um, and that goes, uh, to checking and making sure things happen.
And so that data export failed and now you can know about it and hope and get events on it. And then let's look at the, the next case, the daily data load. Here's, here's another tool, an Airflow job. And so what this does is actually run an Airflow job and builds five tables and sort of think of it as a dimension dot mar order customer some, uh, you know, some, uh, healthcare provider information in a fact table.
So it builds this sort of simple, uh, four dimension fact table, right? And, and what happens is Airflow's loading this, doing some transformation, um, and every time it runs, it makes these tables. And so in its run, here's the last run that that happened. This is what Airflow looked like. It did this sort of step by step process, each step that happened, and there's a bunch of events that happened against it.
So each one of these runs statuses were, were good, right? We found that it ran, that it's good, and if something goes wrong, well, maybe we can go off and and dig into the Airflow job to, to find out what happens. And, and here, um, we're actually running each one. But that, you know, the problem with this is that, you know, it runs, but you don't know if the data's right.
And so you've actually gotta check the data or test the data. And that's what this box at the end is called DataKitchen test. So let's go back and, and look at this and actually say, okay, for this instance, the, uh, of the daily data load,
I've got these five tables. How do I know these five tables are right? That the data's right? Well, we've got a bunch of tests. And so here we've actually got 791 tests that run. And so all of those passed and, and there's a, a number here that have warnings, meaning something's wrong, and there's a bunch of specific tests.
And we've got a new tool that we're releasing next month called TestGen that actually generates tests, uh, for you and, uh, be able to go off and write those, those tests. And I think that's a really important part of helping people is, you know, you can profile, some tests are really need to be developed cuz they're based on your, you know, your customers, uh, et cetera.
But some tests are just based on the syntax and the data. And so for us, we've got this idea of think of testing in sort of four bucket or five buckets, right? One is sort of profiling your data to help you understand it, right? That's one aspect. Um, and then repro profiling it as time goes on.
And the second is like, is the data different than the profile? So you have a column, it's got three values. When you first profile it, you get a fourth value, could be wrong, could be right, uh, you should know about it. And then there's sort of more fill in thelan test. Maybe as a data engineer, you don't know all the business logic. Um, and can you have some simple ways to give that logic to, uh, data Stewart or someone else to sort of fill in the blanks?
And these are more business rules, like what should the increase in our sales be every, every time we update, um, and this is our new product called DataOps Test, Jen. And then it's not enough to actually test just the data. You have to test the things acting upon the data. And so you have to test Python code or test the rest a p i.
And let me just go into that a bit here and, and talk about, um, talk about this. And so we've got the dashboard and model production, right? It's, it's got a Tableau dashboard and a Python model. So you go into the Python model and say, well, did it run? Yeah, it ran and it had some tests that ran against it.
And so I can look at it and see the events and I can say some tests came. And if I go into those tests, we actually have our, our DataOps automation tool that's doing the testing here. And so here's the machine learning model. And so, you know, models are very funky, right? And then they have, they're very specific.
And here we're actually running a model, uh, in a container and we're just testing the sort of data going in it, but we're also testing the, uh, root mean square error of the model. And we're doing that because it has to be custom. The, the data data scientist writes their model, uh, they have specific parameters and they're actually even have their own favorite
00:40:00
language to write in. And so actually interacting with the model here, the data scientist likes to do things in Python. So we're using, uh, a Python script to do it. It doesn't have to be maybe like sql, maybe you have your favorite tool, but the test here is, is saying that the root means square has to be less than 60, and 60 is a number that comes from, um, uh, the data scientist.
And this r m s variable actually comes from that Python code that you saw. And so the idea is that you can't, um, in the system, you have to prove that everything works. So proving that, um, your dashboard is right or your model is right, don't hope, right? Hope is not a strategy trust, but verify, make sure, and the verification is at multiple levels. Did it run, uh, did a metric trap or is the data or the thing acting upon the data? Correct?
And so that's why we bring it all, all together. And so lastly, um, once you have all these data journeys and you've got the overview, you, you get this dashboard that you can actually share with your business users, right? But you can also kind of keep track of what happened and learn from it.
So we have a, a dashboarding functionality here that actually allows you to look at it. And, and so what I've learned, um, managing data and analytic teams is, is trying to, you, you've got, as a leader, you've got sort of two views. H how good is your team and how do you show that up?
And that's looking at is everything on time, right? And here's a metric saying you had 141 data journeys, this, this last week in 81 or on time and 60 were late. So you've got a database of that. Likewise, being able to go in and look at data quality, being able to look at the number of tests. And so, uh, in a month or two, we're gonna replace this with a more generic dashboarding feature because being able to sort of look across all the dimensions and, and being able to go, but uh, being able to go in is also a feature.
But these sort of standard dashboards are a way to help you, uh, kind of understand the data, show off how good you are, and be able to, um, uh, help your team, uh, uh, think about how the work that they do, uh, is good. And so lastly, um, just to say we've got three products now. We have the observability product, which I demoed the TestGen, which runs against databases and generates tests for you is that we're gonna release in a month. And then our DataOps automation.
And so conceptually they look like this together. We've got observability, which is really just a simple rest API that you can program against. Um, it's got an integration agent that actually is really quick to set up. So for instance, last, uh, a few weeks ago we had a customer start up. They put the integration in within, within 15 minutes, they had sort of 60 components, uh, in, in their system, all pumping data in, um, looking at what happened. And then building the data journeys was, was a pretty simple, uh, a process to be able to connect all the pieces together.
And then sort of writing tests. Well, we've got two tools, one for APIs and rest, and one that generate tests. And of course, if you've got your own test tool or written your own tests or want to use one of the other data observability tools that does testing, we're perfectly happy with that. And so, um, kind of like putting it together more in a concept picture here to go to a slideshow just to look at our use cases. So for instance, you've got data coming on the left, maybe using Airflow to load, maybe you're transforming it in D B T, maybe you're doing some predictions in data bricks and reporting in Power BI, right? And here's your data and your infrastructure.
The first thing is we've got a test tool that actually tests your data in your database and that sends it to observability. We've also got the automation, which will test your tools and send it to observability. And then just sort of monitoring things like logs and errors and runtime and schedule also goes. So pulling all this information together and building the data journey layer is, is what our, our software's meant to do. And these, these other boxes are optional products that, um, uh, you know, allow you to sort of connect and test. But if you've got that already, that's just fine with us. And so, like, uh, lastly sort of why should you care?
I think the, the, the top bar here really matters, um, is that if you look at an organization, and we've done surveys with Eckerson and I think Gartner's that, uh, some people are spending a lot of time on fixing broken stuff that doesn't have to be broken in the first place. And so when you're fixing production errors, you're not doing good work. Um, and that is a big eater of teams productivity. And then second, if you can't change things quickly with low risk, then you're not responding to your customers and you're doing work you don't have to do.
And so the idea of observability is sort of drive down production errors, and that's a big part of the sort of productivity benefit that, that Gartner and we've seen. Um, and so lastly, the conclusion. Um, if you're interested, sign our manifesto. There's the, the link at the bottom. Um, you'll be on our email list, but that's, we have, we've had over almost 15, 20,000 people sign the DataOps manifesto surprisingly. Um, and then lastly, there's just a bunch of, uh, links here. We've written two books on DataOps.
00:45:00
Uh, we've written a, a manifesto, um, we've written a bunch of things on observability, um, and, uh, uh, a data journey manifesto as well. So thank you for the time and I see some questions and let me answer those questions here. So the another question is, are you able to access the PowerPoint? Sure, I'm gonna share, share the PowerPoint, um, a link to these links and, uh, a recording of the session here, probably, uh, tomorrow or the next day.
And you'll get an email and then we'll also post it on our website.
And another question, do we provide a free trial? Absolutely, yeah. We provide a free trial. We have a a 90 day free trial on our website. We also, um, try to use our observative tools in production to prove that they work. Um, so I, I hope this was good. We, we talked a lot about, uh, different things about observability and the manifesto and, uh, trying to help you be, be better and, and have your teams have less errors.
And that's really the key, I think, to teams being successful. So thank you again. I'll email out these, uh, these slides, uh, and hope you have a great rest of your day.
Oh wait, okay. Awesome. Thank you.
And so is there a link that the services, the tool should connect to? There's one more question. Yes, yes, there is. We can, we can provide that and we're actually gonna have it and our next release that link's gonna be, uh, in the ui. And one of the nice things about the observability tool is, is, is we have a 15 day sla.
So if there's a commercial tool that we don't yet support, we'll incorporate it in within 15 business days. Um, again, that's it. Thanks everyone.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Questions from this session
What is a Data Journey?
A Data Journey represents the expectations on the paths data takes from source to the insight delivered to a customer. It tracks every level of the stack, from data to servers to software to code, covering components across the toolchain and down the technology stack, and it supplies real-time status and alerts. Its purpose is to tell you whether everything ran on time and without errors, and to name the parts that did not.
How is a Data Journey different from a data pipeline?
A pipeline runs the work. A Data Journey observes it and runs nothing itself. It sits over the pipelines, jobs, and schedulers already in place, holds the expected schedule, durations, dependencies, and quality thresholds, and reports the variance between that expectation and what actually happened.
What is the Data Journey Manifesto?
The Data Journey Manifesto is a statement of principles published at datajourneymanifesto.org, written for teams tired of being blamed for data problems they did not cause. Its principles include knowing the exact difference between what should be and what is, making hope infrequent, treating a customer finding a problem as unacceptable, automating all testing, and treating data production as a factory in the tradition of Toyota, Lean, and Deming.
Why is good data quality at ingestion not enough?
Because most of the path comes after ingestion. Even with perfect initial data quality, the data still moves through ETL, databases, models, dashboards, and exports, and each of those can fail on its own terms. Finding the exact source of a problem across raw data, integrated data, models, reports, servers, software, and code is described as half the battle.
What kinds of data tests does a Data Journey need?
Four kinds are named. Drift and consistency tests are generated automatically from profiling baselines. Business rule tests are parameterized fill-in-the-blank checks that carry domain expertise and double as documentation. Custom SQL and Python tests handle logic too complex to parameterize. Custom API and tool tests cover things like Power BI, Databricks notebooks, and REST endpoints.
Why do Data Journeys need historical data?
Because root cause analysis needs the past. Collecting run history longitudinally is what lets a team analyze, learn, and predict rather than react, and each instance of a journey becomes evidence of whether production errors and missed SLAs are actually going down. The journey also becomes the context that gives each individual event its meaning.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.