On-Demand Webinar · 25 min
Data Observability Demo Day
DataKitchen's DATAVERSITY Demo Day session from November 2023: why complexity in teams, tools, and environments makes Data Journeys unreliable, and how observing errors across the toolchain and down the stack replaces the game of who is to blame.
What you'll learn 6 points
- Complexity across teams, tools, and environments is what makes Data Journeys unreliable — hundreds or thousands of journeys with no enterprise-wide visibility and no end-to-end quality control.
- Errors sit in three layers: data errors, data-and-analytic process errors, and basic IT-level monitoring errors. They appear both across the steps and down the stack.
- Not being able to correlate those errors quickly is what produces finger-pointing, burdensome problem-hunting, lost productivity, and customers who stop trusting the data.
- Observability here means events from every journey in one place: TestGen data quality test results, infrastructure logs and metrics, order of operations, tool status, and end-to-end SLA.
- It attaches to an existing toolchain through agents, so the pipelines themselves do not change.
- The same end-to-end view applies in development, for regression and impact testing, not only in production.
Slides
Transcript
Show chapters and dialogue 4,623 words
00:00:00
Hello everyone. My name's Chris Bergh. I'll be the host of Data Observability Demo Day.
Thank you for attending.
So we're gonna have a about a 30 minute webinar today, and we're gonna do just a few things. I'll have a few introductory slides and I'll walk through our data observability products. Um, we'll be sharing this webinar as a recording and the slides, um, uh, you'll receive an email and, uh, we'll, uh, post it up on our website, um, today or tomorrow. So, again, my name is Chris Bergh, I'm c e o at Head Chef of DataKitchen. Uh, a lot of experience in in data and analytics. Um, and thank you for attending.
So let, let's start. Why do you need to think about observability? Well, I think the biggest thing is failure, and that most people, most data and analytic projects are rife with failure. Uh, the projects fail. They have too many errors. Uh, things don't get into production. Um, and, uh, we did a survey, uh, about two years ago that 78% of data teams are sort of so stressed.
They want a therapist. And, and really it comes down to these three C's, um, complexity, chaos, and crushed. And so I think all of us know that, uh, we have a lot of tools out there. There's a lot of little boxes, uh, uh, Amazon, Google, Azure, modern Data Stack. There are very complicated architectures, uh, very complicated design patterns and lots and lots of different tools that you could use to ingest, store, transform, view, govern, visualize data. And so I don't think that's a, uh, too big of a surprise to anyone who's worked in data and analytics.
And I also think there's, uh, I don't think this is too, too much of a surprise either, either is that many things break. We have sort of staggering error rates, um, and, uh, in terms of that late, that the data's wrong, that something's broken, um, that we've put some code in production that doesn't quite work. And lastly, teams are kind of overworked and under prepared. Um, most data and analytic teams feel that they've got too, too many requests from customers. They've got too much work, and they can't answer very basic questions, uh, about what's happening with their work. Like, is something done or did it get done on time?
Is the data that I got from my sources correct? Um, are, is that dashboard actually showing the thing I want? Is it not blank? Um, you know, are are things arriving on time? Um, did things run in the right order? Um, where is the problem? Um, how do I write a data validation test?
And so the way that we think about it at DataKitchen is, is we're gonna talk about our, our three, uh, two of our products, our DataOps observability, and our DataOps TestGen product. And, um, the reason that we built them, i, i, is really, if you think about the world as what happens with data in place, maybe that's in a database or in a bucket store, um, in a lake, a lake house, and, and you know, it, it could be put in raw tables, it could be put in an integrated schema or, uh, but there's a lot of errors that could be sourced in there.
Maybe the data isn't fresh or it has volume errors or schema errors. Maybe something is wrong. Maybe your sales are down 50%, but errors also happen in use, maybe in terms of your predictive model is not correct, your dashboard isn't showing any data or your exports didn't run last night. And so there's lots of different places where things could go wrong in your data and analytics process just in looking at data in place or data in use.
And this sort of across problem is, is important to find where it is, but it's not just across. It's also down, did something run or did it not run? Did it run on time? Did something run out of order? Are people actually using it? And then lastly, um, is the that process run? Are there errors that came up from my software process?
Are some metrics down things like C P U or disk is cost happening, uh, and going, uh, too high. And,
00:05:00
and this sort of across and down problem is really hard for data and analytic teams to find because you have to look across your tools and down the stack into your data. And if you don't do that, you end up with this sort of chaos and complexity and, and being crushed because you're finger pointing. Um, you have lack of team, uh, productivity and lack of customer data trust.
And what that comes down to for us and the way that we we've built this product is you're missing something fundamental. You're missing this idea of a data journey that helps you solve this across and down problem that sort of brings it all together for you to find out. And, and that's what I'm gonna show today. And so, um, if I go back to our product and start talking about a data journey, and here what's in front of you is a data journey.
And so a data journey is something that doesn't run anything, even though it sort of looks like this looks like a, perhaps a very simplified, uh, workflow process. But this is really a set of expectations about the way the world should be. And in this process, there happens to be an Azure, this happens to be running in Azure. It could run in, in any cloud or on-prem, but, uh, there's an Azure Data Factory job that runs, and then there's a Python model and a, uh, a Delta table notebook and then a Power BI dashboard.
And this order has to happen and it's all running in a Databricks personal compute cluster. And so how do you know that this is right? Well, a lot of data and analytic teams, they just sort of put it out there, um, and then they wait for their business customer to look at the dashboard and then call themselves up, call them up and say, Hey, something's wrong.
And then they run really fast, like heroes to fix it. And I just don't think that that's acceptable. I think you should know where the problem is before your customer sees it, and if possible, correct it. And so that's what this is. This is a set of expectations about the way the world should be.
And so what if we look at these sort of classes of expectations? Well, the first thing is that each one of these boxes, um, should run. It should not have any errors. It should be able to start, it should stop. Um, it should run on time and have a duration. And so if you think about it from just a software process, that software process should run. And then the order of these things matter.
So for instance, you don't wanna
load your Power BI dashboard export before your data and analytics are are done. And here in a case before your notebook is actually put in the, the, the customer segmentation. Um, and so the way that we've broken this up, these data journeys up are into components. And so if I go into these component processing, I can see that the Azure Data Factory job ran, and then the notebooks and the Power BI dashboard.
But what we do is we have these connectors that snap into your infrastructure really quickly within minutes and start pushing out data about what's happening. And so for here, we actually saying that there's two parts to this. It's our data factory job. We pick that up, but we also capture all these different events.
And so events are of different types in our infrastructure. So for instance, some of them are just messages on what's happened. Um, think of those as log files. Some of those actually could be metrics like data written or file written. Some of them are actually run status, uh, events. And, and what's interesting about these run status events is that we actually go in and allow you to very quickly find out what's happening.
So we can go in here and, and, uh, actually launch off into, um, the Azure Data factory job and see the pipeline and be able to diagnose. Uh, it, it, it's happening. And of course, 'cause we don't run anything, we're just setting expectations setting. And then the finally, the actual interesting one is what we call test outcomes.
And that's actually to find out what's happening with the data at rest or data a use. And we're gonna talk a lot about tests in this demo. And what we mean by tests are data quality validation checks that prove that the data is right, the integrated data, right, or the data in use is right. So if I go back, look at this data journey, I've got nu most companies have a, a number of these running. So in this particular project, there happens to be, uh, four of them running. And so one of them is the nightly exports that happen.
One of them is the daily data load, and let's go look at that. And here we can see it ran and there was a problem, and it's also scheduled to run tomorrow. Um, and so let's look at that one. And so in this daily data load, there's a lot of teams that, uh, will break themselves up into sort of ingest into usage teams where maybe the ingest team is, is bringing data in, or the data team is working, and then there's data enablement. So in this case, we've got an Airflow job.
The Airflow job is, is running, and we can actually see, uh, its timeline and we can drill into, uh, the DAG that we built from, uh, automatically out of what happens with Airflow. Um,
00:10:00
but we actually have have built all these tests that run against it. And you can see here the 70 tests against the, the customer dimension. There's hundreds of tests against, um, uh, the fact table to be able to prove that it's right. And so each one of these tests is, uh, is actually a discreet way of looking at the data.
And I'm gonna talk quite a bit about that and, and talk about why it's a challenge for a lot of data and analytic teams to actually test their data, do these data quality validation tests. They, they don't know what tests to write. Um, they don't know how to, uh, when they do write them, they don't know how to tune them and tweak them and get business people involved. And so we're gonna talk, uh, a bit about that, uh, uh, in a minute.
But these data journey are sort of decorated with agents that connect to your infrastructure and tests that actually tests that things are right. And we pull all this, that data together into an overview, and then we actually pull it together into an overview across your business. And so these overview pages I think are really important. Um, and, and why is that? Well, um, in this case, in, in, in, uh, our Azure project, we can see that the nightly exports have run.
We can also see across projects. And some companies have not 10, but hundreds of or thousands of pipelines running. And now it, it becomes a very public information. Your data engineers, your data team's building it, your operations team who's running it, your customers who may be using it are all interested in this. Um, is it done and can I trust it? Question. And so we have a very simple way of, of looking at these data journeys and sort of filtering them by projects.
And so if I go, uh, back here and I look at, uh, all of our projects, I'm gonna go back into our a w s demo project. And, uh, I can see that I've got today's journeys. And each one of these journeys is, uh, we've assembled all this journey information into a database. And it may, may help right now to actually look at our, our architecture. And so, uh, one of the reasons we built this product as a SaaS product, um, and what this diagram shows is that you have your infrastructure, right?
You have your data sources on the right, you have a load process, a transform process, a prediction process, a report project. And in your tools may be these tools, they may be different tools. You have databases, you have your infrastructure. And so we have these agents that lock in and sort of pull messages out of those infrastructures very, very simply.
And they send things like the run status or the schedules or logs or metrics or various events, but we also have the ability to test and check. And so you may have tests, like for instance, is the RO count big enough? Or did the data arrive? And those checks are actually very important for, for you to see, um, and, and to do. And so to do that, we've actually built our own, uh, test engine. And let, let's look at it, uh, here in the product.
So one of our data journeys happens to be what's called a database monitor. So it runs, um, it runs recurringly. And if I go look at this data journey, and especially look at the tests in it, and so we can see hundreds of tests here, 533 tests, but only a few of them have failed. So for instance, let's go look and find out what that, what test failed here on us.
And so here I've got number of rows at or above threshold. And so, uh, that came from our test suite, this test suite scheduled to run periodically against your database, and it's a monitor. And so we hear, see some test information. So how does that test information work? Well, it gets to our observability product through the a p i, but a lot of data engineers are sort of perplexed as how to write tests and, and where to, uh, what do I do? How do I check data?
And so we built a new product called TestGen, which means it, it generates tests and it starts with scanning your database and building a whole, uh, series of profiles across it. And then we have 51 different characteristics of your data that we profile. And so from that profile, we actually find some profile anomalies.
And so for instance, here, this is, you know, you load a new data load, some new database tables. And so for, uh, what we're, we're, we're doing here is kind of finding strangeness in the data when you load it. And so for instance, we've got some cases where some dates are out of range or we've got no columns, uh, are, uh, in the data.
And so that's actually an interesting way to look at it once you profile the data. But if I go back and, and, and profile the data, I can actually start looking at our profiling results. And so there's a lot of different ways to profile data. So for instance, here, we've, we've scanned the zip codes and we've looked at the frequency of values.
We've looked at the, the record counts, missing values, numeric counts, uh, top frequent values, uh, ex and have some typing information. And we also have some functional data types where we try to figure out what it is based on, based on the data. And so we profile your data with these 51 different characteristics. Now, there's a lot of different data profiling tools out there, and I think we've,
00:15:00
we, we have an interesting one. Um, but what's more interesting is, is that question is like, well, how do I actually write a test? So I've got a new database table, I don't really know anything about it. How do I actually check that it's right? And so it's, and, and so you can look at things like it, is it fresh or has it loaded?
Um, but when you start to get into the actual quality of the data, does it make sense? You need to, you need to, uh, uh, uh, sort of a leg up. And that's, that's what we've done here. We take this and we actually generate tests from this, this profiling information. And so we'll go in and, and, uh, uh, build a series of tests that have happened.
And here I can go in and I'm gonna look at all our tests, um, uh, all past tests. And so we, we build a whole bunch of, we build 28 different tests that run against the data. Now, these tests can run when the data arrives, they can be set up to run periodically, like every hour or two hours.
They can be set up to run before your production in the middle of your production, just after your production. Um, and there's various reasons why you may want to do each case based on how your team is organized in general. It's better to find data problems as soon as they happen, um, instead of waiting. And so, uh, what, and but the challenge that we're trying to address in this tool is to be able to go in and find out if there are any data tests or, or help data engineers write data tests.
So we build a profile and then we build each one of these 28 tests. And so, for instance, I'm gonna look at the tests that failed. And so in this, I've got a RO count test that failed. And so we can see the, the history of the RO count, we can see what happened, but we can also change the configuration.
And so what's really interesting about this is, um, that, uh, tests were, were building an engine that intelligently scans the data and intelligently builds a whole bunch of data quality tests. And so, but that's hard that that gets you a lot of places. But in some cases, the world changes. You may need to re-file your database and regenerate the tests, or there may be someone on your team, like a data steward or a business customer who's got a better idea.
And they may say, no, that threshold shouldn't be 50, it should be a hundred, and you can easily go in and update it here. You can also turn the test off, or you can set the test to be a warning or a flag. Likewise, you can disposition tests. So you can go into your list of tests and decide what, what you want to do with those and say, um, you can go in and click them and say, okay, I want to be able to say that this, this is a relevant issue. I wanna say it's not relevant, but keep this check.
I'm gonna be able to deactivate it for future runs. And so, uh, you know, there are teams that want to be able to, uh, help sift and winow based on their business knowledge, all the tests that run. And so lastly, we actually have a, a, a series of tests now that profile and auto generate tests, but we also have tests that go in and, and, um, are based more on the business case that happens. And so in this case, you may have, for instance, you may be tracking sales.
And one of the great ways to look at tests is to look at the test between a current period and a previous period. And so that will, uh, in DA DataKitchen, we've called that a historic a balance test. And so we've got some fill in the blank tests that you can do. And all these tests execute, um, in your data center.
None of your data is moved or shipped off. Um, they're all executed as SQL in the database, so they're very efficient. The profiling is also very efficient. And so getting this running and setting it up can take, uh, again, just a few minutes to be able to go in, depending upon your data size to be able to go in and scan and, and build these tests. And so they show up again in, in, in our, as a data journey. And in, in this case, they're showed up as a database monitor, but they could be fit in any, uh, data journey in any, any production. And, and so lastly, I think one of the most important parts is you're getting all this information about what happens on your data journeys, um, events, test results.
We build it into this nice way to look at it and be able to quickly go across and down and find the problem. But you don't always want to do that, right? Sometimes you'd be able to want to, like, not you, you wanna be able to have notifications, take care of you. Uh, you don't want to sort of manage by inspection. Um, you want to manage by alerts. And so the, the, the challenge that we, we have is like what to alert on and where to alerts.
And so we have a very robust rules engine as part of the product that to just configure alerts. And here's some examples. So for instance, when a run failed to start, or when a run has a late start, or when a run has a late end, um, or when, for instance, a metric,
00:20:00
um, hits 80%, so the Databricks cluster goes up or something's out of sequence, and we can send alerts and email messages. And so a a lot of cases, the, the, the question in notifications is, how do you notify the right person? Um, and so here I could say, well, um, when a task is had that, uh, failed for the Tableau dashboard, I could send it to the, uh, alert or a Slack or a Jira ticket to my BI team@DataKitchen.io, and they could be able to save and, and, and, uh, uh, and add that in so we can get the right blast radius.
And so this rules engine allows you to, to connect to different, uh, tools and send a very robust set of alerts. And so lastly, uh, as part of our tool, we've got, um, dashboards. And these dashboards are ways to analyze the data and look at it. So we've got a test dashboard, we've got a timeliness dashboard to see if things are on time.
We've got an instance dashboard to be able to go and look and analyze data over time to see how your team is doing. And so, um, I guess the, the last thing I want to bring up is sort of why, why you should care. So why do observability at all? And so, um, I think the, the number one thing is that most data and analytic teams are suffering from these three C problems, complexity, uh, lots of chaotic errors.
Um, and as a result, they can't get much done. So they're crushed. And so we think the first thing that you should work on is production errors. And if you reduce the production errors, um, by building data journeys, by decorating it with tasks, with our TestGen gen engine, you will very quickly and very easily drive these errors below.
And then the second problem that you should solve is how do you actually deploy faster and get things into production? And again, our data journeys can help there. They can actually be used in development as a form of regression or impact analysis testing. And if you do both those things, you actually can drive in incredible increases in the productivity of your team.
Um, and we've seen our customers, uh, go have this 10 x improvement. We've also, uh, been had that validated with several reports from, uh, Gartner in terms of the productivity improvements. And so that's really it. In terms of our, our demo today, I just wanna, um, uh, point out some, uh, a summary of our products, our, our observability and TestGen products.
And then I just wanna point to a whole bunch of different resources that you can use to learn more about TestGen and, and, and actually learn more about DataOps. We've had, uh, 3000, over 3000 people take our DataOps certification. We, we have a new version of that that came out a few months ago. Um, and we also have, um, uh, several books that we've written, including one that just came out. And so, uh, let me go back and see if there's, uh, before I finish up, um, I promised you a short and sweet discussion today, and I wanna see if there's, there's any questions.
Okay, there's, so there's, there's one question that, that came from the audience, um, really is about, uh, how hard is it to get this into, uh, really to fix this problem? And I think that's, uh, one of the challenges that data and analytic teams has is, is that you, we end up building systems, we rush to get them production, we underinvest in observability and testing and monitoring, and then there's no time.
And so we built the product very easily to sort of have agents that snap in with a few lines of configuration that start sending information. Our TestGen product runs on-prem in your cloud. It can quickly scan your data, quickly, start building, uh, these data quality tests. And so getting going is a matter of just a few hours, uh, with our, uh, with our products. And so it's, it's very fast and very easy to get going, and then you can start, uh, um, from there quickly assembling data journeys and, and being able to, uh, solve that across and down problem in data and analytics.
And so the, there, there's one more question about, um, uh, of course there's one question about will the, will the slides and the recording be shared? Um, and of course that's, that's true. I will, uh, I'll share the slides, I'll share the recording, um, off, uh, on our website and you will get a follow up email. And, um, there's, uh, sort of one more question, um, about cost. And I think, um, uh, and this is a SaaS product, so it, we have a, a yearly cost.
It's really sort of a function of the number of events in the, and the number of users, um, uh, out for using both our observability and TestGen products. And so we are very happy to share with you, but it's, it's, it's, this isn't a $10 million project. Um, it's, it's something that small teams of five or 10 people can get going on a very
00:25:00
reasonable budget. Okay, I don't see any more, um, recordings. I want to thank you, uh, for your time today on the data observability demo day. And again, I will follow up with, uh, slides and uh, the recording. Thank you. Have a great day.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Questions from this session
What is a Data Journey?
A Data Journey is the whole path data takes from source system to delivered customer value, across every data set, tool, server and pipeline it passes through. An enterprise runs hundreds or thousands of them at once. Treating each one as a single object is what makes end-to-end quality control possible, instead of watching each tool separately and hoping the gaps do not matter.
What kinds of errors does data observability have to catch?
Three layers, and they show up both across the pipeline and down the stack. Resulting data errors: freshness, volume, row count and schema problems in raw and integrated data, then model prediction errors, empty dashboards and short export row counts. Data and analytic process errors: run start and stop failures, schedule errors, order of operations errors and usage issues. Basic IT monitoring errors: log file errors, CPU and disk metrics, and cost issues.
Why is it so hard to find the cause of a data problem?
Because the evidence sits in different places. A business rule failure shows in the warehouse, the run that caused it shows in the orchestrator, and the resource problem behind that shows in infrastructure logs. Without something correlating errors across tools and down the stack, teams fall into finger-pointing and burdensome problem-hunting, productivity goes to the search, and customers stop trusting the data.
What is DataOps Observability?
DataOps Observability is mission control for every Data Journey from data source to customer value. It gives end-to-end visibility across all tools, data and infrastructure, monitors and alerts on the complete toolchain against key metrics, and holds dashboards and historical analytics for the whole estate. The same end-to-end view is used in development for regression and impact testing through a CI/CD process, not only in production.
What is DataOps TestGen?
DataOps TestGen generates and runs data quality tests. It scans and profiles a database to identify bad data, then produces tests from what it found, and the session cites 53 unique data test types combining generated checks with fill-in-the-blank business rule tests. Its results feed DataOps Observability as one of the event streams alongside infrastructure logs and metrics.
Does adding data observability mean changing existing pipelines?
No. Observability is attached through agents that read what the existing toolchain already produces, so the pipelines themselves are not modified and any toolchain is supported. That is what gives rapid time to value: the events arriving in one place include data quality test results, infrastructure logs and metrics, order of operations, tool status and end-to-end SLA.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.