On-Demand Webinar · 53 min

Map and Monitor Your Data Journey

Can you draw a map of every path data takes from source system to production insight? Chris Bergh breaks down a Data Journey, and what it takes to monitor, track, and test the run-time lineage once you have mapped it.

Presented by Chris Bergh

What you'll learn 6 points
  • A data journey is every path data takes from source to the insight delivered to a customer, tracked across all levels of the stack: data, tools, code, and tests. It supplies real-time status and alerts, so a team knows whether everything ran on time and without errors and which specific part did not.
  • At a company in the top 20 by revenue in the United States, the head of data got a call from the CEO about a compliance report that came out empty. He pulled 26 people off their work for a full day, and the cause was one blank field passed through the pipeline. The deploy cycle was six weeks, and a thousand other data journeys sat in the same hope-it-works position.
  • The numbers on data analytic projects: Gartner says 60 percent fail altogether and 87 percent of data science projects never reach production, Eckerson says 79 percent have too many errors, and DataKitchen found 78 percent of data engineers are stressed enough to need a therapist.
  • The model for managing this risk is mission control. NASA and SpaceX run the biggest journey of all by building an interface carrying information about every aspect of the flight, using it to make decisions and communicate to interested parties, storing it for after-the-fact analysis, and alerting automatically.
  • IT hardware monitoring and application performance monitoring are lagging indicators of data problems. APM tools do not check the data, the integrated data, or the reports and models built from it, and they supply no context for the pipelines, jobs, and tools acting on the data.
  • Production data tests fall into five categories: raw data profile qualification and consistency tests, statistical process control, location balance, historical or population balance, and time series anomaly tests. Most tests do double duty as development tests and as production monitoring.

Slides

51 slides

Transcript

Show chapters and dialogue 8,703 words

00:00:00

Hello everyone. My name is Chris Bergh. From DataKitchen and we'll hang on for about 30 seconds give people a time to arrive and then start our webinar.

A few more people are coming in. So if you just hang on for 30 seconds, we'll start.

All right. So let's begin our webinar for today before we start just some practical things. We will distribute the recording of This webinar and the slides for this rep webinar at in an email and on our website after the webinar ends. So look for an email the next day or two. The second is I will be your host and and MC today and walking through some slides and if you have questions, there's a box that says questions at the end and be sure to put those in and I'll try to get a leave time to answer questions at the end of this webinar.

And it should take about about 45 minutes. And so why don't I start here with my slides introduce myself? So, my name is Chris Bergh. I am a CEO and head chef of DataKitchen. And so I am. A sort of recovering software engineer spent 15 years building software at companies like MIT and NASA and Microsoft and startups and then about 2005. I got the data bug and have been working on.

Kind of this problem the sort of how you actually build data and analytics systems that you can change really rapidly that you can run with really low quality and where you can have highly productive teams and very happy customers. And so our company's been in business about 10 years. And so what's our agenda today? Excuse me, so we're going to talk about this idea of mapping and monitoring your data Journey. So first we're going to kind of talk a little bit about data Journeys and that their risks and challenges and why?

I like this idea and sort of why we stole the name from customer Journey mapping and CRM systems. And then we're going to go talk a little bit about sort of observability principles or data Journey observability principles. Then we're going to talk a little bit about our product and conclude an answer some questions.

So I'm going to stop sharing my webcam because it'll probably make your your screen a bit bigger. So I'll stop.

And so what is this idea of a data Journey? Well think of it as the path that data takes. from any Source in your company or organization to any destination in a company and maybe that destination is a dashboard or maybe that destination is a database table or maybe that destination's a model or maybe that destinations another system that's actually using it and we've all heard about the Myriad of sources but that think of it as a path or the journey that data takes And in a data journey is think of it as a an abstraction

00:05:00

of that path, right that goes through all the tools and systems and places where data lies. And the idea is a data journey is a way to not replace those systems but to observe or monitor those systems so that you know, what's going on you sort of know where your data is your reports are, you know, if your customers are going to yell at you, you know, if this sort of system is making sense and so you want to observe that system monitor that system check that system and it really the idea is it's context it's context for what's happening in the complicated world of your of data. And so the question is like why do you need this context? Why why is this why is it important to have yet another sort of idea and our complicated data World well, if we look at our complicated data World, here's an example, so a prospect got a call saying that the report was empty. He had to convene sort of two dozen people.

They spent all day looking for it. They found out the reason that it was wrong and then it took him several weeks to fix and he's got thousand other sort of data Journeys that he doesn't know and and the reality is he had sort of a kind of a streaming batch hybrid system. There's a lot of customers to this system.

There was a lot of fan out. It was very complicated and you know, he could tell the server was running but he couldn't tell if the data was actually good or passing through and so, you know data Journeys go bad and and they go bad for a number of reasons. So let's just go kind of take a step back and say well, what's your basic data Journey?

And so in every data and analytics system data comes in from somewhere. Maybe it's batched or streamed. Maybe it's moved or virtualized but it but it comes in and that sort of green arrow on the diagram on the left sort of represents the idea of a data journey and this picture beneath it sort of taken from from eccerson and it's sort of place in a database. Maybe it's in TL. Maybe it's elt. Maybe it's transform. Maybe it's clean and then think of it as used it's visualized it's modeled. It's sent to other systems reverse etlt and a reverse ETL and then there's some characteristics of it. There's multiple tools. There's multiple data sets. There's multiple paths. There's multiple methods. There's multiple architectures multiple customers and multiple people and so let's kind of take through that step by step. So, Everyone who's on this call is probably in some way has decided to move their data in analytic architecture to the cloud and the cloud providers have given us a whole bunch of new tools to do those things to put data in a bucket store. They have several databases from you know redshifts to bigquery. They have tools to transform data Maybe, it's AWS glue.

Maybe it's a version of Airflow sometimes they have their own visualization tool. Maybe it's looker or quick site. Sometimes they have their own governance tool. Sometimes they have their own and all of them have their own data science platform. And in addition to underlying things like, you know, storing sequence secrets and Version Control. So there's like a grocery store full of tools that you can pick and likewise the sort of startup Community as started to talk about this other idea of the modern data stack and here's a sort of a picture from Andreessen Horowitz is content that sort of talks about all the pieces in the sort of sequel engine and I guess the biggest thing is just look at all the boxes, right look at all the places that data is going and being transformed and being worked on.

There's a lot of boxes and there's a lot of errors in that diagram. And if you look at our sort of Circa 2001 picture on the right of a data warehouse and marks and olap and Mining. There's still a lot of boxes. Maybe not as many but there's still a lot of boxes going on. And so there's this just sprawl and tools that's happening in our organization. And when our customers say the data doesn't look right we've got these which box is the problem.

and likewise we've been

on top of the architecture the tools and the component. They're sort of a using a software term. There's a design pattern of our architectures and I'm going to go through several and talk about how they work. And so a lot of ways in data analytics we produce we get data we produce data and people consume it maybe in a dashboard, but sometimes we have multiple producers and consumers so I build a table and there's five dashboard still thought of it and and sometimes that is coupled meaning I can run it

00:10:00

myself and I know the fan out from it and sometimes I don't and so this producer consumer relationship is hard because there's not an overarching way to run all the steps in some organizations. They are sometimes they're coupled. That means they all run one after the other and there's some master orchestrator on top but oftentimes they're decoupled something runs and then wait a while something else.

House Runs, maybe somebody else has to do something. So this producer consumer architecture has been it's been prevalent for a long time. Another design pattern that's been gaining popularity is think of it as data enablement or the Hub and spoke team and a central data team sort of puts all the data together into a data lake or a data Lake Warehouse or pick your favorite term. And then there are self-service teams on various spokes who are doing the work and maybe they're building dashboards. Maybe they're actually doing some data work themselves. Maybe they're doing some data science, but this Hub and spoke model is a is become a popular architecture in a way to separate the work of sort of centralized teams from line of business teams.

And so that's another design pattern and if you take that to another extreme this this year has been sort of very fashionable to talk about data matches. And in that way those teams are actually self-sufficient their domain teams. So in this graph, I'm I'm a team working on a very specific set of data.

My specific set of customers and I can dig into the data. I'm building products that are continuously updated and we've done some webinars and written some articles on data mesh. And so this is another design pattern and maybe you like the data enablement pattern. Maybe you just like the typical normal producer consumer. Maybe you're really into doing data mesh, but these design patterns are put on top of all the tools that we use and they have the same sort of challenge of being coupled and decoupled systems, but they also have another way to think about it. In addition to that coupling.

Some of them are batch and some of them are streaming. Some of them are event driven sometimes and so a common architecture patterns that I've seen is that the data is ingested in a streaming way and then the serving part the data warehouse models are sometimes sort of fast batch. Sometimes they're not a lot of companies desire.

It be kind of a fully addressed event driven system, but the complexity beguiles them. So they have kind of data loading as a ventriven and and data it data integration and visualization is batch. But so you've got this. Set of design patterns on top of our tools batch versus streaming producer consumer Hub. Spoke data mesh all working on top of how we build systems.

And so when something's wrong, you've got to kind of think about well, what's my design pattern? Is this a data match? Um, where is it in the tool chain? And then you've also got to kind of go into the depth of where the problem is, right? Because the problem is not just an across problem across your tools and across your design patterns.

It's a deep problem. And so let's say something goes wrong. And so you've got to go in and find out where that is and then find out what it's going to affect and so there's a lot of layers here in this diagram. So one layer is the server that it's running on, you know, maybe it's virtualized. Maybe it's an actual server. Another layer is the actual tables themselves, you know, maybe it's an FTP site or that's three bucket or a Tableau report. Another level is the exact tool that you're using to do the work.

Maybe that's problematic. The third is maybe the tools fine. Maybe it's the code that happens to be in that tool that someone puts some SQL into DBT that's causing a problem or there's an IPython notebook. That's messed up. And then there's a layer of kind of think of it as timing or orchestrating. You could be using Kafka you could be using Airflow you could be using cron jobs. You could be manually extracting things.

And then, you know more and more that I'm very happy to do is people are actually checking their data during production and there's open source tools. There's testing and DBT there's test libraries and Airflow. There's python test libraries. We have a test engine ourselves and all these pieces are kind of help. You've got to find out it's across a problem and a down problem on finding out where the where the where the issues are.

So that's why. Companies have a very hard time understanding how to answer basic questions about the

00:15:00

data Journey like Kenya map all the paths that data takes on the data journey in your organization. Can you just write it down somewhere and oftentimes? It's not in anyone's head. It's maybe there's one smart person who knows the whole system. Maybe they're in operations. But people it's you know, there's that old adage touch the elephant. We all are sort of touching touching the elephant and

you know, you know just simple questions like did it run and is it is it on time and was it refreshed with the latest data and you know is my CPU being pinned is my data source Clore quality did a bunch of jobs that have to happen first. get done and then the jobs that have to happen second were completed and how many things ran yesterday and how many things are gonna run today and there's a lot more sort of basic kind of questions and maybe you could call this runtime information or runtime lineage, but oftentimes, these are sort of basic operational questions that Our customers or our customers and and teams can't answer.

So what does this all mean? Well, it means that there's problems right is that this sort of? Myriad of complexity and data Journeys the lack of being able to check data during those data Journeys means that companies are having problems and and that my favorite one is the one that we did a survey last year of that 78% of data Engineers are so stressed. I mean they want a therapist.

is you know, our customers are unhappy projects aren't successful and we're having stress and You know frankly. I'm a bit worried that like the The Glory Days of data and analytics are behind us and and the CFOs are gonna start at questioning. Well, why am I spending some money? I might so much money on my dating analytics team. I haven't seen that yet, but that's happened in other downturns.

So what are the principles here? And and I think the problem is is there right that we've got design patterns of mash and producer consumer and hub and spoke on top of a complicated architecture of lots of boxes and lots of tools and lots of cloud vendors and private software vendors, and we've got a tech stack of code and tools and servers that we run on and when something goes wrong, we've got to kind of find a way through that through that morass of complexity just to find out where the problem is.

Um, and so I just saw one quick question. Will you send us the slides? Yes, we will we'll send the slides in the recording after the after. After I finish with this so tomorrow the next day. So what are the principles so you've got a complicated Journey? Right? Well, how do people handle complicated Journeys? Well, let's take NASA. Um, they had this thing called Mission Control, right? And so what are the characteristics of Mission Control they built a series of uis that brought in every aspect of the flight, you know, is it going in the right direction is the engine hot? It's the crew still alive and they use that as a basis to make quick factual decisions about what's going on.

And then they store that information to learn from it and then overall they also try to get alerts and little red buttons. If something goes out that they manually monitor things but they also get alerts and that idea of a mission control I think is what we need. We need the idea of a mission control for your data Journeys.

And again a data journey is complicated because they're everywhere in your organization and they have relationships to each other and their batch and their streaming and maybe they're data mesh pattern or maybe they're a data enablement pattern but there are a lot of complexity in your organization and you know, there's a lot of tools and and you know, you may have someone who's trying to consolidate on to AWS or Azure or has two or three private tools, or maybe you know you're going to gcp.

And of course, there's companies like Oracle that has still has all a lot of the databases in the world and there's just a lot of categories of tools. It's a 80 billion dollar industry. And so remember your data Journeys go across all your tools and they also go. Kind of across your teams. So this is an example of a typical large Financial Service Company organization, right? They have they have banking they have insurance. They have high net worth individuals maybe brokerage. They have a supporting organization with CIO and maybe the CDO and then they've decided to like have Teams kind of own their own own their own work and so

00:20:00

each team and each group is kind of running their own data Journeys and it's complicated because sometimes you know, there are centralized teams whose data Journey feeds into so for instance project F will feed into different teams. And so there's lots of owners and dependencies and architectures here. And then also our data Journeys themselves are not always their in production, but they're also in development and so we've got to be able to push things from these Journeys and changes from these Journeys from development into production.

And this process of deploying is also varies in different organizations. And so I was talking to a company this morning. They have a hub and spoke architecture and they have their home architecture team is using

Amazon tools like you know code and pipeline Unfortunately. They don't have any automated tests. So they're manually checking things and then they're spoke teams are kind of doing whatever to get things into production. Sometimes they're just dragging code from their desktop. Sometimes they're mailing files to people. Sometimes they're doing it manually. There's just a lot of different ways that people will deploy into production. There's not as if In larger companies, there's oftentimes, you know, their self-service people.

They're just pressing a button on the Tableau dashboard to say push to production. And so there's multiple teams in multiple paths to production and lots of organizations. and so as a result as I've shown before there's just a lot of levels in Tools in many components and you know the data Journeys themselves are not There's not a single thing. There's not one data journey through your organization a data Journey can feed into other data Journeys.

And that also means that this idea of Conway's law is true that you know, perhaps it would be best. If you had one data Journey that encompassed all the value of source to customers in your organization, but you know, you've got different bosses and different teams, maybe so what Conway's law means is that your organizational design kind of equals your Tech design and I think this is very true indeed in analytics and and That pieces of the sort of value stream of delivering insight to your customers are owned by different parts of the organization and they have different relationships, you know that where people provide and I think data matches the the great example of that and as well as the sort of Hub and spoke and producer consumer relationships.

And of course this idea of correlation, right and and if something goes wrong in one journey and here we can see a data Journey that feeds into one with two customers of it. And what happens if it goes wrong, how can you tell what is affected Downstream? And so if you're doing Hub and spokes something breaks in the hub which spokes are gonna be affected. Um, if you're in a data match and one part of the matches has a failure, how does that affect other parts and then we're in the problem. Is it is it bad data is it bad servers and and my experience and sort of doing data engineering and data science and teams. It's yeah, I think a lot of times it's the data or the integration of the data where we find the problem, but sometimes it's in the code acting upon the data Maybe the dashboard wrong or the model needs a tweak and you know rarely, but but sometimes we just run out of server space, you know on our machine in that.

Of the problem or someone just misconfigured something in the environment. So we need to be able to correlate and sort of find out where the problem is because our customers, you know, we sort of run run a factory in our customers expect these things to happen on time. And maybe Factory's not the right metaphor or maybe it's plumbing and our customers just expect the plumbing to work. They expect their data to get there and when it doesn't they're they're upset.

And I think this is a key slide in here in our discussion because the ways that Journeys relate to each other is is complicated because of these architectures The Hub spoke data match producer consumer streaming patch. And so if we look at how each part of the journey connects to each other right like let's just go through a causal one, right the completion of Journey one causes journey to to start.

Um, there may be a temporal one the first journey runs at 2 am finishes and then there's a separate schedule that runs at 5 am. And well what happens if Journey one? Finishes at 7:00 am well, then you're gonna have a problem in Journey too. And this is a comp a very common organization. We've got two scheduleers each running independently of them and something goes wrong. And so and then

00:25:00

there's manual cases where something runs on a schedule but you know, Tom's got to press a button every morning to upload the dashboard and then there's sort of periodic ones for instance you have streaming events that are coming and continuously and then every hour you run a match on top of it. and so how do you know if is the batch running is the stream running is it has the stream stopped publishing events, but your schedule still running.

These are questions that I think we need to answer and then there's just there's full of venture of an architectures. You have a stream and every time you have a new data item the dashboard tree calculated and then there's kind of the idea of inferred connection. So what happens in a lot of organizations and I've seen this is they build a data lake or data warehouse would 50 tables and then there's a thousand dashboards and there's no relationship known between which dashboards are being used which dashboards are using which tables so if a table's wrong, you don't know who to talk to um, and likewise you don't even know if it's the usage is there.

And then from a kind of the relationships, there's the sort of fan in case where I'm getting a data from a bunch of different places and I'm building a data market very common. And then there's the fan out case I built a data Mark. I've got all these people using it or I'm joining it. I've load some raw data, and I've got to update these three different data matches.

And then there's our friend the dag of different Journeys and then there's kind of subcomponents where a journey has to is running but it also causes another journey to go and go off and this kind of organizational relationship is often tribal. This runtime lineage is not anywhere in the organization. It's not actionable and not know and so I think this is what creates the fact that when people have problems 26 people have to get on a phone call and and figure out all day where the problem is because this isn't anywhere in an organization and it's certainly difficult to diagnose.

And what that means is that your Journeys have expectations on them. They have timing expectations. They have data quality or data goodness, or you know dashboard is full dashboard is full with the latest data, you know dashboards have the right values for your total number of sales those kind of metrics. So those expectations should be there and then you should judge the reality of what's Happening against those expectations.

And is there variance and I've got that to this layer of expectations of a data Journey again, it's not very often in anywhere in an organization. Maybe it's in a spreadsheet or a Wiki but it certainly not actionable and judging the expectations of what happens against the reality. And then lastly so you've got these expectations.

And you've got reality and so what's happening is you're getting think of it as you're getting events from your internal system Airflow started tables been filled reports then updated and once you have those events and you're telling if things are right or wrong you should be able to alert and be able to send things to people and perhaps that means I'm sending it to a tool like jira or slack or team. Maybe it means sending it to a person finding the right person. Maybe it means stopping the actual production because the date is so wrong. It's sort of the and on chord And there's a lot of things that you're interested in in terms of events. Did it start? Is it running? Did it stop was its schedule is the data right the test work?

Did your DBT test go good your test of? the actual dashboard work and so these idea of expectations versus reality and then when reality doesn't meet expectations creating events and kind of think of it as an event you want to do something with that event notify people notify the right person because the there's a term that comes out of Software called the mean time the recovery the MTT R mean time to recovery and you want to automate that right you want to be able to say okay. There's a problem. It's a problem with a server. So I have to contact this person. It's a problem with a data set is it's the problem with my data provider. It's a problem with my actual Airflow job.

Well, maybe I got to talk to the Airflow guy and shortening that time to recovery sort of who it is where it is. Where's the location? I think is important part of these systems. And so the mttr meantime to recovery is important but a lot of times we're not most organizations aren't at that. They're not at improving their meantime to recovery that they don't know what the

00:30:00

problems are and their meantime to recovery is really when a user finds a problem and and then they're all scrambling. and then once you start to get this sort of runtime lineage looking over time what happened last week, what is the quality of our system in terms of quality checks, what's our on-time system and think of it as um, you know, you're sort of boss is awesome report how to prove that your team is actually being productive and and the key is to look at it longitudinally over time keep a database of it and then use that to understand what happened and use that to improve and understand and I think that's actually a very key point because our company is really been about trying to get people to do testing and orchestration and and deployment automation from for many years and we've seen a number of customers kind of think about doing DataOps, but then something else comes along and I think by getting and compiling this information about how much of a struggle it is to run Systems, it's a great way to actually get people motivated to actually make the changes.

And we'll talk about that in a little bit. And then the last idea of data Journey as context. And so a lot of people care if things are working right the you may have a production team in another country you build some stuff and you throw it over the wall and they run it.

Well, of course they care. And then you know, you may have your person who's your data engineer your data scientist to written some code and they want to know if it works, you know that there it's important to them and maybe you've got a data team leader who's you know wants to will is the person who's gonna get a call when things go wrong? And then maybe you've got a Data customer and they just want to know is my data arrived and can I trust it? And so this idea of a context of all the data Journeys is shared among all these people.

Um, and one of the hassles of being a data engineer off and is your data customer keeps telling you hey, is it late? Can I trust this data? When's it gonna be there? You know, the data team leaders asking you for metrics and so a lot of people spend time in communication that could be removed by having this idea of a shared sets of data Journeys that they can that they can communicate on and just look at some system first before kind of talking to each other and so I'm a you know, Talking to each talking to people's nice, but giving people tools to allow creative people to spend time doing good work. I think is a good answer.

You know and then honestly people are interested in different things, you know data customers sort of interested in the things that I care about. They're not interested in every pipe one in the organization. Maybe the production Engineers interested in every Pipeline and every dashboard fill and every raw data load, you know, the data developers interested in the projects, they work on and the data leaders interested in the teams that they manage and so this idea of sort of mapping the organization to all these things and giving people tailored views is important.

And then um, you know, of course data Journeys need tests and and this term test is interesting because what we mean by test is an automatic quality check. of the data the integrated data or the things that happen on data, um during the production process or during the development process and so we have tests and the market hasn't really given us a good term for tasks. I just happen to like test some people call them QC checks. Some people call a monitors some calm data quality checks. It's just we just don't have a good common term, but we need these checks and they need to happen everywhere in it's not just is the raw data right because oftentimes you're putting data together and maybe the logic is wrong or maybe you only see the problem in the Raw data when it's been integrated with other data and that happened to me as a data engineer on several occasions. And then maybe your data is all perfect.

And your Warehouse is all perfect. But somebody mocked up the reporter model and and for all the good data work you've done the view of the Isn't right. And so for all those reasons you need tests and you know, we think that tests should happen in production kind of on top of your tools that you should use whatever way that you want to create tests row counts exception statements and SQL python tests Etc. And we've written quite a bit and some books on the different types of testing but the benefits here is that by having less errors in production you get more time to do good things.

And you get more customer Data Trust And So at the

00:35:00

end of this we've got links to our books and we've had webinars on testing and this is a key concept also is that this idea of tests have a dual nature in production, you're running your data varies, but the code that's acting upon data is fixed and your tests are really trying to prove that your production system is working in the way that you want it to but in development the code is very but your test data should be fixed. And in that way you're trying to find out if I made a change to the to the code or the configuration or the model did it have impact on other parts of the system and that double duty I think is an important part of of how we work because you know like that customer I talk to you this morning. They spend a lot of energy and automating the train of deploying new code into production, but no signaling infrastructure, so they still had this whole manual Up, and it's kind of you know that takes a lot of time and then sort of why automate if you're still going to have manual Parts in it.

And you know, I'm a lot of people have been getting the testing religion and so certain data Engineers will have row counts in their ETL tool. Maybe the python people will have some python checks in that they can that they can monitor of their model maybe some people write SQL tests. A lot of orchestration tools will have test Frameworks. So for instance DBT has tests, they're specific open source, testing tools that you can use there's closed Source testing tools. There's python tests. There's just a lot of ways to check and remember you're not just checking data. You're also checking the tool right? And so Tableau can do lots of cool things as well as looker and so all these things of the code and the tool acting upon the data.

You need to check and write tests against every part. And so there's a bit of an eye chart, but I'll just include it in and so it's kind of think of it and it's sort of a concept model of testing and so you want to look you want to do testing on the raw data the integrated data the tool modified data and you want to check the status in timeliness. So these are sort of checks that you want to do and then you want to check the process.

You want to check the data structure and syntax? And then you want to check the context that's based on really sort of based on your domain. You know, how many customers you have? What's their average sales? What's your total number? Sales? What it's what's it's break that what's it's broke out by region and be able to have those things are very specific to your company.

And again sort of when did you testing is kind of across all your whole data Journey from source to value on every tool during production against live production data and during development against test data. And when you test, you know, you want to you want to have notifications you want to air be able to error out you want to have the it's looks weird. I'm not sure if it's rights answer and then you want of course have that it passed or it's logged and then there's just a lot of ways to think about how you do testing. And so maybe you just profile your raw data and qualify it based on. Okay, this column's got three values. So I'm gonna have a test to check make sure it has one of these three values or it's a valid zip code. We've talked a lot about statistical process control. That's sort of variance over time. You could also have variants Overlook location as it pop through your bucket stores or multiple levels in your bucket store your database you could also look at variants across populations. So the previous time that you've given a report the current time.

And then they're sort of a variance the sort of Time series anomaly of various metrics. And so there's a lot of sort of categories of Productions last development data testing that are out there. And so and then there's also another way to think of how do you categorize your test? It could be statistical process or location balance or population balance or time series anomalies. This is the way that I think of it you could also think of it from software development where it's sort of unit tasks which are quick tests that are happen during data during code integration performance tax functional tasks regression tasks impact analysis tasks and you oftentimes you the single test can have dual Duty.

It could be a functional task. That's all so a location balance test. And so these again we don't have the right language yet and data to talk about the Dual nature. And so we just use the term test because it's the simplest way to say. I've got tests and development test in production and we can categorize them in two different ways.

And again for more information, you know, check our books and blogs on that. And let me keep going and so this idea here of being able to.

00:40:00

Check build data Journeys that represent the path of data through all your tools and organizations is a lot like this idea of application performance monitoring and so, you know in my experience tools like New Relic or data dog are great, right? You need to monitor your servers. You need to monitor your network your security. Those are all very important tasks. But what they're missing is is that they don't actually check the data or the integrated data.

And as you know all the years I've been doing data. It's kind of like a couple of percent of the time things go wrong because it's like the server and the network and yeah, it's annoying but oftentimes it's wrong because the data was wrong or we did something as a team along to the data.

And that's really where it comes. That's the high 90s percent of the problem and I I see application performance monitoring doesn't really help you with that. Um, and so because it doesn't have in our argument here at least in this is that it doesn't have the context of what's happening in your organization does have the data Journey the expectations versus reality the alerting all those things. It's missing and that sort of coherent context is really essential to truly observing and monitoring the complexity of your analytic systems.

and so lastly, I think there's another case of like Data governance, and how does this idea of data observe data observability enhanced data governance. And so there's a lot of definitions of data governance, but two of the bigger outputs of data governance or a data catalog, which is a list of columns and types and descriptions and then data lineage which talks about where the data came from and what the steps are and those are really good things and there's lots of tools that can do that. It's a it's a huge Market.

But there's other things that I would like to know is well, I I got the catalog and and should I trust it. My great this is what it means. This is where it comes from but like is it right? And so what's what tests were done and we've had one of our customers was very proud and and everyone of his, you know a big team at a Pharma companies like okay, we run 10,000 tests every week. And here's the list of tests that we run and it gave people confidence that that the data was good and helped him sort of show his awesomeness. And then the other part too is think of it as instead of data lineage think of it as one time lineage.

Is it started? Is it running? When will it be done? You know and and has it been updated? And one of the benefits of putting even combining these things it's really about trying to reduce the hassles of your team. They're not asking you for what's the meaning of the table? They can go look in the catalog. They're not asking you should I trust the data will go look at the test results. They're not asking you when's it gonna be ready? They can see that from your sort of your dashboard.

And so as we go on here, we've got 1242 and I can see that we've got several questions. I want to get to the third part. We just kind of talk about of course. We're a software company and have a product that does a lot of these things and so in fact a lot of the design points of the product are based on the sort of observations that we made over the years and and in some ways the the description of what we're what we have here is is to answer these questions that we've raised and so we have a new product called Data optimizerability. You may have noticed that our website has changed and it really does the set of requirements that have I talked about in the previous. It's about an enterprise-wide view of all the data Journeys all their complicated paths and does sort of monitoring of alerting the tools and data across any tool and you know it what it does is it has this idea of the data journey in it and in you can build a data Journey that represents the complexity of Hub versus spoke or producer consumer or data mesh or any of these patterns into the product and then you can set expectations on it and have rules that alert go off of it and I think that is really what's missing and you know, and and also there are and I think that idea of a data journey and representing the complexities of our design patterns is important and then storing information and sharing it either to business customers or other people and giving the organization of you of that data both to share as well as the kind of understand and analyze and they're sort of a bunch of charts and graphs that talk about history and dashboards.

And and really we've tried to make it very simple. It's basically an API where you can push all this information in and so our you know,

00:45:00

and and the data could come from DBT or Airflow or synapse or any tool that you have in your data and analytics stack and so and any automated test engine and so for us we we have a new product called Data optimability. We've renamed our existing products and DataOps on Automation and they work kind of synergistically right is that um and really we think that the idea of observability is kind of the first product that people should use and our automation product is really a way to feed information into observability. And so there's a lot of new content on it that we'd like to talk about but we think it's a great a great new product that that you all should check out.

And so let me just finish up here and get some questions. So like what's the feeling if you do data observerability that you want to have well, I think first of all, it's less embarrassment right? No one likes to know but we are as an industry relying on our customers to check our results and and that's just embarrassing when they're calling you up and saying the data is wrong and kind of rolling your eyes at you and saying it's and it happened to me. I don't like it. Um, and I think that's that embarrassment actually leads to hassle. Right people are hassling you and not trusting you and and they're constantly not giving you time to do the work that you want to do and not having space to create. So the embarrassment and hassle actually is a huge time suck.

And so for us we actually think you can't focus on getting most custom most people in data and analytics run their shots with a ton of errors and they either they know it or and their customers are unhappy and and most data analytic organizations just have a backlog of a bunch of stuff they have to do and they're not getting enough time to do it.

And so our Point here is start first with observability sort of use and observability software. It kind of find the problems in bottlenecks and reduce the risks of Hassle. And then when you find a problem or a bottle neck automated at a test, um improve how you deploy improve how you orchestrate improve how you build development environments. And so this sort of two step process is what we have have learned over the last four years or five years of kind of pushing or six years now pushing DataOps sort of focus on observability first get the metrics of of what's happening with your data Journeys and use that as leverage to add Automation in your organization and say look we've had all these errors let us write some tests. I know it's going to take some time but it's going to enable us to do things faster and and that way you'll end up getting more things done and and sort of the big the big idea here. Is that most organiz And take a long time to deploy. They run. It's very rare that they have an error-free day. And actually their productivity is low because they're sort of running around fixing things and that leads to sort of high cost teams and unhappy customers and I think there's you can and as well, it's actually unhappy teams as well who are doing the work.

And I think if you apply observability and automation, you can push up on all these things. You can run with less errors. You can deploy things to production quickly with high confidence your team productivity goes up because they're delivering things often and you're actually maximizing both the amount of work you don't have to do and have unhappy customers and we talked a lot about this and other organizations where the sort of almost a factor of 10 Improvement in the productivity of teams because they're doing things that really matter and they're doing with short cycles and low errors.

And so I want to conclude with this slide that has a lot of information here and you'll get this in the deck. It's links to our books. It's links to the manifesto that sort of about data opposite idea. And then we've actually got a whole bunch of new content around that talks sort of a white paper that talks about the sort of principles on this. We've got a sort of a technical white paper on our observability product.

And of course you can go see that in there. So I want to thank everyone. Um for this and I'm just going to take us take a breath here and I'll answer some questions. But if everyone remembers that we will send this DAC we will send the recording out to you via email and it will be available from our website in the next day or two.

And so let me let me look at some questions here.

And so yeah, so the question around the statistics it says the question is that's those statistics look so exaggerated. And so I guess I'm not sure of the statistics but the sort of perhaps that was around the project failure statistics and

00:50:00

unfortunately, that's not exaggerated and it's been sort of proven over and over and over again in different ways that are Ways to help our organizations be data driven. It's the mining. That teams aren't productive that work people aren't getting the analytics Roi and it's sort of a well validated assumption and it's sort of hard to believe given the sort of noise from the market about how data and analytics is taking over the world. But that's sort of the Dirty Little Secret and I think one of the reasons that you've seen a lot of the ideas from software went through the same sort of crisis of confidence and the idea of agility and DevOps observability value stream management domain-driven design.

All these Concepts that are kind of being ported into the data analytics world are really because of that same concept so I don't really think it's understated. Unfortunately. I wish it was

and so

yeah, and so I guess the there's a question about our software and really being able to sort of integrate custom jobs and I think My view is that every job is a custom job right? Because oftentimes you may have something scheduled in Informatica or Airflow. You may have a batch process that runs before Airflow or after Airflow, you know, you could look in sort of the ETL tool run view or Airflow run viewer to extrude run be great you get that that's good, but that's only part of the picture because you may have ingestion processes in front of it. You may have modeling processes or consumption processes off of it. They may be scheduled. They may be not scheduled. And so this kind of think of the greater world of what happens in your the your data journey is important. And so I think that That's what we're trying to represent is instead of.

I'm the automation. I'm the Airflow guy. My world's good. I don't have to worry about what's in front and back the idea of building data Journeys representing that complexity of architectures that complexity of design patterns that complexity of tech Stacks in software will help you find these problems and and allow people to see the full journey and and so working with custom jobs custom orchestrators all the sort of batch tools out there. It's very part of the design part why we build it as an API driven product.

and so I think that's the last question and I see so again, we'll send out the deck. Thank you for your time. Really appreciate you joining in and you might have a great day.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What is a data journey?

A data journey is the full path data takes from source to the insight a customer receives, including every tool, data set, method, and person along the way. It tracks all levels of the stack, from the data itself to the tools, the code, and the tests, and it supplies real-time status and alerts so a team can tell what ran on time, what failed, and where. Most organizations hold this map as tribal knowledge rather than as something active and actionable.

Why is application performance monitoring not enough for data pipelines?

Disk and CPU metrics are lagging indicators of data problems: by the time infrastructure looks wrong, the bad report has already shipped. APM tools do not check the raw data, the integrated data, or the reports and models created from that data, and they do not give a context that shows the pipelines, jobs, and tools acting on it. They also cannot synthesize production runs and development work into one coherent picture.

What kinds of tests belong in a production data pipeline?

Five categories: raw data profile qualification and consistency tests, statistical process control, location balance, historical or population balance, and time series anomaly tests. Testing happens in four places, on raw data, on integrated data, on tool-modified data, and on tool processing status and timeliness, and it runs both in production against live data and in development against test data. Results are graded as log, warning, or error, and most tests do double duty across development and production.

What questions should a data team be able to answer about its pipelines?

Whether the job finished successfully, whether the dashboard or data set is correct and refreshed with the latest data, whether source files arrived on time and at the right quality, what resources the process consumed, and whether job X ran only after every job in group Y completed. Also how many jobs ran yesterday and how long they took, and which pipeline is troublesome with frequent or intermittent errors. Teams that cannot map the paths data takes cannot answer any of these.

How does DataOps observability support data governance?

A governance program answers what the data is, through a catalog of metadata plus management and search tools, and where it came from, through lineage covering origin, changes, and movement over time. Observability adds the two questions a catalog cannot answer: can I trust the data, answered by test results recording what tests ran against the data and artifacts and what they returned, and is the data fresh, answered by run-time lineage recording when the source and integrated data were last updated.

What are the two steps to adopting DataOps?

Observability first, automation second. Observability reduces risk by finding problems and bottlenecks across the toolchain and giving a team enterprise-wide views of hundreds or thousands of data journeys with monitoring and alerting on both tools and data. Automation follows, adding testing and orchestration that fix the problems so they never happen again, which is where cycle time, error-free days, and team productivity improve.

Where to go next