On-Demand Webinar · 32 min
You Need Mission Control for Data Operations
Chris Bergh's session from the 16th annual MIT Chief Data Officer and Information Quality Symposium, July 2022, on what a mission control for data operations is and why a data team needs one.
Transcript
Show chapters and dialogue 5,301 words
00:00:00
The biggest challenge and data analytics is that projects are failing Gardner estimates that 50 to 80% of all dating analytics projects fail despite having great new people who are data scientists create new tools and databases 150 billion dollars in spent every year in Tech is fundamentally comes down to a people and process problem. And that's what data office tries to address. It focuses purely on helping people in their process with their tools with all their data be successful and deliver unparalleled insight to their customers.
here
Good afternoon. Welcome to track five leadership in action learning about different success stories Lessons Learned. I'm Baba Odette partner data management at guide house. And today's session is going to be mission control for data operations and It's really a must-have and some of the sessions going to cover off on a couple key questions one is why is the lack of enterprise-wide what is why does the lack of enterprise-wide view of the multitude of data analytics pipelines in production?
Cause repeated failures and missed deadlines. Why reducing production errors and monitoring data teams development dramatically improves team production and many other types of questions related to this topic. I do want to cover on a couple logistical things. The session is being recorded and the contents will be available and who have app secondly you'll be able to access this session 30 or three months from now and we'll be out on the MIT. Cdoiq YouTube site.
The other thing is regarding questions for those in the room, please if you do have a question kindly raise your hand, I'll give you the mic and then please pose your question. Otherwise those that are virtually can hear your question those that are virtual. Please use the who've app to actually pose your questions and throws in the room. Feel free to use that as well.
And with that said I'd like to I'd like to introduce Chris CEO of DataKitchen. Chris is a really distinguished more than 30 years of research software engineering data analytics executive management experience. He's held multiple executive titles coo CTO VP director of engineering. Here's a Master's in science of Columbia University a BS from the University of Wisconsin-Madison.
He's recognized experts on DataOps published numerous books on the topic the DataOps cookbook. Chris really began his career at mit's Lincoln laboratory and NASA Ames Research Center. So now I'll pass it off to Chris for a session. Yeah. Thank you so much. So
It was interesting looking at the application today the MIT application. They had a survey of like what's your biggest problem in data and analytics and the number one thing was Data quality or you know when I think what it said is people don't trust the data and so in some ways that's a risk like why don't people trust the data.
I'm there we go. And so I guess from my perspective. You know, I've seen about 15 years working in research and software development. And then I joined the data World about 2005 and it was hard. Like I had a whole bunch of people working for me to data science and engineering and things were breaking left and right and I had this thing called the morning dread.
I don't know if any of you have it I went to work and I just didn't like what I would see sometimes I would hide in my car sort of bracing for the day. Oh is something gonna go wrong is the dashboard gonna be late? Did somebody screw us over with bad data? And so it was difficult and
You know as an example just recently talked to a company. He had a data report go wrong. And the CEO of the fortune 15 company called him up to yell at him. It's ahead of a data team had to get 26 people for six hours to find out where it is. And that just seems too much risk to me. You don't want to dread going into work. You'd want to have all your smart people in your team diagnosing what a problem is. And so what really means and what our company is about is trying to address these challenges because it's too hard and it's too frustrating to build insight. And as we go it's my fifth or sixth year presenting here and a lot of data and analytic projects don't work. They fail. In fact, what's interesting is we did a survey last year with data Dot World who want to speaking
00:05:00
in the next meeting and that 78% of data Engineers are so stressed. They want their job to come with a therapist. so why is that like are they mentally ill or just the job just kind of suck and so I think the idea here is that we have a very complicated world that we run in complicated data complicated tools just in the last five years the move to the cloud has brought in a whole level of new tools and Technologies.
And so the assembly line that we use to build Data Solutions from data engineering to data science to data visualization is complicated and sort of if you walk the journey that data takes from its source to is value in a rapport or a model. It's it's complicated right and and it's not just one tool one place. It goes through many places and it's almost like manufacturing insight. And if you're in charge of a data team a lot of people that we've talked to and a lot of our customers before they work with us couldn't answer these basic questions like did the data arrive on time?
You know are my new reports or my new models done on time. Should you trust the data did I check it before you saw it? How many jobs ran yesterday versus today are they late? What's coming today? And is the model that I have in production still accurate or is the dashboard still working? And it seems to me that they're sort of basic production level analytics that were not doing very well as data and analytic leaders And so some ways the theme of this talk is that you should get analytic about your analytic systems and a way to do that is to think of it as risk. And so how to how does how did NASA manage risk? Well, they put everyone in the room. They gather all the information about takeoff and status of various things and put it on a bunch of screens.
And then they kept it in one place over time. So when something goes wrong, they could analyze it. And so this idea of Mission Control is what I'm going to talk about today. And the real point of it is to alleviate the hassle in the pain that you feel in building and delivering analytic systems and how we arrived here is that we've been talking about something.
called DataOps data operations for years, which is really the application of Sort of agile and DevOps and tqm techniques to data and analytics and so we think the first step is really to get the problems to get the errors under control and to really build sort of Mission Control System. And so the next part of this is really to talk about kind of what the characteristics are of the problem and and why it's worth solving and and so I'm going to talk about this as DMC your DataOps mission control and so one perspective is that this manufacturing perspective on data and analytics and whether the data is big or small or comes in in batch or in streaming or whatever. You have a tool chain that you manufacture insight and it goes into a database it goes through data transformation tool. It goes into a model and visualization and just like a car is assembled from its components dashboard or a model is assembled from its sub components. And so you've got lots of tools acting on your data and so but in different A Toyota factory or maybe you have one of these assembly lines you have dozens. If not hundreds of these assembly lines running simultaneously in your organization. And there's a in the bigger the company you are the more of these I'm going to use the term pipelines to talk about this but it's really think of it as the journey from raw data all the way out to customer value. They're everywhere and so come, you know, we talked that company that had the issue with 26 people. They have 2,000 production Pipelines.
and no ideas sort of where they are because why is that it's hard because The operations in your data and analytic team. These pipelines are a reflection of your organizational structure. They live in different places. Some of them are Central. Some of them are Hub and spoke. Some of there are aligned by functions. Some of them are aligned by division.
And so the other aspect that's interesting of these data and analytic systems is that they're not static. So when you manufacture a car, maybe you change the the car assembly line once a month, or maybe if you're Tesla, you do it once a week, but our data and analytic production systems can because they're software they can change quite quickly.
So you have a process to put things into production from development to get things from your data engineer and data scientist head into that Insight system. And unfortunately, it takes a long time for a lot of organizations. It could be weeks or months before something happens before that
00:10:00
20 lines a sequel that new model is actually taken from the desktop or the development environment and put into production. And the also the process of putting things from development into production is not consistent around companies. So they may have multiple environments a Dev a QA uat. Some of them may just push it a button.
They may be a self-service team and they press a button and Tableau that deploys the production there have any process and so the standards in how you get things into production vary. And so one of the ideas of Mission Control is that you should be able to understand what's happening. You've got all these production lines. You've got all running and production. You've got a process of getting things into production. And so one of the ways is like take that information about that and put it in one spot.
And because you put it in one spot the constituencies and operations that people who run operations the people who build the factory the people who manage people who build the factory even business people are using the information are used to it. and that sort of single view of your operations is what we're what we're about because you know, if you're a data engineer, do you really want to answer questions from your customers about is it ready?
No, you want to be able to do good work. If you're a manager you want to be able to answer questions of how is productive as is my team. And when someone asks says a problem in the data and you can say no we ran 4,000 tests on that data.
And if you're a person who's using a production person who's in charge of running these thousands of pipelines you want to know if something's wrong. So this sort of shared context between these constituents in our organization is something that's lacking and and why is that so why is this? If we have an assembly line things are breaking we have customers who are demanding lots of new things but it takes a while to get in why haven't this been solved already? Well, there's a set of technology that it organizations have called application Performance Management and there are tools like data dog or Splunk and they accept things from the server level sort of things like the CPU the disc space process errors and they have logs and traces that they go and it's great. Right because that's really good information. However, it misses the point. It's sort of looking at the floor when you're really should be looking at the assembly line and our systems are built on top of this information. And so they don't actually check that the data is right or that the dashboard is still showing the right information.
They're good for Diagnostics, but not good for getting actual information. And so for us we actually think that you should you need to build something above what you do and for the purpose of this presentation. I'm calling it an observational Pipeline and so it doesn't run things. It just observes what goes on In this picture below you see all those tools.
Right. So imagine a really simple case I get data it goes in an S3 bucket goes in the Snowflake. I use Talon to do data transformation. I've got a python model and then I use click to view the data. So you get four five tools a couple of different locations. Did it all work did it all get through when I get a call who we're what point did it break and so being able to look at the world? And so if you look at this case of a process to get value from data And think are really simple case where data comes in. Maybe you have.
10 files that come in it gets into some buckets it gets into a database and gets into a table report. Then you have a bunch of tools. Maybe you're a fan of the modern data stack and you're using five train and DBT, maybe you've got some jupyter notebooks and some Tableau. So now you've got dozens of data sets. You've got dozens four tools and then you have some ways to run those tools like Apache Airflow which orchestrates or schedules maybe you have a Quran job. And so now things are starting to not go. Well, how do you know where it is?
Is it the data? Is it the transform data? Did Tableau fall over did someone put in a report in the town below that had a calculated field that was wrong. Where is it? Is the server getting too bad. And so this idea of observational pipeline. Um encapsulates that idea abstracts it
00:15:00
and also it's it's useful because it's a shared context for the organization to understand what's going on with their data operations. And what's interesting is that these ways to view the world this observational pipeline is an abstraction of the complexity of production. And so there's a lot of levels to do this did something go wrong in the report a node in the Airflow graft.
Did it go wrong in the data is the server pinned and so we want to the idea here is can you abstract that up? Can you chunk it into reasonable bits of information to get good understanding by people? And then second. Can you actually represent the tribal knowledge of operations? And I don't know how many of you have put things into production where it requires a button push or something scheduled and then something they'll schedule that ones at 3 AM ones at 5am and if they sort of run completely independently of each other and suddenly your customers are saying this is wrong and you're like I don't know is the data wrong did something not running order and also there's relationships between organizations. So we talk to a lot of big companies Central it is having a lot more sort of Hub and spoke model where it's a data enablement organization. And then you have self-service data. So you bring in a data set and it has a relationships to other data sets that it could impact and then there's the idea of a data mesh right the idea of domain driven design and having so you get an operationally complex environment with interrelationships between each piece.
And so that tribal knowledge isn't anywhere or maybe it's in a document or maybe it's in one operations head. And so you start having the death of a thousand. without this and what the idea is that you should build an expectation of the way the world should be this should run at this time followed by something else. It should be of you should test the data to see if it's right and you should be able to tell if it's late and be able to judge the expectations of the world versus the reality of the world.
And then get data about it and be able sorry get data about it to be able to see what's going on and build sort of events and alerts on top of it. And so the idea here is that once you build a data analytics system. Don't trust that it's right and I talk to a lot of people they build something and they Sort of hope the data works. Oh, it's I've delivered it. I checked it and you know uat I have no idea is the data good did it change suddenly and so you need this idea of automated data checks or data quality checks. And so it's not that I don't trust the people who provide data. It's that I don't trust them at all. I've been Sort of screwed over so many times by data providers who forget that you exist with all good intentions and how do you know and it's not like that's going to change and so you need to sort of protect yourself and think about testing in two different ways think about testing the data and the things that are happening to data during production.
And then think about the opposite of that checking from the this code that is acting upon data in development. And I think those are both actually related and so because what most organizations want is to have a lights out operation data is coming in insights going out. I can focus my time on creating new value instead of tending and fixing errors in the past and a way to judge that and my belief is that you should work in an agile manner deliver small bits of insight quickly.
And if you do that small bits of insight quickly where you've automated the world around it you build an observation system on top of it. Then you get more time to deliver value. You can find what customers really want and your productivity goes up and your customer success goes up. So that's you know, if anything I've been talking about this idea that there's an operational world.
That you live in that it's worthwhile to invest in to fix in to automate this idea of automation testing observability. it went like a storm through the software world in the last 10 or 15 years. The average software development team is 28% of their staff devoted to what's called DevOps, which is observability testing deployment Automation. And I think we should do the same and I think we should start with this idea of observability because in a lot of cases this idea of testing is not unique there's we always have check data Road count checks simple checks, and so there's just many sources of checks.
And so from that from that standpoint is
00:20:00
just to do a quick summary. We do have a product that helps this it's called that were in development right now to sort of Build That central view of cross data operations and be able to see every Pipeline and every organization and drill down into its status and then build this idea of the observability pipeline.
That's a representation of the complexity of your world and You know, we've got some other software don't want to go into that. But that's that's it for my my discussion today. And so are there how much time do I have? Okay, I'm getting quite eight more minutes. Any questions for you?
So how many of you have had problems with? production falling down and had to grab a whole bunch of people and so How did you fix it? How did you go about addressing that challenge?
Sorry, could you please so if you've had a production analytics system fall over? How did you go about finding it? And then how'd you go about addressing it? So actually, you know, what for fertility Investments and I manage Advanced State analytics technology team and we have like data quality is a big like, you know, they can like, you know keeps up at night actually like, you know, so Once we identify actually we have a dashboard we use like, you know, the timeliness company completeness accuracy and we run the python programs nightly on the transaction system to make sure that historically verify, you know, like the data now and run analytics now it fails like, you know now identifying just we identify that we have to dig deep into where it failed the lineage and across the data where it failed that's a manual work for now for us. So that's trans your question, but I have a question as well. Have you seen a successful or organization where they have data quality?
and how they are set up actually, you know, so Yeah, I think. I've seen more unsuccessful organizations with data quality because there's there's two aspects I think to the term data quality one is the sort of passive measurement of think of it as the balance sheet and financial terms of your data. Is it good? Is it bad the sort of Dame a dimensions of data quality?
I think that's a wonderful stuff just like a balance sheet is a wonderful way to to judge your company, but I run a company. I'm actually more interested in the cash flow of my company, right are we going to have enough money a year from now to stay in business and enough money next month that the pay employees and so we do by the way, so don't worry about that. But I think of that as a more active more time-based way that's an active measurement of where you are day in and day out and I think that's something that is lacking and I do see organizations making changes on that implementing observability implementing automated data tests. Just like you when the data lands every night you want to know if it's still good right and you don't want to once a year.
Do the balance sheet of your data? It has to tell you because you have to identify if the data is wrong or if someone the report is wrong or if someone did a transformation of the data is wrong because your customer. In my experience, your customer does not care. You are a magicians. Right and The Magicians stopped giving magic correctly and I am mad and I'm going to go stop my feet until it's right and you can either put your head in the sand and ignore that or you can try to do something about it and do something about it in a way that Drives scalability and drives automation so you don't have to worry about it again. And so I think the best answer to these challenges of production problems is to say to your customer. Yep. You are right. Yep. We were wrong and we're going to put a check or we're going to put some automation. We're going to put that into our observability system and say it'll never happen again.
And I don't know I found that to be a much more effective way because you know data is combinatorially complicated and so you can't check everything. I know Anthony, were you gonna bring something up? Oh.
That I have a question around testing. I saw. On one of the slides you had like performance testing and and all the good stuff. How do you manage your Enable that without a particularly for for using production data or production like data, right? How do you do that? Do you rely on it providing a copy that's or obfuscated or do you do that yourself? And
00:25:00
can you explain? Yeah, I think so the question is in development, you're going to check data for performance test. You're going to check your analytic system to see if it still working and let me back up and tell a story as a preface to that. So I have a 22 year old son who just got a job as a software developer. I'm very proud of them at Amazon.
His room was a disaster growing up sort of covered in Legos when he was a teenager. It was just scary like it's scary. He's an engineer and so he took his job at Amazon and he put code into production in Amazon's back in system within two for two weeks. Now, it was a lot of code. It was like five lines of code, but he did it and so why does Amazon trust my 22 year old son with the in my mind is still eight with a room covered in Legos like Is that well because they built a system around him.
To be able to tell if my son made five lines of code. What's the impact of that change on everything else? And so I think what the What is the characteristic of those kind of systems and data and analytics? The first one is you need good solid test data that is like production data, but it's not.
Some organizations you can you know right now we can throw 100 terabyte databases around zero copy clones and Snowflake very easy. Some organizations have rules privacy rules limitations and you need a privacy aware data. So in some companies you're dealing with a four petabyte database. Maybe you don't want to copy and paste it. Right but the idea is can you take a copy of your data and analytics system run it against that and have a tell the 22 year old organization that this change worked and this change in the database table affected the reports or the models because the impact of an individual's change on a great system is what we're not building.
And I think that's we should Benchmark the success of all our teams to be able to have messy room 22 year olds make a change within a month of their first job. And and know that it's going to work not hope that it's going to work. And if you can do that, then it's much easier to make changes. It's much easier to notice if the changes have causes problems. So it's a long-winded answer to a simple question around having good test data is important.
And using that test data to validate the impact of your change whether you're 22 or 62 you should be able to do it and there's too many of us smart people who have to bless changes, you know, the the bottleneck person who has the whole system in their head. They're a problem. Because every organization has a man or a woman who's like the smart person and like if you're gonna make a major change, you know talk to talk to Karen. She's got it all she'll tell you if it works. You don't want to have that because those persons honestly, they get bored and they leave and it's your best people don't want to be answering questions.
They want to be creating.
So some curious because I love the idea of the of the mission control concept. But the thing in my head is like how do you get the situational awareness from some of the systems and applications that aren't necessarily well set up to share that and are kind of reluctant to give you the keys to understand everything that's going on. They think it's some sort of competitive advantage to keep the information from you.
Yeah, I I don't think I think there's two levels of information. This is it you could argue that the information is what is a class of metadata right? It's operational metadata. Is it running has it stopped is the data good. It's not actually the core Jewels because you don't need to pass all your data into the system to make it work. And so I think that's that's part of it is that the instrumentation of the core systems is important and then the other part is the value of of it proving to the people that if you instrument your system you build these tests you will have more time to do good things. Your customers will be off your back and you can iterate quicker. And so I think it's the it's the trade-off in any organization because if you what we're asking people to do is to focus on the system next to the work they're doing And if they focus on that, they're going to get better value and that's that sort of 28% of all people in software development teams now are in DevOps and they're actually more expensive than your average software engineer. And so there's a career path here that I think for people in this sort of DataOps or operation side.
Sure.
In some respects this what you're talking about is a political question not a technical question, right? So, is there a model that you've seen that would allow there to be Central and decentralized views of these things so that we can see, you know, centrally there's lineage but then you can also kind of deploy a local version where
00:30:00
the local people could derive value from that yeah. I think that's I think you have to do that right the Hub and spoke model the sort of centralization versus Freedom States versus federal government rights. That's a characteristic that happens in a lot of organizations and the value chain that delivers analytics and a lot of big companies is no longer centralized companies, like click self-service tools self-service bi self-service data science self-service data engineering as well as the central it function have to work together and I think one of the indications that things aren't going well is is the Hatfields and McCoy. I put the data right? You got it wrong. No, your data is wrong. I got it, right and how do you resolve that? Well, I think having a Central View So can people can look at the same picture not of the 360 degree review your data.
I'm saying the 360 degree of your operations. So they can actually see the journey if you can't see it. You don't believe it and how many organizations have you run into that where the central it the data providers are stuck because they don't know how many reports what's used. They change a schema. They alter a data set and Hell Breaks Loose. It's never happened. Okay.
Zero minutes. Okay, so I think we're done. Thank you so much for the questions. Very good questions, and I appreciate the time.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.