On-Demand Webinar · 1 hr 11 min
Everything You Need to Know About DataOps Solutions
Wayne Eckerson, President of Eckerson Group, and Chris Bergh walk through Eckerson's DataOps Processes and Technologies framework: the capabilities a DataOps program needs, and how to compare vendors that rebranded into the category. Recorded September 2021; updated August 2026.
What you'll learn 6 points
- Wayne Eckerson lists ten symptoms that a team needs DataOps: the data team is flooded with request tickets and burning out, business users do not trust the data, source system changes keep breaking ETL jobs, business users are used to debug data quality issues, analysts recreate existing pipelines with minor variations, SLAs are missed, data scientists wait months for data and compute, the cloud migration is stuck, data fires crowd out predictive analytics, and deploying a single predictive model takes months.
- Eckerson sorts DataOps vendors into four groups: pureplay DataOps, with DataKitchen as the example, pureplay DataOps+ with DataOps Live, data pipeline platforms with Zaloni, and purpose-built tools with Unravel.
- DataOps tools are a separate market from the roughly 100 billion dollar data tools market. Data tools focus on the data and the insight: ETL such as Informatica, BI such as Tableau, analytic databases such as Snowflake, data science such as DataRobot, and governance such as Collibra. DataOps tools focus on process and workflow and act on those tools. Some data tool vendors use the word for a DataOps halo, which is marketing rather than product.
- Gartner found in 2020 that only 22 percent of a data team's time goes to innovation and 78 percent to errors and manual execution. A 2021 DataKitchen and data.world survey found 52 percent of data engineers said errors were a major source of burnout.
- Intel's enterprise analytics engineering manager Gregorio Martinez describes the practice concretely: reusable design patterns so a pipeline can be built quickly and changed reliably, metadata-driven automation so engineers are not adjusting data models every time a schema changes, lean techniques to find and remove bottlenecks, and more than 1,000 tests in the test automation framework.
- Teams adopt DataOps in four stages, each with a different person in mind. Production DataOps lowers error rates through automated testing and observability, for the production engineer. Development DataOps raises cycle time, productivity and collaboration, for the DataOps engineer and the data engineers and scientists. Measure DataOps adds process measurement, for the data team director. Transform with DataOps addresses the whole organization, for the chief data officer.
Slides
Transcript
Show chapters and dialogue 11,489 words
00:00:00
Awesome. Okay. So, I know DataOps is a difficult topic to get your mind around and figure out where to start and what the solutions are, and hopefully we'll address that today. But I thought a concrete way to really go about figuring out what DataOps can be used for is to talk about the symptoms that when you need it, right?
So I've got 10 symptoms that if any of these things are going on in your organization, you probably need some form of DataOps or another. And I'm going to charge you and the audience to come up with an 11th or a bonus symptom, because I'm sure you all have things, and you'll get the gist of where we're going with this as we move into it.
And Chris, I'm going to ask you for a bonus symptom as well. So get ready. I like that part too. All right. So number one, your data team is flooded with request tickets, and it's burning out. Basically, can't keep up with the request flow.
Number two, business users don't trust the data you give them. There's always something wrong with it. They just don't have a lot of confidence for one reason or another. And then they roll their eyes at your team. Yeah. And the messenger gets shot, right? Yeah. Source system changes keep breaking your ETL jobs and data pipelines and cause a lot of rework and fire drills and people getting woken up in the middle of the night. I think that's your favorite, right, Chris?
Oh, yeah. This has happened to me more times than I can remember, and I use more colorful language to describe that than source system changes.
Number four, you use business users to debug data quality issues. Yep. That's just like the hope and pray approach to data pipeline development, right? Exactly. We think it's good enough, we'll throw it over the wall and see if it sticks. Number five, data analysts recreate existing data pipelines with minor variations. In other words, no one's working and collaborating together.
No one sees what other people are doing. Everyone's creating their own patterns for ingestion and transformation and delivery. Just heard about this this morning at a major chemical company. Everyone's in full hero mode. Full hero, yeah. Yeah, everyone's saving the day. Number six, you have difficulty meeting service level agreements. That's kind of like number one in some ways.
Number seven, data scientists, if you have them, wait for months for data and computing resources. We have that with one client right now for sure. Number eight, your company can't figure out how to migrate to the cloud. It's just too complex to figure out how to move all the pieces and parts up there because you don't have a good control over what you do have.
Number nine, you are too concerned with data fires to implement predictive analytics. In other words, the blocking and tackling is consuming all of you and your team's time, and you have no time for value-added activities. And number 10, it takes months to deploy a single predictive model. And again, that's kind of like nine, you just don't have time to do it.
You don't have good control of your pipelines. You can't stabilize those to deliver a predictive model in production for a real-time application. All right, those are 10 symptoms that I actually came up with from our research and from consulting clients. So you in the audience, I'd love you to submit your own through, I guess, the chat window, and maybe Beth can monitor that.
Now, Chris, I asked you to come up with a bonus symptom here, so let's start with you. Well, I'm going to do the Hatfields and McCoys, that your central data team has, who's guarding the data resources, is at war with your self-service team, who in some way, shape, or form, they've provided them with data, but the self-service team's making some of these mistakes, and they're getting blamed.
And so it's the Hatfields and McCoys of the data world, the central team versus self-service teams. Yeah, and probably the color commentary is that the self-service team tried to rely on the data team at one point in time and gave up, right? Exactly. And, yeah, the self-service team thinks the IT data team's idiots, and then the IT team thinks the self-service team is a bunch of cowboys- Yeah ... and they're getting the blame, and I've seen this over and over again.
Yeah. So, I think DataOps is a great way to start to address a lot of these problems. In many ways it delivers on the holy grail of what I've been talking about for almost 30 years in this space, of delivering data and analytics solutions faster, better, cheaper. So, and we do have a good model for DataOps- Oh ... and that's the software engineering field.
00:05:00
And essentially DataOps applies the rigor that software engineers found about 10, 20 years ago, when they embraced DevOps, and but it brings this to the execution and development of data pipelines, not software code. Right? So you can see this first arc was DevOps trajectory. They had their manifesto for agile software development published in 2001.
First DevOps event, 2009. Well, DataOps, we're a little bit slower to the game. We're always the caboose picking up the rear, but we're getting there. We had a DataOps manifesto published in 2017 thanks to Chris Bergh here, my fellow speaker, and the first DataOps event was two years ago. So we're catching up, but there is a big difference between DevOps and DataOps, and that fundamentally is obviously the data.
We have to manage more than just code, software code. We have to manage data as well, and data can be very gnarly, squirrely, hard to pin down. I know Chris can tell you more about that, but from my perspective, and Chris, I don't know if you'll agree with all this or not, so I'd love to get your commentary, but to me, the promise of DevOps, or DataOps, excuse me, are these things, right? That we improve the way we develop, and deliver solutions to users. And a lot of that means how we write the code, both the data and the software code, whether it's reports or dashboards or predictive models, or the orchestration of things. So we want to componentize our code and avoid these big monolithic applications that are hard to troubleshoot and fix, right? So the more we can componentize, the easier it is to manage.
We can build tests for those components and have a better chance of catching errors on the input or the output side of that code block. And also, we want to treat data like code, and just as in the software engineering world, we want to put that code in a repository. We want to do the same with data.
And we want to lock it down, so that when we run the pipeline, we can be very sure that the data is going to be accurate and it's going to flow properly into the target system. Also, we want to automate these pipelines that move data from one or more sources to one or more targets, automate that via metadata so it's not manual.
And essentially, we can push a button and run that pipeline and all the tests associated with it, and see not only in development while we're testing it, but also in production, right? So it can be run the same way each and every time. And then finally, when we do automate, we want to monitor continuously using the tests we built in development as well as additional tests or using observability tools to understand any anomalies that might be occurring in the data itself, in the code, or in the systems supporting both of those things.
So to me, those are the fundamental development practices that you want to develop with DataOps. And the results are pretty stunning. It's faster cycle times, fewer data defects, less rework and churn, better night's sleep for your data team, and the ability to focus on more value-added activities. So really it is faster, better, cheaper, and ultimately, you're going to make your customers happier.
So, Chris, do you agree with this promise of DataOps that I'm putting forward here? I do. In some ways, like Maybe this is pardon my French, but automate and test the s**t out of things. And those things are the work that you do, the models, the visualizations, the transformations, the governance. And so DataOps is kind of misnamed.
It's about those processes that act upon data and trying to get an end-to-end view that's automated and tested and that you can trust. And what makes you better and happier is that you don't have to worry about these things. You can focus on what really matters, which is at the end of the day, making your business customer successful and trying to do that.
So let me play devil's advocate, because what you describe sounds like these are the practices you want to put forward for production data pipelines, not necessarily self-service, which is what you mentioned in your bonus symptom there. So how do these things apply to the self-service world, where you really want to
00:10:00
encourage your data analysts and your data scientists to kind of go out and discover things and build things on their own? Do you expect them to go through these processes as well? I think that's the part that people are into dial in. Whether you're using a really nice graphical UI to do a model or a visualization, or you're actually writing it in Python, you're creating code.
And maybe that code is expressed in JSON or XML, or maybe it is in Python. And so you're in the software business and in the code business. And so, what I think is that what makes it harder than software development, I spent 15 years in software development, is in software, you've got one cycle of iterating on the code with your customer to make sure that it's right and meets theirs.
But in data, you're also iterating on the data to see if it actually is predictive, it actually makes sense. And both those cycles of iteration need to be coordinated together in order for you to deliver faster value to your customer. And I think that's what's hard about
data and analytics is that it's data plus code acting upon data. And whether that code is generated through self-service or high code, low code tools, it's just code. So any self-service user who builds something, let's just say they want to build something and share it. That's where self-service at some point needs to be governed.
You would say then that probably needs to be brought into the DataOps sphere and the code for the tool you're using, whether it's Tableau or Power BI, what have you, should be brought in and managed by a DataOps tool. I would think so, yeah. So let's say you're building a Tableau workbook or a Power BI dashboard.
In there, you could make a calculated field that's got an if-then-else statement, and that shows up in the table or the bar chart. Now, you could argue, where should that if-then-else statement live in your analytics system? Should it be derived by a model? Should it be as an attribute of a dimension of a fact table?
Should it be in the loader script that loads the data? I've seen every case where it goes, and you could have technical arguments either way on where it goes, but that if and else is just code and happens to live in different tools and different systems. And so if it's wrong, or if someone copies that workbook and changes it a bit, it's still going to be wrong. And so a lot of inconsistencies and errors are just from the fact that
we don't treat these artifacts as code, as versioned and tested and automated. Yeah. And so DataOps also corrects if you have a new version of Tableau or ETL tool or a new release and the configuration changes and it might affect the output as well, right? So- Exactly right. Yeah, because it's not the code, it's the version of the tool, right? And both those change. And some companies get so scared to change even small incremental versions of their tools.
And another symptom you could say is you've got something in production, but you're fearful of changing anything. Like, "Oh, it took me so long to get this to work and don't ask me to change." And really it's about the acceptance and saying change is inevitable and build for change. Yeah. And try to spend some time building so you can change whatever part of your system as fast as you can possibly can. And if you build that superstructure, that sort of process how to do it, you're going to end up actually giving more business value to your customer, and you focus on the factory and not on the cars that come out of it. And that happened in software, and that also happened in mass manufacturing and lean manufacturing.
Yeah, it's interesting. I think a lot of data managers think, well, if we want to go faster, we just implement self-service and let the users go faster but what they usually find is that creates more problems than it's worth. That can be part of the solution, right? But at the end of the day, you're building things, whether through a nice UI or through a slightly harder, more technical UI, and both those things are acting upon data. And having a system that can industrialize it and automate it and deploy it and test it, I think is what DataOps is about.
And so it's never about replacing the tools you currently have. And so that's one of the things that we're going to talk about later, is that this is fundamentally a system that works with the huge market of tools that exist throughout the data and analytics economy. All right. So let's dive into our framework for DataOps here, which I think can answer a lot of questions, because there's a lot of FUD in the marketplace, thanks to a lot of vendors who have jumped on the DataOps bandwagon because it was a hot
00:15:00
new buzzword, right, a couple of years ago, that Chris and team really helped evangelize and bring to the forefront of people's minds. But when you look at DataOps, it gets confused oftentimes with data management. So this, not the prettiest slide in the world, but you can see the pipeline, which is the big arrow. It's moving source data through an ingestion process, and through a transformation process, and then into an analytics environment, so that business users or applications, which there's an omission there.
Should be applications can consume this data that was put together in this pipeline, right? Now, up until the DataOps movement began a couple of years ago, we really focused on the bottom of this circle here, these data management technologies and processes. And there's a lot here. There's a lot of work to do to make these data management technologies work.
So data capture, data integration, data preparation, data analytics, more end-user type of activities at the end here. All working around a common repository for storing the data, your data warehouse and data lakes, your analytics sandboxes, and then obviously the computing and infrastructure supporting it all.
So that's what you use to build your data pipeline, you use those data management technologies and processes. Well, DataOps sits on top of all that, kind of like a metadata layer, and it helps you build that stuff and do it better, right? And it does it through these four areas. The development, through continuous integration, environment management, which means putting your code and data through dev test and then production, and obviously, hopefully, doing that in a very agile, responsive way, creating sandboxes very quickly for users to go play in and discover stuff.
Deployment means that once we make a change, as Chris said, some teams are reluctant to make changes because they're afraid things are going to break. But in the software world, these
cloud-first companies, these SAS companies are deploying changes dozens of times a day to production environments. So we want to get there, closer to that in the data world, by feeling free to make changes and do it quickly, to meet customer SLAs. Orchestration means that we can move things through the pipelines on a coordinated, scheduled basis, and we've got kind of the metadata layer or the scheduling tools to really make that happen in an automated way.
And then continuous testing and monitoring. The testing, as I mentioned in the symptoms slide Starts in development, obviously, and then you bring those tests that you create in development into production and maybe add some other monitoring capabilities. There's a whole new set of data observability tools that can monitor the data and the infrastructure and alert you to anomalies and problems.
There's a lot of things inside that circle there, agile technologies, code repository, configuration repository, container management software that is part and parcel of a DataOps process as well. So this is the big picture of how DataOps fits into your world. And when we start to talk about products, which we'll do in the next section, it gets a little fuzzy, but we're going to use this framework to help clarify that.
So Chris, any comments on this framework before we move on? Yeah. What's interesting is the bottom half of the circle is centered around data. Data storage, data lab, it's all about data. The top half of the circle is really centered around code and configuration. It's the intellectual property that your team creates, and the work that your team does. And so I think it's interesting because that's more of what DataOps focuses on, code, configuration, et cetera.
Okay. All right. So let's take this framework and show you how different vendors are implementing parts of it, or all of it. So you get pure-play DataOps vendors like DataKitchen, who's hosting this webinar here, and they're going to focus on the top half of the circle. Right? And that half circle as well, that's not shaded in.
But so they're pure DataOps. They're not going to do any of the data management.
00:20:00
They pretty much presume that you have those data management tools already, which most companies do. So just plug these data management tools into your DataOps platform, pull the data, pull the code, into your DataOps tool so that it can manage it, automate it, schedule it, monitor it, help you build tests for it, et cetera. And then run those tests in development in an automated way.
Make sense, Chris? I'm describing your company here. Oh, perfect. Yep. And, yeah, you did a great job. Okay. All right, so there's some other, quote-unquote, pure-play DataOps vendors that I'll call pure-play DataOps plus, like a company like DataOps.live, which just is a new company. It works only with Snowflake right now, and its purpose is to focus on DataOps.
But it also provides some data management capabilities to help you get started if you don't have any, which is unusual, or you just don't want to get tangled up with the data management tools. You just want an all-in-one solution. Now you can see here, it doesn't support the analytics side of things, but that's not unusual as well. So that's the pure-play DataOps plus.
And Chris, you're never going to go with the plus, right? You think there's a big market. Oh, as a CEO, I sometimes wake up in the middle of the night and say, "We should do that." But I don't want to because I've been in the field too long, and people love their tools, and I don't want to get between Informatica or Airflow or ELT vendors or DBT. People love their tools, and so use the best tool that you're most comfortable with.
Yeah. And you work with all of them. Yeah, we work with all of them. Yeah. Okay. All right, so let's keep going now. We've got two more here. So here's another one. We'll call this a data pipeline platform that is giving you an all-in-one data management solution, with or without the analytics. Sometimes we see with or without. Everyone can do a little visualization, so we're starting to see that become part of these platforms.
And to support the data management, they do have built in to their platform some basic DataOps capabilities, stronger in some areas than others, to help automate the platform. Right? So in this case, Solone is pretty strong in orchestrating its own modules to help you build data pipelines. Right? And offers some capabilities for testing and continuous development and deployment. But it's really a data management platform with some DataOps functionality. Maybe not all that you would like.
And we see this a lot. A lot of data management vendors, frankly, jumped on the DataOps platform and really made life confusing for everyone. There's a lot of FUD, and people began to believe that data management was DataOps when it's really not. DataOps is really the metadata overlay on top of everything. Would you agree, Chris?
I wouldn't call DataOps a metadata overlay. I'd agree metadata ops is an overlay, but it's a meta orchestrator. It's meta something, I don't know if it's metadata. Okay. Meta. Right. I got carried away. Just it's a meta overlay. It's a meta layer. Yeah. No, that makes more sense. All right, and the last example here are pure-play vendors. So a company like Unravel focuses on continuous testing or monitoring, which we now call for some reason, I haven't quite figured out, data observability. There's a lot of these vendors coming to the market right now that will continuously monitor your data or your infrastructure, or your applications, or all three, using AI and ML and surface up relevant alerts, so you can better see what's going on.
And in some cases, supplement the test, well, in all cases, supplement the testing you might, the manual handwritten test that you might write. So these are becoming pretty powerful tools, but they are pure-play. They solve one part of the problem, but again, a lot of these companies are billing themselves as DataOps vendors, too, so just to add more confusion to the mix. By the way, all these examples came from our report, Eckerson Group report, called "Deep Dive on DataOps," where we examine these four vendors in depth.
So Chris, you've tracked a lot of the pure-play vendors out there in each of these
00:25:00
areas. Do you have anything you want to add here as a continuous- Well, having gone five years ago from having people look at me like an alien to talk DataOps and writing the Wikipedia article and the manifesto, I'm happy that people are talking about it, right? Because the marketer in me likes that people are talking DataOps because, all in all, we agree that trying to support our customers deliver value faster with lower errors is a good thing. No one actually disagrees on that.
Yeah. Another good thing to talk about. The engineer in me gets upset when people don't fit into nice categories, but people are going to do that. That happens in every tech word. I remember when I'd read articles about Excel being a big data tool, and my head would explode, right? Because big data was parallel machines with dump queries in between. And so it just happens.
But I think it's helpful for you to clarify for people what really is DataOps and what's not, because oftentimes I get that question. Yeah. So I hope that helps. And Beth, if there are any questions out there, let us know. I can't really... Well, maybe I will, because I'm going to go check the questions, but I'm going to pass this over to Chris now because you're going to key off these charts here and come up with your own analysis.
Yeah. So I've got two slides. And so I think the first takeaway I'd like, people, is that DataOps tools are separate from data tools, because they focus on the process and the workflows and not data. They act upon your tools and integrate with those tools, and they drive business outcomes like lowering errors in production, or faster and less risky deploys, or team productivity, and more insight for your customers.
Now, the stuff on the bottom, it's like a $100 billion market. Right? You've got tools that do ETL and ELT, like Informatica and all the other tools out there, 50 versions of it. You've got self-service and BI tools like Tableau. You've got analytic databases like Snowflake. You've got data science and AI tools, DataRobot. You've got governance tools.
And then you could say each one of those, the cloud vendors have a version of that. And then there's subsidiary ones like streaming tools and specific tools. And I think that's great, and as I said, people love their tools, and- I like Redshift over Snowflake. Some people like Snowflake over Redshift. And a lot of it has to do with our experience and our perception from the market.
But one thing I do know is that you should be able to try out different tools and upgrade the versions of your tools a lot easier than you do now. And so having a data op system allows you to do that. And that's just purely a lesson from the software industry, is that your tech will change, and build for change. Data and analytics is not a field of dreams activity. You don't build it and they'll come.
It's about managing change, a constant series of change over time, and tools, and data, and customer requests.
And so if you go to the next slide.
One of the things that we're doing out of our own self-interest is we want to market around this, and it's forming, right? And
aspects of markets forming are companies getting funded, companies open source projects. And so we sort of bucketed in our analysis sort of the capabilities of DataOps into these buckets. And so in this report, we've got different vendors. And actually in the last three or four months, the observability vendors have gotten a lot of funding. Before that, you play back two or three years, some of the data science vendors have gotten.
And then similar, there's a bunch of vendors who talk DataOps, but really they do governance and other activities. And so I think it's more marketing than product. And I run a company, and I can understand people want to get on the Google keyword flow and get noticed for it. But I also do think that these capabilities that are in our list are part of what you need to do with DataOps.
Is that it? That's it. You can go to the next slide. Okay. I was trying to look at the small print there. Small print.
All right, so we talked about pipelines, and we didn't really define it, only to say that it's the process of moving and transforming data from one of our sources to one of our targets. There are really two major types of pipelines.
00:30:00
One is building stuff and moving it from dev to test to production, and then once it's in production, executing it, so execution pipeline. And we really want to deploy DataOps to both of these and do it in a consistent standard way. I think where I see a lot of our clients get into trouble is that they're doing things differently each time they do them, and they don't have any consistent way of ingesting data or moving stuff from development into production.
So we really want to create what are called design patterns. And one of the best DataOps practitioners I've met is Intel, and their former enterprise analytics engineering manager, Greg, who's still at Intel, just in a different role. He had a lot of great quotes. In fact, I have another slide of quotes from Greg because he gets right to the heart of how they use DataOps to improve what they do, increase their capacity and reduce defects, automate stuff, increase cycle times, all that stuff.
What he says here is, "We create reusable design patterns for every step in that pipeline that enable us to create a data pipeline quickly, change it as needed, and maintain reliability and consistency of the data output." So they do ingestion in the same way every time, or they have the same sorts of tools.
If they're going to do change data capture or streaming, they have a consistent set of tools and process by which to do those on the ingestion side. Same for transformation, and so on and so forth.
Also from Greg, about four more quotes here. All of them are fantastic. Quote, "We apply lean techniques to measure bottlenecks. Then we adjust our processes to remove constraints and minimize or eliminate waste." And this was a statement he gave me in response to, "Well, where do you start?" And he said, "Well, we get our team together, and we just list all the bottlenecks to getting things done that we encounter, and then we prioritize those bottlenecks, and then we come up with ways to unravel those bottlenecks, whether it's new tools, new processes, scripts that we write." And they wrote a whole bunch of scripts themselves to automate stuff like handle changes to source schema and do it automatically so that if the source adds a new column, they can automatically detect that and then add that column to the target database.
And there's no manual intervention at all, no fire drills in the middle of the night. Nothing breaks. It's all automated. So that's the kind of bottleneck that they say, "Hey, we got to fix this. We don't want to be always manually reworking things when the source systems change." So they figured out a way to do it. That's DataOps.
Here's another great one. "We use metadata to automate our pipelines." We talked about that earlier. "We want to focus our engineering resources on innovation rather than rudimentary tasks such as adjusting data models and transforms every time there is a schema change." Well, actually, I just mentioned that, so I was getting ahead of myself. Here's an important one.
"Test automation is a huge part of what we do. Without it, we can't maintain a high level of quality, scale, and speed with which we operate. We have more than 1,000 tests in our test automation framework, and we keep adding tests all the time. We are only as good as our test practices, and we strive to improve here." This might be the most important part of DataOps, is adding tests.
And it helps if your code is componentized, right? So you can have a couple of tests on the input, a couple of tests on the output, and you run them all the time. Yeah. And what's interesting is the dual nature, right? Because when you're in production, your data's varying, but your code that's acting upon the data is fixed, and what your tests are doing is really checking the data.
In development, you want those tests to tell you if you've had any regressions or any sort of-- When you change something, you want to know if something else has broken. And in fact, what you need to do is run those tests again in development, because really what testing means in terms of a data and analytics system is you're always testing data. And so what I see a lot of companies do is they have manual checks, or they give it an eye up, and that's their way of getting things into production. But you should really be able to prove whenever you make a change and pull the pin forward so that one developer, no matter where they are in the full end-to-end value chain of analytics, they make a tweak, they can see the effect of that tweak on the whole system.
And I think that's really powerful for people, and it sort of pulls the pin forward
00:35:00
to the individual contributor. But on the other hand, it empowers them. And what you want to be able to do is, here's a rule of thumb. You hire a 22-year-old. Can they make a change in any part of your data and analytics system, a really simple change, and know if there's a problem? And if they make that change, you know it's going to work.
And right now, that's pretty rare, actually, in a lot of teams, because you've got to go through, if you're having meetings and checks, and you've got the smartest person who's got the whole system in their head, who's got to go look it over, change review boards, you know that you haven't automated this up thoroughly.
Yeah. And developers push back initially against DataOps because it is another hoop they have to step through. They don't have to go through a review board, but they do have to put their code into the DataOps framework. So they don't want to do that. It kind of feels very constraining. But once they see that they don't get woken up at night to fix something that broke or what they changed didn't break anything else, it gives them a huge amount of confidence, and they embrace it. That's what I've discovered.
Yeah. And as my co-founder likes to say, tests are the gift you give to your future self.
I like that. All right, one more quote here. "DataOps enables us to create a lights-out data processing environment that minimizes waste in our data centers and fosters a culture of continuous improvement." Actually, one of the best things that Intel did, the most astounding thing they did Was, in the agile methodology, you spend a half a day out of a two-week sprint doing a retrospective and reviewing what you've done.
Intel, and at this time, Greg was in charge of the sales and marketing agile teams, and they had 13 agile teams supporting sales and marketing functions at Intel. They would spend one sprint out of every four sprints doing a retrospective. That takes courage. That's two weeks where you're just sitting around trying to think about how to make things better, how to automate things.
But it gave them the time to write the scripts to do schema evolution. Right? And if you take one step backwards, they're able to go two steps forward. And this is how DataOps works. And in fact, it's really a practice. It's a mindset of continuous improvement, but you need to make time to do that.
All right. So Chris, I'm going to hand this over to you. You got your own best practice slides, so go at it here. Yeah. And so part of this is to help you understand what DataOps solutions are. And so what I want to do is kind of go through that in detail. And the way that I want to talk about it is, how we recommend our customers adopt the ideas in DataOps. And the first one is really about what happens during day-to-day production. And what we want people to do is, our high level is if you're going to do DataOps, apply DataOps principles to your adoption of DataOps. So find things that can give value quickly, with a small amount of effort, and then work from there to take that evidence and proof points forward.
And so, focusing just on, man, just get errors down in production. Check the data as it's being processed, observe it, and that's your first step. And then the second is you've got this whole process of making changes to existing pipelines. You have a development process. And how do you test it? How do you do regressions? How fast can you move it through?
And we'll talk about that and talk about the vendors that help. And then the third, we're going to talk a little bit about measurement of those processes. And it's just always surprising to me that data and analytic teams are unbelievably unanalytic about what's happening in their organization. Like, they can't tell how many errors they have.
They don't measure the productivity of their staff, or they don't measure their cycle time. And so those process measurements can both help validate your investment in DataOps, but they also can drive behavior changes in people. And so let's go to the next slide, and then I'll talk about it. The first is a story. Now, every data and analytics team's got an Eric, right?
They're the production perfectionist. I
just talked to a company this morning, they got 400 pipelines in production. If Eric's in charge of them, he's got a daily grind. He wants to minimize errors and chaos. And he wants to know across what's happening on all those pipelines, and he wants to know what's wrong before his customer sees it. And there's this whole market emerging, and there's a bunch of vendors on the
00:40:00
right,
some of which are new, and some of which are, like us, have been around for a long time. And so it's an area that the venture community sort of woke up to about 18 months ago. And so you're going to hear more about it because these companies have gotten a boatload of money. And so, for instance, Exceldata just got $35 million in a Series B funding today, I saw it on my tracker.
So let's go into what this means, if we go to the next slide.
And so, for us, it really comes down to a set of best practices. And on those, the best practice-- I actually, Wayne, would you mind if I presented, because I added a few more slides and could I just switch them? We only have 15 minutes left in the whole event. Oh, yeah. You're right. So, okay, I'll go through here.
So we want to get the questions, so I won't add more slides then, you're right. So if we think about it, one is that there's a journey that your data takes from its source system through all those tools that you have, data transformation, integration, visualization, modeling, governance, to value to get to your customer. And what people like Eric want, is they want to know what's happening on that journey. And automatically, they want to get an alert saying something's wrong or something's weird.
And they want to know patterns over time. If this pipeline or this data source or this tool's been problematic, they want to know it, so keeping track of history. And one of the things that we believe is that people love their tools, and they want to be able to use their tools to create tests.
And our argument is that 10 or 15% of the time people should be using is to do this automation and testing. And maybe that could be a separate DataOps engineer, but also the people who are doing the work should do it. And the ways to test data in production are interesting. And so there are some ways, and we've got tools to take data and auto-profile it and say, "Okay, here's the rows and columns in this data, and this third column has got three values," or, "This third column looks like a zip code, so we should apply zip code tests to it." And that sort of traditional data quality checks that our software does and other software does.
And then if you think of this step-by-step process, it actually reminds me and what I learned 15 years ago is that I was running a factory. And those pipelines in production, one of the ways that automobile manufacturers and other manufacturers have been successful is to do something called statistical process control. And that's if you get a million rows from a vendor one day and a million rows next day, if you suddenly get 300,000, you should probably know about it. And then there's other tests that we call location balance, and I think a bunch of tests should also look at the business context.
So it's not just testing the data in its rawest form, it's actually testing the integrated data, and it's actually testing the artifacts that come from the data, like we talked about the Tableau workbook and its if/then/else statement or the model. And so why bother to do all this? Well, less errors in production mean more innovation. And for us, errors are a superset idea.
It could come from the source data quality problem, or it could come from the server's down, or you're late, or somebody put code into production that broke things. And so we see errors as the superset of trying to automate data quality in production. And if you drive less errors, you get more innovation and you get more customer data trust. And honestly, when your system is down, when you have a data problem, just consider the whole data analytic system down. It's a failure.
And you want to decrease that time to failure. And last but not least, it's just less stressful. And I can't tell you how many days, before I learned of DataOps, I'd go into work and just dread the morning dread of, "Oh, it's wrong" or dread the phone call from the sponsor saying, "If you don't fix this, you're fired." And all these things.
I never really liked that stress. And so, part of that reason is if you start implementing these things, you'll have much less stress. Hey, Chris, question for you. Yeah. On the types of tests in the middle there, are those all tests that people would manually create? Yeah, I think some of them can be automated, but by and large, it's sort of a myth that you can do this automatically because you have- That was my next question, because a lot of data observability vendors would say, "Yeah, we do all those tests automatically." Yeah, there's no magic beans, unfortunately, because the difference in data is so great and,
for instance, if you're understanding pharmaceutical data, you have a number of physicians on your target list. That's a number.
00:45:00
And you've got to test every time to see if you've added more on your target list. And that's a number that you should test, and that's a test you should have that's very specific to the type of data, the way that the data's grouped, and no tool's going to automatically create that. Just likewise, there's no automatic software test tool that creates tests for software. And so companies have been able to achieve cycle times of once a second to deploy to production in systems that are as complicated or more complicated than the average data and analytics pipeline without sort of automatically tracing tests. So yeah, and we've got some work that you can help build tests and basic tests automatically, but testing is part of what someone should do because it's a way to look at-- It's very dependent upon the data and the domain.
So I'm not a big believer. It can help, but it's sort of the magic beans that some vendors are selling. So the data observability tools would supplement your handwritten tests, I would say, right? Well, I think my understanding is, and ours like it, we can automatically create tests and we can handwrite them. And I think you've got to do both, and I think it's up to the developer who knows their particular code, and their particular work to write the tests.
Or they could have a data- Right ... ops engineer who'd work with them. Yeah. I've had clients say the developer shouldn't write the code because they're too close to it, and they can't see it, and they're not really vested in it. Yeah. I hear you. I think it's both. I've been in this software and data industry for a while, and I think both work.
I think you've got to have the data engineer and the data scientist write tests, as well as have people outside of it, and it depends on how you want to run your organization. And there's the spirit of Agile, where you build it, you run it, and having small empowered teams do the work. And if you're on that, then by and large, you're going to have to write your test because you build it, you run it, and you're going to see the results of not doing it. If you're in an organization where you throw it over the wall and there's an operations team, and I think that's more of a valid point is that they're probably not going to do it.
But we want to pull the pain forward and have people see the results and stop having their blinders on, and I just wrote the code and my job is done. Yeah. So you would have Agile teams creating tests, but you might have a QA team too to write tests as well, especially for more complicated pipelines.
During our own software development, we do the same thing. Yeah. We have a test automation team, and everyone writes tests. Yeah. Cool. So if you go to the next slide, and the last few minutes here. So let's just talk about a second group, and imagine there's Priya, who uses all those tools, uses Dataiku, or DataRobot, or Informatica, and what's her goal?
And you could call her a data scientist, or a data engineer, or a BI person. She wants to do cool stuff with data. And here's another role, and I call him Chris, for obvious reasons. He's the operations optimizer. His goal is to help Priya do more cool stuff and take what Priya creates and put it in a basket that allows fast cycle time and low error rates. And if you look at the vendor landscape there, there's tools that do sort of pipelining and tools that do environment management and infrastructure as code and tools that actually store source code, and those things are fitting in the development DataOps landscape.
And if you look at the next slide in terms of best practices there, and it's really about what you want is... So if you think of it this way, Priya makes some work in her favorite tool, Dataiku or Talend or whatever, and you want to be able to not have her walk it from her development environment in production. And the first things people are starting to do is make railway tracks.
So like, "Oh, we don't have to walk it. We can move it fast." But I've seen that doesn't work, and I've seen plenty of organizations still take 10, 12 weeks to get something. And you say, "Well, why can't you just put it on the railway tracks and send it down from dev to prod?" Well, because you need the signaling infrastructure on those. You've got to make sure that things work, that you don't have any regressions. And here's where testing and development, and that's where this idea of pulling the pain forward, because the cost of fixing a problem in development is so much cheaper than in a lower environment, in dev or even in production.
And so, being able to have everyone see the effect and kind of collaborate on the shared concept of the end-to-end pipeline, and this idea of, I'm seeing a number of companies say, "Okay, I've got my railway tracks, and we do a couple of code-based unit tests, and oh, we keep finding problems in UAT and
00:50:00
QA environments. Got to go back to development. The developer hasn't worked on it in two weeks. They've got to reset their context, and they've got to go back." And so fixing problems, knowing problems as early as possible, is what matters. And really remember, you're testing data to test your code. And the benefit here is you get less regressions.
You actually means faster deployment, and you kind of allow people to remove their individual blinders, where I make a change in a big system, I don't know if it works, I'm just going to toss it over to someone else. And fundamentally, it's about velocity and risk. You want to lower the risk to change and have the velocity.
You need to have railway tracks and the signaling infrastructure on those tracks to make the movement from dev to prod fast. And lastly, I've just had too many eye rolls, like, you do good work, and then the customer asks 10 follow-up questions, and you're like, "Yeah, that'll be three months from now." And then they roll their eyes at you, and I want that to stop.
I want you to be able to say, "Yes, we can do the first one tomorrow and the next one after that," and what's your priority, and be able to live with the river of follow-up questions that come when you're doing analytics and be able to get them. When you have a really great concept and they say, "Fantastic, can you put that in production for my 3,000 sales reps?" You're able to say yes in a reasonable time, and not, "Oh, you know what? We have to recode it, and that's going to go into our SLDC, and that's going to be six months later."
Hey, Chris, just back on this slide. Are you saying that Chris, Priya, she's the self-service user, basically. Are you saying that Chris over here is actually going to take her work and put it in the DataOps framework for her, or just make it easier for her to do that herself? Well, I'm using Priya as kind of a general archetype, and she could be self-service, or she could be a data scientist or a data engineer who writes SQL.
And I think their goal is to create insight, and the way that they do that insight is expressed in code. And so someone's got to help take that code and help Priya put it in a system that allows her to be able to see if there's problems in development, to be able to deploy faster, both in terms of railway tracks and signals.
Yeah. And so I think Priya's trying to make the VP of marketing or the VP of sales a success. Chris is trying to make Priya a success. I guess my question, though, is, I can see a data engineer putting the code into the data op system, but would you expect a data scientist or data analyst, a Tableau developer to do that?
Honestly, I haven't seen a lot of uptake
in that concept because there's a lot of self-service users who are just so embedded with heroism and hope that things work, and they just change things. They throw it up. But almost everyone I've ever talked to, once they have something running for two weeks or three weeks, they don't want to deal with all the operations, and they're looking for help. And so having someone to help optimize their operations and take that operation burden from them so they can focus is something that they welcome.
So- Is it self-service? Is it helped? That's really the challenge is trying to get people to believe that they need an infrastructure around it. Right. So they're self-service to produce one thing one time. Ad hoc, here's a question, here's your answer. It's another thing to take the answer and productionize it so it runs every week or every day. So you're saying that once you productionize it, then you want to throw it over to Chris over here, to help you put it into the DataOps infrastructure and- Yeah. That's right. But maybe you're a little more incented, and you can use a tool like DataKitchen to do it, or maybe Chris can help you. Either one works.
Yeah. Okay. Hey, Chris, just a time check. We have just under five minutes left, so if you want to- Do we have a lot of questions, or? There's definitely some questions. So-
Wayne, can we go a little longer? Or I know we're a talker. Should I end it right here? Wayne? I could go longer, but let's take some questions if there's some good ones. Okay. So let's get through number two, and then I want to go to the third part, Beth, and then I'll finish. So I'll finish in one minute.
Okay. So Wayne, if you could go to the next part, next slide. The last one is,
I think both Wayne and I talked about the challenges in people adopting DataOps and the belief that you can run production with low errors and change production really quickly. And so part of it is you need to measure the work that you do, and measure the errors in production and measure your cycle time.
And that also helps the team who is running data engineers and data scientists
00:55:00
prove their worth, as well as help adapt people to this world of DataOps. And so if you go to the next slide, we talk about the ability to measure your analytic processes, and that means error rates and cycle times, and data freshness reports and data arrival reports, as well as just trying to look at, well, how much work are people doing?
How fast are they deploying? How many tests are in production? And so these sort of process metrics can be very helpful for teams to understand their system and also to adapt the changes. And so if we go to the next slide
The last is really just helping teams transform, and there's a lot of consulting vendors, including DataKitchen, who are out there hanging their shingles on how do you get an organization to adopt DataOps and do DataOps transformation. And if you go to the next slide, we actually just wrote a book on it. We have a maturity model that Wayne helped us put on our website, and other ways to talk about how organizations to get value and to bring DataOps to a large company or a smaller team. And the next slide.
That's it. Yeah. I think we've talked about all the things, where to start. We've talked about educate your team, you're doing that now. Identify bottlenecks, figure out how to address those. Spend a lot of time thinking about how to improve your bottlenecks, and then monitor your improvement with the tools that Chris just told you.
Componentize your code, build tests for every code block, treat data as code, and invest in DataOps tools.
And if you do that, DataOps puts your data on a solid foundation, does all the stuff we said it can do, speed cycle time, improve quality, increases the capacity of your team to do more. That's what we learned from Intel. They're always striving to make their teams increase their capacity so they could do more,
and reduce cost. Focus on value-add activity and make your customers happy. That's the promise. Faster, better, cheaper.
So Beth, should we take a- Yeah. Thank you both. Yeah. So it's the top of the hour. We are starting to lose some attendance, but we'll try to get through a couple questions here. So here's a good one that I think comes up frequently, "How do you separate DataOps tooling versus enterprise scheduling tools like Control-M?" Wayne, do you want to take that one?
Yeah. Chris probably would be better, but it's the same ilk, right? Scheduling, orchestration, to me, there's a lot of overlap. I think the world now knows about orchestration. That's how they call it. Airflow is one of the more popular open source orchestration tools out there. It's really just automating the movement of data through a pipeline.
But Chris, what am I missing there? You go back, you probably use Control-M. Yeah. A lot of companies, when they have their production work, it's not just data production work, it's other IT tasks right that need to get done, and it's like, "Oh, we need to get a server up and running." And so these general IT orchestration tools they want to standardize on, and I think that's fine.
They have all your schedules in one place. And so software like ours integrates with that. What those tools don't do is actually what we call meta orchestrate all the data and analytic systems, and oftentimes they don't have a test framework. And so Control-M, and other tools like that have a limitation because what you're really trying to do is get a plug-in Airflow and Informatica and all the other tools into a platform to give you that visibility and control.
And who calls the schedule, I just don't think is-- we have a scheduler. I just don't think that's actually the most important part. It really becomes a part of how your enterprise actually deploys activity. No, I- And so they're compatible, they're not competitive. And it's interesting, Chris, I never really thought about that in an orchestration tool, you want to orchestrate conducting the tests or executing the tests. That's probably a big part of the orchestration tool.
Yeah, it has to be test informed because if you're halfway through your production process and you have a huge error, you probably shouldn't go on, right? Why finish? And an error could be, remember, we're not just checking server errors, we're checking that, hey, the data didn't arrive, or the data is so wrong or in such a bad state that going on-- So test informed orchestration, I think, is a really
01:00:00
important way to gain efficiencies. And so you want to find out problems as soon as you possibly can because it gives you time to correct them, and you don't want to have to wait till your 10-hour job is done or your two-hour job is done. You want to know it as soon as possible.
So stupid question, does Airflow support test automation?
It has some plug-ins that can help you. So it doesn't do it natively. But a tool like DataKitchen, you're set up to do not only the schedule and execute the different processes and execute, open the tools and run the tools, run the code and the tools, but also to open the tests, run the tests, look at the results, and then make a decision to stop or keep going.
Exactly, yeah. And send the alerts, and so- Release ... functionally, you need to do all those things. And as an engineer, yeah, you've got a Python, you're a Python programmer, so you can do whatever the heck you want in Airflow, right? And I think it's having that control plane across self-service, backend, Airflow, or whatever tools that you have, makes sense. But companies have centered on Airflow as their enterprise orchestrator or Control-M as their enterprise orchestrator, so working with those, because data and analytics is part of a greater IT ecosystem of things that are running, updating your website, running batch jobs like payroll and operations does all those, and they don't want you to be different.
Right. Yeah, but you'd have to program a lot of the analytics specific stuff into them, right? With scripts. Yeah, you would. You would. But scheduling is pretty standard across them, and so what happens is people will call DataKitchen from Control-M as their enterprise scheduler and say, "Okay, we've got to run this job. Call DataKitchen.
They'll do the work." Nice. Okay. Right. Great. Thank you both for that one. So how does DataOps fit into data sharing like with Snowflake or data virtualization tools? Who wants to take that one? I- Well, those would fit on the bottom half of that framework circle, right? So data management tools. Yeah, exactly. Virtualizing data or doing sort of sharing data, that's data stuff that fits on the bottom of the circle.
And what DataOps has helped change that, if you're going to add data to share, or you're going to change the semantic layer on data virtualization, and you're going to deploy a new update to the semantic layer, or deploy a new update to the data catalog and the semantic layer. How do you do that?
How do you do that in conjunction with the new data set that's being deployed or the new schema that's being deployed? That's where DataOps can help.
Okay, great. So, moving on to the next question. So data catalog/lineage is also an IP asset. I'd argue it's the family data jewels. What's your opinion? Wayne, do you want to start with that one? Yeah. Data catalog, data lineage, it's kind of the centerpiece of a data governance strategy, and like DataOps, data governance is kind of a meta layer that sits on top of your data management stuff.
And the catalog basically catalogs all your data and queries and reports and schema and additional stuff as time goes on, so people can find it, annotate it, collaborate around it, document their tribal knowledge. So that's important, and part of that tribal knowledge is lineage, so understanding where this element came from and what it impacts if you changed it. So
that's important for data governance, it's important for data discovery, it's important for self-service. From a DataOps perspective, I'm going to pass that over to Chris. I assume you would manage that the way you- Well, I know. Yeah. I think the change of the catalog, but also, yeah, having a catalog and lineage is really important.
But those are sort of what is the data and where did it come from? Those are really important questions, and there's two other questions that analysts want the answer to is, is it fresh and can I trust it? Which I think are equally as important. And lineage and a catalog don't answer that. And so what DataOps provide is the answer to that. Hey, it's fresh.
And A, we've run 3,000 tests on it, so you can trust it. And here you can look at it. And so, data lineage and catalog I think fit hand in hand with what I think of as process lineage, and helping people understand freshness and trust is just as important. And that's why Collibra now has a data observability component in their product,
01:05:00
because they see that sort of 360-degree view of your crown jewels extends beyond lineage and catalogs. Right. Okay, great. Thanks. So GUI-based tools generally leave a lot to be desired in my experience. They tend to not be scalable and maintainable or lack options. What are your thoughts on that?
Oh, I don't want to get into that war. It's like, do you use a GUI or do you write code? I'm an engineer and I like code, but I also use GUI-based tools, so there is a real divide in the world, and I think the data and analytics market is really sort of breaking up into low-code tools and high-code tools. And I'm going to leave the one to Wayne to answer that one.
You need both. You need the low-code tools that have an exit for the high code. So, the market cycles back and forth, and BI used to be all low-code, also used to be all GUI-based, and the benefit of that is that it brings more people into the process. It makes it easier to hire folks, train them.
It also helps document the metadata better. So there's a lot of benefits to the GUI-based development tools, but now we're in this cycle where people are moving to Python and using Spark to develop all their pipelines and transformation code, and there's more power in that, obviously. But you have to be trained in Python and what do you do with your metadata?
You can easily bury rules into your code, and when you leave, no one knows what the hell you did. So we'll come back out of the high code, into low code, and I think we're seeing that shift already. Eventually the market will, I think, settle on low-code tools, but ideally they have exits for people who want to write code and see the code, and I see that in a lot of tools now. They give you both.
Okay, great. We have a couple specific questions on tests, but I think I'm going to save those, and we'll follow up directly on those questions. We did get a few extra symptoms. Do you guys want to hear those? We can go over them. Yeah. Yeah. Okay. So number 11, dependency on knowledge trapped in data analyst heads.
Yeah. Another bottleneck. Yeah. They think it's job security, but it's not. Yeah. And number 12, your users are constantly identifying and requesting new data sources. Oh. Yeah. Is that a problem or is that a... That's actually a good thing, isn't it? And they wouldn't be doing that if it was so hard to get the data.
Yeah. Maybe that's one of those good problems, right? Because you can go faster and faster, increase your cycle time capacity, and guess what happens? The business wants to go faster and faster. They'll always want to go a little bit faster than you can go, it seems to me. Yeah. Yeah. Yeah, and I think if you build for speed instead of being defensive, I think you can support that. And that's why analytics is a river.
There's always new questions. There's always a new data set to be analyzed. There's always a different way to do it, and we're trying to search for these nuggets of super valuable insight in data, and when we find them, it's incredibly powerful and in some cases transformative for organizations. But the way that you get there is by trying things out a lot and seeing if they work, and that's why business users are always going to have...
And that's almost my own personal one. I used to push back on them. I was like, "You don't understand. We can't go fast and not break things. You just don't get it. I'm a technical guy. You're not a technical guy. You don't get it." And I've just come to believe that the business people are right. They really just want to ask a lot of questions, and so it's up to us to figure out how to have a world where we can live in that world, and it's not that our business users are stupid and don't understand.
It's that we haven't properly built a system that can support them, and that's what DataOps is about, is really doing exactly what Wayne says. It's never going to end. And if you're not going to suffer needlessly, build a system to handle rapid change and rapid response, and then your life's going to be a lot better and your business users are going to be a lot happier.
01:10:00
Great. Well, I think that's a wrap, so thanks everyone for taking the time to join us today. I know we went way over, so for the brave souls who hung on, thank you so much. An extra big thanks to Wayne for joining us. We loved having you on, and great conversation with Chris, as always.
Yeah. To all the attendees, we'll be sending out the recording and the slides within the next 24 to 48 hours or so, so be on the lookout for those in your email. If we didn't get to your question, we'll follow up with you directly. If you have any additional questions, don't hesitate to reach out to us directly as well. So thanks again everyone, and have a great afternoon and evening.
Thanks everybody. Thank you. Thank you, Wayne. Bye. Bye.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Questions from this session
How do you know your organization needs DataOps?
Look for the symptoms rather than the label. The data team is flooded with tickets and burning out, business users do not trust the data or are being used to debug quality issues, source system changes keep breaking pipelines, analysts rebuild pipelines that already exist with small variations, SLAs are missed, and data scientists wait months for data and computing resources. If deploying one predictive model takes months, that is the same problem.
What is the difference between DataOps tools and data tools?
Data tools work on the data itself: ETL, BI, analytic databases, data science platforms and governance catalogs, a market worth roughly 100 billion dollars. DataOps tools work on the process and the workflow around those tools, integrating with them rather than replacing them, and are measured by business outcomes such as lower production error rates, faster and less risky deploys, and team productivity.
What types of DataOps vendors are there?
Eckerson Group classifies them four ways. Pureplay DataOps vendors do DataOps and nothing else, with DataKitchen as the example. Pureplay DataOps+ vendors add adjacent capability, with DataOps Live as the example. Data pipeline platforms such as Zaloni build DataOps into a broader pipeline product. Purpose-built tools such as Unravel address one slice of the problem.
How is DataOps different from DevOps?
DataOps applies the rigor of software engineering to the development and execution of data pipelines, but the pipeline it manages is different. Software CI/CD does not cover self-service sandboxes, meta-orchestration across many tools, or continuous testing and monitoring of data in production. The two also arrived a decade apart: the first DevOps event was in 2009, the DataOps Manifesto was published in 2017, and the first DataOps event was in 2019.
How does Intel practice DataOps?
Intel builds reusable design patterns so a pipeline can be created quickly and changed without losing consistency, and uses metadata to automate pipelines so engineers are not adjusting data models and transforms on every schema change. It applies lean techniques to measure bottlenecks and then remove them, and runs more than 1,000 tests in its test automation framework, adding more continuously.
Where should a team start with DataOps?
Eckerson splits the starting point in two. Organizationally: educate the team, identify bottlenecks, optimize the processes around them, debrief regularly, and monitor the improvement. Technically: componentize your code, build tests for every code block, treat data as code, and invest in the tooling that supports it, meaning a code repository, continuous integration, test management, orchestration and monitoring.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.