On-Demand Webinar · 1 hr 1 min
DataOps & Analytics: A Recipe for Accelerating Business Value
Mark Marinelli, Head of Product at Tamr, and Chris Bergh of DataKitchen introduce DataOps: why large organizations need it, the benefits teams report, how to get past the obstacles that stall adoption, and real examples. Recorded 2020; updated August 2026.
What you'll learn 8 points
- In a DataKitchen and Eckerson research survey, 79 percent of respondents said they have an unacceptable number of errors and incorrect values in their data analytics products every month.
- Gartner's March 2020 survey, Data Management Struggles to Balance Innovation and Control, found data analytics teams spend only 22 percent of their time on new initiatives and 56 percent on operational execution.
- Deming's finding that 94 percent of causes are common cause is used to make the case for fixing the process instead of looking for a person to blame, alongside Elon Musk's line that the real difficulty is building the machine that makes the machine.
- DataOps runs three orchestrations, not one: the Value Pipeline in production, the Innovation Pipeline that moves changes from development to production, and an Environment Pipeline underneath both.
- Tamr's DataOps framework has three components: process, meaning an agile incremental delivery model; technology, meaning the architecture of the data supply chain and the platform under it; and organization, meaning roles across mixed-skill teams and a working model between technical and business teams.
- Two operating models are offered. Shared services centralizes technical knowledge and resourcing but creates contention over priorities. The advisory model bootstraps projects with experts rather than implementers but leaves development and maintenance on each department. Start with shared services for the first project.
- The recommended way in is to pick one constraint, solve it, measure it and iterate: too many errors, slow deployment, or poor coordination between teams.
- An American transportation company running Nifi, Kafka, an ESB and Informatica across on-premises and AWS, with separate teams in different locations owning creation and operation of the pipelines, started with errors. It used a DataKitchen Recipe to test streaming data for fitness for purpose across the toolchain and route alerts to JIRA, email and Slack.
Slides
Transcript
Show chapters and dialogue 9,907 words
00:00:00
Hello, everyone, and thank you so much for joining us today. My name is Mingo Sanchez, and I'm a sales engineer at Tamr. I hope you're all staying healthy and safe. Tamr and DataKitchen are thrilled to be putting on today's webinar, "DataOps and Analytics: A Recipe for Accelerating Business Value."
In today's presentation, we'll be going over an introduction of what DataOps is and why large organizations need it. After doing so, we'll go over some of the key benefits that organizations experience when they're using DataOps, and how to overcome some of those common challenges associated with these sorts of pipelines. And at the very end, we'll have some time for Q&A to ask our speakers some questions from the audience.
So throughout the presentation, if you have any questions, please feel free to enter those questions in the control panel through GoToWebinar, and we'll address those at the end. As a quick reminder, this session is being recorded, so afterwards, if you want to access the recording, a link will be sent out to all attendees.
Our two speakers are from Tamr and DataKitchen, and the first is Tamr's Head of Product, Marc Marinelli. Marc has over 20 years of experience in enterprise data management and analytics, and as a man of many hats, has held positions in engineering, product managing, and technology strategy. He previously worked at Lucent Technologies, Macrovision, and Lavastorm, where he was CTO.
Our second speaker today is Chris Bergh, CEO and Head Chef at DataKitchen. Throughout his career, Chris has been a COO, CTO, VP and Director of Engineering. As a co-author of the "DataOps Cookbook" and the "DataOps Manifesto," Chris is an expert on DataOps.
And with that, I'll turn it over to Chris and Marc.
All right. Thanks, everybody. I'm going to start off. So this is Chris Bergh speaking. So,
the first slide, it has a picture of a river, and in some ways, at a 10,000-foot level, the data and analytics, and I'm going to use those words to mean everything from data science to data engineering to visualization. We're not building houses anymore on 30-year mortgages. We're kind of delivering a river of value, and a river of insight.
And our customers, in some ways, subscribe to that river instead of buying a mortgage. And so they demand kind of an experience where they can get original insight that's sort of delivered fast, like overnight, like Amazon does. And it's got to be really high quality. People are very intolerant of errors. And of course, it's got to be low cost.
And so that idea that analytics is a subscription experience as opposed to a 30-year mortgage on a house, is not really being met very well because analytic teams are struggling to actually deliver on this subscription experience. And there's a lot of statistics to back it up. So Gartner has this statistic that 50% to 80% of all projects fail. We've done a survey with Eckerson that most analytics have just a huge number of errors in.
And so a lot of people think that data and analytics is really important, and in surveys of CIOs, it ends up being some of the number one or number two things being invested in. Yet there's this challenge of why things aren't
working. And so I guess the question for this webinar is, given a subscription economy, should we be worried if our users are going to cancel our subscription, given there's a recession? And so let that be the front, and let me keep going. So, what we're going to talk about here is sort of an introduction to DataOps, and I'm going to go through that. And so the first idea is almost a philosophical idea. So what you do is much less important than how you do it.
And so let me repeat that one more time. What you do is much less important than how you do it. So how does that apply to data and analytics? Well, let's first of all think of it from a different perspective. If you look at industrial manufacturing and Tesla, he has this quote that the machine that makes the machine or the factory is more important and more complicated than the product.
And Dr. Deming said when you have a problem, it's often not the person problem. It's not a special cause. It's the process around it that's the thing that you should work on.
00:05:00
And those are two aspects of what you do is much less important than how you do it. And so there's a lot of things that we do in data and analytics. We create a model. We choose an algorithm. We find a data feature, create a data pipeline. We train the model. We do data visualization, data governance, data engineering, even the data itself. And so the idea, I think here is that how you do these things, the process of development, the process of deployment and monitoring and iterating and collaborating, is very important.
And so it's a really kind of contrarian idea that we have here, and that there's a huge industry, $100 billion, devoted to kind of selling you tools to do things, to do on top of your algorithms and data features.
And so I'm holding here, just trying to... There we go. And so, the tools, the technology, the data are just not important as the people in the process. And so, why is that? I think in a lot of ways, the data and analytics industry is kind of like the US auto industry in the 1970s. We're making cars, we're making a lot of cars, but the process to actually build a new model for the car takes forever.
And the cars themselves have lots of bugs, and they don't last very long. And so it's kind of slow to add new features. We have a lot of errors in our dashboards. Model deployment is slow, and if you worked in a Ford factory in the '70s, you weren't happy, and I think a lot of data and analytic teams are struggling.
And,
the other challenge that data and analytic teams have is that their time is sort of not well spent.
If you look at this graph that has a percentage of a team's time per week, they're just spending too much time on kind of stuff that doesn't add value. They're not creating new features. They're not adding new data sets. There's just a lot of errors in operational tasks, and that's actually due to the complexity that we've created, the complex organizations and roles, the tool chains, the data, the collaboration across different teams.
And, we're not the only one who agrees with this. We actually created this graphic a couple of years ago, and Gartner re-recently did a survey that sort of validates that people aren't spending the right amount of time on things that they should. They're not innovating. They're just kind of making the train to work.
And so, for many years, I've
worked in software, and then for about 10 years, and I still do, I manage data engineers, data scientists, people who do analytics work. And, as that, in that previous role, I was pretty unhappy, to be honest. Mainly because my data providers who gave data didn't care that I existed. They would forget files and drop columns.
And the data consumers were living in that subscription economy already, and they were constantly asking for innovation and pretty intolerant when me or my team were late. And, it really actually sort of makes for kind of an emotional beaten down and distraught team. And so also teams that just can't create and innovate. And so in some ways, DataOps is more of a movement that try to sort of help teams reclaim control of that situation. And from a definition standpoint, think of it as a set of technical practices and cultural norms and architecture patterns that really focus not on what you do, but how you do it, and specifically how you experiment, the cycle time at which you can get things into your customers' hands. And how can you actually do that in a way that has low errors and a way that can collaborate across all this complicated technology and people and all the different locations they are in the organization. And finally, it also talks about measurement, which is really, can you measure these processes that you have in the organization?
And so, what are those processes? And so, I'm going to talk about three of them here. And so the first one is really the process that you do when you're
running a factory, when you have something in production, it's accessing data sources, it goes through a bunch of, think of it as manufacturing stations. And that could be from access to ETL, to visualization, to data science, to MDM, to governance. And you're assembling data, changing data, creating artifacts, and finally, you're getting the result that your customer wants.
And so one perspective is that these pipelines, this orchestrating data to customer value is a very important process that you need to manage. And what happens in that process is there's just a lot of different great tools in
00:10:00
each category, and people love their tools. And I don't think I ever want to get into an argument between someone who loves R or Python or someone who loves Tableau versus Looker. They're all great tools. But it's the process that those tools live in, and the factory is more important than the actual tool that you happen to use in that factory.
And then if we look at those pipelines themselves in production, they're kind of mapped into how organizations work because where data and analytics lives in a company is varied. Sometimes the IT team will own a data warehouse or a data enablement or a data lake. Sometimes there'll be a data science team who maybe just works for the CEO or some other part of the organization, and then there'll be self-service line of business teams using their own tools.
But at the very end is your customer. And that customer is seeing the result of all these different teams working across all these different tools. And sometimes there's not one sort of factory assembly line, but a lot. And so that is, I think, pretty common in big organizations. And then the other part is, if you look at it, there's another process to deploy ideas from the fingers and the minds of your data scientist and data engineer and business analysts, from their fingers into the hands of their customers.
And that's generally called a deployment process. And that's a lot like what software has done with continuous integration and deployment. And a DataOps perspective is that, think of partly what your team is, is a software team, and you've got to handle that complexity and that code.
And if you look at these two pipelines, the pipeline that's happening in production where data's coming in and you're producing value, maybe it's batch, maybe it's streaming, and another pipeline that you're actually deploying to production, these things almost have to happen at the same time, right? And you just don't want to find out about data problems from your customer, and you want to make it easy and fast to take ideas from your data scientist's head and get into production.
And those things are very opposite and very challenging to do together. And there's a third pipeline here to think about. It's not just the process of running things in production or deploying to production. There's the environments that these things live in, the technical environment. And so, in a lot of organizations, the process to build these things is slow.
How fast can you create a sandbox environment? What's the difference between one environment and another? And so to think of these things as processes or pipelines that are important to invest in, important to manage, and important to run and perfect is a perspective that DataOps has. And so, I'm going to change hands here and give it to Mark.
And so, thank you very much, and take it away, Mark.
All right. Thanks, Chris. So, I'm going to take us through a framework that we've developed at Tamr over the years that we've been doing DataOps as part of our customers' data management evolution, oftentimes a digital transformation initiative, but really holistically changing the way that they're working with data. And, so I'm going to take you through that, both the constituent parts of it and some advice on how to get started in that framework, at whatever degree of maturity you're already at. But first, I'm just going to reframe the problem.
This is an exemplar of the type of headline that we see and don't like to see, wherein we've probably hired a bunch of data scientists. There's now enough supply to meet with demand. We've got great ambitions of how we're going to apply data science, artificial intelligence, machine learning, to evolve our business for competitive advantage, et cetera. And then the record scratch moment comes when we find that they cannot get their hands on enough high-quality, accurate data to produce good, trustworthy, actionable outcomes from their work.
And so to solve that problem, stop doing data science and start doing data engineering, and end up having to go upstream and do a lot of the stuff that we didn't hire them to do, just so that they can find decent data to work with. So I think probably a lot of folks on this webinar can identify with areas where they've had projects that were lofty in their ambition, but were stifled initially by a lack of good quality data, ongoing data, and what everybody needs to build something worthwhile.
If you look at the current challenge, how do we change this? How can we enable consumers of the data in the face of shifting requirements? They're always going to have new requirements. Those analytical applications or machine learning or AI, the world's best customer churn model that we're coming up with,
00:15:00
is always going to have shifting requirements coming from those consumers. Without the analysts and data scientists who are doing all of that work actually having to spend their time on the data engineering that gets them the data that they need, and knowing that we're going to have to do this as quickly as possible, rapidly delivering new value and new data in a scalable way using a ton of data, that's the point here, that changes all the time.
That's just the reality of it. So that's sort of a distillation of our current challenge. The good news is that this is a pretty good analog to a challenge that we've already solved in software development. Just as data management practice is all about delivering the best high-quality data as quickly as possible to the people who need it and satisfying those requirements, software is about delivering features with high velocity in ever-changing requirements.
And also doing so more recently in cloud-based SaaS applications or even traditional data management applications with a large volume of data that are changing and new data arriving every day. So, the way that software was able to solve this and lift developers into focusing all of their energy on developing against these requirements was a combination of the adoption of agile practices, but also the elevation of DevOps, development ops, as its own practice. So there was specialization in the supply chain, really, that was necessary to take these requirements, turn them into software, deliver them continuously and at scale.
So, a lot of what we've learned from how we develop software is applicable in the way that we're developing data applications. And make no mistake, anybody who's doing data management right now, as I've described it, is developing applications. You're an application developer.
So our framework, as many, includes a combination of process, changing the way that you work, technology, choosing the right tools for the job, making sure that you're not just changing who and how is working on this and using technologies that weren't designed for it, but rather picking the best-of-breed technologies that are designed to facilitate this new way of working. And lastly, organization.
What skill sets are brought to bear, and how are the teams structured to do this work?
So starting off with process. We're probably all pretty familiar with the waterfall methodology. I'm calling it the wrong way. And the framework here we'll use in the ensuing slides is data on the left in a variety of different sources, a variety of different formats, and a variety of different levels of quality and trustworthiness, some of which is under your control and some of which is not, because it's living outside of your four walls.
And then your constituent consumers on the right, who are building all of the aforementioned applications and have a variety of different needs and requirements from the data. Getting those data from the mess that it is on the left to all of the wonderful things we want to do with it on the right is where we come in.
The waterfall approach here, we have to take all of those data. We have to think of everything that we think everyone would ever want to do with those data. So I'm talking traditional enterprise data warehouse or master data management. We have to apply a model to those data. We take these data, we homogenize them and coalesce them around a format, third normal form, and figuring out all the relationships among the data.
We codify a set of rules that apply all of the application logic that's necessary to cleanse the data, to change the format of the data for those downstream consumers. We do a bunch of testing to make sure all of this stuff works. And then finally, many months and potentially many millions of dollars down the road, data starts coming out.
And at that point, these folks who need the data can start deriving value, there's your dollar signs, but also start reacting to it and asking questions and maybe even pointing out errors, often pointing out errors. So only at this crest of the waterfall are we now able to assess that we got something wrong far upstream. Now we need to go and fix it.
We need to change that data model, potentially, we need to retrofit those rules, we need to do more testing, we deliver it again, hopefully we got it better. This is enormously labor-intensive. There's a lot of people necessary to interview these consumers to codify their knowledge in these rule sets and the data model itself.
There's a very IT-driven process here. And
00:20:00
people will get it wrong. There's a translation issue sometimes in one of your consumers saying this about the data and then not being around when you actually have to go implement that rule and you make a guess and the wrong thing comes out the other end. So taking a long time to deliver value not only compromises the length of time and the time to value of what we do, but can also, people just leave these projects.
It takes so long for them to derive value and they have to put so much into them that these projects often go off the rails and fail. The majority of these projects fail. So the right way to do this is adhering to agile precepts here, is to start small. Start with a subset of the data.
Start with something that can actually deliver value, not the be all and end all. We know where we want to go, but starting off with a small subset of the problem of the data where we can deliver value. Allow those users to then react to that, start deriving value from it, but interrogating it, asking questions, providing feedback about the quality of the models we've built, the rule sets we've built, et cetera. Then we build on that.
We do incremental delivery of value. Sometimes we get it wrong. There's another explanation point. We've completely messed up this iteration of this. We've got to stop and go back and try again. But at least we are on a biweekly or monthly basis, getting some value out of the data, getting a lot more insight into how the data can be used.
And people will really give their time, less time, because we can automate so much of this over time.
So then, if we've all stipulated that we're going to work agilely, which I think most people in the modern era at least aspire to, then we've got to pick a set of tools that are going to get all that data from left to right. The seven boxes that you see here represent the major components of a modern DataOps data supply chain. We need a catalog, optimally, some way to crawl the data so that everyone who's using the data does not have to go down to the raw sources or know how those sources are formatted necessarily in their native form, but rather there's a one-stop shop where we can say, "Where is all the customer data in my entire organization?
Who knows about it? How do I get my hands on it?" So that's catalog and crawling. Then we have to take these data and move them typically from those raw sources into a lake or into some area where we're going to do analysis. To do that, we need some storage, we need some compute to build whatever applications we're going to use.
And then we're going to apply the cornerstone here of mastering and quality treatment. We're going to remove duplicates, we're going to enrich the data, enforce data quality validations, make sure that the data are adhering to how they should be, and out the other end, we're going to publish these out to those consuming users and applications in a simplified way. So much different from having to go back to those raw sources or even the catalog.
But a nice distillation of here's a way to go find the best supplier data we have in the entire company. Here's a way to go find the best customer data we have in the entire company. Here's the current version. Here's what went into it, and when new data arrive or something changes in our upstream processes, here's a new version.
You should use this instead. There's an outer loop to this that is, one, applying the correct or necessary data protection and governance policies to these data. We want to collaborate. We want all of the people to get their hands on the data that they need, but we need to do so typically in compliance with some regulatory oversight or whatever our data stewardship and governance policies are.
So we need to make sure we do this responsibly. But the other loop is around gathering feedback from the consumers of those data about the quality of those data, about the usage of those data, so that we can make the necessary upstream corrections in the source data or in any of the treatment that we're doing of those data.
So that's, I think- Have been underserved in really meeting the users where they're working with the data and getting lightweight feedback, asking them questions about the data. Is this right or wrong? Is this the most current version of this customer? And systematically being able to collect those data and do something about it. So these are the major components when you're going to go and pick out tools, and I haven't put a bunch of logos on the slide. As Chris said earlier, there are dozens and dozens and dozens of these tools. There are going to be different ones that more befit
00:25:00
your own needs. But in common, there are characteristics of these tools when we fill in all of those seven boxes on the prior slide that we want to make sure we're looking at. One is scale out and distributed. If you can build all this new stuff in the cloud, start there. The economics are great.
The ability to have unlimited but temporary storage and compute for different types of workloads, all of this stuff is wonderful and we should get there as fast as we can. Collaborative. There's a huge application here in everything that I've mentioned so far for machine learning and artificial intelligence as a way to automate a lot of the stare and compare or any sort of data quality or QC checks have been done on these applications as we're building it, just like in an agile software development shop, you're going to offload a lot of your testing to automated tools that are going to be able to remove the necessity of traditional QA of manual testing. Want to make sure that this stuff is open. If I'm going to go and pick the right tools, they should be best of breed, but they need to work with each other.
So I need to avoid tools that have proprietary formats or proprietary ways of doing things. Everything should be exposed via APIs, and I can mix and match and do the orchestration among them. Things need to be built for continuous processing. The data never stop arriving. They never stay in one shape. So we need to accommodate batch and streaming workloads and just from the beginning, know that things are going to change often.
And lineage and provenance should be embedded in all of these tools. For everybody on the right side of this screen to really trust the stuff that's coming out, they need to know oftentimes where it came from, who weighed in on it, and be able to perform forensics if some of the stuff is incorrect and know who to talk to and what data sets to consult as they have to do any remedial action.
The other side of the tooling here is infrastructure. So being able to-- and it's wonderful now, the three major cloud vendors and then the traditional on-prem vendors have big data solutions, some variant of the old Hadoop stack. So you can use EMR and AWS or Google Clouds, GCP, or Microsoft's Azure or something like Hortonworks or Cloudera. There's plenty of these tools packaged up nicely for you to do this sort of infinite scale-out.
Then who does the work? We've organized this in three major categories. Now, there are some people in your organization who are already doing this work and do not have any of these titles, but being deliberate about a separation of concerns and specialization within the different roles that are necessary to get the data from left to right is really important, especially as you scale out.
You want someone to really think of themselves as a dedicated curator of the data, for example, and that they live and die by making sure that the people on the right side of this screen get the data they need, the right data. You want someone to be specialized in data engineering. Let those data scientists go off and be scientists.
Let the data engineers focus their energies on data engineering. So these data suppliers are really your traditional owners of the systems. They need to be consulted when we're going to build anything. They are typically represented at the executive level by a CIO. The data consumers on the right side of the screen, that's any set of users who want to build some analytical application or operational application.
They're going to be represented by the chief marketing officer, chief financial officer at the executive level. And this new class of users with the elevation of the CDO, chief data officer, or the chief analytics officer at the executive level, is representing this new group that is dedicating themselves to doing data engineering curatorship and stewardship, the governance of these data.
Said you've already got some people probably walking the halls that do some of this work. How do you find them? A good way is to look at what tools they're using right now, and what you intend for them to use going forward. So a bit of an eye chart here. You could read it in posterity. But being deliberate about, again, the separations of duties and nominating people within the teams to at least perform this role initially, and then eventually to have this as their position in the whole supply chain.
00:30:00
Structuring these things, there's a couple of ways to bring these applications to light. Line of business wants to stand up a new application, and they need a whole raft of data that haven't been prepared for them before. We can do the shared services model on the right, where we've stood up all of the aforementioned tooling and people, and you just come to us with your problem, and we solve it, and out comes data. And we work with each one of those teams intimately to make sure that we're iterating and getting the right stuff out to them.
There's also an advisory model wherein You're going to select the tools or select the right tools for whatever this application is, and you may provide some oversight in how these applications are built, but you're actually not taking on the entire competency of building all of these apps yourself. Those are for other teams within IT or maybe even the line of business themselves. So depending on where you are in the maturity of your organization, depending on where your budgetary allocation is, the shared services model or the advisory model, these have their advantages and disadvantages and you may kind of fluctuate between the two, but these are the two major models that we've seen.
So people, process, technology. I've given a layout of how to go find the right people, how to find the right tools, or at least what those tools need to be, and how to apply them under the rubric of an agile development cycle. Now, how do you get started?
Process. Agile is key. I can't say that enough. Absolutely choose a model that works for you. Maybe you're already doing Scrum or Scaled Agile, is nice in larger, more complex organizations. But choose a model that works. Make sure that you don't try to entirely contort your business to fit the model, but rather you contort the model to fit your business. People can get pretty religious about Agile, and that can only end up with inefficiencies.
Put Agile to work for you, the way that best befits your business. Then go out and look at the set of available projects. Everybody out there has a hobby horse in every line of business operation. And score them on not just the value to the business of solving the problem, that's the obvious one, but also the availability of the data.
It is better to solve the third most important problem in the business that you can start solving next week because you have access to the people and the data, and you're not subject to PII considerations or something like that. It's better to solve that third most important problem in a couple of months than to solve the most important problem 18 months from now.
So it's really important to find a good balance between accessibility of the relevant data and people and the value of the problem so you can put your first points on the board. And that's defining that high-value, data-rich project that is actually going to be hard to solve, but is going to be that right combination of accessibility and value that it can be solved relatively quickly.
And this is going to be the exemplar now, when you can show to the rest of the organization, "We solved this really hard problem. We did it in a tenth of the time with half the people that we normally would. Who wants to apply this to their problem going forward?" Technology. You'd want to find your end state.
We're not going to get there immediately. We don't want to boil the ocean here. But think of that, like I showed you, those seven boxes, what those tools are going to look like, how they're going to interoperate, how you're going to apply orchestration, and the overarching blueprint for your data management landscape. And then we figure out how we're going to increment there.
Look at the current tools that are available and the current systems, figuring out which one can be replaced immediately, which one has a lot of stickiness and can't be replaced. So we know a rough order from a systematic point of view, where we're going to be able to start taking advantage of new tooling.
Wrap some of these legacy tools in APIs so that they're emulating how those tools or those components are going to work in the future state, in that interoperable, nice free flow exchange of data via APIs. But you can wrap some of these monolithic or legacy systems in those APIs, keep them right where they are, but abstract yourself enough from them so that eventually when you rip them out, everything else doesn't have to change.
And then start with low-touch, high-value proof of concepts to evaluate the new technologies as you're looking at replacing some of those old technologies.
00:35:00
Then getting started at the organizational level, finding who you've got and mapping them to those areas of specialty. You really want to make sure that your data engineers are identified and minted to specialize in data engineering. Again, trying to put a bulwark between the consumers and scientists and data scientists and analysts, from actually having to do any of the data engineering work.
They're going to work very closely together, but there is a real advantage in specialization there. Find people who will be able to fill the newer specializations of data curation and data stewardship. Again, data curation is responsible for getting people the data that they need in the formats that they need it. Data stewards are responsible for making sure that we are collecting all of the necessary feedback about those data and their consumption, and we're fixing the data, or applying governance policies to make sure that the data are being used responsibly. Then you create a quick cross-functional team for one of those first projects that has at least these four capabilities. That might be two people wearing four different hats.
It may be more than that in the beginning. Then choose your operating model. Start with that shared services. You take this on all by yourself for this original project, and make it your problem, and provide a nice layer of abstraction from the line of business so that they can just show up with their requirements and you can give them outputs. And start with shared services.
You may evolve that over time. And make sure that you've got the proper error cover, that the chief data officer or the CIO or the CAO or whomever has overarching responsibility for the success of Leveraging the company's data, and also has the authority to make sure that they can marshal the right forces to get this stuff done. Without their buy-in and oversight, this is going to go nowhere.
So that's my take on both the framework and how to get started. I'll hand over to Chris, who's going to dig in a little bit on their view on where to get started as well. Thanks, Mark. So let me give a case study on where people start, because I imagine everyone here works for a large-sized company.
They've already got analytics in production, and perhaps the reason they're here is they've seen some of these business constraints that drive people towards wanting to take a DataOps approach. Perhaps you've seen a lot of errors in your dashboards or reports. I was talking to a large multinational company just this morning, and they count that they have over 250 problems a month in their production analytics. Maybe that's the one constraint that you'd like to solve.
There's plenty of companies who will take months to deploy 10 lines of SQL or a new predictive model from their development environment into production. Perhaps that's the business constraint. You want to focus on deployment. A lot of organizations have poor collaboration, either between a data engineer and a data scientist or between the people who work in a centralized function and a decentralized function, the hub and spoke model.
Or perhaps your chief data officer just wants to get analytic about the work their analytic team is doing. And so these different constraints are ways to focus what you can actually affect with DataOps. And so I'm going to tell a story about what one company did to do that. And so if you look at these as constraints, you can then work and think of, "I want to fix this constraint.
I want to fix this bottleneck." And this is actually not an idea that goes way back, 50, 70 years to manufacturing. And there's a thing called the theory of constraints, and there's a book called "The Goal," written in the '80s. There's also a book called "The Phoenix Project" or "The Unicorn Project" that talks about pick a constraint, fix it, find the constraint or bottleneck that has the biggest impact, fix it, and then keep going. And you keep hearing this term iteration, and iteration or cycles, I think are really important.
And not only the cycle at which-- Don't start off trying to do everything, pick one thing, fix it, and iterate and improve. And I think that goes for your analytics. It also goes to how you implement DataOps. And so if we look at this customer who had too many errors. They're a large multinational transportation company.
And if you look at their architecture, it's really cool in some ways. It's got data sources on the left, customers on the right. Some of their things are actual transportation of items that are streaming data. Some are systems that are like their ERP system that are more batch. They've got different tech tools. They've got Kafka and NiFi.
They've got an enterprise service bus. They've got Informatica.
00:40:00
And so they've got all these sort of ways that data gets into the system, and then they've got a big data warehouse, and they use notebooks and Tableau, and it's on-prem. And then they're also like a lot of companies, experimenting and having part of their work done in the cloud. And so at the end of the day, their customers, the people, the business people who are using that data are upset because they're getting errors.
Things are late, and they don't know where the error is. Is the error in the actual machine that produced the data? Is the error in something in the Kafka part? Is it in the Informatica part? Did something go wrong in Oracle, or is it the if-then-else I put in my Tableau workbook? And so this complex, error-prone environment also goes into an organizational structure that runs it.
And there's different teams that run these different tools. They work in different parts of the organization. There's IT teams and data warehouse teams and data science teams and cloud teams, and so there's this complex ownership. So one of the things that we say is, how do you prove to yourself that things are going to work before your customer sees it? So don't hope that things are going to work.
Don't hope your source data's right, actually prove it. And so what we did is actually developed in our software a thing called a recipe. And that recipe sits sort of lightly on top of all these tools and technologies and kind of monitors and observes not only the systems, are they up and down, but actually the data itself.
And so that way you can tell if, for instance, a field is wrong or the data itself is incorrect or some processing that happens on the data. And there's techniques that actually come from manufacturing called statistical process control that you can actually look and see. You can actually dig into the data and build rules that say, this data is right or not. Follow the bouncing ball on the data as it goes across systems. But the whole idea here is hope is not a strategy.
Don't hope your complicated data and analytic systems work, and then work nights and weekends to try and fix it. You don't need to carry that hair shirt in data and analytics. You can change it. You can develop software that actually goes in and works and tests, and you can prove to yourself that before your customer sees the data, it's right.
And so if you do that, it's actually focused on errors, then you can actually start alerting when things go wrong. So if one of those data providers drops a data file that they should, you can then get an alert, and then that alert can actually have you call that data provider and say, "What the heck's going on?" And those alerts can be delivered in different ways. But if you start thinking of this like a factory instead of that you can own and improve, and you think of building an operational system, and working on that factory itself, I think you can then work to drive errors down. And the benefit of driving the errors is that you end up having more customer data trust, and actually your team ends up having more time to innovate because they're not running around on nights or weekends trying to fix things, and they've got more time and more trust, and they don't live in fear of something breaking.
Because a lot of times people develop these complicated analytical systems, and no one wants to touch them because no one understands them, and you're just afraid to break them. And so we've talked a lot about DataOps in general, and so I showed these two graphs at the beginning. And one of the things that we all want to do as a data scientist or data engineer, we're involved in trying to solve customer problems. We want to do new things and cool things.
And so we just don't spend enough time on that. And it's because of this sort of the type of complicated, multi-tool, multi-data set, multi-people technology environment that we're stuck with these errors and operational tasks. We're having too many meetings, too much process to be able to get that done. And so what we'd like to do with DataOps is to be able to change that equation so you can spend more time doing new, innovative things.
You can try out a new tool and less time having to get up on Saturday morning because a data provider had forgot to load a file, and now your CEO's yelling at you, and you've got to go figure out and fix it. And if you've been in the data field, everyone knows that situation. And so I think by focusing on error reduction, you actually improve your ability to innovate.
And likewise, focusing on deployment latency, how fast can you get something from the idea in your head into production? Can you do that in hours or minutes instead of weeks or months? There's a side benefit of that. As you learn more as an organization, you get that feedback quicker. And so there's something I've learned over the years is just to be very humble about what I
think my customers want. And it's much better to get something that's 70% right in their hands sooner and get feedback than to spend all that time getting what you think is 100% right, only to find out that it's wrong. So focusing on the cycle time, the
00:45:00
deployment latency, focusing on counting your errors and actually focusing on your team
happiness, and I think are benefits of adopting both our software and a DataOps approach.
And so I'm going to switch back to Mark, but thank you much. Thanks, Chris. So we're going to have time for Q&A, but a parting shot here. We've talked quite a bit about process and people and tooling and the skill sets and the tool sets that are necessary to make this stuff work, but there's also an overarching sort of mindset shift, and kind of abandoning some of the ways that we've gone off the rails before.
One of them is just you can't boil the ocean upfront. Having an a priori understanding of how everybody's going to use all of the data and all of its nuance and what all the data are going to be, and trying to corral all of that from the beginning, not going to work. You have to have a long view and know where you're going with your platform and with your vision for leveraging your data.
But it's really all about quick hits, quick wins, and then incremental improvement over time to get there. A single platform is not going to solve your problem. Those different boxes that I brought up, the different constituent components that provide functionality in the modern data engineering supply chain, are best of breed. Choose the best of each one of them that befits your own needs, and you're going to do more work if you build the infrastructure to pull them all together, to orchestrate and to make them play well together. But they're designed to play well together, to be interoperable.
So taking on that work as opposed to a single vendor or a single platform has a bit more upfront cost, but there's enormous benefit in the optionality that that gives you down the line. When somebody comes up with a better data quality solution than you're using now, or somebody comes up with a better data catalog, it's a lot easier to swap out
your old catalog for their catalog because you've had this nice layer of abstraction between, and a loose coupling between all of these tools. So you're not beholden to one vendor's platform or their roadmap for the next five years, and that's part of the problem we're trying to solve here. If you're going to use open source, of which there are a raft of tools that are immensely capable, just don't underestimate the effort of doing that work. There's a reason why companies like Red Hat or Cloudera were so successful in being able to package this stuff and provide services and training and complementary material that made it easier to use open source.
So caveats there if you're going to take it on and make sure you know what you're getting yourself into. And then the last thing is people. A lot of the people who built these systems and who are part of the problem you're trying to solve are trying, hopefully, to be part of the solution.
But there's just behavioral stuff. There's data hoarding. There's the exposure of issues with the data by sharing the data. Imbuing your organization with a culture of data is immensely important, pretty hard to do, requires some sound leadership. But really the principal thing that's going to drive that is adopting these new practices and putting points on the board, solving the problems a lot more quickly than you've been able to before. New projects take a 10th of the time.
The trust that people have in the data coming out of whatever you're building for them goes up, over time, there's a virtuous cycle there of increased trust and increased collaboration that'll make this work.
So, with that, I think we'll conclude, and Chris and I can take some of the questions, I think, that have been coming in in the chat. Now, Minga, over to you to moderate that. Perfect. Well, thank you, Mark and Chris, for a wonderful presentation. Just as a reminder, this webinar is being recorded, and that recording and the presentation will be shared afterwards with all attendees. So no worries if you want to view this material afterwards. And we'll do our best to address as many of these questions that have come in as we can during the live session.
But any questions that we don't have time to get to in the live session, we'll also follow up with a written response to. So to start things off, a pretty broad question:
00:50:00
organizations need to be able to trust their analytics. How can I make the case that agile DataOps workflows produce better and more accurate results than traditional workflows, especially given that processes are updated so often with DataOps? Sure. I'll take that one. Transparency, traceability are huge. I mean, first get it right, and it's pretty easy for people to determine whether the data that you're giving them are accurate or not. And gather that feedback, make sure that you have a really good instrumentation when people find that something is wrong, because in that first iteration, you didn't get it right, and there was a missing rule or a machine learning model, needed more training data or whatever.
Making sure that you have a systematic way to collect users' feedback, that's just as important as collecting user feedback on a product in an agile software environment, getting that feedback, doing something about it, and then sharing the remediation of that feedback so people know that they actually have some agency over the improvement of their data quality, will get a lot more buy-in. And the more eyes on the problem, the better. You've already got a lot of people using the data. If you can get a lot of them actually helping to improve the data by providing that feedback, that's huge.
And the other part of it is the transparency and visibility into where these data came from. If people know exactly what was done to the data, what sources were used, who weighed in on why the data were shaped a certain way or certain metrics were derived analytically, that's really helpful as well, so that they can have confidence that the right people worked on the problem and the right data were involved, and all the data were involved.
That's another big one, making sure that as people start working with these data, if they say, "Oh, this would be so much more interesting if we had this other CRM database that we didn't use this time around." All of those things, I think, will contribute to people having more confidence that all of the data, all of the people that could possibly make this work, are being involved.
Great. Thanks, Mark. So the next question, Chris, is for you. On one of your earlier slides about data catalog and governance, Wiki was listed as one of those types of data catalog sources. Can you expand on how you've witnessed the use of Wiki in this context?
Yeah. So if you've got a running analytic system, and let's say you just want to add a data file, right? And so you've got to add a table. Maybe there's going to be a predictive model that has that table that's added in a database that's joined with another table. You're going to have a visualization.
And finally, you're going to have to know what's in that table. And so, what we try to think about in DataOps is the flow of changes. And the flow of changes, the description of the data itself could be in a nice tool like a data catalog, or it could actually just be in a wiki that says, "In this table, here's the definition of...
Here's the attributes of the table, and here's where it comes from." And so, from my standpoint, what we're saying is it's really important that people understand where the data is and where it came from. We've got some customers who just happen to do that in a wiki, and focus on making sure that it's right and making sure that the data catalog matches what's in the current database and making sure it's deployable, and at the same time you deploy a schema change.
Perfect. Thanks, Chris. So the next question pertains to the current climate with everything going on in the world right now. So getting organizational buy-in for new projects is pretty tough. Given everything that's going on today, organizations are trying to cut costs wherever they can. How can these organizations justify investing in new technologies to build DataOps pipelines rather than doubling down and focusing on the tools they already have?
I guess I can take that one because I just had that same conversation with a customer just yesterday. And yeah. I think that the reality is that automation and removing work is something that everyone can believe in. And fundamentally, DataOps is an automation idea. Can I automate some of the manual work and some of the rework that's done in order for my team to be more efficient? And we've all had the great acronym life in AI and ML and, well, big data's not an acronym, but we've gone through this period of really insane growth in the data and analytics industry, and I think it is time to work on yield and how well we can do with what we have.
And I do think the idea of doing more with less is really fundamental to DataOps, and I think there is a good
00:55:00
argument, in a recessionary time, to invest in improving the yield on what you have.
Perfect. Thanks, Chris.
So another follow-up question is from the audience about how Tamr and DataKitchen DataOps offerings complement each other. So Mark and Chris, I'm sure you both can speak to this one.
Sure. I'll go first. I think they're very complementary in that Tamr is in the business of a central component in what I laid out of mastering and data quality and that feedback loop that I talked about. But we don't do it all. We need to be orchestrated alongside other tools that are complementary to us to subsume that entire data supply chain that I talked about.
And DataKitchen, with both their expertise and their tooling, I see as a great outfit to bring all of these tools together and solve the whole problem. And Chris, you may or may not agree with that, but that's certainly the Tamr looking out answer to that question. Yeah, no, I totally agree with that. And we're two T stops away from each other, so that also helps.
Being close is definitely something that makes it easier.
So another question that relates to what's going on today in the world, how do DataOps methodologies support remote work? Lots of organizations are being forced to adapt to a primarily remote workforce and may continue doing so in the near future.
I've got one on there. I'll flip it around and say not so much how does DataOps support remote work, but how does remote work support DataOps? I think a really important aspect of having all of these teams work together and institutionalizing their knowledge and codifying it in tools is to capture all of that knowledge systematically.
I kept talking about capturing feedback about the quality of the data. If that stuff can no longer happen by grabbing somebody in the kitchen and saying, "Hey, I think your Salesforce report is broken because you don't have the Cleveland customer base in it," but instead has to be captured somewhere in a system or on a wiki, which is what's going to happen now because we're probably not going to fire up Zooms to have those conversations, all the better if the way that people are having to work now, and the tools that they're leveraging to collaborate lead to more knowledge capture, and less stuff being trapped up in inboxes and hallway conversations than all the stuff that we talked about is going to work better.
Perfect. Thanks, Mark. And I think we have time for one last question before wrapping up. So one final question. If I have technical skills like database and programming skills, but no experience with DataOps, how can I present myself to be a part of a company that follows these principles?
I guess I'll take that. I think maybe I'll speak to it as an analogy. So when I used to manage software teams in 1999, there was someone called a release engineer. And that person, we sort of threw our code over the wall to that release engineer, and they were paid a little less than all us cool Java developers. And if you play that forward now 20 years, that release engineer is now called a DevOps engineer, and they're, in fact, paid, in some cases, more than the software engineers because the process of delivering things into production has become, and the cycle time at which you do it has become so essential to the success of software development teams. And I think that focus on cycle times and errors and the role that goes with it, the DevOps engineer role, I think is important. I think that's going to happen here, and it's already starting to happen in the world of data and analytics, the idea of a DataOps engineer, and it is a way that you can take your sort of skills and database skills and try to transform them into focus less on, I can make a database work better, but more I can make the whole team be able to deploy faster with higher quality and in a more collaborative way.
And I think that actually, I'm very excited about the whole field of DataOps because there's more companies being funded. In fact, there was just another one being funded just yesterday. And there's a lot of innovation going on in this whole field, and I think the perspective is that, and I think it's only going to accelerate.
And so I'm excited about the role of a DataOps engineer and the possibilities for the future of DataOps.
Perfect. Thank you so much, Chris, and thank you, Mark, as well.
01:00:00
And thank you to all our attendees for joining the session today. I hope this was interesting and insightful. We will be sending out the recording and the slides after the fact, as well as answering any questions that we didn't have time to address today during the live session. So thank you all again. Have a great rest of your day.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Questions from this session
How much time do data analytics teams spend on operations instead of new work?
Gartner's March 2020 survey, Data Management Struggles to Balance Innovation and Control, found only 22 percent of a data analytics team's time delivers innovation and new insight, against 56 percent spent on operational execution. The rest goes to improvements and technical debt. Errors and manual production tasks are what consume the difference.
What are the three DataOps pipeline orchestrations?
The Value Pipeline is production: data flowing to customers, which has to be tested and monitored so problems are not reported by the customer first. The Innovation Pipeline is deployment: changes moving from development to production without breaking what is running. Underneath both sits the Environment Pipeline, which creates and manages the environments the other two depend on.
What is the DataOps framework?
Tamr's framework has three parts. Process is an agile incremental delivery model instead of a labor-intensive, monolithic, IT-driven one. Technology is the architecture of tools in the data supply chain and the infrastructure under it: cloud first, loosely coupled, best of breed, continuous, with lineage treated as essential. Organization covers roles across mixed-skill teams and the structure connecting technical and business teams.
How do you get started with DataOps?
Inventory the projects available and score them on data availability against the value of solving the problem, then pick one high-value, data-rich project complex enough to force end-to-end coverage. Inventory the current tool set on cost and skills, decouple monolithic processes by wrapping components in APIs, and build a cross-functional team with data engineers, a curator, a steward and the consumers, with executive alignment behind it.
Which bottleneck should a DataOps program tackle first?
Whichever of the three constraints the team already complains about: too many errors, meaning customers find the data problems first; slow deployment, meaning changes risk breaking production; or poor coordination between data science and analytics teams. Select one, solve it, measure the result, then iterate. Trying to fix all three at once is how programs stall.
What are the common mistakes when starting a DataOps program?
Five. Running a waterfall project measured in years instead of delivering analytic value along the way. Overestimating what a single platform can do. Overestimating what a single vendor can do, rather than aligning vendors on APIs and the expectation that they work together. Underestimating the effort to make open source work. And underestimating the human and behavioral challenges, which are the most common reason projects stall.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.