On-Demand Webinar · 60 min
DataOps & the Cloud: A Match Made in Heaven
Capgemini's Deepak Juneja, VP of Global Data Management and Data Fabric Practice Leader, joins DataKitchen's Chris Bergh on cloud DataOps: why a cloud data platform still needs it, why a single seamless system is hard to assemble out of disparate cloud tools, and what companies that got there did. Recorded March 2021; updated August 2026.
What you'll learn 6 points
- Capgemini's argument for why cloud environments need DataOps: cross-platform orchestration and platform independence, collaboration between data consumers and data providers, data democratization and self-service, and simplifying data operations through automation and lean process.
- Capgemini puts the business value of DataOps in concrete terms: real-time visibility of data and data processing, trustable data with fewer quality issues, hybrid environments provisioned in hours, data and insights delivered in days rather than weeks, and new products and services launched 30 percent faster or more.
- AWS, Azure and GCP ship a powerful collection of data tools with no defined process for using them as a system. The result is data integration without process integration, and it is overwhelming to design a solution from DataOps principles on top of it. What is missing is a DataOps superstructure over the tools.
- DevOps and workflow tools fall short at DataOps for six reasons: no end-to-end meta-orchestrated production pipeline, no environment pipeline, the complex team and data center coordination that data analytics needs, no common system or vocabulary, no process measurement to drive behavior change, and the fact that the work is not DevOps CI/CD but continuous self-service sandboxes, meta-orchestration, integration and deployment, and testing and monitoring.
- A top 10 global health company ran two separate analytic groups: a New Jersey team on a large Spark cluster with a best-of-breed toolchain and high-value drug development data, and an Azure cloud team on ADLS and Databricks with research data sets. Drift between the two schemas caused unexpected delays, errors, mistrust of the data on both sides, and conflict between the teams. Schema and meta-orchestration, system-wide testing and monitoring, and versioning resolved it.
- Production monitoring in this model runs five kinds of test: traditional data quality, statistical process control, location balance, historic balance, and business-based tests. They run automatically, in production, on top of the entire toolchain, sending alerts and keeping a history.
Slides
Transcript
Show chapters and dialogue 10,035 words
00:00:00
So good morning, everyone. Thanks for joining our webinar today. My name is Beth Pfefferle, and I'm the VP of marketing at DataKitchen, and I'll be the host today. So our topic is DataOps in the Cloud, a Match Made in Heaven. We're very excited to have a very special guest join us. We have Deepak Juneja from Capgemini.
And if you haven't heard of Capgemini, they're a global leader in consulting, technology services, and digital transformation. Deepak is the VP of Global Data Management and the Data Fabric practice leader at Capgemini Financial Services. He has more than 25 years of experience implementing technology systems and enterprise data management platforms. He's implemented enterprise data management initiatives involving data strategy, DataOps, data architecture, data warehousing, data quality, data governance, metadata management, data analytics, and more. So that's quite a long list of experience there.
So, welcome, Deepak. We're very excited to have you here today. So, Deepak will be joined by Chris Bergh. Chris is the founder, CEO, and head chef at DataKitchen. He's also a leader in the DataOps movement. He has more than 25 years of research, software engineering, data analytics, and executive management experience. At various points in his career, he's been a COO, CTO, VP, and director of engineering, and he's also the co-author of "The DataOps Cookbook" and "The DataOps Manifesto." So before I hand it over to Deepak, just a few housekeeping items.
So we are recording this session. We'll send out the video and slides to everyone as soon as they're ready. So within the next 24 hours, you should be receiving those. We're also reserving the last 15 minutes of the webinar for questions. So please enter your questions in the Q&A box on the webinar control panel during the course of the webinar, and we'll collect those, and we'll make sure we have time to answer those at the end.
So that's it for the housekeeping. I think we're ready to go. You can take it away, Deepak. Sure. Thanks, Beth. Hello, everyone. Welcome to today's webinar with DataKitchen. As we talked about, I lead the data management, data fabric capability, and particularly supporting financial services client. And I'm here to talk DataOps and specifically on cloud, how it bring agility and speed, and we'll discuss this with Chris after I present a few point of views.
So in today's fast-paced world, right? In order to keep up with market and competition in any of the sector and particularly financial services, it is very extremely important to leverage all the available data as well as the real-time insight for making business decision, right? Here is where DataOps comes into play, right? It brings value, agility, and all.
But let's start with some of the key drivers which you may be aware and what I'm seeing from all the clients across financial services. Start with insight available at speed of business, right? Based on survey conducted by Capgemini, more than 50% to 60% organization wants to make their critical decision driven by data. Organization have realized that if the business decision are driven by insight, there has been a significant performance advantage in terms of revenue, in terms of profitability, and overall, the client experience you provide.
And it is very important to bring this insight at the speed of business, then you can really get the value. In line with the same point, the second trends, what we see is real-time data processing. We have been producing data at batch or iterative mode, but now there is a strong need and desire for real time, as well as delivering the insight.
As we have realized, real-time data processing helps us to bring the right data what we need and drive them part of the critical business processes and available outcomes. So it is very important to drive real-time insights. To achieve all these two things, one of the key thing all organization financial services is trying to become lean and agile for data as well as for analytics. The idea behind is how can bring lean and agility around development, how do we quickly release new capabilities as well as insight and models, but also make data operations agile, whether you want to bring the trust into the data or whether you want to build new models and capability which needs those data.
Now, during the past year, what we have witnessed, COVID pandemic has impacted businesses all over the world.
00:05:00
The shift has accelerated the need of data marketplace, and more importantly, how do you become data-powered revenue-generating organization? So you want to leverage your data and insight, which can be made available to drive additional data as well as monetization of those data. So clients are looking for if they can make data marketplace available where they have internal data, they can bring external data quickly, and then build the insight, and make it available for other organization to leverage, as well as their individual clients to leverage. And finally, all to achieve it, it is very important to reduce error and improve quality.
So that's where clients are looking for all data analytics team, whether it's part of banking or insurance, looking for how we can reduce the error as well as speed up the processing of data, and more importantly, improve the data quality. The trust is becoming very important, whether it's data or the insight you want to make it available.
So next slide, please. So let's talk about what is DataOps and why it is required on the cloud environment. Now, DataOps is made up of three pillars, right? It's not about just technology platform and tools and capability, but it's a environment where you want to bring people with different roles and the processes, as well as having capability around that data platform or DataOps platform, which can bring the required capability and insight. So if you look at what's required from DataOps is to deliver trustable data and enable faster insight. And that can only be achieved by bringing people, process, and technology part of the overall program, as well as establishing that as a practice or center of excellence.
Now let's talk about what is driving DataOps for all these needs, particularly on cloud. As organization are moving towards cloud and meet their digital transformation or cost optimization goal, they're choosing multiple clouds, right, based on best of the breed. And there'll be some of the processes which will still continue to be on-prem. So it's overall a hybrid environment, and you're looking for some kind of a orchestration tool or platform which can enable end-to-end agility and speed.
While cloud platform has their tools available to leverage, but you want to bring platform capability which is across the hybrid environment, and that's one of the need driving DataOps. In the similar lines as we discuss lean and agile process, there's also automation needed where you can leverage for the purpose of driving the release of the new capability, as well as automate some of the data analysis capability or ingestion of the data and bringing other capability to the table.
So that's driving, again, the need for DataOps on cloud. Now one of the critical, as I discussed, people and processes. There's a big demand of collaboration between consumer and provider. So you're looking for a seamless collaboration environment where data consumer and data providers can work together and exchange data as well as build insight and collaborate in understanding what could be the right models to use, what could be the right data to be leveraged, and continue driving better insight and as well as drive the data monetization we discussed.
The last thing about is data consumers will like to have access to all their data, which has been challenging today. So they will like to have a quick access to the data which they are authorized and they have the role and responsibility, but looking for self-service tool which allows them to bring new data, or they can build self-service reporting, analytics, and more about making that data available to downstream consumers so that they can leverage for the purpose.
And also as a part of the data governance and stewardship, you're looking for access to lot of KPI around whether data quality, how are you managing your metadata, or what is happening with your PI information. That's also become important, how you can manage it self-service as there's vast amount of data and it's growing day in, day out.
Moving on to next slide. So let's talk a little bit about what kind of a capability you need. I'll be talking briefly about few points, but there are extensive capability which you need to build and I'll talk about how the journey looks like. To start with, there's need for data analytics as you're building those capability.
There's a CI/CD or continuous integration and delivery of
00:10:00
that capability needed, which can be automated from the point of you're doing the analysis to the point of delivering in the production and actually running end-to-end in production. So some kind of CI/CD capability which make all that automation. But you also want to look at can you still continue to leverage the investment you may have done in your current platform and, ETL capability or data capability, but bring that kind of orchestration which can help you automate development, testing, and finally, releasing in the production.
The next capability, which is important part of this cycle, is about testing automation. How do you make sure you bring the right data when needed, provision it as self-service, and then be part of the overall end-to-end cycle where automation can be done for the unit testing, integration, and overall testing with lot of other platform, which could be on cloud as well as on-prem. You want to have end-to-end automation around testing and trust the data quality as well as other aspects.
The next thing about is as you're bringing data together, and we talked about trust, it is very important when you're building some of the solution around it, you want to do the assessment of the data quality and see how the quality looks like and define end-to-end process where you can communicate some of the issues with the source data, whether it's internal or external, and looking for improvement in the data as time. So knowing what is the quality, tracking that becomes very important, and that's need to be aligned with this end-to-end process.
The next important thing about is provisioning environment. All data scientists or stewards and others who are working with the data responsible for delivering insight or building the trust on the data, looking for quick environment provisioning where they can use for as a sandbox environment or for testing. So automation around it, but thinking about more of a hybrid platform.
And finally, couple of two points last is all about end-to-end data life cycle. How do you manage creation of the new data? How do you manage ... changes to those, and finally, how do you demise the data, archival and all aspects. So how do you bring automation to end-to-end lifecycle? That's the key point about DataOps as a platform, as a tool, which Chris will little bit talk about later, what DataKitchen brings to the table, as where clients have seen this success. And finally, overall automation around monitoring the SLA, the quality of the data in production real-time, so that you have a good understanding what is available, what you can leverage, and how do you take advantage of those.
Moving to the next slide.
Now I want to talk about what is the final value proposition, or what is the value business will achieve with all this. First, business is looking for how do they speed up making new product available in the market, whether it could be new product, whether it could be new models or new services, and they want that to be data-driven. So DataOps helps to achieve that goal, first driving it through the data, as well as using some of these automation to speed up the journey more than 30% to 40%.
The other thing we discuss about, how do we deliver data insights in days and weeks and not in months and more. So how do you speed up that journey? That's the value that DataOps can bring to the table. Finally, quick provisioning of the environment, as we discussed. Surely trust what we can build on the data and make sure you minimize the data quality issues as you track end-to-end process and automation of identifying those, communicating, and then getting that resolution and seeing how the value is improving. And finally, real-time visibility of the data and landscape, which will drive to understand what data is available, which they can quickly leverage for making decision, and overall get the value from a business insight. But overall, DataOps is all about establishing a collaborative environment between data provider, data managers, and consumer, so that they have real-time data available, real-time insight can be generated, and more in a self-service marketplace mode.
Going to the next slide.
So we talk about what is the driver, what capability is needed, but now let's talk about the journey, how it looks like. For sure, you need to start with what is your business goal, what are the key challenges you want to solve? And work with your stakeholder to get their buy-in, which is critical.
You want to make sure as DataOps would impact multiple team, business, technology, as well as multiple teams across
00:15:00
other landscape. You want to make sure all the stakeholder is involved, they understand the capability, and you define a target state platform, what the capability needs, what you want to achieve, how you want to drive it, and building some kind of a journey and roadmap like this would help. The next step would be engaging in some kind of a POC or pilot, where you want to identify one or two use case, where you see the challenges, define what outcome you want to achieve, and then demonstrate the value it can bring so that everybody's comfortable. This is the investment.
This is the time and value will bring all the business value, what I talked about. The next would be now take the POC pilot, productionize it, with the capability. But more important, at this time, you want to establish a center of excellence who has people who have the experience and capability so that they could be the team starting this journey and helping organization to broad it.
And work with multiple business team, technology team to have their involvement so that it's a collaborative environment, as I talked about, to define what DataOps as well as the journey together. Next would be take it from one business area, use another use case. One use case could be, how do I bring data together quickly? Second use case could be how you handle the reporting and BI and self-service. Third use case could be you want to see how I can drive data, good quality data, and quickly for building some kind of a model. So expand it to other areas and other use cases and drive, see what value it can bring.
And then last, once you have established, surely DataOps as a platform and capability, you want to build incrementally as you use. So put a plan, how do you expand the capability around the DataOps, and how do you drive usage across the board? And one of the key points it will be helpful is as you go through the journey, use some kind of a platform and capability which has this.
And Chris will talk about how DataKitchen helps you through this journey and the success he has seen with other clients, as well as some kind of a partner who has gone through this journey and guide you through and help you through the success you would like to achieve. With that, Chris, I'll pass on to you.
If you can talk about your point of view on some of the things I talked about, as well as how you have seen the success with other clients leveraging DataKitchen and overall DataOps as a practice and program. Oh, yeah. Thank you. That's an excellent presentation, and I love this slide about the journey because we've been working on this for years, and customers have different ways to go on this journey.
And so in my section, I'm actually going to talk about this journey, and I'm going to use upstream as a metaphor for it. And at first, I'm going to talk about Why DataOps is a different way to think. And so this is, a lot of times, and like Deepak, I've been doing data and analytics for a long time, and I've seen a lot of different companies.
And a lot of times in the past, my focus has been on kind of the task that's right in front of me. I've got to build a model. I've got to do some data transformation. I've got to build a visualization. I got to get the team working and the data, and it's all, you feel burdened with this backpack of tasks that are full of things that you've got to do.
But I think the challenge is that task focus, as Deepak pointed out, really isn't working. A lot of data and analytic projects of any form are failing. And looking at these quotes from Gardner or Eckerson, I think we've got to realize that by focusing just narrowly on the task, the actions that I have to do, as a team, we're not succeeding. And I think the fundamental perspective of DataOps is that it's an upstream problem, that it's not so much about the model, it's about how you develop the model, how you deploy the model, how you iterate to improve the model, the data transformation, the MDM, how you work within your data fabric to test and collaborate and measure your process.
And so, I think what we're going to see in the next presentation is how people have focused, what sort of principles they've focused on. And it's really, I'm going to use a metaphor of focusing on four principles, which are error rates, which are errors could be caused from poor data quality or you're late or something happened in the processing.
Cycle time, which is the deployment of new ideas from the heads of your data scientists and engineers into production, collaboration and measurement. And those four areas are really the focus of a DataOps program, as Deepak pointed out. Because honestly, if you have a lot of errors, either from poor data quality or something you did wrong, your customers don't trust the data.
They don't trust your team.
00:20:00
And, as a result, your digital initiative, your data transformation fails to happen. And then if it takes forever to get something into the hands of your customers, your team doesn't learn, and you don't actually iterate and then therefore iterate. And I think a lot of organizations, data and analytics is not one team, and so collaboration is incredibly important.
So let's talk about how companies have gone about doing this. So if they have these goals, lowering error rates at the bottom, cycle time, collaboration, and measurement, we can break those down into a further set of components on what people do. And so we're going to talk through some customer examples. And so if you look at that next column, they're looking at production monitoring or data observability, and we're going to be doing a webinar sort of next week on that.
How do you test in development to prove that things work? How you orchestrate all the tools that are involved in data and analytics. How you manage your dev and prod and all your other environments, technically, that you work through infrastructure as code, and that working in the cloud has created a great opportunity to do that.
Actually, how fast can you deploy from development to prod? And then how do you collaborate both within your team and within between teams? And then finally, as a leader, how do you measure all this work that you're doing? And so you could kind of look at an axis where people start on DataOps. So one is what sort of goal am I after? Is my primary goal lowering error rates, or is my primary goal decreasing cycle time?
What sort of things that you have to do technically, and then what tools are you working it with? Is it a database, an ELT, ETL tool, a data science tool, a BI tool, and a governance tool? And so this graph is-- I'm going to start off with one example of an American transportation company, and they really focused on kind of production monitoring and observability. And that's really their focus across this tool chain. And to do that, they had to do some orchestration of multiple tools because they had a really cool and complicated tool chain that looks like this. They had sort of, as Deepak said, companies want real-time insights, so they have streaming data.
They want regular insights, so they have batch, and they want to move to the cloud. And so they had an on-prem environment, a cloud environment, a data warehouse in Oracle, some Tableau, some notebooks, and some ingestion routines coming from vehicles. And the challenge is these were actually owned by lots of different teams. And so, you may have a streaming team own the ingestion, the IT team own the agents that are running in the vehicle.
You have an ETL or a warehouse team owning one part, a data science team, a BI team, and then who are you going to call if something goes wrong? And honestly, things always go wrong. You're always going to have problems in this. And I think we've all had the sort of finger-pointing challenge of like, something went wrong. We don't know where it is, and VP, CEO is breathing down our throat to fix it.
And unfortunately, it's Friday afternoon and you got to call your spouse and sort of beg forgiveness. And I just think that that's untenable, and I think, one of the reasons is that we just don't know the source of the problem. And so one of the things that we do is actually don't change any of your tools, right? You've got cool databases and ETL tools, but sit on top of all those and what we call meta orchestrate those tools.
So connect to those tools so they work, and then observe what's happening. Pull little bits of data, pull little bits of artifacts of the data to make sure that they're right. And overall, you want to make sure that you know that as your system's running, the answers that you're giving to your customer are correct. Don't just hope that they're right. Prove that they're right.
And if something looks weird or something's wrong, send an alert, a Slack message, a Jira ticket, an email. And that drives low errors because you're actually watching the system while it's running and making sure that it's right. And not just like it has CPU time, it's that is the data the right shape and size. And the principles here behind that are really that you need to actually automatically, in production, on top of your entire tool chain- Look at the data and the things you're creating from the data, and send alerts and notifications, and keep track of history.
And to do that, you need to create tests or monitors. You need to actually interact with all these tools. And we've written and talked a lot about all the different test types that are out there. And one that is my favorite is to think of this as a manufacturing line. And as data's going through and you have different tools, things are being assembled from the data.
Well, a thing that manufacturing has used and lean techniques is called statistical process control. If you've got a million rows from a data provider and you've gotten a million rows every week for the past year, and they suddenly send you 10,000 rows, you should know about that. Right? It's a trend break, and that's what statistical process control's about.
So the whole benefit here is by focusing on errors, you actually get more time for
00:25:00
innovation and more customer data trust, and you get less downtime and less stress. And that's the principles of why this company decided to go with focusing on production error rates first. Now, you don't always have to do that. There are other companies where either their data was already good or error rates were secondary, and they were really focused on trying to handle the cloud landscape and trying to handle the fact that they had lots of different teams in different locations. And here's an example.
So they had a team that actually had built a really great Spark cluster, sort of best-of-breed tool chain, on-prem, high-value data, giving a lot of benefits to the organization. And then they had an Azure team that was using ADSL and Databricks and other IT teams, and they were also working in various research sets.
But they had to work together. Because it's a big company, and so big companies have things that are cloud and on-prem. And the challenge is when you have two groups working in two locations, they want to be able to have independence, but the relationship between them is hard. If something's happening in New Jersey and you're using Hive and StreamSets, and you're in one area, and you're using SQL Server and other tools, Azure Data Factory in California, how do you get these people to coordinate together, and how do you know about problems?
And so you can't just throw things over the wall to someone in another group who's dealing with the data. You've got to prove that it works and be able to coordinate. And so one of the benefits of this idea of meta-orchestration and deployment is that you should be able to notice that there's problems, and you should be able to notice those in development.
And so if you're going to make a change here in New Jersey in a data set, in a schema, you want to make sure that you don't find that you've broken anything downstream from that, and you want to do that in the development process, not in production. And so having something that encapsulates this whole system allows independent control in each of the clouds, but allows the team that's actually producing data to see if they've broken anything, to see if things work in development, and be able to check if that's right is really huge.
And that can actually speed deployment. And so the idea here is when you're thinking about a cloud, and there's just a lot of cloud tools out there, and there's a lot of cloud tool chains. And you're looking at Azure and AWS and GCP. If you look at the tools to just do your data work between ETL, data science, version control, DevOps, databases, analytic tools, data catalog, data governance, you've got a whole chain of tools, and they're really powerful.
And unfortunately, there's no defined system to be able to do that. And a lot of it's about the tasks that you have to do, but not about the upstream problems of building a system to handle it. And so from our perspective, and that's why we've built the company and I have a technical background and actually we tried to do DataOps in a similar tool chain with this, and it didn't work.
And so we had to build a DataOps superstructure to handle that. And really, one of the ideas here is what's missing. Because we talk to people and they're like, "Hey, Azure's got DevOps tools. There's Git and AWS. Why do I need to have another tool?" And I think there's six categories of problems. Because first of all, we've talked a lot about this term called meta-orchestration.
Somebody's got to sit on top of this entire tool chain and find out if it's running in production, be able to tell if things are changing in development. We're trying to introduce the world's worst acronym here, and so everyone's heard about DevOps and CI and CD, continuous integration and continuous deployment. We're trying to say it's not about CI and CD, it's about CSMOITAM, which is the world's worst acronym, because there's a lot more complicated, because the data world is different.
Because in software, if you look at it at a really big picture, software's got an application, a product that people are interacting with. And one of the reasons to be agile is to get something in front of them quickly so you can learn. And you still want to do that in data and analytics, a dashboard, a set of charts. But you've also got this other cycle of is the data right? Can you actually predict from that data?
And so the complexities of actually developing data and analytic products mean that CI and CD is not enough. And that has to go with environments. It has to do with how teams work across data centers and coordination, and it also has to do with the fact that data-- I've spent the first 15 years of my career being a software engineer and the last 20-ish now working in data analytics.
00:30:00
And software people and data and analytics people are just different and they have different preferences, and you need a way to do that. And then finally, you need to drive behavior change with measurement. So-
That's why we think this idea of a superstructure that works across that is needed and why we built DataKitchen. And then, if you go look, there's a whole bunch of details here that will-- if you want to check out on the things that you can get from your cloud vendor tool chain, but you actually need to do, and we provide in our platform.
So I give that as sort of a reference if you want to look at it later. So let's talk about another case, a top five bank, and they want to start their DataOps journey. Now, they were actually really concerned with the amount of self-service users out there, and it's a big bank, top five bank. They had thousands of users using data across the organization. And how do you monitor and track their usage?
So they had a centralized IT and data team, which had a really great data warehouse, databases, and they had people using it all across the United States. In Brooklyn, some small business, in Texas was high net worth, and the Northwest is the branch. But there's different needs, right? The New York people need business data.
Maybe they want a SQL database with Tableau. The Texas people want Python access because the high net worth person is a data scientist. And the Seattle person needs DDA data that's anonymized with Power BI. And how does a centralized team provide these resources to people? And the case here is really sort of giving them self-service sandbox environments, give them the data and the tools that are just for them, but watch it and make sure that they're not going off the rails. That they're not perhaps integrating data and producing things that have bias, that they're not taking data out on their laptop.
And so how does the central IT team group, who actually ends up getting blamed when someone on the self-service team gets something wrong, sort of give some governed control over all these sandboxes? And that's a really important thing, right? Because no one likes to get blamed when our prolific self-service users kind of do something wrong.
But the other idea here, and this is a phase two for them, is when your self-service teams are doing something, there's kind of three modes of work. One case, they may do something and there's no path to production. They've done it, it's answered a business question, they throw it away and never do it again. So they've built a dashboard and the business customer's happy, and they answered their one question and it's done.
There's another case where they do want to get it into production, it's useful, and there's sort of two cases there. One is that they've done good work. They've got maybe an Alteryx workbook. They've got a Tableau dashboard, some Python files. How do you sort of wrap that in a basket that supports DataOps? So it's version controlled, so it's stored in one place, so it's tested, so if the central IT team changes something, they can actually see if they've broken it.
And so that's one way, just wrap their existing work in sort of a DataOps blanket. And the other case is because the central IT team may notify it, they may say, "Hmm, doing an Alteryx workflow takes an hour. You know what? We could actually do that in our database. It would take five minutes, and it should actually be an attribute of a dimension that everyone uses," as opposed to this one-off thing.
Or there's a data set that they're using that needs to be pushed centrally, and so it sort of earns the right to be implemented, reimplemented centrally. And that's a good thing. But instead of being blind to the world, you should actually start thinking about it. And that's the principles here, and that really the way in which data and analytic teams work.
On the left here, we sort of our mental model is they're all centralized data engineers, data scientists, all working in one development team. And perhaps there's a production team working with them. But it's actually more on the right. There's more decentralized development, decentralized production, and how do you make that work in the real world?
And we're never going back to having a centralized group because it's just too slow. And the mix of centralization and decentralization is a principle that you really need to handle and really need to focus with DataOps. And so, I think that's a principle that's worth talking about. So the last case here, and this kind of goes to Deepak's slide, is how do you actually get people to do DataOps? And because at some point, we've all delivered data and analytics systems. We get it in production, and man, do we not want to change it because it took a long time. It's fragile.
You're just barely getting by. And when someone says, "Hey, you can build something and you can change it every day or every week, and you can do that in a way where you're not going to get
00:35:00
yelled at and not create a lot of technical debt. There's not going to be a lot of errors and problems." People don't believe you because it's kind of, in their experience, it's fantasy. And so that's what we actually can do in data and analytics with DataOps. You can actually have it where teams in different locations can work together.
You can deploy every day or every hour. You can run a system where you're not working nights and weekends to fix it and getting yelled at. And so how do you make that change in a company that's big? And it's hard enough with a small team, but let's say you're a big financial service company, right?
And you've got a banking division and a high net worth division, an insurance division, a brokerage division, and maybe you've got a whole CIO that's supporting it. And there's different data teams doing different things. And so some of them may be on Amazon, some of them may still be on on-prem, some of them may be on GCP.
The cloud has presented this opportunity of people to sort of go and do really great things independently. But if you're a CIO or a chief data officer, how do you actually and realize that, yeah, you can go fast and not break things. And that's actually the key to your success. Walking upstream, your problem's not going to be solved by buying the next fastest database.
Your problem's not going to be solved by finding that next tool to make your individual contributor take things from their mind and get it into some code or configuration faster. You're going to have to build a system. You're going to have to do DataOps, just like people have done with DevOps and software. Buying another compiler or another app server doesn't make your team deliver web apps faster.
And so how do you do that? Well, in a lot of cases, they're complicated organizations, right? So you may have different teams who are doing data and analytics across a company, and you may have a chief data officer who may work for the CIO or not work for the CIO, who's trying to actually get these teams to follow DataOps principles.
And they don't work for them, or they have influence, but they don't have power. And so how do you get that influence to happen? And how do you actually get everyone to start doing DataOps across the organization? And I liked what Deepak said about sort of starting small and proving value because there are these Doubting Thomases out there, or as we say growing up in the Midwest, people from Missouri, the Show Me State, they have to see it first, and prove that it is of value. And I think this is a really important point of how we do DataOps transformation, because you do need a platform like DataKitchen to do it, but you also need to think pretty intentionally about how your organization is going to manage change.
And that has a lot of different aspects, and some of which are organizational like this, how you talk about it. And thinking, with Capgemini, we sort of work to intentionally bring DataOps to convince people that DataOps is a thing that they can do, convince them that there's a community that can support them and help them build that community, find demonstrable projects, and then roll it out in a COE or a dojo or some form. And this is this sort of thoughtful process to do DataOps transformation.
There's a bunch of organizations around the country and around the world who are on that journey. And I think it's very parallel to what happened in digital teams who are doing websites and IT applications with their DevOps transformation. In fact, it's very parallel to auto manufacturers and manufacturers going from mass manufacturing to lean manufacturing or total quality management.
And so, I think these are really important, really exciting ways because I think in my career, this is the problem. You're not going to get solved by buying a better database or getting a whole new tool chain in your favorite data fabric is not going to solve it. We have a people and process problem, you've got to focus on that.
And so what's the conclusion here? So, as I said, if you look at a lot of companies today, where they are, if you think of this as a graphic equalizer, it takes them months to deploy something new into production. They are having problems all the time. Things are late, data's wrong, and they're just having errors every day.
And I've talked to companies with pride who says they've gone to one major data error a month in a large insurer. And to me, one a month is not great. That's awful. But my own experience is just having errors made my life, running data analytics teams made my life hell. And so how do you get error-free days?
How do you get error-free years where you can run your system without any major error? And then how do you stop the sort of Hatfields and McCoys of self-service and your data engineering team, or your data engineering team and your data
00:40:00
scientists, or even just getting two people on the same team to work faster and better together. And then how do you show your worth as a chief data officer? At the end of the day, we're all tasked with delivering insight, but insight's hard to measure. But how do you measure the actual work that your team's doing?
And if you don't do that, it actually ends up being highly costly and lack of success, and you have a lot of unhappy customers. And that's kind of the dirty truth in data and analytics, is a lot of CDOs are running through their jobs very quickly. A lot of new grads out of master's programs and data scientists are unhappy.
And a lot of our customers still are unhappy and still sort of like my experience 10 or 15 years ago, they roll their eyes at us going, "Oh, you big dumb data person don't know what I mean." And so how do we change the equation, and how do we stop having our data and analytic teams feel beaten down? And I think it is really to believe that you can push up on all four of these things at the same time, with the ideas and software that we have and partners like Capgemini, that you can deploy quickly and can do it with low errors, and you can not have your team at its throat.
And the result of doing that is you actually end up doing much more work with much less effort. And your customers are happier because you're finding the value, the thing that they really want by working in an iterative way. And I think these are really important things, and it is that upstream thought that enables it to happen.
And so that's my sort of section here on our experience of where people have gone and this sort of idea of looking upstream and focusing on cycle time and error rates and collaboration and how people have gone about working the first steps of that in their organization with the tool chains that they have and the organization structure that exists in their company.
So that's it, Beth. Yeah. Thanks, Chris, and thanks, Deepak. Those were both great presentations. So now we have some time for questions. So if you have a burning question, definitely enter it into the control panel box and we'll try to get through as many as we can in the next 15 minutes. So, here's the first question. So I'd like to know a bit on the lean process. Can we talk a bit on that? What lean principles are applied?
Deepak, do you want to start off with that one? Sure. I can. So, I think lean process has multiple factor into it. One, you surely want to get to more of a agile process, whether you want to use, say, for others, there are industry standard process. But more importantly, as we talk through the whole slideware and how do you approach it is more about leveraging tool and capability like DataKitchen and bring that collaboration so that you can automate most of the processes.
So lean is driving towards automation. Lean is driving towards how do you simplify the steps which are needed when you have built, bringing new data, whether you make it self-service or you give access to the capability so that they can distribute the data, download, as well as able to quickly set up. So what are the minimum steps you need for releasing something into production and how do you automate the testing, as Chris was talking about, is to drive that lean process. So it could be part of agile as a process you use, but more as a automation and then simplify the steps you need.
And that's the complication we all get involved. And also the point to make is this becomes more complicated. Surely we're looking for moving to cloud so that it can handle quickly, but because there are so many moving pieces across the hybrid platform, you want to simplify the steps, what you take and all, and define that process which works right for you.
And
DataKitchen can help you get to that process and how best to define that. Okay, thanks. Do you have anything to add to that, Chris? Yeah. If you think about my learning about this, there's a book that was published in the '80s called "The Goal" which is sort of a novelization of a lean transformation of a factory.
And there's a lot of aspects to lean, but one that comes to mind is something called the theory of constraints or find your bottleneck. And, in that manufacturing case, that they found a certain part of the assembly line, was the bottleneck that it was causing, and anything that you were trying to fix that wasn't at the bottlenecks was actually a waste of time.
And they talk about how they paralyzed the machine and did these fixes, and the factory started performing better. And I think that actually goes in software and as well as data and analytics, and you got to ask yourself, where's the bottleneck? And a lot of times the bottleneck isn't the data flow through your system.
The bottleneck is a person who happens to be
00:45:00
involved in technology review boards, checking that person who knows how everything works in your system, and you've got to ask them if you change something, is it going to break? And so bottlenecks or this idea of theory of constraints, which is a lean idea, I think is really important. And ask yourself, where's your bottleneck? Or who's your bottleneck?
And removing bottlenecks actually increases flow, the change at which you can get ideas from your people's heads into production. And so I think, that's another aspect. And leans, has different interpretations, but the theory of constraints and bottlenecks, and the book is called "The Goal" if you're interested in reading, it's actually a pretty good read.
And we have a follow-up to this question. Are there any A3 continuous improvement case studies identified as part of the lean journey?
Do either of you know of any offhand? Do you mean within the scope of DataOps or...
I'm not- I'm assuming that means, yeah. Yeah, within the scope of DataOps, I would assume. Yeah, I think in some ways, the case studies I presented are ways that people have focused on the problems of finding bottlenecks. And in a lot of cases, like I said, how do you get things into production? The bottleneck is a few small people who have the whole system in their head.
And you have meetings where they inspect the work, and they end up spending all their time in meetings and then get bored and frustrated and sometimes end up leaving the company. And so trying to look at the steps that it takes to go from the keyboard of your data scientist or data engineer into production and removing bottlenecks, automating, scripting, adding tasks, managing your environments are the process at which you get the flow of work from your data and analytic team into production.
And that's why this idea of really thinking upstream matters because a lot of people are saying, "Okay, let's go implement a new tool, a new database" and that doesn't really help. You got to focus on these sort of things like think of the manufacturing line of new code or new configuration and your tools and how that gets in, as well as the sort of manufacturing line that goes into production.
And just to add, I think the additional question was, do we see case studies in FS market in Europe? Yes, we do see banking and insurance clients in Europe, particularly trying to leverage this. Again, the purpose is how do you deliver the products, instruments quicker? How do we offer better insight? And using the agile and lean process, right, as we talked about.
So they are credential. If you reach out to us, we can talk about that. Great. Thank you both. So changing gears a bit, is DataOps mandatory for analytics and BI initiatives to be a success? I think just to add- You guys can take that one. ... I would say yes, it is mandatory now. I know we have been talking about this for some time. Gartner has also projected that as a key capability or key trend for the last three, four, five years. I think it is very becoming critical with the pandemic, the change in the need, and how agile and fast you need to come with new models, new products in the market to be competitive. At this time, this is becoming very mandatory for any of the data analytics team to leverage the data quickly. It's a vast amount of data, and as you're moving on the cloud, as Chris talked about, this is the time to take advantage of through this part of the whole process, instead of doing it, setting up a platform on cloud and then trying to figure out, how do I further become agile when I look at the hybrid platform? If you are on one platform, it may be still possible you could be using one cloud, but I don't think anybody's environment would be like that. You'll be on-prem, you'll be multi-cloud, you will have other SaaS services you may use from others.
All that need to work together as one integrated capability, and also, it's a must, it's the time we should all going forward. Great. Thanks, Deepak. Yeah. And Beth, could I speak to that as well? Yeah, sure. Well, it's not mandatory if you like to have these things happen to you. So, and I'll give some examples. If you want to sit on the edge of your bathtub on Saturday during your kid's birthday party, fixing a bug in a pipeline while your kids run around, then yeah,
00:50:00
by all means, it's not mandatory. If you want to walk out of a meeting with your business customer and feel them rolling their eyes at you because you just told them it's going to be six months, and they think it should be six days before they get something, then yeah, by all means, don't do DataOps.
If you want to go to another meeting where your data scientists and your data engineers are bickering and finger-pointing, then by all means, don't do DataOps. If you like going through your review as a manager and not being able to show that you've done anything for the last three to six months, yeah, don't do DataOps.
So, I think you should do it, because I've experienced all those things, and they suck. And don't live that way. And to me, I think it's the realization that there's just a better way to work is
helpful, and that other people have gone through the journey and other companies are seeing it. And so that upstream focus, I think, is really important. And by all means, I think everything that Deepak said, I was being a little cheeky. Everything he said is right. It is the trend and our people are doing it. And so, in that way, it is mandatory. So, Deepak, I was just trying to be a little cheeky there.
That is fair. All right. Great. Well, okay, here's a good question. Some of these measures such as... Wait a minute. Let me start over. Yeah. Some of these, such as measuring collaboration, it's a highly intangible measure. How do you quantify that in a matrix? How do you measure if the team gets really motivated to collaborate?
Would it be things like the number of meetings or peer development exercises being used to measure that? Chris, do you want to start with that one? Because you- Yeah. Well, first of all, individual productivity is hard, right? Because, I think there's different ways to measure team productivity. So for instance, if you do Agile, you've got story points or feature points, and you can actually look and aggregate over time at your Jira feature points and see that people are doing more of them, and they're doing more things focused on actual customer value than fixing bugs or being reactive. I think that's one way to actually judge productivity, and I think collaboration itself is also shows up in the fact that you are able to deploy faster, or the fact that you have less error rates.
And so for instance, a lot of organizations, as I said, the process of getting anything out that's new takes months because there's manual op meetings and hours. So yeah, actually measuring the amount of meetings you have and trying to minimize those is probably a good idea. Not to say that meetings aren't fun and we don't all love going to meetings, but most people who do technical work want to spend five, six hours a day doing technical work, not in meetings.
Deepak, do you have anything you want to add on that? No, I think that's good enough. Thanks. Okay, great. So, the next question is, as you can imagine, we have multiple sources, ETL technology, reporting technology, et cetera. How can DataOps and DataKitchen help visualize the technical lineage from report to source?
Sure. So I'll start, and Chris, you can add. As we are talking about DataOps, its purpose is to have a collaboration with the data provider and data consumers and establish environment across multiple teams so that they can collaborate and work together, right? So the idea is to bring this orchestration capability which can stitch everything together so that you can minimize the steps and the lean process, right?
But the idea behind the platform is not to solve every need of the capability. So this gives you flexibility, whether if you're using an ATL tool for your purpose, you can continue doing that. Or if you're using a data lineage tool, which are available multiple in market and most of the clients have been using it, you can continue using them. If there is a- point of view needed, how best to do that as a lineage and capability. But the idea is to use some of those best-of-the-breed tools which you are leveraging and bring them and stitch them together.
What we have also done is some concept called data fabric, which if you have interest, we can talk. How do you bring all the capabilities and metrics together and lineage? And then DataKitchen bring that value to showcase whether you're looking at the metrics and other aspects. Right? So that's the idea, is not to have one platform do everything in the world, but bring best of the breed and then provide you environment capability platform which helps you achieve that goal. Right? So I'll pass on to Chris, your point of view.
Yeah, I think, Deepak, that's great. And I look at it from process lineage, right?
00:55:00
If you want to avoid meetings, you got to have a thing that people work on, and you have these complex distributed systems that are broken across teams, and you have no way to understand the whole picture. Right? Think of a journey that the data takes from source to value, what tools, what systems, what happened when.
And so process lineage, in some ways, is more important than data lineage. And in fact, a lot of DataOps is really not focused on the data, it's focused on the steps and the tools and the process to happen. And if you're worried, I think about sort of data silos, I think the inverse thinking is focus on people and tool silos first. If you fix those, you're actually going to make it much easier to get your data silos together.
Great. Thank you both. So I think this question is for you, Deepak, is, do we have credentials already of DataOps in the financial services market in Europe? Yeah, I thought briefly I touched about that already.
But I think I see the need, so I'll reach out to John, and we can connect on this. Okay. And then lastly, has Capgemini and DataKitchen done projects together in financial services in Europe? Yes. We are working with multiple clients in Europe and pursuing number of discussion with the clients. We have done some POCs and pilots with them and going through that journey.
Okay, great. And we have just a couple minutes left, so I think this would be a good concluding question. Any quick tips on how to frame the benefits of DataOps for an executive audience?
Chris, you want to take that?
We're failing, and this is the solution . I guess basically my view is admit that you got a problem first. It's sort of like you're an alcoholic, admit that things aren't working. And say... That's where I think it's got to start. You got to start to realize that. And as a leader, the things that you own, which are the people and the process, are the things that you can affect.
And so to me, it's not about... That's the first start, and I think you would actually find out that your customers would agree with you. And that's really hard to hear, but that things aren't working. And so I would start there. And just to add, I think as I talked about some of the business value, right, the value what you need to show to client, things are not working, and Chris talked about the errors, the effort it takes to manage the platform, the time it takes to deliver, bringing new data and all. I think those are key metrics to show.
But more importantly, how do you become data-powered organization or insight-driven, which is the key demand for you to compete with any of the competition. So that's key business goals, so this can be very well tied to the business goals beyond the technology goals we all have for making the platform more agile and linked. Right? So those where you can start, and we can always help you to put that story together.
Okay. Well, this was perfect timing. We're at the top of the hour, and we hit all the questions, so that was really great. I just want to give an extra big thanks to Deepak, to you and your team for joining us today. That was a really great presentation. Chris, thanks as always. And thanks everyone who joined us.
I hope you got a lot out of this. I know I did. We'll be sending out the recording and the slides to you within the next 24 hours, so please be on the lookout for that in your email. If you have any other questions, don't hesitate to reach out to Deepak, to Chris, or I directly.
If you'd like to learn more about DataOps, we actually have another webinar next Wednesday. Chris will be talking about data observability and how to achieve that with DataOps. You'll get the registration link for that as well in your follow-up email. So, thanks again, everyone, for attending, and I hope everyone has a great afternoon and evening.
Thank you. Thank you, guys. Thank you.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Questions from this session
Why do cloud data platforms need DataOps?
A cloud platform gives you a powerful collection of data tools but no defined process for running them as a single system, so you get data integration without process integration. DataOps supplies the missing layer: cross-platform orchestration and platform independence, collaboration between data providers and consumers, self-service with governance, and automation that keeps operations lean.
What capabilities does a DataOps platform need?
Capgemini names six. CI/CD enablement through data and analytics pipeline automation, test automation with self-provisioning of test data, ML-assisted data mapping quality assessment, on-demand environment provisioning across hybrid environments, automated data lifecycle management and governance, and automated monitoring of data drift, data processing, and data quality.
Why is DevOps CI/CD not enough for data analytics?
Data analytics needs continuous self-service sandboxes, continuous meta-orchestration across many tools, continuous integration and deployment, and continuous testing and monitoring. It also needs an environment pipeline, coordination across teams and data centers, a common vocabulary, and process measurement, none of which a software CI/CD pipeline supplies.
What kinds of tests catch data errors in production?
Five kinds: traditional data quality tests, statistical process control, location balance tests, historic balance tests, and business-based tests. Run them automatically in production, on top of the whole toolchain rather than inside one tool, and keep a history of results so trends are visible. The payoff is fewer errors, more customer trust in the data, and less downtime.
What is a self-service data sandbox?
A sandbox is a prepared data analytic environment that a central IT or data group gives to a business team for short-term use, then monitors, governs, and takes back. A top 5 US bank built them for more than 1,000 non-IT users so that legal rules on data usage and lifetime were followed and usage was tracked. Ideas that prove out earn the right to be reimplemented centrally with recipes and tests.
How does an enterprise start a DataOps program?
Capgemini describes five stages: strategize by identifying business goals and defining the target state, engage by picking a use case that can demonstrate value, implement by running a proof of concept or pilot with a partner and platform, expand by converting the pilot to full-scale delivery and standing up a Center of Excellence, and scale up by consolidating the platform and establishing an enterprise adoption framework.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.