On-Demand Webinar · 1 hr 1 min

The Role of DataOps in Data Modernization

Cognizant's Jayaprakash 'JP' Thakur, Senior Director and Head of Data Modernization, joins DataKitchen's Chris Bergh on why DataOps belongs at the center of a modernization effort rather than bolted on at the end of one: continuous delivery of data and insight, faster transitions to new data ecosystems, and the foundation any agile data architecture needs. Recorded June 2021; updated August 2026.

Presented by Chris Bergh

What you'll learn 7 points
  • Cognizant reports that 90 percent of its data modernization customers are asking for CI/CD and DataOps pipelines, while roughly 56 percent of client data teams still manage their data pipelines manually.
  • Cognizant's DataOps work with a leading banking service provider cut cycle time from 12 weeks to two weeks by automating release orchestration, deployment, integration, and continuous testing on Pivotal Cloud Foundry.
  • A Cognizant financial services engagement onboarded 600 application projects onto an integrated CI/CD stack with GitHub and Jenkins, ran more than 100 application builds and deployments a day with no manual intervention, and increased release frequency by 35 percent.
  • Cognizant frames the problem of today's data and analytics ecosystem as the three Cs: complex, confusing, and costly, with too many tools, frequent pipeline changes, and people spread across the organization.
  • DevOps and workflow tools fall short on data work for six reasons: no end-to-end meta-orchestrated production pipeline, no environment pipeline, a process that is not DevOps CI/CD, the team and data center coordination that data analytics requires, no common system and vocabulary, and no process measurement to drive behavior change.
  • Moving to AWS, Azure, or GCP delivers a powerful collection of tools with no defined process for using them as a system. That is data integration without process integration, and it leaves the customer to design the DataOps superstructure themselves.
  • At a top 10 global health company, schema drift between an on-premises analytics team and an Azure cloud analytics team caused delays, errors, and team conflict until schema management, meta-orchestration, testing, monitoring, and versioning were put in place.

Slides

47 slides

Transcript

Show chapters and dialogue 9,757 words

00:00:00

Good morning and good afternoon to everyone today, depending on where you are in the world. Thanks for joining our webinar. My name is Beth Beverly, I'm the VP of marketing at DataKitchen, and I will be the host for the webinar today. So our topic is the role of DataOps in data modernization. We're very excited to have a special guest join us today.

We have JP Thakur from Cognizant, and he's here to share his experiences and insight. So if you haven't heard of Cognizant, they're a global leader in digital transformation and data modernization. And JP is a senior director and head of data modernization, cloud data architecture, and engineering for the communication, media, and technology industries. He's also the global DataOps leader and architect across all industries. He has more than 20 years of IT leadership experience in data engineering, cloud architecture, quality engineering, DataOps, and data analytics.

At various points in his career, he's also worked at Cognizant, Capgemini, Apple, and Microsoft. So welcome, JP. Thanks for joining us. He'll also be joined today by Chris Bergh, who's the CEO, founder, and head chef at DataKitchen. For those who don't know Chris, he's been a leader of the DataOps movement. He has more than 30 years of research, software engineering, data analytics, and executive management experience.

He's been a COO, CTO, VP, and director of engineering. He's also the co-author of "The DataOps Cookbook" and "The DataOps Manifesto." So welcome, Chris. Thank you. So, before I hand this over to JP to start, just a few housekeeping items. We are recording this session, and so we'll send out the video and the slides to everyone after the webinar as soon as they're ready.

So be on the lookout for those. We'll also reserve the last 15 minutes of the webinar for questions. So if you have any questions during the course of the webinar, please enter those in the control panel, and we'll make sure we get through as many of those as we can at the end. So, with that, without further ado, I'm going to hand it over to JP to take it away.

Sure. Thank you, Beth. Thank you, Chris. Good morning, good afternoon, and good evening, everyone, depending on your geographic location across the world. And thank you for joining this webinar today. So our agenda for today is shown here. In the first part of the session, I'll be mainly covering Cognizant's DataOps point of view and how it's important while thinking about overall data modernization. What are the DataOps best practices and guidelines that we have to look for? How our clients were facing challenges earlier, and how DataOps solution has helped them solve their complex data problems.

In the remaining part of the session, we'll talk about DataKitchen's point of view and how the Cognizant and DataKitchen together are helping our clients jumpstart their DataOps journey. So who is the webinar for? So any stakeholder in the data and analytics space who is really concerned about the continuous failure of projects and is frustrated by complex, confusing, and costly path to modernize their data platform would get lot of inputs from this webinar.

We hope this session will be helpful for all of you. So with that, let's move to the next slide, Beth.

Yeah. So today's typical data and analytics ecosystem and three key challenges. What are those? So if you look at today's data and engineering ecosystem, we are facing three key challenges because of the whole ecosystem is very complex, confusing, and costly. It's become a difficult data puzzle to solve, right? We have too many components, applications, complicated data pipelines, no or low data automation, people needed all over the places, many skills required to manage and integrate the systems, right?

Also, the data teams are struggling to extract the value from these complex data means. And B, data and analytics stakeholders have been involved putting this data means together, right? Started with the data engineering on-prem, then created many ETL engines, then streaming, machine learning, predictive, prescriptive analytics, and the surrounding system. We have built over time.

Now, the interconnectivity of all these are become very-- It's become like a everyday challenge, and you all may be facing these challenges, right? And after entry of this cloud and making hybrid environment work, it's really

00:05:00

becoming nightmare for enterprises, right? There are too many silos, hard to manage technology mountains. We are calling them technology mountains. And with the regulation, compliance, governance, privacy challenges are like a big deal, and people are really scared to touch anything that is working as of now. And the value we are getting out of these complex systems are very limited, very less, and very minimal, right? So there are plenty of dollars, business impacts, business benefits, positive influences are kind of left under the table because of this constraint. Therefore, we believe that this needs to be changed, right?

So in summary, how do we really get out of this complex ecosystem and establish a simplified, next-gen, fast, and efficient ecosystem? Okay? And the answer is the DataOps. Let's move on to the next slide.

So let's talk about data silos to DataOps. So we have data silos, complexities across the system, and we have three key challenges today, right? So how do we come out of monolithic, complex architecture and quickly develop and optimize the scalable and responsive architecture, right? How do we get rid of non-helping, lengthy business processes and get into smart and automated processes, right? And how do we eliminate so much of dependencies around and bring in machine driven development approach, right?

How do we simplify our existing and complex delivery methodologies and establish a super agile and automated delivery mechanism with better collaboration and greater synergies? Those are the questions, right? So the answer is DataOps. So let us get out of data silos, data complexities, and get on to DataOps journey along with us, right? So next slide.

So this is our kind of simplified reference architecture. So we call it kind of Cognizant simplified DataOps reference architecture, and we have defined DataOps as a practice that enables an automated orchestration of data pipeline by the process of acquisition, integration, transformation, and consumption to deliver and monitor data continuously with agility and assurance, right?

It's not just a technology change. Here we are talking about people, process, and technology transformation, right? So here we have to focus on three key aspects. The first one is establishing an automated data pipeline. Second is to set up an agile data team, lean process control, and governance around it. And third is that leveraging the power of DevOps to establish a robust DataOps platform, right? So this is a combination of three things, okay?

And you can say that, or you can imagine as DevOps brings in Dev, QA, and Ops team together for any application and services development. Similarly, DataOps bringing all data stakeholders together, like your data engineers, data developers, data architects, data analysts, data validator, data scientists, and the data business users, right? So if you are thinking about implementing DataOps, then you have to think about this architecture shown on the screen.

This is a simplified and optimized DataOps architecture, which can help you make your overall data engineering and operation very efficient and smart, right? And the important thing is Cognizant was the early adopter of DataOps framework and initiative, and this was included as a key component and was part of our Cognizant's data modernization offering.

A lot of our clients have taken advantage, a significant advantage of embracing DataOps while embarking their journey of digital transformation and data modernization, right? So we'll talk about those success stories in the next couple of slides as well. Next slide, please.

So why DataOps? So DataOps is new norms now. High speed digital transformation is now top CXO priority. And digital transformation is not just changing technologies. As I said earlier, it's not just creating the technology mountains, and current ecosystem and making it more complicated and polluted. Rather, it is about simplifying, optimizing, and making it more smart and efficient. So if you look at the industry status today,

00:10:00

90% of our data modernization clients are looking for DataOps processes for faster data delivery. 37% of clients are in acceleration path after adoption. And 21%, we have seen that they are making DataOps as a culture. It's a cultural shift. 1,000% increase in Gartner's analysis inquiry on DataOps. It's very, very high. 500% in Google Search on DataOps. 86% of companies plan to increase their investment in DataOps, and we have seen a lot of traction and interest in the industry. So these are some of the key statistics which are shown here. DataOps is a need of time, and it's a new norm.

And hope all of you are excited after getting this. Move to the next slide.

So we have validated, this is the Gartner's view and explanation. We'll not spend much time here. We have validated our DataOps offering with Gartner's understanding and definition, and it's fully aligned. Gartner also see DataOps as a potential space for new data and analytics possibilities. So these are just for the definition perspective we thought of, kind of be sharing with you so that you can see that it's not Cognizant, it's not DataKitchen, it's not industry. It's a leading industry analyst, also aligned with the definition and the methodology of the DataOps.

Okay, moving to the next slide.

Yeah. So the unified DataOps platform solution overview, we can say. So this is the overall DataOps platform view we have put together. If you have to establish DataOps in your data and analytics ecosystem, then you can imagine this view. This is the five-step approach, where data origination can be at on-prem or through various public clouds.

Especially we are focusing on AWS, Azure, and GCP cloud at this moment. Then you should implement mandatory agile way of project execution. Then you need to kind of have to automate. You need to have automation and CI/CD implementation and orchestration there. And then you think about setting up an automated data pipeline. And finally, you can imagine the high speed data delivery for advanced analytics, reporting, and faster insight. You can think to establish data marketplace or self-service analytics kind of capabilities, adopting data principles and methodologies for your internal users, external users, business users.

So that is how you will see the benefit over time. We have this joint view with DataKitchen, and you can see the DataKitchen logo there in the third layer. Cognizant can still leverage its homegrown accelerator called Intelligent Data Work and complement with DataKitchen's product for DataOps. So overall, you'll be able to establish a faster, flexible, and flawless data delivery for better data analytics and insight.

So this is the overall unified DataOps solution. Next slide, please.

So we have witnessed a lot of benefit of having Cognizant's DataOps approach and framework. Like if you establish a continuous data delivery pipeline, 50% of the cost of delivery goes down in next three years. 40% of productivity gain, we have seen bringing agile principles through this framework and collaboration engine. We have achieved 80% of automation across the data and delivery value stream.

30% and above operation efficiency gain we have seen. So there are a lot of benefits of kind of having this process, these methodologies, and this framework in your ecosystem.

The next slide, please.

So, we have many successful DataOps implementation, especially for the data and analytics clients. We are kind of just pulling here top three, our key implementations. One is for leading large financial organization. We have built a DataOps ecosystem for this client, starting from ground zero, and approximately 600 plus projects have been onboarded through this DataOps framework developed by Cognizant.

00:15:00

And we were able to achieve 35% increase in release frequency, 99.9% automated deployment also seen for this client. Number two is leading a banking service provider. We have implemented DataOps solution on Pivotal Cloud to automate the end-to-end development and deployment for clients' enterprise data platform. Automated release orchestration, deployment, integration, and continuous testing was done on this platform. And overall development and deployment cycle time was reduced from 12 weeks to two weeks.

So this was a great kind of achievement we have seen for this client. Third is the leading life science service provider. So we were able to modernize client's data analytics platform on Azure Cloud and adopted this agile DataOps processes across the line of business. And Cognizant team was able to do end-to-end automation of the data flow processes using DataOps solution and got approximately 40% effort reduction there.

So overall, a lot of benefits our clients have achieved so far, and it's a continuous journey. We have been kind of talking to many clients, many partners, and seeing a lot of traction in this space. And every client nowadays wants to save cost. Reduce their whole kind of time to market or time to production, and that is where the DataOps framework helps a lot. So with that, I will kind of stop here and invite Krish Werle, CEO of DataKitchen.

Krish will talk about DataKitchen's point of view in how Cognizant and DataKitchen together are helping our client to establish DataOps ecosystem. And Krish will talk in detail. Thank you. Krish, over to you. All right, thank you. And thanks for that great introduction to your point of view on DataOps and its role in data modernization and our partnership.

And so I'm going to back up a little bit because we've got a lot of people on the call, and they may not know exactly what DataOps is and how to think about it in relation to all the work that they do. And so we're going to take a little journey backward and sort of talk about the principles. And so, in my career, I've spent a lot of time doing data and analytics, focusing on predictive models or data pipelines or visualizing data or governing data or even the data itself. And they're all really interesting, and as JP said, there's an amazing amount of explosions in tools that you can use.

And so this, I would think, is a task focus. We're all kind of focused on getting the task done. How do I tweak my model or get the new data set? But unfortunately, that task focus, even though you think it's a good thing, it's actually ended up with almost a majority of the data and analytic projects are failing to meet the customer timelines or meet the customer requests.

And this has been going on now for decades. And in varying of whether you're doing big data or small data or smaller unstructured or cloud or on-prem. And so the promise of- Data modernization is that can provide the infrastructure to be enabled to do this. But I think the real challenge is that this task focus is not working, and that there's a bunch of problems that are really upstream to doing the tasks. Sort of the processes that you use to develop something, to get it from the hands of your data scientist or data engineer into production, how well they iterate, how fast they can change things, how well they can monitor things, how tested it is, how collaborative it is, how much you can measure these processes. And these upstream processes, I think, are the heart of what DataOps is about. So it's not really about an individual contributor being able to build a better model. It's about being able to have a whole system that supports that individual contributor, so that they can successfully create a model, get it in the hands of someone, learn, and then iterate and improve.

And that's really that upstream problem is about what DataOps is focused on. And so, some of the challenges that we see in this upstream is that a lot of people have, they build data systems, and as JP said, they're afraid to change, because it took so long to get it to work. And they're afraid to change it because in a lot of ways, their data providers or their systems are breaking, and so they end up with a lot of errors and scrambling to fix them.

And they end up being not wanting to change a system because it took them so long to build it. It's sort of fragile. And then people in data and analytics often

00:20:00

spend way too much time in meetings trying to figure out who owns this problem, how to get it done. And then, the irony of all of us data and analytics people is we don't measure a lot of our work products that we do. We don't measure the velocity of work or the cycle time or other important process metrics.

And so what, really, the sort of conceptual ideas that we focus on in DataOps is how to run your factory, the steps day in and day out, so that you have high customer data trust, so that your data warehouse, data lake, the analytics that are built up on it are error free. And so lowering error rates is a key point.

And then also this idea of cycle time, how fast you can get innovation into your customer's hands is another point in the process. And then the third is how you can have sort of less meetings, not only when you're working with a person who's a data engineer on your same team, but also how you can work across self-service teams. And then finally, how can we get analytic about the processes we use to build and deploy and run analytics?

And so these are the four upstream processes that DataOps focuses on. And notice they're not really anything saying it's about big data or small data or about what type of work you do. These processes, I think, affect all types of work, whether it's visualization or governance or data engineering, et cetera. And if we look at the example of this, and honestly, who wouldn't want to have this? Who wouldn't want to run their data and analytics without any problems in production? No one wants problems.

And who wouldn't want to be able to push a button and get something in the hands of their customer where there's no problems at all when you deploy it? And who wouldn't want to have a great dashboard that shows how awesome their team is working in front of their boss? And who wouldn't want to go to less meetings?

And so, I think by and large, everyone wants these things. And so how do people get at them? And I think let's go through some just concrete scenarios and talk about it. And so the first thing from a background perspective is everyone in data and analytics works in teams. It's no longer kind of a hero culture.

And maybe you're a data engineer, maybe you're a data scientist, maybe you're a self-service person doing BI or data governance, and these are kind of in general the categories. People have different titles associated with them. But all these people, they work with just a huge amount of tools. And so whether you're a data engineer using ELT or ETL, whether you're a data scientist using all the data platforms out there or Python, whether you're running on Redshift or Snowflake, there's 50 types of tools, all of which have great properties to them. And so it's a really massive fragmented tool chain.

And often companies will have more than one, so they'll have several different tools, even in the same department in the company. And so, the second is just even doing the day-to-day work between these teams requires a lot of coordination. So for instance, let's say a data engineer sources some data and they put it in a table, and there's a name and some sales amount.

And then the data science team takes that data once it's landed, and then they apply a predictive model to it somehow. And so maybe they are using FML rules or a random forest, but somehow they segment the customers, and they end up with high and low value. And so, one thing has to happen after the other.

And then the self-service team comes along, and maybe they build a dashboard, but they also sort of mix in more data from an Alteryx flow and kind of categorize these people by where they are geographically, the west team and the east team. And then finally, the poor data governance team's got to go along and sort of do catalogs and lineage on where it came from.

And so just to get what we would consider sort of basic additions to a data, loading the data, segmenting data, visualizing the data, and prepping the data, and then governing it, we've got a bunch of different people involved. And if you look at it from a different way, what's their development process? How do they work from their laptop or servers into production?

The data engineering team has one process. The data science team will have another. The self-service team has a very different process. And data governance and finally, the people who run it. And so we have these variants in how people actually do their work.

But our customers kind of don't care. They just care that it works. And so if you look at the data and analytics pipeline from another way, of course, we source data from a lot of systems. CRM, ERP. And we put it into lakes or into the cloud. But the day-to-day production is really that you run a factory, is that the data itself goes through a series of steps and it's assembled.

So it goes into a raw layer, it goes into a combined layer. Visualizations are applied, models are applied, a data catalog's updated. And if you think about it, and this is what I learned when I, after 15 years in

00:25:00

software, spent and for the last 15 years, I've been focused on data and analytics, is that these pipelines, these steps, are really a factory. And you want to have a factory that produces Toyotas but not AMC Pacers. So you want to lower your error rates. You want to give people control to stop the assembly line.

You want your factory to throw off statistics so you can learn and improve when something goes wrong. But this factory itself isn't so easy. It's not like there's one factory where it all goes. Most of these group is divided up among us in large organizations. Maybe you have a central IT team that develops a data warehouse, and a data science team in another part of the organization, and self-service teams.

And so, the end result is that sort of what's called Conway's Law, that the actual technical work here is broken on organizational divisions. And of course, if your customers aren't getting what they want or they see something wrong, it's a question of whose problem it is. And likewise, we also have a process to bring things from ideas to innovate into production.

And so, that process in data and analytics is sort of uniquely hard and is parallel to what happens in software. It is kind of a bit of a deployment process, but it's actually a lot more complicated. And the question is how do you properly regress or find out that your data and analytic work when you've made a change doesn't have any problems? And that is a real challenge because a lot of people will do their work, kind of throw it over the wall to someone else, and then weeks or months later find out it's problematic.

And so the cycle time of innovation is really important or deployment. And so the challenge we've got to do all these things. We've got to run a nice factory with low error rates. Our customers want us to deploy quickly, and we want to deploy quickly so they can learn. And then we don't want to sit in a lot of meetings.

And also we want to show how awesome our teams are. So all these things, we run this sort of value pipeline and innovation at the same time and have all these challenges. And if you look at this from kind of a different angle, look at it from a more technical angle, it's the same thing that applies.

That on the top here, you may have your production environment. And in that production environment, you have a bunch of tools, but you may have other environments below that, development, test, and you want to be able to automatically and continuously deploy and regress across those. You want to be able to handle the creation of environments to do that testing.

And so this sort of ability to architect for change, I think is an important part of what we're trying to do and an important part of our partnership. And that building, instead of thinking just production-only systems, thinking about how the sort of right to repair and the right to change And so, I think that's a great first step. And so now let's look in the final 15 minutes here.

Let's take those ideas. Okay, DataOps is about error rates and cycle time and collaboration and measurement. And yeah, we work in teams, and we have a lot of complexity in our environment. How have customers kind of gone off and done this, and what have they focused on? And I think the most important part is to be agile about your approach to data monitoring and DataOps.

Find an area to focus on first. Deliver some value, get it in the bank, and then keep working to improve. And if you look at it from an access standpoint, what we found is that a lot of companies are interested in all these things. Who wouldn't want to be able to get a model into production quickly or not spend six weeks deploying 20 lines of SQL?

Who wouldn't want to have less errors and fire drills in production? And who wouldn't want to have less meetings? And so they peck, and they say, "Okay, we're going to focus on one particular of these." And then if you look at this other axis, a particular feature or value proposition of DataOps is on the Y-axis.

And on the X-axis, it's kind of who does it, your database, your ETL team, your data science team, your BI team, your governance. And so let's just talk through one company. And so, this company's a transportation company, and they have data kind of coming off their vehicles, data in batch from internal systems like SAP.

And what was happening is they were having these fire drills every few weeks where two, three dozen people would get on a call after some vice president would say, "Hey, this report looks weird." And they'd try to figure it out. And so one of the challenges is that they built a very cool architecture that had streaming data in, they had a sort of fast batch in, had sort of batch in Informatica, had some models that were applied.

It all sort of ended up being marshaled and integrated into an Oracle data warehouse, and then sort of notebooks and visualization, and then also moving to the cloud. The organizational challenges, who ran these teams, were different.

00:30:00

They were ingest teams, who either were streaming or batch. There was an ETL team. There were IT teams who were in charge of getting the data. There was an enterprise data warehouse team that integrated it, a data science team, a BI team. And so, when something went wrong, no one sort of knew where it was. Was it the fact that the vehicle stopped producing data?

Was it the fact that the ingest routine somehow broke? Was the integration in the facts or dimensions of the warehouse, did somebody misconfigure the Tableau report? And so, that's a real challenge. And sort of if you build systems where you hope that it works, that they're entirely unobserved and unchecked during production, you end up with this situation where you have high error rates and low trust, and then you just don't know who to call. And it's Friday afternoon, and you want to go home and have a nice weekend, but you're instead running a bunch of queries on a phone call with two dozen people trying to figure out where it goes. And believe me, I've been there. It's not fun.

And so I think the solution here is to build something that sits on top of this and observes what's happening. And so, in DataKitchen, we call that a recipe, and it integrates to these tools, and it actually goes in and checks data. So as an example, your vehicles may be sending out 10,000 rows day in and day out, and all of a sudden, you stop getting 10,000 rows for one day.

That should be an alert. You should know that right away. It shouldn't actually end up going through the whole system, getting in a report for your customer, and then the customer calls you. Likewise, if a server goes down or someone happens to change the configuration of Tableau accidentally, you should know about this. And so by testing the data, the artifacts, and the transformations that are created from the data, going in and grabbing little bits of data and checking it, you can actually end up lowering your error rates by doing that because you've sent some alerts and some notifications, and you find out the problem before your customer does. And that's the idea here. Work on lowering errors, because you just don't want to come into the work in the morning and have that pit of your stomach saying, "Oh, I don't know if things are going to go work.

What fire drill are we going to have today?" And, I think that that is a worthy thing to do, and I think the end effect of that is you actually end up having more time to innovate and more time to actually create new value to your customer. And so, really these principles of kind of data monitoring, observability, testing, whatever term you want to call it, is that do it in production automatically, kind of on top of your tools.

And when you find something wrong or something weird, send a notification. And keep track of history, and make it easy, because to do this, you've got to sort of put tests across everything and make it easy that you can check. And if you do that, you get less errors and more time for innovation.

Your sort of mean time to failure and downtime goes down, and then your customer data trusts. And honestly, you end up with less stress and sort of less downtime. And so let's go to the next case. So maybe you're at a company and maybe your source data isn't bad, and you haven't had a lot of errors or maybe you've squeaked by without knowing them or things are good.

And this company actually put a lot of work into their

source data. But they also wanted to work on trying to kind of enable DataOps across the whole drug discovery. And so they had a challenge in that they have people using different tool chains in different locations. And so there's on-premise, there's GCP, Azure, AWS, so there are multiple tools, multiple places, people doing different things.

And so there was sort of no consistency and no sharing. And that ended up being problems because, again, part of the job was being done in one place, and part of the job was being done in the other. And here's an example. So there was maybe a team working on a large Spark cluster that had a best-of-breed tool chain, things like StreamSets, other tools, and then just kind of working with high-value drug development data. And then there was another group who were on Azure, and they're using sort of Databricks, and sort of data lake storage.

And they've got various research data sets, but they want to work together. They want to be able to have a process, a workflow that goes across both places. And so to do that, you need to have some degree of independence, right? Because the cloud team wants to make their changes and also the on-prem team wants to make their changes. But again, your analytic user sees the sum of these two groups. And so how do you tackle when there's data drift or schema drift, and how do you make sure that the cloud people don't complain that the on-prem people don't know what they're doing, and how do you make sure that you can find the problems

00:35:00

before your customer sees them? And that way both groups end up looking good. And the way to do that actually is to kind of what we call meta orchestrate. Orchestrate across all these teams, and apply tests, and assure the quality of the analytics by the sort of system-wide testing and monitoring. And then kind of enable a sort of a continuous delivery where people can make changes on their own but also make changes together.

And that gets really complicated, right? Mainly because of the fact that our clouds have just got a lot of great tools. And so if you want to do ETL or big data, you can use Databricks or Glue or Kinesis or EMR or Cloud Composer or Cloud Data Fusion, Data Factory, Data Pipeline, Data Prep, sort of pick your tool, and there's an entire chain of tools.

And that actually makes it incredibly hard because it's a really powerful set of tools on every cloud, and there's another version for on-prem, but there's no defined process to use all those. And you end up moving data, but you don't actually have an overall process to run it. And what's really needed is some kind of superstructure to sit all over that.

And one of the challenges that a lot of organizations have is they look at DataOps and go, "Ah, it's just DevOps. And so we'll just use our DevOps tool." But the problem is that sort of CI and CD is not enough. And to really properly regress a data and analytics system to find errors in production, you need that sort of end-to-end meta orchestrated view.

And to actually really do that, you need to have an accurate development environment with good test data. And to do that, sometimes you need this complex team and data center coordination. And then lastly, data engineers and data scientists aren't software engineers, and so you need a common system and vocabulary. And then lastly, I believe you need to measure your process.

And so there's some gaps in trying to say, "Okay, I could just use my DevOps tools and workflow tools." And unfortunately, all too often we see people fail. It's a good step in the right direction, but it's not enough. And so that's another reason why we've solved these problems with having our software platform that can do this.

It's really a superstructure for data modernization to enable you to modernize, the ability to go from on-prem to the cloud, but do this in a way that fixes these problems.

And so let me just go to the conclusion here. And so why does this all matter? Well,

a lot of organizations, and we've done surveys with analyst groups and other companies, including Gardner, have written about this, the cycle time to deploy a new data set, a new data transformation, a new model, is really quite long. And some companies are months before they can get something into the hands of their customer. And oftentimes companies, when they have things in production, they're ending up with, it's not that it's error-free years, it's what went wrong today, what problem are we going to fix?

And in fact, a lot of people have moved into the data science and engineering field, and they're finding it frustrating because they're spending a lot of time on meetings and documentation and not getting things done. And then if you ask people about your boss to show how well your team's doing, there's not a lot of measurement. And so a lot of these things end up where the teams are much less productive and much more costly, and your customers are unhappy. And so, what we're saying in the process of doing DataOps is that you can improve on these. In fact, you can improve all of these simultaneously.

And one of the challenges is that if you've been in the data and analytics industry for a while like I have, that we're saying that the DataOps enables you to actually make well-tested, won't break production changes very quickly in whatever part of your data and analytic work that's doing the data transformation and visualization model, and you can run your data assembly line with completely no errors. And so people who've been in the industry for a while push back on that quite a bit because they say, "We're going to change things a lot and I'm going to have low errors. I just got this working.

It's so fragile, I don't want to touch it." And I think the processes and software that we provide help you do that, in fact, actually help you improve with collaboration and process management. And so I think best-in-class companies are doing this. They're able to push up on all these levers. And if they do, actually, their teams are just incredibly more productive.

And what that means is there's lower costs, and your customers are a lot

00:40:00

less unhappy. And so you can do it all with DataOps. And this similar transformation to being able to have best-in-class companies do everything is very similar to the process that happened to software companies. And best-in-class software companies are able to do this. They're able to deploy thousands of times a day with no errors.

They're able to have their teams work well together and enable their teams to be highly productive. And so, in some ways, that's the goal that we want to have and why we've partnered with Cognizant. And we think that the idea of modernizing your data infrastructure, moving to the cloud, trying to be able to apply DataOps principles to part of that, is really a way for us to be successful. And I think the key takeaways I'd like you to take is that DataKitchen's been, and Cognizant have been kind of leading the way here to the industry to get them to do DataOps.

And we built a software platform, and together we've really got a lot of domain and industry expertise. And we've, like this, we've got a lot of thought leadership, and we can get it done. And so this journey to get your team on DataOps, we've got a good partnership to make that happen. And so, that is all the slides I have today.

And I appreciate you listening, and I'm going to turn it over to Beth, and she's going to see if there's any questions, and then we'll answer your questions and then finish. Thank you very much. Yes. Thanks, Chris and JP. Before that, just two points I just wanted to mention here. As Chris said, the Cognizant is also a global data and analytics leader, known for best-in-class limitless data and analytics services for its line and recognized by the top industry analysts.

And Cognizant and DataKitchen's DataOps solution orchestrate data-to-customer value. It speeds up deployment to production, and our flawless automation across the data pipeline makes it kind of unique offering in the industry. And second point I wanted to mention is, I stated this before as well, Cognizant is an early adopter of DataOps methodology, and we have got a level of maturity in this space already.

We have established a strong partnership with the DataOps industry leader, top leaders like DataKitchen, and we offer a range of options to implement DataOps solution for our client. So this is what I just wanted to tell the audience, Beth. You can go ahead now. Yes, no, that's great. Thanks for adding that, JP. So now we do have a few minutes for questions, so if you have any, definitely enter them into the control panel, and we'll get through as many as we can in the next 15 minutes. So, just to start it off, so Chris, you emphasized the need for collaborative DataOps lifecycle management tools with self-healing automated mechanisms placed at heartbeat cycles.

Could you please elaborate more on the unified platform you're using, and how efficient is the integration? Maybe both of you could take a stab at that. Yeah. Well, I think what's interesting is it's not about-- What we mean by integration is not really data integration because there's plenty of tools that you have to do that. It's really about the tools that act on data integration.

And so most companies will have a whole set of ETL, ELT, data science, data visualization, and governance tools, and we're trying to integrate across those to give that end-to-end visibility. And also the ability to understand when things go wrong, both in terms of when you're developing something, that you can look at one point in a system and see its downstream effects of a change, but also in production, where you can see if there's a problem at any point in time.

And so I think that's one of the reasons why we built the platform and why integrating to all your tools. But I think the key point of the partnership is that people who have Cognizant have industry expertise in analytics, and especially industry expertise or especially specific expertise in your company can help make that transition faster.

Great. Do you want to add anything to that, JP? No, I think Chris answered very well. Okay. Great. So the next question is, so data science models move the data into the pipeline through modeling and to the data analyst. Where would the orchestration tasks fit in with all that's going on in the data pipeline?

I don't know if that's-- Is that a Chris question?

00:45:00

Go ahead, I'll take- I'll take a first shot. Well, what's interesting is that no model is an island. And so models are fed with data, whether they're batch or real-time. Models have a delivery mechanism. And maybe it's a batch segmentation that shows up in a charts or maybe the model itself has a API on your website. And so that's the first principle, in fact, well, is that sort of system-wide view of your data and analytic processes.

And so,

fortunately or unfortunately, depending on your perspective, the data science team is in charge of models, and maybe the data engineering team is in charge of data, and maybe the data visualization team's in charge of visualization. So they have three different teams. And so the trick here to make this work is not to just see yourself as an island and focus only on your model. You really need to work across all these teams.

And I think the idea here is not to have more meetings, or more Word documents. It's can you give a common shared abstraction that everyone can look at and say, "Well, this is working, this is not working," that the data engineers and the data scientists and everyone else in the process can see. And I think that eliminates a lot of meetings because it raises what's not visible to visible. And it also allows you to say, "Okay, I tweaked this model.

Now, in a development environment, I caused a regression in my- ... data visualization layer, or I tweaked the data that's going into my model as a data engineer, and oops, I broke the model. See that in development, see that in that common shared abstraction across all those people. And we actually have a name for that, we call it a recipe.

And so that I think leads to more timely execution with the teams because they're looking at something together, and they're not just sort of talking about it up in the air or through Word documents. They can see it living and running. Yeah. And just to add on top of that, so as Chris said, right, there are the three different teams that involved. Your data engineering team, your visualization team, your data scientist team. So some of the clients we have seen, data offsets with the data engineering capabilities, data engineering kind of stream.

Often clients have started thinking about the MLOps kind of capabilities, though DataOps and MLOps are interconnected. You can't imagine to have the MLOps without having the data. You can do it still, but the benefit you'll be getting, it not be that great, right? So some of the clients we are seeing, they start with the data engineering and they end up with doing the data scientist activities with the model provisioning, deployment of the model, versioning of the model, validation of the model, and productionizing those models.

So some of the clients are doing DataOps, and calling, it's a unified DataOps, starting from data engineering to data scientist and AI ML activity. Some of the clients have started saying DataOps or data engineering, and then for data science work, they have started adopting MLOps kind of category. So that's how the trends are in the industry.

Yeah, as long as you put the word Ops on it, it's good. DataOps, Model Ops, MLOps. You just got to put the word Ops at the end. Yeah, and just to add on to this question, right? This is Sagar, part of Chris and JP's group. So as you mentioned, right, how the data science model moves the data into the pipeline and through modeling and into the analysis, right?

So it's the same process, this entire process will be orchestrated with our CI/CD pipeline. So as soon as the data science model moves in data or any changes, right? The data pipeline will orchestrate through the CI/CD, which in terms kicks all the process controls, branching and code strategy and code reviews. And then it will move through the data pipeline triggers via the orchestration, and it will make this data available to the end user, whether it is data analyst, data scientists, or ML data people through all these continuous testing and all controls. So it's same as what happens in the data pipeline, how the overall CI/CD orchestration happens. So similar way, it will happen for any MLOps or data science model wherever there is a change or any triggers.

Great. Thank you all for that one. So I love this next question. JP, maybe you can take this one or start, JP and Sagar. If I already have a DevOps tool, can I afford to have both a DevOps and a DataOps tool?

Yes. So our team is to kind of simplifying the overall data delivery kind of ecosystem. So now what we have seen, the clients have kind of created the tool mountains in their data analytics ecosystem, right? They keep buying the tool, keep implementing the tool.

00:50:00

So first of all, we have to really simplify whatever the tool we are using right now from the DataOps perspective, and then optimize those. And definitely the DevOps, its focus is more on the delivery value pipeline, right? Your for your applications, your version control, your CI/CD, your deployment, your automation, and things like that.

So we need to have the DevOps in ecosystem, and then on top of that, we encourage to set up the Data pipeline, right? And Chris talked about the meta orchestrator, the tool from the DataKitchen. So we need to have DevOps tools in, optimize and kind of, those optimized DevOps tools in. On top of that, we need to have the DataOps tools in. So it's a combination of both.

Great. Chris, do you want to add anything onto that? Well, on a good day, I'm really happy that people are even doing DevOps, applying, because most people are just running their data pipelines wide open and doing everything manually. But, I do think that if you want to succeed, you won't do it with just DevOps tools.

For the reasons that I laid out, they're insufficient to be able to handle the complexities of large data sets and data analytic people. And so, I applaud people who are taking the first step because that's a good thing. But on the other hand, you're not going to succeed.

Okay, great. So now getting into a little bit of detail. So when the end product is a dashboard and many teams are involved to make that end product, such as ETL teams, modeling teams, visualization teams, how do you define a product and manage the CI/CD aspects of it? Do you deploy the entire ETL code for every release, or is there a better way to organize your code into products?

JP, do you want to take that one? I would pass it down to Chris, and talk about a little bit meta orchestrator and then how that framework can be leveraged here. Yeah. So, I think one of the most important things is that you have a way to see all those tools and how they work together in one place, and then see if they're, A, running correctly, and B, if the data is being transformed and put in that report in the model in the right way. So that means test it.

And so that's one thing. And then the question also comes along is like, okay, you have all these people, how do they actually be organized? And so that's a different question. And so I think you can be successful with DataOps while having teams organized in functionally as data engineers and data scientists. Although there is an emerging movement to sort of group people together under data products, where it's a combined team of data scientists and data engineers and people who do visualization. And that is often called, or it's called now a data mesh, where these teams are grouped together around some functional set of data or some specific customer group. But either way, if you organize your team functionally, across roles, or you organize them by data product, either way, the ideas of DataOps are sort of end-to-end, being able to see all your tools so you can have accurate regression tests, so you can deploy quickly, so you can run an errors with low production.

When run in production with low errors, I think all those ideas apply.

Great. Thanks, Chris. So do you have any case studies for folks who received some light benefit from doing DevOps CI/CD, but then needed to take it to the next level with full DataOps CSMOIDEM? Chris, you'll be pleased to hear someone's been listening. The world's worst acronym, someone's using it. So as well as cycle time, throughput, quality metrics, post CI/CD, then post CSMOIDEM. So who thinks they can take that one?

Some case studies on that. I actually know a case study. So we're working with a sort of top 10 insurance company, and they're on their DataOps journey, right? And so, six months ago, you could think of their path to bring new changes to their ETL process and some of their visualizations, like it was unpaved.

And so they had a lot of meetings, just moving the work, the change in their ETL code and biz code. And what they did is they put in, they actually used the CI and CD tool. And so they sort of paved the road, they made a railway. And it got it from

sort of 10 weeks down to about six weeks in terms of the speed. But they're still at six weeks, and what's the reason? Well, they went from a dirt road to railroad tracks, but they have no signals on the railroad track, so they still got to, even though they can automatically move

00:55:00

the rail car down, it can't go very fast because they can't tell if they're going to break anything. So they don't have any automated testing and regression, the stuff that we've talked about. So I think that idea of end-to-end visibility, meta orchestration, environment management that we talk in DataOps is needed. Because it's not just about movement of code, it's about successfully proving that that code works in every environment based on realistic test data.

So you can have a fully regressed system, so you can add signals to that infrastructure. And so those things are, that's why sort of just CI and CD isn't quite enough in the data analytics world, and we invented this world's most terrible acronym, CSMOIDEM.

Yeah. And just to add to that, so a lot of Cognizant clients have started adopting this end-to-end automation and automation from day one. So nowadays, when we are talking about data modernization, and I stated earlier, 90% of our clients are interested and showing excitement about automating from day one. And it's not just the automation, it's the proving something, putting something on production very quickly.

Days are gone when we have to wait for the two months or three months of one feature to be released, and then one new kind of modular option should be available for your end user after six months. Every single day, how do you want to deploy something on production? How quickly, how efficiently, and how accurately?

That's where the automation testing comes into the play. So a lot of Cognizant clients are also showing this kind of interest, and we are doing these kind of things with them.

Yeah. Great. Thank you both. So now I'm going to change topics a little bit for this next question. So could DataOps

be part of data governance, or would it be independent? This is considering that data governance is a wide area, part of a lot of internal processes of a company.

Yeah. So let me start with this, and then I'll hand it over to Chris to head on top of this. So all data management capabilities, right? Your data governance, data quality, lineage, cataloging, all kind of complement DataOps ecosystem. So data, as Chris said, DataOps is not just doing CI/CD for data. It's much beyond than that. So how do the other data management capabilities, like you said, data governance, data quality, data cataloging, data lineage, how do they complement with that particular CI/CD pipeline and ensure that the smooth, accurate, and clean data is going on production that has to be there? So data governance is, yes, is a part of it.

There are a lot of clients are keeping data governance as a separate kind of stream, but eventually they have to converge and they have to go together. Why did mine is successful data products. Chris, you want to go it? Yeah, I agree. And, we think that DataOps is a great addition to data governance. And if you just think of it, one aspect of data governance is, as you said, there's a data catalog and that has data lineage.

So you can go to a place and say, "I want to understand the data, and I want to understand where it came from." Now, that's great. It's just what people who use data want to know two other things. They also want to know, can I trust it, and is it fresh? And those are also equally important things, too.

Can I understand it, and where did it come from? And that idea of is it fresh and can I trust it can actually be built from a DataOps system that sits next to it. What tests were applied, when it's last updated, when it's scheduled to be updated. These things are of incredible importance to people who lives are interested in analyzing data or utilizing data.

And so, likewise the idea of governance as code, and when you make a change to a database table, a database DDL, how does that register, and how does that change at once with your data catalog and governance. And I think these ideas of automation and as code, augmenting a data catalog with data test information and data update information, I think is a good adjunct and an addition to data governance.

Great. Thank you both. Well, I hate to say it, but sadly, we are out of time. We actually did not get through all the questions. There were a few outstanding ones, so if we didn't get to your question, we will follow up with you individually after the webinar. So I want to thank everyone for taking the time to join us today.

A big thanks to JP for speaking and sharing your insight, and to Sagar for helping with all of this behind-the-scenes effort on this webinar, and then to

01:00:00

Chris, as always, for joining us and sharing your insight. So to all the attendees, we'll be sending out the recording and the slides in the next 24 hours, so please be on the lookout for that email. If you have any additional questions, please don't hesitate to reach out to Chris or I, or we can put you in touch with JP or Sagar for any other questions directed toward them. So again, thanks everyone for attending.

It was a great webinar, and have a great afternoon. Yeah. Thank you, Beth. Thank you, Chris. Thank you, Sagar. Thank you. Thank you, everybody. Bye. Thank you. Bye.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What is DataOps?

DataOps is the practice of automating the orchestration of a data pipeline through acquisition, integration, transformation, and consumption, so data is delivered and monitored continuously with agility and assurance. Gartner describes the goal as delivering value faster by creating predictable delivery and change management of data, data models, and related artifacts. Just as DevOps brought development, QA, and operations together, DataOps brings the data stakeholders together.

What role does DataOps play in data modernization?

Modernization projects replace the platform but usually carry the old process with them: errors in dashboards and pipelines, slow deployment of new features, poor collaboration across distributed teams, and no measurement of productivity or SLAs. DataOps addresses those four upstream processes directly, so a move to the cloud reduces error rates and cycle time rather than reproducing them on new infrastructure.

Why do DevOps and workflow tools fall short on data projects?

Six gaps. There is no end-to-end meta-orchestrated production pipeline across the toolchain, and no environment pipeline. The process is not DevOps CI/CD. Data analytics requires team and data center coordination that workflow tools do not model. There is no common system or vocabulary across teams. And there is no process measurement to drive a change in behavior.

What results have DataOps modernization projects produced?

Cognizant reports a leading banking service provider cutting cycle time from 12 weeks to two weeks, a financial services provider running more than 100 automated builds and deployments a day across 600 onboarded application projects with a 35 percent increase in release frequency, and a life sciences provider projecting a 40 percent effort reduction on Azure.

What tests and monitors reduce production data errors?

Five test types carry most of the load: traditional data quality checks, statistical process control, location balance tests, historic balance tests, and business-based tests. They have to run automatically in production, sit on top of the entire toolchain rather than one tool, send alerts, keep history, and be easy to create.

What is meta-orchestration?

Meta-orchestration runs an end-to-end pipeline across the tools that each already have their own scheduler, spanning the database, ETL, BI, data science, and governance layers and spanning development, test, and production environments. It sits alongside storage and version control, history and metadata, permissions, environment secrets, and automated deployment as part of a data architecture built for change.

Where to go next