On-Demand Webinar · 54 min

Optimize Data Privacy and Security with DataSecOps

Peter Lancos and Sonal Rattan of eXate Technologies join Chris Bergh on the what, why, and how of DataSecOps: how it relates to DevOps and DataOps, how it resolves the conflict between data privacy and data access, and how it simplifies test data management. Recorded September 2020; updated August 2026.

Presented by Chris Bergh

What you'll learn 6 points
  • DataSecOps, as eXate presents it in this session, bridges the gap between operations, governance risk and compliance, security, and data teams, the same move DevSecOps made between security and development teams. The aim is to automate privacy by design and cyber controls into both new and existing applications.
  • The list of parties with a claim on how data is used is long: chief data officer, chief information security officer, data protection officer, compliance, legal, technology, operations, the person the data belongs to, and the jurisdiction the data belongs to. Today that policy gets interpreted and applied by the data engineers and data scientists, which is where it breaks down.
  • The regulatory map is the reason automation is needed. California's Consumer Privacy Act came into force in July 2020 and Brazil's LGPD in August 2020, alongside GDPR in the EU, Singapore's PDPA, South Africa's POPIA, Australia's Privacy Act, Canada's Digital Privacy Act, and drafts in India, Chile, Argentina, Kenya, Uganda, and New Zealand.
  • eXate describes its product as an aggregator of privacy enhancing technologies with governance and controls embedded in the DataOps pipeline, covering PII discovery, access management, record retention, data destruction, right to be forgotten, subject access requests, and audit of use, on the argument that no single protection technique fits every case.
  • The first DataKitchen use case is deploying data security rules as code. Authentication and authorization rules move to production at the same time as the new schema, new data, updated model, updated report, and updated catalog they govern, instead of being applied by hand afterwards.
  • The second DataKitchen use case is test data management. Data and analytic teams need test data that is recent, accurate, reflects production, and does not break local privacy rules, and the three challenges are distribution, meaning the time to operationalize it; quality, meaning high fidelity; and security, meaning private information such as credit cards and medical records.

Slides

38 slides

Transcript

Show chapters and dialogue 9,448 words

00:00:00

The introduction. So our agenda for today is I'll do five quick minutes on what is DataOps, and then I'll hand it over to the folks from X8. They'll do the majority of the talking, and then I'll talk a couple of use cases. And so this is set as an introduction to what is DataOps and where does DataSecOps fit within it. And so,

I think the big idea in DataOps is what you do is much less important than how you do it. And a lot of data in analytics, the how is treated as secondary or something less. And I think security fits in with that, as well as a bunch of other ideas. And one of the main concepts here is that there's a machine that makes the machine.

There's the system at which you produce your analytics. And the characteristics of that system is it should be able to produce analytics that are quick, that are high quality, and of course, are safe, and that's where security comes in. And we spend a lot of time in the world of data analytics talking about models and algorithms and pipelines and visualizations, but we don't spend a lot of time talking about how you develop, how you deploy, how you monitor to make sure it's secure, how you iterate and collaborate. And so for us, our perspective is that the people and the process and the operations are actually much more important in some ways than what you do, the tools and technology and data.

And so it has this, what DataOps has this contrarian perspective associated with it. And so when you think about it, well, what does that mean? What do you focus on with this contrarian perspective? Well, I think we focus on four things. One is you want to decrease the cycle time of change. That is, can I make something and put it into production and know that it's going to work, and, of course, know that it's secure and that nothing's going to be wrong? And then second, lower the error rates in production.

As I'm producing an analytic, is it wrong? My customer's not going to trust it. And then, since a lot of data in analytics has a lot of different people involved, and one of the reasons the DevOps movement has started to use the term DevSecOps is that there's a challenge between development and operation, just like there's a challenge between development and security in operations.

And so how can everyone collaborate, the people who are building the database, the model, the operations people, the security people, and how all this collaboration can happen. And then finally, how you can get analytic about your analytics. And so we've coined this term DataOps, which is really the solution to that problem, how teams can measure and iterate and deploy quickly and measure what happens.

And so we've written a book, and I've talked quite a bit about it, and it's becoming more popular. And so, what does that mean? Just to be concrete. So the two perspectives that we want to share is one is that when you're doing data and analytics, it's a manufacturing process, and data's coming in and value is coming out, and you want to run it according to Lean techniques and have low errors across all those different tools that you use. And that those pipelines themselves are kind of owned by lots of different people in the organization, maybe your IT team, your data science team, et cetera. And so you want to run a good Toyota-like factory producing analytics. But then you also want to be able to pick up a piece of that factory and change it, and be able to make changes to production, be able to deploy it quickly.

And so you've got a lot of things that you're doing here in data and analytics. You're having to run a production with low error rates. You're wanting to pick up a piece and change it. You're wanting to collaborate across teams. You're wanting to get analytic about your analytic process. And all these things are actually pretty hard to do.

And so that's where this idea of DataOps can help you accelerate your way to do this. And hopefully, at the end of the day, you deliver more value to your customer and are more customer-focused. And so, we've been talking about this idea for six or seven years. It's gained some traction. And this term Ops is, there's a lot of Ops terms out there that you may have heard, DevOps and DataOps and SecOps. And just to kind of provide some clarity on that as I end my start is that I think there's a general way that people manage organizations where they all share something technically complicated that they're working on.

Maybe that thing that is technically complicated is a manufacturing line, maybe it's a complicated piece of software, or maybe it's a data and analytics production pipeline. And there's a set of techniques that have been developed over the past 100 years that stem from Deming and the Toyota quality methods, Six Sigma. And then they transfer down into how you run a team. And you may have heard of Scrum or Kanban or Agile or Safe Scale Agile. Or you may have heard of manufacturing, things like Lean and Six Sigma and total quality management.

00:05:00

And those are how you manage the people. But then the organization that you apply those principles to have different words. So in IT and software teams, there's DevOps and DevSecOps, and now there's terms like GitOps and AI Ops and Cloud Ops, and there's a lot of Ops terms out there, and perhaps you want to roll your eyes at all of them.

But I think they do have ways to describe the technical environment and process. And likewise, we see DataOps as kind of a broad term that covers how you apply these ideas across all the work that you do. And whether you call it DataOps or DataSecOps, we think that there's just a lot of different areas that you can apply it to.

And so what we're going to talk about today is really about DataSecOps, how you apply security in the process of delivering analytics with high quality and low errors across all the people and processes that you do. And so that's my setup on this. And I tried to promise five minutes, so take it away, Peter.

Thanks, Chris. I appreciate it. Thank you for a very good introduction. I am going to change myself to back being there. You should be able to see my screen now.

Yeah. Yep. Great. Well, thanks again, everyone. So I'm back. Round two on this one here as we go through again. So we are X8, and we're here to talk to you today about DataSecOps. So what we've done at X8 is effectively created a DataSecOps platform that's comprised mainly of two parts. The first part is how do you manage and automate your governance and controls to be able to deploy within your DataOps pipeline? The second one is to build an aggregator of privacy-enhancing technologies.

So when people talk about privacy-enhancing technologies today, people talk about sharding, they talk about consent-based privacy, they talk about zero-knowledge proof, they talk differential privacy, homomorphic encryption, various other forms of encryption. So we've built an aggregator of all of those, and we'll show you very shortly how they all fit together. But first, about Sonal and myself, I will put my screen on just for the introductions. I will put my camera on quickly.

I will take it off shortly after this. Whilst everyone wants to hear about DataKitchen, they don't want to see Peter's kitchen. But just for this part, we'll keep it on. I'm Peter, I'm the co-founder with Sonal. As Beth said, my background is actually from a banking background. I spent most of my career on the business side, but dealing with tech issues, both from a capital markets perspective as well as a digital transformation perspective.

So we spent a few years in how do you digitize a large bank from a commercial banking and investment banking perspective. And we're going to tell you some of our lessons learned on that and where the whole concept of DataSecOps came from. But before we do that, I'll switch it over and let Sonal introduce herself.

Yep. Hi, I'm Sonal, and I'm the co-founder alongside Peter for X8. So where Peter mentioned that he was spending many years in banking doing a variety of roles, in those roles, I was his tech person. So essentially I ran a team of engineers that were looking at solving real business problems, typically around regulatory things, where swap dealer licenses, all of these things were under scrutiny, but we just couldn't get the data to do things. So a lot of the things that we're talking about is where we saw the entire market moving, where as more and more cyber risks come into play, how it actually impacts people on the ground, and how it ended up causing more problems for us.

Creating applications was a lot more simpler when we had access to data, but as you need more and more data to supplement it, that's where we found it really incredibly challenging. And that is a big part of why we believe DataSecOps is the right answer to be able to solve some of these things.

Thanks, Sonal. So in summary, to really, in a nutshell, pull together what it is we do, is we consistently and repeatedly protect data to allow you to eliminate some of the challenges that you typically will see in trying to get your job done. Now, actually coming from a corporate, I always have to start off with a management summary. It's just a habit. Just to give you a very high-level overview of what we're going to discuss today. So I'm going to spend a little bit of time to start with just talking through digital transformation, just very quickly, a bit of a background to set the scene on the reasons why people look to become digital and the challenges with respect to that, and talk a bit about the problem we face, which is how do you get access to the data in a controlled manner?

So from our perspective and our digital journey, working for a large global bank, it wasn't as much as building the apps and

00:10:00

digitizing that was the problem. Where we ran into the problems was more once you start going on your digital journey, you realize that you're taking data out of their nice, safe, and secure source systems where they were all fine and everything was relatively stable. And you're using APIs to pass data all around, and you're starting now to put that data at risk. So the cyber risks and the privacy risks and security risks that Sonal mentioned. So we'll talk a bit about how those are addressed.

And lastly, we're going to spend quite a bit of time on talking about consistently, repeatedly, how does DataSecOps solve the challenges that you will see and give some concrete examples for you and discuss what you might want to see in a DataSecOps platform of how that could be incorporated into your business. So to kick off, firstly, a bit of background.

So digital transformation is a very generic word. A lot of people are using it, especially now in the COVID world. Everybody's wanted to be digital, but now they really want to be digital. So you're still seeing some budget addressed in this area. And the first place that we're seeing is cloud migration is a big issue.

So from a cloud migration perspective, when we speak to people, lots of people want to use the cloud, especially now, because the resiliency, the low cost, everything we talked about, the working from home. But we repeatedly are seeing clients that we deal with not quite ready, especially in regulated industries, to put all of their data on the cloud. And we're still seeing people who are putting non-sensitive data on the cloud, but holding back those sensitive bits.

And this is one of the challenges that you see with respect to this. Next is API adoption. So this is, again, as I mentioned, you're passing this data all around via APIs and creating this new transformation. But while you're doing this, you have to look about how do you do this irrespective of device or jurisdiction, et cetera.

But also what we're seeing is now people plugging our solution, a DataSecOps solution, into an Apigee layer. So governing very quickly, lots of APIs in a very short amount of time. So in order to get that API transformation, you're now looking at a proxy level about how you can address things very quickly. And lastly- One topic that we're quite keen on is hyper-personalization.

We're seeing more and more in this on the people we speak to. And this is due to a change in the way that consumers view the world, and where a lot of people we speak to are looking to go is more the Amazon type of solution. So when you think about how Amazon is reaching out and driving sales, I don't think anyone actually goes on Amazon and searches through 20 pages of Amazon looking for something to buy.

You're typically going to get an email saying, "This is what people like you buy," and, "Because we know all about you," and, "Here's the two to three things you're likely to buy. Which one do you want?" And that is becoming more of a common way to be sold to. A lot of people like to be sold to that way.

So this is the new economy that we're seeing, and that new economy is the information age. The information age is where we are heading to, and when you look at a Gartner report that came out, they're saying in 2022, that's not very far away, companies will be valued on how they use data and on how they monetize that data. It's going to become a part of the valuation of a company. So while that's all fine and dandy now, 2022 is not that far away, and the firms that can unlock their data quickly will be the ones that are the winners in the information age. Now, when you look at the information age as we see it, the race for data is on, and who is winning the race?

Right now, it's the FAANGs. They're the most valuable companies in the world. If you look at Facebook, Apple, Amazon, Netflix, and Google, what is common about them? They've all monetized their data incredibly well. They've unlocked it, and they use it to generate value to the business. Now, research says that this is exactly how you do and what is expected. So we quoted some research from McKinsey that says data-driven organizations are 23 times more likely to get new customers and 19 times more profitable.

That's backed up by Accenture as well, who says that companies who invest in data have a 38% increase in revenue over three years. Now, logically speaking, you would say there is an unlevel playing field here, because if everybody had the same access to data and everybody can easily unlock their data, you wouldn't have this arbitrage of a small amount of companies generating loads of excess revenue. So you have to consider, what is stopping you from accessing your data? That's the real question.

Why can't everybody do what these people do? And there's lots of reasons. Believe me, we completely respect there's loads of reasons why you can't access data, but the one we want to focus on is data regulation, and that is something that drives different aspects

00:15:00

of it. So when you look at data privacy regulation, we have in Europe, I think most people in the world now have heard of GDPR. It is the poster child of data regulation. But then you look at America, California now has its own equivalent of GDPR. And for those of you who are multinational corporations or companies, all of these other countries have data regulations.

Now, these regulations take a lot of different shapes and forms. It could be that consent based, where you can only process a consumer's data if you've achieved their consent, or they have the right to opt out. It could be the fact that you can only process data in a specific country, or the most recent one that us data geeks have been reading a lot about is Schrems II.

Schrems II is the revocation of Privacy Shield, effectively meaning you can no longer send an EU citizen's data to the US because the US won't give them data privacy with respect to protecting their data. So all of these different countries, different rules, different everything, makes it incredibly difficult to deal with this manually. And what you wind up with is the need for an automated solution for dealing with all of these different issues.

Now, when you have regulation, especially in a regulated industry, what happens is you have a lot of people who start opining on what is going on. Now, these people that when we're speaking to clients, you make sure you speak to your CISO, so your information security officer, he has his concerns. Your data protection officer, they want to know how you're protecting the data. Your chief data officer, who wants to know how is it being used. Compliance wants to make sure you're complying with regulation.

Legal wants to make sure you're complying with the laws around this. Operations, how are you going to do it? Technology are the ones who need the data for either doing analytics for just doing their job and ensuring that the applications are developed and work properly. But then, as we said, you also have the person who owns the data in many data regulations have a say in how that data is used, as well as the jurisdiction that the data belongs to. And that list goes on and on.

There are lots of people governing what you need to do when you want to effectively use data. The next bit is actually based on a true story. So we did something called Cylon, for those who are in London, you may know it. It's Cyber London. I think it's one of the largest cyber accelerators in the world.

During that, we had the benefit of getting to speak to loads of CISOs and data protection officers and chief data officers because they're the type of people that hang out at cyber events, and what we learned was exactly this. So you have the CISO will say, "Yeah, absolutely, I care about data privacy and security and all of these great things, but I'm a stakeholder. It's the business's problem." And the business says, "Well, just because I have the budget, I don't know anything about data privacy and all these type of things." So there's a lot of flipper pointing, and they say, "IT, figure it out." So the little penguin in the middle is the one who winds up having to sort it out, and that little penguin, what he sees is Between him as a data engineer or a data scientist, they have all those policies of the people on the side, and they have to interpret and figure out, "What do I do with these policies?" We hear this all the time from the architects and the engineers we speak to, is that developers and architects, they have to read through all sorts of policies, make judgment calls, whereas their skill set is with respect to writing code, not with respect to being experts in legal or compliance or other areas.

And this is exactly where the thinking with respect to DataSecOps and how do we protect this and how it is all put together. Now, luckily for the little penguin, he can follow a logical thought process on how to deal with this. So DevOps, as a very technical crowd, is supported by DevSecOps. Very simple. If you make the correlation with DataKitchen, even though it's not DevOps for data, but DataKitchen is DataOps, that needs to be supported by DataSecOps, just keeping the correlation going.

What DataSecOps allows you to do is to easily and effectively deal with privacy by design as well as your cyber controls and automate that both in existing and in legacy applications. So the whole goal here is how do you incorporate this into your CI/CD pipeline and just have it much like DataOps, but the security aspect of that, and ensure you bridge that gap between your ops team, your GRC team, your security team, and your data team.

So to get the most out of your data, and this is the one thing I want to remember to take away from this, there's one important thing you have to do.

00:20:00

You need to have a party. It has to be a DataSecOps party. It is absolutely a good party bunch. And the invite list for everyone who has to come to that party is the people that we spoke to before. These are all of the people that need to come together to, as a community, deal with these issues and not just leave it on one team and take a policy and throw it over a 20-foot wall and hope somebody reads it and follows it. So what we want to do at the DataSecOps party, which we're thinking we'll have it catered by DataKitchen, so the chefs will be providing the food, and we will make sure that everybody eats well and has a good time.

So when you look at what happens in the party, you have your operations, control, and governance. And this is what a DataSecOps platform would look to entail, and these are the type of things you would want to look at. So the first thing has to do with the harmonization of the data. So what we decided to do was integrate with something called ODPI Egeria. For those who don't know it, it's an open source metadata framework sponsored by IBM, Hadoop, ING Bank is a very big part of it.

What that allows you to do is to harmonize how your data is called in different applications. So it takes different vendor applications and puts everything into the same language so you compare apples to apples. Very similarly, it is looking at client accounts. Some of the people we work with realize that they have pretty much every system is calling client a different name.

So how do you harmonize that so you can bring it together? And lastly, how do you do PII discovery? How do you recommend it for you, carrying on with my Amazon example of how do you recommend what needs to be protected? The real core here, this is the nutshell of it, is the governance and policy with respect to access management.

So all the people who were invited to the party, you recall, all have their own policies, and they all have a separate and different policy about how data needs to be protected, controlled, where it's allowed to go, where it's not allowed to go. By automating that data policy, what you're doing is you're embedding the rules technologically into the platform. So it knows that your outsource center in India can see the attributes one, two, seven, and nine, because that's what they need to do their job, but the rest will be blocked.

Conversely, your team in the Philippines can see field nine, 10, and 12 because that's what they need to do their job. People see the data they need based upon the job, and this would integrate very nicely in your identity access management system, be it active directory or whatever else you're using. That also allows you to deal with record retention issues and things like data ethics. How do you know that you're using the data properly?

How do you know your AI is taking consent into consideration, for example? The next bit, as I mentioned at the beginning, is in a good DataSecOps platform, you need to embrace the fact that there's not one privacy-enhancing technology that is needed to do the job. So you're not going to be able to take one.

We see this a lot in the market. You'll have a firm will have one privacy-enhancing technology, and they will use it as a hammer and treat every problem as a nail in that they're going to make sure that you use that for everything. And as you'll see in a second, we have people who use multiple PETs on a single issue.

Next up is data destruction or revocation. How, if you had a breach or say, a different example, you want to share data with a third party startup, say a fintech, say someone doing AI on the cloud because you want to get insights. How do you, once you do a trial with them, put the genie back in the bottle and say, "Okay, I shared data with them, but I don't want to share it anymore. I'm done. How do I stop sharing it?" That is dealing with the revocation. That is one of the key parts that you have to look at.

Last but not least, audit. A lot of people think of audit as, "Oh my gosh, these are people are going to come tell me everything I'm doing wrong and I'm going to get caned, and it's just so much work." But in reality, we view audit the opposite. When you have the audit, you can prove to people that you're doing the right things with the data.

You can do things like, for example, when you have an insider attack, you can now look and say, "Ah, somebody's doing things based upon the audit of who's looking at data that are not consistent with they normally do. Let me investigate that." So audit, when you have audit, you want to look at things like who looked at data, who tried to look at data and was blocked, why they were blocked, what rules and governance and policies were in place at the time they accessed it.

So you have a lot of different aspects and components in a DataSecOps platform. But to see now where does that fit in to what you're doing? Let's take a minute to look at that. So when you look at your source data, what we've learned obviously from a lot of work here is source data comes in many different ways, shapes, and forms.

Your source data may come from a database. Okay? That database might be an Oracle database, it might be a SQL Server database, it

00:25:00

might be a PostgreSQL database, it might be a MySQL database. And then you're getting a payload. That payload may be in JSON, it may be in XML, or you might have a variety of CSV files, be it fixed length, be it comma delineated, semicolon delineated, et cetera. There's many different types of source data.

But one of the key things you want to look at doing, and again, depending on which PET you use, you may want to have all of that data consistently and repeatedly protected. Meaning that if Peter will be ABC456 across all of those different mediums, all those different sources, until the next day, he'll be XYZ789 because you now have had key rotation in a different day. But it does allow you, if you put all of your snapshot dates onto the same day, to do multi-chain testing because the data will flow through protected.

Those are the adapters that we've built into all of those different applications. When you look at this and those DataKitchen fans, which I'm assuming most of you are since you've dialed in today, this diagram is something you will recognize from their material. So we've borrowed this from them. Once the data goes into your production flow, it has a certain set of rights and rules associated with it.

You want to protect it in a certain way. And you have all of these ongoing customers. These are your call centers I mentioned. These are your offshore centers. This might be your dev team versus your users who are actually using the product. And you will protect production data in a certain way, and you will ensure that the right people see the right deals in production data.

Once you get into a QA environment, you now have different rules applied to it because a software engineer may want to test with mass production data across a variety of different applications and do multi-chain testing. This, again, different privacy-enhancing technologies for different uses. And lastly, in dev, you might be happy just to use synthetic data, which is another privacy-enhancing technology.

But the concept is you set the rules you define. So in your QA environment, as I said, it may be protected from developers, but your users who sign off see the actual end result of the data. In dev, you care less. You may need another privacy-enhancing technology or another rule set when you want to be debugging something in production, and you need to get a copy of that production data in order to debug it. So there's many different ways that source data is protected in all the different environments for all of the different customers. And all of this is defined in the DataSecOps platform above, of harmonization, governance and policy, privacy-enhancing technologies, data destruction, and audit.

So in conclusion, I want to make sure I leave plenty of time for questions. Data should be a real strategic competitive advantage for people, but the risk and regulation have really made it difficult to monetize that asset because you have to follow a lot of rules. There's a lot of rules that need to be ticked before you can actually become that data-driven organization and win the race for data.

And one of our clients referred to security and data provisioning as soul-destroying, which a bit dramatic, but it is hard. So where you can use automation to make something less painful, we are all for that. And by embracing DataSecOps, you can effectively automate your data policy, so your developers do not have to become legal and compliance experts.

You can protect your sensitive data, but you can also stop sharing it if you change your mind and don't want to share it. And lastly, you can unlock your data to give that maximum value to win the race for data and to become the leader in the information age. So by doing this, we're looking to automate that pain point that people-- we've recognized ourselves and other people have told us exist, and that is what we are looking to do.

So I will hand it over back to Chris. Thank you for listening. Make sure we have plenty of time left for questions, and Sonal always loves a few good challenging questions, so keep them coming for her. And Chris, do you want to talk through some use cases?

Yeah. Thank you, Peter. So let's see. Let me share my screen here and make sure I got the right thing shared. Yeah, so I think that's a great introduction, right? And what I see with our customers is that they're-- I'm going to talk about two cases in DataSecOps and kind of go through them in a little detail.

And I think Peter covered them, but I'm just going to go give a simplified case. So the first case is, well, you've got something in production, right? And we need a set of rules that say who can see what in your production environment. In this graph, your production environment could have reports like a Tableau server and an authorization criteria on Tableau. It could have a database server, perhaps

00:30:00

Redshift, and who gets to see and log in Redshift. And you need to manage that work, who can access it, what can they access, and how that happens. But that is an item that changes. You add users to access to Redshift. Someone, a sales rep, changes his or her last name. You want to change that in the security.

Someone leaves the company. So you need to update the production security just like you update the other things in your environment. And so if you look at it from a process change perspective, which is the DataOps perspective, maybe you've made a change and added some new data to the system. Maybe you've made a change and added a new schema to the database, or maybe you've updated a predictive model, or maybe you've added a new report, and subsequently you've got something updated to your data catalog. All those things should be deployed into production.

You should have a development environment, perhaps a test environment, and then into production. And so we think those things are best expressed as code, as configuration. And so we think the data security and authentication and authorization rules should be, in essence, code or configuration that are kept in source code and deployed simultaneously with all the other things.

Because remember, let's take a simple case. You've got a working system that has a database and a report and a model in it, and a data catalog. And you're just going to add a new table, join it with an existing table, and build a report on it, and maybe tweak your model. So you've got all these parts of the system that you've changed.

And so the process of deployment from one environment to the other is a process of managing change in the configuration of those systems and deploying them from one to the other. And we think security applies right with that. And so this sort of data security as code or deploying data security on top of your "infrastructure as code" I think is a good best practice.

And another way to think about it is if you have processes where someone's sitting down at a keyboard and manually typing things in when you're deploying, it's not a great thing to do. You should be able to script it, automate it, and that way it's more repeatable, it's more trackable, and less errors go in.

So you've had a development environment and you make a change. So now let's talk a little bit about the next case, just about how you actually do development and the environment that it takes. And so there's a lot of complexities in development, and this chart has got four columns on it. One is maybe an individual development environment, a team development environment, a test or UAT development, and a production environment.

And there's lots of characteristics. Maybe there's different people that use the environment. So for instance, your production environment may have very restricted access. Maybe it's only a production team. You may have a team development environment or a test environment. There could be different tools. Each tool could have different code or configuration. There could be different versions of the tools in each environment, which creates problems.

You could have different hardware and software, sometimes bigger hardware or software, depending upon the environment. And even you could have different versions of the libraries. And all of these are sort of opportunities for problems as you move something from an individual development environment into a team into production. And so one of the things is that taking that friction out, making it easy to move things from one environment to the other, is what we're trying to accomplish in DataOps. And the last row here on this, data, is actually really important because data that you use to develop your analytics, sometimes you can take a copy of production, but oftentimes that makes everyone nervous. Do you have all the things that Peter said?

Do you have data that is not useful or is illegal to access in some countries? Are you breaking the law? Are you having a security risk by taking some data from production and putting it in test or putting it on an individual development environment? A lot of data scientists like to do their work on their personal laptop.

Is your production data on their laptop, going out the door every day? And is it vulnerable to some smash-and-grab through a window where they take the laptop off and go? And so how do you make sure that that data that you're using to develop your analytics in the specific hardware with the specific libraries is secure? And who knows all the rules about that?

And so we think that actually that row, that last row, the test data, is very important because inaccuracies in the test data versus production data are a reason why it takes-- One of the many reasons, and actually probably one of the most important reasons why it's hard to deploy things from a dev environment into a production environment. And so your dev environment is

00:35:00

actually pretty complicated because there's a lot of pieces that go into it, like that chart I showed. It's hardware and software, it's the right version of the code, perhaps the right people, and of course, the right test data sets. And

almost every customer I talk to has some concern about getting accurate test data, whether they're in high tech, whether they're insurance, of course, whether they're in the healthcare field. And so test data management is complicated, right? Because you want to make sure that the test data has the right distribution, i.e. it looks like the production data.

Maybe it's got the same fields, the same schema. It has the right quality, has the right security applied to it. And so there's cases where you may want to use synthetic data, i.e. made-up data, or cases where you want to filter out data, either because it's too big or because it has personally identifiable information, and you want to make sure that the data set fits to the rules of the people who are using it. And so how fast can you get that data? How much of a work is it? Some companies have test databases that are six months out of date because they haven't made it an easy-to-use sort of button-push process. And how much effort does it take?

And how the accuracy of that test data actually makes it harder for some organizations to build good automated tests, which you heard us talk about DataOps being able to automatically test your analytic processes, both from a regression and in production, is a very important part of being able to iterate quickly. how much storage is required for the test data, use of real data versus fake data. So there's a lot of complexities in building test data, which is why one of the reasons that we partnered with a company like ExaSan that can help with that. But you have to have test data.

And so it's not an option to be able to do this. And so data and analytic teams need to be able to have test data as part of their solution in DataOps. And they need it that it's recent, that it's accurate, that it reflects production and doesn't break any privacy rules. And I like what Peter said, doesn't get you caned, which is, I guess a very British way of saying what teachers do to you when you're a bad person. And I guess at my Catholic school growing up, it was being hit by the ruler. It wasn't a cane.

But- ... so I think that that's it. These are just two simplistic use cases, but they're very important. Test data and deploying the security as code are cases where we see the impact of data privacy, and

making sure that this is part of your DataOps process and why we want to bring the idea of DataSecOps to the front of people's thinking. And so that's it for my section, and now we're going to go to questions and answers.

Yes. So Beth, is there some questions? Yep. Thank you all. That was a great overview. So if you have any questions, enter them in the control panel on your screen, and we'll just go through as many of those as we can. So just to kick it off, I think this question's probably geared more to Sonal.

What new privacy-enhancing technologies are you seeing coming to the market? Yeah. Hi. Thanks, Beth. There are many that are hitting the market, and we're constantly seeing different types hitting, but they're all these single-point solutions around how this is going to solve all of the world's problems. And they get more hype as they get more investment.

So you'll see some of the big backers from the VC community backing things like TEEs, which is an execution environment. They're talking about this is going to be hardware being used to share data with different third parties, do comparisons, do benchmarking, and be able to do it in a safe and controlled manner. But this is using part software, part hardware, where we're looking at putting this data on a given dedicated environment that is ring-fenced, the way it's using all of the security processes within Intel processes, et cetera, to make sure that this data is sectioned off and is not subject to any sort of hacks.

We see a lot of hype around homomorphic encryption. Homomorphic encryption is a type of deterministic encryption where we're looking at you create individual ciphertext, and you have to use the libraries behind it to be able to decrypt it and see if it's the same information. This, as for large volumes of data, is incredibly slow.

So you really will reduce that to using it to maybe one or two fields of data that you need to do a comparison on, and that links back to another set of data.

00:40:00

We're seeing a lot of movement-- Well, we've seen it for many years, actually. It's been around differential privacy, which is where you look to completely anonymize datasets. But by doing that, we are stripping out some of the value from that data. And considering the way that the information age is now looking, most companies have explored anonymizing their datasets, and they've got the insights that they need at that generic high level.

But now, as the market is moving more towards this hyper-personalized state, consent starts playing a bigger part of it. So where you've got customers that are happy to share their information for some other benefit or for the service or whatever it is, and are equally happy to share it with third parties for maybe getting a coupon in return. You want to be able to provide that privacy to the individuals that opt out to say, "I find it really creepy.

I don't want anybody to look at my information." But then you've got the opposite, where people say, "Actually, I want this service, and I'm happy to give you my data," and transact with their data. So differential privacy, we're seeing a bit more of a move away from that, but the technologies that we're putting together will give you that hybrid that you can use both sets.

You will have an overlap, so you are able to use multiple privacy-enhancing technologies to allow you to do more with your data and be able to provide you with a lot more information to train your models with. And as a final thing is the regulation is constantly changing. We're seeing another regulation in Europe that's hitting, which is data altruism.

This is where people actively want to give you your data. So where we're looking at things in healthcare, for example, the view has been, let's just use tools like differential privacy, where we're just going to anonymize the dataset so we're not really honing in on a particular individual, where if you combine it with other data, you might be able to work out who they are.

This is a case where somebody might have a really rare form of cancer and say, "Actually, I don't mind. For research purposes, I'm going to give you all of my information. Please do use it and do the most with it." But models have been created in such a manner that we're not actually able to still utilize that data as well as we would like to. So it's moving more to that hybrid mindset around this.

Great. Thanks, Sonal. A follow-up to that, are any of these technologies more or less helpful for DataSecOps? All of them are. So that's the thing, that's the beauty of it, is that our view is no one size fits all. So you are going to need cases where you need to completely protect data, and homomorphic encryption might be the right way. It might be preparing information to put into an execution environment, but before you do that, you want to run tests on it. So all of these things you'd want that incorporated in because the entire market is changing. If we look at normal asymmetric encryption, as the quantum age is kicking in, things are changing in that market where they're saying that within 5 to 10 years, you're going to be able to start breaking that encryption, maybe even sooner.

So there are debates around about how quickly is that encryption going to become obsolete. So you want the tooling to be able to

change and evolve as the cyber landscape changes, as the different threats increase, or somebody's created a new privacy-enhancing technology that helps solve many of these problems. So that's where we fit in with that orchestration layer to be able to bring all of that in into one platform. So the developers, the data engineers, they're less concerned about what needs to be done because you'd have to stay abreast on all of these topics.

And we find that having every single data engineer do that, not only could it be potentially boring for them, it is another thing that they have to learn on top of Kubernetes, on top of DevOps, on top of all of the other things that you need to do as a data scientist now. This is another thing to add to the mix.

Okay, great. Thanks. So what technical roles would you need to implement DataSecOps and with what skills? So Sonal and Peter, maybe that one's geared- Yep. Sonal, do you want to describe from an implementation perspective what

skills would be required? Yeah. So we've tried to be data scientist friendly and developer friendly because my tech team absolutely despised doing anything to do with protecting data in test environments. Actually spent more time trying to circumvent the controls that were put in place than actually just do what was actually required.

00:45:00

So we're working off the similar sort of mindset around how DataOps is, that it should be configuration-based, it should be limited code base. So we create a manifest that will sit on top of datasets. And if you have predefined contracts, you can use that over many datasets. And so as an example, we managed to protect over 250 APIs for one of our global asset managers, and that was done in two and a half days. So it's all configuration based, and we've created a nice user interface for you to be able to say, "This data needs to be protected." We do that discovery part through different CSPs at different databases, which then provides a simple API call at the end of it to say, "Add this into your pipeline." So we've done quite a lot of work to make it as easy as possible.

Okay, great. Can you talk through an example of how using code and configuration to ensure a sensitive field in dev is properly handled, for example, masked? Should the code config do the masking or alert you to the fact that not all fields are masked? So we don't think that everything needs to be masked, because we do recognize that at a transactional level, where we're looking at certain information, it doesn't make sense to just protect everything.

So predominantly, the regulations are around specific things such as customer data, be it an individual, be it a corporate. It

might be just something that can re-identify that individual, where are the things from a privacy perspective that you really care about. So what our tooling does, it does the discovery. It'll do a recommended part where what we think that you should be protecting. Tools like Egeria, all of these metadata things that you also talk about in terms of DataOps, they already will have classifications and definitions on what is defined as sensitive information. So the tooling will go through that and match up what that information is and be able to then ongoing do the protection. So we have slightly different configurations for production compared to UAT compared to dev, but each one will have its own manifest that we run depending on what environment it is.

So we just append it with the word dev, and the code programmatically picks it up. Okay, great. Thanks, Sonal. So Chris provided some good interesting use cases. Do you have some real life examples of using DataSecOps in practice that- Yeah. Peter? Yeah. Do you want me to do that one? Yep. Sure, no problem. So yeah, we do, exactly. I actually have some really good ones.

So the use case that Chris referred to is a classic one for us, which is how do you protect data in a UAT environment, but being able to perform multi-chain testing? So that is one of the big ones that we have done for a few different people now, whereby they're finding that it really slows things down basically when you have 50 applications in a chain and you tell 50 different development teams to go protect data, guess what? They're going to do it 50 different ways, and now you've just tied your data into knots.

In reality, they probably do it 40 different ways because 10 just won't protect it at all. They'll just put production data in there. So what we built is the ability to, in a test environment, you can set each application to be the same snapshot date, which effectively is your seed. Which would allow you then to have the data protected consistently across all of those applications. Or you can run a unit test and have it done differently in one.

Another area that we've seen is, again, something I mentioned before, is companies that want to use cloud-based AI technologies, but they don't want to give the full data set to them. So we've come up with a few different ways that you can share the data securely with a third-party AI provider, but still get the insights on your data without putting your data at risk. And then the third one I would say, touching on Sonal's point, is with respect to protecting APIs. So we're seeing, again, a lot of interest in how do you protect APIs, because that is now where some of the risk people are seeing it as a concern, that they want to make sure as the data is being passed around, that it is passed around in a protected state.

And we're working with something called least privilege, where the data's blocked until you turn people on. So again, a couple of real-life use cases that we're doing now. Great. Thanks for those, Peter. So we are running up against time now. So there's one final question here, which I think is a good closing question, is "What is the best first step if you've decided to implement DataSecOps?"

00:50:00

Contact us. What's that? I said contact us. That's the best way to do it.

But yeah. So, a lot of companies have already started doing the data discovery component of it, but we do recognize that there is,

especially with things like COVID, the acceleration that we've seen in terms of cyber risk around data is also increasing alongside the need to get more digital friendly, I guess, and start using all of these cloud-based technologies. So, what we've done is we've created a platform. We are in the process of also offering it as a more broader API offering as a SaaS service.

What we're looking to do there is, once you sign up, and you can sign up to have a look at the platform, we can allow you to just drop in your endpoint or your sample payload. You select the attributes that you want to protect. We give you the cURL command, and then you're off, you're ready to go.

So, we don't think this should be an onerous process. We don't think you should have to spend years to really understand how this works. What we're trying to do is make it as a part of something that you can add into your CI/CD pipeline so you don't have to worry about it. So we've taken a lot of that complexity out of it and tried to create user-friendly tools that allow you to just quickly start being able to protect data in different forms.

So we have adapters that allow you to do that. Okay, great. Thanks, Sonal. Chris, any tips from your perspective on getting started?

No, I think that's good. I think for us at DataKitchen, we focus on the process of monitoring things in production for errors and deploying quickly. And as part of that deployment process, having a good set of secure test data and then being able to deploy the changes that you make to your analytic artifacts into production quickly and securely are where we focus.

And so, working with partners like X8 to help make that happen, I think is a good step as people are working through their DataOps journey. And there's a lot to data security. There's also a lot to code security. So we've got customers, and I don't know if that falls under DataSecOps, but some of our customers are creating code in Python or Java to run analytic models.

And because that's code, they're concerned as it's deployed in a container that that code is in fact secure. And so there's data code security that also applies in this. And so I think security is a good part and should be thought about as you deploy data, because there are too many cases, and I think a lot of people have seen the amount of breaches on data and the embarrassments that happen when data teams accidentally leave an S3 bucket open and hackers find it.

So I think security is important. It should be thought of right at the front with data in a data analytics process. All right. Well, thank you all. We are right at 12:30, so perfect timing. I just want to thank everyone for joining us today and joining the webinar. An extra big thanks to Peter and Sonal for joining us and sharing your wealth of insight on this topic. I'm sure everyone learned a lot.

I know I did. So to all the participants, we will send out the recording and the slides in the next 24 hours. So be on the lookout for that in your email. And if you have any additional questions, please don't hesitate to reach out to either Chris or I at DataKitchen directly or to Peter and Sonal at X8.

So that is it. Thank you all again, and I hope everyone has a great afternoon and evening. Thanks, Beth. Thank you. Thanks, Chris. Bye. Thank you. Bye. Bye.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What is DataSecOps?

DataSecOps is the practice of automating data privacy and security controls into data pipelines rather than applying them by hand after the fact. In this session eXate defines it as bridging the gap between operations, governance risk and compliance, security, and data teams. It sits alongside the other DataOps areas such as ModelOps for data science and AnalyticOps for self-service BI, and maps to the data governance and security function.

How does DataSecOps differ from DevOps, DevSecOps, and DataOps?

DevOps uses automation to speed integration, test, and deployment of code. DevSecOps adds security to that, bridging security and development teams. DataOps applies the same automation to data sourcing, quality, and cleansing. DataSecOps is the fourth square in that grid: automation that brings governance, risk, compliance, and security into the data pipeline itself.

What are privacy enhancing technologies?

Privacy enhancing technologies, or PETs, are the techniques that protect sensitive data while keeping it usable, such as masking, anonymization, and tokenization. The argument in this session is that no one type fits all cases, so a DataSecOps platform aggregates several and applies them consistently, together with PII discovery, access management, record retention, data destruction, and audit of use.

Why should data security rules be deployed as code?

Because the thing they govern changes at the same time. A release that adds a schema, loads new data, updates a model, refreshes a report, and updates the catalog also changes who should be allowed to see what. Treating authentication and authorization rules as code and deploying them alongside the rest keeps production permissions matched to production content, instead of leaving a window where new data sits under old rules.

Why is test data a privacy problem?

Development environments need test data that is recent, accurate, and reflects production, and the shortest route to that is a copy of production, which carries credit card numbers, medical records, and everything else a privacy regulation covers. The three test data management challenges are distribution, the time it takes to operationalize the data; quality, the fidelity required for the tests to mean anything; and security, minimizing risk without slowing the team down.

Which privacy regulations does a global data team have to satisfy?

As of this 2020 session the map includes GDPR in the EU, the California Consumer Privacy Act in force from July 2020, Brazil's LGPD from August 2020, Singapore's Personal Data Protection Act, South Africa's POPIA, Nigeria's Data Protection Regulation, Australia's Privacy Act and its 13 privacy principles, Canada's Digital Privacy Act, and China's Personal Information Security Specification, with bills in drafting in India, Chile, Argentina, Kenya, Uganda, Thailand, and New Zealand.

Where to go next