On-Demand Webinar · 51 min

How DataOps Enables a Data Fabric

A Data Fabric promises to simplify data management across multi-cloud and on-premise environments. Chris Bergh works through where the analytics agility actually comes from, which comes first (the fabric or DataOps), and how a DataOps platform supports the initiative. Recorded April 2021; updated August 2026.

Presented by Chris Bergh

What you'll learn 6 points
  • Gartner describes data fabric as a design concept rather than a set of technology components, focused on composability so users can build a flexible, agile, scalable architecture that supplies data to humans or machines.
  • A data fabric is mostly the centralized data infrastructure a company already runs: ETL, databases, governance, storage, lake, warehouse, and stream or batch transformation. What is new is an AI component and a data virtualization or semantic layer.
  • The goal of a data fabric is agility, but agility is a second-order effect of better tools. The primary driver is people and process following DataOps.
  • AI inside a data fabric maps to Level 1 of autonomous driving, hands still on the wheel, not Level 5 crossing Boston in the snow at night.
  • The data fabric stops short of the end of the value chain. It covers store, transform, virtualize, and govern, but leaves out models, visualizations, reports, and self-service, so it is more hub than spoke.
  • A canonical data architecture designs only for production and not for the process of changing production, which is a little like designing a mobile phone with a fixed battery. The result is unplanned work, manual deployment, errors, and bureaucracy.

Prefer to read it? The written version is in DataOps Enables Your Data Fabric.

Slides

57 slides

Transcript

Show chapters and dialogue 9,254 words

00:00:00

Hello, everyone. Welcome, and thanks for joining us today. My name is Beth Pfefferle. I'm the VP of marketing at DataKitchen, and I'll be the host for the webinar today. So today we have Chris Bergh. He's the founder, head chef, and CEO of DataKitchen, and he's going to talk about how DataOps enables the data fabric. So for those of you who are new to our webinars, Chris is the leader of the DataOps movement.

He has more than twenty-five years of research, software engineering, data analytics, and executive management experience. At various points in his career, he's been a COO, CTO, VP, and director of engineering. He's also the co-author of "The DataOps Cookbook" and "The DataOps Manifesto," and a regular speaker on DataOps at many industry conferences. So before we jump right into it, just a few housekeeping items. If you have any questions, just please enter them in the question box on the webinar control panel, and I'll collect all those, and we'll answer as many as we can during the Q&A session during the last fifteen minutes of the webinar.

Also, the webinar is being recorded, so we will send a link with the recording as well as the slides to all of you at the conclusion of the webinar. So that's about all the housekeeping. So with that, I will hand it over to you, Chris. Thank you, Beth. Thanks for the introduction. And actually, I was counting, it's over thirty years of experience, so I'm getting up there.

And so today's topic is this thing called a data fabric. Now, my company doesn't actually build or sell a data fabric, and so we don't have a pony in this fight. And honestly, when I heard the term a few months ago, I started to kind of laugh at it, saying, "Oh, it's just another one of these tech terms." But I want to talk about how we arrived at why a data fabric's important and what it means to you, and how the sort of DataOps approach can actually make your data fabric work go better.

And so that's really the topics of today. And plus, there's some of my own personal commentary in, which should be enjoyable. And so we're going to talk about what a data fabric is and its new fashion trend. We're going to talk about what a data fabric needs. And we're going to talk about this idea of something called the process fabric, which is sort of an addition or next to a data fabric, and sort of walk through an example.

And so how we arrived at sort of the term data fabric, because I'd heard it, and honestly, I thought it was just another name for data virtualization, like sort of Denodo or Delphix. And data virtualization's been around for a while, and I sort of was like, "Okay, it's another vendor trying to put a fancy name on what they do." And so we regularly meet with Gartner analysts, and the Gartner analyst who's covering it said this is actually a really popular topic.

It's one of the top ten most downloaded papers from Gartner was on data fabric, and they're getting a ton of inquiries on data fabric. And it's not just data virtualization, it's kind of a name for almost everything that you do to build analytic data sets. And I was like, "Wow, okay." And so I went off and I read a whole bunch on what Gartner wrote, and we talked to the Gartner analysts.

We ended up talking to the Forrester analysts as well. And so they're very similar definitions. And so from a Gartner standpoint, it really is, think of it as an architecture pattern for the chain of tools that you're going to use to deploy and build data sets. Right? And Gartner talks quite a bit about sort of AI and metadata sort of acting upon this, and I have an opinion on that, and I'm going to talk.

And I think really the goal of data fabric is kind of, for Gartner is faster informed data, is really agility. And when they look at it from a kind of a stack standpoint, it has a lot of pieces that I think I would've called a few years ago the data management stack, like a data catalog or being able to do data delivery or data prep or ETL.

Being able to actually manage a data catalog. It has the data sources in it, and it's the sort of chain of tools that you use to sort of build a, what they say, a flexible, agile, and scalable architecture that will supply data. And so it's a design concept. And so I think that's interesting. It's like, okay, one way, it's sort of a name for the tool chains.

And what Forrester says is something similar. It's kind of dynamically orchestrating all the data sources and data platforms to deliver integrated and trusted data to support various applications. And I think that's a really good point, because I do agree that data consumers, they want data that they can trust, that they want data that's in a form that they can use. And so Forrester's got a slightly different definition. But if you look at the diagram on the left, it's got

00:05:00

data lineage, data quality, data processing, governance, security, metadata, and there's existing categories for here, and it's sort of got the little AI box in each. And so to summarize, it's sort of hot stuff. Gartner and Forrester are talking about it. It's a top downloadable report. They're getting inquiries. And so in my view, what's a data fabric?

Well, it's all the stuff that you have normally done to build a centralized data infrastructure, ETL, databases, governance, store, lake, warehouse. It also includes stream and batching. Plus it's got some sort of fancy stuff in. And then I think there's an AI component that I'm going to talk about. And I actually think it's kind of magic pixie dust.

It's sort of analogous to self-driving cars. And then it's got a data virtualization, which actually I don't think is magic pixie dust in a semantics layer, which I've seen being used at a bunch of companies. And I do have a criticism, and I think it actually is missing sort of the other part of the data value chain, which is visualization, self-service, and science.

It's sort of more about a hub than a spoke. And so why is it important? Well, it's a moniker, it's a term, and sometimes terms matter. And sometimes terms get distorted. And so it's a moniker that covers sort of the latest architectural pattern in doing data management. And for that, I think it's important.

And the second part, I think the goal of implementing a data fabric, in my mind, is not just to implement the cool new tools, right? It is really about agility, the ability to give your customers a data set that they trust, that they can change. And to me, that agility is kind of a second order effect from the better tools. And I'm going to make an argument to you here about the primary driver of agility is sort of people and process following a DataOps process. So if you want to go off and eat lunch, this is my discussion, and all the slides are going to be in this.

Is data fabrics interesting? It's got some new stuff in that we're going to talk about. It's a moniker and the goal of a data fabric, while being lots of new technology that has some benefits in it itself, if you're driving towards agility and being able to make your customers successful, it's the DataOps process that gives that, not the tools.

So what does a DataOps fabric need? So, I'm going to do a couple of things here. So watch, I'm going to go through six sections. So your Fabric's AI and shiny new tools do not make agility. I'm going to talk about that first. Stitch the fabric together. Your fabric's a factory. Agility needs the fast flow of fabric fixes.

Your fabric is delivering a platform for self-service. And measure your fabric. And I'm a little bit sorry about the dad jokes. Actually, I'm kind of not that sorry, because we got to have some fun here. And so I thought about sort of wearing a poncho or some other kind of interesting fabric today in the webinar, but, ultimately, I was dissuaded by our VP of marketing, probably smartly.

So let's talk about this first thing, your Fabric's AI and shiny new tools. And so, if you look at, there's this tool chain or chain of value or factory that we've talked about and sort of data comes into an analytic system. It's transformed, sometimes in batch and sometimes in streaming, which I think is a really good design pattern.

Sometimes it's actually put in a catalog so you can understand where it is. That metadata in the catalog talks about where it is, where it goes, and then there's this layer on top that tends to virtualize it, and sometimes it's called a semantic layer, sometimes it's called data virtualization. And I think the data catalog has, or the idea of a data fabric includes this.

So it's not just a batch architecture, it's a batch and a streaming. It's not just a where all the data has to be in the same database. It can be virtualized across multiple databases, and I think that's a fairly, I think, a very good way to do things. And so you're not going to hear any qualms from me about including streaming in a reference architecture or virtualization.

I think those are good things that people should think about doing. And so, what I am going to do is talk about this sort of magic of AI inside. And now, why did I put "Danger, Will Robinson"? Well, there's a couple reasons. One personal is sort of I spent the first five years of my career working on an AI project to automate air traffic control to improve the sequence and spacing of aircraft landing at terminals all around the United States. It was a NASA Ames project and MIT Lincoln Laboratory, and I loved it. It was a lot of fun, but it was actually really hard. And we had the ability to automate every case and sequence every plane landing, but it turned out that that wouldn't work some of the time. And every time we wanted to get it a little bit better, go from 95% to 96% right, it just took forever, and I spent five years on it.

And now the AI techniques have improved.

00:10:00

But to me, it felt like every time you could do most of the cases pretty well, but the sort of corner cases were really hard. And so what we ended up doing is, instead of telling every air traffic controller exactly what to do, we took out less information and said, "Okay, here's some advice. Just sequence this plane behind another plane." And so we did a little less, and we saw it as a way to aid people.

And I actually see that now in this sort of latest version of AI being an amazing thing. And so I think beware the magic of AI inside. And if you think of it as autonomous driving, and in autonomous driving, there's this level, where level one is some assistance, and all the way up to level five, meaning I could go from Kendall Square to Southie on a winter day through construction at rush hour, and I don't have to touch the wheel. And honestly, nobody can do that now. We're probably at level two, maybe, or level one in autonomous driving, the sort of sunny day case on a freeway, and you got to keep your hands on the wheel. And I think that's great.

I'm not denying that that's AI or that that's of value. I'm just saying that the idea of fully autonomous driving is a ways away, and I'm not the only one. It's Rodney Brooks, professor of AI. And I just think that type of AI is going to work. And I think similarly, we're sort of level one of the AI and the data fabric, and don't think we're going to get to level five. And so don't fire your data engineers today.

They're really important. Don't think that a magic box is going to make all this happen, and you pour your data in, and you get out an integrated, trusted data set that all your data scientists and analysts are going to use. And just don't believe it, because it's not going to happen. And so what that really means is that I think there's a theme here that sort of AI and plus new tools in the data fabric sort of gives agility, gives the ability to build that integrated data set that your customers trust, that you can change at a moment's notice, that can respond to industry trends and their requests.

I don't think that that's true. I think what it really means is that with the agility, the ability to do that and iterate quickly and deploy quickly comes from the people and the tools that you have in a DataOps process. And so it's a very different view than your tools will make you great. It does mean that your people are going to have to work in a different way. And yeah, I run a company that is a DataOps vendor, so perhaps there's bias in there, but this is my lived experience.

And so this is what I believe. So the second thing, stitch the fabric together. And again, sorry for the bad pun. So if we talk about that tool chain of data streaming and data virtualization, and we think about all the pieces together, we've often talked in these webinars about how your tool chain is really an assembly line, and you're building data products, and you've got to be able to monitor them in production to make sure they're right.

You've got to be able to use their tools, do things like version control and meta orchestration. And we've talked a lot about those in different webinars and why you need those. But I also think the data fabric from that perspective, the value chain perspective, is incomplete, because the final parts of the value chain are things like the model and the visualization. And just building an integrated data set, that could be the perfect, most well-formed integrated data set, but there could be a problem in the calculated field in your Tableau workbook.

The model could have come out of its root mean square error compliance. And guess what? You're going to get a call on Friday saying something's off, and you're going to have to work the weekend to fix it. And so to me, we really need to start thinking in data and analytics in the complete value chain and making sure that whether-- And just don't stop at an arbitrary point in the middle saying, "I've done my job.

Hands up, you guys figure out the rest." I think we've got to think about products and not projects, and the final product here is always something that goes to your end customer or your website. And so including that final value, I think, is an important part. And we've talked about that in the stuff that we wrote as a value pipeline. That whole piece, all the way from the journey that it takes from its source, all the way to the value to your end customer.

And so if you look at those dark green boxes that say store and transform and govern and virtualize and model, man, there is an industry out there, right? There are hundreds of companies that do that. And there's been a general trend to the cloud, and that tool chain has actually been available. So you could kind of build a data fabric in the cloud, and I'm not exactly sure whether they have all the pieces.

I'm a little unsure whether you can do data virtualization in AWS and Azure. But I'm sure there's some company selling a tool that you can install and run there. And so they are really a collection of tools.

00:15:00

And there's no defined process or end-to-end system, and there's no way to sort of run it as a value chain that's sort of test informed, and we use this term meta orchestrate. And so one of the arguments that we're going to make is that if you don't have that, if you don't have that sort of process fabric that sits on top of your data fabric, you're going to have some challenges.

And of course, as a vendor who makes one of those process fabrics, we think there's characteristics of it that matter. And so one is the ability to meta orchestrate, because every data fabric is going to have tools. And if you buy from one vendor, they're all going to have different tools which end up running as different software processes, sometimes on different servers.

And as your data flows in through your system, you're going to want to automate tests and monitor those tools. And sometimes we call them observe to make sure the data's right, because if your customers want integrated data sets that they can trust, well, trust comes from verification, and the verification comes from automated tests running in production.

And then, of course, the other part is that you want to be able to have people develop quickly. And so you need ways for them to manage their development sandboxes and ways to collaborate together. And then finally, since we are dealing with a process, a process fabric, measuring that process fabric's important. And so we're going to talk a little bit more about some of these in the next sections.

And so this first one I think is important, and thinking of your fabric as a factory, so automate and test. And I still see a lot of companies who build something and put it into production and then kind of hope the data is flowing through is right and hope their data providers haven't changed anything and hope there's no significant problems, and they rely on their end customers to say, "Hey, something's wrong with this data. Can you check it out?" And I just don't think that's an acceptable way to work. For a lot of reasons. It costs your team time. It's a big hassle.

It lessens your time to actually innovate. And so if you look at it from an end-to-end system, how do you observe or monitor or check the data that's passing through each section of the factory? And you've got to monitor and test every step. And one way to think about it is this idea of statistical process control.

Well, you run a factory, so sample your factory, sample your data set, the size, the shape, the counts, the values, and then do trend charts on it. And look at it over time, and that guy Shewhart in the middle, do a Shewhart chart. And when something goes wrong, make sure that you tell people and alert, because it's just so much better to find problems in your production line. And you want your production line, as I've said, to make really high-quality cars and not have to make American Motors Pacers from the 1980s. And one way to do that is to don't think of testing and monitoring as a manual process or kind of a bag on the side.

Think of it as integral to what you do and make checks at every step of the way. And if those checks have a problem, send an alert. Tell someone right away so they could possibly fix it before you miss your SLA. And we've talked a lot about how to do these checks and one of the things that's interesting about a data fabric and why people have data virtualization is that data bounces from place to place.

It comes from an FTP site. In the cloud, it may be stored in a bucket store, it may end up in a database in one layer. It may end up being cached in a report in another layer. And how do you know that you haven't had any lossy transforms? And so we think of that as just a location balance test.

Make sure that your fabric doesn't have any holes in and things don't drop out. And then another way to look at it is to make sure that you are testing if the goal, as Forrester says, is to give it a trusted, integrated data set. Well, most people who use the data look at it from their perspective, and your customers are going to look at it from their top products, their top regions, their top suppliers, their top manufacturing centers, and they're going to have memorized those numbers.

And so make sure that every time you give them a new update or a significant change, that you check their view against what they've seen before. And we call that a historic balance check, because a lot of times in doing data work, the errors come from small files, and sometimes those small files can make big changes. And so you can't just check that the data's right when it lands and say, "Okay, I've checked some of my data size from my FTP site, it's the same, therefore everything's going to be right." You're having a software-defined process that integrates and transforms data, and sometimes that aggregation or software may be wrong, and therefore you may end up with an error that you're going to have to pull your hair out on a weekend and fix.

00:20:00

And the last thing I think is that these idea of a fabric is that the tools that are in a data fabric tend to have different ways that people want to work with them. And some people like to work at the bottom with a nice UI tool, and maybe they like to use Informatica's graphical UI.

Some people like to do things that are more structured, like Python or SQL. There's different ways that people build their data fabric and different metaphors for how people use and do their work. And so my feeling is that

you should use your favorite tool, and if you like visual UIs, use them. If you like to write SQL, use it. If you like to write language X, use that. And at the end of the day, that is code that needs to be in version control and needs to be checked, and you should test based on the tool of your choice. And so, of course, our software provides that.

And so think of it this way. You're building a data fabric, and so it's running. You've got all the latest tools and your AI is doing some transformations and fantastic. And do you trust that it's right? Well, I don't think so. I think you got to verify, you got to do some checks and do those automatically on top of your data fabric and send alerts and notifications.

And we talked about a bunch of different types of tests and I think we did a webinar a few weeks ago talking about testing in detail and talking about the needs for it. And at the end of the day, it's just the hope that your data providers aren't going to screw you over. The hope that with a lot of people working in different parts in the data fabric, that they're all going to know that everything's going to work, I just don't think is a realistic thing. You need to check it and test it and focus on error reduction. And if you do, you actually end up with more time to innovate and more customer data trust, which is, after all, what you wanted to build the data fabric for. And that goes to me, another point why I see

the data fabric as a long line of tools where the industry hopes that if you adopt the tools, sort of the benefits on the right will happen. And I've seen it happen in big data. I've seen it happen when everyone was moving to Teradata, even Oracle cartridges and data warehouses. I think the idea that the tool will figure out and make you better, I don't believe in. I think you need to, as a leader, build a system that your team and tools can

be better in. And that system, the characteristics of that system, one of them is using the principles of DataOps. So the fourth thing I want to bring up is sort of agility needs the fast flow of fabric fixes. So what's the benefit of that alliteration, the fast flow of fabric fixes? Well, I think agility means that I think I've often been humble about what I think my customers want. And so I've learned that if you put something that's 70% right in front of them and iterate and improve upon them, you are much more likely to get something that they want, and you actually end up saving time.

And so to do that, you need to be able to deploy things from a development environment into production, not in weeks or months, but in hours or minutes. And you need those deploys to be able to not have a lot of problems. And if you do that, and I think if you focus on that, your data team ends up being happier and more productive.

And so in this next section, we're going to introduce the world's worst acronym. And so a lot of people, when they talk about deployment, they think of, "Oh, software's figured this out. It's called continuous integration deployment. And if I just do what software people do, it's right." And I think that's partially true. I think there's a bunch more, and we've got this terrible acronym called not DevOps CI/CD, but DataOps CSMOIDM.

And we talk about sort of continuous sandboxes and meta-orchestration and integration and deployment and testing and monitoring, and all those things have to happen. It's more complicated than what you're typically doing with With software. And I know because I spent 15 years in the software industry, or actually most of my career building software.

And so another idea is that in a development process, when you want to change something, you don't just change the individual piece. You need to see the whole picture. So if I change the middle part of my data fabric, I want to see the effect of that change in not only what is using it, but also I want to make sure I haven't broken my metadata about it. Because if you're going to add a new table or change a table, that should be reflected in the metadata that

00:25:00

represents it. And likewise, if there's a report that uses it or a model that uses it, I want to see the effect of my change in a development environment. And that's what is the whole point of CI and CD, right? From a software standpoint is individual contributors can go from their local change to the global impact. And I think you need an automated system to do that. And it's not just about if I write a simple unit test and toss it over a wall, it's somebody else's problem.

So, there's another phrase that I like in the DevOps world, sort of pull the pain forward. And pull the pain forward so that your individual contributors can see the effect of their changes. And I think these tests that you use in development form a dual purpose on it. So the last point here is, I think a lot of data architecture or data patterns are very focused on what's in production, and can I get my new data fabric? They think about, okay, they only look at the production view of the world. And I think the idea of that is a lot like the right to repair movement.

You should be able to change what you buy. You should be able to pull a battery out and put a new one in. And if you don't do that, you end up with a lot of sort of manual work and bureaucracy, and changing it's hard. And so if you look at a typical data fabric architecture that I've got in front of you, it's very production-centric.

And maybe it's not in green boxes, maybe it's got a lot of complicated lines in it, but nonetheless, there is an architecture that the idea of a data fabric's getting, and it's got these components in, plus the additional two that I talked about. And so what we think is that this idea of thinking about the change, thinking about a system, thinking about the functions of deployment and environment management and orchestration and testing as kind of first-order ideas, they're as important as your data fabric architecture.

And build for that is-- and build for change, build to do the right repair because at some point, there's going to be a new version of your data catalog that you're going to want to transition to, or a new version of a database or a data tool. And you're certainly going to want to change the code that runs in those tools as fast as you possibly can to have impact on your customers. And so build a change, don't build... So, I think these ideas are important. And also this transition to cloud.

A lot of people have big organizations who are interested in data fabric, or even if they go from an old version of their data management to a new data fabric, they're going to have two things running. And how do you make those two fabrics work? I think that's an important concept to think about.

So the right to change, and how do you actually transition or run when you've got your new data fabric next to your old style data management tool chain. And the last thing is, I think if you take the definition of data fabric as building data sets that you give to your customers, well, those customers are going to use that in some ways. And so that has sort of two patterns. If you think of it as a hub, your data fabric provides the data hub of integrated, trusted data sets.

Then there are going to be people who are using that hub. And so this hub and spoke, I think is really an important part to think about. And I think from a central IT standpoint, you need to figure out how the people who are doing the work on the spokes in your team can be able to do that work in a way that doesn't make your brand-new data fabric look bad.

And so you've got to think about what's the development? How are those people going to do that self-service? And so, kind of going to that idea of self-service first is that, if you look at it, a lot of times I've seen a number of companies say, "I want to go to this hub and spoke model," or they call it data enablement.

And you've got self-service teams using sort of Tableau or Looker doing their work. And the sort of data enablement team sort of stops and hands it off and says, "Hey, I'm done." And think of them as they're building the airport. And the people who are actually doing the fighting and flying the fighter planes are suffering because you can't fight a war with just an airport.

You need a whole support and logistics to make sure that happens. And so there's kind of a gap. And you need those two things to work together. And so having your data enablement platform is important, but it's not enough. You've got to consider the spoke, consider those fighter pilots out there doing battle with your business customers, trying to give insight.

And when they create something, how can you turn that into a production asset? What's the path to production for those teams? And how can they create reusable components that you can share? And so I guess I urge you to think about when you're building a data fabric, just don't think about, "Oh, I've got self-service users that are going to do their

00:30:00

stuff, and I don't have to worry about it." I just don't think that makes sense because your business customer at the end of the day sees the combination of the work, and honestly, you're going to get that from an IT team perspective. If things don't get worked, you may end up being blamed for The lack of it, the lack of coordination.

And that also goes true, and I'm going to talk about this example, when you've got two different groups working in two different places, even if they both are IT groups. You may have an on-prem or a cloud solution. So you've got to think about this hub/spoke, this coordination between groups as an important part of your data fabric design, and just don't expect that you've got the data fabric, that you're building trusted, integrated data sets, that your job is done.

And then lastly,

I think measuring the work that you do on your fabric, both from a production metrics, how your streams are going, how your data sets are going, how you're building the integrated data sets, I think is a really important thing, because you're running a factory. How efficient is that factory? Are your current data sources live? Are your current builds running?

How good are your data sources? And then also looking at the work that your teams do with your data fabric tools. How well are they collaborating? How productive they are. What is their test coverage? And getting metric about the work that you do, I think is a very valuable way to actually help your team improve.

And it's ironic that when I've talked to a lot of companies and asked to see their sort of production metrics or their productivity metrics, these are big companies with hundreds of data scientists and data engineers. You'd be surprised at how very little people have invested in this sort of process measurement capability. And I think it's an important part of both showing your value and running an efficient team.

So, as we go forward here, let me talk about sort of what this means, why a data fabric sort of requires a process fabric. So if some of you are thinking about getting a data fabric, fantastic, right? You've got a tool chain, you've got virtualization and streaming. And so we think that in addition to doing the data fabric, you need this thing of a process fabric. You need to be able to think about the processes that are built into your data fabric, that are governed by code, that are created by people who may or may not work in the same location.

And invest in those processes, because you're not going to get agility just by plugging in a new data fabric. That's a fantasy that your sales reps are telling you. The

agility, the ability to deliver small bits of value to your customer and iterate and learn comes from following a DataOps process. So you need some to think about this process fabric. And what we've seen in is this, is there's just a lot of problems that people have when they don't think about their process fabric.

They end up with a lot of errors and a lack of data trust. They've invested all this money in a new data fabric, a new tool chain, and the customers say, "Well, what did I get out of it?" Or they spend two or three years building it, and the customers then don't need it anymore, or what you thought you built isn't what they need.

And so when they run it, we talked to a customer yesterday, they've built a new system and they can't get anything new done because they're spending all their time dealing with execution issues. Things are breaking left and right. And we also talked about how teams and how to make sure that in a hub and spoke model, you deal with the spokes.

And lastly, don't build a new data fabric and expect that you're done. You're not sort of building a house and walking away. You're not doing a project. You're building a product. And so this thought that a data fabric is continually changing, continually adding new data sets, reconfiguring the data build for change, and adding new feature requests and addressing that rapidly is incredibly important and really the value that your customers want, because if they want a trusted data set, well, they want a trusted data set that they can control and change, and make sure that it's right. And so, from our standpoint, we've built a product to enable this about meta orchestration and testing and monitoring and building controlled environments and being able to deploy new work quickly and being able to handle collaboration and measurement.

And so this is the thing that we add to your data fabrics. Again, DataKitchen isn't a data fabric. We're a process fabric that works with your data fabric. So let me just talk through an example, and before we finish up and leave some times for questions. And so it's a global pharma company, and they have a complex data landscape, right?

They're doing drug development, and in this company, they had sort of two groups. One was building kind of a big, or has built a really big data lake with sort of peta-scale, sort of drug discovery

00:35:00

data, a large Spark cluster, and sort of a best-of-breed tool chain. Like they had Stream Sets. They may have had Denodo. They had models. They had just like half a dozen different tools, really cool stuff. And then they had another group, also in the drug discovery, who was working on Azure using the Azure tool chain, plus Databricks, plus another ETL tool. And they actually had to work together.

And there were two smart teams, both working independently, but the result of what happened on the left was fed into what happened on the right. And so they needed to work independently, right? They had their own different tool chain, their own different work processes. And if someone in New Jersey made a change, how could they tell if someone in McLeod, in California, things were broken?

And so in this case, it's almost like there were two hubs in addition to two data fabrics running with each other, and they were kind of peers. And so that lack of trust and that problem with these teams working was paramount. And I think, don't just think you're gonna have one data fabric. You're gonna have multiple fabrics, and you need to connect those fabrics together with a process fabric, and that process fabric has to extend to working across multiple fabrics.

And it goes down to continuously monitoring and end-to-end testing and stopping regressions and alerts and restarts. And they're very sort of basic things, but they're Really, it's important to get these two groups to work better together. And if they don't, then you're just going to end up with challenges. And so let me just kind of conclude our discussion here with one slide that sort of summarizes what I've been saying. So, data fabrics are a cool technology. Gartner, Forrester are talking about them.

You're probably going to be hearing more about them. And of course, the term's going to be used in a lot of different ways. But like any data management technology or data fabric, to really be successful, you're going to need to focus on the processes that your team works in and your tools work in to be successful. And we think the best way to do that is called DataOps rather than sort of waterfall. And so connecting all those tools together. And so one way to think of what DataOps is, is as a process fabric that sits next to or on top of your data fabric or on the bottom of, in this diagram, that connects your people and process and tools together.

And once you have that in place, you can actually start alleviating these bottlenecks of how you can change things and build to that repair, build for change right as part of it. And we've talked a lot about this before in our manifesto, our cookbook, the excerpt from "The Unicorn Project." And so, if you want to learn more about DataOps and DataKitchen, please do these resources.

And the last thing I want to talk about is just give a shout-out to our next one. And we're also talking about another fashionable term out there called the data mesh, which I also think is a very cool idea sort of related to-- and I like it because in a lot of ways, DataOps is sort of bringing ideas from software development into data and analytics, and here's another case of it, sort of a data mesh, which is domain-driven design, or DDD, which has been around software for a while. And so I want to talk about how that works and how it relates to DataOps. And of course, there's going to be some more bad wordplay on how

DataOps can connect your mesh so you don't end up with a data mosh. And hopefully you haven't been upset by my bad wordplay today, but that's it for my presentation. And, as usual, Beth's going to be gathering questions and I'll see if I can answer some of them. So thank you for your time.

Great. Thanks, Chris. That was great. As Chris said, we have a few minutes for questions, so if you have any, please enter them in the question box and we'll try to get through as many as we can. So Chris, first one is actually more of a comment. It's a good comment, and I'll just let you provide your thoughts on it.

So data fabric seems to be the be-all and end-all of data management and analytics. Perhaps listing what, if any, is not part of the data fabric could help understand and sort the hype from the facts. I know you did talk about that a little bit, but if you could elaborate. Yeah. So I think, kind of going through these slides, what I don't think is in the data fabric is anything having to do with data science or building models. And so I think that's a bit of a myth, right? Because it is really having models that make predictions, sitting on that trusted, integrated data set I think is important, as well as sort of visualizations. And so in some ways, I think the data fabric is a name for the set of tasks that big companies' centralized IT teams do. And so they're trying to build trusted

00:40:00

data sets that their customers can use. But also in that whole self-service visualization layer, there tends to be a whole set of data prep tools. It's not just visualization. It could be tools like Alteryx or Trifacta, where you're actually doing good data integration work on top of the work of your central IT team.

So the hub and spoke model, I think is the other part. You've got to be conscious of what's in your spoke, and sometimes that is not even self-service tools. That could be Python or people could be writing a lot of SQL in your spoke. And so it's not just self-service tools. And so that concept of a hub and spoke architecture here, I think is implicit. And the data fabric is really putting a good name around all the technologies that you want to build in your hub.

But don't forget the spokes. Great. Thanks, Chris. So how do you deal with dynamic data when building a data fabric? What are the most important things to consider when building one? Yeah, I guess for me, all data is, there's dynamic in a couple of ways, right? So if I think about what dynamic data means, is that your data provider is going to give you new data sets, is going to change perhaps the meaning of a data set in a rapid way.

And so, how do you handle that? And so in two ways, one is during a production process, you want to notice that that happens. If they add a column, you just want to know that before your customers do or change the meaning of data, and that comes down to automation and testing that we talked about. And then second, if they're adding a new column or giving you a new data feed, how fast can you get that into production? Or how fast can you get that into a state that you can get a customer to look at it and give you feedback?

And that's a different question than in production sometimes. Can I take that data set and do a quick work in a week to do some joins and give a version of production that I can get feedback from my customer saying, "Hey, I looked at the data. I think it fits together this way. What do you think?" And get it sort of 70% right and get feedback.

And I think every time I've done that, customers have loved it because they can see and touch it. And rather than having sort of contract negotiations and requirements documents, how about creating a variation of production that instead of having your 10,000 normal users use, you have one or two users look at that variation. And so that's how you can handle dynamic data by thinking about handling variation, thinking about changing, thinking about automated testing, and these things really are DataOps capabilities in my mind. They're not data fabric capabilities.

I do think the data fabric could help you actually perhaps build the integration faster. Maybe some of that touted AI could Go in and look at a schema change and automatically put a new column in your staging database. That sounds cool, but I don't think it really handles the full extent of what people mean by dynamic data.

Great. The next question is, "Does the data fabric configuration require extensive collaboration with the infrastructure team? If so, automation becomes a challenge in regulatory environments, right?"

Yeah. Well, I think in a lot of regulated environments, there is an infrastructure team who doesn't want anyone to touch production. And so maybe you've got company's financial data, and maybe you have FedRAMP certification. What that means is the development team has to work with data that's sort of like production, but not production-level data. And it has to work with hardware that's sort of like production, but not production hardware. And the movement of your work, your code and configuration from development and production should be as automated as possible.

And so in that environment, which is not only regulatory, there are cases where you have very sort of

strong production teams who don't want to be messed with and want you to, as a development team, give them the stuff, I'm going to run it because I'm going to be the person waking up at 2:00 a.m. if it breaks. And so I think that's a common case. And so, again, thinking intentionally about the flow of work between your development environment, what it takes to have a development environment that's like production, how to run production with low errors and alert, those sort of DataOps capabilities, I think are important.

And again, it doesn't matter if you've got a data fabric or not. Data fabric consists of a set of tools with a new name on it. And so that set of tools, whether you call it a data fabric or data management or whatever the next term is that the analysts are going to use, it still is a set of tools that are governed by code that has to run in an environment where production shouldn't be broken.

Okay. Somewhat along the same lines, "Is there data quality in this model and where does ETL testing land?"

00:45:00

Yeah. And so I think if you look at, the Forrester model, it does talk about data quality as a thing. And, I think it's right here on the left-hand side, data quality. Because a lot of ETL platforms like Talend have a data quality suite on. And so I think what we talked about in our last webinar is that that's great.

Being able to measure statically data quality is fantastic. Being able to profile data and understand what it is and being able to then use that profiling information to help you do your work, I think is great. What I do think is missing is that, okay, let's say you profiled your data and one column has three values in it, or the sum of a column is equal to X amount or the average.

How do you take that and then put that as a test where you know that the next time you get that file again, you're checking that that one column has three values, or the sum of the column was 10,000 and now it's suddenly 10 million. And I think those kind of checks are what we talk about in DataOps, which is the way to help handle the dynamic data changes.

So data quality is, I think, part of it, and to me, that tends to be more of a sort of a cashflow versus a balance sheet where the data quality is kind of your balance sheet. It's a point in time view of your data, and a cashflow is sort of a look at the changes and what happens in your system over time.

And sometimes we've called it data testing and monitoring, we've called it data observability, but it's that actively grabbing bits of data in every part of your process to make sure it's right is what we want customers to do. Okay. Thanks, Chris. So the next question is, "If DataKitchen is the process fabric working with the data fabric tool chain, does DataKitchen contain a data lineage/metadata process within its process?" Yeah. No, we don't, actually. We're not a data lineage tool.

And I think there are tools out there that do it, and one of the reasons why is because we actually think the process lineage is important, what step, what code, what test results happened at every time, and keeping that available. And so I think there's a lot of good tools out there that are part of a data governance suite that have data lineage.

But to me, I look at it in a very practical way. When things go wrong, I want to be able to diagnose what went wrong quickly. And to do that, I want to see what happened. I want to see where the problem was, at what part of the process. I want to see what code acted upon it, what test results.

And I want to be able to, if I get sued two years later, I want to look at that same set of data to be able to help me understand why that happened. And so process lineage, I think is an important-- And that's different than data lineage, right? Where the data came from, what steps happened to the data in the middle.

And so I think just like there's a process fabric, just like the data fabric is related to a process fabric, I think data lineage is related to process lineage. Okay, great. So, I think this is our last question. It's actually two questions that are related, so what comes first, related to timing, what comes first, DataOps or data fabric, and can they be implemented simultaneously?

I

I think they should be implemented simultaneously, but it sort of depends on the size of your team. If you've got one or two people working and you want to get some stuff out, you're going to end up building some little technical debt that you're going to have to fix when you hire more people.

And that sort of DataOps technical debt is something that you have to realize, and it's not the end of the world that you do it, you just have to work through it. If you've got a bigger team kind of building a data fabric using DataOps principles and saying, "Okay, we're going to implement a data fabric, and we're going to focus on deployment and testing as first order things," I actually think it makes the implementation of a data fabric much faster and much closer time to value.

Because one of the sort of cardinal rules I have is that don't spend a year building something and then give it to your customers and expect them to like it. Have ways that you can incrementally start deploying your data fabric, using the current data, using a DataOps process, and iterate quickly, and every week have something new for your customers to do, and to see and give you feedback on. And I think that's a better way to work.

And so how you can do that sort of iterative development methodology, I think you can start off sort of avoiding DataOps principles, but when the team gets to be a sufficient size of more than one or two people, you're going to have to start adding tests and automation. All right. That's great, Chris. Thank you so much.

We have no more questions, so we can get a few minutes back in our day.

00:50:00

Thank you so much, Chris. And thank you everyone else for taking time to join us today. I hope you found this as useful as I did. We will be sending out the recording and the slides in the next 24 hours, so be on the lookout for that in your email. As Chris said, to continue on the fabric theme, we have a data mesh webinar coming up in a few weeks, so we'll include the registration information for that in the follow-up email as well.

And then lastly, if you have any additional questions, please don't hesitate to reach out to Chris or myself at DataKitchen, and we'll be happy to do whatever we can to help you. So thanks again, everyone, for joining us, and have a great afternoon and evening.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What is a data fabric?

Gartner frames data fabric as a design concept rather than a set of technology components, built on composability so a flexible, agile, scalable architecture can supply data to human or machine users. In practice a data fabric is most of the centralized data infrastructure an organization already has, meaning ETL, databases, governance, storage, lake, warehouse, and stream or batch transformation, plus an AI component and a data virtualization or semantic layer. Typical toolchain elements include catalogs such as Alation, virtualization such as Denodo or Delphix, pipeline tools such as Informatica, Talend, or Airflow, and databases such as Redshift, S3, or SQL Server.

What is a DataOps process fabric?

A process fabric is a single user experience spanning the whole DataOps process that stitches multiple separate tools into one system, connecting people, processes, and tools rather than only data. It supplies meta orchestration across tools from one pane of glass, automated testing and alerting in both development and production, automated sandbox creation and management, a shared collaboration system across roles and teams, and measurement of delivery speed and quality. It spans environments and data centers, so a production environment in one cloud and a development environment in another sit under the same platform.

Why doesn't a data fabric deliver agility on its own?

Agility is a second-order effect of better tools; the first-order driver is people and process following DataOps. The tools in a data fabric lack the team and environment awareness needed to promote reuse, and they lack a common collaboration system for multiple teams using disparate technologies. The AI marketed as part of a fabric sits at Level 1 on the autonomous-driving scale, keeping hands on the wheel, not the self-driving data the label implies.

What are location balance and historical balance tests?

Location balance tests confirm that data properties match business logic at each stage of processing, so the same measure holds as data moves from store to transform to model to report. Historical balance tests compare current data to previous or expected values, using history as the reference for whether today's values are reasonable. Both are forms of statistical process control: check upper and lower bounds over time, watch for a trend break, and alert.

How should test results be graded in a data pipeline?

Three severity levels. An error stops the line. A warning is something to investigate later. An info result is a list of changes to be aware of. Tests belong inside the pipeline at every step rather than bolted on the side, and they should answer three questions: are the data inputs free from issues, is the business logic still correct, and are the outputs consistent.

What is C/SMOIDTM, and why isn't DevOps CI/CD enough for data?

C/SMOIDTM is the data analytics counterpart to DevOps CI/CD, and the presenter concedes it is the world's worst acronym. It stands for continuous self-service sandboxes, continuous meta orchestration, continuous integration and deployment, and continuous testing and monitoring. DevOps tools such as Jenkins and Azure Pipelines limit their scope to CI/CD and target software development toolchains, so they cannot orchestrate, monitor, and test the two pipelines data work requires: development and production.

Where to go next