On-Demand Webinar · 55 min

DataOps: The New Normal in Pharma

Chris Bergh walks through how four pharmaceutical companies apply DataOps across R&D, commercial, and manufacturing data: fewer data errors, new analytics delivered faster, and collaboration across teams working in different tools. Recorded June 2021; updated August 2026.

Presented by Chris Bergh

What you'll learn 6 points
  • Worldwide pharmaceutical sales run $1.2 trillion, with North America accounting for 46 percent of the global pharmaceuticals market in 2021. Pharma companies increasingly compete on analytic capability, but internal teams struggle to deliver on-demand insight because the data is complex and sensitive and the teams sit in silos using different tools in different locations.
  • Drug innovation is expensive and uncertain: only a fraction of eligible products win FDA approval, and it costs on average over $1 billion to bring a drug to market. The first few months of a commercial launch decide long-term revenue for a brand.
  • At Celgene the DataKitchen platform integrated hundreds of data sets into a unified star schema with more than 10,000 automated tests, absorbing over 100 schema and data changes per week with very few errors or missed SLAs, run by a staff of seven data and DataOps engineers.
  • A commercial launch analytics team owns roughly 24 recurring reports and activities, from forecast tracking and launch tracking to stocking, field activity, incentive compensation, payer and plan performance, REMS, and a demand-based P&L, delivered hourly, daily, weekly, and monthly.
  • A commercial pharma data mesh splits the business into three domains that each dominate a different lifecycle stage: NPP, meaning non-personal promotion such as email, web, and radio, matters pre-launch and during growth; the Physician domain matters during the first years after launch; the Payer domain, which controls price through rebates, formulary, and tier, matters most in the mature phase.
  • Domain layers stack from raw sourced data through mastered data sets owned by IT, integrated data sets owned by data engineers, and self-service tools owned by analysts. Mastering is its own layer: there are one million physicians in the US, but a company's physician master holds only 40,000.

Slides

68 slides

Transcript

Show chapters and dialogue 9,177 words

00:00:00

Hello, everyone. Thanks for joining us today. My name's Beth Beverly. I'm the VP of marketing at DataKitchen, and I'll be the host for the webinar today. We have Chris Bergh, who's the founder and head chef and CEO of DataKitchen, and he joins us today to talk about how DataOps is the new normal in pharma.

So for those of you who are new to our webinars, Chris is a leader in the DataOps movement. He has more than 30 years of research, software engineering, data analytics, and executive management experience. At various points in his career, he's been a COO, CTO, VP, and director of engineering. He's also the co-author of "The DataOps Cookbook" and "The DataOps Manifesto." So just a few housekeeping items before we jump into it.

If you have questions, please enter them into the question box on the webinar control panel. We'll collect all those throughout the webinar, and we'll answer them during the last 15 minutes of the webinar in the Q&A session. Also, just a reminder, the webinar is being recorded, and so we will send that recording to all the registrants at the conclusion of the webinar.

So I think that covers all the housekeeping. So with that, without further delay, I will just hand it over to Chris.

Oh, Chris, we can't hear you. I think you're on mute. Hello, everyone. I hope you're doing well today. I'm going to stop sharing my screen so you can see the slides in their full glory. So thanks for that great introduction, Beth, and we're going to talk about DataOps and pharma. And so, I've been working closely with the pharma industry now for 16 years.

And so I think it's really just a spectacular industry. And how DataOps works in it is varied and rich and detailed. And so I'm going to talk through that. So first, I'm going to give a little high-level summary of DataOps and what the pharma industry is, and talk about an example of how it works in R&D. And then talk about sort of a process hub in commercial, and then something how that processes needs to be broken into smaller components called a data mesh in commercial.

And then a set of services we have that we have done in commercial. And then talk about how do you get an entire pharma company to adopt DataOps. And then a conclusion. So if you've been to our webinars before, one of the things that we really focus on is this idea of DataOps and that the work that we do, there's lots of problems.

And really the evil that we're trying to solve with DataOps is project failure, and that most data and analytic projects don't work. They're late. They don't get what the customers need. And we're doing all this amazing stuff with models and algorithms and pipelines and data, but it's really an upstream problem. And that we've got to focus on

cycle time and error rates, and really these problems that DataOps solves. And so to define DataOps, it's a set of technical practices and cultural norms and architecture patterns that enable analytic teams to deliver new insight, new data sets, changes to insight, changes to data sets rapidly. And to also run their production lines like Toyota and not American Motors, with very low error rates. And not have people sort of be at their throats and attend a lot of meetings with collaboration.

And then being able to measure the processes. And so we've talked a lot about DataOps, certainly willing to talk more, but today that's sort of our quick intro to DataOps. And then the next one is, so this pharma industry. So what's interesting is, the world's GDP is about $73 trillion, and I checked this morning by asking Alexa. And it turns out worldwide pharmaceutical sales are like $1.2 trillion. So what, it's almost 2% of GDP is pharma sales. And North America was the largest market for about 46% of it in 2021. So it's a big, huge market.

And it's actually really incredibly rich in data.

And in a lot of ways, pharma companies have always competed on the basis of their molecules, small molecules, biologics, their chemistry, how well they affect disease. But increasingly, pharma companies are competing on the basis of their analytic capability. And a lot of internal teams are struggling to deliver the sort of high quality and on-demand insight that their teams require.

Why is that? Well, they've got really complicated data. It's sensitive. It's a complex marketplace. There's a lot of organizational complexity with different teams and different tools, and the data's just hard. And so let's look at how pharma companies are

00:05:00

organized. So there's very different functions in pharma. One is there's research and development, where someone creates the molecule, gets it through clinical trials, and then there's sort of a commercial aspect, marketing and sales, who markets it, sells it, gets it on formulary, has a government in a country put it on its list. And then there's sort of manufacturing that actually makes it, and then finance.

So it's a complete set of capabilities. And a lot of people who work in pharma kind of stay in their lane. If you've been in R&D, you sort of stay in R&D. It's rare that you hop over to commercial or manufacturing. And a lot of people who've been in pharma have stayed in pharma their entire career.

And how companies are organized is very different. In some cases, they're by function, like there's R&D and commercial and manufacturing and finance and supporting. Other times it's by product line. So for instance, there's-- And that's by therapeutic class. So, we worked with Celgene quite a bit. They had hematology and oncology and INI. And sometimes it's a matrix organization.

And how data teams and analytic teams map into that Is as complicated as other cases. Sometimes they're aligned by product line, sometimes they're centralized, sometimes they are aligned by function. And in general, I think that is people who do data in pharma, there tends to be groups that focus on marketing and sales, research and development, and manufacturing in our experience, and so they tend to be more functionally aligned.

And the data is actually really different in each place. So, the goal of an R&D organization is to discover products and bring them to market through clinical trials. So they've got things like genomic data, experimental data, laboratory data, chemistry data, clinical trials data, image data, and lots of unstructured or semi-structured data. And then in commercial, at least in the US and internationally, there's sort of syndicated sales data about sometimes physicians or payers or plans or products.

There can be de-anonymized patient data, claims data, what's called NPP or non-personal promotion marketing data, Salesforce data, and lots of unmastered structured data, real-world evidence data, and the list is very long. And then in the manufacturing side, well, there's operations data, like from SAP, there's IoT data, regulatory compliance data, process control data, GMP data.

And then in finance, it's a little bit more normal. There's some SAP data, HR data, et cetera. So they're very different landscapes in data. And so let's talk about the first one in R&D. And I think one of the most important things from a business perspective is that, and Tamayini Carl says this well, "The most important battles grounds in pharma are compressing the clinical cycle time," that is getting drugs through clinical trials, "and demonstrating the value of the different therapeutics." I think that describes R&D well.

And so, as an example, drug innovation is just really hard. There's a lot of compounds that go through a lot of work to get to phase one, phase two, phase three FDA approved. So it's sort of 50,000 to one here. And it costs almost a billion, sometimes people say even more to bring a drug to market. And so these are very long, expensive, and if you can accelerate this process, if you can use data to help us identify more compounds, accelerate the way through the compounds, you can actually then get to the commercial phase where you can make your money back.

And so, one case, and this is an example, and we're going to give some examples here that when we work with customers, although we're contractually obliged not to disclose who they are. And so here's a case of a customer who wants to bring a self-service DataOps platform for R&D. And one of the challenges with R&D groups is they're varied, right? They're sometimes very technical, a lot of times very technical, very independent, and they want to use their own tools.

So here's a case where they want to use an on-premise stack, a GCP stack, an Azure stack, an AWS stack. And then there's all the tools that varies from ETL tools to viz tools to custom tools and different teams across to use them. So how do you get some consistency and reuse? How do you share best practices? And how do you get these people to start to do DataOps? And here's just a concrete example of that complexity.

So, in this case, imagine you've got a New Jersey data and analytic team, and of course, I've hidden the information here. And there's a large Spark cluster and a best-of-breed tool chain with stream sets and other tools, and it's sort of very high value, very big drug development data. And then there's another group that's also in R&D.

They happen to use Azure, and they've got other tools like Databricks and

00:10:00

other databases, and Azure's got its own set of data analytic tools, and they've got varied research sets. And how do you get these teams to work together? Because they're too independent. And how do you have, for instance, the work that that on-premise team does, how do you coordinate that work with the team that's in California? And just how do you handle things like schema drift, or is the data good?

How do you actually coordinate the work? And so what happens is companies try to solve this with meetings and documentations, and it slows things down, and inevitably, there's problems and team conflicts and people start thinking the other group is a challenge. And so for us, one of the things that we think is important is to be able to have a process that sits across all these tools.

And that works in both places. And that process follows our DataOps principles of being able to test and monitor, have a shared abstraction that has local control, but as well as centralized governance. And allows them to use whatever tool they want to use, but make sure that it's tested and monitored, that the code is in Git, and that we keep track of all the processing for governance and for compliance reasons.

And so that's one example of just within R&D, trying to deal with the varied teams and their tools in an R&D context. And it's hard because scientists are just very smart and very independent. And so, trying to get these teams to work together and have a governed, tested process is hard. And so, just to end out, here's some quotes from Kurt Zimmer, who's head of data engineering and enablement at a CDO summit. And so he talked about distributed DataOps, and he said, one of the things that struck him is sort of like how long it took to do anything, right? How much horsepower just to create something.

And he said something that I agree with, "It struck me that DataOps has the potential of being one of the transformative capabilities. And not just from a tech perspective, but from the lens it brings on how to approach these complex data environments." And he brought up the idea of sort of taking it out of the craft world where we used to sort of build cars by hand with hitting nails, and then we went to mass manufacturing, and now we're on lean manufacturing, and we're trying to move, in a lot of ways, the data and analytics from that.

And now let's talk a little bit about the world of commercial. And so we've got a couple of examples and talks. And so in a lot of ways, DataOps is trying to do what this quote says. How do you get an organization to be nimble and agile? And how do you enable them to be self-service?

And so if we look at the commercial product life cycle,

a drug has a lifespan at which it's protected by patents. And during that lifespan, before it goes off patent, it has a launch phase where it grows very fast and a maturity phase, and then it has a phase where it's either a generic enters or it's going off patent. And really the first few months of a brand can be very important to its life cycle. And so there's an entire industry around how to maximize the commercial success of a product. And you may be a fan of Bernie Sanders or not, but the fact is that the commercial life cycle and the amount of money that a brand creates is poured into research and development of new therapeutics. And so the success of a drug also allows the success of further drugs.

And so in a pharma organization, especially during a launch or during any phase of a product's life cycle, the entire organization demands insight. And so there tends to be an analytic team who is kind of under marketing and sales, may report up to a brand team, may report up to a product line, may report up to a head of sales or marketing.

And a lot of people are actually interested in the data that they do, especially during a launch. And the analytic team is really focused on the product trajectory. How many scripts are written, sort of, physician productivity or penetration, how fast the brand's growing. And the insights have to really come fast and furious. It could be hourly, daily, weekly, monthly.

And the data that they use just is incredibly varied. It comes from marketing data or call data or data that comes from companies like IQVIA or Salesforce, and trying to make sense of that, especially since the data itself isn't always-- It may be about physicians or payers or products, but a lot of data sets are loosely coupled in the sense that they don't cover the same number of physicians. The payer names may be different.

00:15:00

They're about the same or not. And so a lot of analysts are trying to sort through whether the data makes sense, whether it's predictive, and whether it can answer their business questions. And just as an example, and this is a bit of an eye chart, just the amount of questions that come for different reasons.

Like for instance,

sales may have only one thing that they're interested in. Are they making their number? Are they creating a forecast? Are the sales team meeting the forecast? But also has to do with can the products get to market? So are the products being stocked? And then is the field actually working according to the marketing plans?

What's the field activity? Are they going to make their incentive compensation? And then there's the whole part of, since pharmaceutical products in the US aren't sold directly, they're paid for by payers and plans, what kind of claims is happening? And then there's just questions that happen all the time from analytics teams, very quick questions. And then there's what's called non-personal promotion or direct-to-consumer, all those ads that you see, that you hear on the radio, how those are doing.

And then there's just special reports that may come from specialized data sets like REMS data, or there could be P&L and finance reports, or data that comes from specialty pharmacy or real-world evidence data, or longitudinal patient data. So there's a lot of data, a lot of reports that have to be put together. And so some ways, you wonder when people will spend years trying to get good at understanding commercial data.

And so an example of a success at one of our customers, and this is when it was Celgene, there was a brand called Otezla, and it actually became one of the rare things that only happen once or twice a year now. It became a billion-dollar brand. In order to make Otezla successful and keep that graph going up rapidly during its launch, it was about integrating hundreds of data sets, literally hundreds of different data sets.

And being able to sort of master them, kind of analytically, get a quick master, maybe not the world's perfect master, but enough to answer business questions. And do that in a way that had very, very few errors and very few missed SLAs, because when you have thousands of people looking at your data, it really is not great when things don't go right.

And to do that, we had to build a system that has sort of tens of thousands of automated tests and allow hundreds of schema and data changes per week. And built on really a staff of seven, which is incredibly low, of people who did the data work and the data engineering. And then on top of that, there was another sort of staff of 10 or 12 who actually did the analytic work. And so it was a very efficient, very fast, very high-quality way of doing work that gave great insights to the Otezla brand and made material impact.

And so, particularly just from an order of magnitude, I know companies, instead of having their sort of marketing and sales analytics done with what is about maybe 20 people, they have that times 10 for a brand of similar size and are less effective. And so that's where this benefit of DataOps is. Instead of trying to have more people or cheap people, you actually automate the people and automate their work so they're constantly trying to work on the most valuable thing.

And that's the kind of benefits you get from DataOps. And so really the idea here is to allow fast changes to really support investigative analytics, ad hoc questions, and then turn the results of those ad hoc questions into high-quality production deliverables. That's the game in pharma, and certainly during a launch or any phase in a pharma's product.

And so to do that, we think one of the most important things is to allow teams to work together and to focus on the processes that act on the data, in addition to the data itself. And so this idea of how do you get analytic agility, how do you respond during a launch, during the whole life cycle of a pharma product?

And so how do you make that happen? And so some of the challenges with commercial analytics are, some companies, it's not too hard now to get the data in one place, right? They have a cloud, they have data sets that are integrated. And so a lot of IT organizations are kind of saying, "Okay, we're going to get the data in one place for you." The problem is that's sort of 10% of the work.

What you really need to focus on are these rapid processes that act upon data, because there's always the next question. And how do you turn that next question into insight? And sort of that ends up with sort of firefighting and stress and inconsistencies. And if we take just some quotes, a lot of times people, commercial analytics start

00:20:00

talking about it's a last mile problem. And, how do you get data ready for analytics and ready for your analyst to make questions? Because the questions constantly change. And how do you get the people who are analyzing the data for the brand or for sales to not be so reactive, but to be able to get ahead of their customers, and to do that in a way to be able to answer what if questions successful. Because the analytic team is kind of the tip of the spear.

They are always looking for insight. And one of our customers has the metaphor that in commercial pharma, the company you're competing with is Amazon. Not because of AWS, but because Amazon can ship a package across the world, or across the country in one day. And your brand team wants the same insight. They want their questions answered the next day, and then they want that answered question to be given out to the field that next week.

And so how do you have that sort of insane velocity of finding answers to questions, delivering insight, and turning into production value? And for us, the answer to that is a process hub. And so there's a lot of cases where organizations will build a lake that has lots of data, and that infrastructure is in place.

But really what we need, what organizations need, is a place to centralize their processes that act upon the data for the purpose of rapidly answering questions and rapidly turning those into sustainable, high-quality production insight. And so to do that in a way that automates and not adds consulting work, and do that in a way that reduces errors, and do that in a way that allows sharing and reuse. And think of it as a way to curate and manage your processes that act upon data. And to get that sort of commercial pharma domain expertise in one place in your organization instead of on different people's desktops. And to give a place that's sort of controlled and trusted by your commercial pharma team. And there's lots of users of it, right?

It could be people doing your data engineering, there could be outside pharma consultants, there could be data and analytics leaders, there could be people doing the analytics. But if we look at this process hub and take an example of it, the main idea is if you centralize the process, you can generalize the process. And by generalizing the process, you can save time and effort, and therefore that same time and effort you can use to get more insight to your customers. And we do it through all these bullet points, right? Having a hub for multi-team, multi-vendor, multi-location coordination by having governed reusable process components. By automatically detecting errors.

By working on top of your existing data and allowing you to use your favorite kind of tools. If you're a SQL person or an Alteryx person, if you like to use Tableau or Looker. And think of it this way, it's an open, secure platform for anyone who's got permission to access the data and change the processes used to produce insight from that data. And another way to look at it is that your IT organization may have a permanent data lake, but what data and analytic teams need in commercial are abilities to rapidly create insight from that data.

And what our process hub does is store the recipes, the tasks, the processes that act upon the data using whatever tool you want, the production runs of that, and then allowing you to be able to do rapid development. And the way that we do that is sort of build these sort of analytic-ready databases for people to use.

And on that, people can do their weekly, daily production ad hoc work. And so the benefit here is that this part in the middle is built for rapid change, built to turn rapid change into high-quality production insights. So it's built on a process-first principle, that those processes that act upon the data, the tests, the recipes, are shareable and manageable. And as an example of that, here's a case of a biologic launch.

And so in this case, there's 70 data sources, right? Complex cadence of daily updates, weekly builds, daily builds, and then there's a whole bunch of checks on that data because biologic data is very complicated and not clean. And so has the data arrived on time? Do we have enough data? Sort of profiling and creating automatic rules-based quality checks.

Being able to actually take that data and put it in a data catalog for people to have. And then at the end of the day, build sort of processed data, aggregation, stars, views, tailored to a whole bunch of deliverables. A bunch of Tableau workbooks could be built on it, email notifications, exception reporting, standard reports, however you want to use it.

And so, it's rapid not only because the velocity of updates is rapid, but

00:25:00

also the ability to change this is incredibly rapid. So again, dozens of changes to any part of this process a week are allowed with very low errors. And so building this kind of process hub that allows us to do it, allows teams to have these sort of benefits. So the first is that the processes kind of create data designed for analytics.

And if the data is designed for insight, then the commercial pharma analysts can use simple self-service tools to get at it and pivot to the data. And then being able to focus on the process on top of the data allows you to automate it, use it, build a factory for insight. And what that means is just less dollars spent on consultants and staff time.

And it's not about sort of replacing the data lake or the data hub. It's sort of something that sits on top of it. It's not about replacing your favorite tool. If you're an Informatica shop, or you're running on Snowflake, or you're running on Databricks, we work with those tools and sort of about extending the IT investment. And so the main business value is sort of replacing an army of people with a process hub and allow commercial analytics to produce more insight faster at the speed that the customers demand.

And so that's one way to think about what DataOps is, is as a process hub. And so just to boil it, this section down, the main idea, processes that act upon data are as important as the data itself, which is a little heretic idea, but that's, I think, we call it DataOps, but really it's processes that act upon data operations.

And so let me keep going. So another case is we're going to talk about something called a data mesh and how centralizing processes are great, but you don't want to create a monolith, a mesh, or you want to break those processes up into more manageable units. And so here's a quote that backs it up.

"So much of what we do involves business questions that are fire drills. Executive wants answers as quickly as possible. The infrastructure we set up with DataKitchen allows us to mix and match data in new ways so we can quickly get the answer to a question." And so partly what that means is that given the complexity of data, breaking that data set down into smaller components does that. And there's a new word. We've been doing the same idea for 15 years, but there's a new word around it called data mesh. And it actually comes from software development, just like sort of the idea of DataOps. So I'm a fan of it in the sense that building large centralized systems fail.

And they fail because data engineers, pharma data is so complicated that you need to know the data in order to be effective. So you can't jump from one data set to the other. You can't jump from one brand to another brand to go working on biologics. And so the data itself, data domain knowledge matters, and sort of one size fits all doesn't.

And of course, this has led to sort of endemic project failure and endemic sort of, we can't do it in-house, so we're going to hire Zias or Biggoo or Axstria and just write them a check and they're going to do it all. And so I think one of the ideas here is that if you reorganize your team into smaller groups, more discrete components, more discrete teams can produce value.

And this is inspired by sort of domain-driven design and software, being able to get a way of what are called monoliths into smaller, more component-based development. And that's the main ideas here. We can take this complexity-reducing idea and apply it to pharma. And so what exactly do we mean by that? Well, there's different domains in pharma and analytics.

And so during this growth curve, pre-launch, growth, maturity, think of them as an example, three domains. Kind of the non-personal promotion domain, emails, websites, and visits, even radio ads. The physician domain, in the US, that's the good-looking people who go to your doctor. Sales, claim data, anonymized patient data, and then payer data. And so each one of these have a different importance during the growth of a pharma product.

Payer data tends to be more important in the mature phase because trying to extend the size of this curve, but it's not unimportant in the pre-launch phase. NPP data is important pre-launch, but it also can be important. But think of the data domains as these, as a way to group it. And it's just they're very complicated, and they have actually different teams. So there's a non-personal promotion marketing and sales team focused on digital and ads and websites.

There tends to be a lot of people in the physician domain, who are working with salespeople who are knocking on doors. And then there tends to be people focused on payers and trying to reduce

00:30:00

rebates and cost and kind of essentially controlling the price of the product. And so you've got different teams in the business side of pharma who are focused on this. So there's different data sets and there's also different teams. And of course there's some feedback back and forth, but in general, in my experience, most pharma companies, once they get a certain size, they'll have teams devoted to digital or non-personal promotions, to marketing and sales, to payers.

And our belief, if you look at the sort of rich set of data that goes in- Whether it's sub-national payer data, whether it's census data, stocking data, source of business data, longitudinal patient data, claims data, ERP data, data from your website, forecast data, each one of these data sources is largely different and has different perspectives on it, different ways to mix real and projected data.

And they may not match one to one. So your rebate data may not match directly with your sub-national payer data that you go to do. Your formulary data may not match to the payers. You may have different ways of looking at payer data that comes from source of business versus another source. And so what that means is, one of the ways to organize your process hub and organize the teams working on the process hub is to group people into domains. And in this case, this is what we've seen and worked with, is one case is the sort of mastering and small files foundation.

And this way, trying to get at who are the, for instance, the called on universe. Of the one million physicians, maybe there's only 4,000 who the company is really targeting for their particular brand. And then the sort of main data warehouse, integrated data layer, the sort of facts and dimensions. And then the self-service with analyst tools trying to create their own data. And so there's a lot of cases where

self-service analysts are using tools like Tableau. They're mixing in small data sets. They're developing their own segmentation models. And so how do you coordinate this? If you see it from the domain layers here, I've got all my raw source data. I've got sort of master data, maybe my physician data that comes out of my MDM system.

Maybe other domains like target lists, product market baskets. These are all good to have one version of. And then those actually sort of integrated into data sets like a payer domain or a physician domain or an MPP domain, where they're related but independent. And then the brand team may actually be using data from one or more of these.

And so here's an example of the sort of complex relationship of domains that I'm talking about. So if your source of business change, that may mean it's your physician master change, which means the physician dimension and the payer dimension in your integrated data sets change, which would mean your brand team reporting and also probably your field reporting would change. And so there's a coupling between these domains.

And so one of the reasons why it's hard for pharma organizations or any organizations to adopt this sort of domain-driven design is because trying to master the couplings between how these different teams work is hard. And so what that means is when you think about a domain-driven design or a data mesh, you need to think pretty intentionally about how you're going to coordinate, because there's different people working in each one of these things and different processing happening, and so happening at different times. And so by allowing people to work independently, you maximize their control, you maximize their expertise, but you've also got to think intentionally about how to coordinate each domain. And so our software luckily helps to do that.

It allows you to say,

you can query one domain to the other. You can link the processing of one to the other. You can have an event cause one domain to link to another. And then you can even have the development process linked. And so thinking of domains allows you to then create this way for your data engineers and your team to really know the data, to really add value. However, you've got to start thinking about how each team works together and linking of them.

And so, the main idea here is that, and this may be a little complicated, but process and team monoliths fail. All data engineers are not perfectly fungible. All team members are not perfectly fungible. Sort of break them into smaller link domains and win. And I think this way of handling a very complicated data environment in commercial pharma has a really good, at least in our experience, a really good way to match that mix of control, low errors, and high velocity changes.

So the third part of our discussion here in

00:35:00

pharma, and I've got a lot to say, so hopefully you're bearing with me, is one of the ways that DataKitchen has helped is that we actually sell a software product that enables a data mesh, enables a process hub, enables DataOps, but we also have people that help. And so, one of the ways to think of it is that we have some services that we wrap around our software.

And so, we take it from the perspective of, in commercial, what does it take to be the sort of VP of commercial insights? What does it take you to be successful? And our feeling is that, as an analytic leader, you should have your span of control. And from your data sources to your marketing and sales customers, you need to sort of control all the work that happens in there. And for us, what we tend to focus on in green is you need some data engineers to work with your data. You need data analysts.

You need a process hub, and you need DataOps engineering. And so we've offered both data engineers and DataOps engineerings to our customers. And we do that optionally. You don't have to use them. But some customers have found them, and in fact, some customers have worked where our data engineers are essentially replaced. We get them going, we build the process hub, and then our data engineers disappear.

And we stay with our DataOps engineers to do the daily builds. In some cases, we don't have either. But if you think about what an analytic leader needs, they need to control their destinies. They need to control all these things. And I think the best way we think of our services is that we wanna make our customers successful.

As a data engineer, you wanna make your data analysts and data scientists successful, give them superpowers. And as a data analyst and data scientist, you wanna make your business customer successful. And so we've wrapped our software in a bunch of services. And so one way to think of it is that the gorilla in commercial pharma is a company called Zeus, who has been in the field for a while and basically does everything related to commercial analytics.

And from our perspective, we're not trying to be a little Zeus, but what we're trying to be is a partner on your journey to do DataOps. And maybe part of that partnership is you need data engineering for a while, maybe you need it forever. Maybe you, of course, need a DataOps platform, maybe you need DataOps engineering for a while.

And so one way that we do that is through what's called a managed service. And technically, our managed service involves building integrated data lakes, having recipes that build the integrated database, this sort of process hub, and building it out to all the different people. And so, as an example, working with a Cambridge Mash biotech.

We did both data engineering managed service and DataOps managed engineering service, with our software. And so there's a lot of complexities in getting data from shipments and interactions and syndicated data using the Veeva network for mastering, integrating, and doing what's called an analytic master, being stewards of data and building that integrated dataset, which is really, think of it as a commercial data warehouse and master data management for this Cambridge biotech.

And what we did is kind of build these datasets, sort of link stars and dimensions about commercial and of course, normalized tables that actually help for specific analyzed questions. And being the source of the best integrated data, we end up being kind of a feed and the system of record. And so being a hub for other companies to come in and do that work. And so, as an example, one of the things that we do is, in order to get an answer, we need to link records. And if you think of the universe of physicians in the world, there's a million physicians in the US, but maybe your MDM program only covers 40,000. Well, you still need to link all these different datasets in order to get insights. And so one of the ways that we do is build an analytic master, which is kind of related to a master, but it's actually helping companies answer high-level analytic questions.

And also what we do when we wrap our services in kind of a complete process. Since we're believers in Agile, we follow an Agile process. Everything's a ticket. We build a data dictionary in Confluence that's updated every week or every time the database is updated. We track our time, and then, of course, being good data nerds, we track a whole bunch of reports on cost and budgeting, on data error details, data errors by source, data source reports, and pulse reports, tickets open and close, so you can manage our consulting. And the

00:40:00

last thing that we do is we've got, since we've been doing this for a while, we've got a set of accelerators. So existing schemas, testing tools, automatic schema documentation, templated recipes, best practice reports, and automated sort of DevOps and security. So it's not like we're starting from scratch, but the idea here is our goal really is to accelerate your DataOps journey with services and then allow you to take over when you wish.

And so, we think that that's the most important thing, because I've seen too many cases where a company hires an outside consultant, they work good for a few years, and then they wanna pull it in-house, and then it goes from in-house to out-of-house. And I think that one of the reasons we founded DataKitchen was to avoid that.

And so by building this process hub where you own the process, we can man and build some of the process for you and allow you to take it over, I think is a way to sort of break that wheel. And what happens here and what our customers say, and actually one of my favorite customer quotes here is, "Just let me say with respect to DataOps, I'm a believer.

DataKitchen has helped us completely transform our operations. We were able to measure our ROI by showing organizational dramatic cost savings, but we were able to help them better because we were more efficient and do good things quicker." And that's James Royster. So, I've got a lot here. So let me go on to the last discussion topic and talk about...

We talked about the organizational complexities in pharma, right? And that there's a commercial organization, there's an R&D organization, there's a manufacturing organization, a finance, and there's maybe a tech or enabling organization on one side. And so let's say you're a chief data officer at a big pharma company and you're a believer that this DataOps thing is gonna transform your organization.

But there's such different data and there's such different groups, and there's lots of people doing data work across all these organizations. How do you really get your influence, everyone in the organization to start doing DataOps? And so, we've been working with some pharma companies and some other companies, in different industries to make that happen.

And so one of the things is that we believe that DataOps, it needs to be enabled by technology, but it's also a journey that you have to go on. And one of the first things is to educate people about the possibilities of DataOps, that you can create new insight quickly across whatever tool that you happen to use, deploy to production, and run it with low errors, so you can add new datasets and make changes quickly and sort of break that.

And you don't have to have a lot of meetings to do it. And that sort of trio of make changes fast, run with low errors, have low meetings, who doesn't want that? But there's also a lot of cynicism that that can actually happen. And trying to work with an organization to find demonstrative projects, establish a community of interest, demonstrate value, and then expand the scope across all those verticals is something that we do.

And I think the hard part is, really gaining the idea of DataOps. Although everyone would like the idea of, "I want to be able to get more work done faster using my current tools. I want to run things with low errors. I want to have less meetings." Who wouldn't want that? But how do you get sponsorship on this change?

Because it is a people and process change, in addition to building a technical platform to support DataOps. How do you get a roadmap? How do you assess readiness? And on the other half, once you've demonstrated it, how do you get a use case, and then how do you institutionalize that use case in some kind of COE or dojo? And then how do you organize your team?

How do I staff DataOps engineers? How do you measure success? And so all these key DataOps transformations we've helped our customers with, we've got partners that can help as well. And so the main idea here is that DataOps transformation is a journey of people and process and technology. And so, we want to be your partner on that and help you do it.

And so lastly, here's a couple of key quotes. So here's a case. What does success look like? And this is from during a DataKitchen webinar. It's about being able to build consistently across 90 product teams. It's about really driving scale, so you can deploy pipelines across the large number of peoples in different data technologies. And it's about the quality checks, the automation, the reusability and repeatability, and containerization wrapped into it.

And so what's the conclusion here? So I guess this quote kind of sums it up. "Analytics agility leads to business agility. When the data team delivers analytics rapidly and accurately, analytics do a better job supporting decision makers." And I think that's

00:45:00

simple. If I could say that that's the goal of DataOps, that's the goal of DataOps, analytics agility. And so what we've seen across different organizations, across different companies, we see a lot of successful companies going towards doing DataOps. And I think to be competitive, you must as well, because pharma companies that win will increasingly derive competitive advantage from their ability to manage and create value from data.

And what we've seen, I hope, is that companies using DataOps are able to streamline their efficiencies across their entire data organization, from research to commercialization, enabling them to deliver sort of high quality, kind of on-demand insight that consistently leads to successful business decisions. And really, if you think of the key benefits of adopting DataOps and DataKitchen, it's really about eliminating errors and building trust.

It's about reducing cycle time and increasing agility. It's about reducing collaboration and allowing the internal analytic teams not to sort of always have their back against some outside consultants. Bring capabilities back in-house. And so that's all I have today on DataOps in pharma. I'm going to take a breath and, Beth, we're going to answer some questions now.

Yes. Great. Thank you, Chris. Thanks everyone, for listening in. And now we have about just under 15 minutes to answer some questions. So if you have any, just I encourage you to enter them in the question box and we'll get through as many as we can. So, Chris, you talked a lot about, or somewhat about, DataOps engineering. When should an organization consider bringing in a DataOps engineer?

Oh, I think you're on mute, Chris. Yeah, I look at it this way, is one of the core tenets of DataOps is that these operational things, testing, meta orchestration, deployment automation, collaboration, code sharing and reuse, infrastructure as code, are worthy of investment. And so how much investment should a team make? And so I think 15% of the team's effort should be put on that. And you could argue that it's even more, because if you look at software engineering organizations with DevOps, they average sort of 28% of the team's effort is on this.

So 15% isn't too much, and that should be a role that everyone should do some of it, in my mind. If you're a data engineer or data scientist, part of your job should be to write some automated tests. But I think when your team gets of a certain size, more than a few people, I think it's really essential to centralize some of that work in a role.

And so if your team's four or five people, I think that's the time to start thinking of full-time DataOps engineer. Because the whole point is by spending that 15% of your team's time on building a system, you're able to do more work faster, and sort of giving operations and collaborations its due. So I think kind of when your team gets sort of five, six people, that's the time to start having a DataOps engineer full-time on staff.

But really it's about the allocation of effort.

Great. Thanks, Chris. So how can DataOps help support governance practices?

Well, I think it depends on what you mean by governance, right? And so one way to say data governance is the use of things like a data catalog and data lineage. And so there's a number of companies like Collibra and Alation and tools that do that. And so I think one way that a DataKitchen does that is to help make sure that the data catalog and the data lineage are up to date. So if you're going to alter the schema and deploy that using DataKitchen, well, that schema change could end up in a change in your BI, which should end up in a change in your data catalog. So deploying that sort of governance as code is one aspect.

But then if you think about it, what do people want with governance? Well, they want to go to a place that says, "I want to help understand my data. What's there? Where did it come from?" And data catalog tools are great to do that. But there's another part that they want to also understand, "Is my data fresh?

Can I trust it?" And that sort of process lineage comes from a tool like DataKitchen. Here's all the tests that were applied. Here's when the raw data arrived. Here's the last time that data was refreshed. Here's a report of all the tests that were applied. And so I think it goes hand-in-hand, data governance and a system should allow users to go in and understand where their data is, what does it mean, where'd it come from, when's the last time it's refreshed,

00:50:00

and do I trust it? And I think the latter half of things are parts that a tool like DataKitchen can provide. Okay, great. So you mentioned domain-driven data stores and data lakes. What's your take on the pros and cons, and what's your personal recommendation?

I'm more of a fan of domain-driven design, especially in pharma, because of the complexities of data. And so data engineers need to understand the data sets they're working with, and that takes time and experience. So it's much better if they work in the domain that they're working in. And so depending upon the size of your organization, that may be hard, but getting them some expertise.

So that means that working in a data mesh or a domain-driven design way, I think, is better. And of course, it depends on the size of your organization, right? If you're a 200-person pharma company and you're one of the three people supporting commercial analytics, well, it's hard to have domain-driven design when you're only three people.

But as the team grows, and as the brand gets more successful, having more people focused on data sets and aligned to their customers who are using that data sets. Because at the end of the day, having a data mesh design is all about focusing on delivering customer value. And by breaking that monolith into smaller components, giving your teams expertise in the business questions and the data, you maximize your chance of success.

Okay. Awesome. Thanks, Chris. And then we have one more question, and it seems like a good question to end on. What is the best first step if you've decided to implement DataOps?

Well, call DataKitchen, I think. But beyond, I guess that's really it. No, one of the first steps I think is that I'm a believer in sort of who you are. So if I'm an individual contributor, the first thing that you should do to do DataOps is just start doing it today and start writing some automated checks that run production against your data, just to make sure you get your error rate down.

Another thing I often recommend is have a meeting once every three weeks to look at every error that's happened in your system and keep a list and try to fix one error. Have a quality circle. They're very simple things that you can do to start the change from sort of avoiding the shame and blame to improving.

I do think the first step, though, from a manager standpoint, is to try to socialize the idea that these things are worthwhile investing in, and that the benefits of investing in DataOps and data operations will pay value in terms of lowering the cost per question, lowering the latency to get insight, and maximizing your team's productivity.

And I think those things are, from a business standpoint, you can explain to your manager and starting to build the first kind of demonstration project that shows that this actually works. And that what I find is once people start working in a DataOps way, they don't want to go back because that sort of focus on customer value, that focus on fast iteration and automation, means that their life is just much happier. And they're getting more done, and they're having more impact.

And that, I think, is

one way to start. Sort of start really small tomorrow, and then start putting together a demonstration project to get people bought in that this is the way to work. And then also think about a little bit going on that having DataOps is a bit of a mindset change when you're dealing with a larger organization and trying to get people enrolled in that mindset change, sort of buying the religion of DataOps, I think is a good way to happen. And I've seen that happen a lot in sort of the DevOps world, where people pushed back for years in the early aughts that DevOps mattered, that Agile mattered. And now, kind of about 2010, 2015, the world shifted and DevOps became the default. And I'm starting to see that in

pharma, that companies are wanting to go to DataOps processes as default. All right. Well, thank you, Chris. That is all the questions we have today. I want to thank everyone for joining us. I hope you found this really useful. Thanks to Chris again for all his insight. We will be sending out the recording and the slides in the next 24 hours or so. So as soon as those are ready, you can be on the lookout for that in your email.

Don't hesitate to reach out to Chris or I if you have any additional questions. Chris.seberg@DataKitchen.io or beth@DataKitchen.io. So, we hope you enjoyed the webinar, and I hope you have a great afternoon and evening.

00:55:00

Thanks, everyone. Thanks, everybody. Bye. Thank you, Beth.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

Why is pharma data harder to work with than data in other industries?

Pharma is organizationally complex, with R&D, Commercial, Manufacturing, and Finance functions that each hold very different data and rarely trade staff between them. R&D holds genomic, laboratory, chemistry, clinical trial, and image data; Commercial holds syndicated sales, anonymized patient, claims, non-personal promotion, and sales force data; Manufacturing holds operations, IoT, process control, and regulatory compliance data. Different teams use different tools in different locations, on-premise and across multiple clouds, so there is little consistency to share or reuse.

What is a DataOps process hub in commercial pharma?

A process hub sits on top of an existing IT data lake and holds the automated, governed, reusable processes that turn raw feeds into analyst-ready data sets. It gives the commercial analytics team production runs, rapid-development sandboxes, reusable libraries of recipes and ingredients, and automated tests, all controlled by the analytics team rather than by IT. The goal is to lower the cost per business question so ad hoc requests get answered faster without hiring more consultants.

What are the domains in a commercial pharma data mesh?

US commercial pharma splits into three domains: NPP, or non-personal promotion, covering emails, website visits, and advertising; Physician, covering doctor sales, claims data, and anonymized patient data; and Payer, covering plans, rebates, and formulary. Each domain has separate sources and different cycle times, and entities such as physicians overlap across all three. Data does not reconcile cleanly between them: subnational physician data purchased from IQVIA may not match claims data one to one, which may not match payer data, because of supplier differences and timing projection algorithms.

How do data mesh domains communicate with each other?

Domains exchange five kinds of link. A domain query asks when a domain last updated and whether the run succeeded, or asks a domain to prove its data is good with test results. A process linkage hands off control and parameters between domains, an event linkage broadcasts completions, errors, and warnings, and a data linkage covers a shared table such as a common dimension. A development linkage is the ability to re-create another domain in development, read and modify its code, and push a change to production.

What results did Celgene get from DataOps?

James Royster, head of commercial analytics at Celgene, reported dramatic cost savings that let the team measure ROI directly, plus faster response that helped the business capitalize on opportunities. He also described an almost complete absence of data errors that were not caught earlier in the process. A director of market insights described the outcome as a self-service data organization reaching from the marketing department to the sales reps.

How does a large company start a DataOps transformation?

The six-step sequence is Educate, Find, Establish, Demonstrate, Iterate, Expand. Education uses presentations, video, books, and analyst writing; finding a first project comes from in-depth discussions with individual teams guided toward their pain, supported by a DataOps maturity model; establishing a community of interest means wikis, Slack, and alignment with data and agile leaders. Demonstration runs a one-to-two-month pilot with real measurement, and expansion ends in a staffed center of excellence or dojo that sets common infrastructure, tools, and metrics.

Where to go next