On-Demand Webinar · 1 hr 1 min
The Celgene Story: Building a $1B Product Launch Success with DataOps
Rajesh Gill, Associate Director of Commercial Insights at Celgene, on how the team used DataOps to support the Otezla launch. Recorded June 2020; updated August 2026.
What you'll learn 6 points
- Bringing a drug to market costs $2.6 billion, according to the Tufts Center for the Study of Drug Development. The first six to twelve months of a product launch are decisive, because how fast a product grows during launch determines its lifetime revenue.
- The Celgene commercial analytics platform ran on Redshift feeding Tableau Online, with more than 1,000 dashboards serving hundreds of sales people plus marketing and executives. Inputs included syndicated data, sales data, Rx claims, specialty pharmacy, NPP events and campaigns, sales alignments, product hierarchies, and specialty mappings.
- On the DataKitchen platform, Celgene integrated hundreds of data sets into a mastered, unified star schema with more than 20,000 automated tests, absorbed over 100 schema and data changes per week, and had very very few errors or missed SLAs at low total yearly cost for hardware, hosting, software, and staffing.
- The engineering side of a new large data set followed six steps: build a scrappy star in a data mart, send questions to the data supplier while keeping analysts in the loop, add data tests, share the star with the analyst team for feedback, iterate over several Agile sprints, then release a solid star.
- The analyst side ran in parallel: build scrappy dashboards, feed corrections back to data engineering, show early dashboards to users, hold active build and design sessions making as many changes live as possible, and publish production dashboards at 70 percent done.
- The DataOps mindset shift is five substitutions: change fear becomes change velocity, manual operations become automated operations, hope for quality becomes integrated quality, hero mentality becomes repeatable processes, and perfection becomes 70 percent right the first time.
Slides
Transcript
Show chapters and dialogue 10,367 words
00:00:00
So good afternoon, good morning, and even good evening to some of you. Thanks for joining us today. My name's Beth Beverly, I'm the VP of marketing at DataKitchen, and I'll be the host of the webinar today. Today's topic is how Celgene built a billion-dollar product launch success with DataOps. We're very excited to have a special guest join us to share his experience.
Rajesh Gill is currently the associate director of commercial insights at Amgen. Rajesh joined Celgene in 2017 and held various commercial roles, including data strategy and ops, and commercial insights and forecasting. He played a key role in building the analytics capabilities for the inflammation and immunology franchise, using key insights to drive growth for the brand.
In 2019, Celgene was acquired by Bristol Myers Squibb, and the Otezla brand was divested to Amgen for $13 billion, which is where Rajesh now resides. So he will kick off the webinar by sharing the Celgene story. Then he'll hand it over to Chris Bergh, the CEO and founder of DataKitchen, to provide some additional background on DataOps.
We'll finish up the webinar with some Q&A at the end. So before we jump into it, I'd just like to cover a few housekeeping items. As I mentioned, we'll reserve the last 15 minutes of the webinar to answer questions. So as you think of questions during the course of the webinar, enter those into the Q&A box on your control panel, and we'll collect and answer all of those at the end.
Also, we'll be recording the webinar, and we will email a recording to all participants, so please be on the lookout for that email in the next day or so. We'll also send a link to the slides in that same email. So I think that covers everything. With that, you can take it away, Rajesh.
All right. Thank you, Beth. Hello, everyone. As Beth said, morning, evening, afternoon, depending on your geography. Very excited to share my perspective and this use case for this billion-dollar product that we launched back in '13. My name is Rajesh Gill. I'm part of Amgen now, but originally, this product laid its foundation at Celgene.
And my role within this organization was primarily for the inflammation and immunology franchise, leading its data strategy operations. And then eventually, I took roles into analytics and then further into forecasting as well. I have little over 12 years of experience all into the data side of the world. Prior to Celgene and prior to pharma, have dealt with the telecommunication, British Telecom, and some other domains like education domain as well.
Yeah, super excited to share some of the details here. So with me today, would be Chris as well, who'll be sharing towards the second half some details around the DataOps component as well.
Beth, can you move? Yeah, perfect.
All right. We can go to the next slide.
And, Beth, I know everybody's on mute. Can you confirm if you're able to see my screen and if you see the slides moving? The slide did not advance. I'm still seeing the speaker slide. Oh, there you go. Now it's moved. All right. Awesome. I think, yeah, there might be some delay here. So yeah, let's get right into the meat of the conversation here.
So in terms of this slide, so this talks about Celgene, which fortunately or unfortunately is no more in the sense that majority of the Celgene was acquired by BMS, and the product that I'm working on was divested to Amgen. So I'm currently part of Amgen team. But in any case, when all of this started, our goal was really trying to make sure that particularly for the immune and inflammatory sort of conditions, that we are able to reach out to the patients that require this, and then can get benefit out of it.
And again, just a disclaimer, these are all my opinions here, does not represent Celgene or Amgen in any way. So let's get into the presentation here. I think there's some lag there. All right. Awesome. So, the slide I'm talking about is the product launch here. So typically, in the specialty pharmacy side of the world, it takes about 10 to 15 years, maybe more, for a drug to be developed and be launched commercially.
And the average cost to market,
00:05:00
eventually once it is approved, if it is approved through phase one, two, and three, is roughly approximately $2.36 billion. So that's a lot of money. And with that comes responsibility, not just for analytics or DataOps, but across franchise, to make sure that it's commercially viable and we launch it with a big momentum. Because for majority of these specialty drugs, the first six to 12 months are really key in terms of how successful a certain product or a brand would be. The market landscape in which we were playing was a very competitive one.
We had, I think at the time, around six to eight drugs in the market. But Otezla had its own sort of specific benefits with the safety and efficacy profile that it came with. So, at the present, we're talking about roughly a little over a dozen products that it's competing with. But having said that, despite all these challenges ...
we were able to convert this drug into a billion-dollar, and now multi-billion product, hopefully over the next couple of years. And then so in this sort of fast-paced, competitive landscape, there's challenges for the healthcare community in general, the biopharmaceutical manufacturing industry, as well as the marketplace we play a role in. Whether that's around just the fierce competition or the complexity of the marketplace with the pharmacy benefit management companies and the insurance companies, and then just the whole logistics around how do we find the right patients, the right therapy.
And then on top of that, there's always, from a commercial standpoint, continued pressure on the margin, on the managing cost side of it, and then how do we make it viable on a long-term basis? So, with these challenges, I'll get into specifics to our span of control when we're talking about DataOps, data strategy, analytics side of the world.
Some of the key challenges that you see are more around how do we make sure that we have the 360-degree view. So, in a lot of times, and that was one of the key challenge for us as well, was the data sources and the IT systems, whether those are internal or external, were desperate and were in different places. And there was nothing that was joining them together to be able to produce analytics, first of all, and then to then provide that on a continuing basis where the team is not burning out, yet you're able to understand the brand performance versus its forecast. And then eventually as you launch and then there's pre-post launch different needs for each of the brand. And again, that's whether we're talking about here in the pharma industry for this particular launch or any other product launch.
There's differences once you are in the market. So how do we evolve our, not just the infrastructure, but just the overall methodology in terms of how you deliver analytics on top of that? So from a span of control where we all, on this call, are dealing with this from a data analytics or commercial insight standpoint, how do we make sure that we're able to connect with our marketing or sales customers or other cross-functional teams? And then how do we leverage what's available to us here from a data source standpoint and IT systems, and then still be able to tackle all these challenges?
So that was the key challenges as we had launched in. And then six months into it, and this is where we leverage some of the DataOps and the agile methodologies that we work with DataKitchen for this particular brand.
So in terms of what was our goal, so when we started this particular piece for this product, what we wanted to do was, some of the key challenges, like I mentioned, there's data all across the place. So how do we deliver that in a single story? And also to be able to do it in a timely fashion.
The senior leadership cannot wait for the performance on a weekly, monthly, or quarterly basis for
a couple of weeks, let alone couple of months to deliver what's required. So we wanted to do something where all the analytics and the data work that we do, whether that's reporting or the underlying data infrastructure, to be able to have it a fast-moving environment. Our requirements change week to week, even day to day.
00:10:00
So we wanted to make sure that we're not working in a model where you ask for a change in the data table underlying or trying to integrate a component. It takes ages. So we wanted to have a fast environment where we are able to not just do the regular high-quality production deliverables, but also to be able to support the investigative analytics.
So if a business question comes in, whether that's for specific from a patient standpoint or an HCP, the provider standpoint or payer or other components. So we wanted to A, build that 360 degree view, but also to be able to sustain it so that we can, A, deliver the results fast, and B, provide it with the high quality so that three different people, three different stakeholders showing up in a single meeting don't come up with different answers.
That was one of the key challenges early on, and that's what we were able to solve with our implementation here from an agile DataOps standpoint.
And the reason why you see the Amazon box right there, that's something that I've stole from my previous boss, is he always talked about delivering this, and the mindset he had was he was competing with Amazon Prime. And I think that's what DataKitchen delivered for us. Meaning, with Amazon Prime, nobody has patience, right?
We want that product to be delivered within next hour if possible. And in some cities or some countries, that is possible with drones. Won't go into too much of that detail. But the idea is to make sure that we're not waiting for any of the DataOps to be delivered in a month or even couple of weeks. What we were able to achieve was to be able to make these changes on a very fast-paced environment and within a single week, and to be able to deliver those results to the senior management.
Excuse me. So,
the guiding principles going in with these challenges was we wanted to figure out how do we rely on the expertise, not just on the experience. The team that was set up had tremendous experience in this field And then we were still facing these challenges. So we wanted to build on the expertise as well, which I think for DataOps, DataKitchen brings into the picture. And that being the foundation, what we wanted to achieve was make sure that not just the internal team, but then the partners, the vendors, the sub-teams that we work with, they understand what the brand strategy is, where is the product placed, so that when you do the DataOps, when you structure the data sets, when you then run these reports or then you finally do any ad hoc analytics as well, you understand what the business is trying to solve for, what the key challenges are, and where there's product strength or weaknesses so that you can cater to that.
So that was one of the key principles that we went into this project as well. And then the last piece is, the third critical component here is that we all know no plan survives first contact with reality. So we wanted to make sure that when we implement this, we understand that, and we have taken steps to put in place so that when things go wrong, we have plan B or plan C to make sure that there's redundancy built, there's QCs, there's all the processes where we're able to catch a issue or something that's missed from a source standpoint is able to be-- We are able to cater to that.
So that was one of the key principles as well going on. And the last is, just to make sure that we're able to then make rapid decision-making and then resolve these critical issues as they come. So those were the guiding principles as we went into the project for this particular brand from a DataOps and analytics standpoint.
And then, so this particular slide talks about the project overview. So this sort of talks about the landscape of the data that we're dealing with when it comes to the pharma launch that we're talking here. So whether that's your internal sales data in terms of how much the product you have shipped or how much product then eventually goes out from the wholesaler warehouse to then the specialty pharmacies or other places where the product is dispensed to all the syndicated data.
So the syndicated data here is we're talking about the IMS, IQVIA, or Symphony, which are these vendors basically that collect data from the specialty pharmacies, the central systems, the PBMs, and all these entities that can provide them the data. So they contract on manufacturer's behalf, and they build this across product
00:15:00
market level data that a lot of companies then leverage to understand how's your product doing within the whole market, whether that's your market share or other components as you're trying to understand where those other product strengths are versus yours to then make not just a strategic, but then also then the tactical plan. So we had the internal sales data, we had syndicated data. Of course, then all the sales hierarchy and then the reps that are then finally having these customer interactions and then eventually any of the non-personal promotions. So talk about the 360 degree view from a HCP or patient standpoint. We had all of that data, the prescriptions data, the claims data, the specialty pharmacies data, and all of that. So, like I mentioned early on, the challenge was how do we bring all of that together?
And, what we were able to then do was, from a technical standpoint, DataKitchen leverages Amazon Web Services. So we have a Redshift instance that's created for our particular product, which we then leverage for all the data needs. And then on top of that, we built a three-legged stool. I think the third component here is missing.
It's an older slide that I'm presenting here. But we had Tableau Online to push this across the sales or marketing organization within the company. And then our analyst, so my team was able to build that content. And then the third leg to this, so Redshift, Amazon, sorry, Tableau Online, and the third leg to that was Alteryx as well.
So that was also another interesting acquisition that we took early on as well, and we worked in partnership with our DataOps team here, where we were able to do experiments with the data merging or data manipulation in terms of what is required for certain business questions. And then when certain things were ready to go and we know that that's going to be required on a long-term basis, we were able to pass that Alteryx information over to DataKitchen team, and then they would bake that into the underlying data packs for us or data tables for us.
And the reason we took that approach is, A, to make sure that anything that can be standardized is done, and then also to be able to then leverage the platform's QC ability and the availability of the data. So, for the data, everything was done on a weekly basis. So come Monday morning, whether that's our sales organization or internal marketing team, they were all able to see the most recent data.
And then once you push that experimental analytics back to data layer, A, it was faster, and then B, there's less chances of error as well. So talking about the high-quality deliverables, that's how we were able to deliver on that promise. So that's the overall project overview. When in then, right before acquisition, we were talking about 1,000-plus dashboards that we had developed over the years, which one or the other cross-functional team was utilizing for this particular product.
So with that, I'll move on. I'll double-click on some of the sources. I think we touched on this. In the end, whether that's pharma industry or other industries, in this day and age, as we all know, data is not a limitation. We have too much data. So the problem is how do we pick all that data and get some value out of it?
How do we make sure that these are not different data sets and we're presenting insights differently, but then to be able to collate that whole single story together and then be able to present it confidently in front of executive team so that they're confident to make the decisions based on data. And then so, with this next couple of slides, what I'll do is I'll talk about the general approach we took over time as well to build this brand's analytics capabilities.
So in general, as my team would talk to the management, to the executive team from a marketing or brand standpoint, we were getting these questions on a daily basis. So when we launched early on, we had all the scripts data, the prescription data, and then there was the patient-related information that was missing on those scripts.
So we wanted to leverage the claims data, which in our world is essentially you get anonymous patient-level data collected by these syndicated data sources so that you better understand what patient types are more beneficial for your product, where they can utilize
00:20:00
that, and you can target that segment, or the same case from an HCP or provider standpoint as well. What sort of prescriber specialties should our sales team be focusing on? Where should you supplement that face-to-face interaction? If not possible, how do you supplement that with non-personal? So whether that's showing them banner ads or sending them emails with the product information. So, these are the kind of different things that we had to take care of.
And then for this particular example, we'll talk about some of the claims data set that we procured. So this was a new acquisition of the data. And the results were, and this is early on in the launch, the results were required in two weeks. So, in an all a lot of data, and we're talking about big data here, where you have billions of records of the patient's claims, not just within your market, but then any activity that is around the patient. So how do we utilize that data and then deliver results?
So the approach we generally take is we make a scrappy start here with the help of our DataOps team. And then the idea is to then at least have a very draft sketch ready so that you can perform some analysis, understand that data a little better than just utilizing the raw files, and then make tweaks to that.
Or adjustments in terms of the data structure and what are the sort of the key KPIs or metrics that you want to bake into the data layer itself. So we took the large data set on once the NDAs and all of that was signed. So delivering it in two weeks wasn't really realistic because it takes a long time to procure that data in-house and set it up. But at least in that two weeks, we were able to get all of that work done and have a draft data sets ready. So we took that scrappy start and then over time evolved it as the brand needs grew. And then so over time, you see the metrics or the KPIs that are within that star building up so that you can answer not just one specific or set of questions, but then able to leverage that star schema to deliver multiple questions and over time, look at the data differently depending on the brand's need.
And then, so with that, once that scrappy data pack is ready, the star schema, which generally, with our team takes about, a couple of days to be built. Our analysts from there on take control, and they would join to Tableau in this case to that data to be able to then produce some of the insights, and then push it towards the business users.
And then the idea is you tell the business users that, "Hey, we'll have something ready in front of you in a couple of days to a couple of weeks, depending on the need," and then build out from there and then make changes, and then those changes are then pushed into the data layer like I mentioned early on.
So that's what we were able to achieve with this claims data analysis that we did. We were able to create that scrappy star schema, and then on top of that, we build these dashboards, and then as that need evolved over time, we updated the underlying data sources and then enhanced our dashboards as well over time.
So along with all of this, there's always the need of how do we make sure that we can take care of the scalability aspect of it, takes care of the data governance side of the thing, the data quality aspects of it. So in that case, the platform was able to deliver us the data errors, the QCs performed. So I think we have over 1,000 QCs that we have built with the team where if the data is not in a given range or if there's anything missing in some of the mandatory information that we typically utilize, we can bake that into the platform and then we get a weekly update of all of these things in a visual manner as well, to kind of understand the metadata, so the data about data.
So all of that was available to us as well all throughout the journey, which helped tremendously in making sure that whatever we produced was high quality and we are aware of any issues that come our way. Because like I said, there's no plan that doesn't fail. So we always have a ... a clear understanding of what's wrong, if it is wrong, and then we were able to rectify it pretty quickly because of the way we had built and planned for these data deliverables.
And then, in the end, we were not only able to produce our weekly launch trackers or the performance trackers, and to this date, we still produce those outputs as well. But then on top of that, because of those traffic stars, which eventually became
00:25:00
our final star schemas in our world, we were able to, and we are still able to, utilize that in an ad hoc fashion to answer all the business questions that come our way. And then finally, we work on the resource allocation and prediction models as well, utilizing that data. So, whether that's our forecasting side of the world or our advanced analytics sub-teams here, we were able to embrace that same single data platform, so that there's single source of truth, no matter who's looking at that information.
And then eventually, you unlock that value from the data that you buy. And then so in short, summarizing all of this, I think the key things here was really, changing the fear of, hey, there's a new requirement, or we need to buy new data, or we need to link certain aspects, or we need to make changes in the existing one. We changed some of that fear into the velocity of what we do with the DataOps platform, because on a weekly basis, so our sprint was always weekly because of the ever-changing environment and ever-changing space that we play in.
And we were able to deliver that, and promise our clients that, yeah, we'll be able to get in front of you in a couple of days with a draft version of certain things, with this approach. And the other thing which I touched upon early on is, early on when we were producing, for example, even a weekly deck, it used to take three people almost from Monday to Wednesday to be able to produce a slide deck of 30, 35 slides, which represented the whole story of the brand.
We were able to automate that, and then over time, I've seen that being delivered on Tuesdays and now, you can track by 3:00 PM on Monday, we are able to deliver that. And all because of the automation that we have, push certain things into the data layer. The third thing you hear about is the hope for quality.
We all have worked on environments where there's data quality issues or multiple sources providing different results. So, from that hope of quality mentality, we moved into that integrated view. And because it's a single platform which is connected across from data, from a data standpoint, sorry, we're able to deliver on quality as well. And wherever there's gaps or issues, at least we are able to address it, and mitigate those, with the help of the QCs and other processes that we have built.
The fourth piece here talks about the hero mentality. I think the point that there I'm trying to make is that, we all have analysts and team members, including myself early on, where I wanted to take on everything and deliver. Right? What we're trying to move away from that hero mentality so that you have certain repeatable processes, whether I win a lottery tomorrow, and then leave, we have a process in place where certain things, some of your key deliverables are repeatable, because they are all linked into that whole process flow that we have built.
And the last piece is perfection. I think this is more around the organization's culture as well. And different teams, different people operate differently. So I think the idea for us was instead of saying, "Hey, we'll bring you perfection, we'll bring you 100% of what you have asked in three months or in two months or in a given month," we'll get in front of you in a couple of days or a week, and then we'll have the first draft ready, which might not be 100% ready, but at least we'll have 70% right.
And the reason to do that is a lot of times we know what the stakeholders were asking is not exactly what we think it is. So when you're presenting something in front of the team, then they were able to understand that data little better, and then the requirements change as well. So instead of delivering everything 100% and then changing that all over again, by following this kind of methodology, we were able to, A, reduce some of the rework that we had to done, but also then understand the business question a little better as well, and then deliver on that, in a much faster pace.
And then, so the last piece talks about here about the superpower mindset, right? So in our side, I think I touched on this a couple of times. So with DataKitchen's DataOps platform that we have, we're able to partner with them to then deliver to our customers. And then all of that with a confidently saying, "Hey, we'll deliver you the first draft in a couple of days." And I think that's what our business finds very valuable, and appreciates. And then we, in turn, appreciate our data engineering team here, because they're able to bake in those changes, bake in those
00:30:00
business rules, bake in those new datasets, all within a span of a couple of days. And with that, I think this is the summarizing slide. So we talked about all of this, what's in our span of control and the different components that we have to take care, and what resulted in the success. Of course, I'm not taking all of the credit for the billion-dollar brand here.
It's the product itself and then the product placement. But then what enables, as we all understand, the value of the data, in making sure that the key decisions, the timely decisions ... that needed to be made from the data that we are utilizing here was there, and that's not possible without having a, the underlying data sources or the database, the infrastructure, but also the people that you work with, whether that's from a data engineering standpoint, your own data analytics team, and then to kind of educate the business users as well along the way.
Yeah. With that, I'll pass this on to Chris, and I appreciate everybody's time today to go over this with us.
Okay. Thanks, everybody. Hi, this is Chris. Rajesh, thank you so much for taking the time to talk about your experience. And so I'm just going to try to echo some things and share a little bit of my perspective, because I've also been part of a lot of pharma launches, probably maybe a dozen to two dozen over the past 15 years, and also working with other industries.
And I think a pharma launch is a particular case of how hard it is to do data science and analytics today because your customers, especially in a launch, demand original insight. They've just got a lot of questions to answer, and they want the return on it as fast as an Amazon package. And they just are super intolerant of anything going wrong. And they also just don't want to spend a lot.
And so what is possible to do with DataOps? In this squeeze between doing all these four things at the same time, how can you survive? And I think the lessons here are not only from a commercial pharma organization where a product launch is particularly intense, and not only for any marketing department in any industry, but I think anyone who does data and analytics, in the broadest sense, that what you do is kind of much less important than how you do it.
And so when I started actually in 2005, 2006, after spending a career in the software industry and managing teams, I started to work in analytics for pharma industry. And I had people who were putting analytic data sets together, and there were just a lot of problems. And my first temptation was to blame the person putting the data together and saying that they were the problem.
And, actually, that didn't help very much. And what I learned is that the-- I started to read this guy named Deming, and he said it's almost always the process or the system that people work in that's the problem, and not the person themselves. And he said 94% of the time. And I was like, "Wow, that's a very specific number." But also, because I'm a nerdy guy, I liked it, so I believed it.
And what if you actually said that it's very rare that your people actually screw up, but you have to build a system so that people are successful. And so when you do data and analytics, we don't spend a lot of time thinking about that system. We think about models and algorithms and pipelines and visualization and governance and the data, but we don't think about how do you develop it, how do you deploy it, how do you monitor it, how do you create it, how do you collaborate, how do you measure this factory that you work in? And so the process and the people and the operations, kind of the factory, is just much more important than the thing that the factory has output on.
And that's what your customers buy, the output. And so that's why this is kind of a bit of a contrarian perspective. And so the attributes that we look at when we try to, say, build that factory, are to think about how fast can you get something from an analyst's fingertips, a data scientist's fingertips, a data engineer's fingertips, through a computer, into a system, and into production so your end customer can see it. What's the cycle time that you can deploy?
And then also, can they do that in a way that they're not wrong so that people don't find that it's late or something's wrong in the data or something misconfigured. And then in a lot of organizations, there's different teams that are in different locations. And Rajesh talked about the power of the tools like Alteryx and Tableau to be able to very quickly iterate. But then how do you actually get those findings and have them earn their right into a lower level of the stack, maybe in a database?
And then how do you measure your process? And so what I've found is that a lot of teams, and this is across industries, have actually pretty slow cycle time and they have a lot of errors.
00:35:00
Or every day they have dread of going to work and getting that email when something's wrong or late. And they're sort of the Hatfields and McCoys between the data analysts and the data engineers, and they just have no idea how to measure things. And so we did a survey last year with Eckerson that actually sort of followed that up, and most companies are pretty slow.
It's sort of weeks and months to make any changes to their production systems, or they just have way too many errors per month. And so what the result of that is, is that I think a lot of teams have kind of poor productivity and high costs. And the end customers who are actually using the analytics aren't very happy. And maybe it's a good day for people who do consulting because they just say, "Oh, you internal data people or IT people, you guys are dumb. I'm just going to hire the brilliant guys outside." But also, I just think it means that the customers aren't happy. And so I think if we want to stop having people go outside our teams to be able to do work and hire the cool new consultants, I think we have to start thinking about these ideas in a more fundamental way and about how to affect your factory, how to focus on deployment and errors and coordination and measurement.
And Rajesh has done a great job sort of talking about
what happened at Celgene and why all those sets of hundreds of data sources go together. And he showed a version of this diagram before, but I just sort of wanted to dig in a little bit. And in pharma, there's just a lot of different data, and a lot of the data is sort of about a couple of fundamental things. Physicians and payers and plans and products, and they all need to be linked together, and so they can't actually be linked into one ... incredibly complicated data set.
They sort of got to be semi-linked together, and that's known as a star schema. And you've got these entities, like physicians, you're trying to understand what happens with them. And so the people who take that data are really unhappy when it's late, especially during a product launch, because they depend on decisions that are going to happen on Monday morning.
And if you get something wrong, it's just terrible, and I've had that experience in previous companies of having VPs of sales in a pharma company call me up and scream at me, and that's just not fun. And the only way I think you can do that is to take your current production system and put thousands of automated tests, and I mean thousands of automated tests on top of it, because you never can trust your data providers.
They may forget, they may change something, and you just don't want to get surprises, and you don't want your customers finding out problems before you do. And you also want to be able to make changes, so you want to be able to try something out and put in a new data set, change the schema, alter something, and do those seemingly two in-conflict things, run a really, really high-quality factory, but be able to pick up pieces of it and change it at a really high rate. And so where you want to get to with DataOps is that you want to think of this as a graphic equalizer.
And best-in-class companies like Celgene are able to do that. They're able to have fast cycle time with low errors and high collaboration across all the team members and a well-measured process. And if you do all those things, you can actually lower your cost and reduce the amount of unhappy customers that you have. And it's not a trade-off. It's not like if you push up one of these graphic equalizers, all the other ones go down.
You can move all of them up at the same time, and that's, I think, what best-in-class companies do, is they don't see it as trade-offs. They see that they can improve on all these metrics, and they end up getting these results. And so the last thing is, think of what you do as a factory, right? And if you focus on these factory ideas, these lean ideas and agile ideas, what kind of product is your factory building?
Are you building an AMC Pacer, where the plant wasn't too far from where I grew up in Milwaukee, Wisconsin, or are you building a Toyota Corolla? And I think if you can get my answer, the AMC is no longer in business, and Toyota is one of the biggest, if not the biggest car company in the world. And so, we've been talking a lot about this idea of DataOps, and of course, since we're a software company, we have a software product and some expertise to help you achieve these ends.
And then to do that, some of the ways that we've helped people start off is we do something called a maturity model, where we can benchmark your team and look at where you are on these things about errors, and cycle time, and process metrics, and culture, and collaboration, and customer sat. And so that may be a way for you to understand where you are and what the art of the possible is with DataOps. And that's something that we offer, and we're conducting a survey of the industry again this year to find out where people are at.
And the last thing is, if you're interested in just the concept of DataOps, we've written a book. We've given away probably about 10,000 copies. We've had similar number of people sign our Data
00:40:00
Ops manifesto, and the URLs are there. So I just want to thank Rajesh again for taking the time to talk with us, and I think what we're going to do now is answer some questions that may have come up.
Yes. Thank you both so much. So yes, I encourage everyone to enter their questions into the control panel, and we'll try to get through as many as we can in the next 15 or so minutes. So to kick us off, here is a question for Rajesh. Did you need to change your team structure to implement DataOps?
That's a great question. So I wouldn't necessarily say we had to change the team structure. I do want to answer it by saying we actually didn't have to change a lot in terms of how the current team structure was. But I think, yeah, it's just the power of that underlying platform that we were able to leverage with the help of the team here, enabled us to kind of function the same way.
And we were able to take on, to Chris's point, we were getting bombarded with questions as you would expect with the launch, whether that's from your brand or sales or executive or to the market access side of the world, to external agencies that your marketing teams work with. We were getting bombarded with all these people, and I think we didn't have to change the team structure, per se.
We were able to kind of tackle those questions and kind of prioritize those with the help of some of the underlying data that was available to us, the star schemas. And we were able to answer those different questions from that same sort of star schema that we had built, because that provided the whole team with that flexibility.
So yeah, the answer is, I think, no. We didn't have to change the team structure at all. All right. Great. Well, thank you. Now, here's a question. Maybe, Chris, you can pull this up while we answer the next question. Someone would like to see a high-level view or architecture of the DataKitchen platform.
So I don't know if you can show- Okay ... yeah, and maybe while you're looking for that, we can go on to the next question. I have it right here. Oh, you're ready, okay. Okay, let me just bring to front. And this is more of a functional architecture than a technical architecture. But
what Rajesh talked about, about changing the team structure, and there's one thing to change the team structure, and then there's the way to change the people on how to work and have an agile process and manage things so that you can get feedback and get work to your customer quickly. And so part of that is how you use a tool like Jira or Rally or other agile people management tools to do that. And there's ideas like Scrum and Kanban.
And
DataKitchen doesn't particularly help with that from a software perspective, because there's lots of tools to do that. But there's a technical environment that enables you to be agile and to do agile, and that's something that DataKitchen does do. And so it's about taking all your tools with all the data that's passing through those tools and plugging them into a system that helps you be able to do these sort of green things at the top.
One is that in a production environment, and in this diagram, there's data on one side and customers on the other. Given every sort of tool that you have in the world to do ETL and ELT and some data science and visualization, et cetera, how do you orchestrate and monitor and test all those tools so you're running a good Toyota factory with low errors?
And how do you have a virtual Andon Cord and send alerts? And then the other part is that most people have some environment where you develop something and some environment that's production. Some people have more than one. They have a dev and a test and a QA and a UAT. Some people are partly on-prem, partly in the cloud.
But there's this production development environment and how you move the work that you create from one thing in the other, how you move that SQL that you create or that Tableau workbook from dev to production, I think is a very important part to automate. And then those places that you work themselves, the environment, is something else, and that's, in fact, why we call ourselves a kitchen, and there's a kitchen in our software.
So those three green arrows, orchestrate, monitor, and test, deployment, environment creation or management, are the functions, the technical platform that we do. But it really does help to have sort of an agile process to have the team running.
Okay, thanks, Chris. Well, here's a follow-up question on that. How do you implement the continuous integrations and continuous deployment in DataKitchen? Is it possible to integrate DataKitchen with GitLab CI? Oh, the answer is of course, yes. And we actually did a series of the three pipelines of DataOps, and one of them was we talked all about that idea of
00:45:00
automated deployment is very similar to a term in software engineering called continuous integration and continuous deployment. And actually, you can go to our website and watch. We talked a whole hour about how we could work with existing CI and CD tools, how the differences between software CI and CD and data analytics CI and CD, the importance of testing and environments and best practices.
And so I encourage you to go watch that video or download the slides. Okay, great. And since we're touching on tools here, what kind of software do you use to control data versions? We're a big fan of Git, so we use the open source Git tool to store all our versions in. And companies can use their own Git as well.
Okay, great. Thanks. Now here's a question for Rajesh. What was the biggest challenge you had to overcome to be able to start delivering in days?
Can you repeat that? Sorry, Beth. Yeah, sure. What was the biggest challenge you had to overcome to be able to start delivering in days?
Yeah, I think, the biggest challenge was certainly sort of the mindset, right? When we started off, right, we kind of started with different discrete data sets. So, one person is focusing on one and trying to kind of answer that, and the moment that question sort of overlays multiple data sets. That was one of the challenges that we faced early on, and that's what was solved with the platform that we deployed here, with the DataOps.
But I think the other key component is to be able to deliver in days. It's a mindset change. So to Chris's point, the team structure did not necessarily change, but then we did kind of spend time in training our internal teams, right, whether that's analytics, or our DataOps team to kind of deliver on that agile methodology.
And part of the process is once we implemented, for example, Jira in our case, we were able to, A, track everything. And the purpose wasn't to increase documentation and frustrate the team members further, apart from getting bombarded with all those questions. But the idea was let's get the basic information in, and we were able to learn from that, roughly, for example, in our case, at the time, roughly 60 to 70% of the work came in that week.
So, it's a changing environment. You get your requirement that same week, and you're only spending 40 to 50% of your time that's sort of planned, based on the prior weeks, right? So with those kind of learnings that we had implemented with DataKitchen and again, to Chris's point, we're not talking about agile methodologies and kind of how we do it, but then we leverage those kind of skills, those expertise, right, implementing those things elsewhere to train our team to better get on that mode, get in that kind of mindset. And then once we were able to do it, we were able to then take on these challenges, whether that's questioned from our brand leads. So we had set up one-on-ones with our brand leads, and we were able to kind of show those dashboards, whether that's a Jira dashboard that we had put together in a summarized fashion saying, "Here's the 10 things that you have asked last week and we're working on, but then here's 20 other things that you have asked this week, too.
Can we prioritize some of those things?" And people, once they see that in a simplified fashion, they understand, right? There's a lot, and then we need to prioritize those things. So for us, we were able to, A, reduce that workload, making not just our internal team happy because they have a finite stuff of stuff and we're not burning them down with all those questions, but then also kind of making the stakeholders realize the value we bring to the table as well, right?
And it's kind of trickling down effect. For the end customers for us, for internal end customers, they were able to see that value from us- And then in turn, we were able to see that same value from the DataOps platform that we had implemented. So that's what was that challenge and how we tackled it, is basically the mind shift, the culture, so to speak.
Okay, great. Thanks for that answer, Rajesh. Now, how much did you keep manual testing in your development cycle, or did you automate all testing? Maybe Chris can chime in on that as well. Yeah. Well, I think, for Tesla, I think the data side was all automated. But Rajesh, how did you end up automating the Tableau workbooks and the other things that you built, the Alteryx workbooks?
Sure, yeah.
Now it sounds funny, but then when we took on, or rather, when I joined the team, the earlier days, I remember having a set of two consultants working and looking at the dashboards
00:50:00
each Monday morning. So, when it comes to manual, there can't be any more manual work than that, right? Like somebody sitting at the screen and trying to make sense of things that, do they look okay? Last week it was X, now it's Y. Does it seem reasonable? So those are the kind of things that we had implemented long before, pre all of this picture. But then once we implemented those QCs, and it wasn't a single day we had 1,000 QCs. We built in some based on the discussions. We said, "These are the things that generally go wrong, and what we have learned." And the same came from data engineering side of the world as well, right? The data engineers said, "Hey, yeah, we see these 10 different types of errors that are happening in a given period quite frequently." So we build in those QCs, and then over time, as we learn more, there's always a new challenge, there's always a new learning.
So we kind of then, as and when we found those things, we then bake into it. I think at this point, so a couple of, I should say, quarters into that, I think you reach a point where not a lot of errors on a regular basis. I don't remember in one of our weekly...
Yeah, at this point, I think it's been probably a couple of years, I want to say, we have a major glitch in our weekly release. And it's all because of this- Rajesh, could you repeat that? Because that just made... Rajesh, could you repeat that again? That just makes me smile.
Yeah, I'm sure Larry and Steven will be very happy, so would you be. But yeah, early on, I remember our weekly release, it used to be delayed because the SP is not providing us the file, or there's file format changes suddenly happening. Pharma is a challenging world when it comes to data because of the HIPAA compliance, and the kind of data we get, and the sources, and not being able to get the 100% complete picture, right?
Which, if I talk about any other industry like the car building or the telecom industry I worked in, we had 100% of the data. So A, there's always a challenge of how do I utilize 70, 80, 90% of the data and make decisions out of it? And then B, on top of that, if there's frustration coming from just having errors on a regular basis, it just kind of breaks the confidence in the end stakeholder. So what I was touching on was, early on, I think, when we started, we used to see at least once a quarter, we used to have something where things break or we found something. But then because we kept on adding these QCs as we found, we never stopped that.
I think it's been a couple of years, to be honest, where we've had any major glitch. Now, of course, if the source file isn't there, it's not going to be there. You can't do anything. That's kind of out of control. But then how you build your process so that that component can even still be last week's data, for example, or the prior update, and still be able to deliver rest of it, is really the key component here. So, that's what we were able to kind of leverage in terms of controlling these errors and then still being able to perform our normal jobs come Monday morning at 9:00 a.m.
Yeah, and I think that does set the bar, right? Having no major data errors for several years. And I talk to companies in all industries, and just a few months ago, I talked to a major insurer, and they were incredibly happy that they're only having major errors once a month. And so when I find that out, my heart sinks for people who...
Because major data errors are such a soul-destroying thing for a data and analytic team because you rush around, you want to fix it, you blame each other, and it actually takes away the efficiency of your team. Because not only you end up, like with Rajesh, having to spend Monday morning checking things, but it gives you this fear at the pit of your stomach that you're going to get something wrong, and that actually ruins your ability to innovate. And so, I think that fear sort of kills the team from being able to actually create and innovate and add value.
And so error reductions have, it may sound like a really boring thing, but it has all these secondary effects of innovation and trust and helping people to be data-driven. And that's why the focus on the factory and having the kind of success that Rajesh and team have had, of having just making errors sort of not a point of discussion, and being very quiet about it.
It's just such a good feeling, I think. And it's really, I think, a goal that every data and analytic team should have.
Yeah, that's great to hear, Rajesh. Here's another question related to errors, which maybe you've already answered, but, "How easy or difficult was it to automate and reduce error rates in reports and dashboards such as Tableau, given that there could be both data issues and visualization issues?" So, for example, the drill-down on territory does not work the way it used to.
Sure, yeah. Good thing for sure, especially with Tableau, which brings with a new version every other month or so.
00:55:00
I think that has gone down a little, but yeah, we've worked with Tableau where there's a new release, and then with that release, there could be some bugs that might cause these errors. So- From a data visualization standpoint, I think that comes with the development cycle, right? So those errors, whether that's software that you do decide to pick and use, those errors would still remain, and I think Tableau has reached a point where, with the recent releases, the bugs are still there potentially, but the extent of those bugs in terms of impact is not the same as it used to be. They've been maturing their product as well, right? So I don't remember seeing where something broke because of an update or things of that nature.
But then the simple idea is, again, I think it's touching the old point, which is once you kind of figure out what the error is for the very first time, right? The idea is not to make mistakes. We all do make mistakes. Our teams make mistakes. I myself make a lot of mistakes as I do things, but the idea is how do we learn it, right? And then bake it into the data layer, right?
Once we learn something, we push it into the automated world so that, I might forget doing something on Monday morning, but the machine, the platform, it wouldn't. So the idea is to then kind of understand those things, and then push it into the data layer. And the same goes with Tableau or Alteryx as well, right?
The moment we kind of understood certain things were getting complex, or if we're talking about a sub-national and a complicated measure measuring sub-national and territory level information, we kind of then decided on, and the guiding principle is if it is used by multiple people on a regular basis, we can technically push into the data layer. So that's what we did.
Over time, and it's not like, again, a one-time process. Today, you're looking at your approval or rejection rate. Tomorrow, you might want to add something to it. So the idea is to then, anyway, as you learn more things, you then push it into the underlying tables, the data, and it not just helps with the error reduction, and then sort of taking stuff out of your plate, but it also helps kind of, making it faster too, so that you're not worried about Tableau calculating tons of these things and on millions of rows, but it's all baked into, so you can easily kind of lift and shift and present it to the business.
Okay, awesome. Thanks, Rajesh. One question, another question for you is, are you using NiFi for orchestration?
Oh, hey, this is Chris. No, we actually don't use NiFi. We integrate with NiFi, because other people use that as a kind of a tool to do data work. Here, the data work is actually done following an ELT factor, and it's called extract, load, and transform. And all the work is not done in NiFi, but it's done by writing SQL statements on the data side.
Okay, thanks, Chris. Well, we are running up against the hour. We have one more question, and I think it's actually a great question to conclude with, is, Rajesh, how did you get buy-in at the executive level for DataOps?
Sure, yeah. I knew this would be coming. Anyway, it's a great question. I think that's certainly the biggest challenge, right? As you kind of talk about any new change, whether that's addition of anything like that or building something from scratch or even then making small tweaks, right? It's always, "Hey, why are we doing it?" And then how much time we're talking about, how much money we're talking about.
So in terms of the buy-in, right, in our case, what has happened, and it's different across my whole career, right? If I think about it, one thing is you do proofs of concepts, and then you show value, and then you kind of implement it. I would be kind of lying if that's what happened here.
I think in our case, it was an easy buy-in because the way currently system was built in, we had siloed, datasets, and our IT teams were, although they were working hard on it, but they were kind of focusing in on their areas, right? And they were delivering on what was asked. And in terms of the changes that were being asked or the business questions that came our way when we were trying to answer, we were limited in our capability, right, from an analytics standpoint.
So when we were pushing these changes or asking for those improvements, we were told to wait for a month because the traditional sort of waterfall method, right, in terms of delivering something or developing something. So, you're asking your executives to then wait for a month, and then they wait for that. And let's say, whether that's the requirement has changed or just what was delivered isn't exactly what they're looking for, right? It makes it easy for you to kind of implement something like this because then you can technically show the value, right?
01:00:00
Because, again, the idea isn't to present 100% that's available, right? The buy-in came from the very fact that we were able to present at least that 70% picture, right, in front of the executive in a short span of time. And then over time, that became easier because people realized that week over week, we're getting the data.
The reps are getting their volume information. The home office is getting their business questions answered in a short span of time. The velocity increased over time. So hence came in the buy-in year over year and on long term for us. All right. Well, thanks so much, Rajesh. We are up against time. I want to thank everyone for taking the time to join us today.
An extra big thanks to you, Rajesh, for speaking and sharing your story. I know I loved hearing about how Celgene used DataOps, and I hope all of the attendees found it useful as well. As I mentioned at the outset, we'll be sending out the recording and the slides in the next 24 hours, so be on the lookout for that in your email.
And if you have any additional questions, please don't hesitate to reach out to us at DataKitchen. So have a great afternoon, everyone.
Thank you.
Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.
Questions from this session
Why is the first year of a pharmaceutical product launch so important?
How fast a product grows during its launch determines its overall lifetime revenue, which puts the weight on the first six to twelve months. The stakes are set by what came before: bringing a drug to market costs $2.6 billion, according to the Tufts Center for the Study of Drug Development. The launch happens against increasing marketplace complexity, fierce competition, and continued pressure on cost.
What is a scrappy star, and how did the Celgene team use it?
A scrappy star is a first-pass star schema built quickly in a data mart when a new business question requires a large new data set. Data engineering built it, sent questions back to the data supplier while keeping analysts in the loop, added data tests to make speed safe, shared it with the analyst team for feedback, and iterated across several Agile sprints before releasing a solid star. Analysts worked the same way from the other end, building scrappy dashboards and showing them to users early.
Why publish a production dashboard at 70 percent done?
Because feedback from real users on a partial dashboard is worth more than a longer wait for a complete one, and because the remaining 30 percent is usually the part the team guessed wrong. The team ran active build and design sessions, making as many changes live as possible and updating through Tableau Online. Seventy percent right the first time is one of the five substitutions in the DataOps mindset, replacing perfection.
What results did Celgene get from DataOps?
The platform integrated hundreds of data sets, mastered them into a unified star schema, and ran more than 20,000 automated tests. It absorbed over 100 schema and data changes per week with very few errors or missed SLAs, at low total yearly cost across hardware, hosting, software, and staffing. The team supported ongoing production deliverables such as a weekly launch and performance tracker, ad hoc answers for business leaders, and resource allocation and predictive models.
What should an analytics leader keep under their own span of control?
Five areas: agile team management, a DataOps technology platform, data engineers with their tools and database, data analysts with their tools, and data operations. The underlying claim is that a leader accountable for delivering value to marketing, sales, and customers needs the data, the analytic database, the people, the tools, and the delivery process inside that span rather than dependent on another function.
Why does DataOps focus on process rather than on the analytics themselves?
Because what you do matters much less than how you do it. The what is the model, the algorithm, the pipeline, the visualization, the governance, and the data; the how is development, deployment, monitoring, iterating, collaborating, and measuring. Two outside arguments back it: Elon Musk on the greatest potential being in building the machine that makes the machine, and Deming's finding that 94 percent of causes are common cause, which means looking for a person to blame instead of fixing the process misses almost every time.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.