On-Demand Webinar · 55 min
DataOps: The New Normal in Pharma
Chris Bergh walks through how four pharmaceutical companies apply DataOps across R&D, commercial, and manufacturing data: fewer data errors, new analytics delivered faster, and collaboration across teams working in different tools. Recorded June 2021; updated August 2026.
What you'll learn 6 points
- Worldwide pharmaceutical sales run $1.2 trillion, with North America accounting for 46 percent of the global pharmaceuticals market in 2021. Pharma companies increasingly compete on analytic capability, but internal teams struggle to deliver on-demand insight because the data is complex and sensitive and the teams sit in silos using different tools in different locations.
- Drug innovation is expensive and uncertain: only a fraction of eligible products win FDA approval, and it costs on average over $1 billion to bring a drug to market. The first few months of a commercial launch decide long-term revenue for a brand.
- At Celgene the DataKitchen platform integrated hundreds of data sets into a unified star schema with more than 10,000 automated tests, absorbing over 100 schema and data changes per week with very few errors or missed SLAs, run by a staff of seven data and DataOps engineers.
- A commercial launch analytics team owns roughly 24 recurring reports and activities, from forecast tracking and launch tracking to stocking, field activity, incentive compensation, payer and plan performance, REMS, and a demand-based P&L, delivered hourly, daily, weekly, and monthly.
- A commercial pharma data mesh splits the business into three domains that each dominate a different lifecycle stage: NPP, meaning non-personal promotion such as email, web, and radio, matters pre-launch and during growth; the Physician domain matters during the first years after launch; the Payer domain, which controls price through rebates, formulary, and tier, matters most in the mature phase.
- Domain layers stack from raw sourced data through mastered data sets owned by IT, integrated data sets owned by data engineers, and self-service tools owned by analysts. Mastering is its own layer: there are one million physicians in the US, but a company's physician master holds only 40,000.
Slides
Questions from this session
Why is pharma data harder to work with than data in other industries?
Pharma is organizationally complex, with R&D, Commercial, Manufacturing, and Finance functions that each hold very different data and rarely trade staff between them. R&D holds genomic, laboratory, chemistry, clinical trial, and image data; Commercial holds syndicated sales, anonymized patient, claims, non-personal promotion, and sales force data; Manufacturing holds operations, IoT, process control, and regulatory compliance data. Different teams use different tools in different locations, on-premise and across multiple clouds, so there is little consistency to share or reuse.
What is a DataOps process hub in commercial pharma?
A process hub sits on top of an existing IT data lake and holds the automated, governed, reusable processes that turn raw feeds into analyst-ready data sets. It gives the commercial analytics team production runs, rapid-development sandboxes, reusable libraries of recipes and ingredients, and automated tests, all controlled by the analytics team rather than by IT. The goal is to lower the cost per business question so ad hoc requests get answered faster without hiring more consultants.
What are the domains in a commercial pharma data mesh?
US commercial pharma splits into three domains: NPP, or non-personal promotion, covering emails, website visits, and advertising; Physician, covering doctor sales, claims data, and anonymized patient data; and Payer, covering plans, rebates, and formulary. Each domain has separate sources and different cycle times, and entities such as physicians overlap across all three. Data does not reconcile cleanly between them: subnational physician data purchased from IQVIA may not match claims data one to one, which may not match payer data, because of supplier differences and timing projection algorithms.
How do data mesh domains communicate with each other?
Domains exchange five kinds of link. A domain query asks when a domain last updated and whether the run succeeded, or asks a domain to prove its data is good with test results. A process linkage hands off control and parameters between domains, an event linkage broadcasts completions, errors, and warnings, and a data linkage covers a shared table such as a common dimension. A development linkage is the ability to re-create another domain in development, read and modify its code, and push a change to production.
What results did Celgene get from DataOps?
James Royster, head of commercial analytics at Celgene, reported dramatic cost savings that let the team measure ROI directly, plus faster response that helped the business capitalize on opportunities. He also described an almost complete absence of data errors that were not caught earlier in the process. A director of market insights described the outcome as a self-service data organization reaching from the marketing department to the sales reps.
How does a large company start a DataOps transformation?
The six-step sequence is Educate, Find, Establish, Demonstrate, Iterate, Expand. Education uses presentations, video, books, and analyst writing; finding a first project comes from in-depth discussions with individual teams guided toward their pain, supported by a DataOps maturity model; establishing a community of interest means wikis, Slack, and alignment with data and agile leaders. Demonstration runs a one-to-two-month pilot with real measurement, and expansion ends in a staffed center of excellence or dojo that sets common infrastructure, tools, and metrics.
Where to go next
- Install open-source TestGen Apache 2.0, runs in your own database. Docker Compose to a first quality score in about 15 minutes.
- Every on-demand webinar The full recording library.