On-Demand Webinar · 1 hr 3 min

Solve the Analytics Last-Mile Problem with a DataOps Process Hub

World-class engineers and a best-in-class toolchain, and it is still hard to answer an analytic question. Chip Bloche and Chris Bergh argue the bottleneck is process, and walk through a DataOps Process Hub that lets analysts answer stakeholder questions without queueing behind central IT. Recorded November 2021; updated August 2026.

Presented by Chip Bloche, Chris Bergh

What you'll learn 6 points
  • Business analytics teams describe the gap between the warehouse and the answer as a last-mile problem. In their own words: 80 percent of the effort goes into getting data together and 20 percent into insight, and delivering insight to the business is a 90 percent business analytics problem and a 10 percent IT problem.
  • The process hub inverts the table count: instead of a thousand tables from the data lake, each business analytics team gets about a dozen tailored tables it controls, with current data dictionaries and the ability to change data and schema daily or weekly instead of waiting on a multi-month IT cycle.
  • A data hub is a repository for data and a one-way street from raw to processed layers. A process hub holds the processes that act on the data, and it exists because valid data is the product of an effective process. There is no single version of truth without a single version of process.
  • A process hub raises productivity three ways. The Wedge is trusted, analyst-ready data under business analytics control, which lowers the cost per question. The Hammer is full automation from source to delivery, which removes recurring work. The Store is a version-controlled repository holding all business analytics intellectual property for reuse and sharing.
  • DataOps engineering automates eight specific things: production orchestration, production data monitoring and testing, self-service environments, development regression and functional tests, test data automation, deployment automation, shared components, and process measurement.
  • The measured shift is deployment latency from weeks and months down to hours and minutes, with production errors falling from high to low.

Prefer to read it? The written version is in Centralize Your Data Processes With a DataOps Process Hub.

Slides

45 slides

Transcript

Show chapters and dialogue 8,925 words

00:00:00

Welcome, everyone. Thanks for joining our webinar today. My name is Beth Pfefferle, and I am the VP of marketing at DataKitchen, and I will be the host for today's session. So our topic is how to solve the analytics last mile problem with a DataOps process hub. We'll hear from DataKitchen stars, Chip Bloche and Chris Bergh.

Chip has been a DataOps engineering director at DataKitchen since 2018, and he has more than 30 years of experience in data analytics and data engineering. Chris is the founder, head chef, and CEO at DataKitchen. He's the leader of the DataOps movement and is the co-author of "The DataOps Cookbook" and "The DataOps Manifesto." So before they jump right into it, just a few housekeeping items. So we are recording the session, and we will send out a video and the slides to everyone who's registered as soon as those are ready, likely within the next 24 to 48 hours.

Also, the last 15 minutes of the webinar will be for questions. So if you have any questions during the course of the webinar, just enter them into the Q&A box on the webinar control panel and we'll make sure we get through as many of those as we can at the end. So, without further ado, I'm going to hand this over to Chris, who will kick off the webinar.

Thanks for that great introduction, Beth, and welcome everyone. Thanks for calling in. I'm actually really excited about our discussion today because it's targeting at a specific group of people in the data and analytic value chain. And we've been working with them for years, and I'm really excited to talk about it, also very excited to talk with Chip.

And because I'm in a hotel room in New Jersey, don't ask why, I'm going to turn off my webcam and just talk to you because I would want my bandwidth to be there. So, what are we going to talk about today? So let me go to the first screen. So we're going to talk about this term of a business analytics team.

Now, what do I mean by a business analytics team? I'm going to go into that. And how does a process hub help them? I'm going to go into that. And then we're going to actually give some examples of how business analytic teams have used our process hub to make themselves successful, much more productive, and have a much more saner life. I'll give a quick demo if we have time, and then we'd love, as always, for you to answer your questions. So let's start at the beginning.

So who are this term business analytics? And so a lot of data and analytic teams are one team, right? There's a chief data officer or a chief analytic officer, and sometimes they're a pair of the CIOs, sometimes they're separate. And on the right-hand side below them, think of these as a director or VP of the data and the analytics, quote, IT.

They're the people who kind of access the data, build the databases, integrate data. Sometimes they do standardized reporting, sometimes there's a data science team in there. And then there's another team that's sometimes they're called business analytics, and they actually tend to work more closely with customers in the specific lines of business. Sometimes they tend to visualize data more, but they also report on it. Sometimes they do data science, sometimes they integrate data.

So, the actual type of work in these groups is the same, they just have a different emphasis, and oftentimes they're actually separate from each other, which is interesting, that there is no shared CDO or CIO. You have lines of business, a business leader on the left, and you have people working for him or her, right? And they are aligned to that team, and they have to partner with the data team, the data and analytics IT team.

And so in some cases, it's a hub and spoke model, a data enablement model, some people call it a self-service model, but the data and analytic teams is more often integrating data, putting data together, and sometimes doing reporting or modeling. And similarly, the tasks are there with business analytics, so they tend to really be focused on trying to make people in the business successful.

And so that's the challenge here, is these two teams, when they're separate, have to work together. You've got self-service business unit teams, and you've got sort of centralized data and analytic IT teams. And I think if you've been in a big company, you've seen this pattern, and you've seen some challenges with that pattern.

And so we're going to actually focus today on this group, the sort of business analytic teams, and working directly for lines of business, and their unique challenges, because I think there are a set of unique challenges. And let me go into those first. For some business analytic teams, they do one-offs and they're done.

00:05:00

But a lot of business analytic teams have ongoing deliverables where they have a dashboard, they have a PowerPoint, a model that they've got to kind of keep up and running. And they also have to make rapid changes in their own production deliverable. Sometimes business analytic teams have to do a lot of data work.

The data's not ready for them, and they may get a lot of data tables from IT and diverse data sets, and they've got to do a lot of sort of fragile and error-prone and high-effort data work. Sometimes they use tools like Alteryx, sometimes they just use SQL, sometimes they'll pivot around in Excel, but they're doing data work.

And that sort of repetitive kind of production process, that sort of repetitive data work kind of reduces their goal, which is to create original insight for their customer. And that's a real key here, is that the business analytic team is kind of focused on their customer and making them successful because they sit next to their customer. They see their customer roll their eyes when they don't do it.

They field all the follow-up questions. And I think one of the things that they also see, they're the first line when things go wrong. The business analytic team sort of sees when there are data errors or production errors, and as a result, the teams are very stressed, and sort of firefighting and heroism are tough. And a lot of times, those business analytic teams don't have kind of a tailor-made data representation just for them.

They're having to deal with a bunch of raw tables, maybe flat files, and it's not ready for their insight. They've got

perhaps hundreds of tables that they've got to pull together to try and get the sort of 10 tables that matter. And so there's this amount of manual work that goes in, and the sort of tools to help automate, like an Alteryx, they help a little, but they're still doing a lot of manual time. And so they face this sort of awful choice.

They could either kind of throw more staff at it to do sort of putting data sets together for analysis, doing reporting, or they can sometimes work with their IT team, who takes a lot of time to do it. And in this, I'm an IT guy myself, right? And so I think there's sort of two classes of teams who support business analytics, right? And there's one which is, some of them are great.

They're focused on making the business analytic team successful. And that success leads them to work in more of an agile and iterative way. And they judge their own success by the success of the business analytic team. It's their customer. And so, if you're on an IT team, I hope your team's like that. Now, there's other types of data and analytic IT teams who support business analytics who honestly are not so great in my experience.

And I'm not going to say what the ratio is to the other, but they are out there. And what they do is kind of prepare data in some form or artifact, so they throw it over the fence to the business analytic team. And they're kind of not focused on delivering customer value, but more focused on kind of servicing their own internal process, their own internal software development life cycle, following their own standards.

And if they've done that, they've kind of seen success, and they're sort of customer value avoidant and process-centric. And as a result, they take many months to get something done. And when it does get done, it's usually not right and needs a lot of rework. And sometimes the business analyst team end up having to QA IT's data because they don't trust it.

But the business analytic team are partnered with IT. They need them. They need the servers, the data. And so, when the data and analytic IT teams are not so great, the business analytic teams have some challenges. But not to say that all data and analytic IT teams are like that. Definitely don't want to make that presumption.

And so what we're talking about here is a way to use our software, our process of software for business analytic teams and how that works. And so in our years of working with business analytic teams, I've seen a lot of these kind of comments. One is on the upper left, they have a productivity challenge.

We want to do more for our business customers, people in marketing and sales or service or finance or production or manufacturing. Man, they have a lot of business questions they can answer. One of the things I've seen a challenge on business analytic teams is they work very hard. They actually do give original insight to a customer, and then they walk out of the meeting, feeling like a failure, and not because they haven't delivered their job. It's because they got 10 follow-up questions, and they know those 10 follow-up questions are going to take forever to get done.

And so imagine working in a field where even if you do your job well, and having 10

00:10:00

follow-up questions in my mind is a success, not a failure, you feel like a failure. And in a lot of cases, the business analytics team feels like the responsibility of success and insight on the organization relies 90% on them. And in some ways, some people in business analytics have not a good opinion of IT, and they say, "I get cut off at my knees from a data perspective. I'm getting a kind of a special sandwich given to me by IT, and it's not a good one." And as a result, their teams are very reactive with their customers.

The business analytics team's trying to field business questions from marketing and sales people, and they can't get ahead of them. And as a result, they need sort of data in answer-ready format. And the way that they're getting it doesn't solve it, and so they're spending sort of 80% of their time doing stuff they don't want to, and only 20% of the time doing the insight generation, the original cool stuff that gets them up in the morning. And from our perspective, we think our software is the solution to that, right? And it's a way to use our software, and so by the business analytic team. And so I think that IT teams can give a lot of data in a lake or in a warehouse, and a lot of patterns of IT teams now are saying we're data enablement or we're a data hub. And what they do is they get a lot of raw data there.

But in order to make business analytic teams happen, they need to take thousands of files or tables and synthesize them into the 10 tables that matter. Because in a lot of ways, if you can get that simplified representation of the data, that's abstracted away, that has an up-to-date data dictionary, those teams can be really productive. And in the place of sort of multi-month development cycles, those teams need sort of very high iteration rates.

They're building a data representation on the raw data, but they need to simultaneously, what we've talked about with DataOps, is be able to run things in production with low errors, and then pick it up and make a change in a day or an hour, including schema and additional datas. And so one of the ways that we've talked about this is to think about some roles, and I'm going to go into the roles of an analytics engineer or an analytics data engineer and a DataOps engineer and how they work together to focus on making the business analytic team successful.

And it really comes down to kind of, pardon my French, automating the hell out of the process so that the analytics engineers don't spend any time maintaining stuff that they've already got working, that they've built it in a way that can be improved and run with sort of zero effort. And lastly, I think business analytics teams typically have a lot of people.

Typically, people sort of will rotate in, rotate out, sometimes they'll be long. And how do you actually store all the best practices? And sometimes organizations have consulting teams that they bring in. How do you get those consultants to put their work into a process hub? So you've got all the intellectual property that your business analytics team is doing in one place, and models, SQL reports, tap scripts, so you can share it, reuse it, you can curate it and improve it.

And so we think if you've got something like this, what it means is the business analytic team can get faster insight because they've got a data representation that they both control, that can change, and can work for them generating insight. And maybe the insight's charts and graphs or maybe it's a feature set for a data science model.

And what we're going to talk about here is focus on automation and reusing and abstracting processes. And what that means is automation saves time and shifts the equation from doing one-off work to doing repeated work. And it's really about how to extend the-- It's not a replacement for what's been done in a data enablement team or a hub data team. It's on top of it.

It's really enablement for business analytics and not IT. And so can you replace an army of people kind of doing one-off things over and over again with automation, with abstraction, with testing, and then therefore kind of change the operating model of how business analytic teams work. And so let me kind of go further into that.

Some people best like to think in people. So let's just talk about what this is. And so imagine Priya's on the right and she's in a business analytic team, and maybe she has a title data supplier test, or maybe she has the title of business analyst or business intelligence analyst or

00:15:00

customer data analyzer. And in some ways, maybe she's like a viz victor. She's trying to express data in a way that makes sense to her non-technical, non-analytic business customers. And that could be a chart and graph, or a data science model in form chart and graph or a list. And what gets Priya up in the morning is looking at data, understanding data, creating insight, and expressing that insight in the form of charts, stories, graphs, models to her business customers.

And so what we're suggesting is that there's two people who are there to make Priya and all the people like Priya successful. One is an analytics engineer, and think of that as a data doer. They're taking the data that's in an IT hub and building kind of data representations that are made for Priya's success.

And it's really about making a data representation, but also knowing that Priya is always going to have follow-up questions and always going to have to have changed things, so a representation that's flexible and can be changed quickly. And then there's a guy named Chris who's kind of think of him as the operations optimizer. In order to make Ahmed and Priya successful, he's providing kind of the foundation for all that. Think of the technical environment and the setting up the environments and testing. And so

one of the successes in Agile is knowing who your customers is. Chris' customer is Ahmed and Priya. Ahmed is really focused on making Priya successful. And so if we talk about Chris as a DataOps engineer, in some ways it's about collaboration through a shared abstraction because individuals may have a model, some data transformation from an analytics engineer, may have a visualization, may have some new data set.

How do you take those nuggets and put them into pipelines and create automated tasks and sort of run the assembly factory? And then how do you actually work on the process of moving things from production into development? And then how do you measure, and how do you enable self-service? And so DataOps engineering is really about taking nuggets from these people, the nuggets that everyone creates, and putting them in a system, a shared system, and a shared abstraction that allows that sort of scalable, fast, iterative work to develop. And it's just really kind of about automation.

It's about automating production orchestration, data monitoring and testing and production, development and regression and functional testing, test data, deployment automation, shared components, process measurement. It's really about automating things and making sure that everyone can kind of see the process that they're working in and helping people remove their blinders. So if I'm an analytics engineer and Priya's building a visualization and someone else is doing a data science model and IT is providing the data, can I have a place where all that is seen together and I can actually see the problems and work on it and automate it and improve it?

And so DataOps engineering is fundamentally about automation. And I think analytics engineering is really about putting these useful data sets ready for analysts to make it successful. But it's also about not just running to the next fire, because a lot of business analytic teams are constantly firefighting. But it's also about making sure that when you do something, you can then not do exactly the same thing three times over again with three different versions of the doing it.

Can you move assumptions on your business rule or code upstream, and can you generalize it and abstract it instead of taking one file or four files that are doing about the same thing, can you have one file that's parameterized? And can you curate a framework to do that and build out a production toolkit where you're having a set of tools that you can use from one project to the other.

And also just collaborate with everyone in the process, IT,

your data engineers, your data scientists, the people who are doing the visualization. And so partly it's thinking of, okay, as an analytics data engineer, I've got to get the data for my business analyst or my data scientist ready. That's fantastic. But also you've got a job in thinking in systems and thinking in reuse to scale that process. And it's really about making the next urgent deadline easier to meet. And these teams often have very urgent deadlines.

But you've got to make room to improve it so you're not always fighting fires and

00:20:00

running to the next fire. So analytics engineers have that dual role. And let's talk about one more person. Let's talk about the manager of a business analytics team. And so she faces some challenges, right? How's my team doing? Since I'm doing things in a, quote, "production-oriented way," are they working? Are my business customers happy with what I'm doing?

Is my team actually being productive? And how do I have some projects, because a lot of ways a person like Stephanie is managing multiple projects and multiple people. And lots of sort of projects, i.e. answering business questions, they start and stop, and people move from one task to the other and maybe you've got people on your staff, maybe you've got consultants, maybe you've got vendors, they move in and people have different specialties.

And since you're working on these different projects that start and stop, how do you gain consistency? How do you get collaboration? How do you get a longer term view? And how do you curate, maybe think of it as a series of shortcuts and reuse a couple components that can span projects and enable and rotate to enable efficiency?

Because in a lot of ways, once you've done the same data patterns over and over again, you can start to simplify and abstract and generalize. And the idea here is, a lot of industries talk a lot about data hubs, bringing all the data together. But can you have a process hub where you bring the processes that are acting on data, not as a replacement to a data hub, but as an augmentation to it, and bring all the people that act upon data in one place.

And in our mind, managing and curating the processes that are acting on data is just as important. It's the intellectual property. It's a thing that enables scale and productivity and project success. And so that's why we think from a management standpoint, this idea of focusing on automation and process curation and a process hub is as important as the data itself because it's what enable business analytic teams to get leverage to be successful and not just constantly firefighting and constantly having people switch, because I've managed

people who are doing business analytics, and if they're constantly firefighting, sometimes you get heroes and they're really good at firefighting, but they keep creating complicated things over and over again, and eventually they get sick of the firefighting and they leave. And then you're left with all these complicated things that they've created that you have no idea how they work.

And then a poor person's got to pick it up and understand what someone else did without working. And so I think curating these processes in a process hub, breaking them into components, making them testable, making them deployable and automatable is a way to ease the burden and transition from your hero people to everyone else.

And so let's just go through the key value proposition here, and then I'm going to hand it over to the question. So you're in charge of a business analytic team, like Stephanie. And your boss says, "Great, we got some more products coming out. You got to do more. And you know what? Your budget is going to be exactly fixed." And your business customers who are getting the insight are constantly saying you're behind and you got to do more.

So how do you actually do that? If you think of from a business analytic teams, the amount of team and effort that they do, they produce a certain amount of insight every year. And how do you actually increase that? How do you actually increase the amount of business analytic insight generated without continually adding more costs, right?

Which is really a productivity question. How do you drive productivity gains in your business analytic team so that they can do more with the same amount? And for us, it comes to a different perspective, is that by working on a process hub and focusing on automation and testing, that you can then increase the productivity of the people who are doing insight.

And the result is you get more customer insight. So it's a change in the focus. It's saying, "Stop running around doing one-off things. Start generalizing. Start putting all your processes in one place, improving them, curating them, managing them, automating them." And lo and behold, your team's going to become much more productive. And we've seen, from our customers, we've seen incredible success using this on the order of 10 times more productivity from the team.

And you may laugh and say, "Oh, that's just marketing BS," but it really comes from the fact that they're doing what matters now. They're focused on exactly what the customer wants. Their cycles are quicker.

00:25:00

They're able to better tune what they need with the customer. So, it's not only that they're more efficient, but also that they're more efficient at finding out what the real challenge is. And because they've got this process hub that enables them to work in a more agile way. And so that's it for my section now, and I'm going to hand it over to Chip. And Chip's going to actually give some real-life examples of how customers are using it, and I really appreciate Chip's perspective.

So Chip, can you take it away? And just tell me to go to the next slide. Thank you, Chris. You can go to the next slide now. My name's Chip Lock and I work for Chris, but I'm in managed services for DataKitchen. I'm not a product development person or a sales or marketing person.

I'm actually tasked with using the DataKitchen products out in the real world. And sometimes I have to come back to the team with bad news about things that aren't working so well. I'm really excited about the process hub concept because it really is a distillation of some principles and approaches that have been very successful for us and for the teams that we've worked with.

I think there was one right behind that, the process hubs in the wild.

That's it. So

what's really different, I'll give you my perspective of what's really different. I'm going to save some bandwidth too here. What's really different about process hubs in the real world. There's a lot of talk about data hubs, which are a repository or a collection point for data that moves through the enterprise. And often, data layers in enterprises are kind of a one-way street.

You go from raw data to different levels of more processed data. A process hub is a little more complicated. First of all, it's a hub for the processes that act on data So it works with a data hub, but it is a hub for the curation and management of multiple processes, and it really enables a lot of new capabilities.

The idea is that data flows back and forth between multiple teams performing multiple functions, and in the real world, it's way more complicated than those levels. And what happens is very often is a lot of that data flow is because it's not captured in an enterprise system, it's not effectively managed in the enterprise system. It ends up being very ad hoc.

And the data keeps changing, the processes really have to keep changing as well, and it becomes much more chaotic and difficult to manage. And the way I look at it is you can't have a single version of the truth unless you are maintaining, managing, curating a single version of a process.

So you can do the next slide.

Next slide, Chris. Thank you. So, this is an idea of Chris's here that you can really think of three ways that process hubs increase productivity. One is to kind of right the balance between centralized IT and the needs of business analytics teams. The wedge is a way to

maintain control over analyst-ready data sets, and allow for changes to occur. The second principle is the hammer, which is really automation and was what Chris was talking about, automating the hell out of tasks. It really targets that redundant effort and frees up business analytics people to do the work that they do best. And the third is the store, and that's the idea of the process hub.

00:30:00

It's a curated repository for enabling reuse and best practices, enabling better QC, and storing all the information that is critical intellectual property for business analytics.

Next slide. And I kind of think of these as a hierarchy of benefits.

Data curation is important. It provides consistency and convenience to users who are trying to access a single version of each data set. Process automation allows you to have repeatability, and that's valuable. Anyone who's ever taken an analysis through a Jupyter Notebook instead of hand SQL knows that it's really nice to be able to repeat an entire scripted process. But that process also has to be managed.

And what process curation allows you to do is provide transparency and control to other folks on your team and to other teams. It allows you to keep track of traceability of the processes and the data,

data lineage as well, and also it enhances collaboration and reuse. And Chris also talked about abstraction. That's really important because by abstracting your processes into an encapsulated, consistent form within the DataKitchen platform, it allows us to build a kind of plugin tool set that tracks processes and whether they occur, tracks data, whether it moves into the right locations to feed the processes, performs standard kinds of QC testing on those processes, and reports the results in a standard, consistent way.

And that really enables a whole different dimension of automation that makes life much easier and makes teams like ours much more productive. Next slide.

So why can't you leave a process hub to IT?

One of the problems here is that creating and maintaining a process hub in the real world requires business expertise. It's not just a function of infrastructure. It requires analytics expertise and a different level of understanding of the data. And it's also subject to rapid change. The conditions are changing. As Chris was saying, there are always new questions.

You really have to develop a meta-understanding of a process, and the key kinds of questions you want to answer are, what aspects of that process are generalizable to what levels? Next slide.

So alternatively, why can't you leave a process hub to analysts? Well, what analysts are doing every day is very different from what's required because building a process hub requires enterprise priorities,

cross-functional cooperation within the organization, and this can be multiple teams together as well as multiple vendors in multiple time zones. And in a large, vibrant organization, there's a lot more involved than just the next deadline. Process orientation is not necessarily the same thing as a deliverable orientation. And the key here is that we're trying to get beyond the need for people to take heroic steps to get their data out on a regular basis. Those people who are really doing

superhuman and very complicated efforts to do the basic stuff end up solving a problem in the short term, but creating technical debt in the longer term. Next slide.

00:35:00

So, the result is a kind of a vacuum.

On the analyst side, you have people who are often brand-focused or focused on their own particular business priorities, and that's good. It's important, but they can also be siloed. You have people who are deliverable-driven. Again, very important. It's good for the organization, but it can also be short-term. And then you have teams that are oversight-resistant, and I see this most often with vendors.

But we all do. We're all human. I think we all, to a certain extent, prefer less transparency and less second-guessing from people outside of our teams. So there's,

again, that hero focus. There's less incentive to build reusable enterprise features. People just want to get the work done and move on to the next deadline. So, all of this with the pressure with an organization to productionize, to derive value from these analyses in a repeated, reproducible way, and the result is technical debt. Next slide.

What does this look like? I think we've all probably seen it. What you're looking at here is a code from a vendor's laptop. What is it? Where is it used? Who wrote it? It's SQL. It happens to be a bunch of diagnosis codes. But when information like this that is key to a process is hard-coded in a SQL statement that is not managed or controlled, it creates real risks for the organization, and it makes it much more difficult to QC, manage, and maintain it over time.

Next slide.

So how would an analytics engineer approach these kinds of problems in the real world? Some of this is very detail-oriented, and I won't go into the whole thing, but the idea is to break out the steps. And really what you're trying to do is understand what the dimensions of change are within a process, what changes, what remains the same, and

what you're trying to do is to restore the balance really between analytics and IT. There's some things that change frequently, there's some things that never change, and there's a whole range in the middle. If you can identify what you can

model as data as opposed to code, or what you can make a parameter as opposed to a hard-coded

piece on a vendor's laptop, or something that's untouchable because it's sitting in a centralized repository, then you can do yourself a service. And again, we're looking to balance flexibility and consistency here, and ultimately it's about giving analysts the leverage they need to control the complex machine without forcing them to do things that they shouldn't have to do manually.

Next slide.

And this is the kind of stuff that we do see all the time.

Even in

large organizations, or maybe sometimes especially in large organizations, there's some very manual processes that take place for dealing with large amounts of data. And we see huge Excel spreadsheets. We see SQL on people's laptops. Even in software like Alteryx, we see logic that can be cubbyholed in small boxes and formulas rather than made available and easily accessible for users who are looking to repeat some of those same processes in a consistent way.

Can I have the next slide, please?

00:40:00

So I want to give you an example of what a process hub can look like in action, and this example is a target list management process hub for the pharma industry. And it's just an example that's based on something that we've actually done recently. The goal of this system is to identify doctors who are potential prescribers for drugs and treatments that are being marketed by the company. And it's something you can recognize.

It's a very common task within any kind of company.

The company may have multiple brands that they're looking to market and multiple sets of requirements for metrics and approaches to select those potential customers. As part of that kind of process, there's a real interplay between the requirements of the enterprise itself, the requirements maybe of a specific division, and the specific needs of the project itself.

And there's also, over time, an interplay between the analytic requirements that Chris has been talking about and also other operational requirements that come into play as well. It's not just analytics. This data can often have operational customers or consumers at the same time. And adding to that complexity is the fact that these are really mission-critical tasks for an organization like this pharma company that face intense deadline pressure. Next slide.

So probably these kinds of challenges are familiar to you.

How do you make an analytic process that is able to handle these kinds of tasks? Calculate metrics, categorizing segments and filtering for selection, a little predictive modeling, experimentation and analytics, different iterations and versions of your product, and then dealing with changes and requirements and feedback. The key thing is this is not static data.

This is not IT infrastructure. This is more work that business analysts need direct control of in a very flexible way. But at the same time, it's a mission-critical pipeline for the organization, and the data and processes have to be consistent and accurate. So the underlying data sets are always changing. There's master data, demographics that are always being updated.

In pharma, in particular, there are regulatory requirements. There are all kinds of rules for suppressions and email opt-outs. And the QC challenges are significant. The work that's done on this really can't be off the grid. It has to be under some kind of control and part of a consistent overall process. Next slide.

So what does a process hub do? The goal is to lower the cost per question. But it's really finding that balance where you're looking to centralize the process. You're looking to generalize the process, and that's good, but at the same time, you don't want to conquer the process and take necessary controls of the process away from the people who need to be able to use it and operate it and benefit from it.

Next slide. So this is a little diagram of the process hub that we created for a pharma client, and you can see that on the left and right, that there are analytics uses and there are operational uses for the data and processes that go into the hub. For target lists, it's analysts in this case that actually

00:45:00

create the target lists, and then they do analytic work based on the results of the target list. So it's not just a one-way street. There may be common source metrics that are used to generate target lists, and there may be very project-specific metrics that the analysts use to further segment and select for their individual target lists.

There are shared

Services that are performed for all target lists, such as your standard master data management, merge purges with demographic updates, maybe affiliation updates for physicians who belong to medical practices, sales alignment updates. A lot of these tasks, it doesn't matter what the specifics are, these are the things that have to be done for any target list.

And they're also shared products that are consistently created for each of the target lists, regardless of the complexity of how they were individually generated in the first place. And because those products are shared and consistent, they can be used by team members who may be jumping from one product to another, or operations people who work with all different products, or analysts who are trying to build consolidated data sets for multiple products.

There are real advantages here to everybody concerned.

And I think that's the end of it for me. Fantastic, Chip. Thank you so much and for sharing your experience and sharing this example. It's great. So I'm going to take over and finish this out. We were going to actually do a demo, but I'm going to skip that. I always put one in, in hopes that we have time, but Chip and I got yakking so much that we couldn't. And hopefully, as a listener, you found this valuable.

But I do want to talk about one kind of almost a philosophical thing, and it has to do with an experience I had yesterday and the day before. So, I'm working with having our process hub expanded scope at a big company, and some of the business analysts-- or there's a person from IT who's joined the business analytics team, and I gave him sort of a demo of it, and he goes, "This is great. I really like it, but it'll never work in our IT team because we've got to follow our SLDC. We've got our process to follow." And then another business analyst said he's working with IT, and he didn't get what he wanted.

And he said, "Well, why didn't IT get it?" He said, "Well, we had to follow our tech standards." But the business analyst is like, "You didn't give me what I wanted," but they follow their tech standards and their processes. And so in some ways, you can think of this as a metaphor. For many IT teams, their sun, the thing that they get up to in the morning, is their software development lifecycle, their development process, their tech stack.

And what comes second for them is customer value and insight. And there's reasons for that. And so the people on the team, they're being told that this is what is the primary thing. They got to follow their process. If they follow the process, they'll get good results that they can manage. And then on business analytic teams, they're completely different.

They're focused on customer insight and value. That's what they get up for. That's their focus. And honestly, sometimes what comes second are repeatable processes and scale. And so I think this is a challenge, and I think the idea of a process hub allows both teams to win. They're both focused on great processes that scale, that don't create errors, where people aren't killing themselves, and fast delivery of customer value.

And so I think the process hub is a win for everyone, right? Because from an IT perspective, on the left, you can have a robust process. You can improve governance and documentation, and the documentation and governance documents can be kept up to date automatically. And the idea of a process hub can be leveraged by everyone in the data analytic value chain, and it drives cost reductions. And it gives you a single view of everyone who's using your data, all the models and transformations and visualizations that are acting upon. IT's got a way to see it, and they're no longer in the blind, and they have their own opportunity to add value to that.

And we've created a very transparent and open platform that's also secure. And it's all about utilizing the great data infrastructure that IT has created. And from the business analytic team, it allows them to get faster customer

00:50:00

value in a more efficient way, and lowers the cost per question and improves team productivity. And it allows them to build processes that they can control and that they can trust, which means they can speed the delivery of ad hoc analysis, and reduce the ongoing operational cost through automation of doing things that they do day in and day out.

And it does that by repeating errors and allowing not just the business analytic teams, but the people the business analytic teams partner with, that they may bring in a consultant for a few months. And most importantly, it works with the data that IT has provided. It works with their data hub. And so we think that, and maybe this is a philosophical point, that these two different groups aren't so different. They're just looking at a different sun and different focus.

But I think you can focus on both following a sort of DataOps and agile methodology, focus on customer value and focus on improving your processes, so that both teams can win. And so, that's my silly philosophical view of the world. I hope that's not going to make you roll your eyes. And we've written a lot about this.

As Beth mentioned, we've got a manifesto for 18 points. We've got a cookbook, and then we've got also a book that talks about this transformation. And we also have a maturity assessment where you can look at where you are on the scale. And so what I want to do is stop and ask some questions, and I hope that you've gotten the sense that this process hub for business analytic teams improves business analytic teams' productivity.

It solves kind of the last mile problem for them, allows people to do more with less, and allows process hubs to work with data hubs that our IT is creating. And so, Beth, could you, in the last few minutes here, is there any questions from the audience that we can- Yep. Thank you both. That was great and super informative.

So yes, if anyone has any questions, now is the time to put them in the Q&A box, and we'll get through as many as we can. So just to kick it off, Chris, I think this one's for you. Does this sort of data work ever involve moving data to or from legacy mainframe applications?

Well, you sort of saw, I don't know if Chip gave an example of sometimes the work that the business analytic teams does has to end up-- It's sourced from an application, and its result ends up in an application, right? And maybe that is a legacy application, but Chip gave the example of you take data from a CRM system, and you put data back into a CRM system, right?

And so sometimes analytics systems are operational, right? And you've got to transform data and take the results out and get it back into a system. And so I think that's actually a pretty common pattern that

we've always at some conceptual level said you've got operational systems and analytics systems. But it turns out the analytics systems, due to the fact that they're creating insight model data sets unique to them, those can be useful in operational systems. So I think that was the very last example that Chip gave. Okay, great. Thank you, Chris.

Chip, I think this one's a good one for you. How do you integrate data governance into the process hub? Are there any additional data governance tools that I need, or how does that all work together?

Yeah, it's a good question because data governance is so important to the process, and it's so easy to give it short shrift. I think it comes down to the abstraction idea, that by taking these divergent processes and being able to abstract them and run them kind of on a single, consistent, what I think of as a public transportation system, which is part of what the DataKitchen platform does.

You're able to track them in a consistent way and use and build some tools for both everything from QC of individual data points to timeliness checks for receipt of data files that may be coming from external sources, timeliness checks for processes that should be running or scheduled to run on particular times, to tools that you can use to

automatically document your data. We do a lot of profiling, and because we build our data sets in standard ways, we're able to profile them in standard ways.

00:55:00

We're able to report on table structures in consistent ways and create a lot of automated documentation. And because our data is created through these standard systems where the processes are performed, we're able to get automated versions of data lineage documentation, so we can look at process data and trace it back to exactly where it came from, even to the point of the multiple data files and intermediate tables that may have gone into that final data, which is incredibly useful.

It's documentation, in our case, that's updated daily, just based on the operations of the processes. But it's useful for troubleshooting, it's useful for identifying paths that we need to hit for issues that may arise if we're putting in new changes as well. Great. Thank you, Chip. Chris, do you have anything to add to that?

No, that's fantastic, Chip. I think the idea of governance is important, and governance should be not a bag on the side, but be built from the process. And governance should be not just what is it and where did it come from, but is it fresh and can I trust it? And I think all those things, the idea of having a DataOps process hub combined with data governance and giving the ideas that Chip came out are great.

Okay, and then specifically, a follow-on question to that. In the context of data addition, what is meant by the term data dictionary? Chip, do you want to elaborate on that? Well, yeah, data dictionary is the table layout for a particular table, but also much more than that. In

our managed services work, we provide diagrams of the relationships of tables to each other, also data lineage. We also integrate profiling into our data dictionaries, so you know not just the names of the columns in a particular table, but also the kind of data that's in there, the data that appears most frequently,

counts of different numbers of values in a

column. So you're really getting useful information that an analyst can immediately go to and better understand the data. And again, that's updated on a regular basis entirely automatically, so it's not reliant on people to go back and change it if someone adds some new columns into a process table, for instance. Okay, great. Thank you, Chip. So we only have one minute left, so I have one more question.

If we can't buy any new software for a process hub, what's the best way to get started? Is there anything, any tips you could provide for anyone to start working in this way? Chris, do you want to start with that? Well, you can always rent the process hub on a monthly basis from a company like DataKitchen. That's one option, sort of buying it upfront.

The other option is, I think, I'm a big fan of being able to start small. And I think, if you want to start working on taking the work that your team does and curating it, just find cases. One way to help is to get people thinking. I've often told people, form a quality circle and look at all the problems across all the pipelines and all the tools that are acting on those pipelines, and find where there's common problems across those and fix those.

And I think you can do the same thing on a different-- You can have a process circle where you can look at all the work that's being done and all the tools and all the pipelines and say, "Is there an opportunity for us to do-- Are we doing basically the same thing in a bunch of this work, and can we generalize it and make it

shared?" And I think those are steps that can happen just from talking to people in meetings and accomplish the goals or towards the end of having a process hub. One is the reduction of errors in production and common sort of testing frameworks. And then the second is the curation of a process and looking at, I've got four different ways to do almost the same thing.

How can I find a common way to do those and share them across? And I think that might be a good start. Okay. And we actually just had one more interesting question come in, so I'm going

01:00:00

to throw this one to Chip as a way to close it out. Is there a way to avoid the stage when the analytics team starts to burn out under the influence of a huge amount of ad hoc requests? Is it possible to deploy DataOps before this stage, or is the burnout process necessary? Chip, any last thoughts on that?

It certainly isn't necessary. I hope not. Don't do it. Take it from both of us. We've suffered enough. We've got 60 years of suffering between us. Don't do it.

So, yeah, you can. It's great. If you've got to suffer to prove that you're great, then go ahead. But

fix it, because the suffering's going to continue to be there always, and you don't have to do it. And the solution is to change your frame of reference. In addition to focusing on data, focus on the processes and making those as good as you. Put up as much work into making those processes shine as you do in making the data shine. And then you're going to solve those, those headaches and nightmares are going to go away. And we think our process hub is a good way to accomplish that. But by all means, if you want to suffer and wear the hair shirt of suffering, I did that for years. It's not a good look for me.

All right. Well, we are up against time now, so I want to thank you both, Chris and Chip, for this excellent overview of the process hub. Thanks to all the attendees for joining us today. We hope you enjoyed the webinar. We'll be sending out the recording and the slides in the next 24 hours or so, so be on the lookout for that in your email.

And if you have any additional questions about the presentation or Process Hub, don't hesitate to reach out to any of us. You'll find our contact information in the email as well. So thanks again, everyone, and have a great afternoon and evening. Thanks, everybody. It was really fun. Thank you.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What is the analytics last-mile problem?

The last mile is the stretch between the data IT delivers and the insight a business customer actually needs. Analysts receive many raw tables and spend most of their time on fragile, repetitive, error-prone data work rather than on original analysis. Because there is no tailor-made data representation, no current documentation, and no reusable business logic, ad hoc analysis takes longer and training new analysts is difficult.

What is the difference between a data hub and a process hub?

A data hub is a repository, exchange, and collection point for enterprise data, and it moves one way from raw to processed layers. A process hub holds the people and processes that act on data, works alongside the data hub, and handles the fact that data flows between teams and functions rather than simply down levels. The argument for it is that valid data is a product of an effective process, so there is no single version of truth without a single version of process.

How does a process hub increase analytics productivity?

Three mechanisms. The Wedge gives business analytics trusted, analyst-ready data sets under its own control with rapid changes and high reuse, lowering the cost per question. The Hammer automates end to end from source to delivery so recurring production work stops consuming staff. The Store keeps all business analytics intellectual property, including SQL, models, reports, tests, and scripts, in a version-controlled repository, cutting risk and redundancy.

What tasks does DataOps engineering automate?

Eight of them: production orchestration, production data monitoring and testing, self-service environments, development regression and functional tests, test data automation, deployment automation, shared components, and process measurement. DataOps engineers take nuggets of existing code, meaning ETL, SQL, Python, XML, and Tableau workbooks, put them into pipelines, wrap them in tests, and run the resulting factory.

Why can't a process hub be owned by IT alone?

A process hub needs business expertise, analytics expertise, and a meta-understanding of which parts of a process generalize and to what level, and it is subject to constant change as conditions and questions change. Analysts alone cannot own it either: it requires enterprise priorities, cross-functional cooperation across multiple teams and vendors, and a process orientation rather than a deliverable orientation. IT and business analytics work toward different suns, IT toward its development process and tech standards, analytics toward customer insight.

What is the risk of leaving business logic on a vendor's laptop?

Undocumented SQL on individual laptops is tribal knowledge, and tribal knowledge means high cost and slow speed to insight. No one can answer what the code is, where it is used, who wrote it, whether it is tested, whether it is current, or how to find and share it, and it can be lost when a person or vendor rotates off. Storing that logic in a version-controlled process hub makes it reusable across projects and survives team turnover.

Where to go next