On-Demand Webinar · 1 hr 2 min

Differentiation Through DataOps in Financial Services

Simon Trewin, co-founder of Kinaesis, joins Chris Bergh to work through how financial institutions move fast without breaking things: virtual environments and continuous deployment for delivery speed, automated testing across end-to-end pipelines for quality, and what regulation demands of both. Recorded February 2021; updated August 2026.

Presented by Chris Bergh

What you'll learn 6 points
  • Kinaesis builds its financial services DataOps practice on six pillars: a Target driven by a business vision and mapped as user journeys, Instrumentation of the data flow at every step through profiling, data quality and monitoring, Metadata that keeps business definitions and models connected to the data, an extensible Platform that can absorb new business demands, Collaborative analytics across consumers, IT and data owners, and Control through version management, release management and exception handling.
  • Kinaesis defines the target with a five-part process it calls S.C.O.P.E.: Storyboards of the user journey, Content represented with the correct context, Output written down and agreed across stakeholders, Process for interacting with the data pipeline, and Estimate as a cost-benefit analysis of any pipeline change.
  • Data projects in large financial organisations do not naturally iterate or move fast: infrastructure takes a long time to procure, sources are complex and have competing priorities, controls and governance add layers, small data gets big quickly, pipelines are complex, and the methodologies in use were borrowed from software engineering.
  • A large banking group running a BCBS 239 programme was behind schedule with eight months left before non-compliance. Kinaesis added a focused team of five or six consultants, mixing data specialists, risk subject matter experts, analysts and technical specialists, and broke the work into trackable iterations.
  • That engagement produced 300 metrics across seven lines of business, from one month of upfront analysis followed by six months of delivering metrics incrementally. It identified 207 reconciliation breaks, fixed billions of pounds of reporting errors, produced 2,000 rows of metadata and a three-month evidence trail, and left the bank with a new operating model.
  • DataKitchen frames the financial services work as four areas improved iteratively: decreasing the cycle time of change, lowering error rates in production so customers trust the data, improving collaboration within and between teams, and measuring the process. A top 5 US bank applied it to self-service, building governed data sandboxes for more than 1,000 non-IT users with legal rules on data usage and lifetime, monitoring of usage, and a path where an idea earns the right to be reimplemented centrally.

Slides

45 slides

Transcript

Show chapters and dialogue 9,639 words

00:00:00

Good afternoon, everyone, and good morning to some of you as well. Thanks for joining us today. My name's Beth Pfefferle. I'm the VP of marketing at DataKitchen, and I'll be the host for today's webinar. So our topic today is differentiation through DataOps in financial services. We're very excited to have a special guest with us.

Simon Trowen is joining us to share his experiences and insight. Simon is the co-founder and CEO of Kinesis, which is a consultancy specialized in DataOps solutions for financial services, and they're based out of London. He has more than 25 years of experience in risk, data, business, and technology within the investment banking industry, having worked at firms such as Citi, RBC, HSBC, and ING Barings.

Simon's a thought leader in the DataOps space and founder of the DataOps Think Tank. He will also soon be launching an online learning center for budding DataOps engineers, and also soon be releasing his first book on the topic, which we are looking forward to. So welcome, Simon. Thank you. He'll also be joined by Chris Bergh.

Chris is the CEO and founder and head chef at DataKitchen. He's also a leader of the DataOps movement. He also has more than 25 years of research, software engineering, data analytics, and executive management experience. At various points in his career, he has been a COO, CTO, VP, and director of engineering. He's also a co-author of "The DataOps Cookbook" and "The DataOps Manifesto." So before I hand the reins over to Simon, just a few housekeeping items. We're recording this session, and we'll send the video and slides out to everyone as soon as they're ready.

We'll also reserve the last 15 minutes of the webinar for questions. So if you have a question during the webinar, just enter it in the control panel, and we'll make sure we have time to answer those at the end. So with that, you can take it away, Simon. Hi. Hi, everyone. Thanks for joining us on the differentiation through DataOps in financial services.

I'd like to take you through my experience of applying DataOps techniques in the financial services and the types of benefits that you can achieve. Firstly, I'll introduce Kinesis. Secondly, I'll provide you with some background about myself. Thirdly, I'll then explain why I think DataOps is needed in the financial services and what DataOps means to me.

To back up my belief that DataOps really works, I'll provide you with a case study where we helped a client implement a regulatory solution using DataOps methodologies that created a step change in their results.

Kinesis is a 10-year-old consulting and services organization. We are the leading DataOps consultancy in financial services, specializing in delivery of high-performance data architectures, enterprise information management, and high-performance analytics and reporting. We've delivered over 85 projects with an extremely high success ratio, which is not typical of data projects in financial services. We attribute that success to our DataOps approach and our ability to break down challenges and deliver incremental value effectively.

Our mission is to share our knowledge on DataOps to help enable clients and organizations to achieve similar success. To this end, we founded the DataOps Think Tank on LinkedIn, which has over 600 members and contains lots of posts related to DataOps. To add more value for our consulting services, we have recently extended our IP and development of DataOps products to help accelerate data and analytics projects.

These include DataOps training based on our proprietary Six Pillars methodology. This is used to enhance our consulting engagements. Akutex is a new product for us we have built to create a data exchange for small data. KYR is a tool to provide investment managers with customer analytics. Clarity Metadata helps you see business-meaningful insights from metadata.

And finally, IBRA is a product that we offer in partnership with Cloud Risk to value and calculate risk on financial contracts.

I've been working in financial services for over 20 years. There have been many highlights in that time. To name a few, I was business manager for fixed-income credit trading team at BNP Paribas. I led 30 IT developers implementing a large risk and trading platform for credit derivatives at RBC. I led 110 IT staff at Citigroup, managing 46 front-office applications for their structured derivatives business.

More recently, I led teams of consultants at HSBC, Santander, Deutsche Bank, RAC, Lloyds of London, in delivering regulatory and customer-driven data

00:05:00

solutions. I've learnt a great deal from these experiences. One of the big benefits that I've had is to work on both sides of the fence, in IT and in the business. This has enabled me to see the world from many perspectives and has led me to believe that DataOps is the right way to build collaboration, solve business problems through data, and deliver real change.

I started back with data at university, where I won the Addison-Wesley Prize for computer science. I realized that I saw things from a set-based perspective rather than an equation-based perspective. That is typical of computer scientists. This enabled me to grasp concepts around data rapidly and efficiently. My journey to where I am now has been very varied.

At some points, I've wrestled with most challenges around data solutions, and I've refined my approach over the years. Some of the challenges are hard to explain. One particular example comes to mind from a very large bank. Where we were being pressured to make regular releases to the production system. To do this, we needed to get the admin password to make a release on the database. The challenge was that the security team were in Eastern Standard Time and the DBAs were in India.

There was a 30-minute time window each day when we could schedule this to happen, and we were not the only project needing their support. These are things you cannot appreciate until you've been accountable for making things happen in a large organization. On the architecture side, I've been through the journey of normalized databases, to data warehouses, to Kimball, data vault, big data, and most recently, cloud data solutions.

I've delivered these types of solutions in some of the most pressurized environments known. Two examples are front office IT and investment banks, where information that you present is the difference between making money or losing it, and many of the users do not have much patience for losing money. Has anyone seen "Wolf of Wall Street?" Traders can really lose it when you say things can't be done.

The second example is regulatory deliveries, where banks face big fines or face having their operations closed down if they cannot provide accurate data on time and well-explained. I'm going to share with you my perspectives on how you can gain real value through DataOps.

Why do I think DataOps is needed in financial services? The way I'm going to paint a picture is first to present general challenges within financial services. The items from Tribus Consulting provide a good summary of the financial services press over recent months. They all point to leveraging data to identify opportunities, compete, and manage risk.

Firstly, I want to say the progress that has been made in financial services and in government regulation and intervention since 2008 is outstanding. Who would have thought the financial system would be able to ride out COVID-19 so smoothly, given the financial impact? We're not out of the woods by any stretch, as some of the after effects will be delayed, but the focus of the pandemic has been on health and not the breakdown of finances.

What are the challenges in financial services? The regulation that has been applied over 12 years has added layers of complex data processes, and the regulators are getting more sophisticated. COVID has accelerated tap and go as a default payment method for everything. This has sped up payment processing to new levels, bringing in challenges around fraud detection and payments analysis. Big tech firms through online services are gradually disintermediating the incumbents from their value streams, and they need to react to bring their services online faster.

Credit positions of traditional customers are under strain due to COVID-19. There is an expectation on machine learning and AI to improve decisions and predict future crisis. All financial organizations carry a lot of legacy borne out through multiple mergers and acquisitions and having a lot of legacy technology, which is not well understood.

Why DataOps? Firstly, data is important. According to a report published by McKinsey Global Institute, companies that are data-driven are 23 times more likely to acquire customers, six times likelier to retain customers, 19 times more likely to be profitable and successful. This is the same for financial services as it is for any organization. My experience is that data projects in large financial organizations do not naturally iterate, simplify things, or move fast enough.

In my experience, this prevents them from gaining momentum and implementing the right solution. The reasons for this are as follows. Infrastructure takes a long time to procure, although virtual machines and

00:10:00

cloud are changing this slowly. Sources are complex and have other priorities. You need to ask six months in advance often to get hold of data. Control and governance are extremely tight as we are dealing with personal information and money. With the recent regulations around mis-selling through MiFID II and right execution, not to mention GDPR.

Over a long period, a lot of small data has built up in spreadsheets and EUCs that mean the single version of the truth is hard to come by. This has happened because of the need to create many flavors of the same data in different ways, and also to speed up the change cycle.

Things get big quickly. Once a project starts showing success, then it quickly gains supporters who pile on requirements. I remember we built an analytics solution at Citi in 2006, and immediately the whole of fixed income front office wanted one. Complexity exists due to legacy and small data processes that are hard to unpick. More about this in the case study later.

Methodologies are based on software engineering, for example, behavior-driven development that is not suited to data projects.

Why does DataOps approach and tooling help you? DataOps enables you to break down problems into measurable chunks that fit agile methodologies effectively. By using DataOps techniques, you can identify waterfall elements and agile elements of the plan. Get this wrong and you are stuck with a slow release train. How many data projects are there that are stuck in three monthly release cycles?

DataOps helps you to define the full workflow and pipeline relating to data that sophisticated users will want to use. This includes the provision of sandboxes and environments where the user feels empowered. DataOps helps you to differentiate the data and workloads that will iterate at high and low cadence. This is important, otherwise the project will become bogged down.

DataOps enables you to build metadata-driven systems that empower your users rather than leaving them cold. DataOps helps you to design a system that will respond rapidly to changing requirements by setting up frameworks and building in scalability, both horizontally and vertically. DataOps establishes a framework for building collaboration across all stakeholders. 60% of the success of your data analytics project is based on the sponsor.

This was originally written by Kimball in his "Data Warehouse Toolkit" book, and my experience doesn't dispute it. DataOps adds appropriate governance and controls that add value rather than implementing documents that sit on shelves.

So what is DataOps to Kinesis? DataOps is nothing new. We've been doing it for years under different guises. You probably know it as operations development, data management, data visualization, data tooling, data analytics. DataOps pulls together these tools and techniques into a structure that really allows you to move forwards fast. It's a bridge that enables departments and functions to work together cohesively around delivering data to where it is needed and when.

At Kinesis, we break down our methodology into six pillars that, if executed effectively together, will give your projects real impact. These are instrument, metadata, platform, analytics, control, target. Through our training, you'll see how these help to deliver a data strategy rapidly and effectively. DataOps is a wide discipline. It needs to integrate into all the functions around data. The proprietary methodology that Kinesis uses, teaches, and publishes is based on all of the data disciplines and how they need to interact to enable you to build a DataOps culture.

At Kinesis, we leverage the tools and techniques within the six pillars to different levels, depending on the requirements of a project. Different projects require different emphasis, and therefore combine the pillars in different ways. Over the next few slides, I'll break down the impact pillars to describe how they come together to support your DataOps journey.

We've put target at the end of our pillars. However, throughout the project or program, you need to have the target in mind. Up front, you need to define it in a way that allows you to measure the success and the goals all the way through the process. Once defined, this forms the governance for what is delivered.

What is the vision for the solution when it works properly? The vision is more than just a computer system. The vision is a business process, a set of data, and operating models and technology solution. Are your projects defining the target correctly? We find that many people define their target incorrectly. If you focus on reports and calculations, you miss a key ingredient, which is people and process. Then you're prone to building a system that nobody

00:15:00

uses. With DataOps, you can really quickly get to the bottom of the business challenges and define what needs to be done, bringing everyone on the journey. At the end of the pillar descriptions, I'll break down what the difference is between DataOps target pillar and the standard software techniques for business requirements that we see in financial organizations.

One of the most important pillars is instrument. To do anything, you have to know the current state and know which bits are important. To do that, you must measure quality, data shape, structure, profile, lineage, physical size. The importance of instrumentation is to identify and create strategies around your risks early. Risks exist in a data platform due to data quality, feeds appearing at the wrong time, data not available, poor calculations, and data being too large to ship around.

How many people here have only found out issues like these just before UAT when the clock has run out? Wouldn't it be better to know these things three weeks into the project?

The metadata pillar describes how to leverage your metadata efficiently and effectively. We'll show you how to leverage the metadata to build heuristic processes that improve your productivity over time. One example is the metadata that links technical implementations to business meaning. Do you understand what your data really means? Have you considered the scope of a variable, what it was filtered on, the nuance of the calculation, the context of the question that it is answering? Data only really means anything if it's associated with business terms or business meaning.

Have you ever been to a meeting and different departments produce different numbers for the same metric? How much time do you spend in the meeting discussing why the numbers are different? How many times have you implemented a project where it is 60% similar to the previous project, however, you started from scratch because you did not know what you had done before?

Well-structured and defined metadata alleviates this issue. The models that you can develop with knowledge of the metadata techniques enables you to build pervasive solutions that get better over time. The metadata pillar involves lots of good DataOps patterns that help you accelerate your analytics and delivery programs.

The third pillar is platforms. Getting the platform right means that new requirements can be met quickly and efficiently. This could mean plugging in new tools, new visualizations, new data, new models, or new regulations. Has anyone here implemented a big data platform and been underwhelmed by the results? Once implemented, it is often expensive to extend.

The first iteration goes in, and then adding a new column takes another three months. The testing cycle seems to run longer than the implementation. Getting the platform correct with DataOps gives you the flexibility to respond. DataOps platforms help you organize existing data management tools and techniques in patterns that bring flexibility. Through best practice designs and methodologies, we've been able to achieve these results for large-scale regulatory deliveries, reducing time and effort.

The fourth pillar is collaborative analytics. Collaborative analytics is the art of forming a working group around producing data and analytics. This is not at the end of the project when it's in UAT, but throughout. Traditional waterfall does not work with data projects. Why? When working with data, there is a lot of unknowns, data quality, methodology, definition, scope, and business rules. Until you see the outputs, you're not sure if they are right. This isn't because people don't know what they're doing, it's because it's complex, and until you see it, you're not sure.

It makes you wonder why people try to run projects using traditional waterfall. The only way is through structured collaboration and an iterative learning and prototyping and delivery approach. DataOps builds bridges between departments to execute analytics reliably and effectively.

The fifth pillar is control. In my past, in the front office environments, DataOps techniques enabled us to move forwards rapidly with modeling and data changes because we had the governance and control around the processes of changing the system. We could quickly assess the impact of the changes which enabled the business to sign off efficiently and effectively.

We also automated much of what we did. It meant we could move as fast as the world demanded. Within software development, the code is version controlled. In DataOps, software models, metadata, and data should be version controlled. The governance of these additional items is not a cost.

00:20:00

It's a way of minimizing effort by identifying issues early, dealing with them systematically. It is a way to create trust in the numbers by always having a version of the truth that links back to where you were. With the right tools and techniques, you can fit the right governance structure to your problem and deliver results rapidly and effectively.

I'd like to break down the target pillar to the next level to highlight the differences between DataOps methodology and the standard business requirements process that we see in financial organizations.

Kinesis breaks down the target pillar into the scope definition process. We use this because it's better for working on data projects than the business requirements document or BRD. From my experience, the standard BRD process that is used in many organizations misses some key elements relating to building data systems that people will use. I believe that the following scope process provides more clarity on what is required. In my experience, BRD processes describe calculations and reports. They often miss the human interaction parts and the data disciplines of modeling and data architecture.

Often, BRDs are ignored by the developers who source the requirements themselves due to the fact that they do not contain the right information. Why? Because people are good at articulating the data or reports they need. They are not good at telling you how that fits into their processes. Therefore, if you deliver a system based on a BRD, it will be missing the human interaction parts and the deeper data solutions around analytics.

To compensate for this, the users will take the resulting system, extract the data, and work with it offline. This leads to poor user engagement and the eventual breakdown in communication between the business and IT. On the next few slides, I'll explain storyboards, content, output, process, and estimate.

Storyboards are about describing the user journey in conjunction with the users to be able to map out the true vision of what you are trying to achieve. To build up the storyboards, there are a set of questions to ask that help you extract the right information. The storyboards combine with the overall flow of the system and how the users interact with it and articulate the use cases for the analytics. BRDs often miss out these elements, and they are deeper than the standard UX design processes.

By mapping out the user stories, you get sight of the data that is required to meet the requirements alongside the flow of information. Coupled with this, you'll be able to capture the structure of the data that the users need to satisfy themselves that the system is telling the truth. This could include reconciliation reports, completeness and quality dashboards, drill downs, drill throughs, and comparisons to previous numbers.

From my experience, BRDs miss these elements.

The content section of the target pillar captures the information that has come out of the storyboard process. Laying out the content into appropriate groupings enables you to capture the shape of the data and how it contributes to the overall use cases and user story. By laying out the content in the correct way, you'll be able to establish the life cycle of the data and the touch points that it has with the users.

Not all content is needed at all stages of the data pipeline, and carrying it will cause your agility to slow down significantly. The content section of the target pillar will provide you with simple templates to capture this information, categorize it correctly, and establish the foundations for overall design and flow of the system. My experience with data projects in large organizations is that the data mapping documents that form part of the standard BRDs are poorly formed and put together by business analysts and not data analysts.

They often do not incorporate the simplifications in scale that you get with data systems, and they ignore generalizations and patterns.

The output of the system is normally a non-negotiable deliverable. Where the content captures data throughout the life cycle, output captures the specific outputs from the system These are quite often in the form of mandatory reports or user-specified outputs. The output analysis is about decomposing these outputs into fields and groupings of entities, and being able to identify the different types of data that exist in the reports.

From this output, we should be able to trace back the fields to their expected sources and be able to set up the lineage in the system to identify this. At this stage, there is power in playing the solutions back to the users so that they can buy into the overall outputs of the system.

In my experience, typical BRD processes ignore the true meaning of information and also the set-based nature of data. In one example, I can recall having to simplify the specification document that tried to describe every cell of a report, where it was clear that the report was just a pivot

00:25:00

table of the results combined with two dimensions. The formula and the dimensions would have described the whole report.

The process part of scope is to identify two main categories of process that will be involved in the overall solution. The first of these relates to the timeliness of the source data. It captures when it is due, how much of it there is, and by what medium it arrives. At this point, it is important to identify where data in the system needs to be synchronized. Having sources that are correct but based on different snapshots is one of the causes of poor data quality.

It's important to capture this information, as you may find out that the requirement is not as achievable as you'd hoped. The second part of the process is to identify transformations and models required to take the source data to turn it into useful insight. Some of these models will be black boxes that require certain inputs and provide outputs. Others will be things that the data pipeline is executing, like joins, filters, and transformations.

In one consulting engagement that I worked on, the business sponsor asked for the system to show data two hours after the final calculation of risk. The output that was specified relied on a breakdown of specific trade data, which was not available for four hours. It was therefore not possible to produce a risk in two hours.

We could have spent weeks optimizing the risk pipeline to meet the two-hour requirement, but by analyzing the process, we were able to push back on the requirement and focus on higher priority work.

The data that you have captured in the scope process should enable an experienced DataOps engineer to establish a high-level view of the work required to get the data from the sources to the outputs, supporting the content that has been identified along the journey. The captured metadata should enable them to map out the functional process and the data and establish a sensible data architecture.

Knowledge of the existing landscape or tools available will provide the data engineer to then work out a high-level architecture. This is key to establishing the cost of the requirement. Married up with the prioritization process that can be done in conjunction with the end users should enable the right features to get prioritized in the right order in an agile process.

Easiness for DataOps methodology is not too different from the weighted shortest job first process of the SAFe agile methodology. However, it incorporates many of the elements of scope that enable the participants to have a clear picture of the data requirements. The other pillar of the DataOps methodology break down in a similar way to target and enable information to be captured and leveraged to build up your delivery pipelines.

I'd like to take you through a case study where we have used DataOps methodology to great success.

Kinesis was asked to help a financial organization overcome a regulatory burden rapidly and effectively. The data regulation required the bank to deliver a large number of key risk indicators across all of its data silos. One of the key messages from this is that with DataOps, you do not need big teams.

A large banking group undertook a large program to strategically deliver BCBS 239. BCBS 239 is a regulation that holds financial organizations accountable for the data that they use and report on the health of the organization. It is principles-based and requires the organizations to be accountable for the accuracy, completeness, and quality of the reports that the organization use at the executive level.

The original project that this particular bank undertook got too large and therefore quickly bogged down and was not moving at the speed to meet the regulation. With eight months to go, they were facing a scenario of non-compliance. They called on Kinesis to set up a team to help them meet the compliance requirements in a way that they could be extended strategically. The challenges were not small.

Data has to be brought from silos with high complexity and hundreds of small data applications driven by end users. Data was inconsistent and came from legacy systems going back over many years. To make it harder, the metrics that they needed to create were not aligned to the normal metrics that the business used day to day.

Kinesis put in a small focus team of between five and six consultants to provide extra capacity to meet the deadline. The team consisted of data specialists and SMEs at risk, as well as analysts and technical specialists. The DataOps methodology enabled the bank to break the challenge into trackable iterations. Dashboards and metrics and rack statuses kept everyone focused on the key deliverables. Building in flexibility into the solution enabled the bank to hit the changing requirements.

00:30:00

Full transparency of the progress enabled stakeholders to trust the delivery process.

The results of the engagement can be summarized as follows. Project scope was achieved well in time for the deadline. The joint team implemented 300 metrics across seven lines of business where there were no existing governance systems in place and limited tools to leverage. The engagement involved one month upfront analysis, followed by six months of iterative development, delivering metrics.

The team gathered 2,000 rows of additional metadata to support compliance and to report to the chief data office. The team established a fully compliant operating model. Some of the wins on the project, it identified 207 reconciliation breaks against regular production numbers. This represented billions in reporting errors that were fixed over the period. The project enabled three months of evidence to be collected as part of the compliance process.

The client was able to provide evidence to the regulator to keep them happy.

Here are some quotes from the project stakeholders. The COO of Risk said, "We should be able to make a step change in our capabilities." The solution provided them with confidence to tackle other projects across the bank. The CDO was so impressed with the compliance of the solution that they thought they would have to mark it down to be believed.

One of the team members working with us said in a town hall meeting, "Kinesis has restored my faith in consultancies. Other consultancies come in and leave us with some PowerPoints. Kinesis implemented a real solution." This was down to the ability to build momentum and confidence using DataOps methodology.

In summary, I've seen a lot of success of using DataOps techniques and delivering real change in financial organizations. This success has been because the DataOps methodology has enabled us to get data projects to iterate. This has enabled us to engage the business and help us to understand and deliver what was really needed. It has enabled us to incrementally tackle big, complex problems with the process.

As the saying goes, you do not eat an elephant in one bite. The secret of data is to set up processes and approaches that enable you to deal with change over time and build in continuous improvement as part of an evolutionary cycle. This is what DataOps enables you to do.

Some resources Kinesis provides for you to help you on your DataOps programs. Later this year, I'm publishing a book whose provisional title is "The DataOps Revolution." Coming later this month is the first of many online training courses to help with your DataOps skills. Please join us on the DataOps Think Tank group on LinkedIn.

It would be nice to get discussions going. Finally, we publish videos about Kinesis DataOps tools and techniques on YouTube. Please take a look. Thank you for listening. Now I'd like to hand over to Chris Bergh of DataKitchen. Thank you, Simon. That was fantastic. And thanks for speaking with us today. So I'm going to go share my screen, and talk a little bit about

DataOps. And so, hopefully you can see my first slide.

And so, yeah, Simon, you did a great job sort of describing what DataOps is, and I'm going to sort of take it a little bit in a little different area. And DataOps is kind of a set of technical practices and norms that help people kind of focus on, as you said, the cycle time at which they can deploy, or the error rates that they run in production, or that great complex collaboration across all their technology and people and location and measurement of the results. And what we find in financial service companies and others is that when you start to think about, "Okay, I want my team to adopt DataOps," there's a bunch of areas that you could focus on first. And one is, "Man, in production, we're just having a lot of problems.

Things are breaking left and right. Our data isn't correct. Our brokers, our customers are yelling at us." That's one area. Another area is it just takes forever. It takes three to four months to deploy 20 lines of SQL in a data warehouse, and we've run into insurance companies and banks where that's the case. And then, as you eloquently talked about, just the collaboration between sort of centralized teams and self-service teams, and just how to have less meetings and bureaucracy.

And so people kind of think of these as levers at which they should work. And one of the reasons that they kind of arrive at this, as you talked about, Simon, was just the evidence is that

00:35:00

they have these problems. I think they sit across the lunch table with their digital teams who are working on their company website or their company IT systems, and they say, "Yeah, it takes us three months to deploy 20 lines of SQL." And they get sort of snickered at because the website team can deploy every day or can deploy at least every week.

And then the amount of errors that people have, I think, are tough. And so...

Ah, okay. I'm actually showing the wrong thing. So let's see.

There we go.

Yeah, and so they have these problems. And so if I look at kind of how organizations work, and we're just going to kind of talk through two examples of how they work. And one example is, I think everyone who wants DataOps sort of wants the end. They want to have to be able to deploy quickly. They want to run with low errors.

They want their self-service, their centralized teams, their data science teams, their people in different locations to work together. They want to be able to be analytic about their processes. And so, the question is, where do you start? And so one bank kind of started in this place where they said, "Let's just focus on ... the collaboration part of this equation.

Not to say that they didn't want the other parts, but you've got to start somewhere. And so in their case, they started with this idea of a data sandbox or a business-led analytic development environment. And they had just lots of users across a large, complicated bank in the United States, and what those users wanted were kind of short-term use of specific tools.

And what IT wanted was to follow legal rules on sort of data usage and lifetime and monitoring. And they wanted this sort of relationship between all the non-IT users and the IT users to be automated and not sort of slow and manual and error-prone. And so one way to look at this is that they have a centralized IT team, a data team that has the data, has the hardware and software, and this is a United States case. Around the United States, there's different groups that are using the data.

So for instance, in Brooklyn, New York, there may be a small business loan group. In Texas, there may be a high net worth management group. In the Northwest of the United States, in Seattle, there may be sort of a branch regional banking. And they're different people, right? And they want different things. They want different datasets. So for instance, the New York person may want business data intersected with DMB files, and they want a SQL database, and they like to use Tableau to do their work.

However, the person in Texas may be more technical, and they want to actually do Python work on that data. And then the Seattle person may just want the DVA data and want to use Power BI. And so there's different tools, different datasets, and from the IT perspective, that's not the issue. It's more about the issue of, how do I make sure that there's not problems when this happens?

How do I make sure if that person leaves or they do something that's out of compliance, I can pull away those resources? And so what that means is they need kind of a sandbox, right? Where the data and the tools that are acting upon the data are given to them. Think of it as a data analytic environment, and the IT team can monitor and govern these and kind of take it back if something's wrong. And so the fact that there are self-service teams using data and is kind of the way the world is now, right? And the tools are great, and the diversity of tools, and the difference between a Python user and a Tableau user, I think is-- And you're just going to find lots of skills around the organization. But from the view of the central team, they feel like they're at risk, and they get blamed when things go wrong.

So how can they build a system that allows freedom, but with some degree of centralized control? And the other part that I think is really important is-- And that's this sort of self-service sandbox that we're working with them on. And once that happens, so once you give someone a resource, and you give them some data, you give them some tools, they're going to do something with it. And as Simon said, if you think about that work that you do, and maybe that ends up being a Tableau workbook, which is some XML document, or it ends up being Python, which is a file that ends in .py, a text file, a Python file. That's the work that you've done.

And so what's the path to actually make that go into production day in and day out? And maybe there is no path, right? Maybe you're doing it, you're answering a question, you throw it away. But other cases, you create something of value, and maybe that

00:40:00

person who is in the high net worth using Python creates a segmentation that everyone in the group wants to use from that day forward. And that's in a Python file, so how do you get that into production? And if you think about it that way, there's kind of two paths to production that you need to have a complete solution. One is, well, maybe that Python file, which ends up being a segment of a customer base, should actually end up being an attribute of a dimension in a Kimball-like star schema. And that could be fantastic, right?

Maybe that's the template or the logic that has to be re-implemented in the ETL process that runs in part of the central IT group. And what the self-service team done is discover requirements, validated them, and then the central IT group has to re-implement them. And that's a great way to do it. Another way to do it is to take that Tableau workbook, that Python file, and kind of wrap it in a DataOps wrapper.

That is, put it in source code, test it, deploy it. And in our software, we wrap it in a thing called a recipe with tests and orders. And so whether that piece of work that's done in your self-service environment earns the right to go on a path to production or earns the right to be re-implemented, I think thinking about this is an incredibly important thing because this centralization versus freedom argument permeates every large organization. And just like in the United States, sort of federal rights versus states' rights versus this centralization and localization discussion happens everywhere.

And I think it's important that we try to find a solution to that in data and analytics, and this is a bank's attempt to do one. And at the end of the day, it technically just ended up being a request form. I want to request it. And then a bunch of in our software recipes and monitors to make sure that that works.

So if I go on to the next case of financial services is really about transformation. And a lot of organizations are kind of collection of separate organizations. They are lines of business. And so for instance, in large pharma companies, you may have a part where there's a science part that discovers the drugs. You may have a part that's commercial that actually sort of sells and markets the drugs.

You may have another part of the company that is actually involved in manufacturing, and then you may have a supporting function. So you have these four organizations, and oftentimes they're completely separate. And so you may have an organization that is trying to influence all those businesses. And in financial services, I think we all know that there could be banking and brokerage and insurance and high net worth and different parts of the organization. So how does a team, like a chief data office or an enterprise data office, drive change in a big organization?

And so if you look at it, the way a lot of big financial services organizations work, they're sort of a collection of independent businesses. And what happens in that is there's obviously data silos that exist, but there's also team silos. Different data and analytic teams are aligned to banking or brokerage, and sometimes they're aligned because they're literally different IT organizations, and they sort of don't know about what happens in the other organizations. Yet you have this chief data office or these people who are responsible for kind of taking the guild or taking all the people and helping them do their job better.

And so how do you live with just the silos of these teams? And they're sort of working together in their organizations and sometimes they're having standardized tools, sometimes they're having some sharing of best practices. But how does one, as a chief data office or a team who is in charge of saying, "Okay, this DataOps thing is good. We like it.

We want everyone in the organization to make it happen." So how do I influence all these teams when they're busy doing their day-to-day work to evaluate this idea and start doing DataOps? And this is not dissimilar to what has happened in other organizations. Certainly, in the software world with DevOps, they had a very similar problem of how to get this. And it's really about DataOps transformation and influence from a centralized team across all these different lines of business or sub-organizations underneath a larger organization. And so with people like Simon, we try to think about it in a six-step process to sort of bring DataOps to your organization, and especially in organizations where they have multiple lines of business. The idea of educating first and getting people sharing the idea of DataOps and best practice, and then finding

00:45:00

a first project and kind of establishing an initial set of community who are interested in DataOps in each line of business is important. Because in some ways, there's the social proof that this works, and a lot of people are, you could say, from Missouri. They want to see it first and see that it works.

And so being able to demonstrate in short, incremental projects. And so one of the worst things I think you can do with a DataOps rollout is to spend a year building your DataOps system. Apply the iterative methodology to be able to show value and make it work. And then build sort of a center of excellence or a dojo that actually can continue the rollout over after the first six months or a year.

And some organizations set up a team that has initial resources, staffed through the chief data office, who are trying to go into different lines of business and say, "Let's get a demonstration project. Let's get you going. Let's demonstrate some value and start the sort of virtuous cycle of rolling DataOps to their organization." I think that's actually a really good thing to do, and it's a big change for companies, right? Because if they're trying to make change in a large organization, just how do you make that happen?

And thinking about it in a more holistic way, and because it's not just DataOps as an idea, it's the DataOps that happens in your company, right? And how does that work? How do you get standards across how people work? How do you have common, for instance, deployment standards? How do you have common testing standards?

How do you make sure that you can get and see the processes across all these lines of business? And I think it's really a very interesting and important challenge that larger organizations have. And smaller companies have this in a smaller way. And so for us, just to end, we have a software tool that helps people do DataOps that can be that sort of central hub in a large organization or in a smaller team that helps solve these areas.

It helps you reduce errors, helps you decrease cycle time and collaboration. And like Simon, I'm very excited about Simon's sort of Phoenix project-like book on DataOps, and we took a bunch of our blog articles, and I've given, I don't know, 10,000 or 15,000 copies of our DataOps Cookbook and have a similar number of signatures to the manifesto. And also the DataOps Think Tank on LinkedIn has gotten, I think, 500 or 600 members on it, and it's very active. So there's a lot of opportunities for you to start learning about DataOps, start learning about DataOps transformation, and learn about how, from great people like Simon, how DataOps can work in a financial service organization.

And so I'm going to stop there and see if there's any comments or questions, and hand it over to Beth. Thanks, Chris, and thank you, Simon. Yes, so now we have some time for questions, so if you have any, just enter them into the question box here, and we'll get through as many as we can in the next 10 minutes.

So just to kick it off, this one is a follow-up for Chris, related to the self-service sandbox. So who should be managing the self-service sandbox? The IT team? Or if so, would they be responsible for ingesting new data, or would business users be allowed to ingest it themselves? Oh, yeah, that's a great question.

And so big In general, the IT team sort of has the finger on the technical resources and has the finger on the sort of centralized data resource. And so in some ways, they're also taking the blame when security problems go wrong or things go wrong. So they tend to be the organization that wants to make it happen.

But in that case, there are also examples that we've seen where there are local groups who have smaller data sets that they've ingested that they're keeping track of. And maybe it's as small as an Excel file that they're managing. And if you've ever done data, you know that sometimes a small file intersected with large data can create a lot of problems.

And I think it also can help these sandboxes. It can help bring those under some kind of governance and management. And so if you actually give write access to that self-service database and allow them to load that small file in, you at least have visibility of what's going on. And if you actually allow them to save their work as part of a system or a recipe, you have visibility into the code that's acting upon it. And so I think from a centralization standpoint, a central IT team, I think everything is happening in the field.

People are

joining data, intersecting data that you provide them. They're adding new small data sets or sometimes large. They're doing everything that you can do with data. They're integrating it, modeling it, visualizing it, and getting a

00:50:00

chance to govern and control that and have some visibility into that work, I think is important, or the data that they do is an important part of the idea of a self-service sandbox.

Great. Thanks, Chris. So this next question, what is the impact of work with CI/CD in DataOps? Simon, do you want to take that one? Yeah. Well, to be honest, it's a powerful case for the DataKitchen toolkit, to be fair. CI/CD in DataOps is a slightly more challenging requirement. One of the big problems that you have, unfortunately, is that data is sort of mutable.

It's not like sort of web pages that tend to not change. And that therefore means that there tends to be quite a lot of cross-dependencies around CI/CD. You can potentially change, as Chris was saying just now, a small data file, and all of a sudden, all of your production results go wrong in an area that's completely foreign to what you're currently doing.

So what DataKitchen and ourselves

try to promote is the ability to start to test the different parts of the data pipeline through a CI/CD process along the data pipeline. We call it instrumentation, as I went through the slides earlier. And that instrumentation is absolutely key because of these cross-dependencies and the side effects that you get through data. So DataOps is all about being able to contain those dependencies, understand what they are, and also manage them efficiently and effectively, so that you can actually deploy things in a reliable and sensible way that's been tested, ideally through automation, to be able to deliver pipelines sort of pretty rapidly.

Chris, do you have any- I don't know if, Chris, you've got some additional comments on that. Yeah, that's really a very great way to say it, Simon, and the way that we say it is actually really terrible. So we say CI and CD is great for software engineers, but it's not enough for doing DataOps, and you need to do CISMOIDEM, which is another collection, a really awful acronym that Beth hates instead of C- CI and CD is not enough.

And for exactly the same reasons. It's not just code. You've got data that you're working with, so the problem's more complicated. And also the deployment process isn't sort of a one dev team to one ops team. It's sort of a many to many because you've got this self-service and different teams touching the value chain and lots of tools. And so it's just more complicated.

The role of data, the complexity of the teams and organizations, and the fact that people who do data science and engineering aren't software engineers, and the tools that you have for software engineers just don't work for

the people who are doing data. And so I think all those reasons, I think thinking of it in a little bit, instead of just the I and D, the SMOIDEM idea, I think is a little bit more encompassing the reality of what it takes to actually continuously deploy or continuously deliver. And on the positive side, I think some people are trying to do it, right? They're trying to continuously deploy, and they're trying to automate their deployment process, and that's a fantastic thing to do, right?

Instead of having meetings and manually moving code and checklists, trying to apply some automation is exactly the right thing. I think you just have to sort of work back from first principles and say, the

continuous integration and deployment that happens in software is only just a small part of the problem of how you continuously deploy data in analytic systems. And we've written quite a bit about that in our book and on our website and talked about the differences.

Okay. Here's another good question. So you both highlighted data quality and governance as being foundational to DataOps. Can you elaborate on the governance component a bit more? Simon, do you want to start that one? Yeah, sure. So, I think there's a previous webinar by DataKitchen, and someone talking about, sorry, I forget her name, DataGovOps.

I think what's key within governance in large financial organizations tends to almost sort of sprout a life of its own. It sort of creates lots and lots of ancillary documentation. I like to call them tax on people's day-to-day work. Please can you not only do your day-to-day work, but can you fill in this document that describes where you got your data from, please?

What we try to promote with DataOps is actually what should really happen is that the code or the data

00:55:00

pipeline should largely be self-documenting. If you get your metadata processes right, ideally, you can make them lean enough that they self-describe. These types of patterns have existed within DevOps for a long period of time, where code self-documents itself by the comments that are within it. The same can be carried out through metadata within DataOps pipelines. We've successfully deployed a number of different solutions in the past where you actually use metadata to drive the data pipeline. What then happens is you report on that metadata, and it just provides documentation as to what happened.

What it means is that the document is no longer sat on a shelf and needs to be maintained separate to the code, but the actual code itself is documenting itself, taking the tax away. That's kind of the slant on governance that sort of DataOps gives you. Obviously, many of the other parts of governance around who's involved, who's responsible for changing data, who's got permissions to do it, who hasn't got permissions to do it, much of that is the same in DataOps.

The key thing is to try and embed it into a value-add process so that it sort of falls out of what you're doing rather than being something additional to what you're doing.

Chris, anything to add there? Yeah. I'm going to say something ironic because I'm showing a PowerPoint file, but the idea of DataGovOps is try to minimize the amount of data governance you're doing through Excel spreadsheets and PowerPoint in meetings. And try to improve the amount or maximize the amount you're doing as code or automated. So data governance is, you could think of as data governance as code. And Simon gave, I think, an example of that I'll elaborate on.

So like a concrete example is you're going to add a new column to a table, which is going to end up in a report and reflected in a model. And so, well, what are the data sets in that? What's the range of data sets? And have you properly put that in your data dictionary for people to know about? And so one way to do it is, as Simon said, are these largely disjoint PowerPoint, Excel-driven processes that are separate.

But why can't you take that and deploy the change to the table and deploy the change to the metadata of the table, the catalog, the description, how you've looked at the data quality? Your data profiling tool has said there's 15 unique values in this column, and that becomes part of the description that's in the data catalog.

Why can't that be deployed at the same time as code? And I think that'll give you a much better chance of linking and not getting these systems out of whack. And even the other part of sort of the security of that. Some databases allow sort of column-based security, and so do you have the right security?

Is that security deployed as code, as scripts? And so I think it's not to say that data governance shouldn't have meetings and shouldn't use PowerPoints and Excel, and it would be ironic since I'm talking from a PowerPoint here, but I think one of the core intuitions is can you do it in code? Can you write a script?

Can you have that script be embedded in a system that handles things like deployment and testing and monitoring? And if you can do that, then I think it's much more likely that you're not going to have these problems of people not listening to data governance or the data governance world being completely out of whack with what is actually in production today.

Great. Well, we are at the top of the hour, so I think we'll conclude here today. I just want to thank all the attendees for joining us and taking the time. Thank you, Chris, and an extra big thanks to Simon for joining us today and speaking and sharing his insight with us. Thank you for having me.

For all the attendees, we'll be sending out a recording of the webinar and the slides within the next 24 hours or so, so be on the lookout for that in your email. If you have any additional questions about the presentation or any of the content here today, don't hesitate to reach out to Simon or Chris directly.

Chris is cberg@DataKitchen.io, and Simon, you Simon@kinesis.com? Simon.trewin@kinesis.com. Great. Hope you don't mind me putting that out there. No. And then lastly, if you'd like to learn more about the DataKitchen platform, we actually have another webinar next Wednesday where Chris is going to be walking through a platform demo. So we'll send more information out about how to register for that in the follow-up email with the recording.

So that's it for today. Thanks again for attending, and I hope everyone has a great afternoon and evening.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Questions from this session

What are the six pillars of the Kinaesis DataOps approach?

Target, Instrument, Metadata, Platform, Collaborative Analytics, and Control. Target means the pipeline is driven by a business vision mapped as user journeys. Instrument means profiling, data quality and monitoring at every step, with quality information always presented alongside the actual data. Metadata keeps business definitions connected to the data, the platform is extensible for new demands, analytics is collaborative across IT and data owners, and Control adds version control, release management and exception handling.

What is the S.C.O.P.E. definition process?

S.C.O.P.E. is how Kinaesis defines the target of a data project so stakeholders are not disappointed later. Storyboards map user journeys and visions, Content represents facts, reference data, transactions and process metadata with the right context, Output writes down and agrees the reports, analysis and data to be delivered, Process describes how people interact with the pipeline, and Estimate is a cost-benefit analysis of the change.

Why is DataOps needed in financial services?

Regulatory controls have added layers of complex data processes and the regulators keep getting more sophisticated, faster transactions have made the market intensely competitive, big technology firms are moving on incumbent income streams, and large financial organisations carry a lot of legacy. Meanwhile there is an expectation that machine learning will improve decisions. Data projects in these organisations do not naturally iterate or simplify, which is the gap DataOps closes.

How did DataOps help a bank meet BCBS 239?

A large banking group's BCBS 239 programme had grown too large and was behind schedule with eight months to go before non-compliance. A team of five or six consultants broke the challenge into trackable iterations, delivering key risk indicators incrementally rather than in one release. The result was 300 metrics across seven lines of business, 207 reconciliation breaks identified, billions of pounds of reporting errors fixed, and a compliant operating model.

How does a bank give business teams data without losing control of it?

Through governed self-service sandboxes. A central IT or data group gives a business team a prepared data analytic environment with the data sets and tools it asked for, monitors and governs its use, then takes it back or changes it. A top 5 US bank ran this for more than 1,000 non-IT users, following legal rules on data usage and lifetime and tracking usage throughout.

What is DataOps?

DataOps is the set of technical practices, cultural norms and architecture that enable rapid cycles of experimentation and innovation in delivering new insight, low error rates, collaboration across complex sets of people, technology and environments, and clear measurement and monitoring of results. It covers both the technical environment and the people, process and organization around it.

Where to go next