On-Demand Webinar · 1 hr 2 min

How to Build a Winning Data Team

Jesse Anderson, Managing Director of the Big Data Institute and author of Data Teams, on the foundation an organization needs before it hires data scientists: the three teams a data effort runs on, data engineering, data science, and operations, what each one does, and how DataOps changes the way they fit together. Recorded November 2020; updated August 2026.

Transcript

Show chapters and dialogue 10,311 words

00:00:00

Good afternoon and good morning, everyone. Thanks for joining us today. My name's Beth Beverly, and I am the VP of marketing at DataKitchen, and I will be the host today. Our topic is how to build a winning data team, and we're very excited to have a special guest here, Jesse Anderson, who will join us to share his experience and insights.

For those of you who don't know Jesse, he's a data engineer, creative engineer, and managing director of the Big Data Institute. He works with companies ranging from startups to Fortune 100 companies on big data. He has taught over 30,000 people the skills to become data engineers. He's widely regarded as an expert in the field and for his novel teaching practices. He has published three books on data and teamwork and has been covered in prestigious publications such as The Wall Street Journal, CNN, the BBC, NPR, and Wired. So welcome, Jesse.

Thank you for joining us today. So before I hand it over to Jesse, I just want to cover a few housekeeping items. We are recording this session, so we will send out the video and slides to everyone as soon as they are ready, likely by tomorrow morning. We're also reserving the last 15 minutes of the webinar for questions.

So if you have any questions during the webinar, please enter them in the Q&A box on the webinar control panel, and we'll make sure that Jesse has time to answer those at the end. And then lastly, as mentioned on our registration page, we will be giving away a copy of Jesse's newest book, "Data Teams," to 10 lucky webinar attendees. So we'll choose those winners after the webinar and notify you via email if you've won.

And at that time, we'll coordinate with you the best way to get you a copy of your book. So with that, I will hand it over to Jesse. You can take it away. Thank you, Beth, and thank you, everyone, for joining me. This is a talk called Foundations of Data Teams, and really what I want to share with you and help you understand is what do you need to do? What do you need to do to create that foundation?

And I think this foundation is incredibly important because you're eventually going to want to do DataOps. And then we'll talk about how does a foundation, how does a solid foundation put you in a good position for doing DataOps? But before I start getting deeply into things, I want to get to know who you are.

So I'm going to launch this poll and ask what sort of position you have so that I can kind of speak to whatever position you're at. So good. It's updating in real-time, so I can see we've got a lot of management here, a lot of lead developers, and a few individual contributors. Excellent.

Excellent. So, good, we're talking about things that management, executive management, middle management leads need to be thinking about. And probably why you're here is trying to figure out perhaps why you're having some problems. So let's talk about this. Here's our itinerary. We're going to be talking about houses. We're going to be talking about houses as a metaphor for how do we get the teams right? How do we actually deal with this?

How do we actually work with

our teams correctly? Then we're going to talk about disclosures. Whose job is it to even tell you that there's a problem? And we'll talk about the data teams themselves. What are the data teams? What do they need to do? How do we start having those data teams work in terms of DataOps? And what are some of these implications as we go through and create these foundations?

So to start off with, oftentimes when we work with management, especially and through my company, through my mentoring program, the management is asking for or being not very explicit about what they want. They're being implicit, and what that means is the managers or the CEO says, "I want a data team. I want a data science team." What they don't say or what they don't know what to say is what that implicitly means. Saying, "I want a data science team," or, "I want to start doing machine learning," that implicitly says, "I need these data teams. I need data science. I need data engineering.

I need operations." So this is what we have to deal with as executives or management of this, is actually taking that top level, we want artificial intelligence, machine learning, advanced analytics, and we have to say, "No, this is what this means." It's because oftentimes the teams or the management doesn't know this, that they've focused on that exterior. So we see this image there on the right.

That's exactly what happens. If we look at those houses, we think, "Oh my goodness, I would love to live there. Just imagine the beautiful view that's there." We're focused on this exterior. We're not thinking, "Oh, how are they making that? How are they architecting that so that it stays on the hill and doesn't fall down the hill?" That's not what we're thinking about.

00:05:00

We're thinking about this exterior. We're focused on this cosmetic. But if we're going to start creating data products that are actually usable, we actually have to stop focusing on that exterior and that cosmetic and how beautiful that house is and how beautiful that pink on that one house is, or that tan or what have you. We actually have to start focusing on the foundation and think about how are we going to make sure that that house just doesn't fall off the cliff there.

And so in this metaphor, our house is the value that we're creating. Our job or what we're supposed to be doing as data teams is create value. So how do we create that value? That's what we're going to be talking about. And the way, the means in which we create that, the construction, the actual building of that house, this is our data teams.

These data teams are creating our foundations so that we can then focus on things. So if you've ever watched a movie, you probably have seen what's called a facade, and our focus is often on that facade. Where what will happen is when they shoot a movie, they'll be in front of, in this case, Pete's Joke Shop.

Well, Pete's Joke Shop is nothing more than a facade. As you can see behind it, right behind it, there's nothing but boards holding it up. There is literally no structure. This is only for looks. This is only for what we can see. And oftentimes what happens with management is that we're so focused on the model, I want to do AI, I want to do machine learning, that we forget the rest of the structure, and this is a big problem.

Can we live in this house? Could we live at Pete's Joke Shop? Could we even have a joke shop there? Is this even a viable business? Well, of course not. Once you opened up the door, even maybe the door can open up, maybe not, I don't know. But once we open up the door, there's nothing behind that.

There's literally just scaffolding holding this up. And this is often what happens with teams or with companies who just focus on that machine learning. They focus on the facade. They focus on, this is what I want, rather than, this is the structure, this is the rest of the structure I need. So as we think about this and as we think about what needs to happen, we have a picture here of a yurt.

And if you're not familiar with a yurt, this is a house made for nomadic peoples. Nomadic means that they don't stay in the same place. Their houses are actually built so that they could pick it up and go to a different place and live. But do data teams actually need foundations? Do they actually need a solid foundation, or are they just transient? And if you didn't know, if you don't need a house to stand for a long time, you don't need a foundation.

So the question is, do data teams need to be there for a long time? Do they need a foundation? And the answer is yes. This is actually a common question I get from management, and they say, "Do I just need to have this data team, this data engineering, data science? Do I just need to spin it up for two, three, six months?" The answer is no. The reason behind that is the value that these teams will start create will be so intrinsic, so incredibly important to your business, that you can't do anything without them. You'll want them, you'll need them.

The other reason that these aren't short-term teams is that it takes a long time to do this. It takes a while for teams to start creating value within that company. They'll create some initial value, but oftentimes the highest and best, the most value that they can create comes at a later point in time.

And so in order for you to hit that later point in time, you need to actually have those teams in place. They need to have a solid foundation because they're going to stand for a long time. It's going to be from this point forward. From this point forward, we're going to have data engineering, we're going to have data science. And so yes, we do need a solid foundation.

So then the next question comes up that I often get from management is, this is an easy thing. Can I just press a button, just hire a data scientist, and p**f, magic, artificial intelligence, machine learning goodness comes out? Well, maybe somebody thought that. Here we see a different picture where they thought, "I would love the facade of my building to just look better.

I just want to make that just look a little bit better." And here we can see just looking at how many people are involved in this renovation. There's a ton of people. Just to do a simple renovation of the facade, the outside, takes a lot of people. So I really want you to take away from this session or this webinar, that adding in that advanced analytics, it's not just a simple renovation.

It's not just probably what you heard of. You hire a data scientist, and now it pops this goodness. No, you actually do need some significant effort,

00:10:00

some significant renovation. This is not simple at all. So that gets us to disclosure. Let's talk about disclosure. So I know some of you are from all over the world, and probably some of you are from the United States. Well, there's this concept in the United States and elsewhere in the world of when you sell a house, what you're supposed to do in the U.S.

is to say, "Here's all the problems. Here's all the problems that we have." And what often happens there is that the sellers, the people selling you this, that house, they don't want to tell you everything. And they don't want to tell you everything for various reasons. It may drop the price. So if they say, "Well, there's a problem with the roof." Instead of it being 400,000 or $500,000 for this, or 400, 500,000 euros, it's now worth 200,000. $200,000, 200,000 euros less. So that seller takes this cut, and they lose a bunch of money. So they say, "Well, if I don't disclose that, I can get away with it, and I can just not even have to tell you." But there's this problem where this nondisclosure actually has some legal, some monetary implications for us.

There on the right, that's the disclosure form for selling a house in the state I live in, in Nevada. So we have this issue of disclosure. What does this disclosure mean? It means that vendors or people, whose responsibility is it to tell us when there's a problem? Or what are the implications? So whose job is it? Whose job is it to tell you, "Hey, you're going to spin up a team, and you need this, this, and this"?

So that brings us to a question of, whose job is it? Is it a vendor's job to tell you, "Hey, there's a problem. Hey, although you're thinking you just need a data scientist, you actually need the rest of the foundation"? Is it a management job? Is it an individual contributor's job? Is it books?

Well, let's go through each one, and I'm going to put up a poll to see what you think. Who do you think is responsible for telling you this? For example, a vendor. A vendor, if they come in and they actually tell you the truth, quite frankly, they are going to have a problem. They're going to say, "Well, I need to make that sale. I need to make this growth." And some of these companies are growth at all cost. It is, if I tell you the truth about how difficult this is, you're not going to buy.

So they often don't disclose. In fact, we'll talk about this in a second. They actually do the opposite sometimes. Or let's talk about management. Well, management is a pretty good one. But here's the issue. It's often an unknown unknown, quite frankly. What that means is, although management should know about this, they actually don't know.

They don't know to say, "Hey, this is what it means to start doing advanced analytics, machine learning," for example. "This is what it takes to do DataOps," for example. And this is part of why I personally am writing these books, doing these webinars, is I'm really trying to educate management on this is what it requires. This is actually what it takes to do this.

So then we have individual contributors. Looking at the poll, it's kind of interesting. A lot of you think it's individual contributors. So I'm honestly curious about that because I haven't found many individual contributors to know this, quite honestly. This means that a data scientist or a data engineer on the team, perhaps even an operations person, realizes this and tries to push it up.

I actually talk about that in my "Data Teams" book. And guess what? I've seen this. It's one of the most difficult ways to actually affect change. Not going to lie, those of you who are individual contributors on this call trying to think, "How do I make this change?" It's going to be honestly a difficult change.

You're going to have to push upward. Probably in the easiest sense or analyzing this, it's going to be the easiest for executive management to push this change. Probably second would be middle management because they're able to push to both sides, but it's definitely a difficult thing. So others of you marked, some of you marked other, and this would be books, this would be other sources, and this is kind of what I've done.

However, depending on what type of book you read or whose book you read, most of these or many of these books are rah-rah-rah data science. Just here is the value. Here's a bunch of case studies of data science. They don't say, "Well, actually, it takes this, this and this to do it." They don't really say that. They just focus on, here's this awesome thing, and here's the revenue this company created.

They don't actually tell you and say, "Here is

00:15:00

the difficulty of this," or, "Here's what should be done." Then there's other books like mine where I've tried to actually say, "Hey, this is what you have to do." Or other books that are more data science-focused even just say, "Here's how you establish that data science team." They don't say, "Yes, data science is one part. You need the rest of the parts." So whose job? It's a difficult thing.

So let's talk more about is it easy? So you heard me kind of talk about vendors just briefly. One of the issues, quite frankly, around vendors is that they're actually misleading in their disclosures, quite frankly. What they're saying often is, "It's easy. Don't need the rest of those teams. My product or our product is so easy.

You can just have your unskilled people do this. You don't even need to hire somebody new. You don't need to do anything like that." Well, and sometimes they go so far as to say, "Hey, you don't even need that foundation. You could put this in the hands of your frontline software engineer. You can put this in front of your business intelligence person.

They can do it." Well, no, that's not true. It's actually misleading in their disclosure. And quite honestly, some of these vendors have actually fomented this. They've tried to specifically make it so that you think it's easy, that you don't realize the difficulty of doing this. And why do they do this? Well, they're honestly focused on how do we grow? How do we get you to do this? Instead of telling you, "Hey, you can do this. It's going to take you some time.

It's going to take some effort." And it's because they're honestly worried about losing out on those sales. So don't rely on your vendors, frankly, to go through and tell you this, that they're not the ones disclosing, or sometimes they're doing the opposite of that disclosure. So now let's talk about the data teams themselves. So oftentimes, the question around data teams is what do they have? What is made up of this?

So here we see data science, data engineering, and operations. And which one is the most important team? Well, there actually isn't one that's the most important. Each of these three teams is actually equally important. So here is what each team is. We have data science, we have data engineering, we have operations. Each of these forms that triangle, where in a triangle, one part isn't more important than the others. In a second, we're going to actually talk about what it happens and what it looks like when we're actually missing one of those teams. And it's actually a big deal to be missing one of these teams.

It makes what we create sometimes unusable. So what often happens with management, especially with management who's brand new to this AI machine learning sort of thing, they're thinking about, or I've heard this message of, you just need data science. And then in their minds, data science is the most important, is the only one that's needed. And so they just go through that and they say, "This is all we need." But that's not true.

You do need all three of these teams. Each one of these teams is equally important. So I have a question for you. Do you have all three of these teams? And this is an important question. This is whether you're thinking about doing advanced analytics, doing whatever, doing DataOps, this is all required still. So I see about 50% no, 16% saying maybe. So here, let me share some definitions, and maybe that will help you understand this maybe. And part of this will actually depend on the scale of your data, of which ones you need and what you need, which is actually quite interesting.

Data science, what is data science? Data science is doing advanced analytics with statistics. My short definition of a data scientist is somebody who has taken their mathematical and statistical background and has learned how to program and applies that programming and math to particular problems, to analytical problems, to recommendation engines, that sort of thing.

But it's a person with a math and programming background. Now, data engineering, on the other hand, this is where things may get more difficult for you, and this depends on the scale. If you're doing big data, for example, your data engineering needs to match this definition, in my opinion. That definition is a data engineer is a software engineer, once again, a software engineer, who has specialized their skills in big data, in these big data frameworks.

And if you're doing data at that scale, this is important

00:20:00

because I've seen it as a consistent issue, consistent failure mechanism, where a company will take their data warehouse, their DBA teams, and say, "Oh, you do data? Oh, we're going to do big data now. So there you go, it's under your purview." And this is an issue because that lack of software engineering really inhibits them in their ability to create data products, create the right data products. Choose this. This is a big issue.

Depending on if your scale is lower, you're not at that big data scale, maybe your data engineering is still that DBA, still that data warehouse, perhaps. But as we get to medium data, big data sort of scales, that's where we need that difference in data engineering. And then we have our operations. So operations is an operations engineer who specialized in these frameworks that we need to do.

So if we're doing, for example, big data, we'll need an operations person that knows the big data frameworks we're trying to do. Maybe it's Spark, maybe it's Kafka. Somebody has to operate that. And if we don't meet those definitions, depending on the scale we're trying to hit, then as you can see, maybe some of your yeses are actually nos, and some of your maybes may change to nos as well.

It's really important for you as managers, you as leads, to be able to take this honest look, and I would really encourage you to do that. So let's talk about what happens when we don't have data science. Well, we can see in this picture that there is a building. You could say that this building probably, in some places, is clean, clean enough. People could live there if they really wanted to, and it was cool.

But nobody really wants to live there. And this is kind of what happens with data science. Without data science, there's a foundation, there's people, but we're not creating the highest and best usage value. I would argue if you have data engineering and you have operations and you're doing some big data things, I would argue that if you aren't doing data science, then what's the point? There may be some level of value that you've gotten to your business intelligence people, that sort of thing, but I would argue that your highest and best usage is to actually have that data science team.

They're going to be able to do this analytics. There's advanced analytics, that machine learning, that artificial intelligence. This will create the highest value, and then the business will come live there. The business will come live there and say, "I want this data product. This data product is key. It's crucial." And this is kind of what you see as companies get to that level of adoption within data sciences. If you were to walk up to them and say, "Hey, I'm going to take that model offline," they would say, "Are you crazy?

We use that consistently." And if we don't have that data science team, then we get this, "Okay, well, we don't have it, so what value is it?" We need to achieve that highest and best usage for ourselves. So what does it look like when we don't have data engineers? And this is, for those of you who are attending this, maybe one of the most common reasons for problems. This is one of the most commons I see personally is that companies will have data science and they'll have operations, but they don't have data engineering, or they don't have the right data engineering. This is key as well.

So what happens when we don't have data engineering? You can see in this picture, this is exactly what happens. Somebody did not engineer or architect or build that house right. So probably what happened was it snowed a bunch, it snowed a lot maybe, and the roof just caved in. And this is kind of what it looks like literally when I walk into a company that doesn't have any data engineering.

What's happened is that their house has crumbled. Nobody can live there. Nobody can use it. They lack the engineering. And what happens in these sorts of cases is that the data scientists have been their data engineers. And with all due respect to data scientists, they're smart people They are not good data engineers. They are not good programmers or software engineers. And what happens is that they will create systems that have poor engineering, and once we get to any scale, once we get to any level, once we get to any usage, once we get into production, all of these things, the weight crashes, crunches that roof. And oftentimes we use the term technical debt.

What will happen is that the technical debt created by the data scientist will actually be so heavy, be so much, it will crush the system, and you can't actually do

00:25:00

anything with it. You can't put it into production. It doesn't work very well. It breaks all the time. Another reason for you as management to understand is if you are trying to say, "Hey, let's do this, let's do that," and it takes forever for them to do it, this is often a time where there's a massive amount of technical debt created by a lack of data engineering.

And one thing that I just mentioned before is maybe you have the title data engineer or maybe you have the data engineering team. What you need to do is, leads and management, is actually go back and honestly evaluate that team and say, are they actually meeting that definition of data engineers? This is key.

This is crucial. Because, oftentimes what will happen is people will ostensibly have these data engineers and say, but we have this person. They have the title data engineer, but it crumbles. Why? Well, you have the wrong title. And this will often happen with the data scientists where they say, we don't need data engineers. And what has happened is they don't need the data engineers that you are offering them.

They don't know how to program. They absolutely have to know how to program.

Then we have operations. What happens if we're missing operations? Well, here we look at this picture and we think, oh, it's not the greatest place. Somebody lived there, somebody lived there at some point. Somebody could live there if they really wanted to, but it's so messy and so neglected and so unkempt. Nobody would want to live there. And this is exactly what happens when we don't have operational excellence. The business won't come and live there.

The business says, "Ooh, every time I try to use your data product, it's down." Or, "Every time I try to use your data product, it's so messy and with so many problems, I can't even use it." And you know what happens? And I've seen this firsthand. The operational excellence is so poor that what will happen is the business unit will haul off and try to make their own.

So instead of using your data products, they'll just go do it themselves because they can't trust your stuff. They can't trust your data products. And this is what happens when we lack that operational excellence, and to some extent, often, the data engineering. That the data scientists have been creating something, and it's a data product, but it just can't be used well enough.

This is a common, consistent issue. You have to really watch out for this.

So now that we've kind of talked about this, let me kind of weave in the DataOps side of this. As you can see, I'm going back to a previous slide and saying, well, why isn't there a DataOps team? Why don't I call this out? And so I want to talk briefly about that and say, hey, with DataOps... DataOps, I personally believe, is the highest and best usage of our data teams.

Don't get me wrong there. Having written data teams, I came away with that strong belief that the best way for us to use our data teams, the best value that we'll get from our data teams, is specifically with DataOps. However, I believe that DataOps is a more advanced configuration of the teams that if you're just starting off, that doing DataOps straight away may not be the best way to do this. What I believe is that we have to get our foundation right. We have to get our data science, we have to get our data engineering, we have to get our operations.

And talking to people, and I've had these sorts of conversations of when do we get to this point where we're doing DataOps? Well, it becomes when our technical side, our people side of trying to find people, trying to group people, trying to get them with best practices, is not our biggest problem. What I look for in saying we need to start doing DataOps is when we start having friction be the most inhibiting, most limiting thing that we're doing.

Put a different way is when I talk to management, when I talk to teams, what they shouldn't be saying is, "We don't have enough people. We don't have enough teams. We're trying to start this from scratch," for example. What they're telling me is, "We have all this set up, we just can't get enough done. We can't do enough relative to what we have, the people." That's the issue, friction. And so what we're looking at is we do have the teams, we do have the solid foundations, we do have the best practices, but what's inhibiting my adoption, my data quality, my data products the most, is specifically around friction. And if we start doing DataOps, then we won't eliminate friction completely. We'll gradually reduce it.

It will give us a few other problems, but our biggest problem should no longer be friction. We should have another thing bubble to the

00:30:00

top as our biggest issue. But I want to repeat again, I strongly believe that DataOps is the highest and best usage of our teams. However, it's advanced.

So let's talk about the implications for this. So why would we put so much effort into our foundation? Well, it's always the best way to build on top of a solid foundation. Here you can see in this picture, they're not just putting a little bit of effort into this foundation. They're actually putting probably several million dollars just into the foundation itself. You can look at this, you can see that they have a huge crane, and they have a ton of people.

And if you're very sharp-eyed, you can see all the rebar down there and all of the concrete, and that's just what we can see. They've put a ton of time and effort into this foundation, and it's because they're going to build a tall building on top of this. If you're going to build a really tall building, you need to have a very solid foundation. You need to make sure that that foundation is well-built, because it's actually far more difficult to go back and change this.

So what we need to do is with these construction companies, with the architect companies, they look and they say, "Okay, if you're going to build 100 stories or 50 stories, or however many stories, it's actually going to mirror down that and that foundation. Can we even build on that site? Can we do this?" It has to have that solid foundation.

Well, what happens if we don't have a solid foundation? It's going to cost far, far more to go back and fix it later. So, here's an example. You may or may not be familiar with the Millennium Tower in San Francisco. I think it's a pretty canonical example of a foundation problem. So the tower, you can see the pictures there, and you can see that I've taken some screenshots of headlines just to kind of give you an idea.

And here we have the tower itself was $750 million. Well, it's sinking, and it's sinking because of a foundation problem, and it's actually leaning. It's a big problem. So what does this mean? What does it mean that they didn't do the foundation right? Well, for one, they're being sued by their homeowners for $200 million. And the actual houses, the actual apartments that were put in this were all luxury places, very nice, very high-end.

So there's a lot of rich people with an apartment that they can't sell, and so they're suing the builder. And as you can see there on the bottom, it's going to cost at least $100 million to fix this. So the ROI calculation is not going to be very difficult. It's 200 million plus 100 million, and then there is the cost, and these costs just add up.

So looking at this picture, we can see, hey, this is a big problem. If we don't get our foundation right, it's going to cost far, far more to go back and fix it later. So if you are actually really sharp-eyed, this picture is actually the picture from them building the foundation. You can't see the problem in it. I know I can't.

I'm not an architect in that or an engineer, a structural engineer, so I don't see the problem, but that was their foundation as they were creating it. And then here's all the pictures of it later on. This is a huge issue. So let's kind of say, what does this mean to us? If I'm trying to create a team, what does this mean if I'm not actually putting the effort or able to do this right? What this means is it's going to be really costly and difficult to do this later on.

So if you are starting a new team, you should put the effort into getting your teams right, get it right organizationally, or it's going to be far more difficult to fix later on. What does this mean in the real world? Well, in the teams that I have worked with where they have data science and operations, but no data engineering, this means the technical debt created by the data scientist was not just, "Hey, we can pay this down," or, "We can fix this in a week, two, or even a month." This was year, year plus.

These were year-plus efforts to go back and fix things. At the same time as they're trying to create new things, at the same time as they're trying to maintain it, it is a really difficult situation. It would've been far better had they created that foundation with the three teams right, and they would've been far cheaper because they're not going to have to pay a big team to go back and fix the technical debt created by the data scientists.

00:35:00

It is incredibly important to get this right, get it right initially. It may sound like it's more expensive initially, but in the long run, getting your foundation right is always the cheapest way to go. Here's another reason. Let's say you have an existing team and you're wanting to fix it, or you're dealing with the ramifications of that existing team. This may actually help you understand part of what's happening within your team. So what happens is we have velocity, and this velocity is going to increase with our solid foundation.

There in the graph, we see one where the team is lacking foundation in red and over time and velocity. And then we have blue, we have our solid foundation. Well, as we look at that, we can see that when a team lacks a solid foundation, they're always up and down. They get to a point technically, and then their technical debt just crushes them. They go down to virtually no velocity.

Then they build back up, build back up. Technical debt crushes us. It's this true boom and bust thing that happens with those teams. They're never really able to get good velocity behind themselves. They're never really able to do quite what they need to. However, we can compare that with our solid foundation team. What happens with the solid foundation team is that they're able to build up Get to a plateau, and that plateau allows them, with that solid foundation, to build on top of that.

So as we can see, there's never this point in time where they go down and they have zero velocity. They're always increasing, incrementally increasing that velocity because they start with that solid foundation, build on it, go to a next level of technical complexity. Build on it, next level of technical complexity. So this is key, this is really important.

This is maybe as you look at why, or ask yourself why your team is not working well, it could be this. It could be exactly this problem, that you are not building on the solid foundation, so your team's velocity looks like this rather than gradually and incrementally increasing.

So what I'd like you to take out of this is, depending on where you are in the world, you may use the metric system. There's a saying from Benjamin Franklin that an ounce of prevention is worth a ton of cure. If you were to use the metric system, it's a milligram of prevention is worth thousands of kilograms of cure.

And this is really true. If you do this right initially, it's always going to be the cheapest way. But you can always fix, you can always go back and fix, and I've worked with teams where we've had a considerable amount of issues, considerable amount of technical debt, and I've mentored them through that, and we have gotten through that.

So one thing I really want you to take out of this is that these problems don't fix themselves. They don't magically fix themselves. All of a sudden we get this and the team just has this idea. The light bulb comes on and they think, "Oh, wow, here's how we fix this." That doesn't happen. It has to be specific.

It is going to be costly, but it doesn't fix itself. You will have to put concerted effort into fixing these issues, or they will eventually just really kill you. This, that up and down velocity, it really takes you nowhere, unfortunately. So really do think about this. So I didn't get into the reasons or ways in which you might fix this, but I did write a book about this, sharing my pretty vast experience, and it's called "Data Teams". There is the website, datateams.io.

You can go there. You can learn more about the teams. I would say that there's a few parts that I'm really proud of in this book. One is I have an entire chapter on starting a team. Another one is I have an entire chapter on debugging teams, basically saying, okay, given a situation or given a manifestation, what could be the possible issues?

Other things I'm really proud of is these aren't just my ideas. I had other people contribute, both as contributions within the chapters and interviews. I did several case studies with people who don't just have one year of experience. Some of these people had five, 10 years of experience in managing teams. While that may not sound like a lot of experience, relative to the industry, this is a pretty significant amount of experience, and they share openly their experiences and what happened. So, I would highly recommend if you have these sorts of questions, please do read the book.

And with that, I'd like to thank you, and we will open up for questions. Thanks so much, Jesse. That was great.

00:40:00

I love the analogies of the buildings and the houses. Got some comments on that as well. Seems like others enjoyed that. So now it's time for questions, as Jesse said. So if you have any, please enter them in the box here and we'll try to get through as many as we can in the next 15 minutes or so.

We already have quite a few, Jesse, so I'll just go in order. So where do you start if you have a limited budget or you're just getting started? Is there an order to hire or should you not even go there until you're ready to hire all three? So the suggestion I make in the book, and I have it broken out based on size of organization, so in this case, we're dealing with a startup or to a smaller organization, start with a data engineer first.

The reason for that is the data engineer will start creating that infrastructure that's created in these products. Because if you hire a data scientist initially, what will happen is that they'll sit there and idle. So you're paying somebody to wait on data products to be created for them. So what you really want to do is to have that data engineer create those data products so that once you hire that data scientist, they're productive, they're using them right away.

Okay, great. Thanks. Do you see data engineering or operations replacing the traditional DBA role?

I wouldn't say that it's replacing, I would say it's doing a separate thing. So some organizations will have DBAs. Let's say they already have a data warehouse, a DBA organization. What those bigger organizations will do is they'll spin up the data engineering teams, the data teams as a separate part, and usually that's due to the issues, the limitations of the DBAs' technologies.

Conversely, if you look at the newer startups, or even some of the newer companies that have started in the past 10 years, they don't have data warehouse and DBA organizations at all. They have data engineering teams. So this, depending on where you're coming from, this could be happening. So one word of warning I'll give is if you are an individual contributor who's a DBA, there is a, both in my opinion and others' opinions and from people I've heard in the real world, is that DBAs are not going away.

However, their numbers are dwindling. In other words, the companies don't need a 20-person deep team on their DBA team because they're, let's say, in the cloud, or they need a 10-person, they need a five-person team now. And so those 10 people are left with not being able to find a job or having extreme difficulty finding a job. So my strong advice to people and individual contributors is that you need to start now on learning these new skills.

Okay, great. Thanks. So we have a couple questions on team organization or team leadership. So in your opinion, how important is it that the data engineer, data scientist, and ops are under the same team leadership? I think it's really important. By leadership, they may be talking about middle management leadership. I'm going to answer it in this more of a chief level, CXO sort of thing. So the reason I think it's so important to have the same chief level leadership is that if your organization pushes down mandates or pushes down directions, that being under the same chief level person, you'll have a similar or more similar directive on down. And this is a common issue where, let's say, data science is under some kind of analytics organization, and there's data engineering under IT. Now they have very different mandates, and their two mandates may not actually match. So what will happen there is as the data scientists say, "We need this," and the data engineers say, "Well, that's not our priority.

Sorry." That's a big problem. Now we don't have that symbiotic relationship between the data scientists and data engineers. Unfortunately, too common. What we should have is, by having that same C-level person, maybe that's a chief analytics officer, maybe that's a chief data officer, that we're able to have a clear directive on down so that if the data engineers and data scientists have a difference there, that we have the same C-level executive to make that choice and make that difference, or to basically break that deadlock, as it were.

That makes a lot of sense. Is the data engineer responsible for integrating sources to data lakes and warehouses, too? It may be into data warehouses. Generally, what I've seen is that the data lake infrastructure is usually separate from that. There may be data pipelines that are useful to both sides. Generally, what I've seen is that the two are separate.

00:45:00

And it's usually depending on the company's push or mandate, is some companies have this push of we're going off data warehouses, where the data warehouse is considered legacy, and that the data lake is considered what we're our go forward step or go forward design. And so what is happening there is there may be a T, but oftentimes it's a separate project unto itself.

Okay. And so more on data engineers. What attributes make a good data engineer? Is it just a software engineering mindset and skill, or do other aspects come into play? That's a really good question. I talk about it in the book, and I talk about it in-- I'll mention this other book that I wrote for individual contributors.

It's called "The Ultimate Guide to Switching Careers to Big Data." And I talk about this in there as well. And kind of reaching back to that previous question about the DBAs, I have an entire section in there for DBAs. I would encourage you, if you are a DBA data warehouse person, please read that section. It's going to take you some time to do that.

At least read it, decide if you want to go that route. So speaking specifically to that software engineering question, what makes a good data engineer, it's a mix of actually a few things. It's a love or at least an interest in data as well. So personally, throughout my career, I came up as a software engineer and I always had this interest in data.

I was always doing some kind of analysis. I was curious about it. I was curious about how companies did this. And as I talked to other people who were data engineers who really flourished as data engineers, they had this interest, too. And that stands in contrast to some software engineers, or to many software engineers, of, "I don't really care about data. Data is what gets saved in the database.

I don't really care about it once it gets saved in the database." And this is honestly the difference. You can go in, you can say, "Hey," explain to, let's say, a front-end engineer, some back-end engineers, "Hey, this is the value of data," and they may just not care about it. And that's usually that big metric of, yeah, the data engineer understands that software side, but they also say, "Hey, that data is just as valuable, and here's how we need to lay that out, and here's what we need to do with that." So at its core, that love of data is there, as well as the strong software engineering.

So they do need to program. But then there's that whole big data, this distributed systems part, and that's a whole ball of wax unto itself. Quite frankly, most or many software engineers can't hang at that distributed systems. It is complicated. I wrote a post for O'Reilly talking, and with my thesis being big data is 10 times more complex than small data. And so there's some software engineers who are barely getting by on that complexity, just of the small data side or even organizations.

And so as you throw something in front of them that's 10 times more complex, heads explode. So this is definitely an issue. When I work with companies, we'll actually form a data engineering team often. And what we're trying to do is we'll have to go through a lot of their software engineers and actually ask them these questions, "Are you ready to do this? Can you handle this complexity?" These are honest questions and questions that you would ask of your team if you're a manager of, are you even ready?

Can you handle this complexity increase? Thanks. That's really helpful advice. Now I have a question on data analysts. So in our company, they're responsible for the data visualization based on data sets provided by the data scientists. Basically, they cover the whole standard reporting stuff. So where do they fit in? There is a contribution from Harvinder Atwal on that in "The Data Teams" book, and what he says in there, and I agree with, is that not every

problem needs a hammer named data science , quite frankly. What that means is sometimes data science is our go-to, and we think every problem has to be solved with this, but there's still less difficult problems. For example, some of this data dashboarding, for example, some of these more simple reports, well within the purview of a data analyst. And so it may be a really good use of the time of a data analyst to be able to do that, and the data scientist not to have to do that.

It's kind of that Ferrari versus your minivan. Are you going to drive your Ferrari to the grocery store? Are you going to drive your minivan to the grocery store? These are the questions that management should be asking us.

00:50:00

What is the right person for this job? What is the right tool for this job? And sometimes the right person for that job is actually a data analyst because the data engineering and the data science side have prepared that in such a way that it's easy for the data analyst to do. And that's kind of one of the tools or one of the goals that we'll often have with clients, is I wanted to make the data engineers and data scientists to make things so easy that we can have other parts of the organization do this. Sometimes that's a data analyst, and sometimes that's an actual business person, where we may give them, "Here's some queries," et cetera, and they can do that.

Okay, well, here's some more role-related questions. So trying to wrap my mind around the roles in relation to other terminology, some thoughts of mapping them, and wondering if this is correct. Would data architects be part of the data engineering group? Would enterprise data stewards be part of the operations group, and where would data governance experts fit in all of this?

So that's a little bit to unwind there. Yeah. So some of this is around governance. Depending on the type of company, I'm going to take a wild guess that this person asking this is in finance or in insurance. Those are highly regulated industries where they're having to worry about this. So yes, in those senses, yes, there is a data engineering.

So I'm going to go back to that slide.

So here in data engineering, although I focused on data engineering being mostly made up of data engineers, I didn't say it's solely made up of data engineers. This is an issue where what that means is we predominantly want data engineers in that team, but we are going to have other titles or other people.

Depending on the organization, we may actually have front-end engineers in there. And then there's some of those data governance, DBA, data warehouse sort of people. So one of the things I try to say is, hey, with software engineers, they came up through the world in this particular way, and then you have the data warehouse people coming up from a different way. And in those senses, we have the data warehouse DBAs actually thinking about these issues of data governance, of data lineage, that software engineers, quite frankly, don't think of, and some data engineers don't think of.

So what you want them on that team is to come in and say, "Hey, we're actually violating whatever law or whatever rule that the company has. We need to do this, this, and this." And so what you'll have on that data engineering team are some of those other roles, such as a data governance person. Now coming back to that data architect, this is another interesting question around what does your architect look like?

And there's been no writing on it to my knowledge, except for what I wrote in the book. So I have an entire section of, here is what I believe is the issues with architects, or is what we should consider with data architects or architect. So the issue with data architect is often that they have a very specific

DBA sort of data warehouse-focused skillset. And so as we put them into this new world of perhaps these big data tools, this gives us a new level of difficulty that some of those roles won't actually be able to do. And so what we may actually have is a different person, a different type of architect. I kind of make this argument that the architect that we need is a person who actually does know how to program and has programmed in the past, as well as there's this newer thing where sometimes the data engineers themselves are the architects. There is no specific title of architect on the team, and that they'll kind of create the architecture as a group and then implement that. So I would highly encourage you to read that section for the more detail.

Okay, great. So how will one elevate a foundational data team towards DataOps? What you would do is you would make sure that they are getting the right skills, that we have the right people, and that we're pushing in such a way that our major issue now is no longer do we have data scientists, do we have the right data scientists? Do we have data engineers?

Do we have the right data engineers? Do we have this infrastructure? Do we need to spin up this infrastructure? Do we need a spark cluster? All that fun stuff. That the issue that they're hearing back from the business is, "You're too slow. We've got too much friction." And this is what I'm looking to hear is It's not an issue of we need Spark. It's not an issue of we need Kafka.

00:55:00

We have those already. We have them in production already. We are gaining velocity already. So this isn't a zero velocity thing. What's usually happened is they've topped out on velocity. So here, we look at this, and if we were to expand this out to a DataOps sort of issue, we'll hit this plateau, and there's this level that we can't get past that plateau, where we're only very small increments of value being created.

So what we look for in those situations is, is DataOps the right thing to get us to that next level of productivity? Because we're going to take out that friction, we're going to have that holistic team, that team with all the different groups, all the different members that needs to be there to create that data product, and that should actually be the next level of our velocity increasing.

Great. Thanks. Can you expand on the operations skill set? What key skills fall in this area as it relates to big data? For operations, we're looking for people that actually understand some of the issues of this, some of the issues of scale, some of the issues of not just understanding here's how you do this framework or how do you operate this framework, but they also have to understand both the data side and appreciate that data side, as well as the issues that happen at scale.

There's a series of videos I created for the book, and they're all up on the website, the Data Teams, the IO site, where I talk about how do we interview people for these three positions and what do we look for. So in that sense, when we're doing an interview for an operations person, they need to know at the base operational questions,

your basic Linux sort of things. Then we layer on, do you understand how to operate this framework? And then we layer on top of that: do you understand the problems or the issues that can happen as we put these frameworks into production at the scale? One question or one thought may be there is, okay, if we hire this person from this really small team and we're this huge organization, they may not have hit the issues of scale that we have.

And so we'll have to make sure can they actually hit that? Do they understand data skew and what happens during data skew? At that smaller company with, let's say, a terabyte or two terabytes, they never hit that issue of scale, that issue of what happens when we skew incorrectly. And now, as you go to that much bigger company with that much higher velocity of data, now they're having to think about this and may never have even thought about this.

So this is part of what we need to do to look at for those operations engineers. Okay. Thanks, Jesse. So we're running up against the hour, so I'm going to do one last question. So what should be the skills of a leader of a data team? Should he know about all three parts or some of them?

How should he start to fix the conflicts? It depends on how big the team is, how big the organization is.

If you're just starting out, usually what happens is that same frontline manager is responsible for all three parts, so that you would have that same frontline manager, there would be operations, there'd be data science and data engineering under that same person. As the teams gradually grew bigger, they would start splitting out. If you're at a bigger organization, usually there's a specific manager for data engineering, for data science and operations, all separate.

And that depends on that level. So if maybe they're a director level. So at that director level, we may look to have all the teams together, maybe a VP level. So it really depends on what that leader is looking for. So what I would suggest that leader have is, quite honestly, there aren't a ton of people with this experience already.

So what you'll want to do is to be able to say, "Hey, there aren't a lot of people out this. There are other people I can leverage their experience. You can read the books." However, getting help is probably one of the biggest things that you want to do. Great. Well, thank you so much, Jesse.

I think this presentation really struck a nerve because there are actually a ton of questions we haven't gotten to yet. We could probably talk for another half an hour here. So if we didn't get to your question, we have them all, and we will follow up, or Jesse will follow up with you. Jesse, how can people get in touch with you if they have any additional questions? Is it jesse@bigdatainstitute?

.io. Yes. Okay. Yeah. Or if you can't remember, go to bigdatainstitute.io. There's a contact form, and you can fill it out, and we'd love to help you with your big data journey. Great.

01:00:00

Thank you so much. Thanks for taking the time today to join us. I think everyone really appreciated the insight. I know I found it really useful. So to all the attendees, we'll be sending out the slides and the recording within the next 24 hours, so please be on the lookout for that email. We'll also be notifying those of you who will receive a copy of the book, and we'll figure out the best way to get that to you.

If you didn't win a copy, it sounds like it would be highly recommended to get a copy of that to learn more about setting up the foundations of your data team. If you have any additional questions about DataOps or DataKitchen's platform or services, you can contact me at beth@DataKitchen.io, and I'll route that to the right person.

We also have our own book called the "DataOps Cookbook", which you can download on our website if you want to learn more about DataOps. So again, thanks, Jesse, and thanks to all of you for attending. Have a great afternoon. Thank you, everyone. Bye.

Transcribed automatically from the recording's captions. Names of people, products and companies have been corrected; nothing else is edited. Speakers are not identified: the captions carry no speaker labels, and attributing lines to the presenters would put words in their mouths.

Where to go next