Transcript#
This transcript was generated automatically and may contain errors.
Hey there, welcome to the Paws at Data Science Hangout. I'm Libby Herron, and this is a recording of our weekly community call that happens every Thursday at 12pm US Eastern Time. If you are not joining us live, you miss out on the amazing chat that's going on. So find the link in the description where you can add our call to your calendar and come hang out with the most supportive, friendly, and funny data community you'll ever experience.
I would like to introduce our featured speaker for today, which is Max Kuhn, Principal Software Engineer for tidymodels here at Posit. And you probably know Max Kuhn from a lot of other things, and Talks He's Given, and the Carrot Package. Max, would you like to introduce yourself, tell us a little bit about what you do, and also something you like to do for fun?
Yeah, I'm a statistician. I worked in pharma and infectious disease diagnostics for like a long time before joining Posit. Fun fact, at one point, like the way, one of the ways I got the job was, I think Hadley and our owner JJ were talking, and they happened to be in New York City. And Hadley had had my number from like a previous user. And he texted me and said, hey, do you ever come down to New York City? And I was like, really? I actually had Hamilton tickets that day, and I was like three blocks away from them. So I literally was like, I'll meet you at the corner and so and so. And then we talked about modeling stuff. So that was kind of an interesting experience that eventually led to me joining the company.
Yeah, so I'm a statistician, started at Posit in 2016. So I'm coming up in November, I think on 10 years. And I've mostly been concerned with various types of modeling activities, like a lot of data analysis, and visualization stuff, but modeling and data analysis is really where it's at for me.
Origins of the caret package
And I will say, the first time I ever heard your name, or read your name, I guess, was when I was using the carrot package and learning the carrot package. And I feel like a lot of people their introduction to you as a person was probably carrot.
Yeah. So I think it was like around, so I started, so I worked in drug discovery, like early drug discovery at Pfizer. And I think I started there in 2005. And a couple of years before that, I had been working in infectious disease diagnostics and we had a big, like, it was kind of impressive for the, for the time it was at, but a big, like a clinical trial based project to work on predicting sepsis, which is pretty serious and kind of hard to predict. And so we'd been doing a lot of like machine learning, like wide dataset machine learning. And it just occurred to us that like, we don't have a lot of like basic helpers for things like computing sensitivity and specificity. And so I thought I should make a package out of this.
And then I left there and went to Pfizer. And the thing about most big corporations is like, you start there and then like maybe like 80 help desk tickets later, you actually get access to data. So they like day one, bring me this nice laptop, but I can't, I'm going to meetings, but I can't actually do anything. Cause there's no, there's no data. So I decided to just start this R package, which actually had a different name when I started. And I was going to be using it a lot for my job. So I just started writing it. I was doing a lot of computational chemistry stuff then. And yeah, it just sort of evolved. And then when it's like a year or two later, when I finally sent it to CRAN, I think like not long before somebody had used the same package name. So I had to come up with something sort of like at the last minute and landed on Carrot. So yeah, so that's how that came to be. It was maybe like, I think maybe the first CRAN version was maybe like 2007 ish, I think.
Yeah, that was pretty far before. And then, you know, once I started working on tidymodels, there wasn't honestly a lot of time to work on it. So it's sort of like it didn't flounder, but I just wasn't really doing much development for it. Which I think frustrated a lot, frustrates or frustrated a lot of people. But funny enough this year with our internship program, we had sort of like a list of things people could work on in Charlotte, our current intern decided she wanted to work on Carrot. So if you go there and look at like the contribution graph, there's like this big, like no man's land where like nothing really happened. And then it's like a huge amount of contributions now. So we're just like cleaning it up.
Yeah, it's hard to believe this package was written before or had like named spaces before Knitter, like way before like Knitter and things that we just like, it's hard to believe that that was still the case, but so we need a lot of modernization. So we're just basically bringing it up to like where things are at today and just improve it in terms of just keeping it sustainable.
From caret to tidymodels
Well, part of the goal of just keeping it sustainable is because there is a new wear option, right? Cause we now have tidymodels. Do you want to talk a little bit about how tidymodels came about and why, why new tidymodels instead of just updating or modernizing Carrot?
Yeah. So, so Carrot, okay. So like, you know, I learned some, I'm a statistician, right. And we're not known for our like exquisite software engineering skills. And I had learned a fair amount in graduate school. And then that first job I had in diagnostics learned a lot there, but Carrot was something that I sort of did a bit in my spare time. And so it did not, it was not like something where I thought out like, oh, we'll do, here's all the things that are going to be on the menu. It sort of grew and grew and grew. And I didn't expect anybody to use it, but then they started using it. And then I was like, sure, let's add this. And it just like became kind of a monster to maintain because for both me and for Crayon, to be honest with you.
So it has a very traditional like base or type interface and people, you know, a lot of people like that. So I didn't want to like tidy that and just make people use that. And re-engineering it, like for survival analysis, like an example where it was very, very difficult to include survival models in that package, just because of how it was set up. And, you know, with tidymodels, it's more tidy interface. It's we, for better, for worse, make a lot more modular packages. So the idea was to have like a parallel implementation, somewhat kind of driven by like, well, if you like tidy-verse stuff, then you're going to look at Carrot and be like, oof. But then if you're more of a base R person that you would look at tidymodels and not like that. So it's just, it's, it's more enabling in tidymodels to make new stuff.
And I agree. It's so baiting me. I'm like, don't, no, Max, don't, don't rant.
Thank you for the kind words though. And honestly, just about that, like, you know, we think about that a lot. Like the book I'm writing now, I'm writing like both like R and Python companions to it. And I have, I've stopped and started on the Python companion a couple of times because I feel like there's like a lot of basic stuff that's just like super hard to do in Python. And so, you know, Emil and our group who may or may not be here, we've been talking about like a recipes type version for Python. And so like we're, I think we are going to be chipping away at things. At least we find friction on the scikit-learn side. So that's exciting. There's a lot of good things about scikit-learn. So I'm not like down with it or anything like that. But as somebody who did like modeling and machine learning for a living, I feel like you, you have a very different perspective than people who develop like software things who like haven't lived through that.
Writing the new book on tabular data
Yes. Let me get the link for it. So we were going to write a second edition to apply predictive modeling. And it's mostly, I think, logistical reasons why it didn't happen. We haven't had a great relationship with our publisher from the jump. And we kind of got waylaid for about two years about a second edition. And I don't know, Kel doesn't do this as much as, my coauthor Kel doesn't do this as much as I do, but I'll look into something and I'll write a little bit about it and just put a QMD file or whatever somewhere and just like squirrel it away. So kind of had like a lot of things sort of like geared up to do that.
So instead of the second edition, with neural networks becoming so prevalent now, I was hesitant to maybe use this nomenclature, but like tabular data modeling, which is a really weird way of saying like the things we would normally do anyway, is pretty common nomenclature. So it's applied machine learning for tabular data. And it's basically like what we would have done for a second edition. The problem, I mean, the good and bad of it is we're writing it out in the open. So as soon as we write it, it's there.
I don't know if this is bad news, but it's forward thinking it might be an issue where there's really no like boundaries on like page count with a website. So I've got like a bunch of PRs to put in because I finally finished locally like my neural network with like the foundational models and like modern neural networks and stuff like that. And that's probably gonna be like if it doesn't get edited down, we eventually make it a physical copy. That's gonna be like 50 pages at least.
And I don't know, I feel pretty good about it. Quarto is really great. We're using ShinyLive for some things. And I don't know, that's been a very, I really enjoy having, like I'm not just gonna keep writing books in perpetuity, but I enjoy having like something that like forces me to do research and read and stuff like that and look at what's happening.
Deep learning for tabular data
What do you think about the recent improvements in deep learning for tabular data, such as TabPFN, TabICL, TabFM, et cetera, compared to like GBM or XGBoost?
Yeah. They're pretty, they're pretty impressive. You know, like the whole like dominance of boosting is always just bounced with the kind of bothered me because you can do just as well without, like, if you put a little bit of work into it, like you can get like a smaller model usually that does about as well. And then like these tabular models, this we're taking it to the next level. So just to give you a sense of like what this, these are is, I don't know, some, some point in time ago, like in the last, like 10 or 12 years some deep learning people figured out that basically you can approximate like Bayesian inference using like a sufficiently sophisticated neural network that has a few architectural like particulars to it. And so the idea is and so TabPFN was probably the first one of those.
And so what they do is they have this massive neural network that they train, but they don't train on any real data. They have a means to simulate like every possible sort of like tabular data scenario with different numbers of rows or columns or correlations, like things having smooth like connections to the outcome or things that look more like trees and things like that. So they can like really efficiently simulate. I think it's like the last training point, I suppose, like maybe like 250 million datasets. And so they, they train these really large neural networks on these synthetic datasets at once, and you basically get the finished product. And so when you put your data in much like LLMs do these days, what these models do is they take your, you don't really have training data because there's no real estimation that happens, but you give it your training set data and it's sort of like your prompt in an LLM. It gives it context so you can, so it can do better. It has this huge network that could do almost anything. And as soon as you give it your training data, it basically lets it funnel down to concentrate on the parts of the network that are most relevant. And so you basically get an immediate answer without any actual training.
And they do very, very well. I've yet to see them really falter anywhere. So I've just finished like writing like a boatload about this, which I haven't put on the website yet, but just to be honest with you, one thing, and I think I've talked to Frank about this specifically, but in talks and things like that, the thing that confounds me a little bit about this is they perform well in ways that I can't explain. So if you, I have a huge simulation system where I run on a bunch of models where like I give it like one or 10 or 50 just noise columns and a regular neural network just dies when you do it. I mean, this performance is tanks, but and like a tree would not, or GlimNet or something that has like feature selection built into it does not have a problem with that. And I can't explain why these very large neural networks, these TabPFN type models just blow through that successfully. They're incredibly resistant to things like that. And I literally spent like two days both by myself and Claude just trying to figure out how this works and have yet to really give a, like a good answer that I would think like, oh yeah, that's it.
So I've just finished like writing like a boatload about this, which I haven't put on the website yet, but just to be honest with you, one thing, and I think I've talked to Frank about this specifically, but in talks and things like that, the thing that confounds me a little bit about this is they perform well in ways that I can't explain.
So, so yeah, so they don't end up, Morgana, yeah, they don't really end up overfitting because there's no real fitting. As long as you have like, in some other experiments I've run, as long as you have like, but depending on the problem, maybe like 100 or 200 or 500 data points, the model is able to really focus in on like the weights and the neural network that are relevant for your dataset. So I've, I've yet to see them either under or overfit.
And I should also say like of the three that Frank mentioned, so TabPFN is like the original and some of that's, so it's a, you know, it's open source in the sense that we have a Python and we built our interface to the Python package. So we can, we can run it, but they're a company and they have licenses and things like that. And I think they were maybe not bought necessarily, but sufficiently funded to kind of be like bought by SAP. So I think where these models are going is they're just going to get built into databases probably where you have some table of data and it can automatically like just take that data and build like a pretty fast and decently performant model to just fill in the blanks for any outcomes you have.
So TabPFN was sort of the original one, that's a proprietary model. TabICL, ICL stands for in-context learning. That's a, I believe like a completely open source model that has no license associated with it. And then the new one, which is also somewhat proprietary by Google's called TabFM, fm foundational model. And so while we just released a version of Brulé package that has like local torch model for TabICL. So if you don't want to have it like reticulate your way through things and want like a pure R, which to me is like the preferred way of doing it, we basically emulated all the Python code for TabICL. So that's all available. So you start to run it and you have to download like, I don't know, I think it's like a hundred megabytes or so of model files that are the pre-built models and then it works.
Well, and that brings me to TabFM, which is the Google one that just came out, which we're also looking at. It's model files all told, I think are like six gigabytes. So Google is sort of like, well, it's, I don't know, honestly, I don't want to speak bad about it, but it's like, I was just benchmarking on my laptop and it's, I don't know, I feel like it's not built for anything, but Google's hardware. In the sense that just me loading the model is like, if I'm just doing regression, it's like three gigabytes in memory right there.
And all of these models are required to GPU. I mean, you could, for predicting a few things, use your CPU, but the memory just spikes. So it's a little bit unusual that you need actual like GPU hardware to run these things at all.
Day-to-day life as a software engineer
What does a normal day or week look like for you, Max?
Yeah, I'd say the only downside to my job right now is I almost never get new data, right? So I'm used to having like the treasure trove of really cool, like cutting edge new data, but like I don't get any new data. And my day to day is really kind of, you know, I've read interviews with UA and Jenny Bryan, and I think all of us have a very similar outlook on things is, I feel like it's healthy, but it may not sound like it in a way. But like a lot of us have kids. I have two older sons, but two younger stepdaughters. And like, we're kind of free to work when we want. So like Emil is on the West Coast in my group, I'm in Connecticut. So like East Coast of the US and then Hannah's in London. So we have this like Goldilocks zone between like 11 and one where we can actually in real time, like sort of like talk to each other.
And so, you know, there are days where like I'll get up and do kids stuff or whatever I need to do, and then don't get to work till like 11 and then do a bunch of stuff and then kids go home at three. So then, you know, pause, and then maybe do work later that night just to get everything done. So maybe it's not great to be like working at night. But I don't know, it suits me in a lot of ways. Because I can work whenever I want.
Computational chemistry background
What type of computational chemistry did you work on?
Yeah, so like a big pharma company, you know, they're, you have medicinal chemists who know what thing in your body they want to modify, like, make more of this protein or make less of this protein or do X, Y, and Z. And, and so most, most molecules you can write down as an equation, something like called a smile string, which is like connection of like carbons and oxygens and bonds and stuff like that. And so the way it works is, you know, you're working on something like I worked on GLP drugs, GLP-1 drugs for like a long time. And this is like, you know, maybe like 15 years ago. And so, you know, they'll come up with an idea for a compound that might be effective. And sometimes they send that off to synthesis, meaning somebody figures out how to make that, and then they can test it on cells and things like that. But before they do that, they kind of want to get a first pass at, oh, is this going to kill somebody? Is this going to work? Is it soluble? Like, can we make it in solution?
And so we have a ton of these, these chemical formulas in existing laboratory tests. And one thing in computational chemistry is called quantitative structure activity relationship models, QSAR. And that's where you basically get this like string of, of connections between molecules. You can then, oh, I don't do this. We have a lot of software to do this, to then turn those into predictors in a model. And then you, you may have like a dozen, or you may have like a million existing data points. And so it's constantly making models to make predictions on like dozens of things like drug-drug interactions, and so on, so that somebody can type in a molecule they're thinking about making, and they get an initial like dashboard of like, oh, here's what's good about it. Here's what's bad about it. Here's things that are similar to it that we already have, and so on.
So it was basically either, I didn't actually, like I was never on the hook for that many models, but around that time, we were building a whole like framework. It was kind of like pre like Shiny based on Java, where basically you can go in and, you know, point towards a data table in a, in a database somewhere, or upload some, some strings and say, predict these. And then you had a whole list of like, you know, random forest or, or whatnot. And so I, that's honestly why Carrot was built, because I knew we were going to be doing that in having a, like a, a menu of things to make predictions with.
So it's, it's great because the, the people you work with are incredibly engaged. You usually have a ton of data. Once you get the hang of those types of models, they're not all the same, but you pretty much know where to focus and what not to focus on. The downside is kind of interesting because like, you know, I love working with medicinal chemists, but they're really, they're really paternal and maternal. They will have this like molecule that they think is going to be the one. And then they run it through your models and you give it like a, like a, like a risk of having, let's say, like liver injury of being like 10 or 20% is like, no, that's right. And then now they're like, you know, if you give them then like no risk at all they wouldn't question it. But now they're in your office being like, I'm going to submit this every week and see how much it changes.
And there were also times where we had to be sort of like the ambassadors because we had the biologists and the lab people who are running these, these assays, these laboratory tests to generate the results. And then the consumers of the models saying like, well, I don't think the noise should be as high on this, or why is it so high? And then having to mediate between them to say like, yeah, this is as good as it gets, or this could be improved, or just being the person to like, not, it wasn't, it wasn't like any sort of like toxic environment, but basically like a little bit of deescalating people to just set expectations about what models can do and so on. So much of statistics is communicating things to people and being an intercessor.
Inference versus prediction
Yeah, that's a really good question. Yes, I tend to think of, I'm trying to, I always try to find analogies to things and maybe I don't need one. But like, if you think about like linear or logistic regression or something, it could function in many, many different levels, right? And so, you know, being my degrees in biostats and like two thirds, if maybe three quarters of my department was really concerned about like inference. And then the other 25%, let's say, are concerned about like estimating stuff. And I definitely am more interested in the 25% of estimation. To me, it's like a more fun problem.
Yeah, I don't know. I find that like, when I think about these things, I either, I start and either just go in left or right in the direction of either inference or prediction. And from there, like the only time I think they sort of like, there's a bridge that goes between them is, and I promise you, I wasn't really obnoxious about this, but I would see a lot of presentations at where I worked, where people were presenting some logistic regression and they, let's say, you know, for something fairly important, they're looking at, you know, model selection. So they're just looking at P-values and what's the right combination of things to, for whatever definition they had of optimal. And, you know, my question was like, what's the accuracy of this model? And they'd be like, you know, like what? And I'm like, well, is it like 50% accurate or like 80% accurate? And like, well, we're not trying to predict. I'm like, yeah, but like, if it doesn't have much fidelity to the data, like we kind of have to calibrate that to know whether these P-values are important. Because if the model is not even good at even a close ballpark prediction, then like, you know, how, I'm not saying like the P-values would be invalid, but at the same time, like you need to know how well does your model emulate what you're modeling?
And so just the idea that I just, like the only metric is P-value based always kind of bothered me. I'm not trying to make everything a prediction problem, but like, it's the same thing where the other direction where people like, you'll choose like, oh, I should use this many splits in a tree versus those many splits. And somebody would say like, well, they're not really statistically different. Like, well, yeah, they're not. So I don't know. I think there is a situation where those things overlap, but more than like in a diagnostic sense.
Oh, I am 100% down that. Yes. I, when I was in drug discovery, our director said, and a lot of people like revolted about this, like every analysis has to be Bayesian and I get where he was coming from. And, but any sort of inferential stuff I do now, it's Bayes all the way down for me. Like I am completely down.
Career advice
I don't know that I'm even qualified to answer this because my career started in like 1998. So like I'm not a young man. So like, I feel like everything's different now. But I do talk to, I have a lot of friends whose kids are going to graduate school and a fair amount of them are interested in statistics. And, and it surprises me sometimes because I feel like, I feel like we're viewed as these rare and arcane wizards that like, are the only people who know how to kill the dragon. But everybody's afraid to approach us because they're like, don't smite me. Like just by asking a question.
I feel like this might be obvious, but the more people I talk to, maybe it's not. Because we're like, often deep in especially open source software and things like that. But like, talking to a friend's kid, and I was basically saying like, put as much out there on GitHub of yourself as you can. Even even if you know, having something concrete, you can point to that you did, and there's like commits, and things like that. Even if it's not groundbreaking, and it's like, you're just analyzing the Ames housing data or whatnot, like have something that kind of acts as your portfolio. And, and for me, it's like, having both been a hiring manager of statisticians and data scientists and somebody who, you know, almost every year is trying to hire interns, the thing, you know, we get a lot of CVs that just say, Oh, I did this. And I did that. And it's like a few sentences.
And I know some of it, like you can't talk about, especially if you did work company, but you know, I would absolutely try to convince people to do something. It doesn't have the bar is really low, but have something of decent quality. Like, you know, it's not like RM, LS, list equals type code, but like, you know, like a shiny app or something that you can just show people and point to that you did. And you can look at my code and, you know, look that I know how to make commits or whatnot. I feel like that is the number one thing you can do. And honestly, when we, when we go to hire people, even pre AI, that would will it down to like 10% of the people who apply that I can actually see that you've you know, something you've touched code. It's not like we don't believe people. It's just like seeing it is incredibly helpful and trying to figure things out.
And honestly, when we, when we go to hire people, even pre AI, that would will it down to like 10% of the people who apply that I can actually see that you've you know, something you've touched code. It's not like we don't believe people. It's just like seeing it is incredibly helpful and trying to figure things out.
Now, the problem we have now is everybody just like, like we have, I shouldn't say this, but like, you know, when we, we have people apply to internship jobs, for example, like we have to have their GitHub user ID, because they're going to need a GitHub ID. So we asked for it. And if they do it, like, you know, and they have a bunch of tidymodels stuff, I'll look to see if it was like, you know, started the week after the position was like announced, which isn't bad, because they're showing interest. But like, like, it's really easy for recruiters and things like that, that could.
It's really easy for anybody to just, you know, parse this data and look at it. So it's really important to just have something out there showing your interest in something, showing that like, you know how to code or do stuff, even if it's like, even like stat or something like that, just put something out in public that somebody can see it. And do as much of that as you can.
Round of applause. Absolutely. I think the most important thing that Max just said is that out of all the people who apply, it can whittle it down to 10%. That means that if you do the thing, and you talk about it, write about it, put it on the internet, you are automatically putting yourself in the top 10% of candidates for something. Think about that. There's 1000s of people doing the thing. If you're the one who talks about the thing, doesn't matter if you're the brightest, sharpest tool in the shed, best at it, whatever, you are automatically giving yourself a leg up. I highly encourage you to do that. Make a Quarto website and put some freaking Quarto documents up there with your analyses. And it doesn't even matter if they're right, but you can write about what you think about them, what you learned about them.
I love this. This was so much fun. My only regret is that we did not veer into Cosmere territory, Max. But we are at the top of the hour. If you want to talk about the Cosmere with me or Max, find us on Blue Sky and talk to us about it. Alright everybody. This was fantastic. Welcome back. After two weeks being off, this was so fun. I am so glad that I get to spend this time with you every week. And we will see you next week, where we are joined by our very own Hubert Hickman. And we get to learn a lot more about him. You have heard him ask questions for months and months on The Hangout. And he's so much fun, so we can't wait to talk with him. Goodbye, everybody.


