Transcript#
This transcript was generated automatically and may contain errors.
Hey there, welcome to the Paws at Data Science Hangout. I'm Libby Herron, and this is a recording of our weekly community call that happens every Thursday at 12 p.m. U.S. Eastern Time. If you are not joining us live, you miss out on the amazing chat that's going on. So find the link in the description where you can add our call to your calendar and come hang out with the most supportive, friendly, and funny data community you'll ever experience. Can't wait to see you there.
In celebration of five years of the Hangout, I am so excited to introduce you to our featured leader today. We are joined by Hubert Hickman, Senior Data Strategist and Data Operations Lead at the University of Chicago's Data for the Common Good, which is just such a cool phrase, Data for the Common Good. I love it so much. Hubert, welcome, but not really welcome because you've been joining the Hangout for so long, and a lot of people know your face and your name. I would love it if you could introduce yourself, tell us a little bit about what you do, and something that you like to do for fun.
Thank you, Libby, and all the Paws at folks for this Hangout. It's a great place. I am Hubert Hickman. As Libby said, I'm the Senior Data Strategist and Data Operations Lead for Data for the Common Good, also known as D4CG at the University of Chicago. We are the home of the Pediatric Cancer Data Commons, among other initiatives, but the PCDC, Pediatric Cancer Data Commons, is a consortia of consortia, 16 different pediatric cancer types, all rolled into an umbrella, but each with their own data and governance structures, but unified into the PCDC. The data is not ours. We hold the data. It is governed by the consortia themselves. Our data dictionaries are open. You can go and look at them. And so there's a lot of work that goes in within our group to putting those data dictionaries together, getting the data in, wizardry on governance and agreements and data use agreements and that sort of stuff. There's also other initiatives like penetrating brain injuries, rare epilepsies, and other things that are happening.
I, by training, I am a very mathy computer scientist. I am one of those folks who did not complete a PhD. I think there's a number of us here on the Hangout that are in that state. My early career was industrial wastewater treatment, college administration, defense and power engineering. And since 1993, I've been working in clinical and research and dramatics. I got into this area because of a friendship with a liver transplant surgeon here in Omaha. Eventually, one thing led to another. Because of that friendship, I co-founded a company with addition of a grad school buddy in my basement dealing with caring for and transplant patients and helping centers deal with all the logistics and clinical things that were needed. That company started in my basement, stayed there for two years. 17 years later, we had 65 transplant sites across the US, Canada, and Australia. And I think about a third of the transplant patients were in our systems in the US.
After the company sold, in the meantime, in 2003, I worked as an employee at UNMC working on the NICU system, the new things that led to a successful system. In 2012, the EPIC came in, our team was no longer needed for some of the charting for the physician documentation system at UNMC. And so I sort of rattled around a little bit after the company sold with that, got into the research informatics side from a call from another doctor, a different doctor at UNMC that wanted help setting up a coordinate node at UNMC, Patient-Centered Outcomes Research Network, one of the national research networks. And so that got me into the research side, ITV2, the research data warehouse. That led in 2015 to another company finding me that did the same sort of work at NYU Langone, ITV2 primarily in research data. That led, another step led to some initial work with the National Cancer Institute in 2017. That project was called DI3, integrating imagery, image data, multimodal image data, along with research study data. That's the first time I'd actually heard of the Chicago group as I knew it as then because they were also one of the other groups working on this project.
That project took ITV2, research study data, integrated that, that's also where I ran into R because I needed a thing that could write SAS export files for SDTM, study data tabulation model files, and I need a little web server to write extensions for what, for ITV2. And I found this thing called R and Shiny and then, hmm, this will fill the bill, and it did. That led to more work with the National Cancer Institute within the clinical trial reporting program, dealing with structuring of unstructured clinical trial inclusion and exclusion criteria and trial search space. That was my primary focus for the time at NCI. And in early 2024, I moved to the D4CG group. It is a, you know, it is a wonderful group of folks. We're doing really great work in areas where that work is really needed. I think I've written code in 16 different programming languages over the years. I counted them up, and I keep thinking good ones every now and then.
Coming to data science sideways
Well, there's a couple of things I heard in there that I would love to, like, pick back up and just ask you some more questions about. One is you said you didn't complete a PhD, which I think that a lot of people can relate to, and that you came from a more software engineer mathy background. So you've come to data science from, like, the side, right? Not from, straight up from statistics. And the other one is that you did not have a background in cancer research. You got here because you made a friend, right? You sort of found somebody who pulled you in to help you solve a problem. And you said yes to, sure, I'll help you solve this problem that I know nothing about. So could you tell me a little bit about how you met that friend? Did you join a community like this one? What did that look like?
We, in 1990, bought a computer called the Next Workstation. If you remember what Steve Jobs did between getting fired from Apple and getting rehired to Apple, he did Next. And it was a wonderful system, especially compared to the state-of-the-art of today. It had lots of power, a wonderful GUI, networking. It was Unix underneath. So it was, for me, it was like, oh, that's my dream computer. And this transplant surgeon, unbeknownst to me, had bought one of these machines too. And I had set up a little user group, and I think there were like five of us in Omaha. It was not a big community. And that's how I met him. And they had had IBM Enterprise Alliance come in to make a smoke and mirrors demo system for them. So they had had two different unsuccessful runs at building a system to meet their needs. And when we said, yes, oh, yeah, we can do this, we were very naive. We said, oh, three months. It took us two years to build that system out.
So there was a lot of learning and deep learning. There was no clinical background in my history, but you have to learn and listen and watch how people work, what the needs were, and juggle all the other things to make a go of it. And so, yeah, it was a challenge. But fortunately, my friend was a very patient man.
Building data solutions close to users
So you've been building data solutions for many, many years. I'm wondering what advice you would give somebody who is coming in new to the game and is being tasked with building a data solution for somebody? Because my advice is always go sit with them, shadow them, watch them work for a really long time and ask them a bunch of questions. How long did you spend just watching the way that people were interacting with their existing solution for those patients?
Oh, a very long time. I mean, we did interviews upon interviews, watched them work. I made cassette tapes back in the day of all these so I could replay them because I didn't know. It's like I give the analogy of Homer's dog and the Simpsons spot, you know, spot doesn't understand any word other than his name. You know, so it's like blah, blah, blah, spot, blah, blah, blah, spot. So that's the way I felt at first. So you have to be able, I think, to be successful in some of these arenas to where there's sort of this deep domain knowledge needed. You have to be able to learn that domain well enough and step back and make a solution, you know, see what the real problems are. Because that's it's not necessarily the problems they're telling you what's in front of them, but what are the patterns? Where are they inefficient? What do they need to see? And that's whether it's clinical or research, you know, some of those same questions apply. You've got to learn and be willing to learn enough of the domain to know what a system might offer in terms of solving real problems for people, be it research or clinical workflows.
So you have to be able, I think, to be successful in some of these arenas to where there's sort of this deep domain knowledge needed. You have to be able to learn that domain well enough and step back and make a solution, you know, see what the real problems are.
AI as an accelerant for prototyping
I'm wondering if you have seen AI-driven development sort of shift that building bespoke stuff game a little bit.
It has. And that's, you know, AI is an accelerant for good or for evil or for bad things. You know, you can use it in a good way and a bad way. And the bad way, the good way, excuse me, the good way is just what Abigail mentioned, that it makes putting a prototype together really easy. One of the things I'm working on now in D4CG is, you know, we're working on direct from patient data via FHIR and research data. And, you know, some of that data is ugly. And, you know, the EHR data is all over the place and things are here, things are there. It's not a straightforward or easy issue to solve. But once you get the data part of this figured out, you know, what the data should look like, what's the shape, what's the structure of the data, and I still think you need to have a model of that in your head, at least, of what this data, what the data are all about. You know, you can literally, I mean, what I did for one of my, one of my current projects is the first version of this. I just drew pictures for Claude and said, here's the data, here's what I want you to do. And it did a really bad first version of that. Yeah, it wasn't very good at all. But it put it in front of my face to see it. And then you can realize, oh, this is not so good. But then you can go and start cycling and iterating that to make it better. And then you put it in front of people now, oh, wow, this is nice. And so, that thing that would have taken me months and months and months and months and months to write, if I had to do it myself, Claude's perfectly happy doing crazy JavaScript stuff and wiring in shiny reactives and whatnot to make things work. And it's perfectly, you know, it does a good job. Now, you still have to ultimately go and write a lot of LLM code and sort of de-slop it once you're all done with it.
Favorite and least favorite programming languages
Abigail's question was, what was your favorite programming language and your least favorite programming languages and why? And are there any that you wish you could still work in that maybe you don't get to work in anymore? Oh, that's a good question. Least favorite? I would... I'm going to say I do not like working in Java. I've worked in it, but it's just not my favorite. It seems wordy, and I know lots of people love it, and it's efficient. They work in it day on, day out, but it's just not my cup of tea. I'm not a big Java fan. Favorites depends upon the time you ask. I've had many favorites over the years, but mainly those which allow me to get my work done without getting in the way, and I like Python. Both Python and R for that. I do most of the GUI work I do in R. It just makes... I know how Shiny works. The way reactives and interactivity work is straightforward enough that you can get what you need done without requiring lots and lots and lots of work. In Python, I first ran across that in 2001 or so, back in 1.5 Python days, and I could just almost think Python when you started writing Python. And when we did the system in 2003 at UNMC, that was written in Python, and that was not a popular choice at the time. Now Python runs most of the world, but yeah, those are my answers.
Some will never die. Perl, I can write fine, but you can't read it later.
R vs Python for data science tasks
Yeah, I think for the data science-y part, I think R. You can do it in R and Python. A lot of people doing Python. My own preferences are just because it's a little more concise and data frames kind of work consistently within the ecosystem. But it's the conciseness of it all, and plus just the wide range of things that are available. Whatever you need, there's a library for that probably in data science. And I think Python may have caught up a lot of that way, but my personal use of Python is sort of the backend use case, data processing, that versus, okay, let's do presentation. I use Python to manipulate the data coming out to get it into the structure I needed. But yeah, pivot table, all this other stuff, you can see heat maps, graphs, this, that, all the other things. That was not much code to get a lot of functionality out of things. But I would say R. And I think in terms of what people are learning these days, folks coming out of college, I think R is more and more popular for the data science community as well. Because it's open, you can go get it, like SAS or some of the other things that cost money, or you get some version that's not really for commercial use or something.
A typical week
Yeah, I think a lot of what I do on a day-to-day basis, within an average week, I rotate. I work with the data standards group, developers group, I work in their research group. So there's many different groups that are very internal, that I sort of float between and in. And so there's what I would call standard sort of, I need to get this code written to do these things. And I've got features that need to go into an app I'm working on, or there's some backend thing, or I need to update one of the things that's one of the apps that I support. The other part of it is being in what were the kinds of things we should be doing at a higher level in the group, just participating in that, throwing out research ideas and everything, and thinking about what things should come, how, what are out there in terms of opportunities for us to do things, throwing ideas out on the table. And not all of Hubert's ideas are great, but I'm going to throw them out on the table to say, well, what about this idea or that idea, and pull things together. So it's a mix of things. I'm not a always software developer, or data, always data person, or always other person. So it's a mix of different things that I do during the course of a week. And the days do, you know, no days alike in my week because of meetings and other things.
Differences between commercial and public good development
So thanks, Hubert. My question was kind of the, do you, what differences would you think of or see between developing tools for like a startup or a for profit organization versus developing things as you're doing for public good? Does that change how you think about, you know, who the customer is or what the focus of the product is?
It does. You know, when you're working in some sort of academic medical area or academic research area, the focus is not really a product focus. So it's more like, you know, how do we get this research data? There's plenty of things to go tackle and pulling research data together. How do we present it? Not just the sort of statistical analysis to where my statistics are rusty, very rusty after all these years. But there's that aspect of it, but it's still, you know, you do have customers. They're just not like outside paying, you know, paying customers. So you still have to pay attention to what the needs are and where you may make it better. If you're in a small commercial organization, then, then it's a different game in terms of what do we need to do to bring in revenue? What do we have to do? But a small organization is also different than a large organization. And it's, you know, large organizations, some of them are, you know, there's lots of consultancies and this and that, where you design a project and then you're part of the big team and you're fully agile of all the pageantry. And so sometimes you as an individual don't have much leeway in how things get done. I think some of the larger development efforts, you lose the plot because of things like agile. And, you know, you have so many people in there, they're not very near the customers, not very near, whether it be researchers or clinicians or whomever it may be, they're not very close to that customer. And so the further away you are from the people you're building software tools for, whether it's like, you know, a tool like the PCDC, we're very close to those, to those folks who are contributing the data and managing the data and, and handing data to us. So the more distance you get between you and the people who lay hands on your software, then things, I think, tend to unravel.
So the more distance you get between you and the people who lay hands on your software, then things, I think, tend to unravel.
I've been physically threatened by a transplant surgeon one time if the software didn't work right. So, you know, sometimes it's a little bit too close to some of those customers, but, you know, there is being bought into it. And you've got to have some, the other thing to this, if you're close to those people and you have some passion for making sure that your software does the right thing, to where if you're in a big organization, you know, they're 30 steps away. I'm just marking off tickets and here we go. And if it's clunky, they'll use it anyway. But that's, that's not an environment I could live in for very long.
Lessons from starting a company
What was your biggest struggle starting a company? Cause it sounds like the clientele was kind of baked into your motivation to start. So what struggles did you run into outside of that?
Oh God, everything at first. All the things really. And we didn't, we didn't know anything about business. We were soft computer scientists. We, we knew about how to build software, but not much else. So there was a lot of sort of hard learning until we got a few customers and actually started bringing in revenue to initially where the company started doing okay. After that, that took a while. It's not easy doing the things that we had to do, but even early on our demos of software showed us that product wise, we were on the right track. Cause we are very good. You know, people loved what they would see in our software and for what it could do and funding, how to get to the contracts, how to do maintenance contracts, how to charge for things, you know, how to price things, all those things. And eventually, you know, we were fortunate enough to get through all of that.
We did not have much of that entrepreneurial community back in the early nineties. I mean, there wasn't the whole startup culture. It wasn't as widespread back then as it is now. You know, now there's communities, you can find people, it's easier to find people who can give you advice and help. And it was not, you know, and maybe we just didn't look hard enough for them, but it was a lot of bootstrapping and learning things through the school of art and to get there. But we did have that clientele and it's, you know, the organ transplant world was a small world. You know, I think there were 300 and some odd centers in the U.S. at the time. So we, you know, we sort of knew what the customer base, what the customer base was.
Rare disease data and the challenge of small datasets
Yeah. And I'll give a non-math statsy answer to it because I'm really not a stats person these days. But the problems that I think, you know, there's lots of hard problems in cancer. The pediatric cancer area but it's AI will not cure pediatric cancer, but it can help those who can. So part of that is, if you look at the promise of AI and other things that are out there to help not just AI, sort of other machine learning sort of environments, modeling, you know, the kinds of things you have, the data needs can be large. And one of the challenges in our area of pediatric cancer, in particular, is all these cancers are rare. You're in the rare disease space. And so there's, for some of these types of diseases, there's just not much data that exists, you know. So you're not going to have, somebody's not going to build a gigantic LLM to ingest just pediatric cancer stuff because there's not that many patients with these disease states compared to, let's say, diabetes or even organ transplantation or other. So the lack of data is a challenge and how do you address that? And how do you make tools that work in this rare disease space? That's a big challenge and that's not a stats challenge or other thing. It's just, I think the biggest challenge is just bringing the data together, like what PCDC does, but then just working with data where the populations are not going to be large and making conclusions and the statisticians and others can make their judgments. But knowing that the set size, you're not going to have a whole lot of data to work with compared to other diseases where you can sort of wave a wand and get a large population of patients.
Collaborative multicenter data infrastructure
Within PCDC, or within D4CG, because some of this infrastructure is not just PCDC, there is a data standards team that models what our data model should be, and not every cancer is going to be looking at the same data. Now, the model we use, the data is submitted to us at a line level, but we don't have access to their systems. So, the data, we have it, we don't own it. And then, our group, based on the data modeling, there's a lot of sign-offs and other things. There's a lot of logistical infrastructure that's built out in the PCDC for the approval of the data models, of getting the data ingested to bring it into the PCDC. And then, once you go query the PCDC, there is a governance structure that says, does your data request get granted or not? And those are generally the executive committees within the individual disease groups, they make those sorts of decisions like, yes, D4CG group, you can release this data for this data request for this group. We've deemed it proper and valid and everything like that. So, we're taking the data in, we present it for querying purposes, but it's de-identified, you get down, there's a minimum set size, the typical 10 or something you can go query it. In other models, like if you look at PCORnet, or trying, there's other national distributed network models to where the data is federated. So, the data lives at the host site and the queries come in a distributed basis from all these places that are asking questions. So, you can run a query from one of the client sites and get generally like obfuscated counts or something back, you know, sites A, B, C, D, E, that are in this federated network. Here's your stats, here's how many patients meet this criteria that you put out. And they can do data requests and some places have enclaves set up to do it. We're looking at enclaves within PCDC, so that there's sort of a space internally that someone can have a secure workspace to use this data if they don't want to, you know, have to download it and perform their analysis locally.
Ham radio and open source side projects
And yes, Ruby is one of my kids and has been a member of the Hangout who you know, attends every now and then. So, yes, it is a and I will say I'm very proud of Ruby. Ruby just got her Ph.D. from Harvard.
So, ham clock, you know, I'm an amateur radio person. Ham clock is a tool that shows different atmospheric conditions, you know, sunspots, different, there's all kinds of atmospheric measures about how signal propagation will occur. It is a tool that's been around a long time. It's kind of clunky to compile and get going on a Mac. So, I made a native wrapper for this. I made a WX Python app that wraps this, you know, gets and builds everything you need so you don't have to have developer tools or Mac ports or all the other things that you may need on a Mac to develop. So, after you just take this wrapper that I built for it, it includes all the binaries and everything. It's assigned, you know, signed with a developer certificate, goes to the Apple notary. And there was a video, it got out and video, I think, one of the main video was viewed like 7,300 times on YouTube or something and probably more now. So, I have several hundred users for it and it's fun. It's a thing I do that people enjoy using. I know they use it because if I break an update, they let me know you broke something. Why is it not working?
Career advice: durable domain knowledge and a love of learning
I would think to, you know, the world is a place, you know, we, there's a lot of peril and promise in the world. And there's always been a lot of peril and promise in the world. It's different levels of it now. I think my career advice is, you know, learn that durable domain knowledge in your domain to, because we can make AI applications to do anything, but if you don't know what it's doing, because you only have a vague idea of what it's supposed to be doing, then that's an issue. And so you end up with a lot of these AI slop things that developers, oh yeah, I've coded this and six hours. Well, yeah, but what does it actually do? And can you explain it and fix it later when it breaks? So I would say develop that love of learning the durable domain knowledge, be willing to change. I mean, my 16 programming languages or whatever it is, you've got to be able, because the world will, you know, things do change all the time.
Things like R and Python and baseline functionality, those are going to be there, but the libraries will come, some of the libraries may come and go, but, you know, learning how to solve problems and think deeply. I take a lot of walks where I'm not listening to podcasts. I'm not listening to music. I just take a walk to help figure things out and make my, you know, make the brain work. So I think if you can make yourself, it's hard these days, make yourself take that sort of, I'm not actively doing anything, but these problems are still percolating in my brain. You know, there's a lot that could come of that, just taking a walk and just being, and then coming back and taking a look at it. But I would say the deep learning aspects of it is, you've got to be able to learn. There's still transplant patients that need to be taken care of. There's still pediatric cancer patients that need help, and we need to develop new, you know, either regimens for them, new drugs, new combinations of things, new therapies, new this, new that, new things for the clinicians and the people who are providing the care for the patients. So there's a lot of problems yet to be solved, but you can't solve those if you don't understand what it is you're trying to do. And that's what kept me in this for so long. I enjoy solving those problems, and having seen my software make a difference in those, but it doesn't matter what area. I'm medical, but that could be in fraud or advertising and whatever. You have to know, you've got to know that domain.
You can't solve those if you don't understand what it is you're trying to do. And that's what kept me in this for so long.
No, the soft skill thing is important. You've got to be able to listen to people and talk to people, but not assume you know all the answers, because we never do.

