Transcript#
This transcript was generated automatically and may contain errors.
Hi. So, my goal today is to try and share a little bit of what I've learned about working on the tidyverse for the last 20 years and then try and speculate a little bit about what of those things might still hold true today.
Because of course, you know, the elephant in the room is AI, who's going to try and squeeze right past us. And this really, it does feel tough because it feels like AI has kind of come along and just torn up many of the playbooks that we've had, whether it's for the career, or reaching out to the community, or programming, or open source. It's really hard to tell, like, what's still going to remain relevant over the next, you know, the next 20 years.
So I'm going to start by talking a little bit about the tidyverse in case you haven't heard of it. How many of you here have used the tidyverse? Okay, well we'll go pretty quickly then. So, if you've forgotten, the tidyverse is, the goal of the tidyverse is to create a language that you can use to solve data science challenges and with R code, so it's fundamentally tied to the R language.
And it's really connected to my model of data science, which starts with importing your data, then doing some tidying to get it into a form that's amenable for further work, and then you get into the loop of understanding, of visualizing, of transforming, of modeling your data, then you're going to conclude with communicating that work to someone else. And all of this work you are doing in a programming language. You're not pointing and clicking, you're typing, you're communicating to the computer with code what you wanted to do.
And one of the goals of the tidyverse was to help you not get to like a peak of success, where you have to struggle the whole way, but to help you kind of fall into the pit of success. So hopefully, if you do not know, OAP is what Minnesotans say when they have an accident. And this is very much in line with the vision of John Chambers, who created S, the language that led to R, that really the goal is to take users, to allow them to start in an interactive environment, and to kind of gradually like just slide into programming, without really kind of knowing it, or learning too much about it, but eventually over time, they become more and more sophisticated programmers. And it seems to have worked.
Three things that have mattered
So what I kind of come up with is three things that I think have mattered to the tidyverse over the last 20 years, and will hopefully continue to matter for a long time to come. So abstractions, tools, and people. And so for each of these, my goal is to give you a few of the reasons why I think they have been important, to kind of speculate on why I think they will continue to be important, and then to kind of give a few pointers to my work and some of the other folks at Posit along these lines leading into the future.
Abstractions
So I think, like when I think today, like if I was going to sit down and teach someone data science, like what would I teach them? I, you know, I still love programming. I still love to produce code. I pretty strongly believe that reading code is still really important, particularly if you're doing data science. But it's clear that like writing, the ability to write code is much less important than it's ever been.
But what is still important? I think some of these abstractions are really useful to know. So when I kind of look back at my work, you know, I think one of the really important things is like thinking about tidy data. Like this idea that it's really useful to get your data into a form where every column is a variable and every row is an observation. It's probably like easy now to not have tidy data because a coding agent will just grind away at the problem. It will solve it using like whatever horrendous, awful code it has to do to work with whatever structure your data is currently in. But I also believe that if you up front say, hey, make sure my data is tidy, you're going to end up with code that is easier to understand. It's going to be easier for the agent to write, and it's going to be easier for you to verify that it's correct.
The other thing that I think about is about ggplot2 and the grammar of graphics. And certainly today like does it matter if you know the theme syntax for ggplot2 so you can control the size of the font or the number of ticks on an axis or the angle at which those ticks are? That's probably, that's now I think safely information that you can just allow to fall out of your brain and an agent to take over.
But the fundamental idea of the grammar of graphics, I think that is still so incredibly powerful because it gives you a cognitive framework for understanding visualizations, and it gives you a precise technical vocabulary to talk about them. And I think that's going to pay off regardless of whether you're talking to another human, formulating the thoughts in your head, or communicating them to an agent.
But the fundamental idea of the grammar of graphics, I think that is still so incredibly powerful because it gives you a cognitive framework for understanding visualizations, and it gives you a precise technical vocabulary to talk about them.
And so to me like the grammar of graphics, like the core idea is when you produce a visualization, that's fundamentally a mapping from things in your data, variables in your data, to aesthetic properties that you can perceive like color and shape and size. That's mediated by a scale, that's what controls the details of going from a number to a color or an enum to a shape. It's rendered by a geometric mark like a point or a line or a rectangle, optionally transformed by some statistics, maybe you're doing a histogram, maybe you're doing a box plot, you're not displaying the raw data anymore, you're displaying some distribution or summary. And finally, all of that's going to be laid out on a coordinate system, typically Cartesian, but also very commonly if you're working with a map, you don't want to use a Cartesian coordinate system, or if you are working in polar coordinates, which is often also very useful.
So I believe those like, those abstractions are still powerful and important, because they give you a way to break up these like big thorny problems into smaller self-contained pieces. When you have those abstractions, you have a better ability to think about a problem, and you have a better ability to communicate that to an agent.
So abstractions are important, and I think one of the things that's exciting to me is that AI gives us the opportunity to bring those abstractions to new audiences. So one thing that my team has been working on recently, Thomas Peterson, Tian Vandenbrand, and George Stack have been working on is GGSQL. And if you know what ggplot2 is, and you know what SQL is, you can basically guess what ggsql is. It's a SQL-inspired syntax for the grammar of graphics.
So this is what ggsql looks like. It starts out with some, you know, regular SQL code. You're going to generate a data set, and then you're going to visualize it. Here we're mapping the build length to the x position, the build depth to the y position, the body map to the fill color. We're going to visualize that using a point. We're also going to add on top of that an annotation, a line that shows us a slope of 0.4 with a y-intercept of negative 1. We're going to map the fill body mass to the fill color using a bin scale, and then we're going to add some labels. And then we run that with ggsql, we get something that looks like a very classical ggplot2 graphic, but now we have something that's 100% accessible to someone that only works in SQL.
And I'm pretty excited about this, because I think the grammar of graphics is an incredibly powerful framework for thinking about visualization, and this is not something that SQL users have had access to in the past. The other reason we've been working on ggsql is it's also sometimes nice to generate a visualization in a language that's kind of guaranteed to be safe. If you're generating R code to create ggplot2 or Python code to create matplotlib, that code can do anything. Hopefully it's only going to create the plot, but it could be doing whatever to your computer. Nice thing about SQL, you can run this and you can know this isn't going to go and delete files, it's not going to read your passwords and email them to maybe Grok.
Tools
Okay, so I think abstractions have mattered and continue to matter. The other thing, the next thing that I think really matters are tools. And for me, tools have always been this lever that has allowed me to have a much bigger impact than I could without them. And when I think about the sort of tools that we have worked on, this is not just the tools for data science, but we very quickly got to a point where we were developing so many R packages that what we really needed were tools to help us develop R packages.
And so this is what led to the development of a lot of the key kind of software engineering infrastructure for R, like TestThat for unit testing, Roxygen2 for documenting your package, DevTools, which makes it as easy as possible to work with kind of a development version of a package rather than having to install it all the time. And then most recently, Packagedown, which turns your package documentation into a beautiful website. So many, I mean, I think pretty much all of these things you can take for granted with Python. Python has had these for much longer than R ever did.
But these were some of the tools that we realized we needed these software engineering tools in R. And they, you know, fundamentally accelerated the pace of development because with unit tests, you know, we could be much more confident that the code was actually correct. And when fixing bug A, we didn't accidentally produce bug B.
And I'll talk more about this later, but like the goal of software, I guess at least traditionally, has been for people to use it, for humans to use it. And humans need not just to see the code, but they need to understand what it does. They need good documentation. They need aesthetically pleasing websites that they enjoy browsing.
But to me, tools aren't just code. Like another set of tools we have are checklists. Like I think checklists are an incredibly, incredibly powerful tool when you cannot fully automate something. So this is the checklist, which I see is very difficult for you to see. It doesn't really matter. But this is the checklist that we follow before we release a new version of an R package. So this is something that we have tooled up. We have an R function that creates a GitHub issue that has all these checkboxes that force you to work through our standard release process that encourages you to follow all of the best practices that we've discovered over the last 20 years.
One of the maybe surprising things that's part of our release process for major releases, minor releases, is writing a blog post. You are not allowed to release your package until you have written a blog post. And this is because one of our very hard-won learnings was that very reliably when you write a blog post, that's when you discover a bug that you haven't discovered the whole other time of developing the package. So we now always do the blog post before we actually do the test.
And of course we have like a fundamental new tool now, and that's Cloud Code. And I have found Cloud Code to be like tremendously kind of fun and empowering. And here for a few packages that I've kind of been noodling on recently, we've got on the y-axis the number of open issues, the x-axis is time, and you can see the points at which I have like decided to attack one of these packages with me and with Cloud Code. So I've always been a very efficient issue closer. I say that not necessarily an efficient bug fixer, but I am very good at closing issues quickly. I think Cloud has made me even faster at closing issues, but it's also increased the quality. Like now I actually fix more of those bugs rather than declaring them as won't fix. But I think you can see from this, like now I can fairly reliably in two weeks to a month close 100 to 150 issues. And this is me doing this kind of as part of my regular job, squeezing it in between meetings. This has really like fundamentally accelerated my R package development experience.
And I don't know if you can see this, but the kind of magic line for me is 25 issues, because once you get below 25 issues, they all fit on a single page and get up. And I think that's a thing of beauty.
I've also been using this new tool, Cloud Tools, to make other cool tools. So I have this package called Bananarama, which is what I use to generate all of the images in this presentation. It's a YAML specification that allows me to declare a common style for all of the images and then allows me to say, hey, this is what I want. Go away and draw it. And the thing that's particularly important is it allows you to kind of anchor. This is just a Gemini nano banana feature. But you can anchor the prompts with images, increasing the consistency across the images in your presentation.
And this feels like this feels like really fun to me. Like, I am just having a great old time, like, ripping through issues in R packages. But at the same time, it also kind of feels like, well, when you kind of look at the outside world, like, well, there's a lot of, like, crazy shit going on out there. And is it really the best use of my time to be fixing bugs in these R packages? So while I love doing that, I've also been trying to force myself to, like, engage more in the fear and anxiety that AI brings to pretty much all of us.
And so I want to talk about a new tool that I've been working on. I think this is the I guess this is the first tool I've ever developed in at least 20 years that is not explicitly about, oh, this is a cross-platform, cross-language tool like everyone these days. Of course, I am writing a standalone command line interface in Rust. But the idea here is to come up with a new specification for data dictionaries. You know, people have tried to do this many times over the last years. But I think it's time to take another look at it. Because I think it is so important now to encourage people to write down what they know about data. Because if you haven't written it down, your agents aren't going to know about it. And, you know, plot twist. Actually, this is going to help your human colleagues as well.
So the basic idea of DataDict is, like, very, very simple. You're going to declare, like, for each of the tables in your data set, for each of the columns, like, what are they? That's a string. Oh, this is the primary key. And we're going to have some free text description where you can include whatever you know, whatever weirdness and oddities. And we're also going to include a few examples. Because examples are so useful for both humans and AI agents.
And then DataDict.yaml is coupled, as I said, with a command line interface. That's going to allow you to validate your spec, to validate the data. Using some of the really cool tools in Rust for producing beautiful, informative error messages. That allow you to see exactly what's gone wrong and why. And the other thing that's really cool with command line CLIs these days is you can make them completely self-documenting. So that if an agent wants to use this, it's going to call DataDict. It's going to see, oh, actually, there's a couple of skills bundled here. And then it can read those skills and learn how to use DataDict.
So I think one of my, like, one of my early fears when AI was really picking up was, like, how? Like, what's going to, like, this is great for ggplot2, right? There's, like, 20 years of ggplot2 code on the internet. Great for me. What's going to happen with the next generation of tools? How is AI going to know how to generate their code? And I think what we're learning now, the way we deal with that, is if you're producing a new tool today, that tool needs to be self-documenting. And, like, in many ways, that's great for agents, but that's great for humans, too, right? Great to bundle everything you need to learn the tool that you're trying to use.
People
And the last thing that matters, maybe the most important thing that matters, are people. And this has always been, like, front of mind to me. Because if you build it, like, they won't come. Like, you can make the best tool in the world, and if you don't tell anyone about it, like, they're not going to use it.
And I think this is something that, like, folks in academia often feel kind of uncomfortable with, this idea of, like, tooting your own horn or, like, self-promotion. But seriously, if you don't tell people about the awesome work you're doing, they're never going to learn about it. Like, your impact on the world is a product of the quality of the work and the quantity of people that are using that work. And I think you almost have an obligation, if you think you're doing good work, to tell people about it.
Like, your impact on the world is a product of the quality of the work and the quantity of people that are using that work. And I think you almost have an obligation, if you think you're doing good work, to tell people about it.
So I've always believed that, like, marketing is really important. One of the cool things, the things I loved about the R community is our culture of hex stickers. These are, first of all, they are infinitely better than Python stickers, because they tile the plane beautifully. You can put a bunch of stickers on your laptop, and they don't have to overlap each other, which is a crime against nature. And it's just such an amazing way for people to kind of express their personality and their humanity. So many kind of great conversations, connections that have been kicked off because someone sees, oh, hey, I saw on your laptop that you used this package. I used that package, too.
So that kind of community connection, I think, is super-duper important, too. We've also spent a lot of time working on cheat sheets, making it as easy as possible for people to learn a new tool.
But very, very obviously, like, the way that community works has changed. So this is a plot of normalized stack overflow activity, answers, comments, questions, and votes. And you can see that basically stack overflow, like, it's dead. Stack overflow is dead, and AI killed it.
And I think that, like, of all of these things that matter, I don't really know what's going to happen with people, and I don't have any particularly great solutions. But I just want to acknowledge that there is this massive change that's happened. Like, was stack overflow the best community? No. Was it the friendliest community? Absolutely not. Although I will say, like, compared to me growing up in the R help mailing list, the stack overflow was infinitely more friendly than that mailing list, if you can imagine it. So, like, I don't know what's going to happen, but we're going to need to create new mechanisms for community.
I do get the sense that we're sort of starting to, I think, recover from COVID slowly, that there's an increasing kind of desperate need for personal connection, particularly as more and more of us are just interacting with AI all day. I think there's, like, a growing need to connect with other humans, to share, like, our fears and anxieties and hopes and excitements about the future.
Summing up
So, to kind of sum up, you know, the three things that I think have always mattered and will continue to always matter, abstractions, like, breaking down complicated problems into simple primitives that help us understand them, that help us write code to describe them, super, super important. The tools we use, really important. I guess the counter to that is also to not, tools can be intoxicating, too, and you have to always make sure you're stepping back and saying, like, am I having fun or am I doing, like, the right thing, the most important thing? And then, finally, people are the most important thing. We don't know, like, the way communities operate has had a fundamental tectonic, irrevocable tectonic shift. We are going to need to figure out how to deal with that over the next five to ten years.
And then just the very last slide, you know, I just want to make sure that I express, like, I feel both, like, you know, fundamental existential dread about AI, like, what does it mean for me, for my identity as a data scientist, as a software engineer, as a community builder, and this, like, excitement of this empowerment, you know, like, I vibe coded this talk timer iPhone app, which I never could have done before. This, and I don't, there's no resolving, there's no unifying these feelings. All you can do is kind of accept that you're going to feel these tensions pulling you in both directions. There's nothing you can do about that except accept it and, you know, try and take a deep breath and find other people to share those experiences with. Thank you.
Q&A
And while we take questions, one of my unpaid interns is going to distribute 3D printed snakes, if anyone would like a 3D printed snake. Thank you, Hadley. We have a few questions on Slack. The first one is from Henry Schreiner. Can you convert this data dictionary to JSON schema? It seems similar enough that it could be a mechanical translation from that.
Yes, so one of the reasons for choosing, I think, YAML, particularly YAML 1.2, if you really want to get technical, is the sweet spot for a file format that is machine readable, is human readable, and agent readable. Absolutely. It is, you know, there's a one-to-one correspondence with some JSON document. Making people write JSON by hand is cruel and unusual punishment.
Question from Thomas Caswell. Will mailing list make a comeback? Discourse instances, real-time chat only? Yeah, I don't know. I mean, I guess this is an appropriate time for me to plug. Actually, I don't know how to link to my... I don't even know how to show you. It's called... What is it even called? I created this at an early age. But, like, I don't know. Like, this is one of my experiments that, like, yeah, I actually like... It feels weird to say this, but I actually like people sending me emails now. Like, it's kind of nice to get a long-form email that I can read. So I'm experimenting doing with this. I'm trying to explain AI as much as I know about it as we go. But, yeah, like, I think maybe mailing list will make a comeback. Maybe discourse. I'm kind of hopeful that maybe it's time to bring the in-person meetup back. Because I know that that in-person connection, again, is just so valuable and so satisfying.
Well, this is a big in-person meetup.
A question from Christian. Are you worried about how we will learn and discover new things that matter when cognitive outsourcing to AI is so available? A lot of most productive AI users are leveraging their pre-existing skills when using AI.
Yeah, absolutely. I think it's, like, to me, it's very clear, like, from plots. I can't even find the spot again. From, like, the fact of, like, how... Like, this plot. Like, I am a massive winner from AI. Because I have 20 years of experience. I can very, very quickly look at a PR and gut check whether this is correct or not. Like, what does this mean for the next generation? How the heck are they going to learn when there's just such this tempting siren of, like, I can just have AI do that for me? I... Yeah, it's scary.
I am hopeful on the whole that I think this is a problem we can solve. Like, we need to discover new ways of doing this. Doing things. It feels like maybe, like, maybe everyone should have one AI-free day a week. Where you can deliberately turn off your AI and do things like the old-fashioned way. To practice those old skills. But I'm also conscious of, like... Like, I don't want to be the guy that's, like... You know, back in my day, I had to walk to school in bare feet, in the snow, uphill both ways. And now you're gonna have to do that, too, just because I suffered. I'm gonna make you suffer. Like, I very strongly believe that, like, I suffered so you did not have to suffer.
And I always worry where I hear myself saying, like, the things that my curmudgeonly old professors said to me back in the day. That, like, we also don't need to make people, like, retrace our footsteps of learning. We can discover new ways for people to learn. Maybe this audience will get one of my somewhat abstractly related quotes to this. But, like, ontogeny does not recapitulate phylogeny. Like, when you look at an embryo developing, it kind of looks like it's going through all these stages. Like, it looks a little bit like a lizard and it looks like a bird. But that's not really what's happening. And we don't have to make people go through all of the same painful stages we did. We can figure out how to skip that and still have people learn.
And one more question. I don't want you to miss your break. A question from Ken Williams. Great talk. What is the difference between documented and self-documented?
There is no difference, I guess. Yeah, I guess I don't really understand that question. As long as you've written it down, you've documented it. If you share that with other people, that's the whole point of documentation.
