Transcript#
This transcript was generated automatically and may contain errors.
Hello, everyone, and welcome to the TDWI webinar program. I'm Andrew Miller, and I'll be your moderator. For today's program, we're going to discuss from Wrangling to Insight, Human-in-the-Loop AI for Analytics, and our sponsors are Posit and AWS. For our presentations today, we'll hear first from Dion Larson with TDWI, and after Dion speaks, we'll be joined by James Blair with Posit and Shun Mao with AWS for a presentation and panel discussion.
Before I turn over the time to our speakers, please allow me to go over a few basics. Today's webinar will be about an hour long, and at the end of their presentations, our speakers will host a question and answer period, so if at any time during these presentations you'd like to submit a question, just use the Ask a Question area on your screen to type in your question. If you have any technical difficulties during the webinar, you can click on the Technical FAQs area, and you'll receive technical assistance, and if you'd like a copy of today's presentation, you can locate the resource window to download the PDF. Lastly, we are recording today's event, and we'll be emailing you a link to an archived version so you can view the presentation again later if you'd like, or if you'd like, you can share with a colleague.
Again, today we're going to be discussing from Wrangling to Insight, and in the loop, AI for Analytics, and our first speaker is Dionne Larson. She's a research fellow with TDWI and president of Larson & Associates. Dionne is an active data science practitioner and academic. Her research has focused on enterprise data strategy, agile analytics, and data science best practices. She has consulted for several Fortune 500 companies, presented at multiple conferences, including TDWI, and has authored numerous research articles on data science methodology and best practices. With that, please welcome Dionne, and I will hand it over to you now.
Human-in-the-loop analytics landscape
All right. Well, thank you, Andrew. So, let's go ahead and get started. So, our agenda today is first to really look at the landscape of human-in-the-loop analytics. Second, we're going to see a demo from Posit and AWS on how to use HITL. We'll also have a panel discussion, and then we'll open up for Q&A.
So, today we are focused on diving into human-in-the-loop analytics, or HITL. It's a concept that's redefining how we approach automation, analytics, and AI readiness. So, despite all the progress in tools and models, the truth is that data prep still consumes about 60 to 80 percent of analytics project time, and that's where human-in-the-loop assistants really create measurable value. They shorten our time to insight. They help us keep domain experts firmly in control. We're also seeing tremendous energy around AI adoption, but readiness and trust remain kind of uneven. HITL doesn't replace human expertise, as most people think. It really operationalizes it. It helps us strike a balance between automation, accountability, and create a path to scale analytics responsibly.
So, let's take a look at where organizations stand today. A lot of this data comes from TDWI research, and our latest data shows that only 30 percent of organizations have leadership fully aligned on AI strategy and implementation. Another 25 percent of organizations understand that AI's impact is there, but they don't really understand how to operationalize it. This is where guided HITL approaches actually come into place. They help us deliver pragmatic progress. Only 16 percent of organizations are actively executing AI strategies at scale, and the majority are still in planning or doing some kind of proof of concept. Perhaps the most telling is the culture of AI trust. It's still neutral across most companies, and there's a recognition of potential, but there's hesitation around accountability. And human-in-the-loop methods, they help us strengthen things like transparency and governance during this first phase of maturity. They make it possible for leaders to move forward confidently while maintaining oversight and ethical boundaries.
So, if we look past the strategy layer, we'll find some persistent gaps in foundational data practices, especially around quality definitions and accountability. In our research, we found that 38 percent of organizations have no defined data quality strategy, while another 33 percent rely on IT-led approaches without getting wider stakeholder buy-in. Even more concerning, 41 percent of organizations report no company-wide data quality training and half have no shared data definitions across business domains. So, these are areas where human-in-the-loop oversight can make a tangible difference. These types of checkpoints allow experts to review outputs, enforce our shared standards, and then validate that automated processes align with the organizational expectations. In short, HITL becomes the quality control layer that many organizations lack formally. It helps enforce governance in practice before policy catches up.
In short, HITL becomes the quality control layer that many organizations lack formally. It helps enforce governance in practice before policy catches up.
So, now let's look at the technology landscape. While AI platforms are proliferating, the ecosystem is still forming, and it's far from finished. Only 33 percent of organizations have a set and communicated ecosystem strategy, and a mere 12 percent are measuring outcomes against it. But on the upside, we have 54 percent of their organizations report active machine learning or AI support, and 51 percent of organizations are piloting gen AI initiatives. Integration is actually improving as well. 52 percent of organizations have APIs and connectors in place, but only 21 percent report really having a unified data view. We still see silos as a major obstacle. Here again, human-in-the-loop validation adds stability during this transition. It also requires human review of certain things like data joins, transformation logic, AI outputs, and organizations can prevent misinterpretation across data sources while integration and virtualization technologies continue to mature.
So, let's talk about where organizations are actually investing in 2025 and how those priorities actually intersect with HITL. The top analytics and AI priorities include focusing on self-service analytics. 30 percent, 37 percent of organizations basically reported they're focusing on self-service analytics, and they want to support empowering business users with guided, explainable assistance. Also, no surprise, gen AI is also up there. We're looking for human oversight, ensuring responsible implementation. There's also reporting of around 32 percent of focus on data availability and literacy, emphasizing both access and understanding. But then on the data management side, we have priorities where organizations are emphasizing automating routine tasks. About 34 percent of organizations reported this, but they want to do so under human supervision. We also had organizations, around 31 percent of them, focus on catalog and metadata management where they're wanting to improve discovery and lineage tracking. Governance frameworks also ranked up there where they're wanting to look at embedding human checkpoints. And then lastly, advanced analytics support. About 42 percent of organizations looked at pairing AI acceleration with human validation. In essence, HITL analytics is the connective tissue among these priorities, and it helps make automation safe, interpretable, and operationable.
So what does human in the loop analytics really look like in the real world? AI-driven assistance will take on repetitive mechanical tasks such as schema inference, key detection, anomaly flagging, and also code scaffolding. Meanwhile, users and humans have to stay responsible for the context-rich steps such as defining your data transformations, the sampling strategies, handling outliers, and setting what those acceptance criteria are going to be. Every action is captured for traceability from prompts and generated code to decisions and approval. In an HITL ecosystem, for example, each interaction can create a reproducible artifact with full logging and auditability. This ensures transparency not only for analysts but also for governance and compliance teams. It's the best of both worlds where automation is focused on where speed matters and human oversight is placed where judgment counts.
So one of the most critical strategic discussions right now is determining what to automate and what to keep under human control. Tasks that are consistent, rule-based, and low-risk are ideal for automation. This includes some typical things like data type detection and classification, missing value handling, understanding joint paths, maybe even a recommendation for a specific chart or visualization, creating summary statistics as part of exploratory data analysis, or even maybe the first draft of code generation. These are some building blocks where speed brings efficiency without really risking some kind of interpretive error. But other tasks, those things that involve judgment, understanding ethics, maybe even business meaning, these have to stay human-driven. This means that defining metrics, dealing with bias and ethics assessments, interpretation of compliance, and those acceptance thresholds, these really need to be part of what a human does. Until we have things like data quality maturity and training and we are able to catch up, human involvement remains indispensable in these areas. Remember, the goal isn't to automate everything, it's to automate well and to do so responsibly.
So if we bring the big picture together, we're at a moment where we're excited for new things, but we also have to be cautious. Automation is obviously advancing very quickly, but maturity, governance, and trust are evolving more slowly. HITL analytics helps bridge that divide. It allows us to move faster and safer, and we can automate what we can, but we have to anchor all of our decisions in human oversight. So back to Andrew.
Positron and DataBot demo
Terrific. Thank you so much, Dionne. That was a great presentation there. And now it's my pleasure to introduce our first guest speaker today, James Blair with Posit. James is a senior product manager at Posit where he focuses on helping Posit commercial products seamlessly integrate into cloud platforms and environments. He has a background in statistics and data science and finds any excuse he can to write R code and ride his bike, although usually not at the same time. With that, please welcome James, and I will hand it over to you now.
Thank you, Andrew, and thank you, Dionne. I want to just take a moment briefly to provide some context on Posit, and then we'll walk through a quick example of this sort of human-in-the-loop idea and how we're thinking about it in relation to exploratory data analysis and this broader ecosystem of data science and data analytics. On the Posit side, I won't go into too much detail. We do a lot of open source contributions to both the R and Python ecosystems. We'll look a little bit at what some of those contributions are here in a moment, and then we also have a number of commercial products that support the use of those open source tools in large-scale organizations. Across the board on both the open source and the commercial side, we've invested heavily in the AI space with tools like Positron Assistant, DataBot, and QueryChat, which are available in various forms, either R and Python packages, plugins available inside of our IDE called Positron, or just native features of the Positron environment itself.
And what's been really exciting is over the past several months, we've been hard at work making sure that not only do these tools work really well when applied to data-specific, data analytics-type problems, but they also work very well in conjunction with platforms like AWS. So what we want to take a look at today is how these tools, Positron Assistant and DataBot, can combine with services like Amazon Bedrock, S3, and other parts of the AWS ecosystem to really provide a very intuitive way to explore, work with, analyze, and understand data at a much more rapid rate than we've been previously able to do, but without sacrificing reliability, reproducibility, or just the ability to understand and audit the decisions that have been made along the way.
So with that said, I'm going to share my screen here, and we'll take a quick look through an example of this. I'm going to share, this is the Positron developer environment, which is a fairly new tool that we've created here at Posit. Like many modern developer tools, this is built on a foundation of VS Code, but it's specifically built around the needs of data science and analytical work. And when I say that, the main difference in my mind is that Positron's designed around the idea of you have some sort of actively running environment, maybe it's R, maybe it's Python, Positron supports both, and you're using that actively running console to investigate data in real-time. So I'm executing code in real-time, I'm looking at results, based on those results, I'm executing new code, and there's this very reactive, very involved loop that's happening that Positron supports.
What I want to highlight here is on the left-hand side, because again, this is built on the foundation of VS Code, if you've used a VS Code-like tool before, you'll see many things that are familiar over here, but this icon here opens up what we call Positron Assistant. And Positron Assistant is an AI tool that can allow developers, data scientists, statisticians, others in this environment to ask questions, query, provide queries about the code that they're writing, or the data that they're exploring, and it can operate in a number of different ways. Here we can see I indicated, look, I want to use DuckDB, and I'm going to interact with data that's in S3, and it's giving me this response that helped inform some of the code that's been written here on the right-hand side. If we look down at the, or excuse me, if you look at the top left-hand corner, we can see that this is configured to use Amazon Bedrock behind the scenes, and then down here at the bottom, I can choose which specific model from Bedrock I want to interact with. So I have the ability to interface directly with foundational models through Bedrock, and then the choice of which particular model I'm interacting with.
On the right-hand side, we have the code that was generated and provided, and so we can run this, and this gives me a DuckDB interface or connection into some weather data that's stored inside of S3. And now what's really exciting is, now that I've got this connection established, and it's all running through AWS and powered by AWS, I can come in here and I can say, let's open up DataBot, which is this agentic tool that we've developed that is very purpose-built to support interactive data analysis. So now that this is opened up, it'll give me some suggestions of what I might want to do. Do I want to load and explore some data? Do I want to do some sort of statistical analysis or modeling? And in our case, if we look at our environment here, I'm going to collapse some things. I've got this weather data that I've already created this reference to. The data is in S3, but I want to work with this. So I can come into DataBot and I can say, help me explore the weather data.
And with this prompt, because DataBot is operating within the confines of Positron, it understands, hey, it looks like you have this weather table already loaded in your environment. Let's go ahead and take a look at what this includes. Now it's going to write some code and then it's going to ask for permission to run that code. And I can see and inspect what it wants to do. It wants to kind of take a look at what this table is. So we can just go ahead and say, let's allow that. And now one of the interesting things about DataBot in comparison to perhaps some of the other agentic tools that exist today is that it's very closely linked and really built around this idea of this real-time feedback loop of how we work with data.
So it's going to suggest some code here. It said, let's just look at the table. Let's look at the data that we have. And so from this execution, it can tell this looks like it's a DuckDB reference. It looks like you're using this weather data from 2023. Very interesting. Let's go ahead and take a look at how many observations we have. So it's executing this additional query to figure out how many different unique observations exist.
And so DataBot operates in a way that's very similar. In fact, in many ways, as I've used this tool over the past several months, I find that it's nearly identical to the types of things that I would be doing if I were writing the code out myself. It's just operating at a much faster rate than I would be doing if I was the one typing in every single keystroke. And along the way, one of the things that we've been particularly focused on, a couple of things honestly, has been making sure that this leaves a very reproducible, very verifiable audit trail. So I can see exactly what code is being executed. I can see exactly what results are being returned, plots that are generated. I can see. So I have full visibility into the entire process. I can make course corrections if it appears that we're getting off course.
And the other thing that we've done, if you've ever used any of these foundational models for any type of work like this previously, they have a tendency to be very overeager. And with a simple request like, hey, help me explore this data, they might go off for many, many minutes at a time exploring all kinds of different things and come back with an overwhelming amount of information. Here's what I found here. Here's what I found here. Here's what I looked at here. And it can be really difficult to sort through and filter through the noise that some of these tools can tend to create. With DataBot, we've been very particular about letting it explore the data to a certain extent, but then always coming back to the user with what it's found and with some suggestions on what else can be explored.
So here we looked at what are the unique different element values that exist in this data. It gives me a summary of what it found. We have 11.3 million precipitation records. In snowfall, we have 5.6 million records. So it gives me a little bit of a summary. And then it comes in and provides a set of suggestions of what we might want to look at next. And I could choose one of these suggestions. I could say, yeah, let's go ahead and look at snow data and look at seasonal patterns if I wanted to. Or I could come up with my own. I actually want to look at temperature trends or I want to look at something else and supply that. And again, DataBot will go off, execute. In real time, I can keep track of what's happening and then come back to me with what's been discovered, where do we want to go to next. And it's this very collaborative effort to explore data in real time.
The final thing that I'll mention here before we turn the time back over is once I've completed this process, so maybe I've explored different parts of this data and I've realized that there's a few different things that I'm particularly interested in, and I want to preserve this analysis in a more persistent format. Rather than just this conversation that I've had with this tool, how do I preserve what has been learned? And here inside of DataBot, there's a command called report that will easily allow me to say, hey, take what we've explored and let's create a reproducible document that contains not only what we found, but all of the code that we ran to find those results. So at the end of the day, I end up with an artifact that's reproducible. I can run it against this data in six months. And if this data's changed, I can get a new version of that report. And it's all been helped with me staying involved in the loop with this DataBot tool as we've actively explored this data in real time.
We've been really excited about this, as it mentions at the top when you open up DataBot, that this is a research preview. We recognize that there's still lots for us to learn and much for the community to learn about how to use these tools responsibly. But the early returns here have been very promising, as we've made our very best effort to keep the data scientists, the analysts, the statistician, whoever it is, looking at the screen, typing in commands, keeping them intricately connected to this process throughout the entire time. Because at the end of the day, it's the creativity, it's the questions, it's the curiosity that we as humans bring to this process that makes this work as well as it can. And so allowing us to continue to express our own curiosities by prompting this tool to explore data in the ways that we find most meaningful has been a great experience for us as we've built out this tool and a great experience for those who have had exposure to this. So with that, I'll stop my screen share and pass the time back over.
Because at the end of the day, it's the creativity, it's the questions, it's the curiosity that we as humans bring to this process that makes this work as well as it can.
Panel discussion
Terrific. Thanks so much, James. That was a great presentation there. And now it's time for our panel discussion. So please allow me to introduce our final guest speaker today, Shun Mao, with AWS. Shun is a senior partner solutions architect in the artificial intelligence and machine learning independent software vendor partner team at AWS. He has years of experience in data science, analytics, AI, and cloud computing across different industries, including oil and gas and pharmaceuticals. At AWS, he helps strategic AI ML partners build novel products and solutions to bring business value to customers. With that, please welcome Shun, and I will ask Dionne to rejoin us, and I will pass it over to you, Dionne, to begin our panel discussion.
Excellent. Well, welcome back. Let's go ahead and start with a few questions. So I'm just going to direct this to Shun and James, and you guys can kind of step up to see who would like to provide us with the first answer. But our first question is, what audible artifacts do assisted EDA sessions produce for reproducibility?
I'll take the initial stab at this one. I think kind of like what we walked through in the demo, from our perspective, as an organization, we have always been, since the beginning, highly focused on the value of code for a number of reasons. It's flexibility, it's portability, but also this notion of reproducibility. So with tools like DataBot, we've been very focused on, how do we assure reproducibility? And that's taken the form currently of the ability to export and generate reproducible reports directly from interactions with these tools. And DataBot's not the only example of that in the industry, I don't believe, but that's been our perspective, is let's generate the code, let's get the output from that code, and then at the end of the day, let's let the analysts decide which of these things were most helpful and beneficial, and let's wrap all that up into a report that's reproducible and re-executable down the line.
Yeah, I agree with James. Actually, I think, Dianne, in your presentation, you also pointed out that quite a lot of things can be actually preserved in order to reproduce a result, such as the prompt itself, the response, the code itself, and also the explanation of the code. So why the model would generate this kind of code, right? Because a lot of times we actually care more about a reason, not just code itself, right? So those things combined can really help people, help organizations reproduce the whole process. I think also I want to touch upon something behind this thing. So as James has demonstrated, so behind the scene, we're using a large language model. In this case, we're using Amazon Bedrock. So in that layer, we have an extra layer of this recording or logging in order to reproduce. So behind the scene, so if we're using Amazon Bedrock, it actually logs all the model invocations, such as timestamps, prompts, or even number of tokens and the cost also, and including the code it generates, all the things going into the CloudWatch. So you have this full record of the things you have done. And also, we have the AWS CloudTrail for each of every API call. So if later you want to do some governance and compliance, you can go back to check those API call histories and also CloudWatch logs to reproduce things you like to do. So that's an extra layer of the protections or the logging in the lower infrastructure level. So all of this can also be probably great metadata for future projects like the one that was just executed.
All right, let's move to the next question. So how do you detect and prevent hallucinations or data logic errors before code runs?
Yeah, maybe I can start. For this one, I think James' demo is a perfect example of showing that. So I think one of the typical patterns of us, I think maybe back into last year, so when we are dealing with large-length models, we still find there are some areas that the model will give some plausible answer, right? But we don't know whether it's true or not. But now, actually, in this data science domain, when you are using this tool, actually, you have something in your mind. So you have the business problem, you have the data schema, you have the problem you want to solve in mind. And the way that I think we're dealing with this model in data science domain is that it's more interactive, right? So we're not letting the model to do full autonomous. So we are letting the model do each step by step. So we focus on a small problem each time. So in that way, you have more time to think whether the model is really giving you a wrong answer or a hallucinated answer. You have more control over these granular steps and calculations.
Another thing I want to say that you can always have this code review process. So in your organization, you probably already have this Git control or code control processes. So once the model generates all this code, you can still utilize your existing structure of code control that it can pass certain review processes or test processes to make sure that the code itself complies with all the standards that you set up before.
I agree with what Shun said. And I think the other piece to this is this just highlights the undercurrent of this whole conversation of keeping the human in the loop and the value of that process. And I think one of the things that is so valuable of keeping the human in the loop is that we are the domain experts. We have an understanding of what this data represents. And the large language model will develop its own construct of what the data is. But it may be accurate. It may be a little bit inaccurate. And so there have been times where I've been using these tools. And I've realized that the model made some sort of assumption about the underlying data structure as representation that was inaccurate, not to any fault of its own, but just there wasn't enough context. So it just wasn't able to piece together the full picture. And those are opportunities to kind of course correct and say, wait a minute, you looked at this column this way, but it actually has already been averaged. And so we shouldn't be re-averaging it again, or whatever the case is, and then kind of course correct from there. And so I think a big part of this, to Shun's point, is instead of just saying, go do this whole analysis, generate a report, and send the email out, we're saying, let's work on this together. Let's explore together. And that idea of togetherness allows us as the individuals to keep eyes on what's happening and to recognize, hey, wait a minute, that actually doesn't make sense given what we know about this data.
Right. And also, I think another way to actually mitigate this kind of tendency to have a hallucination and also data logic error is that you have all this generated code, and you always start from a sample, a small sample of data. Do a test on these ones before you go to a larger scale of even production data, right? So you have this kind of buffer layer for you to possibly make a mistake and find a mistake and correct it. Yeah. I think that was going to be my point as well. It's really incremental learning. We talk about the large language models giving us output. We talk about human-in-the-loop interaction. But in essence, you're collaborating in your learning, and there'll be continuous improvement over time.
Absolutely. Let's go to our next question. OK. How do you balance speed gains from AI assistants with the need for analysts to learn and grow their skills?
I can start with this one. I think it's such a good question. I think it's one that many of us find ourselves thinking about either directly or indirectly, right? Because I think another part of this is just how do we — kind of what is our role as individuals and as analysts in this sort of AI world that we live in today? And as I thought a lot about this very idea, I think one of the key things that I keep coming back to is what are the skills and what are the values that we as individuals bring to this process?
And I think for a while, when I think about my own experience and career as a data scientist and kind of in this field, there was a point in time where it felt like, okay, the thing that I contribute is my ability to write code and my understanding of code and how this works. And so I can use code to solve problems and that is what makes me useful. And I think, yes, there is usefulness in there. But as we've seen these AI tools emerge and their almost uncanny ability to generate code, it's made me kind of reflect back and say, well, is my value just my ability to translate what I'm thinking into working code or is it something more? And I think a lot of this is while, yes, there's value in being able to write the code, there's I think even more value in my willingness to be curious, my willingness to explore and ask questions about the data that I'm looking at that might otherwise go unasked. And I've realized that the skills that I bring to the table are less about my ability to kind of write code and more about my ability to understand the data and to think creatively around and through what that data represents and to ask questions that eventually lead to meaningful insights and eventually business outcomes and decisions. That's the skill that really matters at the end of the day, in my opinion.
And so I think this in a way is kind of this natural complement where I can kind of offload some of the time I used to spend writing out the code and instead spend more time thinking very critically about what I'm actually learning and what I'm exploring. And we've seen, we've had a couple of really interesting internal experiences with some of these tools where we've worked with very well-known datasets that have been widely scrutinized across the industry, public datasets that have been used in examples and content all over the place. Within just a few moments of using some of these tools, we discovered latent structures and even in some cases, flaws in that data that hadn't been discovered previously. And I believe a big part of that discovery process was we were no longer kind of time constrained by the amount of code that we could write. And so the opportunity to just explore the data in new and creative ways was bigger than it ever had been. And that opportunity led to the discovery of some of these new things that just, after hundreds of hours of people looking at this data, just hadn't really been seen before because the opportunity wasn't there. And this gave us that opportunity. So the short answer, I guess this was a long answer, but my short answer is, I think these things actually complement each other really well. And the speed gains from these AI tools enables us to focus on the other skills that we bring to the table that are truly unique, our creativity, our curiosity, those kinds of things.
Yeah. Well said, James. I think this kind of question really happens in multiple time inside of the human history, actually. So I was thinking about the time when the calculator was invented, right? So before that, everyone is crunching, taking a lot of time crunching numbers and doing all this kind of mathematical calculations. But after a calculator was invented, do you think we don't have mathematicians anymore? But we don't, right? Because the calculator actually frees us from the labor-intensive kind of work. You know the principles of doing this calculation, but it's just repetitive. It doesn't add you a lot of value, right? So I think now in the AI world, it's similar, it's different, but there's also a lot of similarities. So when I deal with AI and large venture models, I feel that now I have more time to ask good questions. So the repetitive work AI can take over, so that same time actually allowed me to explore more fields. So instead of, let's say, accomplish just one task using one approach, now you have more time to say, okay, do you have other ways to solve this problem? Maybe you never thought about that before, right? Now you have the time to think, okay, maybe the second way is maybe a better way I've never thought about before. And now once I implement it, it's another whole world open for me. So I would say this AI world is really opening a lot of opportunities for normal people to learn more, to expand their knowledge base, to expand their capabilities, if you really know how to properly use AI. So also it depends on, for organizations, how do you implement the policies of using AI? So the bad implementation could be, you just ask AI for a code and implement it and run it without even understanding what's going on, right? The good implementation is that you always ask for reasons. You always ask for the background, you always ask for the logic, why the model is suggesting in this way, and maybe you expand that to ask for more routes, more methods to accomplish the same kind of task. So to me, actually, it's a great, great way to really help you develop your own skills and your thinking and abilities, yeah.
Yeah, excellent comments from both science and of course, it even goes beyond this specific discussion around human in the loop analytics. It just basically demonstrates how AI should be integrated across the board and how people still need those critical thinking skills and still need skill sets, but it'll help accelerate and grow those skill sets. All right, let's move on to the next question.
So what rollout pitfalls do you see most often and what practices can help avoid them?
Yeah, I can start on this one, I would say I hear a lot of leaders from organizations, they keep mentioning, so will AI replace all the work that we have in our organization or most of the work we have in our organization? I think the answer is that at least at this moment, we shouldn't let AI to really do autonomous, like full autonomous, it's too soon for us. So I think we should really focus on the business, the business purposes and also the tasks, the use case that your organization have. Do you really need AI? What kind of workload that you need AI to solve? Is this for the productivity increase or is this for the timing? Is it for the increase of the delivery or what? So we shouldn't think about, OK, why we need AI, we need to think about where we want to AI. So that's actually a lot of the confusion I hear from the leaders. And another thing is that, you know, the AI is becoming more and more complex, right? In order to properly use AI, you need to have really a good training policy for your organization, such as how do you do effective prompt for your large length model? How do you do the guardrails? How do you do the audits? How do you do the compliance, right? How do you preserve the data security and all that stuff? So before everything, you need to think about those things in the big picture in order to really, you know, to implement AI in your organization.
I agree, I don't I think you've touched on everything I was I would have touched on. The only other thing maybe would be, you know, we talk about EDA specific type tasks. I think there is this sometimes we we kind of jump ahead of ourselves and think, OK, like I can solve this entire problem for me. And so I can just kind of tune out or focus my attention elsewhere. And while there is certainly speed increases that happen with these tools, I think the whole point of this discussion is it doesn't remove our need to be involved. And so I think sometimes it is pitfall saying, OK, like I can just like offload this entire thing to this machine to go do. But really, it's like let's work together with the machine, and that is kind of the best of both worlds. And so when we talk about the practices to help avoid them, I think part of it is kind of what we showed in the demo where it's, hey, instead of instead of here's a prompt, go work on this for 30 minutes and come back to me, it's here's a prompt. You're going to come back to me in a minute with some new suggestions and things we could look at, and I'm going to stay here lockstep the entire time so that I'm an active participant instead of like a distant observer of what's happening. And I think that that practice of being an active participant in these workflows is the thing that makes the difference.
All right, so based on your experience, which EDA and data prep tasks are the priority to automate? I can I can take this first one, I think there's there's the common kind of adage that 80 percent of data work or whatever, whatever high percentage you want to sign is spent in this sort of like EDA data prep stage, and I think a lot of that is just we need to familiarize ourself with what's there. And as much as we all love to to think about and dream about clean data, in a lot of cases, what we're working with is not clean. There's errors, there's anomalies, there's things that need to be considered, there's caveats. And so I think that process and I found that tools like DataBot and other kind of similar things can be really helpful at just kind of understanding the big picture really quickly and understanding where the rough edges are. Right. Like, oh, in the in the in the weather data that we're looking at. Right. Like I was looking at this the other day and DataBot was was like, look, there's there's a lot of entry points for temperature that are negative forty nine point nine. And I don't know if that's because that's just a really common temperature or maybe that's like a missing value code that we need to investigate. Just things that like I might not have necessarily noticed right away, but within five minutes of working with DataBot, we had come up with, OK, here's a couple of things that just don't seem quite right. We saw this really low temperature spike in January or this really high temperature spike in January that doesn't seem typical. So this could represent noise in the data. We also so I think getting that kind of initial high level overview of the data and then also just becoming familiar with every every piece of data has its rough edges that, you know, it's data quality issues. These kinds of tools, I think, can be really helpful at just getting through that quickly so that then we have the right mental model to start asking the right questions. I think sometimes I've been stuck in the past where I spend hours and hours just trying to figure out, OK, what can I and can't I ask this data based on what it does and does not provide? And now we can get to that point. We can get through that part much, much more quickly and get to the point where we're actually asking the meaningful questions because we have a correct understanding of what we're working with.
Yeah, just one comment, yeah, so I think for me also the visualization also is a great example, we can probably do more automation because, you know, those things are a little bit low risk because it's more for us humans to draw the insights. And sometimes, you know, I have an experience before when I was a data scientist. So it's pretty much a similar piece of code that I always borrow from the previous notebook. But those searching really takes time. In the end, you generate pretty much similar style of visualization, right? And it gives you the nice looking stuff. But it takes a lot of time to really have a lot of this decorations on this visuals, right? You have put the labels where you have put this legends and all the color and all that stuff. I think those detailed work doesn't really add value. The more value comes from the visualization itself, the insight that these visualizations show. So I think we can automate those type of work that will save us a lot of time, right? Because in the end, you just want to look at the insights out of those visuals, right?
Yeah, I think that one's huge. And I'm glad you brought that up because that's been one of my favorite use cases of just exactly that. Like I know what visualization I want to make, but I always have to go back and like re-Google the right syntax to get it to look the way that I want so that I can make the decision I'm making and move on. And I found that AI is a tremendous help there where instead of needing to go and like Google and then figure out how to work through that back into my code, I can just say, wait a minute, actually, let's add this is the y-axis and this is the x-axis or whatever, right? Whatever request, it's a very natural, like help me build a visualization the way that I see it in my mind process. And it's, yeah, I completely agree. I think visualization is a huge area where we see this have a high impact.
Excellent. So a lot of the repeatable, more complex EDA tasks can really benefit from this type of approach. All right, let's go to our next question. OK, what human-in-the-loop controls exist to approve or roll back AI-generated steps?
Yeah, I can I can take this one, I thought it's sorry, I'll be brief. I think we kind of saw it in the demo a little bit where it's always asked for permission before code is executed. And then Shun brought up earlier, and I completely agree with this, that we have a lot of the existing sort of scaffolding in place for some of these things, version control and other tools that are actually complement and work very well with some of these AI tools where, OK, I'm going to build out this process, but everything's going to be tracked in version control. There's logs that I can put in place at the bedrock level. So there's kind of this wide range of ways that I'm going to keep track of what's happening, roll back or revert to things as needed.
I think one of the key advantages that you have with data analysis and data science type work is that it's typically very different in some fundamental ways from traditional software engineering, where I'm not like making this huge, I'm not typically in data science work, making this huge refactor of a massive code base and then recompiling that into a program and then checking to see if there's bugs. Instead, I'm like interacting in real time. And so the way that DataBot does this, I didn't have a chance to kind of show this, but the way that we do it there is I might go down some specific branch of exploration of my data. You know, I'm looking at the weather data and now all of a sudden I'm way in the weeds looking at temperature of a specific region. And I just realized, like, this actually wasn't that valuable. Whatever the case is, it's really easy. DataBot gives you this view where I can take a step back and I can look at the total conversation history and go back to any sort of previous point in that conversation and say, actually, like, let's branch, let's let's take a different direction from here. And I can go back to that point and then like rebranch the conversation into a different direction. So you can end up with this like fairly complex tree of things that I've explored, but I can always kind of navigate back through it if I need to unwind or go explore something else.
Yeah, right. I think on that, I see a lot of enterprises already actually implemented their approval workflow, so they have probably different steps to approve for code or for some kind of decisions for the deployment and all that. So everything still there. You just need to add maybe the AI generated code or data transformations or the processings into that workflow, maybe tagging those AI generated assets differently than normally the human generated assets. And also, plus that, I would say there are also some other controls you can do to roll back whatever AI generated steps, the first one being the data versioning. So, for example, Amazon S3, we have the data versioning capability. So even you've done something, you've made some mistakes, you generate a new version of data, you still have this older version to back you up. You can go back to your older data and redo the work that you've done previously. And notebook versioning also, it's pretty standard. So you generate a new notebook, but you find that maybe something is not quite right, you can go back to your old notebook and still do the work. And same thing with the code itself, the Git control, you have all this full Git history, so you can take a look at what has been changed. Yeah, so I think there are a lot of ways you can implement this rollback and human control.
All right. Sounds like we have traceability steps. We have the ability to use data control, code control. So all of these can help us roll back AI generated steps. Let's look at the next question. And I think this is kind of similar, but I think there's one part of it that we may not have addressed. How do you version and track prompt code changes across projects for governance?
I can start a little bit on this. I think for me, I want to relate to the Amazon Bedrock again here. So Amazon Bedrock, for example, it has a unique feature that supports the prompt management with versioning, aliases, and a flow orchestration. So I think in the AI world, prompt is really the key. Prompt means the ability to ask the correct question. So I think the assets you have in the very beginning are those prompts. For different projects, you may have different prompts. So Amazon Bedrock actually allows you to manage all this prompt you have for different projects. It has proper tagging, has proper logging tools behind the scenes, and it can also have guardrails implemented to make sure you have the correct prompt. You have a legal prompt to fit into your business purposes. So I think this prompt management, to me, is actually a key in this now large-length model world.
I agree. I don't have anything further to add on Shun's comments, but I think that all is what we see on our end as well. All right. So I think let's see what other questions we have here. OK. I think that takes us to the end of our panel discussion. So we'll go back to Andrew.
Audience Q&A
Terrific. Thank you all for that great conversation there. It's time for audience questions here. And we have some great questions from the audience. So we'll start with this one for James. And if anyone else has something to add afterwards, please go ahead and do so. So James, this person asks, how does AI assistance for EDA-type tasks differ from traditional software engineering AI assistance?
Yeah, this has been an interesting question to think about. On the Posit side, our background with the RStudio IDE and developing that for a number of years and really working very closely with this community of data scientists or statisticians or, again, whatever kind of title you want to assign them has, I think, informed us very, very well of what their workflows and what those workflows look like. Many of us at the company worked in a former life as data science practitioners. And there's these subtle differences where, and I kind of alluded to it in a previous comment, but you have, in traditional software engineering, you're often working with static code, large code bases, lots of interdependencies. And so you look at a lot of these, kind of the discourse that exists with generative AI assistance in software engineering, and you see things like challenges with context windows because of large code bases and how do we keep AI on task for large tasks that span across multiple function definitions and calls across different things. And then at the end of the day, you're creating the code kind of based off of these static files and then compiling or building that code into some sort of a final solution.
Whereas data science is far more interactive in a way, like we walked through in the demo where I have a real-time interpreter, I'm writing some R code, I'm writing some Python code or whatever framework I want to use, and then I'm getting immediate feedback that code's executing right away. And then I'm responding to that with some new code that's written or I'm editing the previous code and rerunning it again. So it's this very tight loop that's happening. And I think when you look at tools like DataBot and some of these other things, they're built around this idea of immediate execution. I write code, I execute it right away, I look at the results, and then I make a decision. Maybe that decision is to write more code or generate a visualization or ask a different question, but then that loop just keeps repeating itself. And so we recognize that that's the way that most of these data practitioners work, and that's where they find their most value and their highest level of productivity is in this very rapid real-time feedback loop with some sort of active interpreter. And so the tools that we've built and the tools that we see emerging follow that same design philosophy. We write code, execute it immediately, look at results, and then make a decision based on that. Maybe it's a subtle difference, but I think it's a fairly fundamental difference in the way that these tools are designed to work versus a more, quote-unquote, kind of traditional software engineering style agent.
All right, terrific. We'll move on to the next question here. And Sean, this one's, I think, directed towards you, so if you don't mind stepping up here. This person asks, how do organizations choose LLMs in such a fast-developing space, and what does Amazon offer for this?
Sure, Andrew. Yeah, I think that's a great question. I think a lot of, especially now, a lot of customers are facing this kind of issue, is that with so many model providers on the market, which one actually we should choose? I think the short answer from AWS is that our philosophy is the choice plus flexibility. So we believe that there is no single model that will rule them all. So maybe today, there may be one or two or three models that are better than the others, but who knows? Maybe tomorrow or next month, there are going to be another model providers who have a big breakthrough and break this kind of balance again. So we don't know, actually, which model really is the best over the long run. So that's why we actually have this layered offering from AWS, the first one being the Bedrock, which is a single API access to multiple different model, foundation models. So right now, for example, we have Anthropic Cloud models. We have Lama models. We have AWS-owned Nova models, Mistro, Cohere, AI21. Even open AI models now we support. Also, we provide customers with different modalities of models, such as text-based models, image model, audio model, even video models as well. And plus that, we know that some organizations, they may not use those models directly because they have their own business logics, business backgrounds. They want to tune the model by themselves with their own enterprise data. So we allow this enterprise to have this capability by providing the customization service for them, meaning that they can choose a base model, but feed the model with their own data and fine-tune it towards their own business purposes. So also, if you're an organization or startup who have really strong developers or machine learning engineers, you can choose Amazon SageMaker, which provides even more knobs to tune and customize. You can even build the model from the ground up, from the pre-training stage, or you can do the fine-tuning. You can do model distillation, model hosting, or distributor training. So we allow you to have all these different kinds of choices. So yeah, I mean, the answer is that we provide the customers choice they want to have. We don't provide any single model. We just give them all the choices.
All right, fantastic. I think we have time for one final question here, and James, we'll go back to you for this one. This came in during the discussion, so maybe you can expand on some things that you've already mentioned. This person