Human Centered

Better AI Through Social Science

Episode Summary

NBC News correspondent & former CASBS fellow Jacob Ward leads a discussion with Kristian Hammond, Daniel Ho, and Jennifer Logg on the role social sciences must play in developing safer, more effective, and ethical artificial intelligence technologies.

Episode Notes

Jacob Ward

Kristian Hammond

Daniel Ho

Jennifer Logg

CASBS
@CASBSStanford

Social Science for a World in Crisis

Episode Transcription

Narrator: From the Center for Advanced Study in the Behavioral Sciences at Stanford University, this is Human Centered. Artificial intelligence technologies are full of promise, but fraught with perils. Reports abound of AI and machine learning technologies compromising human safety and well-being, unfairly discriminating and amplifying biases, undermining human autonomy, and exacerbating polarization and the spread of misinformation. So how do we better reap the benefits while avoiding these harms? How do we develop trustworthy, fair, and safe technologies? What would such a discipline look like, and what steps will get us there? Today on Human-Centered, another episode from our World in Crisis series. This episode, which originally webcast June 14th, 2022, is titled Building Social Science into the Foundation of AI Practice. Let's join Kristian Hammond, Daniel Ho, and Jennifer Laug in conversation with Jacob Ward, as together they outline a new field of AI practice that extends beyond engineering and embraces social science.

Jake Ward: Hi everybody, my name is Jake Ward. I was a 2018-2019 CASBS fellow. I'm a technology correspondent at NBC News. While I was a CASBS fellow, I had both the material resources given to me and the incredible support of the center, which made possible a book that I wrote, my first, called "The Loop: How Technology is Creating a World Without Choices and How to Fight Back." It's specifically about the subject we're speaking about today, AI, its long-term generational effects, and what hope there is for trying to build something into how we deploy it, what we know about it, what we can teach both the makers of it and those that it is deployed on to make it better aligned with what we want it to do in this world. This is the 20th episode, as I mentioned, of CASBS's webcast series, Social Science for a World in Crisis. I want to acknowledge before we begin the partners for this episode. Those include the Behavioral Science and Policy Association, the Center for Advancing Safety of Machine Intelligence at Northwestern University, the Stanford Institute for Human-Centered AI, the Psychology of Technology Institute, the Regulation, Evaluation, and Governance Lab at Stanford, and the Rockefeller Foundation. Thank you to all of them for making this possible. Now, you have undoubtedly already read the bios of our panelists in the event promo, and we'll link to that in the chat box in a moment here, but you can link from the chat box, by the way, to resources related to today's event, but I just wanna quickly introduce these folks so we can jump right into talking together. Kristian Hammond is the Bill and Kathy Osborn Professor of Computer Science at Northwestern University. Daniel Ho is the William Benjamin Scott and Luna M. Scott Professor of Law at Stanford University and a CASBS Faculty Fellow. And Jennifer Long is an Assistant Professor of Management at Georgetown University. The reason we've drawn these three together is that they are working together in a CASBS-based project on developing a new field of AI practice. That's what we're here to talk about today. And I also wanna mention one of the co-leaders of that project, former CASBS fellow and current research affiliate, Jim Guzza. And we wanna thank him very much for helping curate today's discussion. So I'm going to, just broadly lay out how we're going to proceed today. I'm hoping that this conversation goes ways that I do not anticipate, but what we do anticipate is more or less 4 topics here. We're going to talk a little bit about how we articulate the current problem, what is lacking in the current practice of AI, what each of these panelists hopes, you know, should be addressed, what we think the priorities should be in addressing that problem. What is needed to get that done, and there are— that is a long list, I have no doubt. And I think everybody here is going to have their own take on the idea of creating a thicker conception of how the social scientists should be integrated into the creation of designing and developing hybrid intelligence systems. So, I'm going to just sort of go around the horn here first, and I'm going to start Jennifer, with you, because you are first on my screen, for no other reason than that. This is the computer shaping my thoughts. And just ask a little bit about the problem, right? Let's talk about the problem as you see it. How would you define where we're at right now?

Jennifer Logg: Sure. I'll give it a go. So there's this idea of the last mile problem that started in operations, but has been applied to data analytics. And really, I describe it as a gap existing between producing and using insights from analytics. So a lot of companies are investing more and more in big data. But what's happening is we're processing large amounts of data faster, and it's not totally clear that those insights are being presented to people in a way that's usable, that's understandable, and that's actionable. So I'm focused on that aspect of the discussion that we'll be covering today. And I come from a subfield of psychology called judgment and decision-making. And if anyone has read or heard of the book Thinking Fast and Slow written by Daniel Kahneman, former CASBS fellow and Nobel Econ Prize winner, he is really the grandfather of that field. And the goal of that field is to try to help people make more accurate judgments. So with the rise of big data, I think it's kind of interesting that we have information from a new source that in previous, in historical context, we've mostly relied on other people for advice. Now we can also, um, potentially produce really high quality advice from algorithms. And the field of judgment decision-making is really interested in how can we improve the accuracy of people's judgments? Well, one way is to listen to algorithmic advice, especially if we know that it's high quality in certain domains. But we really need to understand how people kind of respond to algorithmic output before we can understand that broader picture and before we can connect production and utilization of algorithmic insights.

Jake Ward: Krist, let's talk a little bit about your— how you look at this and think about it. What is the— help us understand how your background plays into this and how you're thinking about the problem.

Kristian Hammond: I mean, for me, the problem goes beyond, certainly goes beyond AI to most of computer science. And that is that we have an incredibly powerful and wonderful tool, I mean, the machine. And, but its relationship with us is what is incredibly broken. That is taking the data that it has, the analysis it can do, the advice it can provide, and finding ways in which to effectively interact with human beings has been beyond us. And in many ways, it's because the model within computer science of what a human being is, is a very old school model of the rational man, that we are making decisions based upon trying to optimize the result for ourselves and reduce the costs. But that's not how we make decisions. And from an engineering point of view, I've always been amazed by this, and that is that if you were building a system that had gears and you didn't pay attention to the nature of the gears themselves, that is what their structure was and when they break and when they work well, what the substance, what their substance is. If you didn't pay attention to that, you'd be a bad engineer. But we're not paying attention to human decision-making. We're not paying attention to human behavior. And so we build systems that by their nature might work really well with the rational man model, but with us, lead us to a place where we overtrust We'll overtrust the machine. Sometimes we'll undertrust the machine. Sometimes we'll get ourselves anchored on something the machine has told us. And we are not attending to human beings as human beings from a computer science perspective. And there's this— always this moment where, especially when I talk to my colleagues and they'll see something's gone bad. And they'll say, well, we told people to attend to these problems. We told people to pay attention. If you don't pay attention, you're going to crash when the control of the vehicle gets handed over to you. And it's like, yeah, but people don't pay attention unless you give them a reason to pay attention. And you've built something where you're ignoring the reality of human behavior, of human reasoning. Of human decision-making. And that means you're just a bad engineer. And for me, it's like, let's make people into better— let's make us into better engineers by attending to the realities of humans.

Jake Ward: And Dan, you were one of the very first people to let me ask a bunch of really stupid questions as an early fellow just around, you know, the legal frameworks around this stuff. You were extremely instructive about this. So how are you thinking about it at this point?

Daniel Ho: Yeah, well, Jake, thanks so much for hosting this. And thanks to Margaret and Jim for convening this group to hold these conversations. This is part of a series called Social Science for a World in Crisis. And I think the crisis that we as a group started thinking about is that as AI systems move increasingly into the social world, into social systems, law, healthcare, government, one common theme and fail point has been that machine learning systems are often, in the way that Krist described it, just unaware of the social context in which they're being deployed. One early example is I spent a number of years actually working on health inspections and food safety. And the kinds of papers that were being put out at that point of time were, let's go and build a machine learning system to predict where sort of food safety violations were likely to occur without recognizing that actually the way that, for instance, inspectors can be assigned are based on zip codes and areas. And the big policy concern within food safety has actually been that the principal driver of variation in health code violations is actually the stringency of the inspector. And turns out when you build out a machine learning system to predict where the— where there's a high likelihood of violations, and if you reallocate inspectors based on that forecast, you're going to be doing the exact opposite of what any expert within that field will tell you is the right thing to do. Because what they're worried about is not the place where Bruce is oversighting violations. They're worried about the place where the health code is being systematically under-enforced. And we see that over and over again. The Patent and Trademark Office built out a technically— quite advanced system using, at that point in time, fairly novel forms of TF-IDF coding some time back to help their 9,000 patent examiners identify forms of priority. And it was technically a very good system for the time that it was built out. It turns out when they piloted it, the only people that could actually take advantage of it were the patent examiners that had computer science experience. And that's the problem that we're trying to tackle here of when this misfires, how do we think about the role of the social sciences to develop forms of hybrid expertise to understand what problems are worth solving, what problems are not worth solving, and what is actually technically achievable and what should be changed about the technical side to actually have this complementarity between human and machine intelligence.

Jake Ward: So let me throw a little bit of recent news into this. I can't help it because of my job, but basically last week we saw a Google engineer named Blake Lemoine put on administrative leave after spending— he'd spent about 7 years at Google working on AI. And he had volunteered to be part of a test group that looked at something called LaMDA, which is Google's auto-chat-generating engine. And it's the, you know, it analyzes trillions of, you know, words and sentences and communications off the internet and is very, very good then at mimicking conversation. And Blake Lemoine basically had what some would call sort of a spiritual experience, basically speaking to this chatbot, and he reported to his superiors, "I think this thing is sentient, and we need to deal with that." And they told him, "No, it's not sentient. It's just doing a really good imitation of sentience and forget about it." And he persevered anyway, and he went public with his concerns. and is now on administrative leave. And so the AI community has been basically talking all week about— I'm sorry to spring this on the three of you, because I'm sure you're thinking about better things than this, but to me and to so many people in the sort of public sphere thinking about AI and how we react to it, and this is what my book is all about, like what happens when we get to a place where AI is not all that smart, but does a fantastic impersonation of being smart. What's that going to do to us, right? And so as we pivot here into this next subject, which is is what are we going to do about the problem that you guys have articulated so well just now. I wonder what reaction you have to, you know, a 7-year engineering veteran of Google basically coming to believe that the chatbot he's being exposed to is really thinking for itself against the expert opinion of basically everybody in the field. and being so convinced by it that he then puts his career on the line, you know, gets blown out of his job for it. I find myself looking at that and thinking, okay, well, if somebody with his experience really can't tell the difference— I mean, it's— I don't know, this isn't quite the Turing test. This is something even beyond that, you know. And I just wonder if any of you have a reaction to that episode and what it tells us about where we are at in just how little we seem to understand about how these systems are built and how broadly we seem to assume that they can do way more than they're designed for. Anybody want to take that one?

Daniel Ho: And I guess I'll offer one, uh, sort of a bit of perspective, which is that, uh, in a sense, this particular episode is not that new. If you think about, uh, Joseph Weizenbaum's ELIZA system that was built out in the 1960s, which was drawing on none of the large language modeling that underpin LaMDA. What we saw is it was a kind of, you know, very simple chatbot meant to kind of reflect back people's sort of comments as they were typing in. And there were quite a number of individuals who were just captivated by the system, much in the same way that this—

Jake Ward: His secretary famously told him, I need you to leave the room because I'm going to confide in this thing. He dressed it up as a therapist and she said, I can't have you in the room, I'm going to have such personal— He ended up quitting the field. He was so terrified by what it ended up doing.

Daniel Ho: Yeah. In a sense, I think it's worth distinguishing how much we've made leaps and bounds in terms of the development of large language models and deep learning. That human frailty is not something new of an individual to be kind of captivated by a system and engage in forms of anthropomorphization when there's no question that Eliza's system was quite simple. So I think it's worth keeping that kind of historical perspective in mind, but I don't know if Krist or Jen have other thoughts on this.

Jake Ward: Like, what is our solution here? How can a thicker conception of how the social sciences could be built into the creation of these systems you know, keep us from having situations like this.

Kristian Hammond: So Weizenbaum didn't quit the field because he was afraid of the systems. He quit the field because he was afraid of us, that we over-attribute like crazy. And my favorite example of this is one of the very first movies ever made was a full-frontal shot of a train coming towards you. And you would— in the early days of film, you would— they would show that film And people would stand up and run away because they only had the interpretation of a moving thing coming towards them as a— it's a train coming towards me, and I'm going to respond to it. We really only have— we don't have that many models of intelligence. And so when something seems compelling, we think that's intelligent. But with large— especially with large language models, they're incredibly good at— they're incredibly good at the structure of language. They're incredibly good at answering questions about things that they've read about in some sense, but they don't incorporate today. They don't incorporate our current lives. So, you know, you can say, oh, Chicago is a great city, and the system will know city. Jake is in a car and he's driving to— I have no idea. And it can answer the— they can answer the questions that are less ephemeral. But anything ephemeral, it's hard to get that information into them. And that's kind of the hallmark of AI, or of intelligence and communication. Communication is supposed to give you new information, but these things aren't really communicators. They're language mockers. They're language mimics, which is great, but they're not complete. There is this— there is a desire for us when we see something that goes beyond just being a machine, there's a desire for us to want it to be more. And so we imbue it with that. But that's not— it's just not the case. If you want to call something sentient, then you've got to really have a firm definition of what sentience means. You have to have a way to test it. You have to wait to understand it. Um, and it's not just a, I talked to it and it really was compelling, because, uh, there are a lot of things that can be compelling, uh, and it, uh, they're not all sentient.

Jake Ward: Jen, do you have any thoughts about this?

Jennifer Logg: This reminds me of, um, something that happened more recently with Google in the last 5 years with Google Duplex. So the idea similar to Apple Siri, as an assistant. And as a user, you could say, "I would like to schedule a haircut," or "I would like to make a reservation at a restaurant," and Google Duplex, this computer application within Google Assistant, would call the business establishment and make that happen. And that came out, and Google had talked about it at their annual meeting. You can find it on YouTube. And headlines started popping up that people really didn't like the idea, the folks who are working at the businesses, thinking that they were talking to a person because they had anthropomorphized, the voice sounded very similar. It was very human-like. And it was a pretty short amount of time when this went public. And then there was backlash about people feeling like they were deceived. And so the fix for that problem was to provide disclosure that once Google Duplex started interacting with the business, where obviously the businesses were having a tough time telling the difference between Jake calling to place a reservation and Google Duplex, that now Google Duplex would say, this is an automated system. But I think it's interesting because when I've talked to people, no one's using that program right now.

Jake Ward: Let me ask you this. So, you know, I hear because I'm in the business of sort of standing in between the expertise that you have and the broad public that is thinking about it. And I hear oftentimes, certainly from people within the companies that are making these tools and these systems, that the public needs to somehow be better educated so that they understand it better. To me, one of the most compelling things about this instance is this is someone who is arguably as well-educated as one could possibly be in how AI systems work without being directly in the working group that built it. And he still had this transformative experience and was entirely convinced that this thing is speaking to him as a sentient thing. His goodbye letter to Google talked about needing, please, you know, it is a sensitive child, essentially, please be good to it. This guy had fully, you know, was really having a powerful experience with this thing. And so I'm thinking, okay, if he doesn't know, then the idea that we're going to somehow create sort of a basic, you know, maybe elementary school core curriculum that's going to teach people the difference, I don't know that I am convinced that that can work. And I wonder what you guys think now, thinking about how the practitioners of this are taught. What are we looking at here? I can imagine, you know, taking into account what Jen studies, taking into account what Krist, you were just saying, taking into account what you were talking about, Dan, you know, do we, you know, somehow take a measure of human reactions to these systems and say, okay, we're not allowed to cross this uncanny valley that we know exists? Or, you know, does it have to do with pulling back on this? I mean, I guess I'm wondering, you know, do we begin to use human reactions to these systems, defining them and measuring them better than we do now, as a throttle or something? You know what I'm saying? Like, as a governor on how fast and effective and authentic these things wind up seeming to us. You know what I mean? So what if we were going to create a curriculum, create a thicker conception of stuff to get in the way of what is happening even with experienced Google engineers? What can we insert into the system?

Kristian Hammond: I mean, I hear you. I don't know if that's what's needed. I mean, simply because you're an engineer doesn't mean you don't get to have spiritual experiences, and he clearly had a spiritual experience. That's very different than, you know, doing a hardcore evaluation of the nature of the mechanism, the causality underlying things, and the reality of the situation. And so there's that. I don't— it's interesting when, you know, the tech companies will say things like, well, you know, the public needs to understand these things better. And I actually see that always as a way of saying, we don't know how to explain these things well. Would someone else please do this for us? And it's like, no, it's your responsibility. You move something into the world, you have to tell— you have to be able to tell us the length and breadth of what it is and what it can do. And if you don't, as a technologist, you failed. And I mean, I think that we've seen time and time again, it's like, no, you attend to the audience, you attend to who's going to use this, you attend to what they need to know. And if you attend to that and you make it part of your job and your mission as technologists, as engineers, as scientists to actually embed your understanding of how things are going to be used in the system itself, so that when people are using it, they're not confused and they're not overwhelmed and they're not misunderstanding things, then you've done your job. But in order to do that, you have to understand the people you're working with. You have to understand the target. You have to understand the audience. And that, I think, is what's important now.

Jake Ward: Jen, you, you know, you study human decision-making, how people respond to this stuff. Tell us a little bit about your impression of this, right? Do we— is it about educating people? Is it about educating the practitioners so that they understand what they're wielding? What are we talking about here?

Jennifer Logg: So as I've been using, um, experiments to test how do people respond to identical advice when they think it comes from an algorithm versus a person, in parallel, I actually started developing a class. I created a class, um, in the last year of my postdoc called The Psychology of Big Data. And that, I think, is pretty relevant to our discussion. It came out of discussions I'd had with C-suite folks, people who were in exec ed. And they said, oh, we would like to know what 5 questions we can ask our data analytics team. So it's a little bit different from what we've talked about. It's decision makers. Who want to know what types of information should we be eliciting, or how should we be trying to pose these questions to our analytics team so that we could get the most out of the resources we're putting into those teams. And then really a month later, I crossed the river and gave a guest lecture in a computer science department, and the computer science master's and PhD students In a less direct way, wanted to know how they could communicate their findings to different types of audiences. So it just became very apparent to me that we have these silos, not only in organizations, but on our campuses, where we have decision makers who want to— they don't know how to interact with engineers and computer scientists. And we have the engineers and computer scientists who are not really sure what information is most useful to those decision makers. So I really think kind of just chipping away at the very tip of this iceberg in parallel to my research, I thought, I just want these people in the same room. So that's why I'm so excited that we're all in the same room today. And we've been working in this working group because I think that's the start of really important discussions. And it's hard to predict things that would fall out of that, but the focus of my class has been similar to the idea of, you know, increasing financial literacy. You can see it as I'm trying to increase evidence-based literacy. So when you're receiving information, what kind of questions should you ask about it? You should ask, what's the sample size? What is— tell me more about the data that went into this. What are the variables that were collected? Is it possible that there's proxy variables, so variables that are strongly correlated with demographics that could help explain potentially some biases in any output? So that at a very basic level is kind of where I'm starting to try to address this problem. And I'm curious to hear Dan's thoughts as well.

Jake Ward: Yeah. Dan, let me throw something at you that I encountered when I was writing my book, which is looking at the ways different societies define harm and how they then look at how technology either exacerbates that harm or can in some way be harnessed to pull people back from that harm. So one example, for instance, is gambling. In the United States, we have an extremely wide-open legal framework around gambling, and it's becoming more and more wide open, right? Online sports betting is now legal in 31 states. California is probably going to legalize it this coming November. But in Canada, where gambling is regulated by the government, there are casino systems there that basically use, as all modern casino systems do, a swipe card. You know, you fill a card with money and then you use that debit card over the course of your time gambling. And what they do in this particular series of casinos is they use a behavioral algorithm, a predictive algorithm to say, Okay, this person has now lost control of themselves. They've gone off the measurable barometer of self-control, and they are going to immolate themselves financially unless we step in. At that point, that casino operator freezes the card, and that person cannot gamble again. The pit boss is instructed to come get them, escort them off the gaming floor, get them a cup of coffee. There's a whole intervention that is built in, and it's all technologically facilitated. And that is, I think, because we have, at least Canada has, agreed that there is a point past which you should not go as a recipient of this stimuli. And I think about myself, I mean, I literally study addictive technology business models for a living, and I get lost in TikTok for hours and hours and hours on end to the point where literally TikTok has to pop up and say, "Hey, buddy, you should go to bed." There's a little video that comes up and says, "It's time for you to go to bed now," which means that somewhere in that system, they know the point at which I have lost control of myself. And yet, I think in this country, we don't really have a clear framework of harms around this stuff. And so, what both Jen and Kristian are talking about makes a lot of sense to me. But, when I think about a world, a wide-open world of deepfakes and, you know, one audience question is asking, you know, what does curriculum matter? If we're being fooled as badly as we are by some of these systems. And I guess I'm wondering, where do you think we are as a society in even beginning to define the possible harms of this stuff, much less come up with some sort of framework for saying, okay, that's too much, this needs to change, and here is how practitioners need to think about the harms that their technology can do?

Daniel Ho: It's a really great question, and I think You know, so many of the current policy interventions focus on this notion of algorithmic impact assessments, really to very forthrightly, before deploying a product, actually thinking about all the potential downstream harms. I think this is sort of part of the reason, Jake, why this group is really, you know, part of the series on Social Science for a World in Crisis, because Many of us in this working group think the social sciences actually have a really important role to play in actually understanding those potential downstream consequences. So think of the recent New Yorker piece that was written that covered this whole question of the impact of social media algorithms on polarization. And part of what comes across in that piece is gosh, it's actually really difficult to isolate the causal effect of ranking algorithms versus a whole bunch of other structural forces that have contributed to polarization, right? And I think that's one of the interesting kind of challenges here. If you, to take your example of gambling, if you could generate the same levels of kind of addictive compulsion around gambling in non-algorithmic ways, right? Is it— you know, we need to know how much is sort of the deployment of algorithms really exacerbating that kind of harm, or is what you're really getting at, Jake, the underlying different conceptions between Canada and the United States in terms of how we treat gambling addiction, whether algorithmically facilitated or not, whether it's lights, cheap host— cheap hotel rooms, and the whole other sort of architectural kind of design choices that try to keep people at the gambling table. And I guess just to let me take that back to kind of the— what this group really has been focusing on has been this notion that one of the most common pitfalls has been that machine learning has tended to really fixate on optimizing model accuracy and not accounting for the human-computer sort of interaction. And, you know, that's fundamentally to us the kind of question about the role of the social sciences. And one of the things that this group has iterated towards is borrowing a really a page out of the playbook by James Landay here at Stanford, who kind of adopted, for instance, this notion of design patterns of, can we think about best practices, encapsulating that in a set of reusable templates to force people to engage and ask those kinds of questions of what are the downstream consequences, what are the kinds of recurring problems we've seen in terms of bias and other forms of algorithmic harms, and is there a way to think about injecting that into the design process that may actually spot these issues earlier on. Or another way to think about it is how do we convert algorithmic impact assessments, which are pretty superficial at this stage, and actually sort of imbue those with the kinds of social scientific insights that someone like Jen and many others on the group can provide to think about these issues like Sandy.

Jake Ward: Do we think, you know, that there is— I'm thinking here about the sort of the spectrum. I mean, over and over again, I basically bump into in both writing my book and in my day-to-day work as a journalist at NBC, I bump into business models and the same, in some cases, literally the same off-the-shelf, you know, AI systems being deployed on these wildly different products and audiences. And yet, you know, so there'll be a spectrum, for instance, from— I did some reporting for the book and on air about a whole category of companies called social casino companies that have created such effective predictive systems for finding and ensnaring people that, that are prone to gambling that they will pay to gamble with virtual currency and with no hope of winning any money back. So it is literally the definition of a loser's game. There is no other outcome other than giving more money in, and yet I interviewed people, and according to the class action lawsuits that are coming out of this, there are thousands of people across the United States who have lost 5 figures and sometimes 6 figures to this. And these are people who cannot afford that money.

Kristian Hammond: That—

Jake Ward: those design elements, the tribal excitement of being in there, the predictive algorithms that tell you who's going to be susceptible to it, all of these things are also being deployed by companies like, for instance, I'm thinking about Noom, the, the, um, diet management system that helps people eat more responsibly. Peloton, which, you know, welcomes— welcome to the Peloton family, they're told every time as they sit down. And the same sort of gamification and tribalism and the rest of it that is, you know, and the predictive stuff is all used to market those things. Basically, the same tools can be used to— can be deployed on people for predatory reasons, and for reasons that can benefit their health and give them longer lives with their loved ones. And I guess I wonder, what role do you think a working group like yours can have in helping? I mean, can we, do we have any hope of pre-embedding our values into these systems through the, you know, something like the work that you are doing? Or this is what something that's one of our audience members is asking, Or do we need to, I don't know, somehow line up some cultural expectations? You know, this is again to the educating the public question, right? But like, can we pre-embed our values when it seems to me our values are so squishy, we don't— we've just barely got our heads around them as a society anyway. Can we pre-embed our values into systems like these through educating the people who make them?

Kristian Hammond: Yes, I think we can. We can certainly embed values into systems. But I think the important part here, and one of the reasons why this group exists, is that in order to do that, you actually have to understand the causation of the system itself, including the human element. And that is that, you know, we can, we can take a look at the world of dark patterns, including, you know, systems that are designed to be, you know, designed to be addictive in nature through engagement models. And pull the science out. That is, forget about the evaluation of what these things are, are making you do, but figure out what the dynamic is in terms of how to interact with an individual to move them in one direction or another. And once you've got that, then you actually have two things in your hands. One is that you can say, if you, if you are designing a system and here is your goal, here is how you go about doing it. And the other is, once you've designed a system and we can look at those goals, then we can actually say, look, you've built this thing for this goal. And if this goal is not a positive— it's not positive for your people, is not positive for your customers, is not positive for society, then we get to hang you up on that because this is the pattern and you know this pattern works. And we can get— we can help you lose weight, um, but we can also get you addicted to, uh, you know, a, a, an ice cream cone every morning, you know. It was like, we can go whatever direction— you can go whatever direction you want, but we will be in a place where we know enough about the causation and we know enough about the patterns that we can call you out on it. Um, and the system— the notion is not that the, the system it has the values already embedded in it. It's that, it's that you have to understand the causation, you have to understand how it works in order to even begin to think about evaluation or doing that embedding.

Jake Ward: Jen, you know, I had an interview at one point with a guy who was a senior person on a team that was doing basically credit scores. I can't give away the name of the company, but he was doing credit scores credit risk assessment, essentially, for making loans to people using automated systems. And I asked him, you know, I had heard secondhand that his company had been trying to get on the right side of history, trying to correct historical inequities when it came to the racial patterns in this kind of loan making. And I asked him, you know, what is the, you know, did you do that? You know, how was that? And he said, "We took a swing at it, but we decided we didn't want to do that. We didn't feel it was our responsibility to do that. Our responsibility is to make the best possible product for our shareholders and get it out there. And if anything, we would have to put our thumb on the scale to such an extent that it would be," he said, "unethical for us to correct these historical patterns," basically. And so he was just sort of going ahead in this kind of, this way of sort of saying, This is how the world is, and we just make the best possible product to work in that world. And I guess I'm wondering, you know, if you had a crack at that guy early in his education, right, what do you think you could have changed about his perspective or thought about in this perspective? Honestly, I'd be interested in anybody's reaction to this, but, Jenna, I'm just super curious what you— how you react to that.

Jennifer Logg: Yeah, I'm definitely not in the practice of trying to get people to change their perspective because the other program of research that I do is overconfidence and people fail to update their beliefs as much as a rational actor would. So I need to preface it with that. I mean, it's a really big problem and I see the problem as bias in, bias out, garbage in, garbage out. And if organizations are not willing to look at their, their historical data to see where mistakes were made before, either biased against certain demographic groups or otherwise, that's another problem in and of itself. That's like, hopefully there's some researchers looking directly at that. It's a little bit outside of my field of inquiry, But I often use Amazon hiring as an example. There was a lot of media coverage that Amazon was using 500 models to predict who the best performers would be in the applicant group that they had to determine who they'd hire. And over and over again, these 500 models were only hiring men and not hiring women. And what could have happened was that Amazon didn't share any of this and no one would know, and they would be doing hiring relying on human judgment. But instead, the media covered this and Amazon shared that what they were able to pinpoint in the historically biased data, which was being fed through the algorithms, and the algorithms were doing their job, of magnifying the tool that it is, um, that there were certain words that were almost perfectly correlated with gender. And these words didn't really matter if you look at them. You would say— saying that you captured value shouldn't really be the make or break of you getting hired at a company if you had worded that sentiment in a different way. And so I share those words with my students and I say this is the difference between getting hired or not with those 500 models, but you can use this now as an applicant, but also the organization can learn where human judgment had produced bias in the first place before it ever got fed into the algorithm. So instead of blaming the algorithm, I think your anecdote really highlights We need the willingness of people in organizations in positions of power who are willing to take a good hard look at the historical data that they have collected. Some might not even have the data collected in a useful way to look at it. So just the first step of looking at the historical data, putting it together in a way where you can find these patterns in the first place. Um, I guess I'm for using algorithms as magnifying glasses. So the algorithm is just going to magnify the input data that's fed to it. And if you really want to know how biased or neutral or high or low quality your data is, magnify it with algorithms so you can go back and rectify it. But that leads to a larger discussion of who is willing to do this. And talking to undergrad students, they're really excited about that. And I hope that with generational changes, more people are willing to kind of take a stand in their company and try to make these changes.

Jake Ward: Do you think there's a way to build that into the early education of these folks before they reach those companies? Because that's the thing that I bump into time and again, right? Speaking to people inside the companies, they talk about, you know, having such limited input into the strategic decision-making around these products. You know, in some cases they can lose their job for really asking how they're going to be deployed. You know, like raising a stink really is, is career suicide a lot of the time inside a lot of these companies. And so I'm wondering, and I'm opening this up to any of you, you know, I mean, what is needed to solve part of the problem that we're talking about here in terms of working the social sciences into this stuff? Is there some early I don't know, I think about like, what would life be like if there was no such thing as a Hippocratic oath, right? And physicians were incentivized by the profits of the hospital to keep you sick over time, you know, just sick enough that you keep coming back, you know what I mean? Like, we have some fairly sophisticated public, sort of, you know, some institutional ideas about how to keep people motivated properly to do things right. Now, of course, we could argue for hours about whether the healthcare system is functioning, but at the very least, I think we can agree, right, the Hippocratic oath is really had an effect. Is there something like that? What are we talking about in terms of specifically the skills that we could teach a set of engineers or things we could work into their education that might change what we're talking about here? Anybody?

Kristian Hammond: I think that right now, if you talk to undergraduates in computer science, they are hungry. They are hungry for ethics. Because they've had that moment where they realize, how much impact the field they are working in is having on the world. And they have no interest in having that impact be negative. And so they want to understand the ethics. And they want to understand how to actually control the work that they do. And the thing that's deadly right now is that if you interview machine learning engineers and you ask them about their data, They'll essentially tell you the truth, and that is, I just was handed the data. There's somebody else who's gathering and collecting, gathering, collecting, managing, merging, making decisions about what features are here, and they're handing it to me, and my job is to optimize the algorithm. And then you ask them, well, what's it going to be used for? It's like, I don't even know that at all. All I know— and that's the problem, is that there's not this idea of of— there's a work stream that is going to give you a product at the end, and you have to, at every single stage of that work stream, you have to have people who are aware of where it came from and where it's going. And in fact, I was trying to explain this to a class that it's— my analogy is the really old movie Spartacus, where at the very end, the Roman soldiers are trying to find Spartacus and Kirk Douglas stands up and he goes, I'm Spartacus, and they're going to crucify him. And one by one, everybody else stands up and says, I'm Spartacus, I'm Spartacus. And that's, in fact, what we have to teach our engineers, that you stand up and you say, no, no, no, I really need to know where it's coming from. I need to know what you've done, and I need to know where it's going, and I need to know the mechanisms at each of these stages. And it's hard to do that right now. But if you— if we actually turn that into, for me, a part of the culture of engineering, that you pull your face away from the screen and look to the world, and you see how you're going to have impact, and it's your job to make sure the impact isn't negative. And we can give you example after example after example of all the bad things that we have allowed to happen because no one pulled their head up. And now it's time to teach everybody to pull their head up and to look to impact.

Daniel Ho: Jake, if I could kind of just take you back to the example that you had of kind of credit risk scoring, I just want to actually also say the law is going to play a role here, right? So the CFPB recently announced that it's going to initiate the process for having nondiscrimination be part of a quality control process. And so I think that would be an important kind of backstop to ensure that there aren't kind of discriminatory products being put out there. But in terms of the broader conversation that we've been having, Jake, you focused us a lot on education as an intervention. You mentioned the notion of the Hippocratic oath. But as many colleagues would remind us, there's no such thing as a professional license for a software engineer. It is not a threat to take away Krist's ACM card, if he even has a physical card, right? That's just no threat in terms of what he's able to do as an engineer. And we have lots of interventions, potential interventions being floated, like checklists, ethics review, notions of contestability. And the one we've kind of been talking about a little bit within this group the notion of design patterns. I think fundamentally what that calls for is a kind of social science that can really understand how AI is embedded and that rigorously evaluates these interventions and their effectiveness in mitigating the kinds of harms that we've been talking about. I'll give you one example from kind of the criminal risk assessment context where much of the algorithmic fairness literature has been obsessed with technical solutions equalized odds, equalization of false positive rates, these kinds of things. But, you know, one kind of consistent finding throughout the research here has been that if you plot the rearrest rate, the sort of one measure of sort of recidivism against the risk score, there's a pretty pronounced gender difference. That is, there is, you know, the scores are, well, are, you know, depending on the score score, they can— they, they're positively— there's a positive relationship, but there's a kind of scalar difference between men and women. And here's where the social science question becomes really important, because one, that leads some people to say, oh, it would be wrong to have different risk assessment scores by gender because we would be over-incarcerating women because they're less likely to be rearrested. But the But that sort of fails to ask the question, why do we see that difference between men and women? And what's actually happening is we have police bias for the same set of factual circumstances, say a barroom brawl. Jen is less likely to get arrested by police than Krist is. Then actually normatively, there's a very compelling argument to be made that actually you're just propagating bias by allowing for that gender adjustment in risk assessment. Scores? Those are the kinds of questions I think are really important to ask and, you know, have not been as much a part of the conversation of algorithmic fairness as they should be. Because as social scientists, you try to really understand and dissect the sources of bias, which is often going to be a much better method to allow you to actually mitigate it than purely technical off-the-shelf solutions.

Jake Ward: Jen, I see you nodding along with this. Do you have anything you want to throw in on that?

Jennifer Logg: I just think that he more eloquently said what I was trying to get to. So we're all focused on the algorithm, but that also to some extent has some consequences of we're relieving organizations and the people who are making these algorithms of responsibility because they're saying, oh, the algorithm is bad. And I think Dan just kind of nailed that point for me.

Jake Ward: It's very interesting because I think we're stuck in this really difficult moment and I think it makes what you guys are trying to do so important that we're, you know, Krist, you had, when we first spoke long before this session, you were talking about Underwriters Laboratories and I'm going to take your story away here and I'm sorry, but you know, that it was this original body that sort of created the idea that like, you know, wires are catching on fire and that's a bad outcome and we have to create some safety standards around it. I feel like we're at this moment right now where we don't know how the wires work. The people who make the wires have no autonomy or agency in communicating with the people who own the wire companies. We can't even really agree that the wires catching on fire are— what kind of bad that is, and how to quantify the negativity of that outcome. Countries like Canada say wires burning are terrible and that's bad. Here in the United States, we are legalizing burning wires all over the place. You know what I mean? It's a very complicated moment we are in. And so maybe just to wrap up here in this session, as we're thinking about what you guys are doing, maybe we could go around and just talk a little bit about this idea that we have spoken about as a group before of just a thicker conception of how the social sciences can be integrated. You know, we have commenters in the audience saying, you know, well, can't we just, you know, emphasize positive human traits, empathy and intellectual curiosity, you know, can we just inculcate, you know, goodness into people, right? But we're talking here, I think, about trying to formalize something beyond just sort of, you know, filtering out, you know, bad people. We have to— there's something— something has to be built into the system here, and it's not quite clear what that is. Maybe one thing we could talk about a little bit is, if anybody wants to touch on this, is this difference between AI, as we've been talking about it, right, and this emphasis on the idea of the algorithm, and what I know you guys are thinking about as this sort of hybrid human-machine intelligence that you guys have been thinking about and talking about. Krist, do you want to jump in on any of this?

Kristian Hammond: Sure. I mean, the notion here is is to say, look, when we are— when we're building systems, let's build systems, not build one device, but actually say we're going to build devices that work together. And one side, one of those devices is going to be people, human beings. And once we say that there's got to be that linkage, that linkage has got to be there, then we have to start considering, well, what is the model that humans have of the thing they're interacting with? What is the outcome of the whole system, not just the algorithm? How well do we trust? How can we collaborate? What needs to be passed back and forth? And you start having those series of questions just because you say, no, they've got to work together. And those are the questions that once you start answering them, give us the information about what we need to do in order to optimize the outcome for the human, not just the outcome of the system itself, of the core algorithm itself. It doesn't guarantee ethics, it doesn't guarantee positive outcomes, but it gives us the substrate of information about how to manage it, that will give us the beginnings of being able to say, look, this is— we're going to— we might need some regulation here. We might need some cultural change here. But the outcome is going to be that we know how— where the knobs and levers are. And, you know, once you've figured out how to get people to gamble, how to get them addicted to gambling, then why don't you figure out how to get them addicted to exercise? I mean, it's like the physics, the physics of it all is what we need first, and then we can decide how to apply that physics to make the world better.

Jake Ward: Great. Jen, tell us a little bit about how you think about this. How can we thicken up our concept of the social sciences and their role here?

Jennifer Logg: Yeah, so something that's come up in conversation with Jim, who you mentioned at the beginning, who's been leading our working group, with Krist has been the idea of decision architecture. I'll just kind of briefly talk about that. We've talked about shaping human behavior. The idea is that any information you present to people— there's not really a neutral way to present information. You're going to have your own perspective. You are going to present it in a certain way, and there's always an alternative to how you could present that. And it struck me at the time, Elke Weber, who's a professor at Columbia, she's now at Princeton, had been doing work related to this and she had been fielding questions and she referred to decision architecture like tools, like we have been discussing here. And she had mentioned, you know, people could use tools for lots of different, goals, good or bad goals. And what struck me about what she said was really taking the perspective— I study people at the individual level, like psychological human cognition at the individual level. Um, if we look at decision architecture and understanding human biases, uh, documented within the judgment decision-making literature and behavioral econ, The thing that's in common there is awareness. So now we know how these tools can be used in one way or another. We know to some extent where some biases are more likely to pop up or what they look like, at least. So at least we can identify them. And I think the strange thing about biases is that you can have a full class on biases, yet we're still going to fall prey to many of them. Just because when we are trying to make decisions, we're often in a time crunch. And so our brains are wired in such a way, time is really the element that starts kind of encouraging the human brain to rely on heuristics such that we over-rely and they turn into biases if they're over-applied or misapplied, right? So if we rely on these shortcuts Too much, they turn into biases. But I think what I've seen from decision architecture and people seeing about biases, people can now name them and they'll say, oh, I noticed that in this negotiation, someone had mentioned a number early on and I anchored to that and I insufficiently adjusted. And I think once we get into that world where people can identify and talk about kind of the psychological responses to the stimuli that is our world, we're in a much better position, right? It sounds a little less scary to me, at least, if we can get people to identify and at least partly understand, even if we will be falling prey to biases, especially when we're in a time crunch. And I hope that we can do that thinking about algorithms as tools as well, and thinking about the different possible outcomes. I, I try to articulate it in a book chapter that I wrote, breaking up preparing to build the algorithm, building the algorithm, and interpreting output from the algorithm. So not just from the end user's perspective, but kind of from the creator's perspective as well. So I think if both sides can think through that more and identify kind of pressure points or points where errors might occur, biases might crop up, having a conversation is the first step.

Jake Ward: That's great. Good. Dan, what do you think?

Daniel Ho: Yeah, I mean, I think, Jake, it's been so interesting to hear your questions because you've pushed us a lot on like very affirmative harms. And I think this is really, I guess one way I might try to divide this is to think about inadvertent harms, which we're seeing a lot of as well, that I think are more easily targeted through things like design patterns, impact assessments, forms of ethics review. But as you've written about so nicely in your work, there are many instances where there's actual intentionality behind the way that products are designed. And, you know, Oliver Wendell Holmes had the sort of phrase of how the law should take the bad man into account, the proverbial Holmesian bad man. Because for every, you know, Microsoft or AWS that issued a moratorium for the use of facial recognition technology, particularly for police, there are the Clearview AIs of the world. And at that point, you know, the design pattern is not going to help mitigate those kinds of more sort of intentional moves that are being made. And I guess one thing that I thought came across really nicely from your commentary is that, you know, there are instances where when we're concerned about those kinds of harms, be it in the criminal justice system, be it in the welfare system, be it in gambling, you know, one way in which to think about that is that at least the kind of value judgment as to the incentive structure that we allow may not be one that requires you to have kind of full technical command of how algorithmic systems are being made. It's much more of a core democratic question of how much do you want for there to be more forms of input and accountability in the criminal justice system, in the social welfare system, or in the kinds of red lines we may wanna draw as to forms of manipulation of individuals when it comes to gambling addiction.

Jake Ward: Well, I really appreciate all three of you and your broader panel for coming together. I think one of the lessons of my time at CASBS and writing the book is just how much we are going to need such a broad set of inputs and perspectives to be thinking about this. This stuff is moving so fast and the companies deploying it are, I think, so poorly incentivized to do the kind of thinking you guys are thinking, are doing in this case. And so I just think it's very holy work and very important. So I want to—

Daniel Ho: If I could add one last thing, I think that one of the commenters just noted, just one thing that I think is worth noting here is The importance of forms of external oversight.

Jake Ward: Yeah, we haven't touched that at all. Yeah, go ahead.

Daniel Ho: It's institutions. Yeah, you know, as we're starting to broaden this kind of conversation, right, one of the commenters notes the really important work by organizations like the Algorithmic Justice League or the ACLU in providing forms of checks and balances in the system. I think that has been hugely, hugely critical to this kind of system. Sorry to—

Jake Ward: No, no, you're absolutely right. And I think that what is so important about what you are describing is that I think the attitude in the United States is wild west around this stuff, especially as compared to Europe. And arguably, you could argue that in Europe, where AI is going to be regulated in all sorts of ways, one of the paradigms they're talking a lot about is just the whole notion of addictive technology being something they want to outlaw, but we live in the world of Peloton and Noom. So is that what we're talking about, right? As Jen has articulated over and over again, right? There are some very smart, you know, ways of nudging behavior in a positive direction. You know, Kristian was talking about that as well. So like, we are in such a raw moment for this stuff, and I think it is so powerful what you guys are doing. So again, I just want to acknowledge the three of you. So Kristian Hammond of Northwestern, Dan Ho of CASBS at Stanford, and Jennifer Logue of Georgetown. Thank you so much for being here. I also want to thank this event's co-sponsors, the Neuroscience and Policy Association, the Center for Advancing Safety and Machine Intelligence at Northwestern, Western, the Stanford Institute for Human-Centered AI, the Psychology of Technology Institute, the Regulation, Evaluation, and Governance Lab at Stanford, and the Rockefeller Foundation. And I also want to just make sure everybody gets a heads up here about the next episode of Social Science for a World in Crisis, the 21st episode, and also some information about how to view previous episodes. All of that is going to pop up on your screen in just a few seconds. So thank you again, the three of you, for the work you're doing and for being here today. Thank you everybody for taking some time and spending it with us. Thanks.

Kristian Hammond: Thanks for having us.

Daniel Ho: Thanks so much.

Narrator: That was Kristian Hammond, Daniel Ho, Jennifer Logue, and Jacob Ward chatting about building social science into the foundation of AI practice. You can learn more about this episode by checking out the show notes. And if you want to learn more about the Center, its people, projects, history, and upcoming events, you can head over to our website at casbs.stanford.edu. And if you want to join the conversation with us on Twitter, we're @CASBSStanford. We're also on all the major podcasting platforms. So if you don't want to miss another episode or you want to check out previous episodes, go ahead and follow us in your podcast app of choice. Until next time, from everyone at CASBS and the human-centered team, thanks for listening.