Susan Holmes, a Stanford professor of statistics and 2017-18 CASBS fellow, chats with host John Markoff about her applied work on the human microbiome, the difficulty with P-Values, the power of heterogeneous data, and her research on Claude Shannon - widely known as the father of information theory and himself a CASBS fellow in 1957-58.
Claude Shannon, former CASBS fellow, and the "father of information theory"
CASBS on Twitter
Shout out to Barbie Mayock, CASBS dining program coordinator, for reading this episode's opening line!!
Announcer: From the Center for Advanced Study in the Behavioral Sciences at Stanford University, this is Human Centered.
Narrator: Today we'll hear a conversation between John Markoff and statistician Susan Holmes. Holmes was a CASBS fellow in 2017 and she focused on the issues of reproducibility when using modern statistics and the complexities involved with using heterogeneous data types. We'll hear her thoughts on the problem with p-values, the importance of messy datasets, and her research on the human microbiome.
John Markoff: I wanted to start by asking about your intersection point with CASBS. I mean, you're, you're sort of nominally within the biological sciences. You were spending time here in a center of social sciences. Is there a point of intersection?
Susan Holmes: Yeah, well, the intersection is my interest in, uh, the teaching the general public about statistics and data science. So the intersection was educational, but also I've written several papers about political science data, and so I had an intersection with the social sciences already. It's just that— and the type of data that I analyze, which has to do with how bacteria interact, they actually— the same techniques from a statistical viewpoint as when we study Facebook networks or Twitter networks. The social network, it's a social network of bacterial communities, but You know, from my abstract viewpoint, it's the same.
John Markoff: But that sort of takes me— I mean, so what was the phrase about lies, damn lies, and statistics? But you— I saw in your conversation with Mike Gatani that you had talked about your skepticism about what are known as p-values.
Susan Holmes: Right.
John Markoff: And I was wondering first, because this is a lay audience, if we could explain what the p-value is and talk about the reliance on that technique and what's wrong with that.
Susan Holmes: I mean, sort of as a walk into —So all social sciences still require statistical analyses to publish p-values, and they are the— what people see it as a measure of uncertainty, and we have this hypothesis testing setup, and if the probability of the data under a certain hypothesis is very small, You say the p-value is smaller than 5%, that's an arbitrary number that was set in the 1920s by R.A. Fisher. And that number has now become like the absolute norm for publishing any kind of sociology result or psychology. And of course, that gave rise to, especially with computers, the ability to hack the system. That is, you can test many, many times and then hide that you tested thousands of times to get the one result. Which people routinely do. Yes. Right, right. And so although in 1920s when R.A. Fisher was doing it, a p-value of 5% corresponded in general to work over a 20-year period where you had a 1 year out of 20 in which you got a result. And so it was a very long-term thing, and he was studying crops, and he considered that, you know, if it was a 1-in-20-year event, then it was special. And this is very different than being able to run it 50,000 times in an hour.
John Markoff: So, you know, I'm a failed social scientist. When I was studying social sciences, you could collect survey data and you could perform statistics on it. What seems to have changed in my sort of observation of the social sciences now now is now it's not just survey data. You can collect census data. And you can get this real-time behavioral data. So you get a lot more data. Right. From your view as a statistician, has that improved the quality of social science? Or are they still relying on the same—
Susan Holmes: No, no, no. It has improved enormously because you can capture multivariate— what we call multivariate phenomena. So phenomena where the perspective— you can measure all the context. By many, many variables, and so it's much richer data, and so you can capture high-resolution subtle effects. And so I think that, of course, it's improved social science. I see the work that some of the people are doing even with Twitter, or, you know, you can follow things which are real. It's just, what I would say is the p-value is wrong because it's trying to summarize this huge complexity with one number. And I'm into this multivariate, where you have multiple things you're measuring. And I just want people to say, make a lot of plots, visualizations, look at your data, but don't think that one number is going to summarize a complex phenomenon. It won't. Well, we— I mean, I believe in graphics. So I like to make what we call principal components or various kinds of dimension reduction. So you can go from measuring 10,000 features on a population to looking at a graphic in 3 dimensions or 2 dimensions with dimension reduction. So we know how to do that with uncertainty quantification, so you get scatter point clouds and you have contours. And even in, you know, New York Times or any kind of newspaper, people are getting much better at doing that. And it's becoming much more standard to say, okay, you have a map, we're looking at a map of some contagion map or something. People, you have the spatial information, but you have color, for one variable and you might have size of the— so you can capture multivariate with the technology that we have, but also because the public has become much more educated. And the problem is the infrastructure of the standard publication in the academic literature hasn't followed. That is, it's not so easy to evaluate the uncertainty. It requires people to make judgment calls. And it's like doctors, they— doctors don't like to be told, oh, there's a score and it's on a spectrum from 2 to 7. They want to know one protocol or the other, which one, where's the, where's the cutoff point. I don't tell me, you know, just tell me yes, no, because I don't want to spend a lot of time hesitating.
John Markoff: And so that's why, sort of heading down a path to your work in biological biostatistics, I saw that you were interviewed at Stanford at a conference and said you really enjoyed working for big messy data sets. I mean, it wasn't so much big data, it was you like messy data sets.
Susan Holmes: Yeah, yeah, yeah. So I don't like this restriction that people have said, oh, you know, now everything has changed because we have big data. As a statistician, if you have hundreds of millions of records, but the records are homogeneous, you have absolutely no problem working with them. You just take subsamples and you can make very good inferences by taking many, many different subsamples I think the big challenges that we are facing have much more to do with heterogeneity, that is that you're measuring many, many different types of different variables. So I take the example on cancer. We— you could take a biopsy and you have an image, and you also have all kinds of measurements to do with the immune system, then you have the measurements of what genes are expressed, what proteins are made. And so you have maybe 12 different types of data. But you don't have 10 million people on which you're doing the measurement. Maybe you have 10. And so the huge challenge is not—
John Markoff: big variable, not big data.
Susan Holmes: Yeah, that's right. It's the features. It's the numbers of variables and the heterogeneity of the variables. That is, the amount of uncertainty on one measurement is very different from on the other measurements. So you don't know which ones are going to give you the info. But you don't want to let anything go. So you want to combine.
John Markoff: And I also saw that you spent a lot of time looking or doing research related to the microbiome. Yeah. And did you— first question is, did you— have you crossed paths with an astrophysicist who's the guy who runs UCSD's laboratory there? His name is Larry Smarr. He's gotten a lot of visibility for sequencing his own microbiome. Yeah.
Susan Holmes: I mean, there are lots of people who— and then Mike Snyder at Stanford. Oh, yes. He's the other example. Yeah. We call that the narcissosome because you only do one person. You do everything, all the data about one person. So you're measuring everything about one person, and the question is whether you can infer— it does give you something. I mean, if you measure things about yourself and you learn things about yourself, it's definitely a sample of size 1. And as a statistician, I would say, you know, that doesn't mean anything. But the person who got the Nobel Prize for Helicobacter, He did the experiment on himself, Marshall, and he gave himself Helicobacter and gave himself an ulcer. Oh my God. And, but proved, and he gives a talk called N equals 1. So, you know, there are pro— there's progress on personalized medicine working on, but that's not what I work on. I work on groups of people. I do studies with David Relman, and we work on the effects of antibiotics on different people. And so what does antibiotics do to you if you take it once, but what does it do to you if you take it several times? And so that's, I mean, that's the microbiome.
John Markoff: There are now at least 2 or 3 companies that are doing 23andMe-like consumer services for the microbiome. And my wife talked me into one of them. And so I have had my microbiome sequenced. I have to admit that I haven't really spent the time to look at what came back. She changed her whole diet as a result of that. But do you think that it's— is that premature to say? Well, we did.
Susan Holmes: It's interesting as a scientist, I love this because what happens is it used to be really hard to recruit participants to studies. And when we got in touch with the Quantified Self group in San Francisco, we were able to find more than enough study participants for our study where we asked the participants to do colonoscopies and take antibiotics probiotics twice, and there were many volunteers who wanted to do that. And in exchange, we give them their data. And but, but, and it's very interesting data, and we needed a large pool of different people because we're not interested just in one. In a lot about the microbiome depends on not only what you eat, but also where did you go, where have you traveled, um, how is your immune system. It's very dependent on a lot of of other factors than just what you eat.
John Markoff: So what is your current research? Do you have—
Susan Holmes: it's all about the methods for the microbiome. So I'm trying to— I've had a lot of success by realizing that when we analyze documents, we do topic analysis. We can find different topics in a document. And what's really important is to say a document can have several different topics. And taking documents on the web of different lengths, we can still say these are the topics. And I've been using those methods for understanding the communities in the microbiota. And so what happens is different strains are like different words, but you can have an overall meaning of the sentence which is the same in these three different sentences, and they're using synonymous words. And this is the big challenge in the microbiome, is the strain-to-strain variation between people is very high. That is, the biggest source of variability in the microbiota is the subject-to-subject variability. So it's very hard to say, okay, this illness is associated to this strain going up or down. It's not one strain, one illness, but it's one topic on one community which is present or not, which is characteristic of a situation. And so you have these strains which are synonymous. And so I use a lot of topic analysis in the microbiome to understand. And the difficulty, of course, the challenge in all of my work has always been the communication problem. It's not the analytics or the math. It's communicating to the biologists these quite complex hierarchical models. So they're Bayesian hierarchical models. But they're the same ones that work for language.
John Markoff: And you talked a little bit about region, geographic region. So that sort of takes you into the social realm. I guess it could also be working class versus upper class and how diet varies or something.
Susan Holmes: Yeah, but it's not actually working class, upper class. What we see, the biggest changes in the studies I've been involved in, and I work with people who work on the Indian microbiome and all over the world, the HAZDA and all different studies we've done, it's the urbanization. That is, that it's not economically exactly, it's the change from the very old cultures, hunter-gatherer, classical farmers, and then all the way to urbanization. So we see a big change with urbanization.
John Markoff: Is that a spectrum of health?
Susan Holmes: Well, I mean, the micro— microbiota as measured by diversity, yes. That is much, much more diverse. And if you took one variable, people always ask me what's one variable, I would say that in the work, one of my collaborators at Stanford is called Justin Sonnenberg, and he's written a lot of interesting papers and a book about this, but it's fiber. We don't eat enough fiber anymore. And so of course, We talk about sugar, sugar is bad, and you know, they're all— but if you, you know, going back to the more traditional societies, there were huge amounts of fiber in the diet, which have been completely lost. So artichokes are good, and broccoli, right? Yeah, it's interesting and difficult because, as you were saying, the socioeconomic factor, there are all kinds of things involved. And it's not only the gut, so I work work a lot on preterm birth and pregnancy, and we use that from vaginal microbiome. And there we do see that, you know, there are certain bacteria which are very protective, and Lactobacillus, for instance, very protective. And so you can predict preterm birth by looking at swabs from the vaginal microbiome. So it's very used as a biomarker as well, but it's causal. We don't know anything about causality.
John Markoff: And okay, and then there's also all of this work about, um, things like personality and intelligence, and, you know, the microbiome was being a factor in all of these sort of very high-level, uh, well, behavioral—
Susan Holmes: talk a lot about the gut-brain axis. And it is true that in all the constituents, chemical ones that we use for our brain come directly from our gut. That is, you know, it's absorbed and it goes directly into the bloodstream and it goes, you know, there's a definite— I just saw a very interesting talk about that axis which had to do with the keto diet. So my son does a keto diet and there are lots of people who do this strange keto. And that came actually from studies of people who had epilepsy. And they found that if they did a specific diet, the keto diet, very, very severe, 50% of the people benefited hugely, practically down to no, having no fits or anything. And so they're understanding, there you have it, there you have the brain-gut axis right there. And a personal note, I have celiac disease and my only manifestations apart from from lacking vitamins and all kinds of things was 45 years of headaches. And so this is wheat, right?
John Markoff: Yes, it's wheat. And so you gave up wheat.
Susan Holmes: And the doctor didn't know, and I was on all these medication for migraines, and then I stopped, and it was— so I know about the gut-brain axis. It's really something which is there. It has to do with inflammation. So overall inflammation, when you have a reaction against something, it depends. The inflammation can go anywhere, but the brain is definitely one of the places. So that's— so n size 1, you only believe about it. Even as a statistician, I believe much more about that than anything else.
John Markoff: What was your path to statistics? How did you find your way to this world?
Susan Holmes: I was a mathematician by training. I was in a PhD program for mathematics. And I was very interested in geometry. And I liked the computer. And I wanted to use the computer. And the chair of the math department said, "We will never have a computer in this building. You have to go and change your subject and go— the statisticians have access to these mainframes." At the time we're still in the world of cards, right? "But the statisticians use them all the time, so you have to change." And so I changed. And I didn't know anything about probability or statistics, but I went to see somebody who would take me on as a PhD advisee, and he gave me a whole stack of books to read over. And I thought, didn't make any sense as a mathematician. Statistics doesn't make any sense because it's an approximate science, and mathematics is very precise. So statisticians accept to be wrong some of the time, and we just try to be useful in a context. So it's very much about decision-making. And very little about actually mathematics, although the tools that we use, geometry in particular, are very mathematical.
John Markoff: So you're completely taught computer science. Did you teach yourself UCSD Pascal? You must have. Yes.
Susan Holmes: Yeah, yeah, yeah, yeah, yeah. But, and what was your first language? Fortran. Yeah. So, so in school, I, uh, we are— the only formal class I ever had was Fortran, and all the Unix I learned I learned it the same way Don Knuth says he learned his, which was there was a computer game on Multics, on Multics, and this, you know, you go into a cave, The Wumpus, or, um, yeah, yeah, what was the other one? Um, there was NetHack and, uh, all those text-based adventure games. And what was funny was in France we had access to a terminal called, uh, it was a very funny thing, the French government To go with your phone, you could have something called the Minitel, which was a little screen which you could open up. When I first visited Stanford, I could email the secretary or anybody on campus to try and find housing from my house, which they found amazing because this is 1988. '88, yeah. Yeah, yeah, yeah. I could set it all up by email from home, which nobody had.
John Markoff: You didn't grow up in France, did you?
Susan Holmes: Well, the first 10 years I was in England. I see, I see. Then afterwards, I did all my schooling in France.
John Markoff: Did you come from a scientific or engineering family background? No.
Susan Holmes: What happened was my father was a doctor, and he sent my sister and I to a school for young ladies. That is, we learned embroidery and knitting and all that kind of thing that only young ladies learned. Then he had a midlife crisis when he was about 40 or so and decided to become a farmer in the South of France. So we moved to the South of France, and I hated it. I didn't like the— I was about 12, I guess, and I decided to work very, very hard in school. I had no academic background, but I had all of this knitting, which was very useful for computer stuff because it's the same, you know, you have to be very patient. But, but then I did learn, you know, I was very motivated not to be a farmer, so I intellectual— I saw this as a way out.
John Markoff: He wasn't growing wine, he was farming.
Susan Holmes: Yeah, that's right, he was farming. He was— he had— but I— getting up at 5 in the morning to milk, you know, goats was just not what I wanted to do.
John Markoff: Where in the south of France were you?
Susan Holmes: It was near Montpellier, it was in the Cévennes. So it was the area where all the people who dropped out in '68 went, and this was '70, '71. And they, you know, there was this whole dropout culture, so it was completely empty. There were no— all the houses were empty because all the people who were farmers had gone back to the city where you could have hot water and bathrooms and things like that. And so there were all these empty villages, and he bought up one.
John Markoff: And did he stay? Was he romantically attached?
Susan Holmes: He gave up on the farm and then he did pottery, which for a surgeon is probably a good a good transition.
John Markoff: I've forgotten his first name. Solomon, head of the statistics department still when you arrived?
Susan Holmes: No, he was a retired emeriti. I knew him. When I first visited, he was active. He was one of the people who created the department, but there were quite a lot of people. Brad Efron at one point was chair, and then Jerry Friedman. Oh yeah. Yeah, he— so Jerry's the person I met who had first invited me to give a course. And Jerry was very much in favor of the French Exploratory Data Analysis School, which was like John Tukey's. And Jerry was originally a physicist, and he'd come to statistics through John Tukey. And John had spent a year at the center, and he was also interested in social science problems as well. John Tukey was definitely a very influential figure here. Tukey and then Shannon as well.
John Markoff: I saw that you had a real interest in Shannon. Was that a research interest of yours?
Susan Holmes: It's a book I'm writing on codebreaking and pattern searching. I'd found out that Shannon had spent a year here in between when he was at the Labs. And when he was afterwards moved to MIT, he spent a sort of hinge year here. And I wanted to look at all the papers that had to do with Shannon. And I remember looking up, I went down to the library to find out, you know, you have to write a proposal of what you would like to study. And he was interested in studying psychology so that he could try and make more intelligent computers. So he was interested in AI.
John Markoff: Did you find where his— Shannon's— that would have been super early, right? 1957. So right during the summer study period when they were coining the term AI.
Susan Holmes: Yeah, yeah, yeah, exactly.
John Markoff: So, so, so, so, so. He wasn't at the meeting though. Shannon didn't participate in that.
Susan Holmes: No, but he was one of the people who really was pushing for that. But afterwards I looked at some of the letters and some of the things he wrote, and he was very, very disappointed with his interaction with the psychologists. He didn't feel that they knew anything that would help him make computers more intelligent. And so, you know, the level of knowledge, uh, you know, he was hoping that all this interaction with the psychologists would push forward his goal of intelligence.
Narrator: Was he, was he right, or was he just sort of really honest?
Susan Holmes: No, no, I think he was right at that time. Yeah, I, I think that there weren't ways of thinking it was a Cognitive psychology at the time was, I won't say quite, you know, just-so stories, but there wasn't an analytic understanding of, for instance, of layers, or of, you know, it was very, very black box.
John Markoff: And still Freud and Freud and at all were still very—
Susan Holmes: Well, and lots of scenarios, but not useful in a from a pragmatic viewpoint and not very much experimental evidence yet.
John Markoff: And where did you stumble across the intersection between Shannon and Turing?
Susan Holmes: Well, that comes from the war. So what I was very interested in was Alan Turing and Jack Good, who's much less known, but he's a statistician who's a hero for people like me. He's a Bayesian statistician. So Turing had a notion of quantity of information and a measurement of how you weigh evidence. And that was done in log base 10, and it was called the ban and the deciban. And that was Bayesian— based on Bayesian statistics. And Shannon had something extremely similar, which later became called the bit, which was log base 2. And which was a measure of weighing evidence and measuring information.
John Markoff: Oh, for channel— for, for look, characterizing channels. Yeah, communication.
Susan Holmes: I mean, Shannon did the work in this, in, uh, language for crypto, and Turing did the work in crypto. They both came from the same code-breaking background, and it was just the way they quantified it. Now, Turing was not allowed to publish So his work became published about 5 years ago, but Shannon published straight after the war. And I wanted to know whether they'd met.
John Markoff: And did you find out? Yeah. And they did.
Susan Holmes: They met, but they weren't allowed to talk about codebreaking. They were allowed to talk about speech recognition. And they were allowed to, because that's what they were both working on as well. That was their sort of in the front. What they were working on. They were allowed to talk about that, but I wanted to know. They met during the war, so they met, I think, '42, '43. Turing spent more than 3 months in Princeton at the time, in New Jersey. I'd met for lunch at the cafeteria quite often Shannon, and they talked.
John Markoff: Is there prior art for the perceptron and deep learning? If you were going to trace the intellectual heritage of that pattern recognition approach to—
Susan Holmes: Well, I mean, there's no— how can I say? The quantification of information is a first step before you start seeing that you have to have a layered system, and the perceptron was slightly after that already. But Turing was definitely somebody who— many people have read various articles that he had about how you could make computers intelligent. And so both Turing and Shannon were very preoccupied with that. And the algorithms or the things that they developed were— they were based on these how do you increase the amount of information that the computer understands?
Narrator: Really quick before we go, can we get an update on— well, first, what is the codebreaking book, and when can we expect to set our eyes upon it?
Susan Holmes: So the book on codebreaking came from a course I teach at Stanford called Breaking Codes and Finding Patterns, which is very much about my work both in biology and in codebreaking during the war. And we're not I've been doing a lot of teaching. I taught pre-med students. I teach a lot of biologists how to analyze their data. I've been very occupied with finishing the book for biologists. Now I've finished it, I've come back to the code-breaking book. The book is not finished yet, but we hope another year or so. It's okay, it's all about the war, so nothing has changed. More information comes up. But this is great.
John Markoff: Thank you. That was fun.
Susan Holmes: Okay, thank you, John. Yeah, yeah.
Narrator: Thanks again to Stanford statistician Susan Holmes. To learn more about her work, be sure to check out the links to her personal site in the episode notes. Human-Centered is a show from the Center for Advanced Study in the Behavioral Sciences at Stanford University. Special thanks this episode to Barbie Mayhawk To learn more about the center, visit our website casbs.stanford.edu or find us at Twitter @CASBSStanford. Thanks for listening.