Skip to main content
Skip to Main Content
Skip to main content
Navigation

Data Science for Us Pesky Humans

Humans have spent centuries probing reality for its secrets. We started with vague guesses and crude measures, and over time, moving to gross math in N dimensions and computer code that processes data. Today, we call this data science, and academics have designed it… for academics. In this session from the 2026 VEX Robotics Educators Conference, Dr. Andreas Stefik, Professor at the University of Nevada, Las Vegas, will discuss a five-year journey into how to make data science easier for us pesky humans. He shares how making it more accessible and easier to understand could help us all get back to a shared language of facts.

VEX Robotics Educators Conference. Hi there. First, yes, Andy Stefik. I'm a professor of computer science, as Jason said, and I just wanna say first, you know, thanks for inviting me. I really do appreciate coming talking to you all. These kind of events are massive and get a lot of kids involved and that's a good thing. So we need more voices in science. And so what I'm gonna talk about today is data science, what that sort of means and what it means for us people. I know sometimes we think about things like AI and other contraptions about how we're feeding things to computers and computers are feeding things to us. But data science is not really like that because data science has to be exactly correct. Why does it have to be exactly correct? Because if you don't and you start approving drugs that don't work, people die and vice versa. So, believe it or not, although I swear I'm not trying to be cheeky, it starts with beer. And it's odd that that's true. But humor me for a moment. It turns out that in 1759, there was this guy named Guinness and he started a company, you'll never guess the name. And at the time, we didn't have these pesky things like germ theory, right? Or you know, pasteurization, or these other ideas that are really important for safety with drinks and food in general.

So there's lots of safety issues.

And of course it was 1759, so getting stuff around was harder. So if you both didn't have germ theory and you didn't have pasteurization, well transportation becomes really tricky. And then there was crop quality issues. And this sounds like it's just some silly little thing, but it was complicated. And so what is a lowly brewing company to do? Well, it turns out Guinness invented statistics. Now you may not know that, it's really odd that that's true, but it really is true. Guinness, like the actual beer company, invented statistics. And that's weird. And it's interesting to think about how did they do that?

So what do you do? Well, the first thing that you do is you hire a bunch of math nerds at Oxford and Cambridge. Now keep in mind, I say hire a bunch of math nerds, but I didn't say hire a bunch of statisticians, that didn't exist yet, right? Guinness, the beer company had to invent it first. It's funny, if you go to Dublin, you go to their little factory, there's actually a little sign. In any case, after they hired a bunch of these math nerds, they then made a nerd pipeline. And what I mean by that is they didn't presume that if they hired just the right person, that they would invent it all. Instead, they had the foresight to think, okay, we're gonna need a lot of people collaborating together. The ones that are more experienced are gonna have to teach the younger ones how to do this.

And then what they did, these nerds helped conduct what are called the barley plot experiments. The details aren't really that important, but the important part is that in the barley plot experiments, they gathered a ton of data and then they waited 20 years. And by waiting 20 years, what I really mean is it kind of took that long to figure out the math 'cause the math was pretty hairy. So what makes it hard? Why is it hard to brew beer or to sew crops? Well, it turns out that a lot of crop stuff is actually a math problem. And when you really think about it, it's an issue of scale. So you couldn't eliminate variations of things like rainfall. How much does that impact the beer? I don't know. What about bird damage? Weirdly, I know it sounds really strange that like they would care about literal bird damage, but they tracked the damage that the birds had to the crops and they're like, what impact does this make? Soil chemistry, temperature, all this stuff affected the crops, which impacted either A, their bottom line, which matters, but also it impacted their ability to get product to the public. And past all of that, they didn't know how to account for all the things that they didn't know how to track.

And this is the weird aspect of data science, because you don't know what you don't know. And so you always end up with this little amount of error. But they figured out a magic trick. And the magic trick, the secret sauce for pretty much all of data science and statistics is math-ifying variation. That's really the bottom line, is you take the concept of variation that is kind of intuitive. You know, climate change is not just a constant. You have night and day, that's variation. You have summer and winter. These are variations. And if you can math that out and how you do that turns out to be pretty hairy, then it lets you rule things out. And this is weird, but with a maybe. And what I mean by that is it's kind of like a pendulum where if you swing it in one direction, you can probably rule something out, not confirm something. That's weird. It has to do with swans. I promise I won't get into it, but if you swing this pendulum, you end up with what you might call maybe thinking, which is hilarious when you think about it.

But if you think about something ruling something out with a maybe, that means you don't know for sure. So you end up with these estimates of like, how sure are you, how big is the effect? These things that we think about today. But the thing is, and this is where data science took a turn for the weird, I guess it's already weird that Guinness invented statistics, but it took a turn for the weird because they then pulled the Lord of the Rings, even though that didn't exist yet. Basically what happened is this guy, his name was William Gosset, he figured out the core statistical formulas and he solved the problem of barley in Ireland. So Guinness went around and they actually bought barley from all over the place and then chose a specific one based on some mathematical formulas. But they knew it was a massive advantage. I mean, think about it, your competitors are just guessing and you have now a math equation that lets you pick a barley and you know it's the best one, or at least you have pretty solid evidence that it's the best one. So they decided to basically keep it secret, keep it safe. Eventually, internal pressure at Guinness went to a board meeting and one guy decided to reveal statistics to the world, but to hide it a bit, they didn't want it connect it to brewing.

They weren't gonna send it to like Brewers Anonymous that I don't know if such a thing existed, but you get my point. And so there was this guy named La Touche, I don't know anything about this guy, but the name is funny. And he was decided by La Touche that such publication might be made without the brewer's names appearing. This would be merely designated pupil or student. So then they got a test, but they're trying to hide it. So it got a really intuitive name. So now at this point, because I think I'm talking about statistics and that runs the risk of being the most boring talk on earth, I want to ask the audience to please just shout out what you might guess the name of this test was. And seriously, just shout it out, I don't care. It's okay. Just guess. Yell something. What do you think it might be? Okay, ideas, what do you think, Jason? Barley test. Barley test. That'd be a good one. Okay. But they gave it a much more intuitive name, T.

Now the best part about T, it very clearly implies what it is. No. And it's a single letter. And when you think about it, that makes perfect sense. If you're a company and you don't want people to know it's related to brewing because you have a massive financial advantage to do so. But on the other hand, without the publication of these materials, literally for decades, we didn't even know who it was that was publishing this. It was just students' thing.

But because of that, Guinness got to keep its advantage. But the whole world changed around this idea of quote unquote maybe thinking.

So a bunch more nerds at various places. Actually really, three people that were near England at the time were probably the biggest ones, but I won't get into it. Eventually they started thinking about this concept of N-dimensional math. Now if N-dimensional math hurts your head, I promise I won't go into it at great detail. But the point is they basically had to turn this into matrices in a funny way. And that turning it into matrices had a really big advantage that you could tackle multiple things at once, like bird damage and rainfall and you could put it all in the same model. Guinness didn't invent that, but their equations were pretty easily generalized if you happen to be a genius mathematician in the early 1900s.

It also led to advances in methodology. Turns out this N-dimensional math is the reason why we know that smoking causes cancer, right? That was done by this fellow named Austin Bradford Hill, who actually used the equations in a very special, funny way along with some methodological advances. And all of a sudden, ta-da, we don't just know that smoking is correlated with cancer. We know that it causes cancer. And you do that by maybe ruling out every little thing that can get in the way. Is it how much you use a pipe? Is it using cigars? Is it gender? Is it your location, is it, et cetera, et cetera, et cetera? And when you read these old papers, you can see how they were using this quote unquote maybe thinking to just go, is it that, is it that, is it that, is it that, is it that, is it that, is it that? And then once you've ruled out enough things you don't know for sure, but you kind of know, there's not much left.

Today, especially starting in the early 90s, this has led to things like evidence standards and reporting guidelines, things that computer science, believe it or not, does not have, evidence, is actually used by most other fields, but for whatever reason, it's very rare in the tech sector and in tech academia even, which is really odd to say, but it's provably true. There's even some evidence that many scholars in certain regions of computer science actually don't use the scientific method at all, even in scholarship, past peer review, which blows my mind. But that's another issue. In any case, while they were doing this, they were making tests for all sorts of situations and it got ridiculous. The first test was called T and it just got worse from there having scholars in charge of all this. And now at this point in the talk, I'm going to read you absurd nerd stuff. And I swear there's a point. I want you to feel how ridiculous it is because if you just look at the list, you're gonna glaze over it and it's not gonna make any sense. But I want you to realize everything on this slide is used for specific evidence-based purposes. And I'm not gonna go through all of them in like a lot of detail 'cause I don't wanna get booed off the stage already.

I don't know why I looked at my non-existent clock, but you get my point. But I do want you to give a sense. So the first one was T, this was Guinness obfuscating stuff. The next one is F and weirdly F is T actually with a little tweak. Z is the number of standard deviations away from something. S is the standard deviation under certain conditions. This little funky thing that's hard to see on the slides, it looks like an E, but it's not. It's actually an epsilon, has to do with error. But scientists don't use it consistently. So sometimes it means different things.

You've got U, H and W, don't even ask. It's this whole calculus thing. This little funny U with a line is actually the mean under some conditions. P has nothing to do with urinals and has to do with maybe thinking. This B here, which is not a B, it's a beta, is basically how strong of a predictor something is under certain conditions. This Y is an outcome variable, what it predicted, this funky N is not actually an N, it's an eta, it's a Greek eta and it has a couple different meanings related to how big something is. And then there's this funky little three dots. Does anyone know what that one is? Just raise your hand if you know what that one is. Okay, good. Because that one's actually make believe. It doesn't mean anything at all. I just tossed it in there. But then there's the tests. They use all these funky things and they make stuff like the Tukey HSD, the Mann-Whitney, the Greenhouse-Geisser correction, the Kruskal-Wallis H, the Bonferroni correction, KMO, Wilcoxon Signed Rank, Spearman, ironically, you know we name it after the mathematicians. Yeah, he didn't even invent it. Welch's paired, homogeneity of variance, Eigenvalues, SS loading, which is very similar to Eigenvalues, but not the same, Levene, Brown-Forsythe, Mauchly's sphericity test, chi-squared, Games-Howell, Pearson correlation, Bartlett's sphericity test, orthogonal varimax rotation, oblique direct quartimin rotation, principal component analysis, factor analysis, which actually uses the same math, but nevermind that, and then regression.

But here's the thing outta everybody that's here, how many of you find this intuitive? How many of you do not find this intuitive? How many of you are horrified?

You're not alone. But the problem is, if it's ridiculous, people will struggle to do things.

Number one, will people understand the data around vaccines If it's written in this language? Obviously there's malice in the world, but that's not what I mean. If you read, how many people have actually read the original COVID-19 paper with 45,000 samples? Probably not many of us, right? And even if you do, unless you're really familiar with the science and how it works, it's gonna be incomprehensible.

Can people really understand the data around the climate unless we're showing them how this works? But why are we talking this way to people?

If we're using robots to collect data, which we do all the time factories, kids know this from VEX, they're trying to do their competitions and stuff like that. If you wanna use variation in the data, well this is how scientists talk about it. But you're not gonna give that to a kid. That's nuts.

If we want kids to read science, it's written like this, all of it.

If we want to have laws passed that follow evidence, which we used to do, then those people have to understand it too. And if we wanna know that AI is really helpful, but also, you know the paper has a limitation section.

You need to know how it works. It's actually a word guesser. It's a funny thing. It's kinda like Siri, but much more powerful. It reads stuff in a certain way, but I can't think, it's complicated. So I have to ask, have we just lost the plot, right? Like as data scientists, oftentimes we want things to be correct, we want the right drugs to be approved and the wrong ones to be rejected. But we also are talking to scientists and we've got these traditions that are hard to change and all this kind of stuff. So I had a question and I told the VEX people that this was a five year journey, but I looked in the mirror and my hair was a little grayer than I thought. So I had three questions for this apparently eight year journey. Number one, can we even clean this up? Like is it even possible? What would it mean to clean this up? Number two, can we make it less scary, right? Like I'm not actually joking when I say is this kind of frightening? I think it actually is. And it's not because it just makes people feel dumb or something like that. It's that it doesn't really make any sense. And that means if you look at it, you have to say to yourself, oh boy, I have to go through, I don't know, years of classes to do this.

And it makes sense that it would be that way. In the field of psychology, just even starting out to understand that stuff is a two year sequence at the PhD level. Does it really have to be? A lot of the stuff that these tests do is pretty simple actually.

And then the final one is, can we make it accessible? But I'm gonna challenge what that means to be accessible in this context because that's part of this crazy journey. So the first one is, can we clean this up? A lot of this stuff seems so obviously wrong that it feels like there might be an answer. Maybe we don't call it T, maybe we call it something else. But the question becomes, well what do we call it? How do we organize it, is there even a better approach? Maybe there actually isn't. It might not be that T is bad, but it might be that cleaning it up really doesn't make any difference. That's a very real possibility. Unfortunately in science that happens all the time. So about in 2018 or so, there was a survey done by this group called Stack Overflow. It used to exist, it was the pre-AI times back when we used rotary phones.

Don't tell that joke again. And at the time, 11.3% of developers were alleged to be mathematicians or statisticians. Hard to know for sure, it's a survey. They have flaws. 45% of all developers at the time were using Python R or MATLAB for their jobs. Pretty high percentage. And so my point in saying that isn't so much that that's good or bad in some way, it's just that a lot of people do data science. It's just part of your job sometimes. And there was an active debate at the time on language style. This isn't just MATLAB, R or Python, it's not like a language fight. It's not that, it's that I'm not the only one that noticed this is a mess. And so a lot of people are like, well what's the problem exactly, how do we fix it? And so there was an active debate and at the time, there wasn't a whole lot of evidence. So I have a test for you all. On the screen, is three different choices. And what we're gonna do, we are going to clap if you think this matters or not. And I'll get there in a second, I like to do this a little bit, especially when it's something like data science 'cause it can be a little dry. So just please, please just humor me.

So option one in the programming language R, you could argue that Python or others are better. We'll get to that.

There's a style. All it does, all this does is calculate the average. There is nothing special about this. It's kind of like clicking the column in Excel and then say take the average. That's really all this does. The first one is solution equals mean(americanStats$RBI).

The second one, Tilde style is solution equals mean Tilde RBI comma data equals americanStats, right paren. And the last one, Tidyverse's solution equals American stats percent greater than percent Summarize left paren, mean left paren, RBI, right paren, right paren. Rolls right off the tongue. Now you might think that if you ran a rigorous experiment, this isn't gonna make a hill of beans of difference. It's just all simple stuff. Or you might think that it does or you might think that something else is the problem. So we're gonna vote, you have three choices and you're gonna clap for which one you think is the right answer. I'm gonna tell you the answer in just a second, but I just want you to guess, okay? So if you think this doesn't make a hill of beans of difference, clap away. (audience clapping) Okay? If you think that this is actually pretty important, clap away, couple. If you think something else is the problem, clap away. Nobody. Oh, one guy. Well you're the right answer actually.

Okay, so we ran an experiment, I was encouraged by a friend of mine in Germany, it's a funny thing. There's actually a castle in Germany dedicated to computer science called Castle Dagstuhl. It's a thing, don't ask. But in any case, at Dagstuhl, this place in Germany, a friend of mine that's a statistician was pretty convinced that this was not necessarily the problem, that it was a problem and it was an active debate in her community, in the R community. And so she's like, Look, Stefik, just run a fricking study for me. Would you? And I was like fine. So we ran a study.

And interestingly, we ran it as what's called a repeated measures randomized controlled trial. Which all that means is it's like you have multiple groups and you are either in group A, B or C and you do a bunch of tasks. That's all it means. And in this particular case, there was 10 tasks and we put 'em in each style. People didn't know what style they were in, they didn't know what we were testing and we randomly assigned them. The reason why for that is a math reason. But there's a second reason too is in that studies that use randomness like that have some face validity. In other words, I didn't like find all the good people and put 'em in the base R group and find all the, you know, and put 'em in the Tidyverse group. They were randomly assigned. I couldn't control it. We have a computer do this for us.

And the tasks were pretty similar. So I mean, sorry, were pretty simple. You know, calculate the mean of a column, calculate standard deviation. We didn't try to get people to calculate the Kruskal-Wallis H test. We just did like means and standard deviations, stuff like that as a simple start. And then we got about 150 people to participate. And once you get 150 people, you do what? Stir, no, sorry, once you have 150 people, you then put them in the assigned groups. Stirring is the AI, right?

Anyway, so interestingly when we recruited, we observed that we had recruited, somewhat accidentally, people with different experience levels in programming. Some people, there was men, women, there was freshman, sophomore, junior, senior, graduate students. Some of them had done some programming, some of them really hadn't. And maybe that's interesting, maybe it's not. And as soon as we ran the task, we figured out that we found nothing until we noticed something.

Now obviously this is super intuitive, right? This uses the same language that I'm arguing is pretty tough and so I'm putting it up there but I will translate it. And at the same time I want to realize that if you're a kid and you wanna understand the debate about data science, unfortunately you have to know how it works before you can even understand and question how it works. Which is very frustrating.

So we thought we found nothing because it turns out whoever argued that this doesn't make any difference at all is correct. It turns out at most, if you look at how people varied like the variation between people, at most, that effect takes up 0.7% of that variance. Basically nothing. It didn't make any difference. But then we noticed this odd thing which I've circled and you very intuitively get it reported as F1148 equals 7.35008, 0.034 obviously very intuitive.

And what that means is if you had a couple years of programming experience, that didn't help either and that's weird, right? That doesn't make any sense. Why would that be true? So you give people a couple years of like C++ training and it doesn't help them almost at all. At most three and a half percent or so of the variance, it doesn't make any sense. So we went back to the drawing board and we're like what the heck? What is the problem? So we thought we'd better throw the baby out with the bath water, she's fine.

But that is an adorable picture.

So what I mean by that is the second question, can we make it less scary? Maybe it's fear, maybe it's something else. Maybe it's organization. We really have no idea. But there were some things we thought we better not toss. And the big one is any of the innovations made by this person whose name is Hadley Wickham, it's the one person I will straight up name drop in the talk just because I think their advances have been quite critical. Hadley Wickham is this brilliant New Zealand data scientist that has all these really common sense ideas about how data science should work. And one of the big ideas that he has is this thing called tidy. And when you really break it down, it's the simplest thing in the universe. And I can't believe no one talked about it before Hadley Wickham, but to my knowledge they didn't. It's really simple, I swear to you. It's really just this. If you have a dataset and you're gonna give it to a statistics library, it should always be structured in the same way. And that's really dumb but also really amazing because all of a sudden if you're using statistics packages, you're feeding it to T, same way F, same way, Kruskal-Wallis H, same way, every time. So it took out one barrier that really needed to go away.

It had nothing, they did use that funky syntax, the syntax didn't help at all. But the idea very likely helped. We don't actually know, I don't think anyone's tested it, but it's kind of obvious that that would happen. Hard to say. But past that, if we're gonna throw the baby out with the bath water, this target, which is what the heck is a Kruskal-Wallis H test anyway? Feels like a good target, right? It doesn't make any logical sense. There's no reason that the mathematician's name would imply anything. It's people's names. It uses single letter variable names, which we know from the literature is pretty hard for people to understand. But also, come on, T doesn't mean anything. H doesn't mean anything, it's arbitrary. So if we're mapping the entire field of statistics and data science to like H or T, that seems like a good target to try to reimagine. But the problem was how the heck do you organize it anyway? Right? Does Kruskal-Wallis H test mean anything? Well, you wouldn't know from the name.

So we did a second study.

And in this one, we had a lot of questions and we sort of in the first study, we made this really narrow thing. And in this one, we kind of just dumped it all out there to try to get a sense of like what might stick. So in other words, a study like this, when you read it you don't say they've covered all the permutations 'cause that's not possible. But it helps you narrow things down in big chunks so you can make more isolated studies later. That's kinda the purpose of this kind of a study. And we were asking questions like how good or bad all the naming conventions, maybe T is bad, but are they all bad? If we could reorganize it from scratch, exactly, how do you do that? And then do you need to re-implement all the math? Turns out, yeah, which sucks. But we did. Do good names even exist for this esoteric stuff? The answer could be no. And in fact in some cases we found evidence that that was true under some isolated conditions. But I'll get to that. And then will the statisticians burn us of the academic stake for trying? And this one sounds like I'm being cheeky, but I promise I'm not. Sometimes when we do studies like this, I end up getting hate mail for awhile from other people in academia and like I always find that really odd, like dude it's called T, you really that tied to it?

But you know, traditions are hard to change.

So we had an idea and we eventually called this say what you do.

It's a simple idea. The idea is if you have a T test, it should say what it does. So audience participation time one more time. Does anyone have a guess what T actually does? I promise it's super simple and easy to understand. Anyone wanna shout something out?

You're in the right direction. An average is involved. And a guess, are you all scared? Okay, it compares the averages of something. That's it.

The math is a little tricky but what it does is not.

So we had a bunch of groups, we had two groups of what we called say what you do. We worked with a statistician and we came up with a bunch of names and then we immediately came up with a bunch of alternative names because we didn't know if any of them would make any sense. The second is we wanted to compare it to real things. We found in that Stack overflow study that 45% of developers are using Python, MATLAB or R. So we figured we'd better test those.

And then we also found, looked at the technical names like maybe Kruskal-Wallis H test is used in textbooks. Turns out none of the languages use that by the way. They made up their own whole set and not consistently either. They just kind of made stuff up. And then the last thing that we do is kind of a weird thing if you haven't heard of it before, it's in the academic literature but it's not used that commonly is called a placebo language. Now the concept of a placebo, you might know, how many of you have heard of a placebo just in general? Yeah, okay. So the idea of a placebo is use something intentionally, either false or intentionally not likely to work. The original one actually came from something called the Nuremberg salt test, which was in Germany in 1834. The idea was that they were actually testing homeopathy, news flash, doesn't work, but like at the time, they tested compared to a highly diluted solution. If you've never read the papers on this, they're actually, well actually they're in German. Okay, well if you've never read the actual translations of these things, they're hilarious because they diluted it like you know, 500,000 times or some absurd amounts like it's water.

So they're comparing water to water but nevermind that. But in a programming language, what a placebo is is less obvious. So what we often do is we generate something randomly, random symbols, random words. And so these particular words here are examples of the types of things that would be a placebo. So gawahi doesn't have any obvious intentional meaning. Or kigujib and cejunas, like you could maybe interpret it as something but it's not intentional. It's just random letters kind of jammed together with some vowels in between. And then we let people, you know, have that be one of the experimental groups.

And so then what we did is we thought well if we're gonna do a broad brush, let's just start with like a survey. What do people think about these things? Because that might not give us any data, but it's a broad enough brush that we might be able to look at a couple things at once and then narrow in other studies. So what we did, we took an idea from ironically election systems and we did something kind of like ranked choice voting. But for statistics terms, the idea is if people really like gawahi, I dunno how you say it, 'cause it's made up, then they can vote for it but they get a limited number of points. That's the idea. So in this particular example, what's on the screen is the word compare ranks is higher. And then some of the other ones like Kruskal-Wallis test or compare means ranked are lower in this particular case. And then we added pesky humans again, stirred.

And in there we got computer science students, engineering students, people at different levels, men and women. There are limits to this. This particular one was predominantly students. There's actually to the best of my knowledge, there's actually a worldwide replication of this study going on right now. No clue what the results are or what's there because I'm not allowed to be involved with it 'cause it's a replication effort. That's science, right?

And also you might imagine like if you're in the Netherlands where they don't speak English the same, then it might be different, or Finland. You would expect this to be different across cultures, that would be expected. But I don't know that, some of the studies that have been done on programming languages that have been replicated in Finland actually did get the same answer. So it's kind of wild. But nevermind that.

So there's a little problem, oh shoot, I forgot to have you vote. Oh well, that's okay.

There's a little problem and this chart is a little bit hard to see. But notice a couple of interesting things, right here is the placebos, you would expect those to do poorly. They're literally randomly generated words but notice quite a few people actually perceive them to be quite effective. That's interesting because it tells you, some people, no matter what you give them are gonna think it's fine. T's okay. But then number two, notice the technical names didn't do great against placebo but like, you know, they're words and they said I'll give the words a few points. Say well you do style names did better but like not slam dunk better. They're definitely the highest on the list. But, you know, there's variation between people. But then check this out, this is the programming languages, Python, MATLAB and R, and these three languages did so poorly that they mostly did not beat random noise in their designs. Now I'm gonna read to you what it says in the actual academic paper and then I'm gonna translate that to English. To investigate, we ran a pairwise comparison using T-test with the Bonferroni correction on each individual concept for each of the language groups and found that for MATLAB, Python and R groups, the existing name was not statistically significantly different from random option in 63%, 62.2%, 30.4% of concepts respectively.

What that means is they did so bad of a job in designing the API that they did worse than the stuff that I put at the beginning. All these weird symbols and stuff, they did worse than that and we were thinking that that was the lowest bar we could imagine. But in fact, they didn't beat random words. So we thought we should try to build this and make something that's built around this say what you do style stuff.

And if in one of the things that we do in my research lab, we build out this particular programming language that's called Quorum. It's based on evidence and it usually works this way. We gather some kind of data or evidence and then we let the cards fall where they weigh. If it looks like the evidence is sufficiently decent, we put it in the programming language. If it doesn't, we scrap it. That's how it works. That's what evidence-based programming means. That idea started around 2013 or so, although there's a few other scholars that attract the etymology differently.

So that leads to a third question. Can we make it accessible and maybe friendly, less scary, whatever? And in this particular case, by the way, my daughter's fine. In this particular case, when you think about accessibility, oftentimes people think about it only for people with disabilities. But that's a misnomer. Accessibility means in a sense universal design. And there's usually four broad categories of it. And they put this in a standard called POUR. How many people have heard of POUR? A couple. So the first one is perceivable. That means if you're a kid and you're blind or deaf or anything, you should be able to perceive it somehow through sound, audio, textual, you know, touch, whatever it is. The second is operable. And most of the time all this really means is that you need mouse and keyboard support. Not only mouse or only, you know, that sort of stuff. But then there's this pesky one called understandable. Now at this point, hopefully it's clear that data science stuff is not understandable. It's only a survey and some other stuff. We'll get to actual tests here in a second. But at least hopefully you at least feel a little bit that this may not be understandable or at least you might have a hypothesis that it may not be understandable.

And then the final one is robust, which has to do with robots. Actually not the same kind of robots that y'all are thinking, but robots connecting to software. It's a funny thing, the name is actually kind of odd, they really shouldn't call it robust but whatever. So we got to building for a couple of years. First I started rewriting all the math from first principles. N-dimensional math is fun just so you know, reading all the original academic papers is even more fun. And finding out the academic papers didn't put the equations in it. Yeah, that's a blast. But nonetheless, we eventually rebuilt it all and some of my grad students helped as well.

Once we've done that, we reorganized it around this idea of say what you do, right? Every test has to be given a name. We sit and argue about it. The data showed it didn't really matter a whole lot which name you picked so long as it implied what the thing does, right? So the two different choices that they had for say what you do were pretty similar. We kind of just picked the top one for each test and just did 'cause it sounded like rational, but nonetheless, it is what it is. We also didn't test every possible permutation of statistical tests in that survey, we tested, I don't remember how many exactly, but a lot of them. But there's a lot of really esoteric tests. And for those we sat around and argued and tried to pick a name together using the same theoretical strategy. The reality is that even in science, although we should test everything, you can't usually. So you make compromises.

Then we built it all and two weird things happened. Turns out some of the math in the actual real programming languages was wrong. So for example, and there's a library called Apache Commons T test, the T test was so wrong. That is this one value approach to infinity. It always returned 0.5. Now think about that. You got a drug approval, it always says 50/50. Oops. Seriously, that's been since fixed. The second, the anova, this compare means at a bigger scale is wrong in the programming language R. We haven't been able to get that fixed yet. We actually emailed the author but unfortunately he died. And so I don't know how that gets fixed. I'm hoping we can figure out how to get it, but I don't know how many papers are published with that equation, but a lot, it could be millions, I really don't know. And the way that the equation is wrong in the actual real programming language is wrong in a really funky way that it's really hard to tell like how wrong and under what circumstances for what kind of papers. So like that bothers me a whole lot that that's wrong. But nevermind that. And then we thought, let's try to make this for kids, right? Like maybe it doesn't have to be two years of a PhD program.

Maybe a kid can do this if we just clean it up like massively. So we took an attempt and on the screen here you can see we've turned statistics into little puzzle solving using good names. But other things too that we know about computer science education, that's in the academic literature. So you like, you can drag over blocks and of course because I'm obsessed with accessibility, we make all this stuff accessible too. Stuff like that. And then on the top here you see a pie chart, on the pie chart, I'll get into a little bit. But you can use this accessibly in addition to seeing it visually. If you as a a user that can see, look at that pie chart, you might think to yourself, I don't get it. What's the big deal? It's just a pie chart and you're right, to you. But if you're a kid for example and you're blind, you wouldn't know that you could see it, you would just interact with it in whatever way that you would expect. So it's almost hidden accessibility.

Under the hood, this is really complex and I won't get into all the details but you can kind of think about it at least in the first versions at this point in the story, it was like a super ugly car with an awesome engine under the hood for accessibility. Over time, that improved. But it had a whole bunch of stuff related to structure and meaning and different aspects about statistics that was built in. And we were testing that at schools all over the place to try to get that right. Funded by the National Science Foundation.

Here's just a couple brief examples of what this looks like today. They just look like charts. If you don't know any better and you squint, you can't tell that they're accessible in theory. But they've got, you know, beautiful little, I mean pretty little, hopefully pretty, little overlays with data and information. If you can't see the screen, you can use a screen reader, you can use braille devices, tablets, whatever you want. They meet as they say in the business the WCAG standard. But that's a whole other thing. Interestingly and the only one that I'll mention past this that was actually kind of tricky to get right is actually not the stuff related to blindness. It's actually color, color's super weird for accessibility because of the strange property that we know from this paper called Colorgorical. It turns out, as you get closer to the accessibility standards, you have to raise how much contrast there is between the colors. And weirdly, and this is strange to me, as you do that, how much people like it goes down. And that's kind of counterintuitive and it actually in a way explains why sometimes there is some small amount of resistance to color. So in the paper, they had a particular set of equations that they thought represented this well.

Now let me ask you, how many people find this pretty, how many people do not? Just by a show of hands. Yeah, so here's the thing, I didn't either. And anytime we showed this to people, regardless of the stuff that we would show, they'd say, why does this have to be so ugly? So I had an idea and it turns out the math behind the way that they did all these calculations was wrong. So I redid the math and this was the best I could do to meet the requirements but also allow a pie chart with this kind of many of colors. How many people find that better? Okay, it's that one guy in the back I'm gonna have to work on. But I mean, it's imperfect. But you can add like little spaces, you can add wedges that pop out. And so, you gotta kind of, sometimes you have to be a little creative for accessibility and that takes a little thinking.

Okay. Now at this point in the story, it's kind of like that Monty Python moment for now for something completely different because then something really strange happened across the entire planet. It turns out, state governments and eventually countries within a couple years, all started changing their laws. At the time, I was just doing accessibility work along with my team and many others in the literature to try to make this stuff better for everybody. You know, data science was confusing, it should not be confusing. I want the public to understand it, these kind of things. And I want any kid to be able to use it just because I don't know, that seems like the right thing to do. But at the same time, Maryland, the state disagreed with that. They said uh-uh, the tech companies have done a terrible job at this. So much of the materials for computer science education at the time was flagrantly inaccessible.

And it's not like some big conspiracy, it's just that there's no financial incentive around it, right? But now there is. So the first, to my knowledge, the first law that got passed in the US, it's complicated overseas, was this one, it's called SB-617. And at the time, the Maryland government was designed, unbeknownst to me, designing a data science sequence that was gonna be a year long sequence for kids to use. But they were using all the stuff that I'm talking about. Like they were using word choices that were no better than random words 'cause it was done in Python, that none of it was accessible. If you had a disability, you were just screwed. There was no chance that you could use this stuff. But the Maryland legislature said that's now illegal, right in the middle of their projects. So I got a panicked phone call and we redid all of the computer science and data science curriculum across the state of Maryland. If you live in Maryland, anybody live in Maryland? This is now free to you and all your kids can participate. And it takes into account all the data that's presented in this talk or at least a good chunk of it.

In general, this means the accessibility has to be automated by law. You can't really use stuff like LLMs for this, especially in data science because if it gets it wrong and it's like a serious thing, that's a real problem. Plus like the actual equations used for data science, they have to be right. Like they can't just be like a guess. LLMs are super useful for many things, just for this, it's horrifying, right?

So at this point, Maryland changes the law, within about two to three years, this is now the law in 30 countries. In the US, it's a regulation. It's a little complicated under the DOJ and actually they just pushed it back a year, but that's not the point. But seriously, within two years, this idea we were just screwing around with became the law in 30 countries, which I did not have on my bingo card. The European Union just passed it last year and Canada has technically had some rules around it since I believe 2001. But they got teeth in just December, like literally last December. So like this has changed very fast for something I didn't think would change in my lifetime.

But the problem is we had really just done a survey. So we'd done this big project with Maryland. We had in theory had all these charts and visualization we had, we thought the understandability of data science as an accessibility with from an accessibility lens, but also just in general. But we just ran surveys, that does not prove that this is actually helpful to people at all. So what we thought to do was let's run another test. Let's just put it to the test, let the cards fall where they may, we don't care.

So we had three kinds of tasks. There's a pick task, you pick a test, you code. We actually have people code in data science and then they have to interpret the output of data science. And I promise you've never, how many people have seen statistics output before? Oh boy, you're in for a treat. So well a treat, here, enjoy statistics. Maybe not, but I promise there'll be one more picture of my kids. In this, this is an example of coding. There's a bunch of code on the screen and they have to write some code related to a dataset.

This is the output. It's a completely unreadable from back there, but I'll just tell you it's a lot. This is the actual output from R. We chose R because it did the best in the survey. So we figured why even bother to test Python? It did, it doubled the number of placebo that it didn't beat. So why even bother to test its conventions? It could be that it's better, but I really doubt it. So the point is, is this big long list of stuff with all sorts of technical information that'll compare and contrast to what we did. First, we took all the text and we put it at the eighth grade reading level, which you can calculate. We couldn't get it below that. And then we left all the technical information, but we formatted it in like tables and like really common obvious things. And so it says stuff like a comparison between two samples was tested, first, the samples were tested for normality and all samples appear to be normal. Results show that there may not be a difference between group one and group two and then some technical information. I still don't think that that's necessarily easy to understand, but the hypothesis was just, is it easier? There may not be a way, for example to make all of statistics easy for all people.

That's okay if that's the result, but on the other hand, it doesn't have to be as bad as it is. So then we again add pesky humans and stir. Thank you. So at this point, we then vote. Does it help or does it not? Okay, so first, how many of you think all the design changes in this kind of stuff probably didn't make much of a difference? Go and vote with your hand. Okay, and how many say that it did?

Okay, so probably about 2/3.

So it turns out it's not magic, which is not surprising, it's science, but it's quite helpful.

It doesn't make it so things do better, but these green lines on the right, that's how much better it did. And one thing that's really interesting about this is if you look at these little blue lines, these are people that gave up at some point in time and we didn't even know to track that, but once we observed it, we're like, oh that's really interesting. So we think it might be related or correlated to anxiety or some other issues. And we did track that formally. So it looks plausible. But again, maybe thinking, right? Like so like it, you push the pendulum over and it looks like you can rule out a couple things but you can't rule out 'em all. So it's like a hypothesis is what I would say. Or maybe even a theory. But the point is this green line, if this was perfect, the green would extend all the way to the bottom. But people still give up. People still get it wrong, it's just they do quite a bit better and they're more sure about themselves.

Again, not magic, but we also took some qualitative data from this. I wanna read a couple of these quotations because I think they're quite funny. My favorite here, and this is from the R group, someone said, I swear I did not make this up. This is really what someone said. They said anxiety went down because I realized it's hopeless for me.

And I swear, I'm not even, well, although this is my favorite quotation, I'm not even exaggerating when I say all the qualitative feedback was this kind of stuff. Oh my goodness, I'm horrified. I don't know what to do. What does all this crazy stuff mean? Et cetera. But the problem is we didn't get all positive feedback either. And it turns out in our study, our survey, some of our naming conventions were still really confusing. So there's one that's called compare counts, which we had mapped to a particular test called chi squared. Details don't really matter. No one knew what that meant. The data showed that and the qualitative data showed the same thing, which means we got it wrong. Then past that, there was other things that got in the way that we didn't realize. For example, we use this word groups because in statistics, you often have groups of people that do stuff. No one had any idea what the groups meant. So even though we said what the test does, we had to be literal in a different way when we redesigned parts of the library after we did this test. Point being, it's science, right? When you do science you look for stuff and then oftentimes you don't get it perfect. That's just how science is, right?

Your robot crashes for a while, that's reality. But we found one last thing that I thought was really interesting.

Almost everyone, almost everyone could describe their data in a very specific mathematical way. And I won't get into the details and I only got the data back from a fourth experiment yesterday, so I can't talk about it yet. But I will tell you, let me say it this way, this is exploitable and you can make different kinds of designs that make data science dramatically easier. That like I don't wanna say fixes the rest of it, but it looks like it fixes the rest of it.

So in any case, there is some hope for data science. I really think we can get this all the way down to K through 12 and make it easier for kids. And that might help us all for reasons. But it turns out that the things that help good names an organization helps a lot. It's not everything, but it helps. You can make it much easier to code if you do this. You actually increase anxiety and confidence. But it's kind of complicated. All of it can be accessible, not just to the point of following the law, but like straight up. Like it really works for people with all kinds of disabilities in meaningful ways. And then people are more likely to engage with good designs. They actually try harder when they do this kind of stuff. And so I wanna leave you with a couple just thoughts about, I know, my daughter played at the Smith Center the other day and this was the best picture I could get. She's only 14. So you know, in any case, I wanna leave you with a couple things to at least consider back whatever your context is. One is kids need more training on science in school and there's certain things that are missing. One is methodology, evidence gathering, not just using a dataset. That's part of the problem.

If you look at something like the Computer Science Teacher Association Standards, you basically give kids a dataset at best, but they mostly don't engage in a 12 year curriculum with evidence at all. And that's really weird. That makes no sense. But they need to learn discovery from evidence. You'll notice anytime we put the data up here, I didn't say here's the answer. I said, hey, what's that? That's how science works.

I think it needs to be at the K through 12 level and should include kids actually reading academic papers. If we expect these kids to grow up to become lawyers and policy makers and educators and whatever, they need to actually be engaged in what scientists do. And we as scientists need to stop talking only to ourselves. We need to go outside. The problem I have, my doctor literally told me to take more vitamin D, but that's not the point, I'm here now. Scientists need to communicate their designs better, especially in data science because it impacts our drug approvals, our budgets, our any of those kind of things. And then the final one, although I don't wanna be the one to do it, we probably need more scientist policy makers at every level, not just federal, state, school districts. We just need people to know how this process works so that they can make decisions that are for all of us. And with that, thank you for your time. (audience applauding) Thank you Stefik. I love listening to you present. Do we have any questions from the group? All right, give me five minutes, I'll be there.

Thanks a lot, such a really well-rounded and really intuitive talk. One quick comment and then two questions to throw at you. One is, I loved the John Oliver's, false balance debate where he brought on 97 climate scientists to compete with against the one. So anyway, you made me think of that and I remember really loving and sharing that a lot. So the two questions are, one is about providing a framework for people to understand data science. And one of the questions I'm asking, and I'd love to hear how you talk about this, is how to organize these really big fields. So data science, statistics, machine learning, AI, math, you know, helping people understand the relationships between these and that there are differences and I'd love to hear how you talk about these big areas 'cause if you don't get the big picture right, sometimes the details don't have anywhere to hang in a mental model or a framework. The second question is, I'd love to hear your take on what the right foundation is for K through 12 students. So for example, should data science replace calculus in high school? I'm asked that a lot as well and I'd love to hear your take, thanks. Sure, I appreciate the questions a lot and I did in fact want to pull that style of joke from John Oliver, just don't say that because I think it's kind of funny how he does that.

But also too that that issue of like saying scientists agree or not isn't actually how science works at all. I don't particularly care whether my colleagues agree or not. I care what the data shows and that's really what the heart of what you're saying is, which I appreciate. So the first question was what again? It was the framework. Oh yeah. So really funny thing about all these areas, and it's really not intuitive in the slightest, but under the hood, all of those areas, I promise you it's really true, are just matrices, all of it, all those statistical formulas, matrices, AI, big matrix, data science, just a matrix. Every single one of those things boils down to a type of math called linear algebra.

Linear algebra is the study of matrices and how you manipulate them. Oh, by the way, computer graphics, matrices, shader programming for really fancy visual effects in movies. Yeah, matrices.

Drug approvals. Yeah, under the hood it's actually matrices. I promise you that that's really true. But the problem is by me saying that it doesn't really help you organize all those things. You tell a kid, oh yeah, well the AI is just a matrix. And they say, what does that mean? How do I use the matrix in this context? So that's tough. So like to me, when I talk about it, I often talk about it as like the backend of decision making. If you use data science or AI or some of those things, you can use them at different parts of the process of decision making. AI is really good for like searching and finding information and then you have to double check it just like you do with a Google search or Wikipedia for that matter, or even a published paper. But it's not good when you want to be more certain, right? Data science and statistics are interesting because the field doesn't agree. I was actually involved in a group called the ACM Data Science Task Force where we tried to figure out definitions for all this kind of stuff. But it turns out some statistics department literally scribbled out their name, wrote data science on top, and then called it a day, which is really unintuitive and doesn't help at all.

But to me that area is statistics. But with some code, that's how I think about it. Because oftentimes in statistics in the early 1900s, they didn't have that. So it was all about the symbols and the math, but that's less true today.

Okay, and then the second question was? Calculus versus data science. Ah. So it's been tried a couple times and I think that I would be a little bit wary of replacing it, but the items that have been tried, I believe there was a group in California that replaced algebra. Do not do that. That is a very bad idea. The data does not support that at all. Algebra is a very important skill for kids to learn for reasons I won't go into. Calculus is fuzzy. Calculus is absolutely necessary for fields like physics. You simply cannot do it without it. That's just how it works. And there's also some areas of data science where you really can't, and all those matrices under the hood, they also use a certain area of calculus to do it. But the question is, does it need to be in high school or college? And that I think is an unsettled debate. I don't think we know, my opinion on it is maybe, but I'm not sure, I can imagine an alternative. I would say though that the one thing that I think should change there is a thing in the college board statistics class, it feels like it was written in 1902, that needs to be updated to modern data science. It literally, actually, ironically, I'm not even joking, it literally stops at T, right?

Maybe even before that, it might stop at chi squared, which was I think late 1800s. So like it's pretty out of date in my opinion. But in any case, I hope that helps, but may not be a full answer. So we don't know for sure.

Well, one more round of applause. Thank you. Thank you Stefik. VEX Robotics Educators Conference.

Share

Like this video? Share it with others!

Additional Resources

Continue the conversation in the VEX Professional Learning Community!

Learn more about the VEX Robotics Educators Conference at conference.vex.com.