Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Tutor Has No Teaching Degree
Transcript
- Lucas: So earlier this month, Khan Academy rolled out a significant update to Khanmigo — their AI tutoring assistant powered by GPT-4. The headline feature is that it can now 'see' what students are working on in real time, not just respond to typed questions. Luna: And I think a lot of parents and teachers hear 'AI tutor' and immediately imagine something like a patient, infinitely knowledgeable helper who never gets tired. Lucas: Right. The marketing leans into exactly that fantasy. But here's the thing: Khanmigo has been in classrooms since 2023, and early research — including a study from the University of Pennsylvania's Graduate School of Education — found that it gives incorrect answers about 26 percent of the time. Luna: Twenty-six percent. That's not just a rounding error. That's nearly one in four interactions where a student is being actively misled. Lucas: And the mistakes aren't always obvious. Sometimes the AI will be confidently wrong about a historical date, or it'll misapply a math formula in a way that a fifth grader wouldn't catch. The student walks away having learned something incorrect. Luna: Which is exactly the opposite of what a tutor is supposed to do. Lucas: Exactly. And what's interesting is that the problem isn't just about accuracy. It's about pedagogy. A good teacher doesn't just know the subject — they know how to scaffold a concept, how to ask a question that leads a student toward the answer instead of giving it away, how to detect confusion in a furrowed brow or a long pause. Luna: Right. An AI can't read a room. It can't see that a student is zoning out or feeling frustrated. Lucas: And yet, school districts across the country are rushing to adopt these tools. In Newark, New Jersey, for example, the district rolled out Khanmigo to every middle school student last fall — about seven thousand kids — as part of a 'personalized learning' initiative. The superintendent called it a game-changer. Luna: But what evidence was there that it actually improves learning outcomes? Did they run a controlled trial first? Lucas: They did not. The district relied on a pilot study from the previous year that showed improved engagement — not test scores, not retention, but engagement. Students spent more time on the platform. Luna: Which is a classic ed-tech trap. More screen time doesn't equal better learning. I remember reading about the Los Angeles iPad fiasco in 2013 — they spent over a billion dollars on devices and curriculum, and test scores actually dropped. Lucas: It's the same pattern. A shiny new tool, a compelling narrative about closing achievement gaps, and then a decade later we find out it didn't work. The difference this time is that AI is far more opaque. When an iPad gave a wrong answer, you could trace it back to a buggy app. With a large language model, nobody really knows why it said what it said. Luna: And that opacity is dangerous in a classroom. Because if a teacher can't understand why the AI gave a particular answer, they can't correct it — and they certainly can't teach the student how to think critically about the AI's output. Lucas: Right. The best case I've seen for the limits of AI tutoring comes from a study at Stanford's Graduate School of Education. They asked Khanmigo to tutor a student on the concept of negative numbers. The AI correctly explained that negative numbers are less than zero. Then the student asked, 'So is negative five bigger than negative two?' And the AI said yes. Luna: Oh no. That's a fundamental misunderstanding. Lucas: It is. And the AI didn't catch its mistake until the researcher pointed it out. The student, meanwhile, had absorbed a wrong rule. That's the kind of error that can compound for years. Luna: It reminds me of something I read about the 'Socratic method' in AI tutoring. Some companies claim their systems use Socratic questioning, but when you actually interact with them, they just ask 'Why do you think that?' over and over. It's a parlor trick, not pedagogy. Lucas: Yes. Real Socratic teaching requires understanding the student's current mental model, choosing the right question to expose a contradiction, and timing that question so the student is ready to hear it. That's incredibly hard. Most human teachers struggle with it. And yet we're expecting a statistical model to do it fluently. Luna: And that gets to the deeper issue. We're not just talking about a buggy product. We're talking about a fundamental mismatch between what AI does — pattern matching on vast amounts of text — and what teaching requires. Lucas: Right. Teaching isn't just information transfer. It's relationship, it's motivation, it's knowing when to push and when to step back. The AI has no theory of mind. It doesn't know that the student is tired, or anxious about a test, or distracted by something at home. Luna: And yet the business case for AI tutors is built on the idea that they can replace or augment human teachers, especially in under-resourced districts. If you're a school in a low-income area with large class sizes, an AI tutor that costs a few dollars per student per year looks like a miracle. Lucas: It does. And I don't want to dismiss that appeal. There are real problems with access and equity in education. But the solution can't be a tool that sometimes teaches kids the wrong thing. Especially when the kids who are most likely to use it are the ones with the least access to a human tutor who can correct the AI's mistakes. Luna: That's a really important point. The students who could most benefit from personalized help are also the most vulnerable to being misled. And they're the least likely to have a parent at home who can say, 'Actually, honey, negative five is smaller than negative two.' Lucas: And so we're essentially running a large-scale experiment on the most vulnerable students without their consent. The Newark district didn't ask parents if they wanted their kids to be part of an AI tutoring trial. They just rolled it out. Luna: Which is exactly the kind of ethics problem we've covered before — from predictive policing to credit scoring. The people affected by the algorithm often have no say in its deployment. Lucas: Exactly. And it's not just Khanmigo. There are dozens of AI tutoring products out there now — from Carnegie Learning's MATHia to Squirrel AI in China. The market is projected to hit $4 billion by 2027. And very few of these products have been rigorously tested in randomized controlled trials. Luna: There's actually a great meta-analysis from the Journal of Educational Psychology last year that looked at 50 studies of AI tutoring systems. The overall effect on learning was positive but small — an effect size of about 0.2 standard deviations. That's less than the effect of reducing class size by five students. Lucas: So the evidence suggests that, at best, these tools offer a modest benefit — and at worst, they can actively harm. And yet the hype cycle keeps accelerating. Sal Khan, the founder of Khan Academy, has said publicly that he believes AI tutors will 'democratize education' and 'provide every student with a personal tutor.' Luna: And he genuinely believes that. But intentions don't matter if the product doesn't work. The road to bad ed-tech is paved with good intentions. Lucas: That's a perfect way to put it. And I think part of the problem is that we're asking the wrong question. Instead of 'Can AI tutor a student?' we should be asking 'Under what conditions, with what safeguards, and for which students might AI tutoring be helpful?' Luna: Right. And that's a much harder question to sell to investors or school boards. It doesn't fit on a slide. Lucas: No, it doesn't. It requires nuance, piloting, iterative design, and a willingness to say 'this isn't ready yet.' And in the current funding environment, nobody gets rewarded for saying that. Luna: Speaking of which — this is the kind of conversation we're able to have because this show is listener-supported and ad-free. We don't have to sell you anything, and we don't have to sugarcoat the challenges. If you find value in episodes like this one, you can support us at buy me a coffee dot com slash fexingo. It's a simple way to keep the show independent. Lucas: Yeah, and we really appreciate those who do. It makes a difference. Luna: So back to the educational piece — I think there's also a question about teacher training. Even if the AI is accurate, teachers need to know how to integrate it effectively. And most schools aren't providing that training. Lucas: That's a great point. A survey from the EdWeek Research Center last year found that only 13 percent of teachers said they had received any professional development on using AI in the classroom. So we're handing them a powerful — and flawed — tool and saying 'figure it out.' Luna: And that's unfair to teachers, too. They're already stretched thin. Now they have to learn how to audit an AI's outputs, decide when to override it, and explain to students why the AI might be wrong. Lucas: Right. So the burden is being placed on the people with the least power in the system: students and classroom teachers. Meanwhile, the companies selling these tools are making promises they can't keep, and the districts are buying them with public money. Luna: I want to be clear: I'm not anti ai in education. I think there are specific use cases where AI could be genuinely helpful — like providing writing feedback on grammar and structure, or generating practice problems for students who need extra drill. But that's a far cry from claiming the AI is a tutor. Lucas: Exactly. It's about honest labeling. If Khanmigo were marketed as 'an ai powered homework helper that can sometimes make mistakes,' that would be one thing. But calling it a 'tutor' implies a level of pedagogical expertise it simply doesn't have. Luna: And it also sets up unrealistic expectations. When the AI fails, students blame themselves. They think they're not smart enough to understand — when really, the AI just gave them bad information. Lucas: That's the insidious part. The AI doesn't know when it's wrong, and the student doesn't know to doubt it. It's a perfect recipe for reinforcing misconceptions. And once a student has internalized a wrong rule, it takes a lot of effort to unlearn it. Luna: So where does that leave us? Are we saying schools shouldn't use AI tutors at all? Lucas: I think we're saying they should proceed with extreme caution. Any deployment should be accompanied by independent evaluation, transparency about error rates, and a plan for what happens when the AI gets it wrong. And above all, we need to stop pretending that AI can replace the human relationships at the heart of teaching. Luna: Because at the end of the day, learning is a social process. You can't automate that. Lucas: You can't. And the sooner we stop trying, the sooner we can build tools that actually help teachers and students — instead of just selling them a fantasy.