Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Tutor Has No Teaching Degree
Transcript
- Lucas: There's a moment from a classroom in Los Angeles earlier this spring that I keep coming back to. A seventh-grade student asks her AI tutor about the French Revolution, and the tutor responds confidently with a string of plausible-sounding claims — wrong dates, wrong cause, wrong key figures. Luna: So the AI basically hallucinated a lesson. And the student wrote it down in her notes. Lucas: Exactly. This wasn't some experimental chatbot in a lab. It was Khan Academy's Khanmigo, deployed in a real middle school history class. The teacher told reporters the system had been used daily for three weeks. Nobody caught the error until a parent spotted it while helping with homework. Luna: That's alarming because the whole pitch of these AI tutors is that they adapt to each student and free up the teacher to focus on deeper instruction. But if the tutor is feeding wrong facts, you're not freeing the teacher — you're just scaling misinformation. Lucas: Right. And the deeper issue here isn't just one bad answer. It's the assumption that a system built to generate fluent text is also built to teach. Large language models like the ones powering Khanmigo are optimized for next-word prediction, not for curriculum alignment or developmental appropriateness. Luna: So the AI doesn't actually know it's teaching seventh-grade world history. It's just producing text that looks like a history lesson. Lucas: That's exactly the gap. A human teacher with a teaching degree has spent years learning how to sequence information, how to anticipate misconceptions, how to gauge whether a student actually understood. The AI has none of that. It has training data that includes millions of web pages about the French Revolution, sure, but also Reddit threads, fan wikis, and — let's be honest — plenty of confidently wrong material. Luna: And the model doesn't have a mechanism to say 'I don't know.' It's designed to always produce an answer. So when it's uncertain, it fabricates. Lucas: A 2025 study from Stanford's Graduate School of Education tested a leading AI tutor against certified math teachers. The AI outperformed the teachers on procedural questions — 'solve for x' — but when the questions required conceptual explanation, like 'why does the quadratic formula work,' the AI gave answers that were technically correct but pedagogically useless. The students didn't learn the underlying concept. Luna: That's the difference between teaching and just telling. A good teacher reads the room. They see a student's face scrunch up, they rephrase the explanation. The AI can't do that. Lucas: And there's a really interesting recent study from February 2026 at the University of Michigan. Researchers gave two groups of college students identical course material, but one group used an AI tutor for practice problems while the other used a traditional textbook with answer keys. On the final exam, the AI tutor group scored 11 percent lower on questions that required application — not just recall, but applying a concept to a new scenario. Luna: So the AI tutor helped them memorize, but not understand. That's a huge red flag for how these tools are being marketed as personalized learning. Lucas: The marketing language is powerful. 'One-on-one tutoring for every student.' 'Unlimited patience.' 'Adaptive to each learner's pace.' And those are real benefits — for some kinds of practice. But the claim that an AI tutor can replace or even substantially augment a human teacher for conceptual instruction is not supported by the evidence we have so far. Luna: Let's talk about the business side. School districts are making procurement decisions based on these promises. How many districts have actually bought into AI tutors right now? Lucas: According to a survey from the Consortium for School Networking released in March 2026, roughly one in four U.S. school districts has piloted or adopted an AI tutoring platform. That's up from about one in ten the year before. The total spending is estimated at around $200 million for the 2025-26 school year. Luna: And who are the main vendors? Lucas: Khan Academy is the most visible because Khanmigo is free for teachers — funded by philanthropic grants. But there's also Carnegie Learning's MATHia, which has been around for years but recently added a generative AI layer. And then there are the big tech entrants: Google's 'Practice Sets' inside Classroom, and Microsoft's 'Copilot for Education' which embeds an AI tutor into Teams for Education. Luna: So it's not just one company. It's a whole ecosystem moving fast. And the common thread is that none of these systems are required to demonstrate pedagogical competence before they get used in classrooms. Lucas: That's the regulatory gap. K-12 software that handles student data has to comply with FERPA, the Family Educational Rights and Privacy Act. But there is no federal standard for educational efficacy of software. No equivalent of the FDA requiring clinical trials before a drug goes to market. Luna: A school district can buy an AI tutor based on a slick demo and a few case studies. There's no requirement to prove that students actually learn better. Lucas: And when something goes wrong — like the French Revolution error — who is accountable? The teacher for not catching it? The district for purchasing it? The developer for building it? Right now, the answer is unclear. The LA incident didn't result in any lawsuit. The district simply paused the program and asked teachers to review all ai generated content before students see it. Luna: Which puts the burden back on the teacher, who already has thirty students and a packed curriculum. So the AI tutor doesn't reduce their workload — it adds a layer of supervision. Lucas: Exactly. And this is a theme we've seen across other AI ethics episodes: the person lower on the totem pole ends up doing the quality control. The bail judge has to check the AI's risk score. The doctor has to override the algorithm. The teacher has to fact-check the tutor. Luna: It's worth noting that some developers are trying to address this. I saw that Khan Academy recently published a transparency report about Khanmigo's accuracy, and they claim that for math problems, the correct answer rate is over 95 percent. But for open-ended history questions, it's closer to 80 percent. Lucas: Eighty percent means that one in five answers is wrong or misleading. In a classroom of thirty students using the tutor daily, that's six students getting bad information every day. Over a semester, that adds up. Luna: And those errors aren't random. The same Stanford study I mentioned earlier found that the AI tutor was more likely to give incorrect answers to students who asked questions in non-standard English, or who used simpler vocabulary. So the students who already struggle are the ones most likely to be misled. Lucas: That's a bias amplification effect. The students who need the most accurate information are the ones getting the worst of it. And because the AI sounds confident and authoritative, they're less likely to question it. Luna: Look, I want to pause on something here. This kind of deep dive into real-world AI failures is exactly why listeners tell us they value this show. And it's why a small number of people choose to support us directly. Lucas: Yeah, it's not something we talk about much, but Fexingo stays ad-free because a handful of listeners chip in monthly through buy me a coffee dot com slash fexingo. That's literally what funds making this many episodes possible. Luna: If today's conversation gave you something useful — maybe a new angle on the AI tutor you've heard about — that's exactly the kind of thing that keeps us going. No pressure, just a quiet acknowledgment that listener support makes this sustainable. Lucas: Alright. Back to the classroom. Another case I want to bring up is from a pilot program in Georgia last fall. A high school used an AI tutor for English literature. The system was supposed to help students analyze Shakespearean sonnets. But it consistently misattributed sonnets to the wrong themes — calling Sonnet 18 'about mortality' when any English teacher knows it's about love and immortality through verse. Luna: So the AI was teaching literary analysis that was just factually wrong. And again, the students didn't know. Lucas: The teacher in that class told me she spent the first ten minutes of every class correcting the AI's previous day's output. She said it was like having a student teacher who never learned the material. Luna: And that's the human cost. Teachers are already stretched. Now they're supervising an AI that's supposed to help them, but actually creates more work. Lucas: There is a policy response emerging. The National Education Association passed a resolution in January 2026 calling for all AI tools used in K-12 classrooms to undergo independent efficacy testing before deployment. They want something like a 'seal of pedagogical assurance' from the Department of Education. Luna: Is that actually gaining traction? Or is it just a resolution? Lucas: It's early. But two states — California and New York — have introduced bills that would require any AI tutoring platform used in public schools to submit to a review by a state board of educators. The bills are modeled on the EU AI Act's high-risk classification for education. They haven't passed yet, but they've gotten bipartisan support. Luna: So there is a path forward. It's just not moving as fast as the technology is being adopted. Lucas: That's almost always the story with AI ethics. The deployment curve is steep, and the oversight curve is flat. What gives me some hope is that the conversation is shifting from 'can AI tutor?' to 'should AI tutor — and under what conditions?' That's a more productive framing. Luna: Because the answer isn't never. There are real applications where AI tutoring helps. Drill practice, vocabulary, repetitive problem sets where the student just needs more reps. But conceptual teaching, critical thinking, interpretation — those are still human domains. Lucas: The best version I've seen is what some schools call 'flipped classroom 2.0.' The AI handles the lower-level practice, freeing the teacher to spend class time on discussion and projects. But the AI is never the primary instructor. It's a tool, not a replacement. Luna: That requires clear boundaries and training for teachers. And transparency for parents about what the AI does and doesn't do well. Lucas: And accountability when it fails. If we're going to put AI in the classroom, we need to know who's responsible when a student learns something wrong. Right now, that responsibility is diffused across vendors, districts, and teachers. That's not a recipe for trust. Luna: So where does that leave us? If I'm a parent listening, what should I be asking my child's school? Lucas: Three questions. One: what is the AI tutor being used for — drill and practice, or conceptual instruction? Two: what is the error rate for the specific subject and grade level? Three: what is the teacher's role in vetting the AI's output? If the school can't answer those clearly, they're not ready to use it. Luna: Good framework. And for school administrators listening, maybe the question is: are you buying a tool for students, or are you buying a tool that looks good in a board presentation? Lucas: That's the uncomfortable question. Because the pressure to adopt AI is real — from vendors, from parents who read about it, from the desire to seem innovative. But the evidence we have right now says that a poorly implemented AI tutor is worse than no tutor at all. Luna: And the students who pay the price are the ones who can least afford to lose a year of learning. Lucas: Exactly. So we need to get this right. Not fast — right.