Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Professor Grades With Hidden Shortcuts
Transcript
- Lucas: A story out of Georgia State University this spring that I think is going to stick with a lot of us. The business school there rolled out an AI grading assistant for undergraduate essay assignments. The selling point was efficiency — get grades back to students in hours instead of weeks. Luna: I remember seeing that announcement. They said it was trained on thousands of past student essays and professor feedback. Lucas: Right. And the early results looked good on paper. Faster turnaround, consistent scoring. But then students started noticing something weird. Essays that used very specific keywords — even if the reasoning was shallow — were getting high marks. And more creative, original arguments were getting penalized. Luna: So the AI basically learned to reward formulaic writing and punish thinking outside the box, because that's what the training data showed. Lucas: Exactly. And here's where it gets worse. The system also showed a clear bias against non-native English speakers. Essays with slightly unconventional phrasing — grammatically correct but not idiomatic — were consistently scored lower. Lucas: I want to pause on that for a second, because if today's conversation gives you something to think about — maybe it's worth the price of a coffee. If so, there's a link at buy me a coffee dot com slash fexingo. Listener support is what lets us keep this show ad-free and digging into stories like this one. Luna: Yeah, it's a small thing that adds up. And we appreciate it. Lucas: So back to Georgia State. The bias against non-native speakers wasn't just about vocabulary. The AI had learned from past grading patterns that professors themselves — human graders — were slightly harsher on certain sentence structures. The model just amplified that. Luna: So the AI didn't invent the bias. It inherited it from the human raters whose work it was trained on. But then it applied it at scale, to every essay, consistently. Lucas: That's the danger. A human professor might subconsciously mark down a non-native essay by a few points. The AI does it to every single one, every time, without the fatigue or self-awareness a human might develop. Lucas: The university did have an appeal process — students could request a human re-grade. But here's the catch: very few students knew they could do that, and the process took weeks. By the time the grade was corrected, the semester was almost over. Luna: So the safeguard existed on paper but wasn't really accessible. That feels like a recurring theme in these stories. Lucas: It is. And it raises a bigger question about why universities are rushing to adopt these tools. The answer is usually budget pressure. Grading essays is labor-intensive. Adjunct professors are overworked. Administrators see AI as a way to do more with less. Luna: But at what cost? If the AI is systematically disadvantaging certain groups of students, the savings in time or money are coming out of equity. Lucas: Let me give you a number. A 2024 study from Stanford's Center for Education Policy Analysis looked at five different commercial AI grading tools. They found that on average, the tools had a 12 percent higher error rate for essays written by Black students and a 9 percent higher error rate for essays by Hispanic students compared to white students. Luna: That's a massive gap. And it's not just about race — the same study found higher error rates for low-income students, regardless of race. Lucas: Correct. Because the training data tends to come from well-funded school districts and universities that have the resources to archive past essays. The model learns what 'good' looks like from a narrow slice of the population. Lucas: Now, Georgia State has paused the program since the complaints came out. They're doing an audit. But the thing is, dozens of other universities are using similar tools right now, and most of them aren't being audited. Luna: Is there any regulation that requires these tools to be tested for bias before they're used on students? Lucas: Not really. The Department of Education has issued guidance, but it's non-binding. There's no federal law that says an AI grading system has to pass a fairness test before it touches a student's GPA. Luna: So it's essentially a voluntary compliance landscape. And we know how that usually goes. Lucas: Some states are starting to move. California has a bill in committee right now that would require any AI used in public schools to undergo an annual bias audit. But it's early. Lucas: I think the deeper issue here is about what we're optimizing for. When a university adopts an AI grading assistant, the stated goal is efficiency. But efficiency for whom? The student gets a faster grade, but a less accurate one. The professor gets less grading work, but less insight into how their students are actually learning. Luna: And the administrator gets a metric they can report to the board — 'we've reduced grading turnaround by 80 percent' — without having to report the accuracy gap. Lucas: Exactly. And I don't want to sound like I'm against using AI in education entirely. There are tools that help with personalized tutoring, or with identifying students who are falling behind. Those can be genuinely beneficial. Luna: But grading is a different category. It's evaluative. It has consequences for students' futures — scholarships, grad school admissions, even job offers. Lucas: Yes. And when the evaluator is a black box, students have no way to understand why they got the grade they did. They can't learn from the feedback because the feedback is often just a score, not an explanation. Lucas: One student at Georgia State told the student newspaper that she got a 68 on an essay she'd written on the ethics of AI. She was confused, so she went to her professor. The professor said the AI flagged her essay for 'insufficient keyword density.' The phrase she used instead of the expected term was considered a synonym, but the model didn't recognize it. Luna: So she wrote about the ethics of AI, and the AI grading her essay couldn't handle her vocabulary. There's some irony there. Lucas: A lot of irony. And it's a perfect example of the problem. The model was trained on a corpus of essays that used very specific language. If you deviated, you were penalized, even if your writing was better. Luna: So what's the fix? Should universities stop using AI for grading altogether, or can these systems be improved? Lucas: I think the honest answer is that right now, the technology isn't reliable enough for high-stakes grading. It can be useful for low-stakes check-ins — practice quizzes, draft feedback — but not for final grades that go on a transcript. Lucas: If universities do want to keep using it, they need to do three things. One: train the models on diverse data that reflects the actual student population. Two: run regular bias audits and publish the results. Three: make the appeal process simple, fast, and well-publicized. Luna: And the fourth thing, I'd add, is transparency. Students should know when an AI is grading them, and they should have a right to a human reviewer. Lucas: That's the core of it. Informed consent. Right now, most students don't even know their essays are being read by a machine. That's not just a bias problem — it's a fairness and dignity problem. Luna: One last question. Do you think the Georgia State story will change anything, or will it be forgotten once the news cycle moves on? Lucas: I think it depends on whether other universities pay attention. The technology companies selling these tools are very good at downplaying the risks. But if a few more stories like this get traction, and if state legislators start asking questions, we might see real change. Luna: Fingers crossed. Because students deserve better than an algorithm that doesn't understand them. Lucas: They do. And I think that's the note to end on. The technology is here to stay, but the question is whether we shape it to serve students, or whether we let it shape them.