Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Professor Grades You Invisible
Transcript
- Lucas: Luna, I want to start with a number that stopped me cold. A 2025 study from Stanford's Center for Education Policy Analysis found that non-native English speakers receive grades that are on average 15 percent lower on automated essay scoring systems than human graders give them for the exact same essay. Luna: Fifteen percent — that's nearly a full letter grade in most college rubrics. And this isn't a hypothetical; these systems are being used at scale right now. Lucas: Right. The study analyzed over 50,000 essays submitted across five large public universities. The automated scoring engines — built by companies like Turnitin, McGraw-Hill, and smaller edtech startups — consistently flagged certain sentence structures, vocabulary choices, and idiomatic phrasing as 'non-standard' and docked points. Luna: And by 'non-standard,' they really mean 'different from the training data,' which was overwhelmingly essays from native English speakers at selective universities. Lucas: Exactly. The training data itself is the problem. Most commercial grading AI is trained on a corpus of several hundred thousand essays, but the distribution is heavily skewed. One dataset I saw from a major vendor was 82 percent essays from students at institutions that are less than 15 percent non-native enrollment. So the model learns that 'good writing' looks a particular way — and penalizes anything that deviates. Luna: So a student from Mumbai who writes 'the reason for this phenomenon is multifold' — which is perfectly idiomatic in Indian English — might get flagged for non-standard phrasing. Lucas: That's exactly the kind of case. The Stanford researchers found that the penalty was most severe for students whose first language was Mandarin, Spanish, or Arabic. But even students from non-standard English dialects — African American Vernacular English, for example — saw a measurable penalty. Luna: And what did the universities do with these scores? Were they used for high-stakes decisions? Lucas: That's the disturbing part. In a 2024 pilot at Arizona State University, automated essay scoring was used as the primary grade for first-year composition courses — not just a diagnostic tool. The pilot covered about 3,000 students. And when researchers later compared the automated scores to human-graded evaluations, they found that first-generation students were 23 percent more likely to receive a failing grade from the AI than from a human instructor. Luna: Twenty-three percent. So a student who might have passed with a C- from a human gets an F from the algorithm. That changes their academic trajectory — maybe they lose a scholarship, get put on probation. Lucas: And it's not just Arizona State. Georgia State University, University of Texas at Austin, and several community college systems have been piloting similar tools. The pitch from vendors is efficiency — grade 50 essays in two minutes, free up faculty time. But the cost is borne by the students who don't fit the model's narrow definition of good writing. Luna: Is there any regulation requiring transparency here? Like, do students know they're being graded by an algorithm? Lucas: That's a critical point. In most cases, no. Arizona State's pilot did include a disclosure in the course syllabus, but in a small print note. Many institutions don't disclose at all. The vendor contracts often include confidentiality clauses — universities aren't allowed to reveal which specific model they're using or how it's calibrated. So students can't even appeal the grade in a meaningful way. Luna: So it's a black box. The student gets a low score, they assume they wrote poorly, and they don't know the evaluator is biased against their linguistic background. Lucas: And that's where the ethics get really uncomfortable. Because the algorithm isn't just grading — it's shaping student behavior. If a non-native speaker learns to write in a way that pleases the algorithm, they're essentially being trained to mimic a narrow, privileged style of English. The AI is enforcing a linguistic monoculture. Luna: That's a different kind of bias than what we usually talk about, like racial or gender bias in hiring algorithms. This is about cultural and linguistic identity. And it's happening in a space — education — that's supposed to be about expanding opportunity. Lucas: A couple of dollars a month is genuinely what keeps these going — buy me a coffee dot com slash fexingo, if you've gotten something out of them. Luna: Yeah, listener support is what keeps this show ad-free and independent. It's a small thing that makes a real difference. Lucas: And it means we can keep digging into stories like this one. So back to the grading algorithms — there are some potential fixes, but they require a fundamental shift in how these tools are built. Luna: What would that look like? More diverse training data? Lucas: That's part of it. But the Stanford researchers argue that the real fix is moving away from a single 'correct' answer. Instead of training an AI to match a single human grader's score, you train it on a distribution of human scores — say, ten human graders per essay, capturing the range of reasonable evaluations. Then the AI outputs a range, not a single grade. Luna: So instead of saying 'this essay is a B-minus,' the system says 'this essay is in the B-minus to B-plus range, depending on the reader.' That gives the student and instructor more context. Lucas: Exactly. Some edtech researchers call this 'calibrated peer review with AI augmentation.' The AI doesn't replace the human — it surfaces patterns and flags outliers for human review. But that's not the business model most vendors are selling. They want to sell a replacement, not an augmentation. Luna: And the cost savings argument is powerful for cash-strapped universities. But at what cost to equity? Lucas: There's a second issue too: temporal bias. The training data for many commercial grading AIs was collected between 2010 and 2018. That means the model is essentially frozen in a particular moment of what 'good writing' looked like. But language evolves. The way students write now — even native speakers — includes abbreviations, references to digital culture, different rhetorical patterns. Luna: So a student who writes a perfectly coherent essay but uses em-dashes or sentence fragments for rhetorical effect could be penalized because the training data didn't see that as standard. Lucas: Right. And this isn't just a theoretical problem. The study found that even native English speakers saw a modest penalty — about 3 percent — if their writing style was more conversational or used informal structure. So the AI is actually narrowing the definition of acceptable academic writing. Luna: What about feedback? One of the promises of AI grading is that it can give students instant, detailed feedback. Is that happening? Lucas: It varies wildly. The best systems do provide sentence-level feedback — 'this sentence could be clearer,' 'consider adding a transition here.' But the feedback is itself generated by a language model, and it can be unreliable. The Stanford study found that about 20 percent of feedback comments from one commercial system were either irrelevant or incorrect. So a student might be told to 'add more evidence' to a paragraph that actually contains three relevant citations. Luna: So the feedback loop reinforces bad habits or confuses students. And again, no human in the loop to catch it. Lucas: There is some positive movement. The European Union's AI Act, which we talked about in episode four, classifies educational AI as 'high risk.' That means providers have to do conformity assessments, ensure human oversight, and allow for appeals. But in the US, there's no equivalent regulation. A few states — California and New York — have introduced bills that would require disclosure when AI is used in grading, but none have passed yet. Luna: So for now, it's really up to individual universities to decide whether to use these tools responsibly. What's the best practice we've seen? Lucas: The University of Michigan has a policy that any AI grading tool must be validated against a diverse sample of student writing from within that university — not just a generic national corpus. They also require that AI grades be treated as 'advisory,' not final. And they publish an annual transparency report showing disaggregated outcomes by language background, race, and first-generation status. Luna: That's the gold standard — transparency, validation, and human oversight. But how many universities are doing that? Lucas: Very few. When I looked into this, I found that of the 30 largest public university systems in the US, only four have any kind of published policy on AI grading. Most are just using vendor tools without independent evaluation. And the vendors, understandably, push back on releasing their evaluation data because they consider it proprietary. Luna: So we're in a situation where a technology that can systematically disadvantage certain groups is being deployed without the safeguards we'd expect for any high-stakes assessment. Lucas: And it's happening at scale. The global market for AI in education is projected to reach $20 billion by 2027, and automated grading is a major component. The adoption is only going to accelerate as universities face budget pressure and growing enrollment. Luna: What can a student do if they suspect they've been unfairly graded by an AI? Lucas: First, check the syllabus — some universities do disclose it. If you think the grade is wrong, ask for a human re-evaluation. Many departments have formal appeal processes. And if they don't, you can file a complaint with the university's academic integrity office or even the Department of Education's Office for Civil Rights, arguing that the AI's bias constitutes discrimination on the basis of national origin. Luna: That's a heavy burden to put on students, especially first-generation or international students who may not know their rights. Lucas: It absolutely is. And that's why this needs to be a faculty and administration issue, not an individual student responsibility. Faculty need to demand transparency from vendors. Administrators need to require impact assessments before deploying these tools. And accreditation bodies should include AI grading practices in their reviews. Luna: Any sign that's happening? Are accreditors paying attention? Lucas: The Southern Association of Colleges and Schools, which accredits over 800 institutions in the US South, issued a draft standard last year that would require member institutions to 'ensure that automated assessment tools do not systematically disadvantage any student population.' It hasn't been adopted yet, but it's a sign the conversation is moving. Luna: That's encouraging. But until that's the norm, we're going to keep seeing cases like Arizona State. Lucas: Yeah. And that's the thing about AI ethics — it's rarely about malicious intent. It's about deploying a tool without understanding its blind spots. The grading AI doesn't hate non-native speakers. It just doesn't know what it doesn't know. And the students pay the price for its ignorance. Luna: That's a good place to leave it. Thanks, Lucas. Lucas: Thanks, Luna. Talk next time.