Latest / AI Ethics with Fexingo: Bias, Safety, and Responsible Artificial Intelligence / When Your AI Art Teacher Fails to Recognize Your Style
Transcript
- Lucas: You've probably seen those AI art generators that can whip up an image from a sentence prompt. But what about an AI that grades your art? I'm not talking about a multiple-choice test on art history – I'm talking about an AI that evaluates the creativity and technique of a student's original artwork. Luna: That sounds like a minefield. How do you quantify something as subjective as artistic expression? Lucas: Exactly the question. And there's a real case from a university's digital arts program that shows just how messy it gets. Last semester, they piloted an AI grading tool for a 300-level digital painting course. The AI was trained on a dataset of professional digital artworks – mostly from popular illustrators on social media, which skewed toward a certain polished, high-contrast style. Luna: So it learned that 'good' art looks a specific way. What happened? Lucas: Well, a student named Maria submitted a series of works with loose, impressionistic brushstrokes – deliberately unfinished edges, visible texture. Her style was more like an oil painting than the crisp vector art the AI had seen. The AI gave her an average score of 62 out of 100. The human teaching assistant, who graded the same works blind, gave her an 88. Luna: That's a twenty-six point gap. Did the system flag any specific issues? Lucas: It did. The AI's rubric included criteria like 'edge definition', 'color harmony', and 'composition balance'. On edge definition, Maria's works scored low because the AI interpreted soft edges as lack of precision. Color harmony was also penalized because she used muted, earthy tones rather than the vibrant palettes common in the training set. Luna: So the AI was basically rewarding a narrow aesthetic. If you're a student who experiments with muted tones and loose strokes, you're automatically at a disadvantage. Lucas: Right. And here's the thing – the course instructor had specifically encouraged students to explore different styles. Maria's approach was exactly what the course was trying to foster. But the AI couldn't tell the difference between a deliberate artistic choice and a technical mistake. Luna: What happened when she appealed? Did the university have a process for that? Lucas: They did – but it wasn't straightforward. The AI's output was presented as a 'suggested grade', and the teaching assistant was supposed to review it. But in practice, the TA told me that about 70 percent of the time, they just accepted the AI's score because it felt objective and they were short on time. Maria's case became a test. Luna: So the TA had to manually override. Did they change the system? Lucas: Eventually, the instructor reviewed all the AI grades and adjusted them. But the damage was done – Maria had already received a midterm report with that 62, and it shook her confidence. She told me she started second-guessing her style, wondering if she should just paint what the AI wanted. Luna: That's the real risk. Not just a bad grade, but chilling the very creativity you're trying to teach. Lucas: Yeah. And if today's conversation gave you something to think about, here's the thing – these episodes stay ad-free because listeners choose to support them. You can do that at buy me a coffee dot com slash fexingo. That's buy me a coffee dot com slash fexingo. No pressure, just if the show is useful to you. Luna: It's a small way to keep conversations like this going without ads or sponsors. Lucas: Exactly. So back to Maria's story – the university eventually decided to retrain the AI on a more diverse dataset that included sketches, paintings, and works from different cultural traditions. But the process revealed a deeper issue. Luna: Which is that art grading is inherently subjective. Even humans disagree. So how do you design an AI that's fair? Lucas: One approach is to use multiple human raters to create a consensus baseline, then train the AI to match that consensus. But that's expensive and time-consuming. The university tried a cheaper method: they fed the AI examples of student work from previous semesters along with the final grades given by the instructor. Luna: But that just replicates whatever biases the instructor had. If the instructor favored a certain style, the AI would learn that too. Lucas: Exactly. And the instructor's grades aren't necessarily the gold standard – they're one person's opinion. So the question becomes: can AI ever be objective about art? Or are we just encoding one subjectivity into an automated system that looks authoritative? Luna: I think it's the latter, unless we build in mechanisms for disagreement. Like, why not have the AI flag works where its confidence is low and route those to human review? Lucas: That's exactly what some researchers are proposing. The AI acts as a first pass, but anything outside its comfort zone gets escalated. The problem is, that still requires human labor, and the whole point of using AI is to save time. Luna: Right, but if you're saving time at the cost of fairness, what's the point? I'd rather have accurate grades than fast ones. Lucas: Agreed. And there's another layer here – the AI's training data had a Western-centric bias. Most of the professional works it was trained on were from European and North American artists. So styles common in, say, East Asian ink painting or African textile patterns were essentially invisible to the system. Luna: That could have huge implications for students from those backgrounds. Imagine being told your cultural aesthetic isn't 'correct' by a machine. Lucas: Yeah. The university actually had one student who worked in a style inspired by indigenous Australian dot painting. The AI gave her low scores for 'composition balance' because the dots were arranged in irregular patterns. The human evaluator gave her top marks for originality. Luna: So the AI is essentially penalizing cultural expression. That's not just unfair – it's harmful. Lucas: And it's a pattern we see across many AI applications. When the training data lacks diversity, the system's outputs reflect that. The fix isn't just technical – it's about who decides what's 'good' in the first place. Luna: So what's the path forward? Should universities just stop using AI for grading creative work? Lucas: I don't think we need to ban it, but we need guardrails. For example, some schools are experimenting with AI that provides formative feedback – like 'your color palette is harmonious but you might want to increase contrast here' – without assigning a numeric grade. Luna: That seems smarter. The AI becomes a tool for learning, not a judge. Lucas: Right. And in Maria's case, after the controversy, the department shifted to a hybrid model: the AI gives suggestions, but the final grade is entirely human. The AI's feedback is used for revision, not evaluation. Luna: How did students respond to that change? Lucas: Positively. Maria told me she felt less anxious about experimenting. She even started using the AI's suggestions as a starting point for pushing her style further – like, if the AI said her edges were too soft, she deliberately made them softer to see what would happen. Luna: That's a great outcome. The AI becomes a sparring partner, not a gatekeeper. Lucas: Exactly. And it reinforces a key lesson: AI in subjective domains works best when it augments human judgment, not replaces it. The moment you automate a creative evaluation, you risk narrowing what's considered valuable. Luna: And that's a risk for students, but also for the art world at large. If AI grading becomes widespread, it could shape what kinds of art are produced – just because they score well. Lucas: That's the long-term concern. We might end up with a generation of artists who unconsciously optimize for the algorithm, the way some YouTubers optimize for watch time. The very definition of creativity could shift. Luna: So it's not just a technical problem – it's a cultural one. And it requires educators, artists, and technologists to work together. Lucas: Right. And on that note, I think the key takeaway is that if you're deploying AI in any creative or subjective field, you need to stress-test it against diverse examples, include human oversight, and be transparent about its limitations. Luna: And listen to the people it's evaluating. Maria's story shows that sometimes the best feedback comes from the ones the system was designed to assess. Lucas: Absolutely. Thanks for listening to this episode of AI Ethics with Fexingo. We'll be back next week with another angle on responsible AI.