There is a version of AI in debate that everyone should refuse, and it is the one being built fastest. An AI that watches two people argue and tells you who was right.
It is easy to see the appeal. A machine is consistent, it does not get tired, and it has no dog in the fight. Why not let it settle things?
Because being right about a contested question is not a thing a model can do, no matter how good it gets. On the questions worth debating, there is no answer key to check against. An AI that declares a winner is not being objective. It is laundering one set of values as if it were arithmetic, and everybody in the room can feel it.
So the interesting question is not whether AI belongs in debate. It is which job you give it.
A referee and a judge do different work, and the difference is the whole thing.
A judge decides who deserves to win. A referee makes sure the contest is fair enough that the result means something, and then gets out of the way so the people who are entitled to decide can decide.
In a live debate, the people entitled to decide are the audience. That is not a technical limitation waiting to be solved. It is the point. A verdict has authority because of who casts it, and a room full of people who watched the whole thing has a claim to that verdict that no observer, human or machine, can take from them.
AI as referee keeps that intact. AI as judge quietly steals it.
Once you draw that line, the amount of genuinely useful work left over is enormous. None of it involves deciding who is right.
Keep the fight fair. Catch the spam, the coordinated pile-ons, the obvious trolling before it drowns the signal. Check the vote for manipulation, so a result reflects the room rather than a bot farm. Flag the bad-faith moves, the repeated interruptions and the personal attacks, and surface them neutrally so a human moderator acts on evidence instead of a hunch.
Keep it honest. When a debater states a statistic, link it to a source, or note plainly that no source exists. Catch the obvious misquote. This is the subtle one, because it looks like judging and is not. Flagging that a claim is unsupported is refereeing. Declaring the claim false is judging. The referee hands the audience better information and lets them do the deciding.
Keep it flowing. Timekeeping, turn-taking, round management. Reading the flood of audience questions and clustering them, so the ten different phrasings of the same question become one, and the best version actually gets asked instead of lost.
Keep it accountable, without taking the gavel. A referee can hold up signals: this exchange stayed respectful, that one got toxic, this debater kept dodging the question. It can even offer its own read on who landed more blows. But offered is the operative word. That read is one more input the audience can see and weigh or ignore. It is never the result.
The strongest case for an AI judge, and why it still loses
The best argument for going further is that crowds are biased. A room votes on charisma, on which side it walked in agreeing with, on who was funnier. An AI judge, the argument goes, would be more consistent and more fair.
The first half is true. Crowds are biased. Anyone who has watched a debate go to the better performer rather than the better case knows it.
But the fix for a flawed verdict is not to hand the verdict to something with no standing to give it. It is to referee the crowd well: strip the bots, structure the question so the vote is about the motion and not the vibes, give the audience better information before they decide. You can make a human verdict more legitimate. You cannot make a machine verdict legitimate at all, because legitimacy was never about accuracy. It was about who has the right to say.
An AI that judges does not remove bias. It replaces a bias you can see, the room's, with one you cannot, the model's, and then dresses it as neutrality. That is a worse deal, not a better one.
Everything above is one attempt at drawing the line. The referee side is deep and mostly unbuilt, and the exact border between refereeing and judging is going to get argued over for years, which is fitting.
So this is a genuine question, not a rhetorical one. If AI is going to run the arena so the humans can win it, what else should it be doing? Where would you draw the line differently? There are jobs on the referee side nobody has thought of yet, and there are probably things on this list that belong on the other side of the line.
We would rather build this with the people who are going to use it than decide it in private. Tell us what a good referee should do. The verdict, as always, stays with you.