요약
I built this fun benchmark to pitch LLM models against each other in Oxford-style debate.The format is inspired by Intelligence Squared. The side who flips most votes win. Comments URL: https://news.ycombinator.com/item?id=47700187 Points: 1 # Comments: 0
본문
How each model behaves as a judge persuadability, bias, and self-judging patterns