Candidate supply and answer selection shape the value of LLM judging in multi-agent systems
Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She, Jin-Tai Yu, and Zhiyuan Yuan
We isolate candidate generation, answer recognition, and terminal selection in multi-agent reasoning. Correct answers are often generated but lost during consensus; a judge-guided selection rule can rescue correct minority answers without changing the candidate pool.
15,336
questions used to map judge reliability
81,390
fixed candidate-pool replays
63.82% → 70.82–70.95%
accuracy after changing only answer selection