Company
The origins of LLM Council
LLM Council draws on two separate strands: earlier human and spatial-network research as design inspiration, and recent experiments that directly test LLM debate, panels, and aggregation. The first does not validate the second, and the AI evidence remains task- and setup-dependent.
2013
At Arizona State University’s W. P. Carey School of Business, Ash Tiwari and Dr. Timothy J. Richards compared how peer recommendations and anonymous Yelp reviews affected human restaurant demand. The work began as a 2013 conference paper and was published in Agribusiness in 2016.
The peer-reviewed article found peer networks substantially more effective in driving restaurant preferences, even after controlling for the endogeneity of peer ratings, and found negative reviews more influential than positive reviews. It did not evaluate group deliberation, decision accuracy, or language models.
- Domain: human restaurant choice, ratings, and revisitation intent
- Finding: peer recommendations had more influence than anonymous reviews
- Finding: negative reviews had more influence than positive reviews
- Limit: influence on preference is not evidence of better decisions or AI performance
2015
Research with Dr. Yueming (Lucy) Qiu and Dr. Yi David Wang, published in the peer-reviewed journal Energy Efficiency, used a spatial fraction logit model and an instrumental-variable approach to study commercial green-building certification in four US states.
The paper found strong spatial correlation in certification diffusion and identified a role for split incentives between building owners and renters. It did not show that buildings deliberate, nor did it establish a universal principle for language models.
Its role in the LLM Council story is an analogy about network effects and diffusion — not direct evidence that model-to-model deliberation improves an answer.
Human and spatial-network sources
Conference paper: Tiwari, A. & Richards, T. J. (2013). “Anonymous Social Networks versus Peer Networks in Restaurant Choice.” AAEA Annual Meeting.
Conference presentation: Qiu, Y., Tiwari, A. & Wang, Y. D. (2013). “Voluntary Green Building Certification: Economic Decision or Following the Trend? A Spatial Approach.” 32nd USAEE/IAEE North American Conference.
Peer-reviewed journal article: Qiu, Y., Tiwari, A. & Wang, Y. D. (2015). “The Diffusion of Voluntary Green Building Certification: A Spatial Approach.” Energy Efficiency, 8(3), 449-471.
Peer-reviewed journal article: Tiwari, A. & Richards, T. J. (2016). “Social Networks and Restaurant Ratings.” Agribusiness, 32(2), 153-174.
2019-2021
At Queensland University of Technology, the research focus shifted to language models: how to take fragmented, noisy information from multiple sources and synthesize it into coherent, meaningful insights.
The language model research explored ELMo, BERT, GPT-2, and GPT-3; used attention mechanisms for signal weighting; coverage mechanisms for completeness; and multi-model quality assessment with methods such as BERTScore, BLEURT, and OpenMEVA.
That work concerned language-model summarisation and evaluation, not multi-agent debate. Attention, coverage, and synthesis later informed the product by analogy; they are not direct equivalents of council scoring, consensus, or dissent.
2023-2025: direct LLM evidence
Direct, task-specific evidence arrived from the broader AI field. Peer-reviewed papers reported gains for multiagent debate on selected reasoning and factuality benchmarks (ICML 2024), ChatEval on two text-evaluation benchmarks (ICLR 2024), forecast aggregation across 31 binary questions (Science Advances 2024), and a 20-model council for emotional-intelligence evaluation (NAACL 2025).
“Replacing Judges with Juries” is an arXiv preprint; across the six datasets it studied, its panel of smaller evaluators outperformed a single large judge, showed less intra-model bias, and cost less. These studies support testing multi-model workflows, not a blanket guarantee: outcomes depend on the task, model pool, prompts, aggregation method, latency, and cost.
LLM Council applies this workflow in a product. Its performance should be judged on product-specific evaluations and user outcomes, not inferred from human peer effects or assumed for every council run.
Product lineage
LLM Council was developed as a product in 2025. Its lineage has two distinct parts: design inspiration from the founder’s earlier human-network, spatial, and summarisation work; and direct technical evidence from the broader multi-agent LLM research community. The earlier studies did not invent or validate LLM councils.
Collaborators
Research collaborators include Dr. Timothy J. Richards, Dr. Yueming (Lucy) Qiu, Dr. Alan Woodley, and Dr. Richi Nayak.
Further reading
Peer-reviewed journal article — Social Networks and Restaurant Ratings: https://onlinelibrary.wiley.com/doi/abs/10.1002/agr.21449
Peer-reviewed journal article — The Diffusion of Voluntary Green Building Certification: https://link.springer.com/article/10.1007/s12053-014-9303-5
Peer-reviewed conference paper (ICML 2024) — Improving Factuality and Reasoning through Multiagent Debate: https://proceedings.mlr.press/v235/du24e.html
Peer-reviewed conference paper (ICLR 2024) — ChatEval: Multi-Agent Debate for LLM Evaluation: https://openreview.net/forum?id=FQepisCUWu
Peer-reviewed journal article (Science Advances 2024) — Wisdom of the Silicon Crowd: https://www.science.org/doi/10.1126/sciadv.adp1528
ArXiv preprint — Replacing Judges with Juries: https://arxiv.org/abs/2404.18796
Peer-reviewed conference paper (NAACL 2025) — Language Model Council: https://aclanthology.org/2025.naacl-long.617/