It's Not the Number of Models: Reasoning Diversity Improves LLM Forecast Accuracy

16 September 202618 views

Researchers proposed assembling a "crowd" of language models based not on their number but on the nature of their reasoning: clustering models by behavior and selecting typical representatives made it possible to get by with three participants where simple voting required twenty-five. This approach not only improved prediction quality on two benchmarks but also sharply reduced inference costs.

It's Not the Number of Models: Reasoning Diversity Improves LLM Forecast Accuracy

Crowd wisdom hits a wall of sameness

When you need to predict something about the future — market behavior, the outcome of an event, the fate of a technology — a simple thought comes to mind: the more models you ask, the more reliable the answer will be. The logic seems ironclad: different systems make mistakes in different ways, so averaging will smooth out random errors.

But there's a catch in this scheme. If all the members of the "crowd" reason in a similar way, their votes aren't independent evidence — they're the same opinion, multiplied by copying. Ten nearly identical answers add no information compared to one; they merely create a sense of confidence and increase the inference bill.

This is precisely the problem that the paper "Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction" is built around — it was prepared by Nirupam Chetlapalli, Yiming Liao, Min-Chun Chen, and Keke Chen. The arXiv preprint 2608.24001 appeared on August 25, 2026, a revised version came out the very next day, and the paper itself was submitted to IEEE BigData 2026.

Simply "more models" doesn't solve the problem

The authors' key observation: different large language models can behave redundantly alike. They're trained on comparable data, gravitate toward similar phrasing and typical reasoning moves. As a result, mechanically expanding the list of participants doesn't yield the diversity that was the whole point.

Here it's important to distinguish two concepts. Quantity is how many models we query. Diversity is how differently they arrive at their conclusions. The second doesn't follow automatically from the first: you can assemble twenty participants and get twenty variations of the same thought.

The method: models are selected by reasoning style, not by answers

What behavioral selection looks like

The authors propose a framework they call behavior-oriented. The point isn't to compare models by the correctness of their final answers, but to understand how they think.

The scheme is as follows:

  1. Each model is run through a set of independent tasks unrelated to future prediction. What's collected isn't so much the answers as the chains of reasoning — the reasoning traces themselves.
  2. Based on these traces, a behavioral profile is built: what steps the model takes, how it breaks down the task, where it makes mistakes, what arguments it resorts to.
  3. Models are grouped into clusters by behavioral similarity. The K-means++ algorithm is used for this.
  4. One representative is taken from each cluster — a medoid, that is, the model most typical of its group.
  5. It's this compact "crowd" of medoids that delivers the collective forecast.

The idea, if you think about it, isn't new — this is what's done in statistics when, out of correlated features, one representative is kept so as not to duplicate the same information. What's new here is that the "feature" becomes the language model's manner of reasoning rather than a numerical characteristic.

What it was tested on

The experiment is built on 25 models. Seven development benchmarks were used to describe and separate the models by behavior, and two separate benchmarks — this time for future prediction — to evaluate the quality of the resulting "crowds." This separation is fundamental: the tasks used to build the profile don't coincide with the tasks used to measure the result. Otherwise it would be training with peeking, and the numbers would mean almost nothing.

The result: three models beat twenty-five

The loudest thing in the paper is the outcome of the comparison. A "crowd" of just three medoid models, assembled through behavioral clustering, performed better than ordinary voting across all 25 models at once. And this holds for both prediction benchmarks.

The savings turned out to be impressive: the number of model calls dropped by 88%, and inference costs by roughly 80%. For practice, this is perhaps even more important than the accuracy gain itself. Prediction systems often hit a wall not on answer quality but on its price: if every question requires querying a couple dozen models, the solution becomes unprofitable long before it becomes inaccurate.

This leads to a non-obvious conclusion: the composition of the "crowd" weighs more than its size. Twenty-five models aren't an automatic advantage over three if those twenty-five essentially repeat each other.

Not all diversity is equally useful

The authors draw one more careful conclusion that's easy to miss. It turns out that what matters isn't maximizing diversity as such, but representativeness — how well the selected participants reflect the different poles of behavior found among the models.

The difference is fundamental. You can gather three models that differ from each other randomly and over trifles — and get noise instead of signal. Or you can select one from each of the genuinely different behavioral groups — and get a compact set that covers the solution space. The first is motleyness, the second is heterogeneity. The second is what works.

A parallel with ordinary expertise suggests itself here. A good committee isn't the one with more people, but the one where different approaches and points of view are represented. Five specialists from different fields are usually more useful than fifteen graduates of the same department.

What this means in practice

For those building prediction services on language models, the paper's conclusions offer a fairly concrete recipe:

  • Don't chase the length of the model list. Expanding the pool without analyzing behavior means paying for duplicate answers.
  • Measure behavior first, then choose the composition. Running models on unrelated tasks and clustering gives a clear picture of who in the set is genuinely different.
  • Save deliberately. Cutting the number of calls by an order of magnitude while maintaining or improving accuracy changes the economics of a product, not just its technical characteristics.
  • Remember the limitations. All of this was tested on a specific set of 25 models and two prediction benchmarks; the numbers shouldn't be transferred one-to-one to other domains and other models.

The paper's main idea sounds almost human: the value of collective opinion rests not on the number of votes, but on the fact that those votes are genuinely different. Applied to language models, this means that what needs to be measured isn't a list of names but a style of thinking — and the team should be assembled based on that.

Frequently asked questions

It's Not the Number of Models: Reasoning Diversity Improves LLM Forecast Accuracy