Testing 67 Models: Combining LLMs Rarely Beats the Best Single Model
A study tested 67 large language models and found that combining multiple models (ensembling) rarely outperforms the single best model. The results suggest that instead of complex orchestration, simply allocating tasks to the best individual model yields better performance. This challenges the common assumption that model collaboration always leads to superior outcomes.