Testing 67 Models: Combining LLMs Rarely Beats the Best Single Model
A study testing 67 large language models found that combining multiple LLMs rarely outperforms the best single model. The research challenges the assumption that model orchestration or ensembling reliably improves performance over using a single top-performing model.