Total: 1
Despite large language models (LLMs) achieving impressive performance on benchmark tasks such as medical question answering, their real-world utility remains limited. We argue that while benchmarks play a valuable role in developing methods and filtering promising models during development, they fundamentally cannot establish deployment readiness. Many models topping benchmark performance have failed to perform seemingly related clinical tasks effectively in practice, while others with modest benchmark performance have demonstrated meaningful clinical benefits. We detail the limitations of benchmark-centric evaluations of deployment readiness. We argue that we should use benchmarks only to identify candidate methods or models, not to justify deployment. We call for increased use of other measurement instruments of deployment readiness, such as prospective studies, and policy changes that align incentives with clinically grounded evaluation.