OpenAI Questions the Reliability of SWE-Bench Pro Benchmark
OpenAI has published a critical analysis pointing out serious deficiencies in SWE-Bench Pro, a popular benchmark for evaluating coding skills of language models. The study raises doubts about the validity of comparisons between models and the true usefulness of these metrics.


