OpenAI withdraws recommendation for SWE‑Bench Pro coding benchmark
OpenAI announced it will stop recommending SWE‑Bench Pro for coding assessments. The decision was detailed in a blog post titled “Separating Signal from Noise in Coding
OpenAI announced it will stop recommending SWE‑Bench Pro for coding assessments. The
decision was detailed in a blog post titled “Separating Signal from Noise in Coding
Evaluations.” OpenAI said the benchmark no longer aligns with its current evaluation
goals. The company highlighted concerns about distinguishing meaningful results from
noise. It indicated a shift toward alternative metrics for measuring model performance.
Researchers and developers using SWE‑Bench Pro are advised to review the guidance. OpenAI
will provide recommendations for other suitable benchmarking tools. The change may
influence how the community validates AI coding capabilities.