OpenAI withdraws recommendation for SWE‑Bench Pro coding benchmark

OpenAI announced it will stop recommending SWE‑Bench Pro for coding assessments. The decision was detailed in a blog post titled “Separating Signal from Noise in Coding

OpenAI announced it will stop recommending SWE‑Bench Pro for coding assessments. The decision was detailed in a blog post titled “Separating Signal from Noise in Coding Evaluations.” OpenAI said the benchmark no longer aligns with its current evaluation goals. The company highlighted concerns about distinguishing meaningful results from noise. It indicated a shift toward alternative metrics for measuring model performance. Researchers and developers using SWE‑Bench Pro are advised to review the guidance. OpenAI will provide recommendations for other suitable benchmarking tools. The change may influence how the community validates AI coding capabilities.