Semgrep Reports GLM 5.2 Outperforms Claude on Cybersecurity Benchmarks
Semgrep has published results from its Cyber Benchmarks evaluating AI models on code security tasks. The benchmark suite tests how well large language models identify vulnerabilities in source code.
Semgrep has published results from its Cyber Benchmarks evaluating AI models on code security tasks.
The benchmark suite tests how well large language models identify vulnerabilities in source code.
GLM 5.2, a model from Chinese AI company Zhipu AI, scored higher than Anthropic's Claude. The
results suggest rapid progress in open and non-US models on specialized security tasks. Semgrep's
benchmarks are widely referenced in the application security community. The findings highlight
growing competition among AI models in niche technical domains. Security teams may need to reassess
which models they rely on for code review. The full methodology and detailed scores are available in
Semgrep's blog post.