Benchmark Shows AI Agents Can Write Ruby Yet Struggle to Navigate Codebases

Researchers evaluated five AI models on thirteen Ruby on Rails codebases. The benchmark measured the ability to write Ruby code. All

Researchers evaluated five AI models on thirteen Ruby on Rails codebases. The benchmark measured the ability to write Ruby code. All models succeeded in generating syntactically correct code. However, the agents failed to effectively navigate existing codebases. The results highlight a gap between code generation and code understanding. The study underscores challenges in applying AI to complex projects. Findings may guide future improvements in AI tooling. The report is publicly available for further analysis.