85.3 GFLOPS Achieved Optimizing FP32 Matrix Multiplication on Single AMD Zen 3 Core

A repository details an optimization of FP32 matrix multiplication. The code targets a single core of AMD’s Zen 3 architecture. Benchmarks

A repository details an optimization of FP32 matrix multiplication. The code targets a single core of AMD’s Zen 3 architecture. Benchmarks show a peak performance of 85.3 GFLOPS. The implementation focuses on low‑level instruction tuning. Results demonstrate the potential of single‑core compute on modern CPUs. The project includes source code and build instructions. It serves as a reference for developers seeking high‑performance math kernels. The author invites community testing and further refinement.