2025-12 - Detected unexpectedly low reasoning performance of GPT-5.2 at low, medium and high reasoning effort (xhigh works fine): The purpose of this project is to test LLM reasoning abilities with ...