상세 보기
Performance Evaluation of OpenAI’s o4-mini on CSAT Mathematics: Multimodal Reasoning and the Reasoning–Verification–Reanalysis Loop
- Oh, Sejun;
- Park, Jeonghwan;
- Baek, Hangyeol
WEB OF SCIENCE
0SCOPUS
0초록
Recent advances in multimodal large language models (LLMs) have expanded their potential for solving complex mathematical problems; however, their reasoning capabilities in real, curriculum-aligned examinations remain underexplored. This study evaluates the mathematical reasoning performance of OpenAI's o4-mini model on the mathematics section of the Korean College Scholastic Ability Test (KCSAT), using official 2023-2025 exams and mock tests as benchmark data. The o4-mini is a multimodal LLM capable of processing both textual and diagrammatic information, enabling assessment on problems that require integrated visual reasoning. Quantitative analyses reveal that o4-mini demonstrates strong performance on standard problems and moderate success on high-difficulty items, while qualitative analyses uncover a distinctive "Reasoning-Verification-Reanalysis (RVR) loop" in its reasoning process. This iterative pattern-where the model reasons, self-verifies intermediate steps, and reanalyzes discrepancies-reflects metacognitive-like self-monitoring that can improve problem-solving accuracy. These findings provide new insights into the cognitive dynamics of multimodal LLMs, highlighting both their potential and limitations in high-stakes mathematical assessments, and suggesting pathways for enhancing AI reliability through structured reasoning and self-verification mechanisms.
키워드
- 제목
- Performance Evaluation of OpenAI’s o4-mini on CSAT Mathematics: Multimodal Reasoning and the Reasoning–Verification–Reanalysis Loop
- 저자
- Oh, Sejun; Park, Jeonghwan; Baek, Hangyeol
- 발행일
- 2026-02
- 유형
- Article
- 저널명
- IEEE Access
- 권
- 14
- 페이지
- 16376 ~ 16388