Copyright (c) 2026 Tecnociencia

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
This study presents a textual benchmark to compare nine large-scale commercial language models for solving Classical Mechanics problems. A total of 42 self-contained items, without explicit dependence on diagrams, were extracted from MIT OpenCourseWare open-source materials from 2005 and 2016. The evaluated models came from three vendors: OpenAI, Anthropic, and Google Gemini. Each problem was presented using the same structured prompt, including a final answer, a summary of the reasoning, equations, units, and confidence level. The complete run yielded 378 usable responses. The correction reference was built in two stages: numerical consensus for 22 items and manual curation for 19 cases lacking consensus; one item remained ambiguous due to residual dependence on missing figures. The pipeline achieved 100% final success, although 29 truncated outputs required parser-level recovery. In heuristic agreement with the curated reference, Gemini achieved the best aggregate performance by vendor (43.7%; 55/126), while gemini-3.1-pro-preview achieved the best result at the model level (61.9%; 26/42). Anthropic produced the longest and slowest responses, and gemini-3-flash-preview was the fastest model. The findings show that LLMs already support reproducible assessments in physics without replacing experts.