Back to papers
March 19, 2026cs.AIcs.CLIntermediate
Reasoning over mathematical objects: on-policy reward modeling and test time aggregation
Pranjal Aggarwal, Marjan Ghazvininejad, Seungone Kim, Ilia Kulikov, Jack Lanchantin, Xian Li, Tianjian Li, Bo Liu, Graham Neubig, Anaelia Ovalle, Swarnadeep Saha, Sainbayar Sukhbaatar, Sean Welleck, Jason Weston, Chenxi Whitehouse, Adina Williams, Jing Xu, Ping Yu, Weizhe Yuan, Jingyu Zhang, Wenting Zhao
AI-Generated Summary
This paper tackles how AI language models can better solve complex math and science problems that require deriving formally structured mathematical expressions (like equations or proofs) rather than just picking a number or multiple choice answer. The researchers created a new training dataset called Principia, developed improved training methods using AI judges that learn during training, and showed how to use extra computation at test time to aggregate multiple solution attempts for better results.
HF Upvotes
3
Difficulty
Intermediate
Categories
cs.AI, cs.CL
AI Tags
mathematical reasoninglanguage modelsreward modelingtest-time scalingSTEM applicationsLLM evaluationon-policy training