M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models
Researchers have introduced M2A, a novel paradigm designed to synergize mathematical and agentic reasoning capabilities within large language models (LLMs). Addressing the misalignment between intrinsic logical reasoning for closed-world problems and multi-turn interactive reasoning for external environments, M2A utilizes model merging in parameter space. This method identifies feature subspaces critical for agent behavior and merges mathematical reasoning task vectors along their null space, thereby injecting reasoning capabilities without perturbing existing agent behaviors or requiring additional gradient updates like Supervised Fine-Tuning (SFT) or Reinforcement Learning (RL). The approach offers a simple merging coefficient to control reasoning length. Experimental results in real-world coding agent settings demonstrate that M2A effectively extends agentic reasoning depth and significantly improves performance. Specifically, when applied to a fine-tuned Qwen3-8B model, M2A increased the SWE-Bench Verified resolved rate from 44.0% to 51.2% without retraining. This advancement highlights a efficient pathway for enhancing LLM reasoning stability and performance through parameter-space manipulation rather than traditional training methods.
Wire timeline
M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models
Researchers have introduced M2A, a novel paradigm designed to synergize mathematical and agentic reasoning capabilities within large language models (LLMs). Addressing the misalignment between intrinsic logical reasoning for closed-world problems and multi-turn interactive reasoning for external environments, M2A utilizes model merging in parameter space. This method identifies feature subspaces critical for agent behavior and merges mathematical reasoning task vectors along their null space, thereby injecting reasoning capabilities without perturbing existing agent behaviors or requiring additional gradient updates like Supervised Fine-Tuning (SFT) or Reinforcement Learning (RL). The approach offers a simple merging coefficient to control reasoning length. Experimental results in real-world coding agent settings demonstrate that M2A effectively extends agentic reasoning depth and significantly improves performance. Specifically, when applied to a fine-tuned Qwen3-8B model, M2A increased the SWE-Bench Verified resolved rate from 44.0% to 51.2% without retraining. This advancement highlights a efficient pathway for enhancing LLM reasoning stability and performance through parameter-space manipulation rather than traditional training methods.
cs.AI updates on arXiv.org