DeepTutor: An Open-Source Agentic Framework for Personalized AI Tutoring
Researchers have introduced DeepTutor, a fully open-source agentic framework designed to enhance personalized education through Large Language Models (LLMs). Addressing the limitations of static pre-training knowledge and existing Retrieval-Augmented Generation (RAG) systems, DeepTutor unifies citation-grounded problem tutoring with difficulty-calibrated question generation. The system features a hybrid personalization engine that combines static knowledge grounding with dynamic learner memory, allowing it to adapt continuously to individual student needs. This architecture supports adaptive learning workflows, interactive books, and proactive multi-channel tutoring agents. To evaluate its efficacy, the team developed TutorBench, an interactive benchmark based on university-level curricula across five domains, alongside an LLM-based first-person interactive evaluation protocol using profile-driven student simulators. Comprehensive evaluations, including human-alignment and ablation studies, demonstrate the framework's robustness. Results indicate that DeepTutor improves personalized metrics by an average of 10.8% and strengthens general agentic reasoning across five backbone models by 29.4%, marking a significant advancement in AI-driven educational tools.
Wire timeline
DeepTutor: An Open-Source Agentic Framework for Personalized AI Tutoring
Researchers have introduced DeepTutor, a fully open-source agentic framework designed to enhance personalized education through Large Language Models (LLMs). Addressing the limitations of static pre-training knowledge and existing Retrieval-Augmented Generation (RAG) systems, DeepTutor unifies citation-grounded problem tutoring with difficulty-calibrated question generation. The system features a hybrid personalization engine that combines static knowledge grounding with dynamic learner memory, allowing it to adapt continuously to individual student needs. This architecture supports adaptive learning workflows, interactive books, and proactive multi-channel tutoring agents. To evaluate its efficacy, the team developed TutorBench, an interactive benchmark based on university-level curricula across five domains, alongside an LLM-based first-person interactive evaluation protocol using profile-driven student simulators. Comprehensive evaluations, including human-alignment and ablation studies, demonstrate the framework's robustness. Results indicate that DeepTutor improves personalized metrics by an average of 10.8% and strengthens general agentic reasoning across five backbone models by 29.4%, marking a significant advancement in AI-driven educational tools.
cs.AI updates on arXiv.org