Large Language Models for Sequential Decision-Making: Improving In-Context Learning via Supervised Fine-Tuning
A new research paper submitted to arXiv explores the application of Large Language Models (LLMs) in sequential decision-making contexts, specifically addressing Markov Decision Processes (MDPs), Partially Observable MDPs (POMDPs), and Ambiguous POMDPs. The authors propose a framework that utilizes supervised fine-tuning (SFT) on offline, oracle-labeled trajectories to enhance the in-context learning capabilities of pretrained LLMs. Theoretically, the study interprets fine-tuned attention layers as implicit estimators of optimal Q-functions, deriving suboptimality bounds that distinguish between estimation error and training bias. Empirical results demonstrate that this SFT approach significantly reduces optimality gaps compared to baseline methods, particularly in complex environments with longer horizons or partial observability. This methodology offers a promising avenue for leveraging abundant offline data in critical sectors like healthcare, enabling more effective policy imitation and decision-making without requiring extensive online interaction. The findings highlight the potential of adapting general-purpose LLMs for specialized, high-stakes sequential tasks through targeted fine-tuning techniques.
Wire timeline
Large Language Models for Sequential Decision-Making: Improving In-Context Learning via Supervised Fine-Tuning
A new research paper submitted to arXiv explores the application of Large Language Models (LLMs) in sequential decision-making contexts, specifically addressing Markov Decision Processes (MDPs), Partially Observable MDPs (POMDPs), and Ambiguous POMDPs. The authors propose a framework that utilizes supervised fine-tuning (SFT) on offline, oracle-labeled trajectories to enhance the in-context learning capabilities of pretrained LLMs. Theoretically, the study interprets fine-tuned attention layers as implicit estimators of optimal Q-functions, deriving suboptimality bounds that distinguish between estimation error and training bias. Empirical results demonstrate that this SFT approach significantly reduces optimality gaps compared to baseline methods, particularly in complex environments with longer horizons or partial observability. This methodology offers a promising avenue for leveraging abundant offline data in critical sectors like healthcare, enabling more effective policy imitation and decision-making without requiring extensive online interaction. The findings highlight the potential of adapting general-purpose LLMs for specialized, high-stakes sequential tasks through targeted fine-tuning techniques.
cs.AI updates on arXiv.org