Efficient Estimation of Kernel Surrogate Models for Task Attribution
Researchers from arXiv have introduced a novel method for task attribution in modern AI agents, such as large language models, which are trained on diverse tasks simultaneously. The study addresses the computational infeasibility of leave-one-out retraining by proposing kernel surrogate models that capture nonlinear interactions between training tasks, unlike prior linear models. The authors establish a connection between linear surrogate models and influence functions through second-order analysis. They developed a gradient-based estimation procedure leveraging first-order approximations of pretrained models, achieving accurate surrogate estimates with less than 2% relative error without repeated retraining. Experiments in mathematical reasoning, in-context learning, and multi-objective reinforcement learning demonstrate that kernel surrogate models achieve a 25% higher correlation with ground truth compared to linear baselines. Furthermore, when applied to downstream data selection, these models yield a 40% performance improvement. This approach enables more accurate and scalable task attribution, offering significant advancements in understanding how individual training tasks influence target performance in complex AI systems.
Wire timeline
Efficient Estimation of Kernel Surrogate Models for Task Attribution
Researchers from arXiv have introduced a novel method for task attribution in modern AI agents, such as large language models, which are trained on diverse tasks simultaneously. The study addresses the computational infeasibility of leave-one-out retraining by proposing kernel surrogate models that capture nonlinear interactions between training tasks, unlike prior linear models. The authors establish a connection between linear surrogate models and influence functions through second-order analysis. They developed a gradient-based estimation procedure leveraging first-order approximations of pretrained models, achieving accurate surrogate estimates with less than 2% relative error without repeated retraining. Experiments in mathematical reasoning, in-context learning, and multi-objective reinforcement learning demonstrate that kernel surrogate models achieve a 25% higher correlation with ground truth compared to linear baselines. Furthermore, when applied to downstream data selection, these models yield a 40% performance improvement. This approach enables more accurate and scalable task attribution, offering significant advancements in understanding how individual training tasks influence target performance in complex AI systems.
cs.AI updates on arXiv.org