Notabletraining methods

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

Rishabh Agrawal, Hejie Cui, Shasha Li, Shanchan Wu, Sercan Ö. Arık

Published
Sep 29, 2026 — 17:55 UTC

Problem

This work addresses the limitation of existing methods that learn from corrections which do not alter the execution of tasks. The authors propose a novel approach to improve the learning process of large language models (LLMs) by leveraging targeted multi-turn self-distillation. This paper is a preprint and has not undergone peer review.

Method

The core technical contribution is the Advisor Self-Distillation (AdviSD) algorithm, which employs a shared-parameter model architecture. The loss function integrates outcome-based reinforcement learning with self-distillation techniques. The data utilized for training includes Qwen3-8B advisors specifically designed for the Gemini and Claude models. Notably, the training compute details are not disclosed. Additional mechanisms in the method include:

  • Reflection: This component proposes corrections to the executor's responses.
  • Advisor Scoring: The advisor evaluates the executor's responses both with and without the provided advice.
  • Supervision Selection: Decisions for supervision are made based on the magnitude of the difference between the advisor's suggestions and the executor's outputs.

Results

AdviSD demonstrates significant performance improvements over the advisor-GRPO baseline across multiple benchmarks:

  • On the BFCL-v3 benchmark, AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points.
  • In the EnvScaler benchmark, AdviSD shows an improvement of 3.9-5.1 score points over advisor-GRPO.
  • The trained advisors exhibit generalization capabilities, successfully transferring to out-of-domain tasks and across different executor versions and model families.
  • AdviSD also surpasses matched-count random selection in performance.

Limitations

The authors note that corrections which provide less impactful advice may hinder the learning process. This limitation could restrict the effectiveness of the AdviSD method in scenarios where the quality of advice is variable. Additionally, the lack of specified training compute may raise questions regarding the scalability of the approach.

Why it matters

The implications of this work are significant for the development of more effective LLMs. By enhancing the learning process through targeted self-distillation, AdviSD could lead to improved performance in various applications of LLMs, particularly in scenarios requiring nuanced understanding and execution of complex tasks. This method may pave the way for future research into more sophisticated advisor-executor frameworks, potentially leading to more robust and adaptable AI systems.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI