Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction
Overview
Recovering multimodal capabilities under extreme visual-token reduction through on-policy self-distillation and a progressive token-budget curriculum.

arXiv 2026 · Preprint.

Authors
Undergraduate Student
I am Ruixuan Yang, an undergraduate in the School of Mathematics and Statistics at Xi’an Jiaotong University (XJTU). I’m interested in Model Efficiency, Multimodal LLMs (MLLMs), Post-training & Reasoning, etc.
I am currently a visiting student in Prof. Yulun Zhang’s group at Shanghai Jiao Tong University and Prof. Huan Wang’s ENCODE Lab at Westlake University.