TL;DR: No Pain, More Gain: Iterative Merging for Effective Multi-Teacher On-Policy Distillation
How can one student recover the capabilities of multiple teachers?
Multi-Teacher On-Policy Distillation trains a student on its own generated samples using feedback from specialized teachers. IM-MOPD starts with a uniform model merge, then progressively adds teacher task-vector updates for domains the student has not recovered well. Across five domains, it improves average capability recovery over the tested uniform-merge and supervised warm-up baselines.
TL;DR: Persona-Pruner: Sculpting Lightweight Models for Role-Playing
Can a smaller model preserve a specific persona?
Persona-Pruner uses a persona description to identify a specialized subnetwork within a language model. In the reported experiments, it preserves role-playing quality better than the pruning baselines while retaining general capabilities.
TL;DR: Cross-lingual Transfer of Reward Models in Multilingual Alignment
Can reward models transfer across languages?
The study finds that English-trained reward models can transfer effectively to other languages. It examines changes in model representations and shows how this transfer can support multilingual instruction following.
TL;DR: ORPO: Monolithic Preference Optimization without Reference Model
Can fine-tuning and preference alignment share one training stage?
ORPO adds an odds-ratio preference objective to supervised fine-tuning. It favors preferred responses without a separate reference model or an additional alignment stage, with experiments across models from 125 million to 7 billion parameters.
Across my work on preferences, reward models, and distillation, I’m interested in how training signals shape model behavior—and how specialized capabilities can be combined.