Back to Research papers
Research paper index

Unsupervised Post-Training of Foundation Models: A Survey

Yijie Xu, Qianyi Cai, Huizai Yao, Yili Wang, Tianfu Wang, Cehao Yang, Xingbo Yao, Zhiyu Guo, Aiwei Liu, Xuming Hu, Weiyu Guo, Hui Xiong

arXiv:2608.24982Published August 25, 2026Updated August 27, 20260 citations
  • cs.CL
  • cs.AI
  • cs.CV
  • cs.LG
  • cs.MM
  • foundation model

Abstract

Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rather than an external oracle. We catalog 80 strict UPT methods and organize them by the object that supplies the update signal: a prediction statistic, a sample relation, a self-generated target, or an internal evaluator. Beyond inventory, we show how the choice of internal signal and task structure determines whether post-training improves the model or recursively amplifies error. An orthogonal Input Visibility $\times$ Update Persistence view maps deployment regimes and defines a unified framework for UPT selection and evaluation.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.