Back to Research papers
Research paper index

Direct Preference Optimization for Chatbot Fine-Tuning: An Empirical Study

Dezhi Yu, Yvonne Qiu, ShuoJia Fu

arXiv:2606.12881Published June 11, 2026Updated June 12, 20260 citations
  • cs.CL
  • cs.LG
  • reinforcement learning

Abstract

We present an approach to fine-tuning large language models using Direct Preference Optimization (DPO), a reinforcement learning technique. Our experimental results demonstrate that DPO simplifies the training pipeline, improves computational efficiency, and achieves competitive performance. The evaluation using BLEU, ROUGE, and cosine similarity metrics indicates effective learning and convergence, though further investigation is needed to address observed training instability.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.