LLM Alignment Techniques: Stochastic Optimizations in LLM Post-training and Reasoning
This talk explores approaches to improving large language model (LLM) post-training and reasoning through stochastic optimization techniques. The first part introduces ComPO, a preference alignment method using comparison oracles in stochastic optimization. The work addresses likelihood displacement issues in traditional direct preference optimization. The second part proposes the spectral policy optimization, a framework that overcomes GRPO’s limitations with all-negative-sample groups by introducing response diversity with AI feedback. Both approaches demonstrate significant improvements across various model sizes and benchmarks, representing important advances in LLM post-training via stochastic optimization.
Room 928, Cheng Yu Tung Building, CUHK Business School
Prof Xi Chen
Professor,
Department of Technology, Operations, and Statistics,
New York University Stern School of Business,
United States