BSO: Safety Alignment Is Density Ratio Matching
Tien-Phat Nguyen, Truong Nguyen, Thin Nguyen, Duy Minh Ho Nguyen, Ngoc-Thanh Dinh, Trung Le

Preference optimization, safety alignment, and token-based methods for aligning foundation models with human intent.
To build principled alignment methods that steer powerful models toward human preferences and safety constraints without sacrificing capability.
This direction develops theory and algorithms for aligning large language and multimodal models with human values. I aim to develop helpful, secure, and smart foundation models aligned with multiple human preferences, values, and cultures.
Paper TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching accepted at ICML 2026
Paper f-Divergence Self-Play for Tabular Anomaly Detection via Large Language Models accepted at ICML 2026
Paper Sharpness-Aware Minimization in Logit Space Efficiently Enhances Direct Preference Optimization accepted at ICLR 2026
Paper CTPD: Cross Tokenizer Preference Distillation accepted at AAAI 2025
Paper Token-Level Self-Play with Importance-Aware Guidance for Large Language Models accepted at NeurIPS 2025
Tien-Phat Nguyen, Truong Nguyen, Thin Nguyen, Duy Minh Ho Nguyen, Ngoc-Thanh Dinh, Trung Le
Vuong Hoang Tran, Van Linh Ngo, Dang Nguyen, Thin Nguyen, Phuoc Nguyen, Mehrtash Harandi, Trung Le
Truong Nguyen, Tien-Phat Nguyen, Linh Ngo Van, Duy Minh Ho Nguyen, Khoa D. Doan, Trung Le
Truong Nguyen, Van-Phi Dat, Ngan Nguyen, Van-Linh Ngo, Trung Le, Hong-Thanh Nguyen
Tue Le, Hoang Tran Vuong, Quyen Tran, Linh Ngo Van, Mehrtash Harandi, Trung Le