DPO-ST. [ACL 2024] Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning
54DiffAug. [EMNLP 2022] Differentiable Data Augmentation for Contrastive Sentence Representation Learning. https://arxiv.org/abs/2210.16536
40MsAT. [ACL 2023] Learning Multi-step Reasoning by Solving Arithmetic Tasks. https://arxiv.org/abs/2306.01707
24Speech-BT. [EMNLP 2025] From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
13