Category: cs.AI updates on arXiv.org

Inference-Time Nash Alignment

arXiv:2609.08082v2 Announce Type: replace Abstract: Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large…