Category: AI

Inference-Time Nash Alignment

arXiv:2609.08082v2 Announce Type: replace Abstract: Preference-based fine-tuning methods such as RLHF and DPO require substantial compute and large…