← All topics

Tag

DPO

Every DPO story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.

4 items · 4 issues

Beats AI & ML

Week 38 · 2026

Read the issue →
  • A Zeroth-Order Paradigm for LLM Preference Alignment

    The researchers present ComPO, a new alignment paradigm that uses comparison oracles to extract directional information from preference pairs. This method effectively mitigates the risks of likelihood displacement and model verbosity seen in traditional direct alignment methods.

    AI & ML arXiv Source ↗

Week 36 · 2026

Read the issue →

Week 35 · 2026

Read the issue →

Week 25 · 2026

Read the issue →
  • APPO: Agentic Procedural Policy Optimization

    This paper proposes APPO, a method to enhance the training of autonomous LLM agents through more granular credit assignment. It moves beyond trajectory-level rewards by identifying and branching at specific high-impact decision points within the reasoning process.

    AI & ML HuggingFace Papers Source ↗