Tag
PPO
Every PPO story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.
Week 36 · 2026
Read the issue →-
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
Researchers have developed SecOPD, a fine-tuning technique that uses token-level feedback to protect AI agents from prompt injection attacks. This method significantly outperforms previous state-of-the-art defenses by precisely identifying and penalizing malicious tokens during training.
Week 35 · 2026
Read the issue →-
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
Researchers have developed SecOPD, a fine-tuning technique that uses token-level feedback to protect AI agents from prompt injection attacks. This method significantly outperforms previous state-of-the-art defenses by precisely identifying and penalizing malicious tokens during training.
Week 34 · 2026
Read the issue →-
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
This paper presents Agentic ESOpt, a novel approach for fine-tuning LLM agents in long-horizon tasks using evolutionary strategies. The method proves more scalable and memory-efficient than RL-based alternatives like PPO and GRPO as task complexity increases.
Week 25 · 2026
Read the issue →-
APPO: Agentic Procedural Policy Optimization
This paper proposes APPO, a method to enhance the training of autonomous LLM agents through more granular credit assignment. It moves beyond trajectory-level rewards by identifying and branching at specific high-impact decision points within the reasoning process.
Free weekly digest
Get next Saturday’s issue in your inbox
The week’s most relevant AI, security, blockchain, and engineering stories — curated, summarised, and reviewed by humans. No spam, unsubscribe anytime.
Subscribe — it’s free