Topic
agent
Posts around this finer topic. 2 posts.
- Jun 27, 2026Agentic RL 综述:工具调用、信用分配与训练稳定性——从 RAP 到 AEPO
梳理 Agentic RL 从树搜索到可训练策略的演进,涵盖 Planner-R1、TORL/ToolRL/ARTIST、GiGPO/ARPO、RAGEN/RAGEN-2 及 AEPO,聚焦 reward 设计、credit assignment 与 reasoning collapse 三大核心问题。
- Apr 24, 2026Training-Free Prompt Optimization:从经验库到问题重构
关于 training-free prompt optimization、GRPO、经验库与 3DrawAgent 的一些思考