Topic
llm
Posts around this finer topic. 4 posts.
- Jun 27, 2026Agentic RL 综述:工具调用、信用分配与训练稳定性——从 RAP 到 AEPO
梳理 Agentic RL 从树搜索到可训练策略的演进,涵盖 Planner-R1、TORL/ToolRL/ARTIST、GiGPO/ARPO、RAGEN/RAGEN-2 及 AEPO,聚焦 reward 设计、credit assignment 与 reasoning collapse 三大核心问题。
- Jun 05, 2026OPD 如何重构后训练的不可能三角
后训练希望学习信号同时准确、稠密、易得,但现实里很难三者兼得。OPD 提供了一种新的组织方式——在学生自己的轨迹上,把教师、验证器、环境反馈组织成密集监督。
- Apr 24, 2026Training-Free Prompt Optimization:从经验库到问题重构
关于 training-free prompt optimization、GRPO、经验库与 3DrawAgent 的一些思考
- Feb 03, 2026LLM八股学习与手撕
Notes on LLM algorithms