Year
2026
Posts published in this year. 10 posts.
- Jun 27, 2026Agentic RL 综述:工具调用、信用分配与训练稳定性——从 RAP 到 AEPO
梳理 Agentic RL 从树搜索到可训练策略的演进,涵盖 Planner-R1、TORL/ToolRL/ARTIST、GiGPO/ARPO、RAGEN/RAGEN-2 及 AEPO,聚焦 reward 设计、credit assignment 与 reasoning collapse 三大核心问题。
- Jun 05, 2026OPD 如何重构后训练的不可能三角
后训练希望学习信号同时准确、稠密、易得,但现实里很难三者兼得。OPD 提供了一种新的组织方式——在学生自己的轨迹上,把教师、验证器、环境反馈组织成密集监督。
- Apr 24, 2026Training-Free Prompt Optimization:从经验库到问题重构
关于 training-free prompt optimization、GRPO、经验库与 3DrawAgent 的一些思考
- Mar 24, 2026远程连接服务器时的 AI 编程工具实践与配置指南
探讨在远程连接服务器时使用 Cursor、Copilot、Claude Code 等 AI 工具的优缺点及网络转发配置方案
- Mar 10, 2026Task Vector in Multimodal In-Context Learning 论文阅读笔记
Notes on task vectors, function vectors, in-context vectors, and multimodal in-context learning.
- Feb 12, 2026Leetcode学习笔记
Notes on Leetcode
- Feb 10, 2026Self-Distillation论文阅读
Notes on Papers about Self Distillation
- Feb 03, 2026LLM八股学习与手撕
Notes on LLM algorithms
- Feb 01, 2026Claude Code 配置与 CC Switch 代理接入完全指南
从零配置 Claude Code 走第三方 API,含 CC Switch 本地代理、CLI 与 VS Code 插件统一接入、排查命令和迁移清单
- Jan 31, 2026AAAI2026参会记录
Log of experience in AAAI2026