跳到正文
Jones Ray

ScholarPulse 日报 2026-08-21

2026-08-21 学术简报:2 篇。AI4AI-Bench基准测试表明LLM代理在递归自我改进中仅能微弱提升训练算法性能,最佳系统得分0.250,仅接近原算法与最优距离的25%。

今日速览

序号标题来源日期主题推荐等级
1AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-ImprovementarXiv2026-08-20RAG高
2Break It Down, Pass It On: Cross-Task Skill Transfer in LLM AgentsarXiv2026-08-20RAG高

重点论文与技术动态

1. AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

一句话结论

AI4AI-Bench基准测试表明LLM代理在递归自我改进中仅能微弱提升训练算法性能,最佳系统得分0.250,仅接近原算法与最优距离的25%。

该研究构建AI4AI-Bench评估LLM代理设计训练算法的能力,结果显示平均得分0.166,最佳系统0.250,多数代理未改变模型学习方式,少数改变的平均得分0.226显著高于其余0.126,但整体改进空间巨大。

核心内容

方法与数据

价值判断

摘要 Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement. That process is the training algorithm: a better objective or update rule improves the compute\mbox{-}capability exchange rate for every subsequent run, including the one that produces the next agent. Whether RSI is feasible therefore turns on whether an agent can design training algorithms. No benchmark isolates that ability: existing suites are won by collecting data or by tuning hyperparameters, and none tells a change to how a run is executed apart from a change to how the model learns. We present AI4AI\mbox{-}Bench, 10 frozen research repositories spanning 10 training algorithm families. In each task, an agent has 4 hours on one B300 to rewrite the training algorithm; its code is then rerun from scratch for up to 12 hours and scored by a fixed evaluator hidden from the agent, against the repository's original algorithm under the same procedure. Because the 10 metrics are incommensurable, every task is mapped onto one scale on which $0$ is an uninformative model, $0.1$ is the algorithm the repository ships, and $1.0$ is the task optimum. Across 29 configurations of 6 systems on all 10 tasks the mean score is $0.166$, and the best system reaches $0.250$: even the strongest closes under a fifth of the distance between the algorithm that was already there and the optimum. The submissions show where that distance went: most never change how the model learns at all, and the minority that do average $0.226$ against $0.126$ for the rest. More reasoning effort mostly buys the willingness to go there, taking that minority from $8\%$ of submissions to $64\%$ and the mean score from $0.094$ to $0.196$. We release the task suite, the evaluators and every scored submission, so that the measurement can be repeated as these systems change.

2. Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents

一句话结论

子任务级技能和文本格式技能能显著提升LLM代理在跨任务技能转移中的可靠性。

一段话:本研究通过系统比较任务级与子任务级技能诱导及文本与代码技能格式,发现子任务级技能平均提升代理性能(高于无记忆基线),任务级技能则降低性能;文本技能转移效果优于代码技能。进而提出技能效用分数,结合具体性与抽象性,能一致预测技能转移成功率,且计算仅需技能和任务描述,无需任务执行。

核心内容

方法与数据

价值判断

摘要 Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them. When agent-induced skills transfer reliably across tasks remains an open question. We conduct a comprehensive and controlled study of how the way skills are induced shapes their transfer across tasks. Specifically, we compare task-level with subtask-level skill induction and text with code skill formats, the two axes along which existing methods differ. Task-level skills mostly reduce the agent's performance below its no-memory baseline while subtask-level skills raise it above on average, and text skills transfer better than code skills. To further understand our findings, we examine two complementary properties of the induced skills: specificity, which measures how closely a skill matches real tasks, and abstractness, which measures how evenly its relevance spreads across tasks. Neither property alone predicts task success, but their combined effect does, which we propose as a skill utility score. The score correlates consistently with task success when skills are transferred, and subtask-level and text skills score higher. Computing skill utility only needs the skills and task descriptions but not any task execution, so our score serves as a practical diagnostic of a skill memory before any new task runs.