跳到正文
Jones Ray

ScholarPulse 日报 2026-07-02

2026-07-02 学术简报:2 篇。LLM代理系统在仓库兼容性救援任务中表现有限,但通过系统互补可提升效果,需结合实际验证以确保可用性。

今日速览

序号标题来源日期主题推荐等级
1RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility RescuearXiv2026-07-01AI-Agent高
2Optimal Resource Utilization for Autonomous Laboratory OrchestratorsarXiv2026-07-01AI-Agent高

重点论文与技术动态

1. RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility Rescue

一句话结论

LLM代理系统在仓库兼容性救援任务中表现有限,但通过系统互补可提升效果,需结合实际验证以确保可用性。

核心内容

方法与数据

价值判断

摘要 Open-source libraries and tools are widely reused, but compatibility maintenance is expensive. Once maintainers leave, useful repositories can stop working as runtimes and dependencies evolve. We study whether LLM agents can adapt old repositories to modern environments, a task we call compatibility rescue. Unlike bug repair, compatibility rescue starts from a repository that worked in its original environment but fails after ecosystem drift. RepoRescue gives agents only the repository and its failing modern environment; the agent must diagnose the failure, locate affected code, and produce a source-code rescue that restores the historical test suite. We build RepoRescue from 193 Python and 122 Java repositories, each verified to pass historically and fail after modernization. We evaluate five deployed agent systems on Python and three on Java. Beyond full-patch pass rate, we rerun patches after removing test-file edits to measure source-only repair, add a runtime-enforced regime that blocks test edits, and validate practical use for repositories whose suites pass after rescue. We find that Claude Code systems sometimes edit failing tests even when prompted not to; with runtime blocking, Kimi still rescues 41.5% of repositories. Systems are complementary: their union reaches 62.7%, exceeding the best single system by 10.9 points. Difficulty concentrates in cross-file coordination: on 14 repositories requiring coordinated whole-codebase changes, GPT-5.2 through Codex passes all 14, while every Claude Code system passes at most two. Finally, a passing suite is only an initial signal: among 34 unmaintained Python candidates whose suites pass after rescue, 22 work in realistic scenarios and 12 pass bug-hunt with patches that address the compatibility failure. RepoRescue benchmarks compatibility rescue with source-only auditing, runtime enforcement, practical validation, and reasoning labels.

2. Optimal Resource Utilization for Autonomous Laboratory Orchestrators

一句话结论

该研究提出两步方法优化自主实验室资源利用,显著提升金属有机框架合成效率。

一段话:在自主实验室中,AI代理建议实验批次后,规划执行任务需充分利用多仪器硬件资源(考虑容量与吞吐量差异),但面临真实硬件约束挑战。本研究通过约束规划与状态依赖系统,实现调度优化和稳健执行,有效减少总时间。

核心内容

方法与数据

价值判断

摘要 In autonomous laboratories, AI agents suggest the next batch of experiments to do. However, planning and executing those tasks taking full advantage of the available resources is a completely different question. This can be challenging when dealing with real-world hardware constraints, especially so when there are multiple instruments with different capacities and throughputs. Here we demonstrate a 2-step method to address resource utilization for our autonomous platform for metal-organic framework synthesis. First, we use constraint programming to find optimal schedules. This finds schedules that minimizes the total time while still satisfying the limitations and capacities of the hardware. Secondly, we use a system of status dependencies for each task, which allows for the robust execution of the optimal schedules.