跳到正文
Jones Ray

ScholarPulse 日报 2026-09-13

2026-09-13 学术简报:2 篇。LLM judge 在专利撰写任务中可作为有效的迭代优化信号,但其评估与专业律师判断之间存在系统性校准偏差,可靠性高度依赖具体指标。

今日速览

序号标题来源日期主题推荐等级
1Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting AgentsarXiv2026-09-11AI-Agent高
2Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?arXiv2026-09-11AI-Agent高

重点论文与技术动态

1. Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents

一句话结论

LLM judge 在专利撰写任务中可作为有效的迭代优化信号,但其评估与专业律师判断之间存在系统性校准偏差,可靠性高度依赖具体指标。

核心内容

方法与数据

价值判断

摘要 arXiv:2609.13422v1 Announce Type: new Abstract: LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision. Across multiple inventions and drafting-agent configurations, judge-guided revision consistently improves judge-assessed quality, while unguided revision tends to saturate. Notably, iterative judge feedback enables a low-reasoning agent to approach the performance of a substantially more expensive high-reasoning agent. Stronger models and increased reasoning generally improve judge-assessed drafting quality, while domain-specific agentic workflows provide further gains. We validate the judge against independent evaluation by a professional patent attorney and find meaningful but strongly metric-dependent agreement and systematic calibration differences. These results highlight both the utility and limitations of LLM judges as evaluators and optimization signals for complex professional workflows.

2. Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

一句话结论

零样本LLM多智能体框架在农业长周期任务中可达到与RL智能体相当的管理效果,且在环境发生偏移时展现出更强的适应能力。

核心内容

方法与数据

价值判断

摘要 arXiv:2609.13436v1 Announce Type: new Abstract: Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as the environment changes. Existing approaches either require substantial data and retraining, or primarily focus on agents operating in the virtual world. In this work, we explore the feasibility of building a self-adaptive physical AI agent that manages long-term physical tasks in a zero-shot manner and adapts to environmental changes without human intervention. We design a multi-agent framework that integrates planning, tool calling, observation, and verification, and evaluate it on agricultural tasks against reinforcement learning (RL) agents under different weather patterns. Our results show that zero-shot LLM agents can achieve comparable management outcomes to RL agents under the same weather pattern and adapt more effectively than RL when evaluated under a shifted environment, highlighting a promising path toward self-adaptive physical AI agents.