跳到正文
Jones Ray

ScholarPulse 日报 2026-08-23

2026-08-23 学术简报:2 篇。金融合规评估应聚焦于规则接地的行动和证据使用,而非单一合规分数。

今日速览

序号标题来源日期主题推荐等级
1ReguSim: Evaluating LLM Agent Rule Grounding in Financial CompliancearXiv2026-08-20AI-Agent高
2Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysisarXiv2026-08-20AI-Agent高

重点论文与技术动态

1. ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

一句话结论

金融合规评估应聚焦于规则接地的行动和证据使用,而非单一合规分数。

核心内容

方法与数据

价值判断

摘要 LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence. In trader runs with DeepSeek V4 Pro and Gemini 3.5 Flash, visible rules reduce but do not eliminate rejected actions, and incentive or persona framing shifts behavior. A bridge study shows that trader rationales can mislead an independent monitor unless enforcement evidence is shown. In monitoring, simple structured baselines either match or exceed prompt-only LLMs. The results frame financial compliance evaluation as an audit of rule-grounded actions and evidence use, rather than a single compliance score.

2. Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis

一句话结论

Brain Researcher 通过规则化分析流程显著提升神经影像数据分析的严谨性与准确性。

核心内容

方法与数据

价值判断

摘要 AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria. We present Brain Researcher, an agentic research harness operating in a neuroimaging researcher's computational environment under rules for admissible analyses, required checks and claim scope. In benchmarks, Brain Researcher increased first-choice tool-selection accuracy across seven models by 70.2 percentage points (23.3% without it versus 93.6% with it) and verifiable grounding from 4.6% to 22.0%. In collaborator-led and self-evolving studies, multiverse analyses exposed analytic-choice sensitivity, and scientific review classified claims as accepted, qualified, revised, blocked, rejected or deferred. By linking decisions to evidence and provenance, Brain Researcher embeds methodological judgment within the workflow, not after it.