模块化交易智能体的循证审计
文章 arXiv papers · 作者: Ali Atiah Alzahrani
总结
本文提出一种针对模块化智能体的逐项主张审计方法。这类智能体会规划、行动、检查和改进,提出该方法的动机是单一综合任务得分存在局限。审计会记录每项结论的证据,将结论判定为有支持、无支持、未解决或未评估,并说明结论适用的边界。其目的是明确评估能够证明智能体及其组件的哪些方面。
三种方法为审计提供证据:预言机策略衡量给定行动集合下可实现的改进;逐次替换为理想组件有助于定位价值损失,同时考虑下游遮蔽效应;另有一项独立测试,用于检查验证器的得分是否确实支持归于它的界限。在一个包含隐藏状态的合成市场中,审计发现,完美状态信息的测得价值会随行动集合变化,情景生成器会丢失大部分状态信号,而运行时验证器可被绕过且结果没有明显变化。这些发现仅涉及一个智能体和一种环境;审计协议才是更广泛的贡献。
核心观点
- 仅凭综合任务得分,可能无法确定哪个组件导致了结果,也无法明确验证器证明了什么。
- 审计会为每项主张附上证据、四类判定之一和适用范围边界。
- 预言机策略会根据明确界定的行动集合估算可实现的收益。
- 替换组件有助于定位价值损失,但下游影响可能使发现仍无法确定。
- 合成市场中的发现仅适用于所研究的智能体和环境。
标签
全文
# Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents # Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent's own checks certify. We introduce a claim-specific verification audit for modular agents that plan, act, check and refine. Instead of scoring the agent, the audit scores the evidence: each conclusion is recorded with the evidence behind it, one of four verdicts (supported, unsupported, unresolved or not evaluated) and the boundary within which it holds. Three tools supply that evidence. Oracle policies measure attainable improvement under an explicitly stated action set, so that a low value can be traced to the evaluation rather than to the environment. Replacing one component at a time with a perfect counterpart locates lost value, with null results read as unresolved whenever a downstream component could mask them. A separate test asks whether the verifier's score identifies the quantity it is read as bounding. Applied to a constrained portfolio-allocation agent in a synthetic market with known hidden regimes, the audit shows that the value of perfect regime information depends on the action set used to measure it, that the scenario generator discards most of the regime signal while better local fidelity does not improve decisions, and that the runtime verifier can be bypassed with no visible change in outcomes. The contribution is the protocol and the evidential distinctions it enforces; the empirical findings are specific to the agent and environment studied.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。