コンテンツへスキップ
ライブラリの全資料

モジュール型トレードエージェントの根拠に基づく監査

記事 arXiv papers · 著者: Ali Atiah Alzahrani

サマリー

この論文は、単一の総合タスクスコアに依存する限界を踏まえ、計画、実行、確認、改善を行うモジュール型エージェントを主張単位で監査する方法を提案しています。監査では各結論の根拠を記録し、「裏付けあり」「裏付けなし」「未解決」「未評価」のいずれかの判定を付け、結論が当てはまる範囲を示します。評価によってエージェントとその構成要素について何が明らかになるのかを明確にすることを目的としています。

根拠には3つの方法を用います。オラクル方策は指定された行動集合で達成可能な改善を測定します。理想的な構成要素に一度に1つずつ置き換える方法は、後続段階で効果が覆い隠される可能性を考慮しながら、価値の損失箇所を特定します。また、検証器のスコアが、そこに帰属させた上限を実際に裏付けるかを別のテストで確認します。潜在レジームを含む合成市場では、完全なレジーム情報の測定価値が行動集合によって変わり、シナリオ生成器がレジームのシグナルの多くを失い、実行時の検証器が結果に目立った変化を生じさせずに回避され得ることが監査で判明しました。これらの結果は1つのエージェントと環境に関するもので、より広い貢献は監査手順です。

主なアイデア

  • 総合タスクスコアだけでは、どの構成要素が結果を生んだか、検証器が何を証明するかを特定できない場合があります。
  • 監査では、各主張に根拠、4段階の判定、適用範囲を付けます。
  • オラクル方策は、明示的に定義された行動集合に対する達成可能な改善を推定します。
  • 構成要素の置き換えで価値の損失箇所を特定できますが、後続段階の影響により結論が未解決となる場合があります。
  • 合成市場での結果は、調査対象のエージェントと環境に固有です。

タグ

全文
# Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents


# Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents









When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent's own checks certify. We introduce a claim-specific verification audit for modular agents that plan, act, check and refine. Instead of scoring the agent, the audit scores the evidence: each conclusion is recorded with the evidence behind it, one of four verdicts (supported, unsupported, unresolved or not evaluated) and the boundary within which it holds. Three tools supply that evidence. Oracle policies measure attainable improvement under an explicitly stated action set, so that a low value can be traced to the evaluation rather than to the environment. Replacing one component at a time with a perfect counterpart locates lost value, with null results read as unresolved whenever a downstream component could mask them. A separate test asks whether the verifier's score identifies the quantity it is read as bounding. Applied to a constrained portfolio-allocation agent in a synthetic market with known hidden regimes, the audit shows that the value of perfect regime information depends on the action set used to measure it, that the scenario generator discards most of the regime signal while better local fidelity does not improve decisions, and that the runtime verifier can be bypassed with no visible change in outcomes. The contribution is the protocol and the evidential distinctions it enforces; the empirical findings are specific to the agent and environment studied.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。