Evidence-Based Audits for Modular Trading Agents
Summary
This paper proposes a claim-level audit for modular agents that plan, act, check, and refine, motivated by the limits of relying on a single aggregate task score. The audit records the evidence for each conclusion, assigns a verdict of supported, unsupported, unresolved, or not evaluated, and states the boundary within which the conclusion applies. It is intended to clarify what an evaluation establishes about an agent and its components.
Three methods contribute evidence: oracle policies measure attainable improvement for a specified action set; one-at-a-time replacement with an ideal component helps locate lost value, while allowing for downstream masking; and a separate test checks whether a verifier's score actually supports the bound attributed to it. In a synthetic market with hidden regimes, the audit finds that measured value of perfect regime information varies with the action set, a scenario generator loses much of the regime signal, and a runtime verifier can be bypassed without visible outcome changes. These findings concern one agent and environment; the protocol is the broader contribution.
Key ideas
- Aggregate task scores alone may not identify which component caused a result or what a verifier certifies.
- The audit attaches evidence, a four-way verdict, and a scope boundary to each claim.
- Oracle policies estimate attainable gains relative to an explicitly defined action set.
- Component replacement can locate lost value, but downstream effects may leave findings unresolved.
- The synthetic-market findings are specific to the studied agent and environment.
Tags
Full text
# Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents # Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent's own checks certify. We introduce a claim-specific verification audit for modular agents that plan, act, check and refine. Instead of scoring the agent, the audit scores the evidence: each conclusion is recorded with the evidence behind it, one of four verdicts (supported, unsupported, unresolved or not evaluated) and the boundary within which it holds. Three tools supply that evidence. Oracle policies measure attainable improvement under an explicitly stated action set, so that a low value can be traced to the evaluation rather than to the environment. Replacing one component at a time with a perfect counterpart locates lost value, with null results read as unresolved whenever a downstream component could mask them. A separate test asks whether the verifier's score identifies the quantity it is read as bounding. Applied to a constrained portfolio-allocation agent in a synthetic market with known hidden regimes, the audit shows that the value of perfect regime information depends on the action set used to measure it, that the scenario generator discards most of the regime signal while better local fidelity does not improve decisions, and that the runtime verifier can be bypassed with no visible change in outcomes. The contribution is the protocol and the evidential distinctions it enforces; the empirical findings are specific to the agent and environment studied.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.