رفتن به محتوا
همه اسناد کتابخانه

ممیزی‌های مبتنی بر شواهد برای عامل‌های معاملاتی ماژولار

مقاله arXiv papers · نویسنده: Ali Atiah Alzahrani

خلاصه

این مقاله ممیزی در سطح ادعا را برای عامل‌های ماژولار پیشنهاد می‌کند که برنامه‌ریزی، اقدام، بررسی و اصلاح می‌کنند؛ انگیزه آن محدودیت‌های اتکا به یک امتیاز کلی برای کل وظیفه است. ممیزی شواهد هر نتیجه‌گیری را ثبت می‌کند، یکی از چهار رأیِ پشتیبانی‌شده، پشتیبانی‌نشده، حل‌نشده یا ارزیابی‌نشده را اختصاص می‌دهد و محدوده‌ای را مشخص می‌کند که نتیجه‌گیری در آن صدق می‌کند. هدف آن روشن‌کردن این است که ارزیابی درباره یک عامل و اجزای آن چه چیزی را نشان می‌دهد.

سه روش شواهد فراهم می‌کنند: سیاست‌های اوراکل، بهبود دست‌یافتنی را برای مجموعه کنش مشخصی اندازه می‌گیرند؛ جایگزینی تک‌به‌تک با یک مؤلفه ایده‌آل به یافتن ارزش ازدست‌رفته کمک می‌کند، با درنظرگرفتن امکان پنهان‌ماندن آن در مراحل بعد؛ و آزمونی جداگانه بررسی می‌کند آیا امتیاز یک راستی‌آزما واقعاً کران منتسب به آن را پشتیبانی می‌کند. در بازاری مصنوعی با رژیم‌های پنهان، ممیزی می‌یابد که ارزش اندازه‌گیری‌شده اطلاعات کامل درباره رژیم با مجموعه کنش تغییر می‌کند، مولد سناریو بخش بزرگی از سیگنال رژیم را از دست می‌دهد و می‌توان راستی‌آزمای بلادرنگ را بدون تغییرات مشهود در نتیجه دور زد. این یافته‌ها به یک عامل و محیط مربوط‌اند؛ سهم گسترده‌تر پژوهش، پروتکل آن است.

ایده‌های کلیدی

  • امتیاز کلی وظیفه به‌تنهایی ممکن است مؤلفه عاملِ نتیجه یا موضوعی را که راستی‌آزما تأیید می‌کند مشخص نکند.
  • ممیزی برای هر ادعا شواهد، یکی از چهار رأی و مرز دامنه اعتبار را ثبت می‌کند.
  • سیاست‌های اوراکل، سود دست‌یافتنی را نسبت به مجموعه کنشی که صریحاً تعریف شده برآورد می‌کنند.
  • جایگزینی مؤلفه می‌تواند ارزش ازدست‌رفته را مشخص کند، اما اثرات مراحل بعدی ممکن است یافته‌ها را حل‌نشده بگذارند.
  • یافته‌های بازار مصنوعی ویژه عامل و محیط بررسی‌شده‌اند.

برچسب‌ها

متن کامل
# Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents


# Verify Claims, Not Scores: Evidence-Based Verification of Modular Agents









When developers change one component of an agent, such as its controller, a learned model or its verifier, they usually judge the change by an aggregate task score. That score cannot tell whether improvement was attainable, which component lost value, or what the agent's own checks certify. We introduce a claim-specific verification audit for modular agents that plan, act, check and refine. Instead of scoring the agent, the audit scores the evidence: each conclusion is recorded with the evidence behind it, one of four verdicts (supported, unsupported, unresolved or not evaluated) and the boundary within which it holds. Three tools supply that evidence. Oracle policies measure attainable improvement under an explicitly stated action set, so that a low value can be traced to the evaluation rather than to the environment. Replacing one component at a time with a perfect counterpart locates lost value, with null results read as unresolved whenever a downstream component could mask them. A separate test asks whether the verifier's score identifies the quantity it is read as bounding. Applied to a constrained portfolio-allocation agent in a synthetic market with known hidden regimes, the audit shows that the value of perfect regime information depends on the action set used to measure it, that the scenario generator discards most of the regime signal while better local fidelity does not improve decisions, and that the runtime verifier can be bypassed with no visible change in outcomes. The contribution is the protocol and the evidential distinctions it enforces; the empirical findings are specific to the agent and environment studied.

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: abstract CC0

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.