למידה מקוונת ללא הקרנות, עם זיכרון וחרטה דינמית
סיכום
המאמר עוסק באופטימיזציה קמורה מקוונת עם זיכרון, שבה כל הפסד תלוי הן בהחלטה הנוכחית והן בהחלטות קודמות. תלות כזו יכולה למדל מצבים שבהם פעולות עבר משפיעות על התוצאות בהווה. המחברים מציגים אלגוריתם למידה מטא־בסיסי ללא הקרנות, שנועד למזער חרטה דינמית, מדד ביצועים ביחס לרצפי החלטות משתנים לאורך זמן. הימנעות מהקרנות מכוונת לצוואר בקבוק חישובי נפוץ בלמידה מקוונת.
השיטה משלבת את Online Frank-Wolfe עם Hedge. המאמר מיישם אותה לבקרה של מערכות ליניאריות המשתנות בזמן וכפופות לרעש תהליך בלתי צפוי, ובונה בקר עם זיכרון וחרטה דינמית חסומה ביחס למדיניות משוב ליניארית מיטבית המשתנה בזמן. מדווחים על סימולציות במערכות ליניאריות שאינן משתנות בזמן לצורך תיקוף. התיאור מציין גם ארביטראז׳ סטטיסטי וחיזוי סדרות זמן כיישומים מניעים, אך אינו מציג עבור תחומים אלה אסטרטגיית מסחר, הערכה בשוק או תוצאות ביצועים מספריות. לפיכך, הראיות המוצגות כאן נוגעות לבקרה בסימולציה ולא לתשואות מסחר שהוכחו.
רעיונות מרכזיים
- למידה מקוונת עם זיכרון מייצגת הפסדים התלויים בהחלטות הנוכחיות והקודמות.
- האלגוריתם המוצע נמנע מפעולות הקרנה ומכוון למזער חרטה דינמית ביחס להחלטות משתנות.
- השיטה משלבת את Online Frank-Wolfe ואת Hedge.
- בקר מיישם את השיטה על מערכות ליניאריות עם רעש תהליך בלתי צפוי.
- התיקוף מתואר בהקשר של בקרה בסימולציה, ולא של ביצועים בשוקי מסחר.
תגיות
הטקסט המלא
# 2301.00497 # Efficient Online Learning with Memory via Frank-Wolfe Optimization: Algorithms with Bounded Dynamic Regret and Applications to Control Projection operations are a typical computation bottleneck in online learning. In this paper, we enable projection-free online learning within the framework of Online Convex Optimization with Memory (OCO-M) -- OCO-M captures how the history of decisions affects the current outcome by allowing the online learning loss functions to depend on both current and past decisions. Particularly, we introduce the first projection-free meta-base learning algorithm with memory that minimizes dynamic regret, i.e., that minimizes the suboptimality against any sequence of time-varying decisions. We are motivated by artificial intelligence applications where autonomous agents need to adapt to time-varying environments in real-time, accounting for how past decisions affect the present. Examples of such applications are: online control of dynamical systems; statistical arbitrage; and time series prediction. The algorithm builds on the Online Frank-Wolfe (OFW) and Hedge algorithms. We demonstrate how our algorithm can be applied to the online control of linear time-varying systems in the presence of unpredictable process noise. To this end, we develop a controller with memory and bounded dynamic regret against any optimal time-varying linear feedback control policy. We validate our algorithm in simulated scenarios of online control of linear time-invariant systems.
מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.