带记忆与动态遗憾的无投影在线学习
文章 arXiv papers · 作者: Hongyu Zhou et al.
总结
本文研究带记忆的在线凸优化,其中每次损失取决于当前决策和先前决策。这种依赖关系可用于建模过去的行动影响当前结果的情形。作者提出一种无需投影的元基学习算法,旨在最小化动态遗憾,即相对于不断变化的时变决策序列来衡量表现。避免投影是为了应对在线学习中常见的计算瓶颈。
该方法将在线弗兰克—沃尔夫算法与Hedge算法相结合。本文将其用于控制受不可预测过程噪声影响的线性时变系统,并构建带记忆控制器;相对于最优时变线性反馈策略,其动态遗憾有界。文中报告了在线性时不变系统上的模拟验证。描述还将统计套利和时间序列预测列为应用动机,但没有提供这些领域的交易策略、市场评估或数值表现结果。因此,此处提供的证据针对的是模拟控制,而非已验证的交易回报。
核心观点
- 带记忆的在线学习描述依赖当前和过去决策的损失。
- 所提算法避免投影操作,并以相对于变化决策的动态遗憾为优化目标。
- 该方法结合了在线弗兰克—沃尔夫算法和Hedge算法。
- 一个控制器将该方法应用于受不可预测过程噪声影响的线性系统。
- 所述验证针对模拟控制,而非交易市场表现。
标签
全文
# 2301.00497 # Efficient Online Learning with Memory via Frank-Wolfe Optimization: Algorithms with Bounded Dynamic Regret and Applications to Control Projection operations are a typical computation bottleneck in online learning. In this paper, we enable projection-free online learning within the framework of Online Convex Optimization with Memory (OCO-M) -- OCO-M captures how the history of decisions affects the current outcome by allowing the online learning loss functions to depend on both current and past decisions. Particularly, we introduce the first projection-free meta-base learning algorithm with memory that minimizes dynamic regret, i.e., that minimizes the suboptimality against any sequence of time-varying decisions. We are motivated by artificial intelligence applications where autonomous agents need to adapt to time-varying environments in real-time, accounting for how past decisions affect the present. Examples of such applications are: online control of dynamical systems; statistical arbitrage; and time series prediction. The algorithm builds on the Online Frank-Wolfe (OFW) and Hedge algorithms. We demonstrate how our algorithm can be applied to the online control of linear time-varying systems in the presence of unpredictable process noise. To this end, we develop a controller with memory and bounded dynamic regret against any optimal time-varying linear feedback control policy. We validate our algorithm in simulated scenarios of online control of linear time-invariant systems.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。