コンテンツへスキップ
ライブラリの全資料

メモリと動的リグレットを扱う射影不要のオンライン学習

記事 arXiv papers · 著者: Hongyu Zhou et al.

サマリー

この論文は、各損失が現在の判断だけでなく過去の判断にも依存する、メモリを伴うオンライン凸最適化を扱っています。こうした依存関係は、過去の行動が現在の結果に影響する状況をモデル化できます。著者らは、変化する時系列の意思決定列を基準に性能を測る動的リグレットの最小化を目指す、射影不要のメタ・ベース学習アルゴリズムを導入しています。射影を避けることで、オンライン学習でよくある計算上のボトルネックに対処します。

この手法は、Online Frank-WolfeとHedgeを組み合わせています。予測不能なプロセスノイズを受ける線形時変システムの制御に適用し、最適な時変線形フィードバックポリシーを基準とする、メモリ付きで動的リグレットが有界なコントローラーを構築しています。線形時不変システムのシミュレーションを検証として報告しています。説明では統計的裁定や時系列予測も応用の動機として挙げていますが、これらの分野についてトレード戦略、市場評価、数値的な成績結果は示されていません。ここで提示されている証拠は実証済みの取引リターンではなく、制御のシミュレーションに関するものです。

主なアイデア

  • メモリを伴うオンライン学習では、現在と過去の判断に依存する損失を扱います。
  • 提案アルゴリズムは射影操作を避け、変化する判断列に対する動的リグレットを対象とします。
  • この手法はOnline Frank-WolfeとHedgeを組み合わせています。
  • 予測不能なプロセスノイズを受ける線形システムの制御に、この手法を用いたコントローラーを適用します。
  • 説明されている検証は制御のシミュレーションに関するもので、取引市場での成績ではありません。

タグ

全文
# 2301.00497


# Efficient Online Learning with Memory via Frank-Wolfe Optimization: Algorithms with Bounded Dynamic Regret and Applications to Control









Projection operations are a typical computation bottleneck in online learning. In this paper, we enable projection-free online learning within the framework of Online Convex Optimization with Memory (OCO-M) -- OCO-M captures how the history of decisions affects the current outcome by allowing the online learning loss functions to depend on both current and past decisions. Particularly, we introduce the first projection-free meta-base learning algorithm with memory that minimizes dynamic regret, i.e., that minimizes the suboptimality against any sequence of time-varying decisions. We are motivated by artificial intelligence applications where autonomous agents need to adapt to time-varying environments in real-time, accounting for how past decisions affect the present. Examples of such applications are: online control of dynamical systems; statistical arbitrage; and time series prediction. The algorithm builds on the Online Frank-Wolfe (OFW) and Hedge algorithms. We demonstrate how our algorithm can be applied to the online control of linear time-varying systems in the presence of unpredictable process noise. To this end, we develop a controller with memory and bounded dynamic regret against any optimal time-varying linear feedback control policy. We validate our algorithm in simulated scenarios of online control of linear time-invariant systems.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。