본문으로 건너뛰기
라이브러리 문서 전체

메모리와 동적 후회를 고려한 무투영 온라인 학습

기사 arXiv papers · 저자: Hongyu Zhou et al.

요약

이 논문은 각 손실이 현재 결정과 이전 결정 모두에 좌우되는 메모리 기반 온라인 볼록 최적화를 다룹니다. 이런 의존성은 과거 행동이 현재 결과에 영향을 주는 상황을 모형화할 수 있습니다. 저자들은 변하는 시계열 의사결정에 대한 성과를 측정하는 동적 후회를 최소화하도록 설계된 무투영 메타 기반 학습 알고리즘을 소개합니다. 투영을 피하는 방식으로 온라인 학습의 일반적인 계산 병목을 해결하고자 합니다.

이 방법은 Online Frank-Wolfe와 Hedge를 결합합니다. 논문은 예측 불가능한 과정 잡음이 있는 선형 시변 시스템의 제어에 이 방법을 적용해, 최적 시변 선형 피드백 정책에 비해 동적 후회가 제한되는 메모리 기반 제어기를 구성합니다. 검증 자료로 선형 시불변 시스템의 시뮬레이션을 보고합니다. 설명은 통계적 차익거래와 시계열 예측을 동기 부여가 되는 응용 분야로 언급하지만, 해당 분야의 트레이딩 전략이나 시장 평가, 정량적인 성과 결과는 제시하지 않습니다. 따라서 여기에 제시된 근거는 입증된 트레이딩 수익이 아니라 모의 제어에 관한 것입니다.

핵심 아이디어

  • 메모리 기반 온라인 학습은 현재와 과거의 결정에 좌우되는 손실을 나타냅니다.
  • 제안된 알고리즘은 투영 연산을 피하고 변하는 결정에 대한 동적 후회를 줄이고자 합니다.
  • 이 방법은 Online Frank-Wolfe와 Hedge를 결합합니다.
  • 제어기는 예측 불가능한 과정 잡음이 있는 선형 시스템에 이 방법을 적용합니다.
  • 검증 대상은 모의 제어이며, 트레이딩 시장 성과가 아닙니다.

태그

전문
# 2301.00497


# Efficient Online Learning with Memory via Frank-Wolfe Optimization: Algorithms with Bounded Dynamic Regret and Applications to Control









Projection operations are a typical computation bottleneck in online learning. In this paper, we enable projection-free online learning within the framework of Online Convex Optimization with Memory (OCO-M) -- OCO-M captures how the history of decisions affects the current outcome by allowing the online learning loss functions to depend on both current and past decisions. Particularly, we introduce the first projection-free meta-base learning algorithm with memory that minimizes dynamic regret, i.e., that minimizes the suboptimality against any sequence of time-varying decisions. We are motivated by artificial intelligence applications where autonomous agents need to adapt to time-varying environments in real-time, accounting for how past decisions affect the present. Examples of such applications are: online control of dynamical systems; statistical arbitrage; and time series prediction. The algorithm builds on the Online Frank-Wolfe (OFW) and Hedge algorithms. We demonstrate how our algorithm can be applied to the online control of linear time-varying systems in the presence of unpredictable process noise. To this end, we develop a controller with memory and bounded dynamic regret against any optimal time-varying linear feedback control policy. We validate our algorithm in simulated scenarios of online control of linear time-invariant systems.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.