유인 가능한 동적 스펙트럼 위험 측도를 활용한 심층 강화학습
기사 arXiv papers · 저자: Anthony Coache et al.
요약
이 문서는 에이전트가 시간 일관적인 동적 스펙트럼 위험 측도를 최적화하는 위험 민감 강화학습 프레임워크를 제시합니다. 조건부 유인 가능성을 이용해 엄밀히 일관적인 점수 함수를 구성하고, 추정 과정에서 페널티로 사용합니다. 심층 신경망으로 위험 측도를 추정하며, 저자들은 이러한 신경망으로 이 측도군을 임의의 정확도까지 근사할 수 있음을 증명합니다.
이 틀에는 중첩 전이를 추가하지 않고 전체 에피소드에서 학습하는 액터-크리틱 알고리즘도 포함됩니다. 저자들은 모의 데이터와 실제 데이터를 사용해 통계적 차익거래와 포트폴리오 배분의 두 응용 사례에서 중첩 시뮬레이션 접근법과 비교합니다. 설명에는 수치 결과, 구현 세부 정보, 실제 데이터셋에 관한 정보가 없어 설계와 명시된 평가 범위는 알 수 있지만 성능 향상이나 강건성을 판단할 수는 없습니다. 기대 보상과 함께 손실 분포 통제가 중요한 포트폴리오 의사결정에 관련된 방법입니다.
핵심 아이디어
- 강화학습에서 시간 일관적인 동적 스펙트럼 위험 측도를 최적화하는 틀입니다.
- 조건부 유인 가능성은 추정 페널티로 사용하는 엄밀히 일관적인 점수 함수를 제공합니다.
- 심층 신경망으로 위험 측도를 추정하며 측도군의 근사 결과를 제시합니다.
- 액터-크리틱 방법은 전체 에피소드를 사용하고 추가 중첩 전이를 피합니다.
- 비교는 모의 데이터와 실제 데이터를 이용한 중첩 시뮬레이션, 통계적 차익거래, 포트폴리오 배분을 다룹니다.
태그
전문
# Conditionally Elicitable Dynamic Risk Measures for Deep Reinforcement Learning # Conditionally Elicitable Dynamic Risk Measures for Deep Reinforcement Learning We propose a novel framework to solve risk-sensitive reinforcement learning (RL) problems where the agent optimises time-consistent dynamic spectral risk measures. Based on the notion of conditional elicitability, our methodology constructs (strictly consistent) scoring functions that are used as penalizers in the estimation procedure. Our contribution is threefold: we (i) devise an efficient approach to estimate a class of dynamic spectral risk measures with deep neural networks, (ii) prove that these dynamic spectral risk measures may be approximated to any arbitrary accuracy using deep neural networks, and (iii) develop a risk-sensitive actor-critic algorithm that uses full episodes and does not require any additional nested transitions. We compare our conceptually improved reinforcement learning algorithm with the nested simulation approach and illustrate its performance in two settings: statistical arbitrage and portfolio allocation on both simulated and real data.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.