This study compares Double Deep Q-Learning with K-means-partitioned mixtures of experts and parameter-matched dense networks for BTC/USDT order execution. The evaluation uses five-minute mean-aggregated Binance limit order book data and focuses on…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
303 documents
MintEval evaluates whether language models translate trader instructions into code that behaves like the intended strategy. It builds reference strategies from composable components, turns them into colloquial instructions, and asks models to implement them.…
This paper studies portfolio choice over an individual’s life when income is stochastic and investments include stocks, a bond, and life insurance. The objective accounts for consumption, death benefits, and terminal wealth. A convex trading constraint…
This study tests whether proximal policy optimization can learn a broker’s trading-speed decisions in a continuous-time broker–trader game with an analytical solution. It derives a discrete reward from the continuous-time objective and checks the…
The document describes using the Graph-based Coalition Structure Generation algorithm, GCS-Q, to cluster assets represented as a signed, weighted graph of return correlations. The motivation is that common clustering approaches can lose information when they…
The paper describes an automated stock-trading approach that combines three actor-critic reinforcement learning algorithms: Proximal Policy Optimization, Advantage Actor Critic, and Deep Deterministic Policy Gradient. The agents are trained to learn trading…
The study uses deep reinforcement learning to turn a short-term trading forecast into limit-order placement and inventory decisions in a limit order book. It trains an agent in a simulated NASDAQ equity environment built from historical order-book messages,…
The paper introduces Trading Deep Q-Network (TDQN), a deep reinforcement learning approach for choosing stock market positions over time. It adapts the DQN method to trading and sets the Sharpe ratio as the performance measure the strategy seeks to maximize.…
The paper proposes a single-directional Transformer model, SERT, for pricing large-cap US stocks and applies pre-trained Transformers to stock pricing and factor investment. It compares these approaches with standard and encoder-only Transformers across…
The paper challenges published claims about AI and machine-learning trading agents. It argues that some earlier evaluations relied on a limited number of test sessions and market scenarios, leaving too little evidence to judge the agents’ performance…
The document introduces a method for inferring lead-lag networks from the actions of individual agents, then applies it to trader-resolved foreign exchange data. It presents network persistence as an explanation for why one trader’s activity can help predict…
The study asks whether diversity among language model families persists in trading after agents are selected into submitting orders. In synthetic markets containing a fixed mixture of three model families, different presentations of news alter which families…
The document outlines a credit spread forecasting approach that combines ensemble learning with feature selection based on mutual information. Credit spreads are framed as useful inputs for bond investment decisions, and the proposed model aims to improve…
The document outlines a numerical method for learning a risk-neutral measure over simulated paths of spot and option prices up to a finite horizon. The setting includes convex transaction costs and convex trading constraints. The learned measure is the…
This paper describes an asymptotic expansion method for stochastic filtering, aimed at approximating a conditional distribution in nonlinear settings. It transforms the problem into momentum space using a Fourier transform and approximates nonlinear terms…
This article introduces moving average reversion (MAR), a multi-period alternative to the single-period mean-reversion assumption used by some online portfolio strategies. It proposes On-Line Moving Average Reversion (OLMAR), which applies online learning…
This work uses machine learning to find minimal equivalent martingale measures for simulated markets containing tradable instruments, such as a spot asset and options on that asset. It extends the approach to markets with trading frictions by seeking…
The paper studies a fundamental conflict in time-series model validation: training runs need enough observations, test folds should cover enough of the sample, and training must precede each test point. It formalizes these aims through bounds relating…
The paper tests whether Google search activity can help predict which S&P 100 stocks will outperform the index median the next day. It combines lagged financial variables with search query volumes and trains gradient boosted decision trees to classify those…
This work develops a reinforcement learning approach for optimizing policies under risk-aware performance criteria. It evaluates policies using rank-dependent expected utility, which lets an agent express different preferences over gains and downside…
This study examines how classifiers that can abstain from prediction can shape trading decisions. A standard direction classifier always implies a market position, whereas a selective classifier can leave the portfolio uninvested when it declines to predict.…
This study explores autoencoders for statistical arbitrage in US stocks. In a conventional two-stage workflow, a pricing or principal-component model identifies a synthetic asset and a separate mean-reversion strategy produces trading signals. The authors…
This study models financial prediction as a population of agents that follow trends over different prediction horizons. It examines how the population changes in stochastic, time-varying environments represented by real currency exchange-rate series. The…
The study separates factor discovery from factor approval. An agent may propose candidates and create diagnostic probes, while a frozen statistical referee evaluates each proposal using market outcomes observed only after submission. The referee uses…