Ranking US Stocks by Lexical Density in Company Filings
Summary
The document describes a stock-selection strategy using language measures calculated from companies’ 10-K and 10-Q filings. Lexical richness reflects vocabulary variety, lexical density measures the share of information-carrying language, and specific density captures the concentration of finance-related terms. The dataset described tracks these measures across thousands of US stocks. The strategy sorts the 500 most actively traded US stocks by lexical density and specific density, buys the highest-ranked tenth, sells short the lowest-ranked tenth, and rebalances monthly.
The source paper’s abstract reports that lexical richness performed weakest in the 500-stock universe but improved when the universe expanded to 3,000 stocks. It says combining lexical density and specific density in the 500-stock universe produced a Sharpe ratio of 0.688. The page also reports a backtest beta of -0.029 and suggests the strategy held up in bear markets based on an equity-curve inspection. These are reported results, not proof of robustness; the document gives limited detail about costs, sample construction, and implementation assumptions.
Key ideas
- Lexical richness, lexical density, and specific density capture different language features in company filings.
- The strategy ranks the 500 highest dollar-volume US stocks using lexical and finance-specific density scores.
- It buys the top decile and shorts the bottom decile, rebalancing each month.
- The source reports improved results from combining the two density measures, including a Sharpe ratio of 0.688.
- The reported low negative beta and bear-market behavior require independent validation and fuller testing details.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.