Skip to content
All library documents

Building Custom Stock Indices and Comparing Their Statistics

Article Quant Q&A · Author: atrain

Summary

The discussion outlines a project to build geometric-average indices for momentum and value stocks, then compare their behavior using correlations, volume and price-range measures, momentum, and moving averages. It weighs Python and R as tools for collecting market data and performing statistical analysis. The answers describe Python’s pandas for data handling and basic statistics, statsmodels for more advanced work, and R’s built-in statistical capabilities alongside packages such as quantmod and Quandl for market data retrieval.

The main methodological caution is survivorship bias: a data source containing only securities that remain listed can distort a historical index. The thread is an initial overview rather than a complete implementation guide. It does not specify index construction rules, define how to classify value or momentum stocks, compare package capabilities in depth, or evaluate data quality and licensing. Historical constituent coverage and careful source selection are essential for credible analysis.

Key ideas

  • Python and R can both support statistical analysis of custom stock indices.
  • Python’s pandas handles data and basic statistics, while statsmodels offers additional statistical methods.
  • R can provide similar analysis, with packages available for retrieving market data by ticker and date range.
  • Using only currently listed securities can introduce survivorship bias into historical indices.
  • Index rules and historical constituent data need careful specification.

Tags

Full text
# Building custom indices; getting data from web; stats analysis; Python or R?


# Building custom indices; getting data from web; stats analysis; Python or R?












I would like to build a couple of custom indices. I would like to be able to enter ticker(s) into an input and have ohlc, volume, qualitative ...data downloaded from yahoofinance, google finance, finviz etc over x period. From this I would like to build a geometric average indices for high momentum stocks and value stocks. I would then like to conduct analysis of these indices as they relate to each other. What stocks have highest/lowest correlation over x period, volume/range analysis, momentum over x period, sma for pairs trade....Is this a job for python or R? Do you have any suggestions on what packages/resources I would need to do this kind of analysis? I appreciate your help.

## Answer by Richard at NorgateData (score 1, accepted)

https://quant.stackexchange.com/a/17183

In addition to the above answers - You should be very careful that you do not introduce survivorship bias in your creation of indices and choose your data source carefully to remove such bias. For example, Yahoo Finance only contains currently-listed securities.

## Answer by TheBlackCat (score 2)

https://quant.stackexchange.com/a/17178

Both R and Python can do this very nicely.

For Python you would need the `pandas` package and its dependencies. `pandas` has a lot of basic statistics, but for more advanced statistics like it looks like you want to do, you can use the `statsmodels` package, which can work directly with `pandas` data types. It can also download the `csv` files directly off the website if given a url, even from `https` sites. Further, it can download the sort of stock data you want for you, just by giving it a stock ticker and a date range. You can download a python distribution like anaconda or python(X,Y) which will have `pandas` and `statsmodels` built-in, so no additional installation is necessary.

R doesn't need any additional packages. It can do roughly the same things as `pandas` and `statsmodels` for your purposes. It can also download `csv` files off the web if given a url, but apparently chokes on https files (which `pandas` doesn't), although you may not even be downloading any files through these programs. You can use another tool to do this in R, though, and it will likely only add an additional line or two of code. With additional packages, such as `quantmod` or `Quandl`, it can also download stock data using a ticker and date range.

## Answer by Andrew (score 0)

https://quant.stackexchange.com/a/17177

It is essentially a statistical exercise, so I would choose R.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.