Finding Index Constituents and Downloading Their Price Histories in R
Summary
The document addresses how to obtain the component tickers of an index, then download their price histories with R's quantmod package. The suggested workflow is to retrieve a constituents table from a web source, parse the ticker symbols, request data for those symbols, and collect closing prices from the resulting series into a merged dataset.
The example focuses on the Dow Jones Industrial Average and demonstrates scraping a table from a financial website with an HTML parsing package before passing the extracted symbols to quantmod. It offers a practical data-acquisition pattern for analyzing a basket of securities rather than only a single index or stock. The document does not compare data vendors, show validation of the scraped symbols, or discuss adjustments and missing observations. Constituents and website layouts can change, so the table source and parsing assumptions may need maintenance; the example is not a general interface guaranteed to work for every exchange or index.
Key ideas
- Index component tickers can be sourced from a web-based constituents table.
- An HTML parsing workflow can extract symbols before requesting market data.
- Quantmod can retrieve multiple ticker histories into an environment for later processing.
- Closing-price series can be combined for cross-sectional analysis.
- Scraped constituents and page layouts may change and require validation.
Tags
Full text
# How to extract all the ticker symbols of an exchange with Quantmod in R?
# How to extract all the ticker symbols of an exchange with Quantmod in R?
I am using the Quantmod package in R for some data analysis. Now I can downbload price history of particular stocks or index with the following code:-
```
library(quantmod) # Loading quantmod library
getSymbols("^DJI", from = as.character(Sys.Date()-365*3))
```
I want to download all the ticker symbols that are composite of a particular Index such as DJI for example. What will be the best way to do that through R?
Thanks a lot in advance.
## Answer by Robert (score 0, accepted)
https://quant.stackexchange.com/a/29763
You need the components. You could use yahoo or cnn site, and read the table from that. Then get the tickers. Read all to an environment and work with the data.
```
library(XML)
urlt <- "http://money.cnn.com/data/dow30/"
doc.html = htmlTreeParse(urlt, useInternal = TRUE)
tables <- readHTMLTable(doc.html,as.data.frame=FALSE)
length(tables)
tables[[2]]
tables <- readHTMLTable(doc.html,
stringsAsFactors=FALSE,which = 2)
ticker=sapply(1:length(tables$Company),function(xs)
strsplit(rawToChar(charToRaw(text[xs])),"Â",fixed=TRUE)[[1]][1]
)
#ticker <- as.vector(as.character(ticker))
library(quantmod)
StockData <- new.env()
data <- getSymbols(ticker, env = StockData)
do.call(merge, eapply(StockData, Cl)[ticker])
```Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.