Skip to content
All library documents

Quandl Wiki Data Provenance and Validation

Article Quant Q&A · Author: antonio

Summary

The document discusses the provenance of Quandl Open Data, also called Wiki Data, and asks whether the data are purchased from exchanges or derived by comparing and transforming other free sources. It presents the dataset as a possible testbed for researchers validating code or models with free historical data, while recognizing that its sourcing process is not fully clear from the description in the question.

An answer says data are contributed by users through broker-provided feeds and checked against free public sources. Contributors and users may notice discrepancies and investigate them, giving the process a community validation element. The answer suggests broker identities may not be disclosed, so the data are wiki-like in their contribution model without revealing the underlying quotations. This is a limited account of provenance and quality control, not an independent audit or a guarantee of accuracy; the document supplies no measured error rates or production-readiness evidence.

Key ideas

  • The question raises uncertainty about how Quandl Open Data obtains and constructs its historical records.
  • The answer describes broker-sourced user contributions that are checked against free public sources.
  • Community users may identify discrepancies and investigate their causes.
  • The source description does not disclose broker identities or provide measured accuracy evidence.

Tags

Full text
# Source of Quandl Open Data


# Source of Quandl Open Data












I am interested in Quandl Open Data, from Quandl.com

These data are also denoted as Wiki Data since it relies on users to flag errors. In particular on their website, they say:

> This new data source is different because it is “original”; the data is manufactured by us and Quandl users. The definitive version of the data actually lives on Quandl and not elsewhere.

Unfortunately it remains unclear how the term "original" should be intended, that is whether they buy data from exchanges and/or they compare and transform other free data sources.

Since these data has a straightforward API, many quants here might find it a viable testbed, to validate some code or models before going in production and it is likely that the followers of this community are involved with these free historical data and can share their knowledge.

## Answer by antonio (score 2, accepted)

https://quant.stackexchange.com/a/11376

> data is sourced via users from brokers and then validated against free public sources. There are many people watching and using the data, so if there are any differences between WIKI data and other sites, we usually discover very quickly and figure out what is going on

Perhaps the brokers sell data as part of the service so they wouldm't be happy to get their name published. So the model is Wiki-like in the sense of user contributions, but not in the sense of disclosing "quotations".

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.