Storing Revised Economic and Financial Time Series with Vintage Dates
Summary
The discussion addresses how to store daily or lower-frequency economic and financial observations that may be revised after their initial release. Its central method is to keep the observation’s reference period separate from the date each version becomes available. A record can include a series identifier, reference date, availability start date, optional end date, and value. To reconstruct a historical vintage, select versions available by the requested date and retain the latest applicable version for each reference period.
The replies also suggest using a versioned time-series store or an RDF and graph database, with metadata and vocabularies for organizing series and units. These are architectural suggestions rather than a comparative evaluation: the discussion provides no benchmarks or implementation results. The schema is presented as broadly implementable in ordinary databases, while the alternative tools and ontology choices are individual recommendations. The key lesson is to model publication time explicitly so that revised data can be queried as it was known at a past date, and to attach metadata that supports discovery and consistent interpretation.
Key ideas
- Store an observation’s reference period separately from when each version becomes available.
- To reconstruct a historical vintage, select the latest version available by the requested date for each reference period.
- A series identifier, value, and availability interval can represent successive revisions compactly.
- Versioned stores and graph databases are alternative approaches mentioned for handling revisions and metadata.
- Units, sources, and product details can be recorded as metadata to improve organization and search.
Tags
Full text
# database for economic & finance timeseries # database for economic & finance timeseries I am looking for a technical solution to store economic and financial timeseries (nothing intraday for now, just daily/weekly/yearly) - Most timeseries database I find do not seem to take into account the fact that timeseries can be revised. For instance Industrial Production for Sept might be revised 2 months after the initial release once the administration has reviewed all the information they receive. For that reason we need 2 time dimensions: the as of date and the publication date, so that we are able to do time travel. - I also would like to be able to tag properly each timeseries with meta data (unit, source, product, etc) so that my timeseries are properly organised and searchable thank you for your suggestions ## Answer by Helin (score 2) https://quant.stackexchange.com/a/68538 There are many ways to deal with this. One is to model the data with the following (sample) schema, which can be done with pretty much any database: - `series_id` - `date` - this is the reference period for an observation - `start_dt` - this is when the observation for `date` becomes available - `to_dt` (optional) - this is when an observation for `date` is no longer active - `value` - this is the value for period `date` To get a time series as of a particular vintage date $t$, you simply retrieve all the observations with a `start_dt` that's less than or equal to $t$ and then filter for the last observation for each `date`. This method is pretty efficient from a storage perspective (because you're mostly saving the deltas in the time series). An alternative is to look at the `VersionStore` of arctic, which can deal with this very gracefully as well. You can definitely attach metadata of your choice to each series. A good reference is Developing Time-Oriented Database Applications in SQL, which has very detailed coverage for time series revisions. ## Answer by Con Fluentsy (score 0) https://quant.stackexchange.com/a/68535 The only Time Series data bases which I know of are similar to Rob Hyndman's R repository which is a database of finalized figures for modeling experiments, the only public one which I know that is available and is revisable, is Quandl which has a very large amount of free and paid private ,commercial & industrial, financial and government data. You could try contacting Quandl just do a google search. Otherwise I can offer no other ideas. ## Answer by hroptatyr (score 0) https://quant.stackexchange.com/a/68539 As for the data model, I'm using RDF in conjunction with industry vocabulary and ontologies. Most data is published in RDF anyway. As for the database, any triplestore or graph database should do. I'm using OpenLink's Virtuoso. To your points: - AGLDWG dataset ontology provides vocabulary for this use case. More generically you could also use pav or just roll your own terms. - Same vein as 1. I use qudt for units and skos to build hierarchies.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.