Skip to content
All library documents

Designing a Financial Time Series Catalog and Update System

Article Quant Q&A · Author: hroptatyr

Summary

The document outlines requirements for managing millions of financial time series from many vendors. The workflow includes scheduling retrievals, detecting stale series, applying back-adjustments and other corrections, notifying research users, and switching to fallback vendors when a primary source is unavailable. It frames end-user access and collaboration as core design concerns.

One response proposes representing each series as a document with metadata and a reference to its stored data, then building searchable views, update subscriptions, and versioned references so past simulations can remain tied to the data they used. Another stresses that storage and access choices depend on use cases such as reporting, statistical analysis, or simulation, as well as data frequency and format. The suggestions are architectural ideas rather than a complete implementation plan; the document gives no performance benchmarks, and notes that flexible views may require faster native processing for demanding queries.

Key ideas

  • A time series platform should track retrieval schedules, stale data, corrections, and fallback sources.
  • Versioned data references can preserve the inputs used by earlier research and simulations.
  • Searchable metadata and update subscriptions can help users find data and learn about changes.
  • Choose storage formats and interfaces according to reporting, research, and simulation needs.
  • The proposed document database approach may need faster native query functions for demanding views.

Tags

Full text
# time series management system


# time series management system












I'm happy how we store a single time series but we somehow lack a system that glues them all together. I'm talking about a few million time series coming from ~50 data vendors and representing maybe a million contracts.

It appears to me that hydrologists(!) have a pretty decent framework (KiTSM) but it takes A LOT of imagination to apply their (GIS-based) system to financial time series.

I imagine something with a comprehensive set of command line tools for batch processing, a neat web interface for tagging and generating custom compilations and maybe something to allow users to subscribe certain compilations, as well as some bindings for the big systems, maybe a matlab/R/octave/SAS/you-name-it plugin.

I'm not particularly fixated on exactly these features, my (daily) work flow in detail:

- schedule data retrieval plans (somewhat like cron jobs) and monitor them, i.e. get a list of time series that haven't been updated this morning

- get all back-adjustments and other corrections to time series

- inform research groups that currently use those time series about corrections

- provide fail-over time series/data vendors on request, e.g. resort to CSI settlement prices when CME's DataMine service is down

Something like this is best given into the hands of the end-users because they probably have a better idea of what might suit their specific needs.

Does a tool like this exist?

## Answer by Dan (score 1)

https://quant.stackexchange.com/a/2430

I like couchdb + couchapp for this. Each timeseries is a doc with a reference to a file somewhere, and you can just update with metadata as you go.

It's nice because all of your web views / interfaces are just a pure JavaScript / HTML + the js map/reduce view. Each one is small and self-contained, and doesn't require a separate app running somewhere.

In addition, you build some views for finding datasets according to your criteria. Everything is REST, so it's easy to wire up an app to query this.

To run a simulation / analysis job, you search for the datasets, pick up the right files, and run. Simulation results can then be stored in couch, with custom views to see results. Because entries are versioned, you can store a reference to a given piece of data, and then if you update it later, future searches pick up the newest version but old results are still valid.

Finally, couch lets you subscribe to an event stream of updates. So you can write something that listens for interesting updates / datasets and notifies the right people very easily.

Couchdb is good for availability but some pure js views are not extremely performant. For those, you can write native erlang map/reduce functions that are faster.

## Answer by kfmfe04 (score 0)

https://quant.stackexchange.com/a/2413

Like most software projects, your decision should be based on what your researchers intend to use the data for. Do they need mark-to-market for real-time or end-of-day reports or is it just for statistical analysis? Do they need bid-ask data? Tick data or minute bars? Is the data periodic or aperiodic?

For statistical analysis, the data format should be amenable to S-Plus, R, Matlab, or whatever statistical software is being used. For reports, your data might need to be stored in a column-based RDBMS or a kdb-type database. For simulations, it will depend on what the simulations are written in (C++, Java?) - a binary format may be fastest here, but it wouldn't be very portable to other uses. You may end up storing your data in multiple formats, but then you have to deal with keeping all your data in-sync.

The success of your project (measured by how much your data is used) will totally depend on how accessible your data is to your users. So, as you briefly touched upon, the aptitude of your users towards software/data will factor into your decisions. Talk to all your users to assess their needs - you need to discern their requirements clearly before you design your project into a sub-optimal direction.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.