Skip to content
All library documents

Decoding Mixed ASCII and Binary Exchange Market Data

Article Quant Q&A · Author: msh210

Summary

The document explains how to convert a fixed-width exchange data feed that mixes text and binary fields into a readable tabular format. It recommends using the feed specification as a field-by-field map, then processing each record as bytes and converting values according to their documented types. Language-specific packing and unpacking utilities can help interpret binary numbers; the response mentions Python and Perl as possible approaches.

It also cautions that converting every field to text is not always the best first step. Some calculations may be more efficient on the original byte representation, depending on the feed and the values required. The document offers no complete parser or field layout for the particular NYSE product, so the specification remains necessary and the exact implementation depends on the language and message format. Its focus is data handling for order imbalance records rather than a trading signal or strategy.

Key ideas

  • Use the exchange feed specification to identify each fixed-width field and its encoding.
  • Process records as binary data and convert numeric fields according to their documented representation.
  • Packing and unpacking utilities can help interpret byte sequences in different programming languages.
  • Some operations may be more efficient before converting numeric values into text.
  • The document does not provide a complete parser for the specific feed.

Tags

Full text
# NYSE binary data, convert to ASCII


# NYSE binary data, convert to ASCII












The data product "TAQ NYSE Order Imbalances" from the New York Stock Exchange is in a format that is described pretty well in sections 4.8, 4.9, 4.10, and 5 of the document "NYSE Order Imbalances Client Specification", version 1.12, q.v. Briefly, it's a mix of ASCII and binary: stock symbols, for example, are in plain text, but padded by null bytes, while numerical fields are in binary digits represented by a byte with that binary value. All fields are fixed-width, so data rows simply follow one another.

Does anyone know how to convert this to, say, a comma-separated file?

## Answer by Matt Wolf (score 1)

https://quant.stackexchange.com/a/7220

I think you do not need to be a "Systems Programmer", certainly not an experienced one, to solve this problem:

1) Focus on the header, its your legend to the file structure. It describes the format and essentially already tells you how to decode the following messages.

2) Depending on your choice of language you then process each message in binary format and convert each item to the numeric format. In C#, some use "BitConverter" but obviously C# is not the language of choice here. If you can tell me which specific language you use to make the conversion then that would be helpful. A lot of people use Python to convert this kind of stuff to a higher level text based format such as csv or any delimiter-delimited structure.

3) Before you convert you may want to think carefully whether you may want to perform operations on the byte array representations of your numeric values (I am not familiar with your mentioned specific feed, though some feeds only output the "alpha" rather than full spread, for example, thus you need to perform add/subtract operations which can in certain cases be more optimal to perform on the byte array itself). Here is an example : https://stackoverflow.com/questions/3641274/c-sharp-int-byte-conversion

Here are couple Python examples just to show you how a simple byte[]-> int conversion could be done:

https://stackoverflow.com/questions/386753/how-do-i-convert-part-of-a-python-tuple-byte-array-into-an-integer

https://stackoverflow.com/questions/444591/convert-a-string-of-bytes-into-an-int-python

P.S.: It won't help you but I find mixed message formats very inefficient but that is not your fault. Most efficient streams only ship byte arrays, nothing else. A symbol should anyway never be in string format internally, but rather be assigned an int32 or int64 code. Mapping internally is much faster than converting each symbol of each message from byte array to string. Also, even if the symbol is decoded in ASCII that is very inefficient and blows up message size.

## Answer by Joshua Ulrich (score 1)

https://quant.stackexchange.com/a/7228

I wrote the pack R package (based on Perl's pack function), to do this for opentick (now defunct) data. You can look at the opentick package (in the CRAN archives) to see how I used it.

I just noticed that you said you're comfortable with Perl in your SO post. In that case, I'd recommend you use Perl's unpack function.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.