Skip to content
All library documents

Choosing R, Python, or Scala for Financial Machine Learning

Article Quant Q&A · Author: Nikos

Summary

The document compares R, Python, and Scala as tools for financial machine learning, with particular attention to unsupervised methods. One answer favors R for in-memory numeric or categorical data and exploratory algorithm development, Scala for data too large for one machine when using Spark, and Python for text or image work. Another favors Python for production use and points to its machine learning and statistical libraries.

For unsupervised trading research, the responses name affinity propagation, DBSCAN, and Dirichlet-process k-means as clustering approaches that do not require a preset cluster count. Topic models such as latent Dirichlet allocation and online hierarchical Dirichlet processes are suggested for text. These are practitioner opinions and examples rather than a benchmark: the document supplies no comparative speed, productivity, or trading-performance evidence. Language choice depends on data size and type, available libraries, and production requirements.

Key ideas

  • R is recommended in one response for in-memory numeric or categorical analysis and experimentation.
  • Scala with Spark is suggested for data that exceeds one machine's memory.
  • Python is recommended for production and for text or image workflows.
  • Affinity propagation, DBSCAN, and Dirichlet-process k-means are examples of clustering without a fixed cluster count.
  • The document offers personal recommendations rather than a measured language comparison.

Tags

Full text
# What is the machine learning language of choice in this industry for unsupervised learning


# What is the machine learning language of choice in this industry for unsupervised learning












I was wondering from those with commercial machine learning financial experience, what the machine learning language of choice in this industry in the most general sense.

Also, what would be the best language for adaptive, unsupervised machine learning, in order to inform trading decisions.

When I say best I mean a balance of speed and developer productivity, and to some extent existing libraries, although with generic deep learning tools this might weigh in less.

## Answer by Bob (score 4)

https://quant.stackexchange.com/a/18667

R.

The others in-play are Python and, increasingly, Scala. But if you're trying to create or test a machine learning algorithm for a new problem, it's R.

Update July 2, 2017: This answer came up in my feed because of an upvote, so I suppose its worth updating.

These days there are a few key deciding factors in what language I choose for a problem. It the problem is to perform an analysis of data that can't fit in-memory on one machine, then I'm probably going to use Scala for Spark. If it can fit in-memory and the data is of a form where the relevant fields are either numbers or categories, then I'm going to use R. If the data contains text or images, then I'm more likely to use Python (unless I intend to use one of the R text modelling packages.)

## Answer by Vadim Smolyakov (score 4)

https://quant.stackexchange.com/a/34926

Python. Scikit-learn is a powerful machine learning library written in Python. In addition to libraries such as stats models and tensor-flow. Python is a stable production-level language (also used by Quantopian).

For unsupervised learning, you may want to look into affinity propagation, DBSCAN or Dirichlet-Process K-Means algorithms that do not require the knowledge of clusters ahead of time. For text data, LDA and on-line HDP are useful in learning the topics.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.