Skip to content
All library documents

Dictionary-Based Financial Sentiment Analysis with N-Grams in R

Article QuantInsti blog

Summary

This article presents a basic text sentiment workflow for financial commentary using R. It applies a bag-of-words approach: prepare positive and negative term lists, clean an earnings-call transcript, tokenize it into words and n-grams, and compare the resulting term-document matrix with the dictionaries. Matched positive and negative terms are counted to produce a net sentiment score. The example analyzes management commentary from Eicher Motors’ fourth-quarter 2015 earnings call.

The model reports 14 positive matches and 4 negative matches, yielding a score of 10, which the authors interpret as positive management sentiment. They compare this interpretation with reported quarterly results and the stock’s movement on the announcement day. This is a single illustrative case, not a test of predictive trading performance: it does not establish that the score forecast returns or could be traded profitably. The dictionary was assembled from a small set of earlier company transcripts plus industry terms, so results depend on vocabulary selection and may miss context, negation, or meaning that changes across usage.

Key ideas

  • A dictionary method scores text by matching terms classified as positive or negative.
  • Text cleaning and n-gram tokenization prepare commentary for comparison with sentiment dictionaries.
  • The example applies the method to Eicher Motors management commentary from one earnings call.
  • A positive document score is compared with company results and the share price reaction.
  • The example does not test whether sentiment scores predict returns or produce a trading edge.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.