Dictionary-Based Financial Sentiment Analysis with N-Grams in R
Summary
This article presents a basic text sentiment workflow for financial commentary using R. It applies a bag-of-words approach: prepare positive and negative term lists, clean an earnings-call transcript, tokenize it into words and n-grams, and compare the resulting term-document matrix with the dictionaries. Matched positive and negative terms are counted to produce a net sentiment score. The example analyzes management commentary from Eicher Motors’ fourth-quarter 2015 earnings call.
The model reports 14 positive matches and 4 negative matches, yielding a score of 10, which the authors interpret as positive management sentiment. They compare this interpretation with reported quarterly results and the stock’s movement on the announcement day. This is a single illustrative case, not a test of predictive trading performance: it does not establish that the score forecast returns or could be traded profitably. The dictionary was assembled from a small set of earlier company transcripts plus industry terms, so results depend on vocabulary selection and may miss context, negation, or meaning that changes across usage.
Key ideas
- A dictionary method scores text by matching terms classified as positive or negative.
- Text cleaning and n-gram tokenization prepare commentary for comparison with sentiment dictionaries.
- The example applies the method to Eicher Motors management commentary from one earnings call.
- A positive document score is compared with company results and the share price reaction.
- The example does not test whether sentiment scores predict returns or produce a trading edge.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.