Choosing Labels for a Supervised Market Sentiment Index
Summary
The document considers how to train a model to estimate broad US market sentiment from headlines, social media posts, and economic indicators. The proposed output is an ordinal sentiment signal with negative, neutral, and positive classes, potentially reduced to positive and negative classes by dropping neutral observations. The central obstacle is finding historical sentiment labels that align with the proposed inputs, especially for long time periods.
It raises the choice between supervised learning, which needs a defensible target series, and unsupervised methods that might avoid relying on such labels. No historical data source, modeling method, results, or validation procedure is supplied. The proposal also leaves key design questions unresolved, including how sentiment should be defined, how labels relate to market outcomes, and how text and numeric data from different periods can be made comparable. These gaps matter because a model can reproduce its chosen label without measuring economically useful sentiment.
Key ideas
- The proposed index estimates sentiment for the US market rather than a single company.
- Candidate inputs include news headlines, social media, and economic indicators.
- A supervised classifier needs historical labels aligned with its inputs.
- The suggested label schemes use three classes or omit neutral observations for two classes.
- An unsupervised approach may avoid labeled targets but is not specified or evaluated.
Tags
Full text
# Target variable for a supervised learning approach for market sentiment index # Target variable for a supervised learning approach for market sentiment index My goal is to produce a signal going from -1 (negative) to +1 (positive) which corresponds to a sentiment index for USA. The index will be computed both based on headlines (taken from some free resources like FinViz), and twitter, and numerical data (economic indicator are likely linked to sentiment). It will not be a sentiment on a stock but instead on a geografic area. The problem is that for my training step I need the "answer". I would need, for the past, a resources which tells me which was the us sentiment from, say, 60s onwards. Clearly I would also need the headlines (input), for the same period. It is a 3-class classification problem (3 classes: -1, 0, +1), which could probably be simplified as binary (-1, +1) removing the neutral sentiment. Does anybody know any resource for such past sentiment data? Does this approach seem a good one? To avoid, struggling with the target variable, do you suggest an unsupervised approach? if yes, which one?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.