Перейти к содержимому
Все документы библиотеки

Финансовые текстовые признаки: от словарей до трансформеров

Статья Machine Learning for Trading

Сводка

В этой главе рассматриваются способы представления финансовых текстов: от словарей и подсчёта слов до TF-IDF, статических эмбеддингов, рекуррентных сетей и трансформеров. Объясняются компромиссы: более простые методы быстры, интерпретируемы и адаптируемы к финансовой тематике, а контекстные модели лучше справляются с неоднозначностью и отношениями между фрагментами текста. В практическом рабочем процессе используются предварительно обученные модели, при необходимости — адаптация к предметной области, а также дообучение под конкретную задачу для создания сигналов, таких как тональность, неожиданность повествования, тематическая экспозиция и структурированные события.

В главе подчёркивается, что качества модели самого по себе недостаточно, чтобы сделать признак пригодным для торговли. Исследователям нужно согласовать временные метки публикации и изменений, сопоставление сущностей, даты отсечения обучающих данных и правила агрегации, чтобы избежать смещения заглядывания вперёд, а затем оценивать признаки на подходящих торговых горизонтах с учётом охвата и времени событий. Атрибуция токенов помогает проверять поведение модели. Материал описывает методы и рабочий процесс, а не одну стратегию или количественный результат доходности; полезность признаков всё ещё требует тестирования с учётом доступности данных на конкретный момент времени и тщательной диагностики.

Ключевые идеи

  • Словари и подсчёт слов эффективны и интерпретируемы, но не учитывают контекст, отрицания и изменение значений слов.
  • Статические эмбеддинги отражают сходство на основе совместной встречаемости, но каждому слову соответствует одно представление.
  • Последовательные модели учитывают порядок слов и некоторые дальние связи, а трансформеры используют механизм внимания для создания контекстных представлений.
  • Финансовые текстовые модели можно адаптировать и дообучать для таких задач, как классификация тональности и извлечение событий.
  • Для торговых решений текстовые сигналы требуют корректной обработки данных с учётом времени и оценки на подходящих горизонтах.

Теги

Полный текст
# Chapter 10: Text Feature Engineering


# Chapter 10: Text Feature Engineering

The chapter establishes the baseline methods that made large-scale financial text analysis possible: lexicons, bag-of-words, and TF-IDF. It matters because it shows both why these methods remain useful and why they are not enough for modern trading use cases: they are fast, interpretable, and domain-adaptable, but they cannot represent context, synonymy, negation, or changing meaning across uses.

## Learning Objectives

- Distinguish lexical features, static embeddings, sequential models, and Transformers in terms of the information each representation preserves and loses
- Explain how Transformer self-attention produces contextual embeddings and why this resolves key limitations of earlier NLP methods, including polysemy and long-range dependence
- Apply a practical financial NLP workflow that combines pre-trained checkpoints, domain adaptation when needed, and task fine-tuning for classification or extraction tasks
- Design text-derived features such as sentiment, narrative surprise, or structured event signals using point-in-time-safe timestamps, model cutoffs, and aggregation rules
- Evaluate text-derived signals using horizon-aware diagnostics, coverage-aware analysis, and event-time alignment rather than benchmark accuracy alone
- Use token-level attribution and related diagnostics to audit, debug, and stress-test NLP features before deployment

## Sections

### 10.1 Lexical and Statistical Models

This section establishes the baseline methods that made large-scale financial text analysis possible: lexicons, bag-of-words, and TF-IDF. It matters because it shows both why these methods remain useful and why they are not enough for modern trading use cases: they are fast, interpretable, and domain-adaptable, but they cannot represent context, synonymy, negation, or changing meaning across uses.

### 10.2 Static Embeddings

This section explains the first major leap beyond counting words: learning dense semantic representations from co-occurrence. It matters because it introduces the core intuition behind modern representation learning and shows that the same logic can extend beyond text, as in asset embeddings learned from portfolio holdings. At the same time, it makes clear why static embeddings are still only an intermediate step: one vector per word is not enough for finance, where context changes meaning constantly.

### 10.3 Sequential Models

This section gives readers the missing bridge between static embeddings and Transformers. It matters because it explains what RNNs and LSTMs solved, what they could not solve, and why the field moved on. The reader should care because this is the architectural turning point: once long-range dependence, sequential computation, and scaling become bottlenecks, Transformer-style attention stops being a technical curiosity and becomes the practical default.

### 10.4 Transformers

This section is the chapter's conceptual center of gravity on model architecture. It explains self-attention, multi-head attention, positional encoding, BERT-style encoders, finance-specific checkpoints, and the practical consequences of domain adaptation, fine-tuning, and model choice. Readers should care because this is where the chapter moves from general NLP history to the modern tools that actually power financial text classification, embedding generation, and representation-based feature engineering.

### 10.5 The Modern Feature Extraction Workflow

This is the chapter's real practical core. It explains that a strong text model is not yet a tradable signal unless timestamps, entity resolution, revisions, aggregation rules, model training cutoffs, and evaluation protocols are all made point-in-time safe. It also connects representation models to real feature families such as sentiment, narrative surprise, topic exposure, structured event extraction, and interpretable diagnostics. Readers should care because this section is what turns NLP from a benchmark exercise into a research and production workflow suitable for systematic trading.

## Running the Notebooks

```bash
# From the repository root
uv run python 10_text_feature_engineering/<notebook>.py

# Test mode (reduced data via Papermill)
uv run pytest tests/test_chapter_notebooks.py -v -k "10_text_feature_engineering"
```

### Docker image split (chapter-specific)

Three notebooks require the `ml4t-py312` image because `gensim` has no Python 3.14 wheel:

```bash
docker compose --profile py312 run --rm py312 \
  python 10_text_feature_engineering/01_word2vec_training.py
```

| Notebook | Image |
|---|---|
| `01_word2vec_training` | `ml4t-py312` (gensim Word2Vec) |
| `02_asset_embeddings` | `ml4t-py312` (gensim Word2Vec) |
| `03_sentiment_evolution` | `ml4t-py312` (gensim GloVe loader) |
| `04_bert_finetuning` | `ml4t-gpu` (PyTorch GPU recommended) |
| `05_financial_ner_finetuning` | `ml4t-gpu` (PyTorch GPU recommended) |
| `06_finbert_cross_dataset` | `ml4t-gpu` (PyTorch GPU recommended) |
| `07_news_return_signals` | `ml4t-gpu` (PyTorch GPU recommended) |
| `08_text_feature_evaluation` | `ml4t` (CPU-only IC + quintile diagnostics) |
| `09_filing_text_signals` | `ml4t-gpu` (PyTorch GPU recommended) |

### Runtime callouts

> `02_asset_embeddings`: ~5–6 min (gensim skip-gram on ~500 13F portfolios; CPU-bound).
>
> `09_filing_text_signals`: ~7 min on GPU (FinBERT sentence-level scoring across S&P-500 MD&A filings, GPU recommended).

## References

- **Allen Huang et al.** (2020). [FinBERT—A Deep Learning Approach to Extracting Textual Information](https://doi.org/10.2139/ssrn.3910214). *SSRN Electronic Journal*.
- **Ashish Vaswani et al.** (2017). [Attention Is All You Need](http://arxiv.org/abs/1706.03762). *arXiv:1706.03762 [cs]*.
- **Benjamin Warner et al.** (2024). [Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference](https://arxiv.org/abs/2412.13663v2).
- **Dogu Araci** (2019). [FinBERT: Financial Sentiment Analysis with Pre-trained Language Models](https://doi.org/10.48550/arXiv.1908.10063).
- **Edward J. Hu et al.** (2021). [LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/abs/2106.09685v2).
- **Jacob Devlin et al.** (2019). [BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding](https://doi.org/10.18653/v1/N19-1423). *Association for Computational Linguistics*.
- **Jeffrey Pennington et al.** (2014). [GloVe: Global Vectors for Word Representation](https://doi.org/10.3115/v1/D14-1162). *Association for Computational Linguistics*.
- **Leland Bybee et al.** (2023). [Narrative Asset Pricing: Interpretable Systematic Risk Factors from News Text](https://doi.org/10.1093/rfs/hhad042). *The Review of Financial Studies*.
- **Leland Bybee et al.** (2024). [Business News and Business Cycles](https://doi.org/10.1111/jofi.13377). *The Journal of Finance*.
- **Nils Reimers and Iryna Gurevych** (2019). [Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks](http://arxiv.org/abs/1908.10084). *arXiv:1908.10084 [cs]*.
- **Paul C. Tetlock** (2005). [Giving Content to Investor Sentiment: The Role of Media in the Stock Market](https://doi.org/10.2139/ssrn.685145).
- **Qianqian Xie et al.** (2024). [Finben: A holistic financial benchmark for large language models](https://proceedings.neurips.cc/paper_files/paper/2024/hash/adb1d9fa8be4576d28703b396b82ba1b-Abstract-Datasets_and_Benchmarks_Track.html). *Advances in Neural Information Processing Systems*.
- **Rajeev Bhargava et al.** (2023). [Quantifying Narratives and Their Impact on Financial Markets](https://doi.org/10.3905/jpm.2023.1.472). *The Journal of Portfolio Management*.
- **Scott M Lundberg et al.** (2017). [A Unified Approach to Interpreting Model Predictions](http://papers.nips.cc/paper/7062-a-unified-approach-to-interpreting-model-predictions.pdf). *Curran Associates, Inc.*.
- **Sepp Hochreiter and Jürgen Schmidhuber** (1996). LSTM can solve hard long time lag problems. *MIT Press*.
- **Shijie Wu et al.** (2023). [BloombergGPT: A Large Language Model for Finance](https://arxiv.org/abs/2303.17564v3).
- **Stephen Robertson and Hugo Zaragoza** (2009). [The Probabilistic Relevance Framework: BM25 and Beyond](https://doi.org/10.1561/1500000019). *Found. Trends Inf. Retr.*.
- **Tim Loughran and Bill Mcdonald** (2011). [When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks](https://doi.org/10.1111/j.1540-6261.2010.01625.x). *The Journal of Finance*.
- **Tim Loughran and Bill McDonald** (2020). [Textual Analysis in Finance](https://doi.org/10.1146/annurev-financial-012820-032249). *Annual Review of Financial Economics*.
- **Tomas Mikolov et al.** (2013). [Efficient estimation of word representations in vector space](http://arxiv.org/abs/1301.3781). *arXiv preprint arXiv:1301.3781*.
- **Xavier Gabaix et al.** (2025). [Asset Embeddings](https://doi.org/10.3386/w33651).

Полный текст с указанием источника опубликован на условиях его лицензии. Лицензия: MIT

Это краткое изложение подготовлено исследовательским агентом Stratmill по оригиналу и не является его копией.