본문으로 건너뛰기
라이브러리 문서 전체

금융 텍스트 특성: 어휘집에서 트랜스포머까지

기사 Machine Learning for Trading

요약

이 장은 사전과 단어 수 계산부터 TF-IDF, 정적 임베딩, 순환 신경망, 트랜스포머까지 금융 텍스트 표현 방법을 살펴봅니다. 간단한 방법은 빠르고 해석하기 쉬우며 금융에 맞게 조정하기 좋지만, 맥락 기반 모델은 구절의 모호성과 관계를 더 잘 처리할 수 있다는 장단점을 설명합니다. 실무 흐름에서는 사전 학습 모델, 선택적 도메인 적응, 작업별 미세 조정을 사용해 감성, 서술상의 놀라움, 주제 노출, 구조화된 이벤트 등의 신호를 만듭니다.

모델 품질만으로 특성을 거래에 쓸 수 있는 것은 아니라고 강조합니다. 룩어헤드 편향을 피하려면 공개 및 수정 타임스탬프, 엔터티 매핑, 모델 학습 종료 시점, 집계 규칙을 맞춰야 합니다. 그런 다음 데이터 범위와 이벤트 발생 시점을 고려해 적절한 거래 기간에서 평가해야 합니다. 토큰 기여도는 모델 동작을 점검하는 데 도움이 될 수 있습니다. 이 자료는 단일 전략이나 정량적 성과 결과가 아니라 방법과 작업 흐름을 설명합니다. 특성을 실제로 쓸 수 있는지는 특정 시점 기준 테스트와 세심한 진단이 필요합니다.

핵심 아이디어

  • 어휘집과 단어 수 계산은 효율적이고 해석하기 쉽지만 맥락, 부정 표현, 단어 의미 변화를 놓칩니다.
  • 정적 임베딩은 단어 간 동시 출현 기반 유사성을 포착하지만 각 단어를 하나의 표현으로 나타냅니다.
  • 순차 모델은 어순과 일부 장거리 관계를 반영하고, 트랜스포머는 어텐션으로 맥락 표현을 만듭니다.
  • 금융 텍스트 모델은 감성 분류 및 이벤트 추출과 같은 작업에 맞게 조정하고 미세 조정할 수 있습니다.
  • 텍스트 신호를 거래 의사결정에 활용하려면 특정 시점 기준 데이터 처리와 거래 기간을 고려한 평가가 필요합니다.

태그

전문
# Chapter 10: Text Feature Engineering


# Chapter 10: Text Feature Engineering

The chapter establishes the baseline methods that made large-scale financial text analysis possible: lexicons, bag-of-words, and TF-IDF. It matters because it shows both why these methods remain useful and why they are not enough for modern trading use cases: they are fast, interpretable, and domain-adaptable, but they cannot represent context, synonymy, negation, or changing meaning across uses.

## Learning Objectives

- Distinguish lexical features, static embeddings, sequential models, and Transformers in terms of the information each representation preserves and loses
- Explain how Transformer self-attention produces contextual embeddings and why this resolves key limitations of earlier NLP methods, including polysemy and long-range dependence
- Apply a practical financial NLP workflow that combines pre-trained checkpoints, domain adaptation when needed, and task fine-tuning for classification or extraction tasks
- Design text-derived features such as sentiment, narrative surprise, or structured event signals using point-in-time-safe timestamps, model cutoffs, and aggregation rules
- Evaluate text-derived signals using horizon-aware diagnostics, coverage-aware analysis, and event-time alignment rather than benchmark accuracy alone
- Use token-level attribution and related diagnostics to audit, debug, and stress-test NLP features before deployment

## Sections

### 10.1 Lexical and Statistical Models

This section establishes the baseline methods that made large-scale financial text analysis possible: lexicons, bag-of-words, and TF-IDF. It matters because it shows both why these methods remain useful and why they are not enough for modern trading use cases: they are fast, interpretable, and domain-adaptable, but they cannot represent context, synonymy, negation, or changing meaning across uses.

### 10.2 Static Embeddings

This section explains the first major leap beyond counting words: learning dense semantic representations from co-occurrence. It matters because it introduces the core intuition behind modern representation learning and shows that the same logic can extend beyond text, as in asset embeddings learned from portfolio holdings. At the same time, it makes clear why static embeddings are still only an intermediate step: one vector per word is not enough for finance, where context changes meaning constantly.

### 10.3 Sequential Models

This section gives readers the missing bridge between static embeddings and Transformers. It matters because it explains what RNNs and LSTMs solved, what they could not solve, and why the field moved on. The reader should care because this is the architectural turning point: once long-range dependence, sequential computation, and scaling become bottlenecks, Transformer-style attention stops being a technical curiosity and becomes the practical default.

### 10.4 Transformers

This section is the chapter's conceptual center of gravity on model architecture. It explains self-attention, multi-head attention, positional encoding, BERT-style encoders, finance-specific checkpoints, and the practical consequences of domain adaptation, fine-tuning, and model choice. Readers should care because this is where the chapter moves from general NLP history to the modern tools that actually power financial text classification, embedding generation, and representation-based feature engineering.

### 10.5 The Modern Feature Extraction Workflow

This is the chapter's real practical core. It explains that a strong text model is not yet a tradable signal unless timestamps, entity resolution, revisions, aggregation rules, model training cutoffs, and evaluation protocols are all made point-in-time safe. It also connects representation models to real feature families such as sentiment, narrative surprise, topic exposure, structured event extraction, and interpretable diagnostics. Readers should care because this section is what turns NLP from a benchmark exercise into a research and production workflow suitable for systematic trading.

## Running the Notebooks

```bash
# From the repository root
uv run python 10_text_feature_engineering/<notebook>.py

# Test mode (reduced data via Papermill)
uv run pytest tests/test_chapter_notebooks.py -v -k "10_text_feature_engineering"
```

### Docker image split (chapter-specific)

Three notebooks require the `ml4t-py312` image because `gensim` has no Python 3.14 wheel:

```bash
docker compose --profile py312 run --rm py312 \
  python 10_text_feature_engineering/01_word2vec_training.py
```

| Notebook | Image |
|---|---|
| `01_word2vec_training` | `ml4t-py312` (gensim Word2Vec) |
| `02_asset_embeddings` | `ml4t-py312` (gensim Word2Vec) |
| `03_sentiment_evolution` | `ml4t-py312` (gensim GloVe loader) |
| `04_bert_finetuning` | `ml4t-gpu` (PyTorch GPU recommended) |
| `05_financial_ner_finetuning` | `ml4t-gpu` (PyTorch GPU recommended) |
| `06_finbert_cross_dataset` | `ml4t-gpu` (PyTorch GPU recommended) |
| `07_news_return_signals` | `ml4t-gpu` (PyTorch GPU recommended) |
| `08_text_feature_evaluation` | `ml4t` (CPU-only IC + quintile diagnostics) |
| `09_filing_text_signals` | `ml4t-gpu` (PyTorch GPU recommended) |

### Runtime callouts

> `02_asset_embeddings`: ~5–6 min (gensim skip-gram on ~500 13F portfolios; CPU-bound).
>
> `09_filing_text_signals`: ~7 min on GPU (FinBERT sentence-level scoring across S&P-500 MD&A filings, GPU recommended).

## References

- **Allen Huang et al.** (2020). [FinBERT—A Deep Learning Approach to Extracting Textual Information](https://doi.org/10.2139/ssrn.3910214). *SSRN Electronic Journal*.
- **Ashish Vaswani et al.** (2017). [Attention Is All You Need](http://arxiv.org/abs/1706.03762). *arXiv:1706.03762 [cs]*.
- **Benjamin Warner et al.** (2024). [Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference](https://arxiv.org/abs/2412.13663v2).
- **Dogu Araci** (2019). [FinBERT: Financial Sentiment Analysis with Pre-trained Language Models](https://doi.org/10.48550/arXiv.1908.10063).
- **Edward J. Hu et al.** (2021). [LoRA: Low-Rank Adaptation of Large Language Models](https://arxiv.org/abs/2106.09685v2).
- **Jacob Devlin et al.** (2019). [BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding](https://doi.org/10.18653/v1/N19-1423). *Association for Computational Linguistics*.
- **Jeffrey Pennington et al.** (2014). [GloVe: Global Vectors for Word Representation](https://doi.org/10.3115/v1/D14-1162). *Association for Computational Linguistics*.
- **Leland Bybee et al.** (2023). [Narrative Asset Pricing: Interpretable Systematic Risk Factors from News Text](https://doi.org/10.1093/rfs/hhad042). *The Review of Financial Studies*.
- **Leland Bybee et al.** (2024). [Business News and Business Cycles](https://doi.org/10.1111/jofi.13377). *The Journal of Finance*.
- **Nils Reimers and Iryna Gurevych** (2019). [Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks](http://arxiv.org/abs/1908.10084). *arXiv:1908.10084 [cs]*.
- **Paul C. Tetlock** (2005). [Giving Content to Investor Sentiment: The Role of Media in the Stock Market](https://doi.org/10.2139/ssrn.685145).
- **Qianqian Xie et al.** (2024). [Finben: A holistic financial benchmark for large language models](https://proceedings.neurips.cc/paper_files/paper/2024/hash/adb1d9fa8be4576d28703b396b82ba1b-Abstract-Datasets_and_Benchmarks_Track.html). *Advances in Neural Information Processing Systems*.
- **Rajeev Bhargava et al.** (2023). [Quantifying Narratives and Their Impact on Financial Markets](https://doi.org/10.3905/jpm.2023.1.472). *The Journal of Portfolio Management*.
- **Scott M Lundberg et al.** (2017). [A Unified Approach to Interpreting Model Predictions](http://papers.nips.cc/paper/7062-a-unified-approach-to-interpreting-model-predictions.pdf). *Curran Associates, Inc.*.
- **Sepp Hochreiter and Jürgen Schmidhuber** (1996). LSTM can solve hard long time lag problems. *MIT Press*.
- **Shijie Wu et al.** (2023). [BloombergGPT: A Large Language Model for Finance](https://arxiv.org/abs/2303.17564v3).
- **Stephen Robertson and Hugo Zaragoza** (2009). [The Probabilistic Relevance Framework: BM25 and Beyond](https://doi.org/10.1561/1500000019). *Found. Trends Inf. Retr.*.
- **Tim Loughran and Bill Mcdonald** (2011). [When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks](https://doi.org/10.1111/j.1540-6261.2010.01625.x). *The Journal of Finance*.
- **Tim Loughran and Bill McDonald** (2020). [Textual Analysis in Finance](https://doi.org/10.1146/annurev-financial-012820-032249). *Annual Review of Financial Economics*.
- **Tomas Mikolov et al.** (2013). [Efficient estimation of word representations in vector space](http://arxiv.org/abs/1301.3781). *arXiv preprint arXiv:1301.3781*.
- **Xavier Gabaix et al.** (2025). [Asset Embeddings](https://doi.org/10.3386/w33651).

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.