跳至正文
返回文库全部文档

用于 NLP 研究的金融语料库情绪数据

文章 《交易机器学习》

总结

本资料介绍 Financial Phrasebank,这是一个带标签的金融新闻句子集合,可用于训练或评估自然语言处理模型。人工标注者将情绪标记为正面、中性或负面。语料库提供四种标注一致性门槛:最严格的子集仅保留所有标注者意见一致的句子,较宽松的门槛则包含更多示例,但标签共识程度较低。本文将该数据集定位为金融情绪分类基准,并列举词向量训练、情绪追踪和 Transformer 微调等用途。

资料说明数据集在磁盘上的不同版本、加载器行为,以及获取源文件并将其转换为 Parquet 格式的另一种流程。它还记录学术引用,以及非商业使用、署名和相同方式共享的许可条款。这是数据集指南,而非交易策略,也不构成情绪信号能够预测收益的证据。实际限制包括标签一致性与样本数量之间的权衡,以及对商业用途的限制。

核心观点

  • Financial Phrasebank 提供标记为正面、中性或负面的金融新闻句子。
  • 标注者一致性越高,子集越小,但标签也越一致。
  • 该语料库支持情绪分类基准,以及 NLP 模型的训练或评估。
  • 该数据集可用于非商业研究,但须遵守署名和相同方式共享要求。
  • 仅使用该语料库无法证明情绪标签能够预测市场收益。

标签

全文
# Text Reference Corpora


# Text Reference Corpora

Labeled text corpora used as training or evaluation data for NLP models.
Unlike `sec/` (filings we produce) or `news/` (news archives we mirror),
these are published academic datasets.

## Datasets

| Corpus | Size | Use case | Loader |
| --- | --- | --- | --- |
| Financial Phrasebank (Malo et al. 2014) | ~2,300–4,800 sentences depending on agreement level | Sentiment classification benchmark | `load_financial_phrasebank` |

## On-disk Layout

```
$ML4T_DATA_PATH/alternative/text/financial_phrasebank/
├── sentences_allagree.parquet    # 100% agreement, ~2,264 rows (default)
├── sentences_75agree.parquet     # ~3,453 rows
├── sentences_66agree.parquet     # ~4,217 rows
└── sentences_50agree.parquet     # ~4,846 rows
```

Total disk footprint: under 500 KB. First load triggers a one-time
HuggingFace download (~1-2 seconds).

## Financial Phrasebank

Academic sentiment benchmark: sentences from financial news labeled
positive/neutral/negative by human annotators. Four agreement levels
are published (100%, 75%, 66%, 50%); the `allagree` subset is the most
reliable but smallest.

**License**: Creative Commons Attribution-NonCommercial-ShareAlike 3.0
(`CC BY-NC-SA 3.0`). Free for academic and non-commercial research;
attribution to Malo et al. (2014) required.

### Download

```bash
# The loader downloads from HuggingFace on first call; no manual step required.
uv run python -c "from data import load_financial_phrasebank; df = load_financial_phrasebank(); print(df.shape)"
```

To pre-populate the cache manually:

```python
from huggingface_hub import hf_hub_download
import zipfile, polars as pl
from pathlib import Path

DATA = Path("$ML4T_DATA_PATH/alternative/text/financial_phrasebank")
zip_path = hf_hub_download("takala/financial_phrasebank",
                           "data/FinancialPhraseBank-v1.0.zip", repo_type="dataset")
label_map = {"negative": 0, "neutral": 1, "positive": 2}
rows = []
with zipfile.ZipFile(zip_path) as z, z.open(
    "FinancialPhraseBank-v1.0/Sentences_AllAgree.txt"
) as f:
    for line in f.read().decode("latin-1").strip().splitlines():
        sentence, label = line.rsplit("@", 1)
        rows.append({"sentence": sentence.strip(), "label": label_map[label.strip()]})
DATA.mkdir(parents=True, exist_ok=True)
pl.DataFrame(rows).write_parquet(DATA / "sentences_allagree.parquet")
```

### Loading

```python
from data import load_financial_phrasebank

# Default: 100% agreement subset (most reliable, ~2,264 sentences)
df = load_financial_phrasebank(agreement="100")

# Or lower agreement levels for more training data
df = load_financial_phrasebank(agreement="50")
```

**Reference**: Malo, P., Sinha, A., Korhonen, P., Wallenius, J., & Takala, P.
(2014). *Good debt or bad debt: Detecting semantic orientations in economic texts.*
Journal of the Association for Information Science and Technology, 65(4).

## Consumers

- Chapter 10 NB 01 (word2vec training)
- Chapter 10 NB 03 (sentiment evolution)
- Chapter 10 NB 04 (transformer fine-tuning)

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。