본문으로 건너뛰기
라이브러리 문서 전체

NLP 연구를 위한 Financial Phrasebank 감성 데이터

기사 Machine Learning for Trading

요약

이 자료는 자연어 처리 모델의 학습이나 평가에 쓰이는 금융 뉴스 문장 라벨 데이터 모음인 Financial Phrasebank를 소개합니다. 사람 평가자가 긍정, 중립 또는 부정 감성 라벨을 붙입니다. 데이터셋은 평가자 간 합의 수준에 따라 네 가지 기준으로 제공됩니다. 가장 엄격한 하위 집합에는 모든 평가자가 동의한 문장이 포함되며, 기준이 느슨할수록 라벨 합의가 낮은 사례가 더 많이 포함됩니다. 이 자료는 데이터셋을 금융 감성 분류의 벤치마크로 소개하고, 단어 벡터 학습, 감성 추적, 트랜스포머 미세 조정 등의 용도를 제시합니다.

이 자료는 데이터셋의 디스크 저장 형식별 변형, 로더 작동 방식, 원본 파일을 가져와 Parquet 형식으로 변환하는 대안 절차를 설명합니다. 학술 인용 정보와 비상업적 이용, 저작자 표시, 동일조건 변경허락 라이선스 조건도 기록합니다. 이 문서는 데이터셋 안내서이지 트레이딩 전략이나 감성 신호가 수익률을 예측한다는 증거가 아닙니다. 실무상 한계로는 라벨 합의도와 표본 수 사이의 절충, 상업적 이용 제한이 있습니다.

핵심 아이디어

  • Financial Phrasebank는 긍정, 중립 또는 부정으로 라벨링된 금융 뉴스 문장을 제공합니다.
  • 평가자 간 합의도가 높을수록 더 작고 라벨이 일관된 하위 집합이 만들어집니다.
  • 이 데이터는 감성 분류 벤치마크와 NLP 모델 학습 또는 평가에 활용할 수 있습니다.
  • 데이터셋은 저작자 표시 및 동일조건 변경허락 요건을 지키는 비상업적 연구용 라이선스로 제공됩니다.
  • 이 데이터만 사용한다고 감성 라벨이 시장 수익률을 예측한다는 점이 입증되지는 않습니다.

태그

전문
# Text Reference Corpora


# Text Reference Corpora

Labeled text corpora used as training or evaluation data for NLP models.
Unlike `sec/` (filings we produce) or `news/` (news archives we mirror),
these are published academic datasets.

## Datasets

| Corpus | Size | Use case | Loader |
| --- | --- | --- | --- |
| Financial Phrasebank (Malo et al. 2014) | ~2,300–4,800 sentences depending on agreement level | Sentiment classification benchmark | `load_financial_phrasebank` |

## On-disk Layout

```
$ML4T_DATA_PATH/alternative/text/financial_phrasebank/
├── sentences_allagree.parquet    # 100% agreement, ~2,264 rows (default)
├── sentences_75agree.parquet     # ~3,453 rows
├── sentences_66agree.parquet     # ~4,217 rows
└── sentences_50agree.parquet     # ~4,846 rows
```

Total disk footprint: under 500 KB. First load triggers a one-time
HuggingFace download (~1-2 seconds).

## Financial Phrasebank

Academic sentiment benchmark: sentences from financial news labeled
positive/neutral/negative by human annotators. Four agreement levels
are published (100%, 75%, 66%, 50%); the `allagree` subset is the most
reliable but smallest.

**License**: Creative Commons Attribution-NonCommercial-ShareAlike 3.0
(`CC BY-NC-SA 3.0`). Free for academic and non-commercial research;
attribution to Malo et al. (2014) required.

### Download

```bash
# The loader downloads from HuggingFace on first call; no manual step required.
uv run python -c "from data import load_financial_phrasebank; df = load_financial_phrasebank(); print(df.shape)"
```

To pre-populate the cache manually:

```python
from huggingface_hub import hf_hub_download
import zipfile, polars as pl
from pathlib import Path

DATA = Path("$ML4T_DATA_PATH/alternative/text/financial_phrasebank")
zip_path = hf_hub_download("takala/financial_phrasebank",
                           "data/FinancialPhraseBank-v1.0.zip", repo_type="dataset")
label_map = {"negative": 0, "neutral": 1, "positive": 2}
rows = []
with zipfile.ZipFile(zip_path) as z, z.open(
    "FinancialPhraseBank-v1.0/Sentences_AllAgree.txt"
) as f:
    for line in f.read().decode("latin-1").strip().splitlines():
        sentence, label = line.rsplit("@", 1)
        rows.append({"sentence": sentence.strip(), "label": label_map[label.strip()]})
DATA.mkdir(parents=True, exist_ok=True)
pl.DataFrame(rows).write_parquet(DATA / "sentences_allagree.parquet")
```

### Loading

```python
from data import load_financial_phrasebank

# Default: 100% agreement subset (most reliable, ~2,264 sentences)
df = load_financial_phrasebank(agreement="100")

# Or lower agreement levels for more training data
df = load_financial_phrasebank(agreement="50")
```

**Reference**: Malo, P., Sinha, A., Korhonen, P., Wallenius, J., & Takala, P.
(2014). *Good debt or bad debt: Detecting semantic orientations in economic texts.*
Journal of the Association for Information Science and Technology, 65(4).

## Consumers

- Chapter 10 NB 01 (word2vec training)
- Chapter 10 NB 03 (sentiment evolution)
- Chapter 10 NB 04 (transformer fine-tuning)

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.