Dữ liệu phân tích cảm xúc Financial Phrasebank cho nghiên cứu NLP
Tóm tắt
Tài liệu tham khảo này mô tả Financial Phrasebank, một tập hợp câu tin tức tài chính có gắn nhãn dùng để huấn luyện hoặc đánh giá mô hình xử lý ngôn ngữ tự nhiên. Người gán nhãn phân loại cảm xúc tích cực, trung tính hoặc tiêu cực. Ngữ liệu có bốn ngưỡng đồng thuận: tập con nghiêm ngặt nhất giữ lại những câu mà tất cả người gán nhãn đều đồng ý, còn các ngưỡng lỏng hơn bao gồm nhiều ví dụ hơn với mức đồng thuận nhãn thấp hơn. Tài liệu xác định bộ dữ liệu này là chuẩn đánh giá cho phân loại cảm xúc tài chính và nêu các ứng dụng như huấn luyện vector từ, theo dõi cảm xúc và tinh chỉnh transformer.
Tài liệu giải thích các biến thể của bộ dữ liệu trên ổ đĩa, cách hoạt động của bộ nạp dữ liệu và quy trình thay thế để truy xuất, chuyển đổi tệp nguồn sang định dạng parquet. Tài liệu cũng ghi lại trích dẫn học thuật và điều khoản cấp phép phi thương mại, ghi công và chia sẻ tương tự. Đây là hướng dẫn về bộ dữ liệu chứ không phải chiến lược giao dịch hoặc bằng chứng rằng tín hiệu cảm tính dự báo lợi suất. Các giới hạn thực tiễn gồm sự đánh đổi giữa mức đồng thuận nhãn và số lượng mẫu, cùng các hạn chế về sử dụng thương mại.
Ý chính
- Financial Phrasebank cung cấp các câu tin tức tài chính được gắn nhãn cảm xúc tích cực, trung tính hoặc tiêu cực.
- Mức đồng thuận giữa người gán nhãn cao hơn tạo ra tập con nhỏ hơn với nhãn nhất quán hơn.
- Ngữ liệu hỗ trợ chuẩn đánh giá phân loại cảm xúc và huấn luyện hoặc đánh giá mô hình NLP.
- Bộ dữ liệu được cấp phép cho nghiên cứu phi thương mại, kèm yêu cầu ghi công và chia sẻ tương tự.
- Chỉ sử dụng ngữ liệu này không chứng minh rằng nhãn cảm xúc dự báo lợi suất thị trường.
Thẻ
Toàn văn
# Text Reference Corpora
# Text Reference Corpora
Labeled text corpora used as training or evaluation data for NLP models.
Unlike `sec/` (filings we produce) or `news/` (news archives we mirror),
these are published academic datasets.
## Datasets
| Corpus | Size | Use case | Loader |
| --- | --- | --- | --- |
| Financial Phrasebank (Malo et al. 2014) | ~2,300–4,800 sentences depending on agreement level | Sentiment classification benchmark | `load_financial_phrasebank` |
## On-disk Layout
```
$ML4T_DATA_PATH/alternative/text/financial_phrasebank/
├── sentences_allagree.parquet # 100% agreement, ~2,264 rows (default)
├── sentences_75agree.parquet # ~3,453 rows
├── sentences_66agree.parquet # ~4,217 rows
└── sentences_50agree.parquet # ~4,846 rows
```
Total disk footprint: under 500 KB. First load triggers a one-time
HuggingFace download (~1-2 seconds).
## Financial Phrasebank
Academic sentiment benchmark: sentences from financial news labeled
positive/neutral/negative by human annotators. Four agreement levels
are published (100%, 75%, 66%, 50%); the `allagree` subset is the most
reliable but smallest.
**License**: Creative Commons Attribution-NonCommercial-ShareAlike 3.0
(`CC BY-NC-SA 3.0`). Free for academic and non-commercial research;
attribution to Malo et al. (2014) required.
### Download
```bash
# The loader downloads from HuggingFace on first call; no manual step required.
uv run python -c "from data import load_financial_phrasebank; df = load_financial_phrasebank(); print(df.shape)"
```
To pre-populate the cache manually:
```python
from huggingface_hub import hf_hub_download
import zipfile, polars as pl
from pathlib import Path
DATA = Path("$ML4T_DATA_PATH/alternative/text/financial_phrasebank")
zip_path = hf_hub_download("takala/financial_phrasebank",
"data/FinancialPhraseBank-v1.0.zip", repo_type="dataset")
label_map = {"negative": 0, "neutral": 1, "positive": 2}
rows = []
with zipfile.ZipFile(zip_path) as z, z.open(
"FinancialPhraseBank-v1.0/Sentences_AllAgree.txt"
) as f:
for line in f.read().decode("latin-1").strip().splitlines():
sentence, label = line.rsplit("@", 1)
rows.append({"sentence": sentence.strip(), "label": label_map[label.strip()]})
DATA.mkdir(parents=True, exist_ok=True)
pl.DataFrame(rows).write_parquet(DATA / "sentences_allagree.parquet")
```
### Loading
```python
from data import load_financial_phrasebank
# Default: 100% agreement subset (most reliable, ~2,264 sentences)
df = load_financial_phrasebank(agreement="100")
# Or lower agreement levels for more training data
df = load_financial_phrasebank(agreement="50")
```
**Reference**: Malo, P., Sinha, A., Korhonen, P., Wallenius, J., & Takala, P.
(2014). *Good debt or bad debt: Detecting semantic orientations in economic texts.*
Journal of the Association for Information Science and Technology, 65(4).
## Consumers
- Chapter 10 NB 01 (word2vec training)
- Chapter 10 NB 03 (sentiment evolution)
- Chapter 10 NB 04 (transformer fine-tuning)Hiển thị toàn văn kèm ghi nguồn theo giấy phép của tài liệu gốc. Giấy phép: MIT
Bản tóm tắt này do tác nhân nghiên cứu của Stratmill biên soạn từ tài liệu gốc; đây không phải bản sao của tài liệu.