Chuyển đến nội dung
Tất cả tài liệu trong thư viện

Tập dữ liệu tin tức tài chính cho nghiên cứu cảm xúc và lợi nhuận

Bài viết Machine Learning for Trading

Tóm tắt

Tài liệu tham khảo này mô tả hai kho dữ liệu tin tức tài chính dùng cho phân tích cảm xúc, phát triển đặc trưng văn bản và thử nghiệm mối liên hệ giữa tin tức với lợi nhuận. FNSPID liên kết tiêu đề tin với mã cổ phiếu và bao phủ một giai đoạn lịch sử rộng; ngoài toàn bộ bộ sưu tập còn có một mẫu nhỏ hơn. Kho lưu trữ Bloomberg chứa văn bản tin tức và chuỗi dữ liệu tài chính có cấu trúc trong một giai đoạn ngắn hơn, được giới thiệu như nguồn thứ cấp cho nghiên cứu cảm xúc, chủ đề và độ vững chắc giữa các bộ dữ liệu. Cả hai được phân phối qua một trung tâm dữ liệu công khai; tài liệu nêu nội dung cơ bản và các tùy chọn tải dữ liệu.

Hướng dẫn cũng giải thích các giới hạn thực tế quan trọng. FNSPID yêu cầu ghi công theo giấy phép được liệt kê, trong khi văn bản Bloomberg chỉ dành cho nghiên cứu và không được phép phân phối lại vì mục đích thương mại. Dữ liệu có cấu trúc của Bloomberg được xử lý riêng với trình tải tin tức của nó, và hướng dẫn khuyến nghị FNSPID cho các quy trình mới. Các kho dữ liệu này cung cấp đầu vào nghiên cứu, chứ không phải bằng chứng rằng đặc trưng tin tức dự báo lợi nhuận; người dùng vẫn cần xác định và xác thực tín hiệu, dấu thời gian và thiết kế đánh giá của riêng mình.

Ý chính

  • FNSPID liên kết tiêu đề tin tài chính với mã cổ phiếu và hỗ trợ nghiên cứu cảm xúc tin tức cũng như mối liên hệ giữa tin tức với lợi nhuận.
  • Kho lưu trữ Bloomberg kết hợp văn bản tin tức với dữ liệu tài chính có cấu trúc để phục vụ các thử nghiệm thứ cấp.
  • Các kho dữ liệu khác nhau về phạm vi, quy mô và cách tải văn bản cùng dữ liệu có cấu trúc.
  • FNSPID yêu cầu ghi công, còn kho lưu trữ Bloomberg chỉ dành cho nghiên cứu.
  • Các bộ dữ liệu cung cấp tư liệu nghiên cứu nhưng tự chúng không xác lập rằng tin tức dự báo lợi nhuận.

Thẻ

Toàn văn
# Financial News Archives


# Financial News Archives

Two news corpora used for sentiment analysis, text-signal engineering,
and news-return experiments. Both distributed via HuggingFace Hub.

| Dataset                 | Records                                      | Coverage   | Disk   | Source                                              |
| ----------------------- | -------------------------------------------- | ---------- | ------ | --------------------------------------------------- |
| FNSPID                  | 15.7M headlines, 4,775 S&P 500 companies     | 1999-2023  | ~50 MB (1M sample); ~3 GB full | HuggingFace `Zihan1004/FNSPID` |
| Bloomberg news archive  | ~470k news records + structured fin-data     | 2006-2013  | ~940 MB | HuggingFace (mirrored archive)                      |

Both downloads are HuggingFace public datasets — **no API key required**,
but a free HuggingFace account (`huggingface-cli login`) makes downloads
faster and more reliable.

## FNSPID

Large-scale financial news dataset linking headlines to stock tickers.
Used in Ch10 for news-sentiment features and return-attribution
experiments.

- **Source / citation**: Zhao et al., "FNSPID: A Comprehensive Financial
  News Dataset in Time Series," *arXiv:2402.06698* (2024).
  GitHub: https://github.com/Zdong104/FNSPID_Financial_News_Dataset.
- **License**: the HuggingFace dataset page lists `cc-by-4.0` — free for
  any use (including commercial) with attribution to the FNSPID paper.
- **Size on disk**: 1M sample ~50 MB (default); full ~3 GB.
- **Runtime**: ~30 seconds for 1M sample; ~10 minutes for the full pull.

### Download

```bash
# 1M sample (default, recommended — ~50 MB)
uv run python data/alternative/news/fnspid_download.py

# Larger samples
uv run python data/alternative/news/fnspid_download.py --sample 2000000
uv run python data/alternative/news/fnspid_download.py --sample 0   # full ~15.7M rows

# Preview
uv run python data/alternative/news/fnspid_download.py --dry-run
```

Output under `$ML4T_DATA_PATH/alternative/news/fnspid/`:

```
fnspid_1000k.parquet      # 1M sample (default)
fnspid_2000k.parquet      # when --sample 2000000 is used
fnspid_full.parquet       # --sample 0
```

### Loading

```python
from data import load_fnspid

news = load_fnspid()                                               # newest sample on disk
aapl = load_fnspid(symbols=["AAPL", "MSFT"],
                   start_date="2020-01-01", end_date="2023-12-31")
```

Schema: `symbol`, `timestamp`, `title`, `body`, `source`, `url`.

### Consumers

- **Ch10**: `07_news_return_signals.py`, `08_text_feature_evaluation.py`.

## Bloomberg News Archive

Bloomberg news headlines and bodies combined with structured financial
data series, distributed via a HuggingFace-mirrored archive. Used as a
secondary corpus for sentiment / topic experiments and cross-dataset
robustness checks.

- **Source**: HuggingFace mirrored Bloomberg archive.
- **License**: Bloomberg owns the underlying text; the mirrored archive
  is distributed under HuggingFace terms **for research use only**.
  Commercial redistribution is not permitted. See the dataset card on
  HuggingFace before using for anything beyond personal study.
- **Size on disk**: ~940 MB total (news ~460 MB, structured ~480 MB).
- **Runtime**: ~3-5 minutes.

### Download

```bash
uv run python data/alternative/news/bloomberg_download.py
```

Output under `$ML4T_DATA_PATH/alternative/news/bloomberg/`:

```
bloomberg_news.parquet                   # headline + body corpus (~460 MB)
bloomberg_financial_data.parquet.gzip    # structured financial fields (~480 MB)
.cache/huggingface/download/             # HF download staging (safe to wipe)
```

### Loading

```python
from data import load_bloomberg_news
df = load_bloomberg_news(start_date="2010-01-01", end_date="2013-12-31")
```

The loader covers ``bloomberg_news.parquet`` (headline + body corpus) and
returns the article publication time as canonical ``timestamp``.
The structured-financials file is still read directly when needed —
Bloomberg corpora are secondary to FNSPID; prefer ``load_fnspid()``
when building new pipelines.

### Consumers

- **Ch22**: ESG RAG-vs-fine-tune comparison (``06_esg_rag_vs_finetune.py``).

Hiển thị toàn văn kèm ghi nguồn theo giấy phép của tài liệu gốc. Giấy phép: MIT

Bản tóm tắt này do tác nhân nghiên cứu của Stratmill biên soạn từ tài liệu gốc; đây không phải bản sao của tài liệu.