コンテンツへスキップ
ライブラリの全資料

センチメントとリターン分析のための金融ニュースデータ

記事 Machine Learning for Trading

サマリー

この資料では、センチメント分析、テキスト特徴量の開発、ニュースとリターンを結び付ける実験に使われる2つの金融ニュースコーパスを紹介します。FNSPIDは見出しを株式ティッカーに結び付け、広い過去期間をカバーします。全コレクションに加えて、より小規模なサンプルも利用できます。Bloombergのアーカイブには、より短い期間のニューステキストと構造化金融系列が含まれ、センチメント、トピック、データセット間の頑健性を調べるための補助的な情報源として紹介されています。どちらも公開データセットハブで配布されており、資料では基本的な内容と読み込み方法を説明しています。

このガイドでは実務上の重要な制約も説明します。FNSPIDには記載されたライセンスに基づく帰属表示が必要ですが、Bloombergのテキストは研究用途に限定され、商用での再配布はできません。Bloombergの構造化データはニュース読み込み処理とは別に扱い、新しいパイプラインにはFNSPIDを推奨しています。これらのコーパスは研究用の入力データであり、ニュース特徴量がリターンを予測する証拠ではありません。利用者は独自のシグナル、タイムスタンプ、評価設計を定義して検証する必要があります。

主なアイデア

  • FNSPIDは金融ニュース見出しを株式ティッカーに結び付け、ニュースセンチメントやリターン帰属の研究に利用できます。
  • Bloombergのアーカイブは、ニューステキストと構造化金融データを組み合わせ、補助的な実験に使えます。
  • コーパスごとに、対象範囲、規模、テキストと構造化データの読み込み方法が異なります。
  • FNSPIDには帰属表示の要件があり、Bloombergのアーカイブは研究用途に限定されています。
  • データセットは研究資料を提供しますが、それだけでニュースがリターンを予測すると立証するものではありません。

タグ

全文
# Financial News Archives


# Financial News Archives

Two news corpora used for sentiment analysis, text-signal engineering,
and news-return experiments. Both distributed via HuggingFace Hub.

| Dataset                 | Records                                      | Coverage   | Disk   | Source                                              |
| ----------------------- | -------------------------------------------- | ---------- | ------ | --------------------------------------------------- |
| FNSPID                  | 15.7M headlines, 4,775 S&P 500 companies     | 1999-2023  | ~50 MB (1M sample); ~3 GB full | HuggingFace `Zihan1004/FNSPID` |
| Bloomberg news archive  | ~470k news records + structured fin-data     | 2006-2013  | ~940 MB | HuggingFace (mirrored archive)                      |

Both downloads are HuggingFace public datasets — **no API key required**,
but a free HuggingFace account (`huggingface-cli login`) makes downloads
faster and more reliable.

## FNSPID

Large-scale financial news dataset linking headlines to stock tickers.
Used in Ch10 for news-sentiment features and return-attribution
experiments.

- **Source / citation**: Zhao et al., "FNSPID: A Comprehensive Financial
  News Dataset in Time Series," *arXiv:2402.06698* (2024).
  GitHub: https://github.com/Zdong104/FNSPID_Financial_News_Dataset.
- **License**: the HuggingFace dataset page lists `cc-by-4.0` — free for
  any use (including commercial) with attribution to the FNSPID paper.
- **Size on disk**: 1M sample ~50 MB (default); full ~3 GB.
- **Runtime**: ~30 seconds for 1M sample; ~10 minutes for the full pull.

### Download

```bash
# 1M sample (default, recommended — ~50 MB)
uv run python data/alternative/news/fnspid_download.py

# Larger samples
uv run python data/alternative/news/fnspid_download.py --sample 2000000
uv run python data/alternative/news/fnspid_download.py --sample 0   # full ~15.7M rows

# Preview
uv run python data/alternative/news/fnspid_download.py --dry-run
```

Output under `$ML4T_DATA_PATH/alternative/news/fnspid/`:

```
fnspid_1000k.parquet      # 1M sample (default)
fnspid_2000k.parquet      # when --sample 2000000 is used
fnspid_full.parquet       # --sample 0
```

### Loading

```python
from data import load_fnspid

news = load_fnspid()                                               # newest sample on disk
aapl = load_fnspid(symbols=["AAPL", "MSFT"],
                   start_date="2020-01-01", end_date="2023-12-31")
```

Schema: `symbol`, `timestamp`, `title`, `body`, `source`, `url`.

### Consumers

- **Ch10**: `07_news_return_signals.py`, `08_text_feature_evaluation.py`.

## Bloomberg News Archive

Bloomberg news headlines and bodies combined with structured financial
data series, distributed via a HuggingFace-mirrored archive. Used as a
secondary corpus for sentiment / topic experiments and cross-dataset
robustness checks.

- **Source**: HuggingFace mirrored Bloomberg archive.
- **License**: Bloomberg owns the underlying text; the mirrored archive
  is distributed under HuggingFace terms **for research use only**.
  Commercial redistribution is not permitted. See the dataset card on
  HuggingFace before using for anything beyond personal study.
- **Size on disk**: ~940 MB total (news ~460 MB, structured ~480 MB).
- **Runtime**: ~3-5 minutes.

### Download

```bash
uv run python data/alternative/news/bloomberg_download.py
```

Output under `$ML4T_DATA_PATH/alternative/news/bloomberg/`:

```
bloomberg_news.parquet                   # headline + body corpus (~460 MB)
bloomberg_financial_data.parquet.gzip    # structured financial fields (~480 MB)
.cache/huggingface/download/             # HF download staging (safe to wipe)
```

### Loading

```python
from data import load_bloomberg_news
df = load_bloomberg_news(start_date="2010-01-01", end_date="2013-12-31")
```

The loader covers ``bloomberg_news.parquet`` (headline + body corpus) and
returns the article publication time as canonical ``timestamp``.
The structured-financials file is still read directly when needed —
Bloomberg corpora are secondary to FNSPID; prefer ``load_fnspid()``
when building new pipelines.

### Consumers

- **Ch22**: ESG RAG-vs-fine-tune comparison (``06_esg_rag_vs_finetune.py``).

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。