본문으로 건너뛰기
라이브러리 문서 전체

감성 분석과 수익률 연구용 금융 뉴스 데이터

기사 Machine Learning for Trading

요약

이 자료는 감성 분석, 텍스트 특성 개발, 뉴스와 수익률의 연관성 실험에 쓰이는 금융 뉴스 코퍼스 두 가지를 소개합니다. FNSPID는 헤드라인과 주식 티커를 연결하며 오랜 기간을 포괄합니다. 전체 모음과 함께 더 작은 표본도 이용할 수 있습니다. Bloomberg 아카이브에는 더 짧은 기간의 뉴스 텍스트와 구조화된 금융 시계열이 있으며, 감성·주제 분석과 데이터셋 간 강건성 실험을 위한 보조 자료로 제시됩니다. 두 자료 모두 공개 데이터 허브를 통해 배포되며, 문서는 기본 구성과 로딩 방법을 설명합니다.

가이드에서는 중요한 실무상 제한도 설명합니다. FNSPID는 명시된 라이선스에 따라 출처를 표시해야 하며, Bloomberg 텍스트는 연구용으로 제한되어 상업적으로 재배포할 수 없습니다. Bloomberg 구조화 데이터는 뉴스 로더와 별도로 다루며, 새 파이프라인에는 FNSPID를 권장합니다. 이 코퍼스는 연구 입력 자료이지 뉴스 특성이 수익률을 예측한다는 근거가 아닙니다. 자체 신호와 타임스탬프, 평가 설계를 정의하고 검증해야 합니다.

핵심 아이디어

  • FNSPID는 금융 헤드라인과 주식 티커를 연결해 뉴스 감성과 수익률 귀속 연구를 지원합니다.
  • Bloomberg 아카이브는 보조 실험을 위해 뉴스 텍스트와 구조화 금융 데이터를 함께 제공합니다.
  • 두 코퍼스는 수록 범위와 규모, 텍스트 및 구조화 데이터를 불러오는 방식이 다릅니다.
  • FNSPID는 출처 표시가 필요하며 Bloomberg 아카이브는 연구용으로 제한됩니다.
  • 이 데이터셋은 연구 자료를 제공하지만 뉴스가 수익률을 예측한다는 점을 단독으로 입증하지는 않습니다.

태그

전문
# Financial News Archives


# Financial News Archives

Two news corpora used for sentiment analysis, text-signal engineering,
and news-return experiments. Both distributed via HuggingFace Hub.

| Dataset                 | Records                                      | Coverage   | Disk   | Source                                              |
| ----------------------- | -------------------------------------------- | ---------- | ------ | --------------------------------------------------- |
| FNSPID                  | 15.7M headlines, 4,775 S&P 500 companies     | 1999-2023  | ~50 MB (1M sample); ~3 GB full | HuggingFace `Zihan1004/FNSPID` |
| Bloomberg news archive  | ~470k news records + structured fin-data     | 2006-2013  | ~940 MB | HuggingFace (mirrored archive)                      |

Both downloads are HuggingFace public datasets — **no API key required**,
but a free HuggingFace account (`huggingface-cli login`) makes downloads
faster and more reliable.

## FNSPID

Large-scale financial news dataset linking headlines to stock tickers.
Used in Ch10 for news-sentiment features and return-attribution
experiments.

- **Source / citation**: Zhao et al., "FNSPID: A Comprehensive Financial
  News Dataset in Time Series," *arXiv:2402.06698* (2024).
  GitHub: https://github.com/Zdong104/FNSPID_Financial_News_Dataset.
- **License**: the HuggingFace dataset page lists `cc-by-4.0` — free for
  any use (including commercial) with attribution to the FNSPID paper.
- **Size on disk**: 1M sample ~50 MB (default); full ~3 GB.
- **Runtime**: ~30 seconds for 1M sample; ~10 minutes for the full pull.

### Download

```bash
# 1M sample (default, recommended — ~50 MB)
uv run python data/alternative/news/fnspid_download.py

# Larger samples
uv run python data/alternative/news/fnspid_download.py --sample 2000000
uv run python data/alternative/news/fnspid_download.py --sample 0   # full ~15.7M rows

# Preview
uv run python data/alternative/news/fnspid_download.py --dry-run
```

Output under `$ML4T_DATA_PATH/alternative/news/fnspid/`:

```
fnspid_1000k.parquet      # 1M sample (default)
fnspid_2000k.parquet      # when --sample 2000000 is used
fnspid_full.parquet       # --sample 0
```

### Loading

```python
from data import load_fnspid

news = load_fnspid()                                               # newest sample on disk
aapl = load_fnspid(symbols=["AAPL", "MSFT"],
                   start_date="2020-01-01", end_date="2023-12-31")
```

Schema: `symbol`, `timestamp`, `title`, `body`, `source`, `url`.

### Consumers

- **Ch10**: `07_news_return_signals.py`, `08_text_feature_evaluation.py`.

## Bloomberg News Archive

Bloomberg news headlines and bodies combined with structured financial
data series, distributed via a HuggingFace-mirrored archive. Used as a
secondary corpus for sentiment / topic experiments and cross-dataset
robustness checks.

- **Source**: HuggingFace mirrored Bloomberg archive.
- **License**: Bloomberg owns the underlying text; the mirrored archive
  is distributed under HuggingFace terms **for research use only**.
  Commercial redistribution is not permitted. See the dataset card on
  HuggingFace before using for anything beyond personal study.
- **Size on disk**: ~940 MB total (news ~460 MB, structured ~480 MB).
- **Runtime**: ~3-5 minutes.

### Download

```bash
uv run python data/alternative/news/bloomberg_download.py
```

Output under `$ML4T_DATA_PATH/alternative/news/bloomberg/`:

```
bloomberg_news.parquet                   # headline + body corpus (~460 MB)
bloomberg_financial_data.parquet.gzip    # structured financial fields (~480 MB)
.cache/huggingface/download/             # HF download staging (safe to wipe)
```

### Loading

```python
from data import load_bloomberg_news
df = load_bloomberg_news(start_date="2010-01-01", end_date="2013-12-31")
```

The loader covers ``bloomberg_news.parquet`` (headline + body corpus) and
returns the article publication time as canonical ``timestamp``.
The structured-financials file is still read directly when needed —
Bloomberg corpora are secondary to FNSPID; prefer ``load_fnspid()``
when building new pipelines.

### Consumers

- **Ch22**: ESG RAG-vs-fine-tune comparison (``06_esg_rag_vs_finetune.py``).

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.