رفتن به محتوا
همه اسناد کتابخانه

ممیزی غربال تیترهای ESG و مقایسه خروجی طبقه‌بند با قراردادهای RAG

نوت‌بوک یادگیری ماشین برای معامله‌گری

خلاصه

این دفترچه خط‌لوله تیترهای ESG مبتنی بر کلیدواژه را بررسی می‌کند و خروجی ساختاریافته طبقه‌بند را با الزاماتی مقایسه می‌کند که دستیار تولیدِ تقویت‌شده با بازیابی باید برآورده کند. فهرست نخست کلیدواژه‌ها تیترها را انتخاب می‌کند و فهرست دوم برچسب‌های محیط‌زیستی، اجتماعی یا حاکمیتی را اختصاص می‌دهد. اعمال دسته‌بند بر کل مجموعه انتخاب‌شده، سوگیری شدید به سمت محیط‌زیست را نشان می‌دهد. دفترچه این عدم‌توازن را به دسترس‌پذیری بیشتر اصطلاحات محیط‌زیستی در فهرست انتخاب و به ترتیب دسته‌بندی نسبت می‌دهد؛ این ترتیب در موارد هم‌پوشانی، تطبیق محیط‌زیستی را در اولویت می‌گذارد.

سپس برای طبقه‌بندی احساسات با FinBERT از نمونه‌گیری طبقه‌بندی‌شده استفاده می‌کند، زمان بارگذاری مدل را از زمان استنتاج جدا می‌سازد و قرارداد شواهد RAG را شرح می‌دهد که مستلزم نقل بخش‌های مستند و خودداری از پاسخ است. هیچ پاسخ RAG تولید نمی‌شود، بنابراین مقایسه‌ای از کیفیت، تأخیر یا عملکرد وجود ندارد. ترکیب اندازه‌گیری‌شده بازتاب همین آرشیو خاص و رده‌بندی دست‌ساز است؛ این ترکیب ثابت نمی‌کند که غربال‌گری کلیدواژه ذاتاً سوگیرانه است. برای ارزیابی تعمیم‌پذیری به برچسب‌های مستقل و فهرست‌های جایگزین کلیدواژه نیاز است و نباید نسبت‌های نمونه طبقه‌بندی‌شده را با نسبت‌های مجموعه خبری منبع اشتباه گرفت.

ایده‌های کلیدی

  • پیش از نماینده تلقی‌کردن برچسب‌های دسته، ترکیب مجموعه انتخاب‌شده با کلیدواژه را اندازه بگیرید.
  • اصطلاحات انتخاب و اولویت دسته‌ها می‌توانند عدم‌توازن بزرگی در پوشش ESG ایجاد کنند.
  • طبقه‌بندی تضمین می‌کند دسته‌های اقلیت در نمونه حضور داشته باشند، اما نسبت دسته‌ها در نمونه را تغییر می‌دهد.
  • برای برآورد توان عملیاتی طبقه‌بند، هزینه یک‌باره بارگذاری مدل را از زمان استنتاج جدا کنید.
  • دفترچه الزامات شواهد RAG را مشخص می‌کند، اما هیچ پاسخی یا نتیجه مقایسه‌ای تولید نمی‌کند.

برچسب‌ها

متن کامل
# ESG Screening and RAG: Implemented Classifier vs Interface Contract


# ESG Screening and RAG: Implemented Classifier vs Interface Contract

**Docker image**: `ml4t-gpu`

**Book Reference**: Chapter 22, Section 22.8 (Applications and Strategic Choices)

Two ways to do ESG analysis, and only one of them is implemented here.

1. **Keyword screening plus a pretrained sentiment model.** Built and run:
   headlines are selected by an ESG keyword list, labelled E, S or G by a
   second keyword list, and scored by FinBERT. It produces a row per headline
   that a factor model can consume.
2. **A RAG assistant.** Not built here. What is recorded instead is the
   contract such an assistant would have to satisfy - narrative answers,
   cited spans, abstention - so that the comparison stays a comparison of
   output types rather than an invented benchmark.

The screening path turns out to demonstrate its own principal limitation
without being asked to. Section 2 measures it: over the Bloomberg archive,
the keyword screen that selects "ESG-relevant" headlines selects almost
nothing but environmental ones, and two properties of the screen itself
account for a good deal of that - its selection terms and the order its
categoriser tests them in.

**Learning objectives**

After working through this notebook you will be able to:

- Say what a fixed taxonomy costs, from a measurement rather than an
  assertion.
- Measure what a keyword screen actually selected, and name the properties of
  the screen that account for the composition it returned.
- Separate a model's inference cost from the one-off cost of loading it.
- State the evidence contract a RAG producer has to meet before its output
  can be compared with a classifier's.

**Prerequisites**: the Bloomberg financial news archive
(`data/alternative/news/bloomberg/`). No RAG answers are generated here.

## 1. Setup

```python
"""ESG Analysis: RAG vs Fine-Tuning - Comparing classification and retrieval approaches."""

import hashlib
import json
import time

import plotly.graph_objects as go

# Core imports
import polars as pl
import torch
from plotly.subplots import make_subplots
from transformers import AutoModelForSequenceClassification, AutoTokenizer, pipeline
from transformers.utils import logging as transformers_logging

# ML4T configuration
from data import load_bloomberg_news
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, show_plotly_with_alt

transformers_logging.set_verbosity_error()
```

```python
MAX_HEADLINES = 0  # total headlines to classify; 0 means 3 * PER_CATEGORY
MAX_QUESTIONS = 0
PER_CATEGORY = 7  # per-category default when MAX_HEADLINES is 0
SEED = 42
REQUIRE_GPU = True
FINBERT_MODEL = "ProsusAI/finbert"
FINBERT_REVISION = "4556d13015211d73dccd3fdd39d39232506f3e43"
```

```python
set_global_seeds(SEED)
if REQUIRE_GPU and not torch.cuda.is_available():
    raise RuntimeError("FinBERT inference requires the ml4t-gpu service with CUDA.")
INFERENCE_DEVICE = 0 if torch.cuda.is_available() else -1
print(f"FinBERT device: {'cuda:0' if INFERENCE_DEVICE == 0 else 'cpu'}")
```

## 2. What the keyword screen actually selects

The screen is two keyword lists. The first decides which headlines are
ESG-relevant at all; the second sorts those into Environmental, Social and
Governance. Both are below, and the second is applied to the whole selected
pool before any sampling, because what the pool contains decides what a
sample can show.

```python
news_df = load_bloomberg_news()
required_columns = {"timestamp", "headline"}
missing_columns = required_columns - set(news_df.columns)
if missing_columns:
    raise ValueError(f"Bloomberg news loader violates canonical schema: missing {missing_columns}.")

ESG_SELECTION_PATTERN = (
    r"(?i)(climate|carbon|emission|sustain|ESG|renewable|diversity|governance|environmental"
    r"|green.bond|pollution|social.responsibility|net.zero|solar|wind.energy|deforestation"
    r"|water.scarcity|labor.rights|board.independence|executive.compensation)"
)
esg_news = (
    news_df.filter(pl.col("headline").str.contains(ESG_SELECTION_PATTERN))
    .filter(pl.col("headline").str.len_chars() > 30)
    .sort(["timestamp", "headline"], descending=[True, False])
)
print(
    f"Archive spans {news_df['timestamp'].min():%Y-%m-%d} to {news_df['timestamp'].max():%Y-%m-%d}"
)
print(f"Headlines in archive: {news_df.height:,}")
print(f"Headlines the ESG screen selects: {esg_news.height:,}")
```

### The E, S and G taxonomy

A second keyword list sorts a selected headline into one category. The terms
are the ones an analyst would reach for first, and the order matters:
environmental is tested before social and social before governance, so a
headline about a board's climate policy counts as environmental.

```python
ESG_CATEGORY_TERMS = {
    "Environmental": [
        "carbon",
        "emission",
        "environmental",
        "environment",
        "solar",
        "sustainab",
        "net-zero",
        "net zero",
        "climate",
        "renewable",
        "wind",
        "pollution",
        "deforestation",
        "water scarcity",
        "green bond",
        "clean energy",
    ],
    "Social": [
        "worker",
        "safety",
        "labor",
        "labour",
        "diversity",
        "data breach",
        "customer",
        "human rights",
        "community",
        "social responsibility",
    ],
    "Governance": [
        "board",
        "ceo",
        "chief executive",
        "compensation",
        "shareholder",
        "governance",
        "audit",
        "disclosure",
    ],
}


def categorize_esg(text: str) -> str:
    """Assign one ESG category by keyword, or Other when no term matches."""
    lowered = text.lower()
    for category, terms in ESG_CATEGORY_TERMS.items():
        if any(term in lowered for term in terms):
            return category
    return "Other"
```

### What the pool contains

Apply the taxonomy to every selected headline before sampling any of them.

```python
esg_news = esg_news.with_columns(
    pl.col("headline").map_elements(categorize_esg, return_dtype=pl.String).alias("category")
)
pool_counts = esg_news.group_by("category").len().sort("len", descending=True)
pool_counts
```

This screen is called an ESG screen and what it returns is environmental
news. The counts are not close: social and governance headlines together are
a rounding error against the environmental ones.

Two properties of the screen itself account for a good deal of that, and both
are visible in the code above rather than inferred.

**The selection list is lopsided.** Its environmental terms are common single
words - *climate*, *carbon*, *solar*, *renewable*, *emission* - and a
headline needs only one of them. Its social and governance terms are mostly
compound phrases that a headline has to contain intact: *social
responsibility*, *labor rights*, *board independence*, *executive
compensation*. The ordinary words those topics are actually written in -
*board*, *pay*, *workers*, *safety* - appear in the categoriser's lists and
not in the selection pattern, so a headline about a workforce dispute is
never selected to be categorised in the first place.

**The categoriser resolves ties towards environmental.** It tests
environmental terms first, so a headline about a board's climate policy is
environmental and not governance.

What that does *not* establish is that keyword screening is intrinsically
environmental. Deciding that would need alternative selection lists and a set
of headlines labelled by someone other than this notebook, and neither is
here. What it does establish is the thing worth carrying: a keyword screen's
output composition is a property of the list, the imbalance can be large
enough to make the label wrong, and it costs one group-by to check.

`MAX_HEADLINES` is a budget for the whole sample, split evenly across the
three categories, so a cap set for a fast test caps what a fast test runs.

It also decides how to sample. On these proportions a uniform draw of twenty
would be overwhelmingly environmental and would more likely than not contain
no social or governance headline at all. The sample below is stratified
instead, taking the same number from each category so all three are present -
which is itself the admission that the screen could not supply them.

```python
esg_categories = ("Environmental", "Social", "Governance")
if MAX_HEADLINES > 0:
    # Quotient and remainder, so a cap below the category count yields that many
    # headlines rather than one per category.
    quota, extra = divmod(MAX_HEADLINES, len(esg_categories))
    quotas = [quota + (1 if index < extra else 0) for index in range(len(esg_categories))]
else:
    quotas = [PER_CATEGORY] * len(esg_categories)

selected_headlines = pl.concat(
    [
        group.sample(min(quota, group.height), seed=SEED)
        for category, quota in zip(esg_categories, quotas, strict=True)
        if quota and (group := esg_news.filter(pl.col("category") == category)).height
    ]
).sort(["timestamp", "headline"], descending=[True, False])
ESG_HEADLINES = selected_headlines["headline"].to_list()
selection_sha256 = hashlib.sha256(
    "\n".join(
        f"{row['timestamp'].isoformat()}|{row['headline']}"
        for row in selected_headlines.iter_rows(named=True)
    ).encode()
).hexdigest()
print(f"Stratified sample: {len(ESG_HEADLINES)} headlines")
print(dict(selected_headlines.group_by("category").len().sort("category").iter_rows()))
print(f"Selected-row SHA-256: {selection_sha256}")
```

### Sentiment-based headline classification

Uses FinBERT as a sentiment proxy. In production, a dedicated ESG
classifier (e.g., FinBERT-ESG) would replace this with a fine-grained
taxonomy output.

### Load the sentiment classifier

Production requires the pinned FinBERT snapshot. A missing dependency fails
closed instead of silently changing the model behind the reported results.


```python
def load_finbert_pipeline():
    """Return the pinned FinBERT sentiment pipeline."""
    tokenizer = AutoTokenizer.from_pretrained(
        FINBERT_MODEL,
        revision=FINBERT_REVISION,
        local_files_only=True,
    )
    model = AutoModelForSequenceClassification.from_pretrained(
        FINBERT_MODEL,
        revision=FINBERT_REVISION,
        local_files_only=True,
    )
    classifier = pipeline(
        "sentiment-analysis",
        model=model,
        tokenizer=tokenizer,
        device=INFERENCE_DEVICE,
    )
    return classifier, FINBERT_MODEL
```

### Classify the sample

Apply FinBERT, attach the E/S/G label, and time the load separately from the
inference. A throughput figure that includes the one-off cost of loading the
weights is a statement about the batch size it happened to be measured at,
not about the model - at twenty headlines the load dominates, and a reader
sizing a nightly job off it would be out by an order of magnitude.

```python
def classify_esg_headlines(headlines: list) -> pl.DataFrame:
    """Classify headlines with FinBERT, timing the load and the inference apart."""
    load_start = time.time()
    classifier, model_used = load_finbert_pipeline()
    load_seconds = time.time() - load_start

    parameter_device = next(classifier.model.parameters()).device
    if REQUIRE_GPU and parameter_device.type != "cuda":
        raise RuntimeError(f"FinBERT parameters are on {parameter_device}, not CUDA.")
    print(f"Sentiment model in use: {model_used} on {parameter_device}")

    inference_start = time.time()
    results = []
    for headline in headlines:
        result = classifier(headline[:512])[0]  # truncate to the model's maximum
        results.append(
            {
                "headline": headline,
                "model": model_used,
                "category": categorize_esg(headline),
                "sentiment": result["label"],
                "confidence": result["score"],
            }
        )
    inference_seconds = time.time() - inference_start

    per_headline_ms = 1_000 * inference_seconds / max(len(headlines), 1)
    print(f"Model load:  {load_seconds:.2f}s, paid once per process")
    print(f"Inference:   {inference_seconds:.2f}s for {len(headlines)} headlines")
    print(
        f"             {per_headline_ms:.1f} ms each, {len(headlines) / inference_seconds:.0f} per second"
    )
    print(
        f"Load is {load_seconds / (load_seconds + inference_seconds):.0%} of the wall clock at "
        f"this batch size, and a smaller share of it at every larger one."
    )

    return pl.DataFrame(results).with_columns(
        pl.lit(per_headline_ms).alias("inference_ms_per_headline"),
        pl.lit(load_seconds).alias("model_load_seconds"),
    )
```

One row per headline, with a category, a label and a confidence. That shape
is what a factor pipeline can consume, and it is the structural advantage
classification holds over a narrative answer however good the narrative is.

```python
# Run classification
print("=== Approach A: Pretrained FinBERT Inference ===\n")

classification_results = classify_esg_headlines(ESG_HEADLINES)
print("\nClassification Results:")
classification_results
```

The `confidence` column is FinBERT's softmax over three sentiment classes. It
is not a probability that the label is right, and it says nothing at all
about the ESG category beside it, which came from a keyword match with no
uncertainty attached.

## Approach B: RAG Interface Contract

A RAG implementation must answer open-ended questions with retrieved evidence,
citations, and explicit abstention. This notebook records that contract without
inventing answers, citations, confidence, or latency.

```python
# Sample ESG questions for RAG
ESG_QUESTIONS = [
    "Summarize the company's strategy for reducing Scope 2 emissions and list any stated targets.",
    "What key performance indicators does the company use to track progress on sustainability?",
    "Identify governance concerns related to executive compensation or board independence.",
]

if MAX_QUESTIONS > 0:
    ESG_QUESTIONS = ESG_QUESTIONS[:MAX_QUESTIONS]
```

### Declare the evidence contract

Each question defines the evidence a live assistant must return before its
output can enter a comparison.

```python
def build_rag_contract(questions: list[str]) -> pl.DataFrame:
    """Return validation fields required from a future live RAG run."""
    return pl.DataFrame(
        {
            "question": questions,
            "required_output": ["grounded narrative"] * len(questions),
            "required_evidence": ["source id + quoted span"] * len(questions),
            "required_checks": ["citation support + abstention"] * len(questions),
            "measured_here": [False] * len(questions),
        }
    )
```

This is a specification and not a result. `05_10k_rag_assistant` implements
the retrieval half; a live generation run would still have to be scored for
citation support and abstention before either could be set against the
classifier above.

```python
# Run RAG analysis
print("\n=== Approach B: RAG Interface Contract ===\n")

rag_contract = build_rag_contract(ESG_QUESTIONS)
print("\nRequired RAG evidence:")
rag_contract
```

`measured_here` is false in every row, and that column exists so the table
cannot be read as a result.

## Comparison: Classification vs RAG

This table separates the implemented screen from requirements that a future
RAG producer must satisfy.

```python
# Build comparison table
comparison = pl.DataFrame(
    {
        "Dimension": [
            "Primary Output",
            "Scalability",
            "Flexibility",
            "Verifiability",
            "Knowledge Updates",
            "Latency",
            "Best Use Case",
        ],
        "Implemented Screening": [
            "Keyword category + sentiment label",
            "High (batch processing)",
            "Low (fixed taxonomy)",
            "Indirect (confidence)",
            "Update rules or replace model",
            "Measured here, load and inference apart",
            "Systematic factor construction",
        ],
        "RAG Contract": [
            "Requires narrative + citations",
            "Not measured here",
            "Open-ended by design",
            "Requires cited source spans",
            "Requires corpus provenance",
            "Not measured here",
            "Due-diligence interface specification",
        ],
    }
)

print("\n=== Approach Comparison ===\n")
comparison
```

The Flexibility row is the one this run puts a number behind. A fixed
taxonomy is low-flexibility in the sense measured in section 2: the screen
could not supply social or governance headlines in proportion, and no
reordering of the keyword list would have made it.

## Decision Framework

Use this framework to select the appropriate approach:

| Question | Choose Classification | Choose RAG |
|----------|----------------------|------------|
| Need to process 1000s of documents? | **Yes** | No |
| Need numeric time series? | **Yes** | No |
| Need to explain the reasoning? | No | **Yes** |
| Need to ask follow-up questions? | No | **Yes** |
| Need real-time processing? | **Yes** | Maybe |
| Knowledge changes frequently? | No | **Yes** |

```python
# Performance comparison
print("\n=== Performance Statistics ===\n")

# Classification stats
if classification_results.height > 0:
    print("Classification Approach:")
    print(f"  Headlines processed: {classification_results.height}")
    print(f"  Unique categories: {classification_results['category'].n_unique()}")
    avg_confidence = classification_results["confidence"].mean()
    if avg_confidence is not None:
        print(f"  Average confidence: {avg_confidence:.2%}")

# RAG contract status
print("\nRAG Contract:")
print(f"  Questions specified: {rag_contract.height}")
print("  Answers generated: 0")
print("  Latency measured: No")
```

Two of the three ESG categories are present in the sample only because the
sample was stratified to include them.

## 4. The screen, in two pictures

The sample's own category counts are equal by construction, so charting them
would draw the stratification rather than anything about the data. The left
panel shows the pool those categories were drawn from, on a log axis because
the counts span three orders of magnitude. The right panel is FinBERT's
sentiment over the stratified sample.

```python
category_counts = classification_results.group_by("category").len().sort("len", descending=True)
sentiment_counts = classification_results.group_by("sentiment").len().sort("len", descending=True)

fig = make_subplots(
    rows=1,
    cols=2,
    subplot_titles=("Selected pool by category", "Sentiment in the stratified sample"),
    horizontal_spacing=0.16,
)
fig.add_trace(
    go.Bar(
        x=pool_counts["category"],
        y=pool_counts["len"],
        marker_color=COLORS["blue"],
        text=pool_counts["len"],
        textposition="outside",
    ),
    row=1,
    col=1,
)
fig.add_trace(
    go.Bar(
        x=sentiment_counts["sentiment"],
        y=sentiment_counts["len"],
        marker_color=COLORS["amber"],
        text=sentiment_counts["len"],
        textposition="outside",
    ),
    row=1,
    col=2,
)
fig.update_layout(
    title="What the ESG keyword screen selects, and how FinBERT scores a sample",
    height=430,
    showlegend=False,
    margin=dict(t=90),
)
fig.update_yaxes(title_text="Headlines (count, log scale)", type="log", row=1, col=1)
fig.update_yaxes(title_text="Headlines (count)", row=1, col=2)
show_plotly_with_alt(
    fig,
    "Two bar panels. Left, the selected pool by ESG category on a logarithmic count axis: "
    "the environmental bar runs off the top of the others by more than an order of "
    "magnitude, with uncategorised, governance and social following far below it in that "
    "order. Right, sentiment over the stratified sample on a linear axis: neutral is the "
    "tallest bar, positive next, negative slightly below it.",
)
```

```python
print("\n=== ESG Analysis Comparison Summary ===")
print(f"Classification: {classification_results.height} headlines -> numeric scores")
print(f"RAG: {rag_contract.height} question contracts -> no generated answers")
print("\nBoundary: execute and evaluate a cited RAG producer before comparing performance")
```

```python
completion_record = {
    "selected_rows_sha256": selection_sha256,
    "headlines": classification_results.height,
    "rag_questions": rag_contract.height,
    "models": classification_results["model"].unique().sort().to_list(),
    "model_revision": FINBERT_REVISION,
    "device": "cuda" if INFERENCE_DEVICE == 0 else "cpu",
    "category_counts": category_counts.to_dicts(),
    "sentiment_counts": sentiment_counts.to_dicts(),
}
print(f"COMPLETION_RECORD={json.dumps(completion_record, sort_keys=True)}")
```

## Key takeaways

1. **Check what a keyword screen actually selected before naming it.**
   Section 2 applies the categoriser to the whole selected pool, and this
   screen returns environmental news by a margin that makes the label "ESG"
   misleading. Two causes are visible in the screen itself: its environmental
   terms are common single words while its social and governance terms are
   compound phrases a headline must contain intact, and its categoriser
   breaks ties towards environmental. Whether that generalises to keyword
   screening as such is a question this notebook does not answer - it would
   need other lists and independent labels. The check that catches it is one
   group-by.

2. **A sample cannot show what its pool does not contain.** A uniform draw
   of twenty from this pool would be overwhelmingly environmental and would
   more likely than not contain no social or governance headline at all. The
   stratified draw is what puts all three categories in front of the
   classifier, and stratifying is an intervention that has to be declared,
   because the resulting category counts are then a property of the sampling
   rather than of the news.

3. **Time the load separately from the inference.** At this batch size the
   one-off cost of loading the weights is a large share of the wall clock, so
   a throughput figure computed over both describes the batch size rather
   than the model.

4. **Classification produces portfolio-ready outputs**, which is the reason
   to keep it: a row per document with a label and a score feeds a factor
   model directly, and a narrative answer does not, however well cited.

5. **The RAG side of this notebook is a contract, not a result.** No answers
   were generated, so no latency, no quality and no comparison is reported
   for it. `05_10k_rag_assistant` implements the retrieval half;
   `04_ragas_evaluation` is where the citation and abstention checks live.

**Next**: [`07_institutional_holdings_graph`](07_institutional_holdings_graph.ipynb)
builds graph-structured features from 13F filings.

**Book reference**: Section 22.8, on choosing between RAG and fine-tuning.
![notebook output](figures/p1_1.png)

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: MIT

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.