رفتن به محتوا
همه اسناد کتابخانه

موازنه‌های حل‌گر رگرسیون لجستیک L1 و تلورانس

کد یادگیری ماشین برای معامله‌گری

خلاصه

این یادداشت پیکربندی، انتخاب حل‌گر saga و تلورانس همگرایی سخت‌گیرانه‌تر را برای رگرسیون لجستیک L1 با برچسب‌های سه‌کلاسه توضیح می‌دهد. saga را با liblinear مقایسه می‌کند و به سنجش زمان اجرا روی پنل ریزساختار نزدک و مقادیر گزارش‌شده زیان لگاریتمی و دقت خارج از نمونه اشاره می‌کند. اجراهای اندازه‌گیری‌شده از نظر سرعت و این سنجه‌های پیش‌بینی saga را ترجیح دادند، درحالی‌که نویسندگان می‌گویند زمان اجرای کل پنل همچنان نامطمئن است.

انتخاب تلورانس همچنین بر این اثر می‌گذارد که آیا ضرایب کوچک دقیقاً صفر می‌شوند یا نه. چون تنکی دقیق هدف پیمایش L1 است، یادداشت تعداد صفرهای دقیق را با تعداد ضرایبی که فقط نزدیک صفرند مقایسه می‌کند. در تلورانس بازتر، برخی ضرایب بسیار کوچک اما ناصفر بودند؛ تنظیم سخت‌گیرانه‌تر این شمارش‌ها را همسو کرد و در تنظیمات اندازه‌گیری‌شده جواب‌های تنک‌تری به دست داد. این پیکربندی برای این حالت با جریمه قوی‌تر، تلورانس کوچک‌تر را به‌کار می‌برد.

این مشاهدات از مجموعه‌داده و مقایسه حل‌گر مشخصی می‌آیند و تضمینی کلی نیستند. هزینه محاسباتی در اندازه کامل صراحتاً مشخص نشده و برآورد پنل بزرگ از شواهد محدود زمان‌سنجی برون‌یابی شده است. بنابراین باید رفتار زمان اجرا و تنکی را در محیط هدف بررسی کرد.

ایده‌های کلیدی

  • حل‌گر saga برای رگرسیون لجستیک چندکلاسه L1 انتخاب شده و در پنل مورد اشاره به‌طور قابل‌توجهی سریع‌تر از liblinear اندازه‌گیری شده است.
  • اجراهای گزارش‌شده saga در همان مقایسه، زیان لگاریتمی و دقت خارج از نمونه بهتری نیز داشتند.
  • تلورانس سخت‌گیرانه‌تر می‌تواند صفرهای دقیق را از ضرایبی که فقط بسیار کوچک‌اند متمایز کند.
  • زمان اجرای کل پنل نامطمئن است، زیرا برآورد آن به شواهد محدود مقیاس‌پذیری متکی است.
  • عملکرد حل‌گر و تنکی را باید روی مجموعه‌داده و محیط واقعی بررسی کرد.

برچسب‌ها

متن کامل
# logistic_l1_C0.001.yaml


```yaml
# L1 logistic regression. The solver is `saga` rather than `liblinear`, and the tolerance
# is set explicitly rather than left at scikit-learn's 1e-4 default.
#
# Two reasons, and the first one is not optional. These labels are three-class (-1, 0, 1),
# and scikit-learn 1.8 makes multiclass `liblinear` a hard error; #740 already moved our
# floor to 1.7. `OneVsRestClassifier(liblinear)` would reproduce the current objective
# exactly - liblinear multiclass IS one-vs-rest - and would keep the problem below.
#
# The second is that `liblinear` does not finish. It is single-threaded coordinate descent
# and scales about N^1.4 here. Measured on nasdaq100_microstructure's `fwd_dir_15m` panel:
#
#     rows      liblinear 1000/1e-4     saga 200/1e-2
#     400,000     144.4s  converged      11.0s  converged
#   1,200,000     716.9s  converged      44.8s  converged
#
# which extrapolates to roughly eight hours per configuration at the full 16.9M rows against
# about twenty minutes. That is not a projection: `06_linear` ran 7h23m at 100% of one core
# on 2026-09-05 and was killed with two of thirteen configurations still unfinished, both of
# them these L1 ones.
#
# saga is also better out of sample at every C measured here: log loss 1.0273-1.0276 against
# liblinear's 1.0293-1.0294, and accuracy 0.415-0.420 against 0.404-0.410.
#
# `tol: 0.001` here rather than the 0.01 the weakly-penalised configurations use, because
# this is where the penalty binds and exact sparsity is the point of the sweep. At 1e-2 saga
# leaves coefficients stranded NEAR zero instead of AT zero, which `coef_ != 0` then counts
# as live. Measured on the same panel, exact zeros against coefficients below 1e-8, out of
# 198:
#
#     C        liblinear 1e-4     saga 1e-2        saga 1e-3
#     0.001      128 / 128         148 / 149        157 / 157
#     0.01        40 /  40          24 /  52         87 /  87
#     0.1          8 /   8           3 /   4         24 /  26
#
# The 24-against-52 at C=0.01 is the defect: twenty-eight coefficients below 1e-8 that are
# not zero. At 1e-3 the two counts agree and saga is *more* sparse than liblinear at every C
# here, so the tighter tolerance is not a concession - it is what makes the L1 solution an
# L1 solution.
#
# Cost at 1.2M rows: 226s, 352s and 287s for C=0.001, 0.01 and 0.1 against liblinear's 30s,
# 202s and 499s. **The full-panel cost of this arm is not established** - the 16.9M-row
# extrapolation is uncertain because it rests on a single scaling estimate taken from the
# tol=1e-2 timings. Watch it on the first run rather than assuming it is small.
model_class: LogisticRegression
params:
  C: 0.001
  max_iter: 200
  penalty: l1
  solver: saga
  tol: 0.001

```

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: MIT

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.