مواد پر جائیں
لائبریری کی تمام دستاویزات

L1 لاجسٹک ریگریشن: سولور اور ٹالرنس کے انتخاب کے اثرات

کوڈ Machine Learning for Trading

خلاصہ

یہ ترتیب سے متعلق نوٹ تین جماعتی لیبلز پر L1 لاجسٹک ریگریشن کے لیے ساگا سولور اور چھوٹی کنورجنس ٹالرنس کے انتخاب کی وضاحت کرتی ہے۔ یہ ناسڈیک مائیکرو اسٹرکچر پینل پر رن ٹائم پیمائشوں اور آؤٹ آف سیمپل لاگ لاس و درستگی کے رپورٹ شدہ نتائج کا حوالہ دے کر ساگا کا لِب لِنیر سے موازنہ کرتی ہے۔ ناپے گئے رنز میں رفتار اور پیش گوئی کے ان پیمانوں، دونوں پر ساگا بہتر تھا، جبکہ مصنفین کے مطابق پورے پینل کا رن ٹائم غیر یقینی ہے۔

ٹالرنس کا انتخاب اس بات کو بھی متاثر کرتا ہے کہ چھوٹے کوایفیشنٹس بالکل صفر بنتے ہیں یا نہیں۔ چونکہ L1 سویپ کا مقصد بالکل صفر کوایفیشنٹس حاصل کرنا ہے، اس لیے نوٹ بالکل صفر کوایفیشنٹس کی تعداد کا موازنہ محض صفر کے قریب کوایفیشنٹس کی تعداد سے کرتی ہے۔ زیادہ ڈھیلی ٹالرنس پر کچھ کوایفیشنٹس نہایت چھوٹے تھے مگر صفر نہیں؛ کم ٹالرنس والی ترتیب نے ان گنتیوں کو ہم آہنگ کیا اور ناپی گئی ترتیبات میں زیادہ کم گنجان حل پیدا کیے۔ اس مضبوط جرمانے والے معاملے میں ترتیب چھوٹی ٹالرنس استعمال کرتی ہے۔

یہ مشاہدات ایک خاص ڈیٹا سیٹ اور سولور کے موازنے سے ہیں، عمومی ضمانت نہیں۔ مکمل حجم کی کمپیوٹنگ لاگت واضح طور پر ثابت نہیں، اور بڑے پینل کا تخمینہ محدود وقت پیمائی سے اخذ کیا گیا ہے۔ اس لیے ہدفی ماحول میں رن ٹائم اور تنکی پن کے رویے کی جانچ ہونی چاہیے۔

اہم خیالات

  • حوالہ شدہ پینل پر ساگا سولور کثیر جماعتی L1 لاجسٹک ریگریشن کے لیے منتخب کیا گیا اور لِب لِنیر سے خاصا تیز ناپا گیا۔
  • اس موازنے میں رپورٹ کردہ ساگا رنز کے آؤٹ آف سیمپل لاگ لاس اور درستگی بھی بہتر تھے۔
  • زیادہ سخت ٹالرنس عین صفر اور محض بہت چھوٹے کوایفیشنٹس میں فرق واضح کر سکتی ہے۔
  • پورے پینل کا رن ٹائم غیر یقینی ہے، کیونکہ اس کا تخمینہ محدود اسکیلنگ شواہد پر منحصر ہے۔
  • سولور کی کارکردگی اور تنکی پن کی اصل ڈیٹا سیٹ اور ماحول میں تصدیق کریں۔

ٹیگز

مکمل متن
# logistic_l1_C0.001.yaml


```yaml
# L1 logistic regression. The solver is `saga` rather than `liblinear`, and the tolerance
# is set explicitly rather than left at scikit-learn's 1e-4 default.
#
# Two reasons, and the first one is not optional. These labels are three-class (-1, 0, 1),
# and scikit-learn 1.8 makes multiclass `liblinear` a hard error; #740 already moved our
# floor to 1.7. `OneVsRestClassifier(liblinear)` would reproduce the current objective
# exactly - liblinear multiclass IS one-vs-rest - and would keep the problem below.
#
# The second is that `liblinear` does not finish. It is single-threaded coordinate descent
# and scales about N^1.4 here. Measured on nasdaq100_microstructure's `fwd_dir_15m` panel:
#
#     rows      liblinear 1000/1e-4     saga 200/1e-2
#     400,000     144.4s  converged      11.0s  converged
#   1,200,000     716.9s  converged      44.8s  converged
#
# which extrapolates to roughly eight hours per configuration at the full 16.9M rows against
# about twenty minutes. That is not a projection: `06_linear` ran 7h23m at 100% of one core
# on 2026-09-05 and was killed with two of thirteen configurations still unfinished, both of
# them these L1 ones.
#
# saga is also better out of sample at every C measured here: log loss 1.0273-1.0276 against
# liblinear's 1.0293-1.0294, and accuracy 0.415-0.420 against 0.404-0.410.
#
# `tol: 0.001` here rather than the 0.01 the weakly-penalised configurations use, because
# this is where the penalty binds and exact sparsity is the point of the sweep. At 1e-2 saga
# leaves coefficients stranded NEAR zero instead of AT zero, which `coef_ != 0` then counts
# as live. Measured on the same panel, exact zeros against coefficients below 1e-8, out of
# 198:
#
#     C        liblinear 1e-4     saga 1e-2        saga 1e-3
#     0.001      128 / 128         148 / 149        157 / 157
#     0.01        40 /  40          24 /  52         87 /  87
#     0.1          8 /   8           3 /   4         24 /  26
#
# The 24-against-52 at C=0.01 is the defect: twenty-eight coefficients below 1e-8 that are
# not zero. At 1e-3 the two counts agree and saga is *more* sparse than liblinear at every C
# here, so the tighter tolerance is not a concession - it is what makes the L1 solution an
# L1 solution.
#
# Cost at 1.2M rows: 226s, 352s and 287s for C=0.001, 0.01 and 0.1 against liblinear's 30s,
# 202s and 499s. **The full-panel cost of this arm is not established** - the 16.9M-row
# extrapolation is uncertain because it rests on a single scaling estimate taken from the
# tol=1e-2 timings. Watch it on the first run rather than assuming it is small.
model_class: LogisticRegression
params:
  C: 0.001
  max_iter: 200
  penalty: l1
  solver: saga
  tol: 0.001

```

ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: MIT

یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔