コンテンツへスキップ
ライブラリの全資料

多クラスロジスティック回帰のSAGAと許容誤差の選択

コード Machine Learning for Trading

サマリー

この設定ノートでは、L1正則化を使う3クラスのロジスティック回帰で、scikit-learnのliblinearソルバーではなくSAGAを使う理由を説明します。新しいscikit-learnバージョンでの多クラス対応に触れ、Nasdaq-100のマイクロストラクチャー実験ではSAGAの処理が大幅に速かったと報告します。報告されたテスト外の対数損失と正解率も、検証した設定ではSAGAの方が良好でした。

このノートでは、許容誤差を0.001に設定することで、L1のスパース性が係数の厳密なゼロとして反映される点に注目します。許容誤差が大きいと、一部の係数がゼロではなくゼロに近い値のまま残り、選択された特徴量の数が歪みます。許容誤差をより小さくした設定では、報告された比較においてliblinear以上のスパース性が得られました。これらは特定のパネルに対する実装とベンチマークの結果であり、一般的な保証ではありません。全データでの実行時間は明示的に不明で、所要時間の予測は不確かなスケーリング推定に依存します。

主なアイデア

  • 新しいライブラリバージョンでliblinearの多クラス適合が失敗する場合に、SAGAは3クラスのロジスティック回帰設定に対応します。
  • Nasdaq-100のマイクロストラクチャーパネルを使った実験では、報告されたテスト外指標においてSAGAがより速く、やや良好でした。
  • 許容誤差を厳しくすると、真にゼロのL1係数と小さな非ゼロ値を区別しやすくなります。
  • 報告された実行時間から、データセット全体での所要時間は判断できません。

タグ

全文
# logistic_l1_C0.1.yaml


```yaml
# L1 logistic regression. The solver is `saga` rather than `liblinear`, and the tolerance
# is set explicitly rather than left at scikit-learn's 1e-4 default.
#
# Two reasons, and the first one is not optional. These labels are three-class (-1, 0, 1),
# and scikit-learn 1.8 makes multiclass `liblinear` a hard error; #740 already moved our
# floor to 1.7. `OneVsRestClassifier(liblinear)` would reproduce the current objective
# exactly - liblinear multiclass IS one-vs-rest - and would keep the problem below.
#
# The second is that `liblinear` does not finish. It is single-threaded coordinate descent
# and scales about N^1.4 here. Measured on nasdaq100_microstructure's `fwd_dir_15m` panel:
#
#     rows      liblinear 1000/1e-4     saga 200/1e-2
#     400,000     144.4s  converged      11.0s  converged
#   1,200,000     716.9s  converged      44.8s  converged
#
# which extrapolates to roughly eight hours per configuration at the full 16.9M rows against
# about twenty minutes. That is not a projection: `06_linear` ran 7h23m at 100% of one core
# on 2026-09-05 and was killed with two of thirteen configurations still unfinished, both of
# them these L1 ones.
#
# saga is also better out of sample at every C measured here: log loss 1.0273-1.0276 against
# liblinear's 1.0293-1.0294, and accuracy 0.415-0.420 against 0.404-0.410.
#
# `tol: 0.001` here rather than the 0.01 the weakly-penalised configurations use, because
# this is where the penalty binds and exact sparsity is the point of the sweep. At 1e-2 saga
# leaves coefficients stranded NEAR zero instead of AT zero, which `coef_ != 0` then counts
# as live. Measured on the same panel, exact zeros against coefficients below 1e-8, out of
# 198:
#
#     C        liblinear 1e-4     saga 1e-2        saga 1e-3
#     0.001      128 / 128         148 / 149        157 / 157
#     0.01        40 /  40          24 /  52         87 /  87
#     0.1          8 /   8           3 /   4         24 /  26
#
# The 24-against-52 at C=0.01 is the defect: twenty-eight coefficients below 1e-8 that are
# not zero. At 1e-3 the two counts agree and saga is *more* sparse than liblinear at every C
# here, so the tighter tolerance is not a concession - it is what makes the L1 solution an
# L1 solution.
#
# Cost at 1.2M rows: 226s, 352s and 287s for C=0.001, 0.01 and 0.1 against liblinear's 30s,
# 202s and 499s. **The full-panel cost of this arm is not established** - the 16.9M-row
# extrapolation is uncertain because it rests on a single scaling estimate taken from the
# tol=1e-2 timings. Watch it on the first run rather than assuming it is small.
model_class: LogisticRegression
params:
  C: 0.1
  max_iter: 200
  penalty: l1
  solver: saga
  tol: 0.001

```

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。