Pular para o conteúdo
Todos os documentos da biblioteca

Escolha de SAGA e tolerância para regressão logística esparsa

Código Machine Learning for Trading

Resumo

Esta nota de configuração explica por que uma regressão logística de três classes penalizada por L1 usa SAGA em vez do solver liblinear do scikit-learn. Cita compatibilidade multiclasses em versões mais recentes do scikit-learn e relata que SAGA foi muito mais rápido nos experimentos medidos de microestrutura do Nasdaq-100. A perda logarítmica fora da amostra e a acurácia relatadas também favoreceram SAGA nas configurações testadas.

A nota se concentra em definir a tolerância como 0.001 para que a esparsidade de L1 se reflita em coeficientes exatamente iguais a zero. Com uma tolerância mais frouxa, alguns coeficientes permaneceram apenas próximos de zero, distorcendo a contagem de variáveis selecionadas. Nas comparações relatadas, a configuração mais rigorosa produziu contagens de esparsidade pelo menos tão altas quanto as do liblinear. São constatações de implementação e benchmark para um painel específico, não garantias gerais: o tempo de execução com todos os dados é explicitamente desconhecido, e sua projeção depende de uma estimativa de escala incerta.

Ideias principais

  • SAGA dá suporte à configuração de regressão logística de três classes em que o ajuste multiclasses com liblinear pode falhar em versões mais recentes da biblioteca.
  • Experimentos medidos em um painel de microestrutura do Nasdaq-100 constataram que SAGA foi mais rápido e teve resultados um pouco melhores nas métricas fora da amostra relatadas.
  • Uma tolerância de convergência mais rigorosa ajuda a distinguir coeficientes L1 realmente iguais a zero de valores pequenos, mas diferentes de zero.
  • Os tempos relatados não estabelecem o tempo de execução no tamanho completo do conjunto de dados.

Tags

Texto completo
# logistic_l1_C0.1.yaml


```yaml
# L1 logistic regression. The solver is `saga` rather than `liblinear`, and the tolerance
# is set explicitly rather than left at scikit-learn's 1e-4 default.
#
# Two reasons, and the first one is not optional. These labels are three-class (-1, 0, 1),
# and scikit-learn 1.8 makes multiclass `liblinear` a hard error; #740 already moved our
# floor to 1.7. `OneVsRestClassifier(liblinear)` would reproduce the current objective
# exactly - liblinear multiclass IS one-vs-rest - and would keep the problem below.
#
# The second is that `liblinear` does not finish. It is single-threaded coordinate descent
# and scales about N^1.4 here. Measured on nasdaq100_microstructure's `fwd_dir_15m` panel:
#
#     rows      liblinear 1000/1e-4     saga 200/1e-2
#     400,000     144.4s  converged      11.0s  converged
#   1,200,000     716.9s  converged      44.8s  converged
#
# which extrapolates to roughly eight hours per configuration at the full 16.9M rows against
# about twenty minutes. That is not a projection: `06_linear` ran 7h23m at 100% of one core
# on 2026-09-05 and was killed with two of thirteen configurations still unfinished, both of
# them these L1 ones.
#
# saga is also better out of sample at every C measured here: log loss 1.0273-1.0276 against
# liblinear's 1.0293-1.0294, and accuracy 0.415-0.420 against 0.404-0.410.
#
# `tol: 0.001` here rather than the 0.01 the weakly-penalised configurations use, because
# this is where the penalty binds and exact sparsity is the point of the sweep. At 1e-2 saga
# leaves coefficients stranded NEAR zero instead of AT zero, which `coef_ != 0` then counts
# as live. Measured on the same panel, exact zeros against coefficients below 1e-8, out of
# 198:
#
#     C        liblinear 1e-4     saga 1e-2        saga 1e-3
#     0.001      128 / 128         148 / 149        157 / 157
#     0.01        40 /  40          24 /  52         87 /  87
#     0.1          8 /   8           3 /   4         24 /  26
#
# The 24-against-52 at C=0.01 is the defect: twenty-eight coefficients below 1e-8 that are
# not zero. At 1e-3 the two counts agree and saga is *more* sparse than liblinear at every C
# here, so the tighter tolerance is not a concession - it is what makes the L1 solution an
# L1 solution.
#
# Cost at 1.2M rows: 226s, 352s and 287s for C=0.001, 0.01 and 0.1 against liblinear's 30s,
# 202s and 499s. **The full-panel cost of this arm is not established** - the 16.9M-row
# extrapolation is uncertain because it rests on a single scaling estimate taken from the
# tol=1e-2 timings. Watch it on the first run rather than assuming it is small.
model_class: LogisticRegression
params:
  C: 0.1
  max_iter: 200
  penalty: l1
  solver: saga
  tol: 0.001

```

Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT

Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.