Skip to content
All library documents

ICML 2022 Methods for Faster Neural Networks, Forecasting, and Decision Trees

Article BigQuant

Summary

This overview summarizes six ICML 2022 papers selected by G-Research staff. The topics include convex formulations for training two-layer ReLU networks, linear-time attention, branch-and-bound optimization of shallow decision trees, structured matrices for efficient neural-network training, attention sharing for time-series domain adaptation, and hierarchical shrinkage for tree models. The article describes the central idea behind each method and, in several cases, the reported advantages over common alternatives.

The evidence is reported secondhand and without detailed experimental settings, datasets, or quantitative results, so the claims cannot be independently assessed from this summary alone. The described methods also have practical limits: globally optimal tree search becomes difficult at greater depths, and architectural or optimization benefits may depend on hardware and data. For quantitative researchers, the most directly relevant topics are forecasting across domains and regularization of tree-based models; the other papers provide broader machine-learning techniques that may support research workflows rather than trading strategies directly.

Key ideas

  • A convex reformulation can make training certain two-layer ReLU networks less dependent on nonconvex optimization choices.
  • FLASH changes attention computation to reduce sequence-length-related time and memory costs.
  • Quant-BnB searches for globally optimal shallow decision trees, but deeper trees quickly become difficult to solve.
  • Monarch structured matrices offer alternative parameterizations intended to speed neural-network training.
  • A shared attention module and adversarial domain discriminator are proposed for forecasting across related time series.
  • Hierarchical shrinkage regularizes tree models at multiple levels and is reported to improve generalization and explanation quality.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.