跳至正文
返回文库全部文档

决策导向学习中的雅可比秩约束

文章 arXiv papers · 作者: Aojie Yuan et al.

总结

本研究分析预测器的几何结构如何限制决策导向学习;在这种学习方式中,模型训练与下游决策目标相联系。研究使用稀疏指数跟踪,区分优化器所用的协方差信息与预测器可用的参数更新方向。秩为一的预测器雅可比矩阵会使非零的逐样本梯度共线,而谱界则描述近似共线性。论文还刻画了批量更新子空间,并给出反例,说明局部秩属性本身并不能决定共同极小值或共线的批量更新。

实验检验这些几何约束是否会影响决策质量。在报告的股票配置中,决策导向学习相较于均方误差的提升很小;其他实验发现,最短路径和背包任务的遗憾值降幅较大,但修正后的结果仅在背包任务中仍然成立。模型容量、坐标缩放和训练设置都会影响优化结果。在所测试的架构中,金融前向目标对照和匹配的神经网络比较均未显示决策导向学习整体占优。这些发现仅适用于所考察的模型和任务:雅可比结构说明可用的学习方向,但要证明实际价值,还需要评估留出数据上的决策质量。

核心观点

  • 预测器的雅可比矩阵描述决策导向学习可用的参数更新方向。
  • 秩为一的雅可比矩阵会使非零的逐样本梯度共线,但批量更新不一定具有这一性质。
  • 局部秩约束本身并不意味着存在共同极小值。
  • 实验收益因任务而异;在所列比较中,修正后的证据仅在背包任务中仍然成立。
  • 坐标缩放会改变优化行为,因此评估仍需关注留出数据上的决策质量。

标签

全文
# Jacobian Rank Collapse in Decision-Focused Learning


# Jacobian Rank Collapse in Decision-Focused Learning









Decision-focused learning (DFL) trains predictors through downstream objectives, but a different loss need not provide an independent parameter-update direction. We characterize this restriction through the predictor Jacobian, using sparse index tracking to distinguish the covariance entries read by the optimizer from the parameter directions available to learning. Rank-one Jacobians make nonzero per-example gradients collinear; a conditional spectral bound describes near-collinearity. A batch-subspace characterization and counterexamples show why these local statements imply neither common minimizers nor collinear batch updates. Experiments examine when geometry translates into decision quality. Across 38 one-parameter equity configurations, DFL gains over MSE remain below 1.8%; a 385-parameter conditional predictor also has pointwise rank one. In validation-tuned shortest-path and knapsack experiments, full-capacity SPO+ reduces mean regret by 11.6% and 10.6%, respectively; only knapsack survives correction across eight comparisons. The capacity contrast persists on fresh datasets across batch orders and training budgets. Holding expressivity fixed, invertible coordinate scaling lowers spectral effective rank and ordinary SGD gains; compensating for the scaling restores the original trajectories. Financial forward-target controls separate forecast accuracy from decision quality; a matched neural comparison finds no aggregate DFL advantage in the tested architecture. These findings distinguish local rank restrictions, coordinate-dependent optimization and predictive accuracy. Predictor geometry helps explain available learning directions, while held-out decision quality remains the test of practical benefit.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。