Reproducing CIFAR-10 ResNet Results with Shortcut Variants
Summary
This article describes a reproduction of residual network experiments on CIFAR-10, focusing on two changes from an earlier run: image augmentation and shortcut design. It pads training images by four pixels on each edge and randomly crops them back to the original size, while leaving test images unchanged. It notes that the training mean must be regenerated for the padded data.
For changes in feature-map channels, it explains a zero-padding shortcut (Option A) that uses average pooling and adds zero channels, and a projection shortcut (Option B) that uses a 1×1 convolution and adds learned parameters. The article provides implementation details for adding a custom channel-padding layer to Caffe, then reports training settings and says Option A broadly matches the paper while Option B produces the strongest reported results. The evidence is limited to the author’s reproduction; plots and detailed result tables are absent from the supplied text, and the source itself flags one result for separate consideration.
Key ideas
- Training augmentation pads each CIFAR-10 image before randomly cropping it back to its original dimensions.
- Zero-padding shortcuts expand channels without adding learned parameters, while projection shortcuts use a 1×1 convolution.
- A custom Caffe layer is needed to implement channel padding in the setup described.
- The reported experiment finds that the zero-padding results broadly match the paper and projection performs better in this run.
- The supplied text does not include detailed numerical results for independently assessing the comparison.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.