Skip to content
All library documents

Deep Neural Network Pretraining and Model Testing with darch

Article MQL5 articles

Summary

The article explains how to build, train, and test deep neural networks with the darch package in R. It contrasts unsupervised layer-wise pretraining using restricted Boltzmann machines or autoencoders with training a multilayer network from initialized weights. It also surveys weight initialization, activation, optimization, regularization, and stopping settings, and describes how pretraining and supervised fine-tuning can use separate datasets or be run together.

The experiments compare models trained with and without pretraining on prepared classification data. The reported classification error with pretraining is around 30% with a variation of about four percentage points across the sets. The author observes signs of overfitting in the model without pretraining and chooses pretraining for later experiments. Results are described as only modestly better than a basic model, and the article attributes weak performance partly to default or near-default settings. It does not establish that pretraining will improve other datasets or trading outcomes; hyperparameter optimization and comparisons with other libraries are left for follow-up work.

Key ideas

  • Unsupervised pretraining can initialize hidden layers before supervised fine-tuning.
  • The darch package supports several initialization, activation, training, and stabilization options.
  • Pretraining and fine-tuning can be separated to reuse one pretrained model across labeled datasets.
  • In the reported experiment, training without pretraining showed signs of overfitting.
  • The tested models showed limited gains, and their settings were not extensively optimized.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.