Towards a Very Fast Feedforward Multilayer Neural Networks Training Algorithm

^**This paper presents a novel fast algorithm for feedforward neural networks training. It is based on the Recursive Least Squares (RLS) method commonly used for designing adaptive filters. Besides, it utilizes two techniques of linear algebra, namely the orthogonal transformation method, called the Givens Rotations (GR), and the QR decomposition, creating the GQR (symbolically we write GR + QR = GQR) procedure for solving the normal equations in the weight update process. In this paper, a novel approach to the GQR algorithm is presented. The main idea revolves around reducing the computational cost of a single rotation by eliminating the square root calculation and reducing the number of multiplications. The proposed modification is based on the scaled version of the Givens rotations, denoted as SGQR. This modification is expected to bring a significant training time reduction comparing to the classic GQR algorithm. The paper begins with the introduction and the classic Givens rotation description. Then, the scaled rotation and its usage in the QR decomposition is discussed. The main section of the article presents the neural network training algorithm which utilizes scaled Givens rotations and QR decomposition in the weight update process. Next, the experiment results of the proposed algorithm are presented and discussed. The experiment utilizes several benchmarks combined with neural networks of various topologies. It is shown that the proposed algorithm outperforms several other commonly used methods, including well known Adam optimizer.

eISSN:: 2449-6499
Lingua:: Inglese

Frequenza di pubblicazione:: 4 volte all'anno
Argomenti della rivista:: Computer Sciences, Databases and Data Mining, Artificial Intelligence

Feed RSS della rivista

Towards a Very Fast Feedforward Multilayer Neural Networks Training Algorithm

Pubblicato online: 23 lug 2022

Pagine: 181 - 195

Ricevuto: 03 gen 2022

Accettato: 06 giu 2022

DOI: https://doi.org/10.2478/jaiscr-2022-0012

Parole chiaveneural network training algorithm, QR decomposition, scaled Givens rotations, approximation, classification

© 2022 Jarosław Bilski et al., published by Sciendo

This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 3.0 License.

Parole chiave
neural network training algorithm, QR decomposition, scaled Givens rotations, approximation, classification