Position-Encoding Convolutional Network to Solving Connected Text Captcha

Text-based CAPTCHA is a convenient and effective safety mechanism that has been widely deployed across websites. The efficient end-to-end models of scene text recognition consisting of CNN and attention-based RNN show limited performance in solving text-based CAPTCHAs. In contrast with the street view image and document, the character sequence in CAPTCHA is non-semantic. The RNN loses its ability to learn the semantic context and only implicitly encodes the relative position of extracted features. Meanwhile, the security features, which prevent characters from segmentation and recognition, extensively increase the complexity of CAPTCHAs. The performance of this model is sensitive to different CAPTCHA schemes. In this paper, we analyze the properties of the text-based CAPTCHA and accordingly consider solving it as a highly position-relative character sequence recognition task. We propose a network named PosConv to leverage the position information in the character sequence without RNN. PosConv uses a novel padding strategy and modified convolution, explicitly encoding the relative position into the local features of characters. This mechanism of PosConv makes the extracted features from CAPTCHAs more informative and robust. We validate PosConv on six text-based CAPTCHA schemes, and it achieves state-of-the-art or competitive recognition accuracy with significantly fewer parameters and faster convergence speed.

eISSN:: 2449-6499
Language:: English

Publication timeframe:: 4 times per year
Journal Subjects:: Computer Sciences, Databases and Data Mining, Artificial Intelligence

Journal RSS Feed

Position-Encoding Convolutional Network to Solving Connected Text Captcha

Published Online: Feb 23, 2022

Page range: 121 - 133

Received: Oct 06, 2021

Accepted: Oct 12, 2021

DOI: https://doi.org/10.2478/jaiscr-2022-0008

Keywords
deep neural network, position encoding CNN, text-based CAPTCHA recognition, character recognition

© 2022 Ke Qing et al., published by Sciendo

This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Position-Encoding Convolutional Network to Solving Connected Text Captcha

Published Online: Feb 23, 2022

Page range: 121 - 133

Received: Oct 06, 2021

Accepted: Oct 12, 2021

DOI: https://doi.org/10.2478/jaiscr-2022-0008

Keywordsdeep neural network, position encoding CNN, text-based CAPTCHA recognition, character recognition

© 2022 Ke Qing et al., published by Sciendo

This work is licensed under the Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Keywords
deep neural network, position encoding CNN, text-based CAPTCHA recognition, character recognition