From text to threats: A language model approach to software vulnerability detection

In the rapidly-evolving landscape of software development, the detection of vulnerabilities in source code has become of paramount importance. Our study introduces a novel knowledge distillation (KD) technique aimed at enhancing vulnerability detection in software codebases. Using benchmark datasets such as SARD, SeVC, Devign, and D2A, we assess the prowess of the KD method when applied to different classifiers, specifically GPT2, CodeBERT, and LSTM. The empirical results are revealed a marked improvement in the performance of these classifiers upon the implementation of the KD technique, particularly with the GPT-2 model demonstrating the most promising outcomes. This work underscores the potential of integrating transformer-based learning models, like GPT-2, with knowledge distillation for more efficient and accurate vulnerability detection.

eISSN:: 2956-7068
Język:: Angielski

Częstotliwość wydawania:: 2 razy w roku
Dziedziny czasopisma:: Computer Sciences, other, Engineering, Introductions and Overviews, Mathematics, General Mathematics, Physics

Kanał RSS czasopisma

From text to threats: A language model approach to software vulnerability detection

Article Category: Original Study

Data publikacji: 31 paź 2023

Zakres stron: 23 - 34

Otrzymano: 04 wrz 2023

Przyjęty: 26 paź 2023

DOI: https://doi.org/10.2478/ijmce-2024-0003

Słowa kluczoweKnowledge distillation, DistilVulBERT, software vulnerability detection, language models, codebase analysis

© 2024 Marwan Omar et al., published by Sciendo

This work is licensed under the Creative Commons Attribution 4.0 International License.

Słowa kluczowe
Knowledge distillation, DistilVulBERT, software vulnerability detection, language models, codebase analysis