AIセキュリティポータル K Program
VulLibGen: Generating Names of Vulnerability-Affected Packages via a Large Language Model
Share
Abstract
Security practitioners maintain vulnerability reports (e.g., GitHub Advisory) to help developers mitigate security risks. An important task for these databases is automatically extracting structured information mentioned in the report, e.g., the affected software packages, to accelerate the defense of the vulnerability ecosystem. However, it is challenging for existing work on affected package identification to achieve a high accuracy. One reason is that all existing work focuses on relatively smaller models, thus they cannot harness the knowledge and semantic capabilities of large language models. To address this limitation, we propose VulLibGen, the first method to use LLM for affected package identification. In contrast to existing work, VulLibGen proposes the novel idea to directly generate the affected package. To improve the accuracy, VulLibGen employs supervised fine-tuning (SFT), retrieval augmented generation (RAG) and a local search algorithm. The local search algorithm is a novel postprocessing algorithm we introduce for reducing the hallucination of the generated packages. Our evaluation results show that VulLibGen has an average accuracy of 0.806 for identifying vulnerable packages in the four most popular ecosystems in GitHub Advisory (Java, JS, Python, Go) while the best average accuracy in previous work is 0.721. Additionally, VulLibGen has high value to security practice: we submitted 60 <vulnerability, affected package> pairs to GitHub Advisory (covers four ecosystems). 34 of them have been accepted and merged and 20 are pending approval. Our code and dataset can be found in the attachments.
Cleaning the nvd: Comprehensive quality assessment, improvements, and analyses
Afsah Anwar, Ahmed Abusnaina, Songqing Chen, Frank Li, David Mohaisen
Published: 2021
Reliable third-party library detection in Android and its security applications
Michael Backes, Sven Bugiel, Erik Derr
Published: 2016
Recent advances in retrieval-augmented text generation
Deng Cai, Yan Wang, Lemao Liu, Shuming Shi
Published: 2022
CodeT: Code Generation with Generated Tests
Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, Weizhu Chen
Published: 2023
Automated identification of libraries from vulnerability data
Y. Chen, A.E. Santosa, A. Sharma, D. Lo
Published: 2020
Emerging trends: A gentle introduction to fine-tuning
Kenneth Ward Church, Zeyu Chen, Yanjun Ma
Published: 2021
Towards the detection of inconsistencies in public security vulnerability reports
Ying Dong, Wenbo Guo, Yueqi Chen, Xinyu Xing, Yuqing Zhang, Gang Wang
Published: 2019
Github-advisory
Published: 2024
Github-advisory-review
Published: 2024
Checking app behavior against app descriptions
Alessandra Gorla, Ilaria Tavecchia, Florian Gross, Andreas Zeller
Published: 2014
Generalized zero-shot extreme multi-label learning
Nilesh Gupta, Sakina Bohra, Yashoteja Prabhu, Saurabh Purohit, Manik Varma
Published: 2021
How well does chatgpt understand go?
Jonathan Hall
Published: 2023
Automated identification of libraries from vulnerability data: Can we do better?
Stefanus A Haryono, Hong Jin Kang, Abhishek Sharma, Asankhaya Sharma, Andrew Santosa, Ang Ming Yi, David Lo
Published: 2022
HierarchyNet: Learning to Summarize Source Code with Heterogeneous Representations
Minh Huynh Nguyen, Nghi DQ Bui, Truong Son Hy, Long Tran-Thanh, Tien N Nguyen
Published: 2022
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, Pascale Fung
Published: 2023
Inferfix: End-to-end program repair with llms
M. Jin, S. Shahriar, M. Tufano, X. Shi, S. Lu, N. Sundaresan, A. Svyatkovskiy
Published: 2023
Vulcan: Automatic extraction and analysis of cyber threat intelligence from unstructured text
Hyeonseong Jo, Yongjae Lee, Seungwon Shin
Published: 2022
Large language models are few-shot testers: Exploring llm-based general bug reproduction
Sungmin Kang, Juyeon Yoon, Shin Yoo
Published: 2023
Look-ahead bias: What it means, how it works
Will Kenton
Published: 2024
Bonsai: diverse and shallow trees for extreme multi-label classification
Sujay Khandagale, Han Xiao, Rohit Babbar
Published: 2020
T test as a parametric statistic
Tae Kyun Kim
Published: 2015
OVANA: An approach to analyze and improve the information quality of vulnerability databases
Philipp Kuehn, Markus Bayer, Marc Wendelborn, Christian Reuter
Published: 2021
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, Douwe Kiela
Published: 2020.5.23
Retrieval-augmented generation for code summarization via hybrid GNN
Shangqing Liu, Yu Chen, Xiaofei Xie, Jingkai Siow, Yang Liu
Published: 2021
Chronos: Time-aware zero-shot identification of libraries from vulnerability reports
Yunbo Lyu, Thanh Le-Cong, Hong Jin Kang, Ratnadira Widyasari, Zhipeng Zhao, Xuan-Bach D Le, Ming Li, David Lo
Published: 2023
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegr-eff, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang
Published: 2023
Generation-augmented retrieval for open-domain question answering
Yuning Mao, Pengcheng He, Xiaodong Liu, Yelong Shen, Jianfeng Gao, Jiawei Han, Weizhu Chen
Published: 2020
Security threats classification in blockchains
Jamal Hayat Mosakheil
Published: 2018
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, S. Johnston, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, C. Olah
Published: 2022
Fast lexically constrained decoding with dynamic beam allocation for neural machine translation
Matt Post, David Vilar
Published: 2018
Transfer learning for sentiment analysis using BERT based supervised fine-tuning
Nusrat Jahan Prottasha, Abdullah As Sami, Md Kowsher, Saydul Akbar Murad, Anupam Kumar Bairagi, Mehedi Masud, Mohammed Baz
Published: 2022
Leveraging automated unit tests for unsupervised code translation
Baptiste Roziere, Jie Zhang, Francois Charton, Mark Harman, Gabriel Synnaeve, Guillaume Lample
Published: 2022
Analysis of variance (ANOVA)
Lars St, Svante Wold
Published: 1989
LibDB: An Effective and Efficient Framework for Detecting Third-Party Libraries in Binaries
Wei Tang, Yanlin Wang, Hongyu Zhang, Shi Han, Ping Luo, Dongmei Zhang
Published: 2022
An empirical study of usages, updates and risks of third-party libraries in java projects
Ying Wang, Bihuan Chen, Kaifeng Huang, Bowen Shi, Congying Xu, Xin Peng, Yijian Wu, Yang Liu
Published: 2020
Understanding the threats of upstream vulnerabilities to downstream projects in the maven ecosystem
Yulun Wu, Zeliang Yu, Ming Wen, Qiang Li, Deqing Zou, Hai Jin
Published: 2023
A no-regret generalization of hierarchical softmax to extreme multi-label classification
Marek Wydmuch, Kalina Jasinska, Mikhail Kuznetsov, Róbert Busa-Fekete, Krzysztof Dembczynski
Published: 2018
Few-sample named entity recognition for security vulnerability reports by fine-tuning pre-trained language models
Guanqun Yang, Shay Dineen, Zhipeng Lin, Xueqing Liu
Published: 2021
Generate rather than Retrieve: Large Language Models are Strong Context Generators
Wenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal, Chenguang Zhu, Michael Zeng, Meng Jiang
Published: 2022.9.21
Atvhunter: Reliable version detection of third-party libraries for vulnerability identification in android applications
Xian Zhan, Lingling Fan, Sen Chen, Feng We, Tianming Liu, Xiapu Luo, Yang Liu
Published: 2021
Libid: reliable identification of obfuscated third-party android libraries
Jiexin Zhang, Alastair R Beresford, Stephan A Kollmann
Published: 2019
Self-Edit: Fault-Aware Code Editor for Code Generation
Kechi Zhang, Zhuo Li, Jia Li, Ge Li, Zhi Jin
Published: 2023
Detecting third-party libraries in Android applications with high precision and recall
Yuan Zhang, Jiarun Dai, Xiaohan Zhang, Sirong Huang, Zhemin Yang, Min Yang, Hao Chen
Published: 2018
Share