首页 | 本学科首页   官方微博 | 高级检索  
     


Unsupervised software repositories mining and its application to code search
Authors:Gang Hu  Min Peng  Yihan Zhang  Qianqian Xie  Wang Gao  Mengting Yuan
Affiliation:1. Department of National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University, Wuhan, China;2. School of Computing, National University of Singapore, Singapore
Abstract:Software repositories are crucial resources for many software tasks, including code retrieval and annotation. Programming forums provide questions and answers (Q&A) from software developers, containing abundant code-description posts for exchanging knowledge about programming issues. However, most posts provide personal opinions of users that are often not adequately confirmed or outdated. Mining software repositories in such open and unrestricted forums is challenging. Since the posts can be arbitrary and noisy, it is difficult to get unified labels for supervised noise elimination. Different from existing mining approaches, this paper proposes Code-Description Mining Framework (CodeMF), an unsupervised framework to eliminate noisy posts and extract high quality software repositories from programming forums. CodeMF treats all social features of the posts as discrete-time signals for kernel principal component analysis and further performs wavelet transform feature fusion to find the delicate changes (noises in temporal signals). We conduct comprehensive experiments on StackOverflow. Experimental results demonstrate that CodeMF can effectively reduce running time and improve precision via mining high-quality software repositories for various programming languages, especially for the large-scale codebases. To further illustrate the effect of CodeMF applied in software tasks, we introduce it to improve the performance of query-expansion code search. Meanwhile, for SQL and C# programs, compared to the state-of-the-art query-expansion method QECK, the improvement of QECKCodeMF is 2% and 6% on Recall@10, and 4% and 14% on mean reciprocal rank, respectively.
Keywords:code search  feature fusion  software repositories  wavelet transformation
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号