302 lines
11 KiB
Markdown
302 lines
11 KiB
Markdown
file:: [COLA-D-23-00096_1699177983001_0.pdf](../assets/COLA-D-23-00096_1699177983001_0.pdf)
|
||
file-path:: ../assets/COLA-D-23-00096_1699177983001_0.pdf
|
||
|
||
- AI-based Clustering of Similar Issues in GitHub’s Repositories Article Type: Full Length Article Keywords: Similar Issues; GitHub; Machine Learning; Maintenance; Clustering Corresponding Author: Hamzeh Eyal Salman Mutah University Karak, JORDAN Corresponding Author Secondary Information: Corresponding Author's Institution: Mutah University
|
||
ls-type:: annotation
|
||
hl-page:: 1
|
||
hl-color:: green
|
||
id:: 654cf55c-0711-4113-8546-4963e635cb90
|
||
- I-based Clustering of Similar Issues in GitHub’s Repositories
|
||
ls-type:: annotation
|
||
hl-page:: 3
|
||
hl-color:: green
|
||
id:: 654cf565-f328-4e8b-9126-0ca88fe5ce9a
|
||
- The attractive repositories on Github receive a large number of issues daily.
|
||
ls-type:: annotation
|
||
hl-page:: 3
|
||
hl-color:: green
|
||
id:: 654cf582-ed5c-4879-849b-0e06165c581e
|
||
- Assigning similar issues individually to different developers for validating and fixing introduces inconsistencies when asynchronously independent developers fix them, in addition to slowing the fixing process.
|
||
ls-type:: annotation
|
||
hl-page:: 3
|
||
hl-color:: purple
|
||
id:: 654cf595-c370-437b-9b2f-dabaf84d8c0e
|
||
hl-stamp:: 1699542423639
|
||
- grouping similar issues into clusters and assigning each cluster to the same and appropriate developer/team speeds up the fixing process.
|
||
ls-type:: annotation
|
||
hl-page:: 3
|
||
hl-color:: purple
|
||
id:: 654cf5c4-718b-4d37-aaea-b4658b5ed7ad
|
||
- machine learning algorithm-based approach has been proposed to support issue management on GitHub by grouping similar issues together
|
||
ls-type:: annotation
|
||
hl-page:: 3
|
||
hl-color:: purple
|
||
id:: 654cf5d6-940f-4078-a6fe-65af2bca9afb
|
||
- adding new features, reporting bugs to be fixed, or asking a question about the capability of the software.
|
||
ls-type:: annotation
|
||
hl-page:: 3
|
||
hl-color:: green
|
||
id:: 654cfab8-9162-4963-a9c3-30dd587e4443
|
||
hl-stamp:: 1699543739600
|
||
- These issues differ in their nature and quality (asking for support, for improvement of a functionality, bug reporting). Therefore, the issues management in such cases will be harder
|
||
ls-type:: annotation
|
||
hl-page:: 3
|
||
hl-color:: green
|
||
id:: 654cfb34-605f-4a2b-9758-c4a94e7d410b
|
||
- (i) to speed up the process of fixing issues and reduce the turnaround time for fixing,
|
||
ls-type:: annotation
|
||
hl-page:: 4
|
||
hl-color:: green
|
||
id:: 654cfb5d-d6e3-4bec-9e2b-bfa61a354b5f
|
||
- ii) when related issues are grouped together, this speeds up the assignment process by finding the appropriate reviewer(s) for a cluster instead of25 individual issues, especially, if the assignment process is done manuall
|
||
ls-type:: annotation
|
||
hl-page:: 4
|
||
hl-color:: green
|
||
id:: 654cfb6a-d45c-4079-874d-2e5b2b0633bf
|
||
- M. Borg, L. Jonsson, E. Engstrom, B. Bartalos, A. Szab’o, Adopting automated bug assignment in practice: A longitudinal case study at ericsson, ArXiv abs/2209.08955.
|
||
ls-type:: annotation
|
||
hl-page:: 21
|
||
hl-color:: green
|
||
id:: 654cfb75-1118-48da-92f2-0850c197b241
|
||
- Hamzeh Eyal Salman
|
||
ls-type:: annotation
|
||
hl-page:: 3
|
||
hl-color:: green
|
||
id:: 654cfb86-d7a3-4fec-870a-b6a38d62a5ee
|
||
- o consolidate similar issues into a single, well-defined issue that covers all the variations and aspects of the problem
|
||
ls-type:: annotation
|
||
hl-page:: 4
|
||
hl-color:: green
|
||
id:: 654d0919-93c4-4619-81ba-14a90b26610c
|
||
- The difference between similar and duplicate issues is that the former refers to a set of issues with some common characteristics or are closely related, but they are not30 necessarily exact duplicates.
|
||
ls-type:: annotation
|
||
hl-page:: 4
|
||
hl-color:: green
|
||
id:: 654d0928-2741-4db7-9e95-5a6eda44bf53
|
||
- . Liao et al. [7] studied the impact of labeled or tagged issues on their management using six popular repositories
|
||
ls-type:: annotation
|
||
hl-page:: 4
|
||
hl-color:: green
|
||
id:: 654d0956-2e2f-4b33-90ba-0d2785d223e8
|
||
- GitGub
|
||
ls-type:: annotation
|
||
hl-page:: 4
|
||
hl-color:: red
|
||
id:: 654d098f-d572-4f9d-af6e-d8c141179f66
|
||
hl-stamp:: 1699547537664
|
||
- ull-request (PR)
|
||
ls-type:: annotation
|
||
hl-page:: 4
|
||
hl-color:: green
|
||
id:: 654d09d4-103e-48cd-bdcc-af0c0437e17d
|
||
hl-stamp:: 1699547606864
|
||
- ssue reporting
|
||
ls-type:: annotation
|
||
hl-page:: 4
|
||
hl-color:: green
|
||
id:: 654d09da-6d06-4e05-9f35-4707b04e02eb
|
||
- Figure 1:
|
||
ls-type:: annotation
|
||
hl-page:: 5
|
||
hl-color:: yellow
|
||
id:: 654d0a15-8279-4622-85d1-a8d2abb01f6d
|
||
hl-stamp:: 1699547670418
|
||
- he
|
||
ls-type:: annotation
|
||
hl-page:: 5
|
||
hl-color:: red
|
||
id:: 654d0a34-dfd6-453b-9354-0d8407dfb7d1
|
||
hl-stamp:: 1699547701974
|
||
- ssues1,2,3, 4
|
||
ls-type:: annotation
|
||
hl-page:: 5
|
||
hl-color:: red
|
||
id:: 654d0a7d-0b3b-4802-8d5b-dcc99ac20fe1
|
||
- Figure 2:
|
||
ls-type:: annotation
|
||
hl-page:: 6
|
||
hl-color:: red
|
||
id:: 654d0aa6-3ed1-4ec6-85a0-548508d471ae
|
||
- cookie handling
|
||
ls-type:: annotation
|
||
hl-page:: 5
|
||
hl-color:: green
|
||
id:: 654d0b36-f779-484b-920d-0576fc0ab546
|
||
hl-stamp:: 1699547960656
|
||
- there is no research work in the literature addressing the problem of clustering similar issues into groups in GitHub.
|
||
ls-type:: annotation
|
||
hl-page:: 7
|
||
hl-color:: yellow
|
||
id:: 654d0b5e-8aa3-4eb4-ac26-83e9fc224d13
|
||
hl-stamp:: 1699548001102
|
||
- natural language processing (NLP
|
||
ls-type:: annotation
|
||
hl-page:: 7
|
||
hl-color:: green
|
||
id:: 654d0bc1-6a33-4f0d-bbfc-a360a082d835
|
||
hl-stamp:: 1699548099501
|
||
- information retrieval (IR)
|
||
ls-type:: annotation
|
||
hl-page:: 7
|
||
hl-color:: green
|
||
id:: 654d0bc7-ac78-4001-932d-c868c79e9f9c
|
||
- machine learning (ML)
|
||
ls-type:: annotation
|
||
hl-page:: 7
|
||
hl-color:: green
|
||
id:: 654d0bcb-ee33-4375-b1f3-f80a968cfe7a
|
||
- The idea behind using IR and NLP is to find textual overlapping between bug report queries and already existing report queries in the repository
|
||
ls-type:: annotation
|
||
hl-page:: 7
|
||
hl-color:: green
|
||
id:: 654d0c2a-31c0-4c27-8c7f-193094faea16
|
||
- Jacard, cosine, and dice measur
|
||
ls-type:: annotation
|
||
hl-page:: 7
|
||
hl-color:: green
|
||
id:: 654d0ce7-8016-4942-b414-3bd21ea81e0d
|
||
- ug reports i
|
||
ls-type:: annotation
|
||
hl-page:: 7
|
||
hl-color:: green
|
||
id:: 654d0cf0-42bc-4786-b1fc-31032e6e8a58
|
||
- This ranked list allows triggers to compare the incoming bug report with this list to discover the duplication
|
||
ls-type:: annotation
|
||
hl-page:: 7
|
||
hl-color:: green
|
||
id:: 654d0d14-426e-4710-bb84-37f6f5dfb4ca
|
||
- uplicate bug report detection with a combination of information retrieval and topic modeling, i
|
||
ls-type:: annotation
|
||
hl-page:: 22
|
||
hl-color:: green
|
||
id:: 654d0ef4-c76f-4b08-916d-0ebb44b6a959
|
||
- In this study, title and body features are only considered during the clustering process while Label is ignored as it is repetitive and missed in many issues.
|
||
ls-type:: annotation
|
||
hl-page:: 8
|
||
hl-color:: green
|
||
id:: 654d0f20-3e2b-4b6d-bddd-015d4a3ad297
|
||
- Then,175 T
|
||
ls-type:: annotation
|
||
hl-page:: 8
|
||
hl-color:: red
|
||
id:: 654d0f3f-322b-4db1-9b86-509a6d740969
|
||
hl-stamp:: 1699548993168
|
||
- As mentioned above, these important features are the title and body of each issue.
|
||
ls-type:: annotation
|
||
hl-page:: 8
|
||
hl-color:: green
|
||
id:: 654d0f66-4f5e-4b82-a369-653543dd3c36
|
||
hl-stamp:: 1699549038638
|
||
- ground-truth dataset.
|
||
ls-type:: annotation
|
||
hl-page:: 12
|
||
hl-color:: green
|
||
id:: 654d0ffb-b347-4fbe-82ed-ff7c00d2e88e
|
||
- refuted
|
||
ls-type:: annotation
|
||
hl-page:: 12
|
||
hl-color:: yellow
|
||
id:: 654d100f-1b43-44ec-ab81-2e0005a65b48
|
||
hl-stamp:: 1699549201228
|
||
- owever, such a dataset does not exist yet.
|
||
ls-type:: annotation
|
||
hl-page:: 12
|
||
hl-color:: yellow
|
||
id:: 654d1019-abd0-47e7-9f81-34a8f1a8b786
|
||
- duplicate clusters of issues, as 100% duplicate issues are similar issues, too
|
||
ls-type:: annotation
|
||
hl-page:: 12
|
||
hl-color:: red
|
||
id:: 654d1026-9678-4b52-b875-57d0c7e594a0
|
||
hl-stamp:: 1699549224448
|
||
- ese clusters are diverse to ensure free-bias evaluation
|
||
ls-type:: annotation
|
||
hl-page:: 13
|
||
hl-color:: yellow
|
||
id:: 65537124-859a-4460-991f-5652b10380e5
|
||
- RQ1: To what extent are the identified issue clusters correct?
|
||
ls-type:: annotation
|
||
hl-page:: 13
|
||
hl-color:: green
|
||
id:: 655373bd-50f7-4cb1-8767-9cbfa678ca5e
|
||
- to what extent the member issues of a cluster are related to each other (similar).
|
||
ls-type:: annotation
|
||
hl-page:: 13
|
||
hl-color:: green
|
||
id:: 65537486-2805-4b02-896d-869221fe3882
|
||
- RQ2: To what extent are the identified issue clusters complete?
|
||
ls-type:: annotation
|
||
hl-page:: 14
|
||
hl-color:: green
|
||
id:: 65537540-a5a6-4caa-ab48-c3c7664eb7fa
|
||
- AgglomerativeClustering
|
||
ls-type:: annotation
|
||
hl-page:: 15
|
||
hl-color:: green
|
||
id:: 6553757f-f7a0-4aa0-90d7-0a0159cc3f28
|
||
- As a summary, the AHC algorithm can extract relevant similar issue clusters from GitHub repositories depending on only two issue features: title and description.
|
||
ls-type:: annotation
|
||
hl-page:: 20
|
||
hl-color:: green
|
||
id:: 655375f7-0c9e-4efa-b858-da6f1bcf804f
|
||
- to return a number of relevant issue clusters equal to the number of ground-truth clusters
|
||
ls-type:: annotation
|
||
hl-page:: 20
|
||
hl-color:: green
|
||
id:: 65537602-737f-487c-96dc-5148243c5ac2
|
||
- credability
|
||
ls-type:: annotation
|
||
hl-page:: 20
|
||
hl-color:: red
|
||
id:: 65537619-ffb9-4fd6-aea0-a46bd19bb3b6
|
||
- Internal threat
|
||
ls-type:: annotation
|
||
hl-page:: 20
|
||
hl-color:: green
|
||
id:: 6553764c-c160-4efc-a31b-dcbb5b54cc4b
|
||
- External threat
|
||
ls-type:: annotation
|
||
hl-page:: 20
|
||
hl-color:: green
|
||
id:: 65537656-c324-42e1-bd1f-dcf886a36063
|
||
- n this paper, a machine learning-based approach is proposed to support the issue tracking system on GitHub via grouping similar submitted issues together
|
||
ls-type:: annotation
|
||
hl-page:: 21
|
||
hl-color:: green
|
||
id:: 65537688-628f-476d-b11f-582180e09e81
|
||
- clustering process in this research work is guided by two important features extracted from the issue’s report. These features are the title and body of each issue
|
||
ls-type:: annotation
|
||
hl-page:: 8
|
||
hl-color:: yellow
|
||
id:: 6554f0d6-ae81-4f8c-bfab-a7041da86d96
|
||
hl-stamp:: 1700065570293
|
||
- Figure 4: An example of dendrogram tree with cutting line.9
|
||
ls-type:: annotation
|
||
hl-page:: 11
|
||
hl-color:: yellow
|
||
id:: 6554f1eb-38bf-447e-85dc-b41d6594246e
|
||
- without a prior assumption about the number of clusters
|
||
ls-type:: annotation
|
||
hl-page:: 12
|
||
hl-color:: green
|
||
id:: 6554f25c-2197-46bf-a790-e29a265917d1
|
||
- In this research work, this cutting point is set after many tries to find255 reasonable similar issue clusters based on title and description matching
|
||
ls-type:: annotation
|
||
hl-page:: 12
|
||
hl-color:: green
|
||
id:: 6554f27b-e237-409f-ae81-1966cf5cac44
|
||
- A cluster from the dataset that shares a maximum number of issues with the extracted cluster is the target and is called a ground-truth cluster
|
||
ls-type:: annotation
|
||
hl-page:: 14
|
||
hl-color:: green
|
||
id:: 655528b8-2fd3-4cb7-a9eb-f989e4abdc91
|
||
- The higher Recall value you have, the higher completeness cluster is extracted, and vice-versa.
|
||
ls-type:: annotation
|
||
hl-page:: 14
|
||
hl-color:: blue
|
||
id:: 655528ec-ada3-400c-80b1-503e711d3a1a
|
||
- For each extracted cluster (EXC),
|
||
ls-type:: annotation
|
||
hl-page:: 14
|
||
hl-color:: green
|
||
id:: 65552908-9bcb-4bc8-ab76-323e776f4928 |