Files
logseq/pages/hls__COLA-D-23-00096_1699177983001_0.md
T
2025-06-02 17:15:13 +02:00

302 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
file:: [COLA-D-23-00096_1699177983001_0.pdf](../assets/COLA-D-23-00096_1699177983001_0.pdf)
file-path:: ../assets/COLA-D-23-00096_1699177983001_0.pdf
- AI-based Clustering of Similar Issues in GitHubs Repositories Article Type: Full Length Article Keywords: Similar Issues; GitHub; Machine Learning; Maintenance; Clustering Corresponding Author: Hamzeh Eyal Salman Mutah University Karak, JORDAN Corresponding Author Secondary Information: Corresponding Author's Institution: Mutah University
ls-type:: annotation
hl-page:: 1
hl-color:: green
id:: 654cf55c-0711-4113-8546-4963e635cb90
- I-based Clustering of Similar Issues in GitHubs Repositories
ls-type:: annotation
hl-page:: 3
hl-color:: green
id:: 654cf565-f328-4e8b-9126-0ca88fe5ce9a
- The attractive repositories on Github receive a large number of issues daily.
ls-type:: annotation
hl-page:: 3
hl-color:: green
id:: 654cf582-ed5c-4879-849b-0e06165c581e
- Assigning similar issues individually to different developers for validating and fixing introduces inconsistencies when asynchronously independent developers fix them, in addition to slowing the fixing process.
ls-type:: annotation
hl-page:: 3
hl-color:: purple
id:: 654cf595-c370-437b-9b2f-dabaf84d8c0e
hl-stamp:: 1699542423639
- grouping similar issues into clusters and assigning each cluster to the same and appropriate developer/team speeds up the fixing process.
ls-type:: annotation
hl-page:: 3
hl-color:: purple
id:: 654cf5c4-718b-4d37-aaea-b4658b5ed7ad
- machine learning algorithm-based approach has been proposed to support issue management on GitHub by grouping similar issues together
ls-type:: annotation
hl-page:: 3
hl-color:: purple
id:: 654cf5d6-940f-4078-a6fe-65af2bca9afb
- adding new features, reporting bugs to be fixed, or asking a question about the capability of the software.
ls-type:: annotation
hl-page:: 3
hl-color:: green
id:: 654cfab8-9162-4963-a9c3-30dd587e4443
hl-stamp:: 1699543739600
- These issues differ in their nature and quality (asking for support, for improvement of a functionality, bug reporting). Therefore, the issues management in such cases will be harder
ls-type:: annotation
hl-page:: 3
hl-color:: green
id:: 654cfb34-605f-4a2b-9758-c4a94e7d410b
- (i) to speed up the process of fixing issues and reduce the turnaround time for fixing,
ls-type:: annotation
hl-page:: 4
hl-color:: green
id:: 654cfb5d-d6e3-4bec-9e2b-bfa61a354b5f
- ii) when related issues are grouped together, this speeds up the assignment process by finding the appropriate reviewer(s) for a cluster instead of25 individual issues, especially, if the assignment process is done manuall
ls-type:: annotation
hl-page:: 4
hl-color:: green
id:: 654cfb6a-d45c-4079-874d-2e5b2b0633bf
- M. Borg, L. Jonsson, E. Engstrom, B. Bartalos, A. Szabo, Adopting automated bug assignment in practice: A longitudinal case study at ericsson, ArXiv abs/2209.08955.
ls-type:: annotation
hl-page:: 21
hl-color:: green
id:: 654cfb75-1118-48da-92f2-0850c197b241
- Hamzeh Eyal Salman
ls-type:: annotation
hl-page:: 3
hl-color:: green
id:: 654cfb86-d7a3-4fec-870a-b6a38d62a5ee
- o consolidate similar issues into a single, well-defined issue that covers all the variations and aspects of the problem
ls-type:: annotation
hl-page:: 4
hl-color:: green
id:: 654d0919-93c4-4619-81ba-14a90b26610c
- The difference between similar and duplicate issues is that the former refers to a set of issues with some common characteristics or are closely related, but they are not30 necessarily exact duplicates.
ls-type:: annotation
hl-page:: 4
hl-color:: green
id:: 654d0928-2741-4db7-9e95-5a6eda44bf53
- . Liao et al. [7] studied the impact of labeled or tagged issues on their management using six popular repositories
ls-type:: annotation
hl-page:: 4
hl-color:: green
id:: 654d0956-2e2f-4b33-90ba-0d2785d223e8
- GitGub
ls-type:: annotation
hl-page:: 4
hl-color:: red
id:: 654d098f-d572-4f9d-af6e-d8c141179f66
hl-stamp:: 1699547537664
- ull-request (PR)
ls-type:: annotation
hl-page:: 4
hl-color:: green
id:: 654d09d4-103e-48cd-bdcc-af0c0437e17d
hl-stamp:: 1699547606864
- ssue reporting
ls-type:: annotation
hl-page:: 4
hl-color:: green
id:: 654d09da-6d06-4e05-9f35-4707b04e02eb
- Figure 1:
ls-type:: annotation
hl-page:: 5
hl-color:: yellow
id:: 654d0a15-8279-4622-85d1-a8d2abb01f6d
hl-stamp:: 1699547670418
- he
ls-type:: annotation
hl-page:: 5
hl-color:: red
id:: 654d0a34-dfd6-453b-9354-0d8407dfb7d1
hl-stamp:: 1699547701974
- ssues1,2,3, 4
ls-type:: annotation
hl-page:: 5
hl-color:: red
id:: 654d0a7d-0b3b-4802-8d5b-dcc99ac20fe1
- Figure 2:
ls-type:: annotation
hl-page:: 6
hl-color:: red
id:: 654d0aa6-3ed1-4ec6-85a0-548508d471ae
- cookie handling
ls-type:: annotation
hl-page:: 5
hl-color:: green
id:: 654d0b36-f779-484b-920d-0576fc0ab546
hl-stamp:: 1699547960656
- there is no research work in the literature addressing the problem of clustering similar issues into groups in GitHub.
ls-type:: annotation
hl-page:: 7
hl-color:: yellow
id:: 654d0b5e-8aa3-4eb4-ac26-83e9fc224d13
hl-stamp:: 1699548001102
- natural language processing (NLP
ls-type:: annotation
hl-page:: 7
hl-color:: green
id:: 654d0bc1-6a33-4f0d-bbfc-a360a082d835
hl-stamp:: 1699548099501
- information retrieval (IR)
ls-type:: annotation
hl-page:: 7
hl-color:: green
id:: 654d0bc7-ac78-4001-932d-c868c79e9f9c
- machine learning (ML)
ls-type:: annotation
hl-page:: 7
hl-color:: green
id:: 654d0bcb-ee33-4375-b1f3-f80a968cfe7a
- The idea behind using IR and NLP is to find textual overlapping between bug report queries and already existing report queries in the repository
ls-type:: annotation
hl-page:: 7
hl-color:: green
id:: 654d0c2a-31c0-4c27-8c7f-193094faea16
- Jacard, cosine, and dice measur
ls-type:: annotation
hl-page:: 7
hl-color:: green
id:: 654d0ce7-8016-4942-b414-3bd21ea81e0d
- ug reports i
ls-type:: annotation
hl-page:: 7
hl-color:: green
id:: 654d0cf0-42bc-4786-b1fc-31032e6e8a58
- This ranked list allows triggers to compare the incoming bug report with this list to discover the duplication
ls-type:: annotation
hl-page:: 7
hl-color:: green
id:: 654d0d14-426e-4710-bb84-37f6f5dfb4ca
- uplicate bug report detection with a combination of information retrieval and topic modeling, i
ls-type:: annotation
hl-page:: 22
hl-color:: green
id:: 654d0ef4-c76f-4b08-916d-0ebb44b6a959
- In this study, title and body features are only considered during the clustering process while Label is ignored as it is repetitive and missed in many issues.
ls-type:: annotation
hl-page:: 8
hl-color:: green
id:: 654d0f20-3e2b-4b6d-bddd-015d4a3ad297
- Then,175 T
ls-type:: annotation
hl-page:: 8
hl-color:: red
id:: 654d0f3f-322b-4db1-9b86-509a6d740969
hl-stamp:: 1699548993168
- As mentioned above, these important features are the title and body of each issue.
ls-type:: annotation
hl-page:: 8
hl-color:: green
id:: 654d0f66-4f5e-4b82-a369-653543dd3c36
hl-stamp:: 1699549038638
- ground-truth dataset.
ls-type:: annotation
hl-page:: 12
hl-color:: green
id:: 654d0ffb-b347-4fbe-82ed-ff7c00d2e88e
- refuted
ls-type:: annotation
hl-page:: 12
hl-color:: yellow
id:: 654d100f-1b43-44ec-ab81-2e0005a65b48
hl-stamp:: 1699549201228
- owever, such a dataset does not exist yet.
ls-type:: annotation
hl-page:: 12
hl-color:: yellow
id:: 654d1019-abd0-47e7-9f81-34a8f1a8b786
- duplicate clusters of issues, as 100% duplicate issues are similar issues, too
ls-type:: annotation
hl-page:: 12
hl-color:: red
id:: 654d1026-9678-4b52-b875-57d0c7e594a0
hl-stamp:: 1699549224448
- ese clusters are diverse to ensure free-bias evaluation
ls-type:: annotation
hl-page:: 13
hl-color:: yellow
id:: 65537124-859a-4460-991f-5652b10380e5
- RQ1: To what extent are the identified issue clusters correct?
ls-type:: annotation
hl-page:: 13
hl-color:: green
id:: 655373bd-50f7-4cb1-8767-9cbfa678ca5e
- to what extent the member issues of a cluster are related to each other (similar).
ls-type:: annotation
hl-page:: 13
hl-color:: green
id:: 65537486-2805-4b02-896d-869221fe3882
- RQ2: To what extent are the identified issue clusters complete?
ls-type:: annotation
hl-page:: 14
hl-color:: green
id:: 65537540-a5a6-4caa-ab48-c3c7664eb7fa
- AgglomerativeClustering
ls-type:: annotation
hl-page:: 15
hl-color:: green
id:: 6553757f-f7a0-4aa0-90d7-0a0159cc3f28
- As a summary, the AHC algorithm can extract relevant similar issue clusters from GitHub repositories depending on only two issue features: title and description.
ls-type:: annotation
hl-page:: 20
hl-color:: green
id:: 655375f7-0c9e-4efa-b858-da6f1bcf804f
- to return a number of relevant issue clusters equal to the number of ground-truth clusters
ls-type:: annotation
hl-page:: 20
hl-color:: green
id:: 65537602-737f-487c-96dc-5148243c5ac2
- credability
ls-type:: annotation
hl-page:: 20
hl-color:: red
id:: 65537619-ffb9-4fd6-aea0-a46bd19bb3b6
- Internal threat
ls-type:: annotation
hl-page:: 20
hl-color:: green
id:: 6553764c-c160-4efc-a31b-dcbb5b54cc4b
- External threat
ls-type:: annotation
hl-page:: 20
hl-color:: green
id:: 65537656-c324-42e1-bd1f-dcf886a36063
- n this paper, a machine learning-based approach is proposed to support the issue tracking system on GitHub via grouping similar submitted issues together
ls-type:: annotation
hl-page:: 21
hl-color:: green
id:: 65537688-628f-476d-b11f-582180e09e81
- clustering process in this research work is guided by two important features extracted from the issues report. These features are the title and body of each issue
ls-type:: annotation
hl-page:: 8
hl-color:: yellow
id:: 6554f0d6-ae81-4f8c-bfab-a7041da86d96
hl-stamp:: 1700065570293
- Figure 4: An example of dendrogram tree with cutting line.9
ls-type:: annotation
hl-page:: 11
hl-color:: yellow
id:: 6554f1eb-38bf-447e-85dc-b41d6594246e
- without a prior assumption about the number of clusters
ls-type:: annotation
hl-page:: 12
hl-color:: green
id:: 6554f25c-2197-46bf-a790-e29a265917d1
- In this research work, this cutting point is set after many tries to find255 reasonable similar issue clusters based on title and description matching
ls-type:: annotation
hl-page:: 12
hl-color:: green
id:: 6554f27b-e237-409f-ae81-1966cf5cac44
- A cluster from the dataset that shares a maximum number of issues with the extracted cluster is the target and is called a ground-truth cluster
ls-type:: annotation
hl-page:: 14
hl-color:: green
id:: 655528b8-2fd3-4cb7-a9eb-f989e4abdc91
- The higher Recall value you have, the higher completeness cluster is extracted, and vice-versa.
ls-type:: annotation
hl-page:: 14
hl-color:: blue
id:: 655528ec-ada3-400c-80b1-503e711d3a1a
- For each extracted cluster (EXC),
ls-type:: annotation
hl-page:: 14
hl-color:: green
id:: 65552908-9bcb-4bc8-ab76-323e776f4928