Nathalia Barbosa, Paulo Borba et Leuson Da Silva
Article de revue (2026)
|
Libre accès au plein texte de ce document Version officielle de l'éditeur Conditions d'utilisation: Creative Commons: Attribution (CC BY) Télécharger (729kB) |
Abstract
Semantic conflicts occur when a developer introduces changes to a codebase that unintentionally affect the behavior of changes integrated in parallel by other developers. Since merge tools used in practice cannot detect this type of conflict, complementary tools have been proposed, such as SAM (SemAntic Merge), based on the SMAT approach and relies on the generation and execution of unit tests in Java. Despite showing good conflict detection capabilities, SAM presents a high rate of false negatives (existing conflicts not signaled by it). Part of this problem is due to the natural limitations of unit test generation tools, specifically Randoop and EvoSuite. To understand if these limitations can be overcome by large language models (LLMs), this work proposes, and integrates into SMAT, LUCIA (LLM-based Unit-tests for Conflict Identification & Analysis), a new test generation tool based on the Ollama framework to interface with various local LLMs, including the Llama, Gemma, DeepSeek, and Qwen families. We then explore these models' capability to generate tests, using different interaction strategies, prompts with different contents, and different model parameter configurations. We evaluate the results with two distinct samples: a benchmark with simpler systems, used in related work, and a more significant sample based on complex systems used in practice. Finally, we evaluate the effectiveness of LUCIA in detecting conflicts, comparing the selected LLMs with each other and against the original test generation tools used by SAM: Randoop, Randoop Clean, EvoSuite, and Differential EvoSuite. Our evaluation shows that LLMs effectively detect semantic conflicts, with Llama and Qwen outperforming the other models. While no single model could surpass traditional tools, multiple models combined achieved superior performance in conflict detection. Overall, the new extension identified three additional conflicts undetected by SAM, representing a 200% improvement in identifying unique conflicts compared to the single-model approach (Code Llama 70B) used in our previous work. These results reinforce previous findings that LLMs can generate unit tests that are effective in detecting semantic conflicts, while demonstrating the superior effectiveness of multi-model approaches.
Mots clés
| Matériel d'accompagnement: | |
|---|---|
| Département: | Département de génie informatique et génie logiciel |
| Organismes subventionnaires: | CNPq |
| Numéro de subvention: | 408817/ 2024-0, 401032/ 2025-6 |
| URL de PolyPublie: | https://publications.polymtl.ca/80247/ |
| Titre de la revue: | Journal of Software Engineering Research and Development (vol. 14, no 1) |
| Maison d'édition: | Sociedade Brasileira de Computação |
| DOI: | 10.5753/jserd.2026.7565 |
| URL officielle: | https://journals-sol.sbc.org.br/index.php/jserd/ar... |
| Date du dépôt: | 11 août 2026 09:52 |
| Dernière modification: | 13 août 2026 15:45 |
| Citer en APA 7: | Barbosa, N., Borba, P., & Da Silva, L. (2026). Assessing the effectiveness of Large Language Models in detecting semantic conflicts. Journal of Software Engineering Research and Development, 14(1), 25 pages. https://journals-sol.sbc.org.br/index.php/jserd/article/view/7565 |
|---|---|
Statistiques
Total des téléchargements à partir de PolyPublie
Téléchargements par année
Provenance des téléchargements
Dimensions
