<  Retour au portail Polytechnique Montréal

Assessing the effectiveness of Large Language Models in detecting semantic conflicts

Nathalia Barbosa, Paulo Borba et Leuson Da Silva

Article de revue (2026)

Document en libre accès dans PolyPublie et chez l'éditeur officiel
[img]
Affichage préliminaire
Libre accès au plein texte de ce document
Version officielle de l'éditeur
Conditions d'utilisation: Creative Commons: Attribution (CC BY)
Télécharger (729kB)
Afficher le résumé
Cacher le résumé

Abstract

Semantic conflicts occur when a developer introduces changes to a codebase that unintentionally affect the behavior of changes integrated in parallel by other developers. Since merge tools used in practice cannot detect this type of conflict, complementary tools have been proposed, such as SAM (SemAntic Merge), based on the SMAT approach and relies on the generation and execution of unit tests in Java. Despite showing good conflict detection capabilities, SAM presents a high rate of false negatives (existing conflicts not signaled by it). Part of this problem is due to the natural limitations of unit test generation tools, specifically Randoop and EvoSuite. To understand if these limitations can be overcome by large language models (LLMs), this work proposes, and integrates into SMAT, LUCIA (LLM-based Unit-tests for Conflict Identification & Analysis), a new test generation tool based on the Ollama framework to interface with various local LLMs, including the Llama, Gemma, DeepSeek, and Qwen families. We then explore these models' capability to generate tests, using different interaction strategies, prompts with different contents, and different model parameter configurations. We evaluate the results with two distinct samples: a benchmark with simpler systems, used in related work, and a more significant sample based on complex systems used in practice. Finally, we evaluate the effectiveness of LUCIA in detecting conflicts, comparing the selected LLMs with each other and against the original test generation tools used by SAM: Randoop, Randoop Clean, EvoSuite, and Differential EvoSuite. Our evaluation shows that LLMs effectively detect semantic conflicts, with Llama and Qwen outperforming the other models. While no single model could surpass traditional tools, multiple models combined achieved superior performance in conflict detection. Overall, the new extension identified three additional conflicts undetected by SAM, representing a 200% improvement in identifying unique conflicts compared to the single-model approach (Code Llama 70B) used in our previous work. These results reinforce previous findings that LLMs can generate unit tests that are effective in detecting semantic conflicts, while demonstrating the superior effectiveness of multi-model approaches.

Mots clés

Matériel d'accompagnement:
Département: Département de génie informatique et génie logiciel
Organismes subventionnaires: CNPq
Numéro de subvention: 408817/ 2024-0, 401032/ 2025-6
URL de PolyPublie: https://publications.polymtl.ca/80247/
Titre de la revue: Journal of Software Engineering Research and Development (vol. 14, no 1)
Maison d'édition: Sociedade Brasileira de Computação
DOI: 10.5753/jserd.2026.7565
URL officielle: https://journals-sol.sbc.org.br/index.php/jserd/ar...
Date du dépôt: 11 août 2026 09:52
Dernière modification: 13 août 2026 15:45
Citer en APA 7: Barbosa, N., Borba, P., & Da Silva, L. (2026). Assessing the effectiveness of Large Language Models in detecting semantic conflicts. Journal of Software Engineering Research and Development, 14(1), 25 pages. https://journals-sol.sbc.org.br/index.php/jserd/article/view/7565

Statistiques

Total des téléchargements à partir de PolyPublie

Téléchargements par année

Provenance des téléchargements

Dimensions

Actions réservées au personnel

Afficher document Afficher document