Banca de DEFESA: MARCOS ALEXANDRE DE MELO MEDEIROS

Uma banca de DEFESA de DOUTORADO foi cadastrada pelo programa.
DISCENTE : MARCOS ALEXANDRE DE MELO MEDEIROS
DATA : 12/08/2026
HORA: 18:00
LOCAL: Google Meet
TÍTULO:

Improving Bug Localization and Repair using Crash Report Mining and Large Language Models

 

PALAVRAS-CHAVES:

bug localization, bug fixing, mining crash reports, large language models


PÁGINAS: 87
RESUMO:

Crash reports and stack traces are widely used by developers to investigate the root causes of software failures. However, identifying the faulty source code based on these reports is often complex, especially when dealing with large volumes of data. Recent research has proposed methods for grouping crash reports by similarity and using stack trace information to assist in bug localization. Although these techniques have shown promising results in retrospective studies, mostly with open-source projects, little is known about their practical impact in real-world software development environments in daily life. This thesis investigates the use of a stack trace-based crash report mining approach to support bug localization and program repair in an industrial context. Over 18 months, the bug localization approach was applied in collaboration with development teams responsible for three large-scale Java web systems. More than 750,000 crash reports were grouped, and more than 130 real bug fix tasks were opened and analyzed. Our findings indicate that the approach effectively reduces the cognitive effort required to locate buggy components, achieving over 90% recall and approximately 82% mean average precision. Developer feedback revealed that factors such as critical business requirements and error reproducibility influence the prioritization of crash groups and that the ranked list of suspicious files and methods remains useful even when the root cause lies outside the predicted scope. The thesis also explores how Large Language Models (LLMs) can complement crash report-based bug localization by providing automated fix suggestions for software defects. A retrospective study of 38 bugs from large-scale Java Enterprise web-based systems assesses the structural characteristics and effectiveness of LLM-generated localization and repair responses across different prompt configurations and models. We found that a single code fragment from a single class is the most frequent response type, and that both the prompt combination and the model choice influence the structural characteristics of responses. Comparing four prompt combinations across five LLM models on a subset of our dataset, we identified the best-performing model and the contextual information that most effectively supports bug localization and repair. Evaluated on the full dataset, the best-performing model achieved 71% localization accuracy and 52\% repair accuracy, demonstrating that combining crash report mining with LLM-based repair can effectively support bug localization and fixing in industrial maintenance workflows. Building on these findings, we conducted a third, ongoing study that applied the approach to the day-to-day bug-fixing workflow of three development teams. Preliminary results show 40\% localization accuracy and 20\% valid repair. However, developers reported that the suggestions were directly useful or helped guide their investigation in most cases, indicating a need for further validation.


MEMBROS DA BANCA:
Externo à Instituição - CHRISTOPH TREUDE - UA
Externo à Instituição - DANIEL ALENCAR DA COSTA - UO
Interno - 1671962 - EDUARDO HENRIQUE DA SILVA ARANHA
Externo ao Programa - 2274774 - EIJI ADACHI MEDEIROS BARBOSA - UFRNInterna - 1709820 - ROBERTA DE SOUZA COELHO
Externo à Instituição - RODRIGO BONIFACIO DE ALMEIDA - UnB
Presidente - 1644456 - UIRA KULESZA
Notícia cadastrada em: 28/07/2026 14:12
SIGAA | Superintendência de Tecnologia da Informação - (84) 3342 2210 | Copyright © 2006-2026 - UFRN - sigaa05-producao.info.ufrn.br.sigaa05-producao