Lucas de Castro Rodrigues Pereira,
Maykon Marcos Junior,
Guilherme de Brito Santos,
Isabela Cristina Sabo,
Thiago Raulino Dal Pont,
Andressa Silveira Viana Maurmann,
Luísa Bollmann,
Maite Fortes Vieira,
João Gabriel Mohr,
Cristian Alexandre Alchini,
Bruno Cassol da Silva &
Aires José Rover
Abstract
This research paper explores the effectiveness of OpenAI’s GPT-4o model in extracting factors and structuring judgments (text data) regarding Brazilian Consumer Law. We constructed two datasets: an unstructured one, comprising judgments on air transport service failures (e.g., flight delays, cancellations, and baggage loss), and a structured dataset created by legal experts manually extracting relevant factors. Two prompts-a raw and a refined version-were tested using two experimental setups. The first setup, Singular, involved 900 judgments with three requests per document. The second, Factor-based sets, also used 900 judgments but partitioned the prompts into three segments: pro-factors, con-factors, and dimensions. Metrics such as accuracy, F1-score, precision, recall, and RMSE were used based on the value type (numerical or categorical). The Singular setup presents the best results, with the refined prompt achieving approximately 90% accuracy and 60% F1-score. In this experiment, individual factor analysis showed moderate accuracy for “Airline assistance” factor, likely due to the lack of clear parameters defining adequate assistance in cases like flight delays or cancellations. However, individual analysis of the F1-score, precision and recall revealed very low values for factors such as “Right to regret and repayment claim” and “Downgrade”, highlighting the model’s difficulty with some unbalanced class distributions. This study demonstrates the potential of LLMs in structuring legal datasets and aiding professionals in extracting factors from legal texts without fine-tuning. Limited model access restricted further experimentation.