Abstract
Interpretive captions play a crucial role in museum contexts by providing verbal information that complements and enhances visual artworks, thereby facilitating communication between visitors and the museum. However, despite their significance, there remains uncertainty about how these captions, generally perceived as secondary to the artworks, vary in their enhancing roles. In response to this gap, this study employs the multimodal corpus SemArt, taking a corpus-driven approach to explore the correlations between interpretive captions and classical European paintings. By applying the Latent Dirichlet Allocation method, themes and keywords were extracted from the interpretive captions and analyzed within a visual communication framework to uncover sub-genre variations and different text-image relations. The findings reveal seven sub-genres within interpretive captions and two types of text-image enhancement relations: content-enhanced and form-enhanced. This research highlights the benefits and challenges of using exploratory quantitative analysis to facilitate multimodal research, paving the way for broader applications of unsupervised machine-learning techniques in digital humanities.