Abstract
This article investigates sonic data generated with a diffusion model for its capacity to trouble the logic of representation. An image–sound sequence from the author’s animation, Carnival, is posited as a computational (dis)ruptor/rupture to sound identification and attribution processes. The animation’s sound component was generated during image-to-sound experiments with Stable Audio Open, a diffusion-transformer (DiT)-based open-source text-to-audio model trained on Creative Commons-licensed audio. The author proposes a piece of terminology, ‘timbral indeterminacy’, enlisting analyses of visual indeterminacy in traditional and machine-learning-based visual art, to describe a passage of sound collage. As a perceptual marker, timbre enables sound source identification and assists a listener to distinguish different sounds from one another. Timbral indeterminacy is produced when an image-sound sequence from Carnival initially appears coherent, but on closer listening, the sonic component becomes ambiguous. The animation’s sonic identity is co-constituted by multiple pasts and sounding object identities, thus challenging processes of categorization and attribution. Carnival’s sonic referents have been scooped out; in their vanishing, the listener is haunted by non-representational tensions.