Song DONG

INT–δ: Painting to Poem

Instructor: Panagiotis Michalatos
Collaborative Project (This project was developed through a fully collaborative two-person workflow. All conceptual and technical decisions were made jointly through continuous dialogue, iteration, and testing.)
Teammate: Jing Liao 
Tools & Technologies: Mask R-CNN, CLIP, ArtEmis, custom texture classifier, Python + PyTorch, poem layout engine

Introduction
Painting to Poem is a computational system that transforms visual artworks into spatially structured poems. The pipeline analyzes a painting’s composition, emotion, and texture, translating these qualities into the arrangement of language.

Painting to Poem — image 1 Painting to Poem — image 2Painting to Poem — image 3Painting to Poem — image 4Painting to Poem — image 5System overview
The system builds a multimodal bridge between image and text through several layers of computational analysis:
  • Segmentation & Form Extraction
        Mask R-CNN identifies figures and major compositional elements, creating spatial “containers” that determine poem layout.
  • Semantic Resonance
        CLIP matches segments with poetic lines from a curated corpus, selecting text that semantically aligns with visual content.
  • Affective Interpretation
         ArtEmis analyzes the emotional tone of the painting—such as calmness, melancholy, awe—and filters poetic material by mood.
  • Texture Classification
         A custom classifier identifies visual textures (soft, turbulent, dense, airy) and maps them to text behaviors, such as spacing, repetition, or structural density.

Together, these layers enable the system to generate two poetic modes:
Type A, where text forms the silhouette of detected figures;
Type B, where text becomes a texture-like field embedded across the painting.
Painting to Poem — image 6Painting to Poem — image 7Painting to Poem — image 8Painting to Poem — image 9Painting to Poem — image 10Painting to Poem — image 11Painting to Poem — image 12Painting to Poem — image 13
Poem Generation Modes
Type A — Text as Figure
In Type A outputs, the poem follows the boundaries of segmented figures or objects. The text adopts the silhouettes extracted by Mask R-CNN, producing poem-shaped forms that echo the painting’s subjects. The poem becomes the figure.

Type B — Text as Texture

In Type B outputs, the poem behaves like a visual texture. Text fills regions of the painting according to their detected texture categories, generating rhythm, density, or repetition patterns that mirror the artwork’s surface qualities. The poem becomes the material.
These two modes allow the system to respond flexibly to different styles of paintings—figurative, abstract, textured, or atmospheric—producing distinct poetic interpretations.
Painting to Poem — image 14 Output Gallery
The system produces a wide range of visual poems shaped by the unique properties of each painting.
Because each artwork carries its own composition, emotional tone, and textural cues, the resulting poetic forms vary dramatically—from delicate contours to dense textual fields.

This gallery showcases the diversity and interpretive depth of the system’s outputs.

Painting to Poem — image 15
Painting to Poem — image 16Painting to Poem — image 17Painting to Poem — image 18
Painting to Poem — image 19Painting to Poem — image 20Painting to Poem — image 21Painting to Poem — image 22Painting to Poem — image 23Painting to Poem — image 24Painting to Poem — image 25Painting to Poem — image 26Painting to Poem — image 27Painting to Poem — image 28
Reflection
Painting to Poem challenges conventional approaches to AI interpretation by shifting from description to poetic transformation. Instead of summarizing what a painting depicts, the system engages with how it feels, how it is structured, and how it behaves visually. The project positions poetry as a computational material shaped by vision models and text-generation engines.

This work reframes machine perception as a creative act—one that produces new expressive forms rather than objective analysis. By letting paintings generate poems through multimodal inference, the project invites viewers to encounter artworks through an alternative sensory and conceptual channel, expanding the possibilities of cross-modal human–AI collaboration.