Context-aware multimodal AI navigates hidden pathways in five centuries of art evolution.

The rise of multimodal generative AI transforms the intersection of technology and art, offering richer insights into large-scale artworks. While significant research has focused on their creative potential, their ability to represent artworks in latent spaces remains underexamined. We use generative AI, specifically Stable Diffusion, to analyze 500 y of Western paintings by extracting two types of latent information with the model: formal aspects (e.g., colors) and contextual aspects (e.g., sub
The rise of multimodal generative AI transforms the intersection of technology and art, offering richer insights into large-scale artworks. While significant research has focused on their creative potential, their ability to represent artworks in latent spaces remains underexamined. We use generative AI, specifically Stable Diffusion, to analyze 500 y of Western paintings by extracting two types of latent information with the model: formal aspects (e.g., colors) and contextual aspects (e.g., subjects). Our findings reveal that contextual information exhibits stronger vector alignment and orientation with conventional artistic periods, styles, and individual artists than formal elements. Also, we show how artistic expression aligns with historical shifts using contextual keywords extracted from paintings. Our generative experiment, infusing prospective contexts into historical artworks, validates this vector alignment and orientation by synthesizing artworks consistent with the stylistic patterns of target periods. This study demonstrates how multimodal AI expands traditional formal analysis by integrating temporal, cultural, and historical contexts to quantify the latent structure of cultural knowledge.




