Repository logo

Infoscience

  • English
  • French
Log In
Logo EPFL, École polytechnique fédérale de Lausanne

Infoscience

  • English
  • French
Log In
  1. Home
  2. Academic and Research Output
  3. Journal articles
  4. Probing the latent hierarchical structure of data via diffusion models
 
research article

Probing the latent hierarchical structure of data via diffusion models

Sclocchi, Antonio  
•
Favero, Alessandro  
•
Itzhak Levi, Noam  
Show more
August 1, 2025
Journal of Statistical Mechanics: Theory and Experiment

High-dimensional data must be highly structured to be learnable. Although the compositional and hierarchical nature of data is often put forward to explain learnability, quantitative measurements establishing these properties are scarce. Likewise, accessing the latent variables underlying such a data structure remains a challenge. In this work, we show that forward-backward experiments in diffusion-based models, where data is noised and then denoised to generate new samples, are a promising tool to probe the latent structure of data. We predict in simple hierarchical models that, in this process, changes in data occur by correlated chunks, with a length scale that diverges at a noise level where a phase transition is known to take place. Remarkably, we confirm this prediction in both text and image datasets using state-of-the-art diffusion models. Our results show how latent variable changes manifest in the data and establish how to measure these effects in real data using diffusion models.

  • Files
  • Details
  • Metrics
Loading...
Thumbnail Image
Name

Sclocchi_2025_J._Stat._Mech._2025_084005.pdf

Type

Main Document

Version

Published version

Access type

openaccess

License Condition

CC BY

Size

2.25 MB

Format

Adobe PDF

Checksum (MD5)

adab2ce708d4b0e879237808bb5280d7

Logo EPFL, École polytechnique fédérale de Lausanne
  • Contact
  • infoscience@epfl.ch

  • Follow us on Facebook
  • Follow us on Instagram
  • Follow us on LinkedIn
  • Follow us on X
  • Follow us on Youtube
AccessibilityLegal noticePrivacy policyCookie settingsEnd User AgreementGet helpFeedback

Infoscience is a service managed and provided by the Library and IT Services of EPFL. © EPFL, tous droits réservés