DSReg Untangles Frozen AI Representations Without Labels or Retraining
A new regularizer called DSReg peels apart the mixed latents inside JEPA-style world models, recovering each individual world variable with no decoder and no labels.
- DSReg provably recovers individual world latents from JEPA encoders without decoders, reconstruction, or labels (paper).
- Builds on LeJEPA's linear identifiability by fitting a sparsity-selecting rotation on top of a frozen encoder.
- Requires only Structural Diversity: no two latents leave the same dependency footprint on observations.
- Main theorem is formally verified in Lean 4 and strictly weakens prior structural sparsity conditions.
- Scales to 8192 latents on a single 48 GB GPU with graceful recovery degradation.
- Code available on GitHub; project page at dsreg.github.io.
DSReg rotates mixed JEPA features into individual latent factors
Yujia Zheng, David Klindt, Randall Balestriero, and Bernhard Schölkopf propose DSReg, a lightweight regularizer that disentangles representations from a frozen LeJEPA encoder. Under the paper’s assumptions, the method recovers individual latent factors without labels, a decoder, or an end-to-end reconstruction objective.
Representation learners often preserve the information needed for downstream tasks while distributing each real-world factor across many coordinates. A supervised linear probe can recover useful combinations, but coordinate-level editing and sparse control work better when one axis corresponds to one factor. DSReg addresses that alignment problem by finding a rotation of the learned representation.
Why JEPA coordinates remain mixed
Joint-Embedding Predictive Architectures train an encoder to predict representations of related observations, such as different views or moments from the same scene. The cited LeJEPA result establishes linear identifiability in the Gaussian-latent setting: the learned representation contains the true latent state up to an invertible linear transformation.
Linear identifiability preserves information while leaving the coordinate system ambiguous. Position, color, velocity, and texture can remain mixed across dimensions, even when a linear map could recover them with sufficient labels. DSReg narrows that ambiguity to signed permutation, meaning coordinates may be reordered or flipped in direction but are no longer arbitrary mixtures.
Distinct footprints identify the factors
DSReg relies on an assumption the authors call Structural Diversity. Each latent factor must affect a distinct set of observed variables, such as pixels or sensor readings. That set forms the factor’s dependency footprint and corresponds locally to the nonzero entries in a Jacobian column.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.