<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Conference paper | Andrej Leban</title><link>https://andleb.netlify.app/publication-type/paper-conference/</link><atom:link href="https://andleb.netlify.app/publication-type/paper-conference/index.xml" rel="self" type="application/rss+xml"/><description>Conference paper</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Mon, 29 Dec 2025 00:00:00 +0000</lastBuildDate><image><url>https://andleb.netlify.app/media/icon_hu0b7a4cb9992c9ac0e91bd28ffd38dd00_9727_512x512_fill_lanczos_center_3.png</url><title>Conference paper</title><link>https://andleb.netlify.app/publication-type/paper-conference/</link></image><item><title>Energy-Tweedie: Score meets Score, Energy meets Energy</title><link>https://andleb.netlify.app/publication/et/</link><pubDate>Mon, 29 Dec 2025 00:00:00 +0000</pubDate><guid>https://andleb.netlify.app/publication/et/</guid><description>&lt;p>Classical Tweedie’s formula links Gaussian corruption, squared-error denoising, the posterior mean, and the score of the noisy data. We present the &lt;em>Energy–Tweedie&lt;/em> identity, which generalizes this correspondence from a mean-based relation to a distributional one and holds for any Gibbs (&amp;ldquo;energy-based&amp;rdquo;) noise distribution. Each noise distribution induces a kernel scoring rule,
whose path derivative evaluated at the denoising posterior gives the noisy-data score.
In the Gaussian noise case, this reduces exactly to classical Tweedie’s formula.&lt;/p>
&lt;p>The identity has three main consequences:&lt;/p>
&lt;ul>
&lt;li>The score can be estimated from &lt;em>samples&lt;/em> from a denoising posterior model (i.e., a conditional generative model).&lt;/li>
&lt;li>All the parameters of the noise distribution (within the Gibbs family) can be recovered from corrupted data in a principled fashion.&lt;/li>
&lt;li>It provides a score-based perspective on diffusion approaches based on scoring rules; among other consequences, existing score-based samplers can thus be used to generate from such models, with the path through the (multidimensional) noise-parameter space a free design choice at sampling time - illustrated by the MNIST samples below.&lt;/li>
&lt;/ul>
&lt;figure id="figure-mnist-samples-generated-using-various-sampling-paths-through-noise-parameter-space-with-the-same-model">
&lt;div class="d-flex justify-content-center">
&lt;div class="w-100" >&lt;img alt="MNIST samples generated using various sampling *paths* through noise-parameter space with the same model." srcset="
/publication/et/mnist_final_samples_by_path_hu225c3f2dcedfbfd7d4300daa6385bc1b_125533_eeba098d0c117e9e46fe1fec01b4eef4.webp 400w,
/publication/et/mnist_final_samples_by_path_hu225c3f2dcedfbfd7d4300daa6385bc1b_125533_1caa2d18846b2f216fcf14fe4b8163ad.webp 760w,
/publication/et/mnist_final_samples_by_path_hu225c3f2dcedfbfd7d4300daa6385bc1b_125533_1200x1200_fit_q75_h2_lanczos_3.webp 1200w"
src="https://andleb.netlify.app/publication/et/mnist_final_samples_by_path_hu225c3f2dcedfbfd7d4300daa6385bc1b_125533_eeba098d0c117e9e46fe1fec01b4eef4.webp"
width="760"
height="163"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;figcaption>
MNIST samples generated using various sampling &lt;em>paths&lt;/em> through noise-parameter space with the same model.
&lt;/figcaption>&lt;/figure></description></item><item><title>Distributional Autoencoders Know the Score</title><link>https://andleb.netlify.app/publication/dpa/</link><pubDate>Mon, 17 Feb 2025 00:00:00 +0000</pubDate><guid>https://andleb.netlify.app/publication/dpa/</guid><description>&lt;p>The Distributional Principal Autoencoder (DPA) (&lt;a href="https://arxiv.org/abs/2404.13649" target="_blank" rel="noopener">Shen and Meinshausen, 2024&lt;/a>) is a recently introduced class of autoencoders that combines a deterministic encoder with a stochastic decoder trained to reconstruct the full conditional distribution associated with each encoding. Jointly optimizing successive encoding widths produces a principal- component-like ordering of the latent coordinates. Our paper establishes the theoretical structure underlying this method:&lt;/p>
&lt;ol>
&lt;li>
&lt;p>the encoder level sets align &lt;em>exactly&lt;/em> with the data score in the normal directions, &lt;em>and&lt;/em>&lt;/p>
&lt;/li>
&lt;li>
&lt;p>extra latent dimensions beyond the data manifold become &lt;em>completely uninformative&lt;/em>, revealing the intrinsic dimension.&lt;/p>
&lt;/li>
&lt;/ol>
&lt;p>These hold simultaneously and circumvent the usual &lt;em>reconstruction/disentanglement&lt;/em> trade-off in unsupervised learning. Thus, we extend the analogy to PCA made in the original DPA work by proving that, instead of finding principal linear subspaces, DPA learns nonlinear manifolds shaped locally by the data density, with a clear, testable dimensionality criterion — conditional independence.&lt;/p>
&lt;p>The first result also leads to strong performance on molecular simulation data: when the data follow a Boltzmann distribution, the learned encoding aligns with the (unknown) force field. We demonstrate that this allows the method to recover an approximation of the &lt;em>minimum free-energy path&lt;/em> for the Müller–Brown potential (a common benchmark) in a single fit, with the potential to speed up chemical simulations.&lt;/p></description></item></channel></rss>