Trajectory completion on the Müller-Brown potential: dashed - a test trajectory, solid - the continuation generated by the largest modelThis is my final project for STATS 700: LLMs and Transformers at the University of Michigan in Fall 2024.
In molecular dynamics, one is usually after collective variables (CVs) - a few coordinates that capture the slow, important motion of a system while ignoring the fast vibrations. Most approaches to discovering them treat the simulation snapshots as i.i.d. samples. The question here was whether modern autoregressive architectures could instead make use of the trajectories themselves, in the way they have recently been used for protein structure generation.
I re-implemented a BERT-style Transformer encoder for continuous trajectories (a linear projection replaces the token embedding), with two heads sharing the same embeddings: one predicts the next step of the trajectory, the other classifies the metastable state the trajectory belongs to. The idea is that having to satisfy both objectives at once should push the embeddings towards the underlying physics rather than shortcuts. The data are Langevin simulations on the Müller-Brown potential - a standard two-dimensional test system with two minima separated by an energy barrier - which made it possible to generate as many trajectories as needed and to control how hard the problem is.
The medium and large models generate physically plausible trajectories and classify the states with up to 99% accuracy. The more interesting finding is in the attention patterns: the models learn a coarse discretization of time, attending to a handful of key steps in each part of the trajectory, and the larger the model, the fewer such steps it needs. The main limitation is that the encoder never transforms the input coordinates themselves, so this coarse-graining in time is a step towards a CV rather than a CV itself; combining the architecture with a spatial encoder and working in its latent space would be a natural next step.