Engineering Papers⌕ Search

Engineering topics

Whitelam, Stephen

Publications and source records attributed to Whitelam, Stephen.

Using the Metropolis algorithm to explore the loss surface of a recurrent neural network

In the limit of small trial moves the Metropolis Monte Carlo algorithm is equivalent to gradient descent on the energy function in the presence of Gaussian white noise. This observation was originally used to demonstrate a correspondence between Metropolis Monte Carlo moves of model molecules and overdamped Langevin dynamics, but it also applies in the context of training a neural network: making small random changes to the weights of a neural network, accepted with the Metropolis probability, with the loss function playing the role of energy, has the same effect as training by explicit gradient descent in the presence of Gaussian white noise. We explore this correspondence in the context of a simple recurrent neural network. We also explore regimes in which this correspondence breaks down, where the gradient of the loss function becomes very large or small. In these regimes the Metropolis algorithm can still effect training, and so can be used as a probe of the loss function of a neural network in regimes in which gradient descent struggles. We also show that training can be accelerated by making purposely-designed Monte Carlo trial moves of neural-network weights.

Casert, Corneel↗

Learning protocols for the fast and efficient control of active matter

Exact analytic calculation shows that optimal control protocols for passive molecular systems often involve rapid variations and discontinuities. However, similar analytic baselines are not generally available for active-matter systems, because it is more difficult to treat active systems exactly. Here we use machine learning to derive efficient control protocols for active-matter systems, and find that they are characterized by sharp features similar to those seen in passive systems. We show that it is possible to learn protocols that effect fast and efficient state-to-state transformations in simulation models of active particles by encoding the protocol in the form of a neural network. We use evolutionary methods to identify protocols that take active particles from one steady state to another, as quickly as possible or with as little energy expended as possible. Our results show that protocols identified by a flexible neural-network ansatz, which allows the optimization of multiple control parameters and the emergence of sharp features, are more efficient than protocols derived recently by constrained analytical methods. Our learning scheme is straightforward to use in experiment, suggesting a way of designing protocols for the efficient manipulation of active matter in the laboratory.

74 ATOMIC AND MOLECULAR PHYSICS↗

Free-energy estimates from nonequilibrium trajectories under varying-temperature protocols

The Jarzynski equality allows the calculation of free-energy differences using values of work measured from nonequilibrium trajectories. The number of trajectories required to accurately estimate free-energy differences in this way grows sharply with the size of work fluctuations, motivating the search for protocols that perform desired transformations with minimum work. However, protocols of this nature can involve varying temperature, to which the Jarzynski equality does not apply. Here, we derive a variant of the Jarzynski equality that applies to varying-temperature protocols, and show that it can have better convergence properties than the standard version of the equality. We derive this modified equality and the associated fluctuation relation within the framework of Markovian stochastic dynamics, complementing related derivations done within the framework of Hamiltonian dynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Nonequilibrium formulation of varying-temperature bit erasure

Landauer's principle states that erasing a bit of information at fixed temperature T costs at least $k$ B $T$ ln 2 units of work. Here we investigate erasure at varying temperature, to which Landauer's result does not apply. Here we formulate bit erasure as a stochastic nonequilibrium process involving a compression of configuration space, with physical and logical states associated in a symmetric way. Erasure starts and ends at temperature T, but temperature can otherwise vary with time in an arbitrary way. Defined in this way, erasure is governed by a set of nonequilibrium fluctuation relations that show that varying-temperature erasure can done with less work than $k$ B $T$ ln 2. As a result, erasure and the complementary process of bit randomization can be combined to form a work-producing engine cycle.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

On the origin of cooperativity effects in the formation of self-assembled molecular networks at the liquid/solid interface

In this work we investigate the behaviour of molecules at the nanoscale using scanning tunnelling microscopy in order to explore the origin of the cooperativity in the formation of self-assembled molecular networks (SAMNs) at the liquid/solid interface. By studying concentration dependence of alkoxylated dimethylbenzene, a molecular analogue to 5-alkoxylated isophthalic derivatives, but without hydrogen bonding moieties, we show that the cooperativity effect can be experimentally evaluated even for low-interacting systems and that the cooperativity in SAMN formation is its fundamental trait. We conclude that cooperativity must be a local effect and use the nearest-neighbor Ising model to reproduce the coverage vs. concentration curves. The Ising model offers a direct link between statistical thermodynamics and experimental parameters, making it a valuable tool for assessing the thermodynamics of SAMN formation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Learning stochastic dynamics and predicting emergent behavior using transformers

We show that a neural network originally designed for language processing can learn the dynamical rules of a stochastic system by observation of a single dynamical trajectory of the system, and can accurately predict its emergent behavior under conditions not observed during training. We consider a lattice model of active matter undergoing continuous-time Monte Carlo dynamics, simulated at a density at which its steady state comprises small, dispersed clusters. We train a neural network called a transformer on a single trajectory of the model. The transformer, which we show has the capacity to represent dynamical rules that are numerous and nonlocal, learns that the dynamics of this model consists of a small number of processes. Forward-propagated trajectories of the trained transformer, at densities not encountered during training, exhibit motility-induced phase separation and so predict the existence of a nonequilibrium phase transition. Transformers have the flexibility to learn dynamical rules from observation without explicit enumeration of rates or coarse-graining of configuration space, and so the procedure used here can be applied to a wide range of physical systems, including those with large and complex dynamical generators.

97 MATHEMATICS AND COMPUTING↗