Big Microstructure Datasets for Materials Informatics: Using Statistically Conditioned Generative Models to Curate Big Datasets
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
SSA oral presentation
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Los Alamos National Laboratory has been performing nuclear criticality experiments since 1946 at the Pajarito site, starting the Los Alamos Critical Experiments Facility in 1948. A transition period occurred between 2004 and 2011 as operations moved to the National Criticality Experiments Research Center (NCERC), where criticality experiments are now performed. Criticality experiments are essential for determination and verification of nuclear data used in calculations and modeling—such as radiation transport codes—throughout the industry, enhancing nuclear criticality safety. In addition to nuclear data validation and benchmarking, the remotely operated critical assemblies at NCERC are used for a variety of experiments and training classes supporting criticality safety.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
We introduce the first generative model trained on the etlass dataset. Our model generates jets at the constituent level, and it is a permutation-equivariant continuous normalizing flow (CNF) trained with the flow matching technique. It is conditioned on the jet type, so that a single model can be used to generate the ten different jet types of etlass. For the first time, we also introduce a generative model that goes beyond the kinematic features of jet constituents. The etlass dataset includes more features, such as particle-ID and track impact parameter, and we demonstrate that our CNF can accurately model all of these additional features as well. Our generative model for etlass expands on the versatility of existing jet generation techniques, enhancing their potential utility in high-energy physics research, and offering a more comprehensive understanding of the generated jets. Published by the American Physical Society 2025
Machine learning-based unfolding has enabled unbinned and high-dimensional differential cross section measurements. Two main approaches have emerged in this research area; one based on discriminative models and one based on generative models. The main advantage of discriminative models is that they learn a small correction to a starting simulation while generative models scale better to regions of phase space with little data. We propose to use Schrödinger bridges and diffusion models to create , an unfolding approach that combines the strengths of both discriminative and generative models. The key feature of is that its generative model maps one set of events into another without having to go through a known probability density as is the case for normalizing flows and standard diffusion models. We show that achieves excellent performance compared to state of the art methods on a synthetic Z + jets dataset. Published by the American Physical Society 2024
Diffusion generative models are promising alternatives for fast surrogate models, producing high-fidelity physics simulations. However, the generation time often requires an expensive denoising process with hundreds of function evaluations, restricting the current applicability of these models in a realistic setting. In this work, we report updates on the CaloScorearchitecture, detailing the changes in the diffusion process, which produces higher quality samples, and the use of progressive distillation, resulting in a diffusion model capable of generating new samples with a single function evaluation. Here we demonstrate these improvements using the Calorimeter Simulation Challenge 2022 dataset.
Simulating stochastic differential equations (SDEs) in bounded domains, presents significant computational challenges due to particle exit phenomena, which requires accurate modeling of interior stochastic dynamics and boundary interactions. Despite the success of machine learning-based methods in learning SDEs, existing learning methods are not applicable to SDEs in bounded domains because they cannot accurately capture the particle exit dynamics. We present a unified hybrid data-driven approach that combines a conditional diffusion model with an exit prediction neural network to capture both interior stochastic dynamics and boundary exit phenomena. Our ML model consists of two major components: a neural network that learns exit probabilities using binary cross-entropy loss with rigorous convergence guarantees, and a training-free diffusion model that generates state transitions for non-exiting particles using closed-form score functions. The two components are integrated through a probabilistic sampling algorithm that determines particle exit at each time step and generates appropriate state transitions. Here, the performance of the proposed approach is demonstrated via three test cases: a one-dimensional simplified problem for theoretical verification, a two-dimensional advection-diffusion problem in a bounded domain, and a three-dimensional problem of interest to magnetically confined fusion plasmas.
Mechanistic, multicellular, agent-based models are commonly used to investigate tissue, organ, and organism-scale biology at single-cell resolution. The Cellular-Potts Model (CPM) is a powerful and popular framework for developing and interrogating these models. CPMs become computationally expensive at large space- and time- scales making application and investigation of developed models difficult. Surrogate models may allow for the accelerated evaluation of CPMs of complex biological systems. However, the stochastic nature of these models means each set of parameters may give rise to different model configurations, complicating surrogate model development. In this work, we leverage denoising diffusion probabilistic models (DDPMs) to train a generative AI surrogate of a CPM used to investigate in vitro vasculogenesis. We describe the use of an image classifier to learn the characteristics that define unique areas of a 2-dimensional parameter space. We then apply this classifier to aid in surrogate model selection and verification. Our CPM model surrogate generates model configurations 20,000 timesteps ahead of a reference configuration and demonstrates approximately a 22x reduction in computational time as compared to native code execution. Our work represents a step towards the implementation of DDPMs to develop digital twins of stochastic biological systems.
In support of analysis for the biennial Integrated Energy Policy Report, the California Energy Commission and the National Renewable Energy Laboratory have partnered to study the growth of distributed energy resources in California. This study involves the use of National Renewable Energy Laboratory's Distributed Generation Market Demand model, available at https://www.nrel.gov/analysis/dgen/, to project statewide adoption of distributed photovoltaics and paired storage. Key outcomes of the collaboration include: • Improved representation of California building stock, load profiles, historical adoption, and tariffs, including the net billing tariff, in the dGen model; • Trained CEC staff members to use and adapt the dGen model for their specific needs; • Developed a methodology for representing emerging consumer segments to potentially adopt distributed energy resources, including low-income, multifamily, and renter-occupied buildings; • Forecasted solar photovoltaic and paired storage growth in California using a common set of modeling parameters. This report describes the multiyear effort, which includes a discussion of: • Methodology and data employed in adapting the Distributed Generation Market Demand model for California to forecast solar photovoltaic and storage statewide through 2040; • Steps taken to modify the base model to forecast solar photovoltaic adoption in emerging market segments such as multifamily or renter-occupied homes or both; • Future enhancements of the model.
Hadronization is a critical step in the simulation of high-energy particle and nuclear physics experiments. As there is no first principles understanding of this process, physically-inspired hadronization models have a large number of parameters that are fit to data. Deep generative models are a natural replacement for classical techniques, since they are more flexible and may be able to improve the overall precision. Proof of principle studies have shown how to use neural networks to emulate specific hadronization when trained using the inputs and outputs of classical methods. However, these approaches will not work with data, where we do not have a matching between observed hadrons and partons. In this paper, we develop a protocol for fitting a deep generative hadronization model in a realistic setting, where we only have access to a set of hadrons in data. Our approach uses a variation of a Generative Adversarial Network with a permutation invariant discriminator. We find that this setup is able to match the hadronization model in Herwig with multiple sets of parameters. This work represents a significant step forward in a longer term program to develop, train, and integrate machine learning-based hadronization models into parton shower Monte Carlo programs.
The Data and Graph Generation for Modeling Adversary Activity (MAA) project developed a methodology along with scalable graph modeling and generation tools to produce realistic large-scale background activity graphs with embedded adversarial activity pathways. The technical report presents PNNL methodology, released datasets, lessons learned, and recommendations to develop graph analytic algorithms for structure-only and attributed knowledge graphs.
Motivated by the high computational costs of classical simulations, machine-learned generative models can be extremely useful in particle physics and elsewhere. They become especially attractive when surrogate models can efficiently learn the underlying distribution, such that a generated sample outperforms a training sample of limited size. This kind of GANplification has been observed for simple Gaussian models. We show the same effect for a physics simulation, specifically photon showers in an electromagnetic calorimeter.