Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “foundation models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Foundation AI Models for Science

Foundation Models (FM) are AI models that are designed to replace a task or an application specific model. These FM can be applied to many different downstream applications. These FM are trained using self supervised techniques and can be built on any type of sequence data. The use of self supervised learning removes the hurdle for developing a large labeled dataset for training. Most FM use transformer architecture utilizes the notion of self attention which allows the network to model the influence of distant data points to each other both in space and time. The FM models exhibit emergent properties that are induced from the data. FM can be an important tool for science. The scale of these models results in better performance for different downstream applications and these applications show better accuracy over models built from scratch. FM drastically reduces the cost of entry to build different downstream applications both in time and effort. FM for selected science datasets such as optical satellite data, can accelerate applications ranging from data quality monitoring, feature detection and prediction. FM can make it easier to infuse AI into scientific research by removing the training data bottleneck and increasing the use of science data.

Manil Maskey↗

Inception of a Spaceflight-specific Mouse to Human Expression Profiling Translation Model

Rodents are foundational model organisms often utilized due to their seemingly analogous morphologies and biological responses to humans. However, recent studies have demonstrated that murine model data are limited in their applicability, particularly in inflammatory disease. In space studies, accurately predicting human response from mouse data is critical due to extreme limiting factors in both rodent and human spaceflight research. With successful prediction, spaceflight ailments can be predicted and prevented while respecting the constraints of the spaceflight industry and minimizing danger to humans. To do so, novel methodologies must be developed that predict human response from murine data after considering biological differences between rodents and humans in spaceflight. After considering terrestrial models, we determined that a spaceflight-based expression profiting translation tool should be created to accurately capture predictions of human gene expression in spaceflight from mouse data. To prepare to build this model, we organized known human spaceflight risks, chose analog human diseases as training data categories, then identified existing RNASeq disease datasets from GEO as potential training data. In addition, we classified existing Genelab mouse differential gene expression datasets for use as experimental data.

Translation↗

Effects of Composition and Oxidation States on the Structures of Chromium-Containing Sodium Silicate Glasses: Molecular Dynamics Simulations using Machine Learning Interatomic Potentials

Chromium represents a significant challenge for the vitrification of high-level nuclear waste into silicate and borosilicate glasses due to its low solubility and variable oxidation states, which can limit the waste loading due to promotion of crystallization or phase separation during processing. In this study, we modeled chromium containing silicate glasses using molecular dynamics simulations with three machine learning interatomic potentials (MLIPs), MACE, CHGNet, and PFP were employed, to gain insights on glass composition and oxidation states on the structures of these glasses. One of the goals is to evaluate their ability of these MLIPs to accurately represent the general structure of silicate glasses and chromium local environments as a function of chromium oxidation states. Density Functional Theory (DFT) based calculations and experimental data such as neutron structure factors were used to validate the structural models. It was found that the foundation models of all three MLIPs are able to reproduce general structural features of the sodium silicate glass structure consistent with experimental and DFT data, but only CHGNet and PFP can accurately capture the oxidation states and local environment of chromium: tetrahedral for Cr6+ and octahedral for Cr3+. Furthermore, we studied the effect of varying Cr3+/ Cr6+ (Cr3+/Crtotal) ratio and total chromium content using PFP. Our results show that Cr6+ enhances network polymerization by reducing non-bridging oxygens through Na? charge compensation required due to the formation of chromate (CrO42-) species, while Cr³? acts as a network modifier that disrupts connectivity. System size effects on the structural characteristics and chromium environments were also tested using the PFP potential. This work highlights the importance of careful validation on the precision, transferability, and potential of MLIPs for modeling glasses containing transition metal elements that can exist in multiple oxidation states. It is also encouraging to see the foundational models are all three MLFFs are able to reproduce the basic sodium silicate glass structures, while suggesting additional training or refining is needed to improve the description of more complex systems containing transition metals.

Puga, Christina L.↗

pnnl/SNAP

In this work, we detail two uncertainty quantification (UQ) methods that provide complementary information. Readout ensembling, by finetuning only the readout layers of an ensemble of foundation models, provides information about model uncertainty. Amending the final readout layer to predict upper and lower quantiles replaces point predictions with distributional predictions, which provide information about uncertainty within the underlying training data. We demonstrate our approach with the MACE-MP-0 model, applying UQ to both the foundation model and a series of finetuned models. The uncertainties produced by the ensemble and quantile methods are demonstrated to be distinct measures by which the quality of the NNP output can be judged.

Pope, Jenna (Bilbrey) [Pacific Northwest National ↗

Scaling Laws of Graph Neural Networks for Atomistic Materials Modeling

Atomistic materials modeling is a critical task with wide-ranging applications, from drug discovery to materials science, where accurate predictions of the target material property can lead to significant advancements in scientific discovery. Graph Neural Networks (GNNs) represent the state-of-the-art approach for modeling atomistic material data thanks to their capacity to capture complex relational structures. While machine learning performance has historically improved with larger models and datasets, GNNs for atomistic materials modeling remain relatively small compared to large language models (LLMs), which leverage billions of parameters and terabyte-scale datasets to achieve remarkable performance in their respective domains. To address this gap, we explore the scaling limits of GNNs for atomistic materials modeling by developing a foundational model with billions of parameters, trained on extensive datasets in terabytescale. Our approach incorporates techniques from LLM libraries to efficiently manage large-scale data and models, enabling both effective training and deployment of these large-scale GNN models. This work addresses three fundamental questions in scaling GNNs: the potential for scaling GNN model architectures, the effect of dataset size on model accuracy, and the applicability of LLM-inspired techniques to GNN architectures. Specifically, the outcomes of this study include (1) insights into the scaling laws for GNNs, highlighting the relationship between model size, dataset volume, and accuracy, (2) a foundational GNN model optimized for atomistic materials modeling, and (3) a GNN codebase enhanced with advanced LLM-based training techniques. Our findings lay the groundwork for large-scale GNNs with billions of parameters and terabyte-scale datasets, establishing a scalable pathway for future advancements in atomistic materials modeling.

Li, Chaojian [ORNL] (ORCID:0000000340309777)↗

Privacy-Preserving Federated Learning for Science: Challenges and Research Directions

This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific AI models, in particular, foundation models~(FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy-- an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.

Kim, Kibaek [Argonne National Laboratory (ANL)]↗

Is tokenization needed for masked particle modeling?

In this work, we significantly enhance masked particle modeling (MPM), a self-supervised learning scheme for constructing highly expressive representations of unordered sets relevant to developing foundation models for high-energy physics. In MPM, a model is trained to recover the missing elements of a set, a learning objective that requires no labels and can be applied directly to experimental data. We achieve significant performance improvements over previous work on MPM by addressing inefficiencies in the implementation and incorporating a more powerful decoder. We compare several pre-training tasks and introduce new reconstruction methods that utilize conditional generative models without data tokenization or discretization. We show that these new methods outperform the tokenized learning objective from the original MPM on a new test bed for foundation models for jets, which includes using a wide variety of downstream tasks relevant to jet physics, such as classification, secondary vertex finding, and track identification.

conditional generative models↗

Point cloud-based diffusion models for the Electron-Ion Collider

At high-energy collider experiments, generative models can be used for a wide range of tasks, including fast detector simulations, unfolding, searches of physics beyond the Standard Model, and inference tasks. In particular, it has been demonstrated that score-based diffusion models can generate high-fidelity and accurate samples of jets or collider events. This work expands on previous generative models in three distinct ways. First, our model is trained to generate entire collider events, including all particle species with complete kinematic information. We quantify how well the model learns event-wide constraints such as the conservation of momentum and discrete quantum numbers. We focus on the events at the future Electron-Ion Collider, but we expect that our results can be extended to proton-proton and heavy-ion collisions. Second, previous generative models often relied on image-based techniques. The sparsity of the data can negatively affect the fidelity and sampling time of the model. We address these issues using point clouds and a novel architecture combining edge creation with transformer modules called Point Edge Transformers. Third, we adapt the foundation model OmniLearn, to generate full collider events. This approach may indicate a transition toward adapting and fine-tuning foundation models for downstream tasks instead of training new models from scratch.

Araz, Jack Y. [Stony Brook Univ., NY (United State↗

Generalist multimodal AI: A review of architectures, challenges and opportunities

Multimodal models are expected to be a critical component to future advances in artificial intelligence. Here, this field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural language processing (NLP) and vision. It is widely hoped that further extending the foundation models to multiple modalities (e.g., text, image, video, sensor, time series, graph, etc.) will ultimately lead to generalist multimodal models, i.e. one model across different data modalities and tasks. However, there is little research that systematically analyzes recent multimodal models (particularly the ones that work beyond text and vision) with respect to the underling architecture proposed. Therefore, this work provides a fresh perspective on generalist multimodal models (GMMs) via a novel architecture and training configuration specific taxonomy. This includes factors such as Unifiability, Modularity, and Adaptability that are pertinent and essential to the wide adoption and application of GMMs. The review further highlights key challenges and prospects for the field and guide the researchers into the new advancements.

Artificial intelligence (AI)↗

Modeling the behavior of concentrated aqueous HNO 3 using machine learning interatomic potentials

We develop two multi-defect machine learning interatomic potentials (MLIPs) trained at the BLYP-D2 and PBE-D3 density functional theories using the DeepMD-kit, allowing for the investigation of structural and thermodynamic properties of nitric acid over a wide range of concentrations via molecular dynamics (MD) simulations. We directly compute the degree of dissociation, α, and pK a from MD simulations, revealing that HNO 3 behaves as a weaker acid at higher concentrations, noting that our standard-state pK a value is in excellent agreement with the experimental one. In general, good agreement is observed with experimental results such as α and density outside the training dataset, with only modest deviations at low-to-medium concentrations. We benchmark our custom multi-defect DeepMD MLIPs against foundational models MACE-MP0 and MACE-OFF23. The foundation models capture some aspects of HNO 3 /NO 3 − solvation in concentrated nitric acid but show noticeable density errors and miss subtle structural features relevant to spectroscopy, whereas the bespoke DeepMD MLIPs yield more compact solvation shells, reproduce density-concentration trends, and run ∼12–15× faster than MACE-MP0. Although classical FFs are still more efficient and match experimental densities better, they lack chemical reactivity and thus cannot predict α or pK a , underscoring the need for system-specific reactive MLIPs beyond universal MLIPs.

Dinpajooh, Mohammadhasan [Pacific Northwest Nation↗

SPUS-Small-PDE-U-net-Solver

Small PDE U-Net Solver (SPUS) is a compact and efficient foundation model (FM) designed as a unified neural operator for solving a wide range of partial differentialequations (PDEs). SPUS leverages a lightweight residual U-Net-based architecture as a foundation model architecture. To enable effective learning in this minimalist framework, SPUS utilizes a simple yet powerful auto-regressive pretraining strategy which closely replicates the behavior of numerical solvers to learn the underlying physics. SPUS is designed to be pretrained on a diverse set of fluid dynamics PDEs from public benchmark datasets.

Siddik, Abu↗

MPEX AI Digital Twins

All magnetically confined plasma fusion power plant concepts (Tokamak, Spherical Tokamak, Stellarator, Mirror, ...) must exhaust the heat and plasma from the core confinement region to the material walls. The primary channel for this exhaust is through a plasma divertor which directs plasma along open magnetic field lines to a material target. The Material Plasma Exposure eXperiment (MPEX) illustrated in Figure 1, is a high-power, steady-state linear plasma device designed to produce the plasma material interaction (PMI) conditions of the divertor of future magnetic confinement fusion power plants: energy flux 20MW/m 2 , ion fluence 1031/m 2 , pulse duration 106 sec. These goals of plasma exposure in MPEX are well beyond those achieved in magnetic fusion experimental devices. Successfully achieving these high power steady state conditions for long pulses requires operational control of the heating and particle sources and the plasma flux to the walls and target. The MPEX AI Hot Spot Controller, proposed in this project, will help achieve the operational milestones of MPEX. The MPEX device will begin commissioning at the end of FY26. A smaller proto-MPEX was operated for 14,666 plasma discharges and will resume operation in September of 2025 as proto-MPEX-lite, with reduced capability, to test a new window for the Helicon plasma source. The proto-MPEX data has undergone surrogate modeling with machine learning methods (R. Archibald, 2022 IEEE International Conference on Big Data). This proto-MPEX data will be used to begin development of the AI digital twins described in this white paper. The scientific mission of MPEX is to qualify materials of different composition for use in the high energy and plasma flux conditions of a fusion power plant. The materials exposed in MPEX will in some cases be exposed to high neutron fluxes at other ORNL facilities to measure the changes to their PMI properties. The targets exposed in MPEX will be transported under vacuum to a Surface Analysis Station (SAS). The SAS will be equipped with the following diagnostics: Focused Ion Beam (FIB) for trench milling, 100-400 angstrom resolution scanning electron microscope (SEM), surface mapping x-ray spectrometer, high resolution camera, and a future upgrade to a laser induced breakdown spectroscopy quadruple mass spectrometer (LIBS-QMS). The MPEX experiments will generate diverse pre- and post-exposure measurement data of detailed material properties down to the crystal grain level in 3D for post-exposure assessment of PMI damage (e.g. cracking, melting, erosion and redeposition of the material). Physics models for the PMI, and how the material composition and manufacturing impact its performance under high energy plasma exposure, need to be validated with MPEX data to guide the selection of new candidate materials. Our vision for the MPEX AI Digital Twins project is to supply experimental and physics model simulation data to train Artificial Intelligence (AI) models for data processing, analysis, operational control, PMI and materials simulation to maximize the scientific output of the MPEX device. Ultimately, an AI digital twin of MPEX material assessment metrics for tested and synthetic material types with simulated PMI will be trained by the AI Modeling Teams on the experimental and physics simulation data submitted to the American Science Cloud by this project. A purely empirical search for the best material is inefficient given the finite number of samples that can be tested on MPEX. In order to expand the material properties database for training the MPEX Material Assessment AI Digital Twin, and to gain physics understanding of the PMI processes, physics models of the material properties and PMI processes are required. The physics simulations provide detailed simulation data, like impact angles for plasma ions, sputtering yields, transport of the ionized sputtered target material in the plasma, and redeposition locations. This simulation data expands the measurement data for deeper physics understanding. The experimental data is essential to validate the PMI and material structure simulation models. The validated models can then be used to generate new simulation data of MPEX material assessments for synthetic material compositions that have not been exposed in MPEX. These predictive simulations, plus the whole experimental dataset, will be used to train the MPEX Material Assessment AI Digital Twin allowing a rapid generative AI search for new materials with reduced PMI damage by interpolating the domain of the training set. These new optimum materials can be simulated with the physics codes and/or tested in MPEX. The ability of AI neural networks to interpolate multi-dimensional parameter spaces and generate virtual data is exploited for a more efficient search for optimum materials. The advent of the Transformational AI Models Consortium (TAIMC) is an opportunity to engage with state of the art private and public AI developers to achieve the goals of the AI digital twins and AI accelerated physics models proposed in this project. Our partners at ORNL from the Advance Scientific Computing Research (ASCR) organization will collaborate in accelerating the integrated plasma material interaction simulation framework. This simulation framework will provide a platform for generating simulation data across a range of physical fidelities, including hybrid methods that produce multi-fidelity results. This data will be leveraged for AI model development, both for generation of surrogates and the automation of simulation campaigns. A part of the research below will include collaborative efforts with the TAIMC to (i) adapt data storage approaches to ensure AI-readiness, (ii) provide a protypical exemplar to inform and exercise constructed workflows, and (iii) generate and share data, using the TAIMC unified AI data standard, for foundational models that will be trained from multiple sources across the DOE complex. We will also collaborate with the TAIMC, as well as the planned AI modeling teams, to develop approaches for reducing the cost of data generation. These include tailored multi-fidelity approaches as well as fine-tuning strategies to augment general, large-scale foundational models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Compressing Vision Transformers in Geospatial Transfer Learning with Manifold-Constrained Optimization

Deploying geospatial foundation models on resource-constrained edge devices demands compact architectures that maintain high downstream performance. However, their large parameter counts and the accuracy loss often induced by compression limit practical adoption.In this work, we leverage manifold-constrained optimization framework DLRT to compress large vision transformer–based geospatial foundation models during transfer learning. By enforcing structured low-dimensional parameterizations aligned with downstream objectives, this approach achieves strong compression while preserving task-specific accuracy. We show that the method outperforms of-the-shelf low-rank methods as LoRA. Experiments on diverse geospatial benchmarks confirm substantial parameter reduction with minimal accuracy loss, enabling high-performing, on-device geospatial models.

Snyder, Thomas [Yale University]↗

SAM-I-Am: Semantic boosting for zero-shot atomic-scale electron micrograph segmentation

Image segmentation is a critical enabler for tasks ranging from medical diagnostics to autonomous driving. However, the correct segmentation semantics — where are boundaries located? what segments are logically similar? — change depending on the domain, such that state-of-the-art foundation models can generate meaningless and incorrect results. Moreover, in certain domains, fine-tuning and retraining techniques are infeasible: obtaining labels is costly and time-consuming; domain images (micrographs) can be exponentially diverse; and data sharing (for third-party retraining) is restricted. To enable rapid adaptation of the best segmentation technology, we propose the concept of semantic boosting: given a zero-shot foundation model, guide its segmentation and adjust results to match domain expectations. Here, we apply semantic boosting to the Segment Anything Model (SAM) to obtain microstructure segmentation for transmission electron microscopy. Our booster, SAM-I-Am, serves as a post-processing engine that extracts geometric and textural features of various intermediate masks to perform mask removal and mask merging operations. We demonstrate a zero-shot performance increase of (absolute) +21.35%, +12.6%, +5.27% in mean IoU, and a -9.91%, -18.42%, -4.06% drop in mean false positive masks across images of three difficulty classes over vanilla SAM (ViT-L).

36 MATERIALS SCIENCE↗

V-HAMSTeR v1.0.0

V-HAMSTeR is a bioinformatics software tool designed to predict the hosts of viruses directly from genomic sequences. It can be used by researchers to predict animal, prokaryotic, plant, protist or fungal viral hosts including viruses that may be fragmented or discovered in environmental metagenomic datasets. Features & Uses: The software employs a novel dual-stream deep learning architecture that dynamically fuses implicit sequence embeddings from a genomic foundation model with 13 explicit, handcrafted biological features (e.g., coding density and strand switch rates). To ensure maximum reliability, V=HAMSTeR deploys a 5-fold deep ensemble calibrated via Joint Temperature Scaling, providing users with statistically rigorous confidence probabilities. It also features an automated sequence chunking and mean-pooling module to seamlessly process variable-length contigs. Advantages Over Similar Technologies: Existing tools (e.g., IPEV, RNAVirHost) typically rely on either basic k-mers or isolated neural networks. V-HAMSTeR's hybrid architecture captures both broad genomic context and specific biological motifs that standalone foundation models often miss. Furthermore, unlike competitor tools that struggle with incomplete data or exhibit extreme overconfidence, V-HAMSTeR is explicitly benchmarked and mathematically calibrated for fragmented assemblies (1kb–10kb). This makes it uniquely robust, accurate, and trustworthy for the messy reality of real-world environmental viromics.

Grigson, Susie [Lawrence Berkeley National Laborat↗

Cumulus clouds - Interactions between laboratory experiments and observations as foundations for models

Early Woods Hole cumulus observations conducted with the aid of an aircraft suggested that buoyancy dilution by entrainment was a major brake upon tropical cumulus growth. The mechanism by which entrainment occurred, however, was not well understood. Ludlam and Scorer (1953) postulated that buoyant bubbles were the building blocks of cumulus clouds and that it was aerodynamic drag which caused the tops to cease rising. The present investigation is concerned with laboratory experiments and analyses which have been conducted to clarify remaining questions. Attention is given to bubbles in water of uniform density, bubbles released into stably stratified fluids, bubbles released into a two-layer fluid, some observational questions, and buoyant plumes and thermals. The considered experiments provide some insight concerning the mechanism involved in the conversion of buoyancy into motion, taking into account a simpler fluid situation.

Simpson, J.↗

Lessons Learned from Medical System Foundation Development for Long-Duration Lunar Orbit and Lunar Surface Missions

The Human Research Program (HRP) Exploration Medical Capability (ExMC) Element has been tasked with the development of Medical System Foundations for Level of Care IV for both short-duration lunar orbital missions and, subsequently, long-duration lunar orbital and surface operations missions. These Medical System Foundations serve as a framework to aid in early medical system design and mission planning. The content of both Foundation models is similar, consisting of a concept of operations, functional decomposition, clinical content (medical conditions, capabilities, and resources), technical requirements (interface, non-functional, and functional), and traces between these components and to the NASA standards documents and parent-level (Program- and Vehicle habitat system-level) requirements. Additionally, the development of both Foundations employed systems engineering principles and a model-based systems engineering (MBSE) approach. Throughout the development of these Foundations, ExMC has strived to improve the efficiency and robustness of its processes and to be more responsive to change (i.e., in design reference mission parameters and assumptions) and to stakeholders’ feedback. The most significant improvements made between the short- and long-duration Foundation models during this transformation process are the following: • Replacement of the traditional document-based ConOps with a model-based ConOps according to MBSE principles, which facilitated more efficient understanding of the material and the consolidation of all relevant information into a centralized location. • Utilization of an agile approach with tasks organized into sprints. This approach enabled solicitation of more frequent usability feedback from stakeholders, incorporation of more human factors reviews into the sprints, and more efficient tasking of team members. This presentation will discuss the journey of developing both Foundation models, as well as the lessons learned and resulting improvements made between the Short- and Long-Duration models.

M Kaetzer↗