Value of Use Cases in AI: Insights from a Foundation Model Approach
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Foundation Models (FM) are generalized Artificial Intelligence (AI) models that are designed to replace a task or an application-specific model and can be used for many downstream applications. These FM can be built on any sequence data and are trained utilizing self-supervised approaches. The obstacle of creating a sizable labeled dataset for training is removed by using self-supervised learning. Most FM employ transformer design that takes advantage of the idea of self-attention, allowing the network to represent the impact of distant data points on one another in space and time. The FM models show emergent qualities that are induced from the data. FM can become a valuable tool for Earth science researchers. Due to the size of these models, downstream applications built fine-tuning these FM perform better and exhibit greater accuracy than models created from scratch. FM significantly lowers the entry barrier in terms of both the time and effort required to develop various downstream applications. For some scientific datasets, such as optical remote sensing data, FM can speed up processes like classification, object detection and prediction. By eliminating the training data bottleneck and maximizing the usage of science data, FM can make it simpler to integrate AI into scientific research. Initial results for three different FMs will be presented.
Foundation Models (FM) are AI models that are designed to replace a task or an application specific model. These FM can be applied to many different downstream applications. These FM are trained using self supervised techniques and can be built on any type of sequence data. The use of self supervised learning removes the hurdle for developing a large labeled dataset for training. Most FM use transformer architecture utilizes the notion of self attention which allows the network to model the influence of distant data points to each other both in space and time. The FM models exhibit emergent properties that are induced from the data. FM can be an important tool for science. The scale of these models results in better performance for different downstream applications and these applications show better accuracy over models built from scratch. FM drastically reduces the cost of entry to build different downstream applications both in time and effort. FM for selected science datasets such as optical satellite data, can accelerate applications ranging from data quality monitoring, feature detection and prediction. FM can make it easier to infuse AI into scientific research by removing the training data bottleneck and increasing the use of science data.
Rodents are foundational model organisms often utilized due to their seemingly analogous morphologies and biological responses to humans. However, recent studies have demonstrated that murine model data are limited in their applicability, particularly in inflammatory disease. In space studies, accurately predicting human response from mouse data is critical due to extreme limiting factors in both rodent and human spaceflight research. With successful prediction, spaceflight ailments can be predicted and prevented while respecting the constraints of the spaceflight industry and minimizing danger to humans. To do so, novel methodologies must be developed that predict human response from murine data after considering biological differences between rodents and humans in spaceflight. After considering terrestrial models, we determined that a spaceflight-based expression profiting translation tool should be created to accurately capture predictions of human gene expression in spaceflight from mouse data. To prepare to build this model, we organized known human spaceflight risks, chose analog human diseases as training data categories, then identified existing RNASeq disease datasets from GEO as potential training data. In addition, we classified existing Genelab mouse differential gene expression datasets for use as experimental data.
Chromium represents a significant challenge for the vitrification of high-level nuclear waste into silicate and borosilicate glasses due to its low solubility and variable oxidation states, which can limit the waste loading due to promotion of crystallization or phase separation during processing. In this study, we modeled chromium containing silicate glasses using molecular dynamics simulations with three machine learning interatomic potentials (MLIPs), MACE, CHGNet, and PFP were employed, to gain insights on glass composition and oxidation states on the structures of these glasses. One of the goals is to evaluate their ability of these MLIPs to accurately represent the general structure of silicate glasses and chromium local environments as a function of chromium oxidation states. Density Functional Theory (DFT) based calculations and experimental data such as neutron structure factors were used to validate the structural models. It was found that the foundation models of all three MLIPs are able to reproduce general structural features of the sodium silicate glass structure consistent with experimental and DFT data, but only CHGNet and PFP can accurately capture the oxidation states and local environment of chromium: tetrahedral for Cr6+ and octahedral for Cr3+. Furthermore, we studied the effect of varying Cr3+/ Cr6+ (Cr3+/Crtotal) ratio and total chromium content using PFP. Our results show that Cr6+ enhances network polymerization by reducing non-bridging oxygens through Na? charge compensation required due to the formation of chromate (CrO42-) species, while Cr³? acts as a network modifier that disrupts connectivity. System size effects on the structural characteristics and chromium environments were also tested using the PFP potential. This work highlights the importance of careful validation on the precision, transferability, and potential of MLIPs for modeling glasses containing transition metal elements that can exist in multiple oxidation states. It is also encouraging to see the foundational models are all three MLFFs are able to reproduce the basic sodium silicate glass structures, while suggesting additional training or refining is needed to improve the description of more complex systems containing transition metals.
In this work, we detail two uncertainty quantification (UQ) methods that provide complementary information. Readout ensembling, by finetuning only the readout layers of an ensemble of foundation models, provides information about model uncertainty. Amending the final readout layer to predict upper and lower quantiles replaces point predictions with distributional predictions, which provide information about uncertainty within the underlying training data. We demonstrate our approach with the MACE-MP-0 model, applying UQ to both the foundation model and a series of finetuned models. The uncertainties produced by the ensemble and quantile methods are demonstrated to be distinct measures by which the quality of the NNP output can be judged.
Atomistic materials modeling is a critical task with wide-ranging applications, from drug discovery to materials science, where accurate predictions of the target material property can lead to significant advancements in scientific discovery. Graph Neural Networks (GNNs) represent the state-of-the-art approach for modeling atomistic material data thanks to their capacity to capture complex relational structures. While machine learning performance has historically improved with larger models and datasets, GNNs for atomistic materials modeling remain relatively small compared to large language models (LLMs), which leverage billions of parameters and terabyte-scale datasets to achieve remarkable performance in their respective domains. To address this gap, we explore the scaling limits of GNNs for atomistic materials modeling by developing a foundational model with billions of parameters, trained on extensive datasets in terabytescale. Our approach incorporates techniques from LLM libraries to efficiently manage large-scale data and models, enabling both effective training and deployment of these large-scale GNN models. This work addresses three fundamental questions in scaling GNNs: the potential for scaling GNN model architectures, the effect of dataset size on model accuracy, and the applicability of LLM-inspired techniques to GNN architectures. Specifically, the outcomes of this study include (1) insights into the scaling laws for GNNs, highlighting the relationship between model size, dataset volume, and accuracy, (2) a foundational GNN model optimized for atomistic materials modeling, and (3) a GNN codebase enhanced with advanced LLM-based training techniques. Our findings lay the groundwork for large-scale GNNs with billions of parameters and terabyte-scale datasets, establishing a scalable pathway for future advancements in atomistic materials modeling.
This paper discusses the key challenges and future research directions for privacy-preserving federated learning (PPFL), with a focus on its application to large-scale scientific AI models, in particular, foundation models~(FMs). PPFL enables collaborative model training across distributed datasets while preserving privacy-- an important collaborative approach for science. We discuss the need for efficient and scalable algorithms to address the increasing complexity of FMs, particularly when dealing with heterogeneous clients. In addition, we underscore the need for developing advance privacy-preserving techniques, such as differential privacy, to balance privacy and utility in large FMs emphasizing fairness and incentive mechanisms to ensure equitable participation among heterogeneous clients. Finally, we emphasize the need for a robust software stack supporting scalable and secure PPFL deployments across multiple high-performance computing facilities. We envision that PPFL would play a crucial role to advance scientific discovery and enable large-scale, privacy-aware collaborations across science domains.
In this work, we significantly enhance masked particle modeling (MPM), a self-supervised learning scheme for constructing highly expressive representations of unordered sets relevant to developing foundation models for high-energy physics. In MPM, a model is trained to recover the missing elements of a set, a learning objective that requires no labels and can be applied directly to experimental data. We achieve significant performance improvements over previous work on MPM by addressing inefficiencies in the implementation and incorporating a more powerful decoder. We compare several pre-training tasks and introduce new reconstruction methods that utilize conditional generative models without data tokenization or discretization. We show that these new methods outperform the tokenized learning objective from the original MPM on a new test bed for foundation models for jets, which includes using a wide variety of downstream tasks relevant to jet physics, such as classification, secondary vertex finding, and track identification.