Big Microstructure Datasets for Materials Informatics: Using Statistically Conditioned Generative Models to Curate Big Datasets
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Abstract not provided.
Computational tools provide a unique opportunity to study and design optimal materials by enhancing our ability to comprehend the connections between their atomistic structure and functional properties. However, designing materials with tailored functionalities is complicated due to the necessity to integrate various computational-chemistry software (not necessarily compatible with one another), the heterogeneous nature of the generated data, and the need to explore vast chemical and parameter spaces. The latter is especially important to avoid bias in scattered data points-based models and derive statistical trends only accessible by systematic datasets. Here, we introduce a robust high-throughput multi-scale computational infrastructure coined MISPR (Materials Informatics for Structure–Property Relationships) that seamlessly integrates classical molecular dynamics (MD) simulations with density functional theory (DFT). By enabling high-performance data analytics and coupling between different methods and scales, MISPR addresses critical challenges arising from the needs of automated workflow management and data provenance recording. The major features of MISPR include automated DFT and MD simulations, error handling, derivation of molecular and ensemble properties, and creation of output databases that organize results from individual calculations to enable reproducibility and transparency. In this work, we describe fully automated DFT workflows implemented in MISPR to compute various properties such as nuclear magnetic resonance chemical shift, binding energy, bond dissociation energy, and redox potential with support for multiple methods such as electron transfer and proton-coupled electron transfer reactions. The infrastructure also enables the characterization of large-scale ensemble properties by providing MD workflows that calculate a wide range of structural and dynamical properties in liquid solutions. MISPR employs the methodologies of materials informatics to facilitate understanding and prediction of phenomenological structure–property relationships, which are crucial to designing novel optimal materials for numerous scientific applications and engineering technologies.
Although the convergence of high-performance computing, automation, and machine learning has significantly altered the materials design timeline, transformative advances in functional materials and acceleration of their design will require addressing the deficiencies that currently exist in materials informatics, particularly a lack of standardized experimental data management. The challenges associated with experimental data management are especially true for combinatorial materials science, where advancements in automation of experimental workflows have produced datasets that are often too large and too complex for human reasoning. The data management challenge is further compounded by the multimodal and multi-institutional nature of these datasets, as they tend to be distributed across multiple institutions and can vary substantially in format, size, and content. Furthermore, modern materials engineering requires the tuning of not only composition but also of phase and microstructure to elucidate processing–structure–property–performance relationships. To adequately map a materials design space from such datasets, an ideal materials data infrastructure would contain data and metadata describing (i) synthesis and processing conditions, (ii) characterization results, and (iii) property and performance measurements. In this work, we present a case study for the low-barrier development of such a dashboard that enables standardized organization, analysis, and visualization of a large data lake consisting of combinatorial datasets of synthesis and processing conditions, X-ray diffraction patterns, and materials property measurements generated at several different institutions. While this dashboard was developed specifically for data-driven thermoelectric materials discovery, we envision the adaptation of this prototype to other materials applications, and, more ambitiously, future integration into an all-encompassing materials data management infrastructure.
Small Punch Test (SPT) uses a thin disk of material to predict mechanical properties. While SPT has existed for decades, it has been used largely as a qualitative evaluator of mechanical properties. Recent advances in computational modeling have enabled the extraction of uniaxial stress-strain response from the measured SPT load-displacement data. Due to small sample volumes and unidirectional testing, SPT is conducive to high-throughput automation and ideally suited to extract properties from high-cost materials. Aerospace alloys have been of recent interest to the Additive Manufacturing (AM) community due to AM’s unique ability to fabricate complex designs not possible, or extremely arduous, with conventional manufacturing. In this research, SPT, coupled with Materials Informatics and computational modeling, is used to develop relevant Process-Structure-Property relationships to decrease the cost and time of process optimization for AM aerospace alloys, namely Inconel 718, Inconel 625, and Niobium C103.
A materials informatics framework to explore a large number of candidate van der Waals (vdW) materials is developed. In particular, in this study a large space of monolayer transition metal halides is investigated by combining high-throughput density functional theory calculations and artificial intelligence (AI) to accelerate the discovery of stable materials and the prediction of their magnetic properties. Here, the formation energy is used as a proxy for chemical stability. Semi-supervised learning is harnessed to mitigate the challenges of sparsely labeled materials data in order to improve the performance of AI models. This approach creates avenues for the rapid discovery of chemically stable vdW magnets by leveraging the ability of AI to recognize patterns in data, to learn mathematical representations of materials from data and to predict materials properties. Using this approach, previously unexplored vdW magnetic materials with potential applications in data storage and spintronics are identified.
Two-dimensional layered materials, such as transition metal dichalcogenides (TMDs), possess an intrinsic van der Waals gap at the layer interface, allowing for remarkable tunability of the optoelectronic features via external intercalation of foreign guests such as atoms, ions, or molecules. Herein, we introduce a high-throughput, data-driven computational framework for the design of novel quantum materials derived from intercalating planar conjugated organic molecules into bilayer transition metal dichalcogenides and dioxides. By combining first-principles methods, material informatics, and machine learning, we characterize the energetic and mechanical stability of this new class of materials and identify the fifty (50) most stable hybrid materials from a vast configurational space comprising ∼105 materials, employing intercalation energy as the screening criterion.
With the increased emphasis on reducing the cost and time to market of new materials, the need for analytical tools that enable the virtual design and optimization of materials throughout their processing - internal structure - property - performance envelope, along with the capturing and storing of the associated material and model information across its lifecycle, has become critical. This need is also fueled by the demands for higher efficiency in material testing; consistency, quality and traceability of data; product design; engineering analysis; as well as control of access to proprietary or sensitive information. Consequently, at NASA Glenn Research Center a robust information management system that manages the digital thread across the full material life (i.e., capture, analysis, maintenance, and dissemination of data) cycle directed at the design of ‘fit-for-purpose materials’ is under development. To this end the Application Table has been incorporated within NASA Glenn Research Center’s ICME Information Management framework within the ANSYS Granta MI tool. The Application Table provides a place where material and structural application information/requirements can be linked to marry the “design-the-material” (structural engineering) and the “design-with-material” (material science) paradigms and thereby enable application-driven design and optimization of materials and structures. In additional several associated toolsets, specifically: AIMAOS (Automated Information Management Across Organizations and Scales), Py MILab, and JARIMIS (Just A Rather Intelligent Material Interrogation System) are also under development to assist in the judicious automation of this process. AIMOAS offers users an interactive graphical user interface for connecting material information management systems with both commercial and in-house simulation tools at various length scales to enable such automation in the handoff across scales and maintenance of material digital twins and the digital thread. Py MILab, is an automatic framework for the capture, analysis, maintenance, and storage of material test data. Py MILab uses a modular approach for capturing raw data, analyzing the data, and storing the data in a database, interfaced by neutral file structures, to promote plug-and-play capabilities for various analysis types. Finally, JARIMIS is an expert system that integrates various materials informatics tools (e.g., MicroNet, Surrogate ML models, ANSYS Granta MI, etc.) to enable inverse design of materials and facilitate the application of machine learning (ML) and data science with human in the loop decision making to rapidly discover and optimize new materials.
Nuclear fuel performance is critically dependent on understanding the evolution of fuel properties under operational conditions, a complex challenge driven by chemical changes and substantial radiation damage during fission. Traditionally, property evolution has been determined via empirical data collected following irradiation. However, these empirical correlations are limited in their applicability beyond the specific conditions in which they were obtained. This study explores a novel approach to address this challenge by applying materials informatics to develop a machine learning random forest (ML-RF) model that captures the effects of fission products on fuel compounds. The model predicts formation enthalpy (ΔH f ) by leveraging extensive quantum materials property data and correlating it with material descriptors such as composition, atomic and site features, and crystal lattice properties. This ML-RF model enables rapid interpolation across the compositional and structural spaces covered by the training data, thus supporting high-throughput screening and energetic ranking of candidate phases. The model demonstrates the ability to predict ΔH f with a mean absolute error (MAE) of approximately 0.1 to 0.2 eV/atom across a wide range of compounds, including key nuclear fuel systems (U-O, U-N, U-C, U-Si, and U-Mo). For example, it was used to assess shifts in stoichiometry for UO 2 (O/M) and UN (N/M) fuels, revealing their distinct tendencies in chemical potential variation and enabling preliminary convex hull analyses. Furthermore, the model provides insights into how individual fission products affect fuel properties. Results indicate that larger fission products (e.g., Nd, Pu, Ce) have a more pronounced impact on UO 2 , while lighter ones (e.g., Zr) strongly influence UN. Here, the model developed in this work can be used to support the Accelerated Fuel Qualification approach by facilitating preliminary evaluations prior to extensive materials modeling and experimentation. To this end, the trained model has been made available to the fuel community to support ongoing fuel development efforts.
The interest in high entropy ceramics (HECs) has increased steadily due to their superior properties. However, the prediction of their formation still poses challenges for the discovery of new systems. Here, we discover a rational rule for designing single-phase high entropy metal diborides (HEBs) using data-driven approach. The machine learning (ML) model is trained on data collected via high-throughput experiments (HTEs). K nearest neighbor (KNN) model shows an experimental validation accuracy of 93.75%. By implementing interpretable ML method, we demonstrate that a mismatch of the bonds between boron and transition metals (δ B-TM ) dominates the formation of HEBs. We propose an empirical rule that HEBs favor forming a single phase when δ B-TM < 3.66; otherwise, multiphase. The rule has a high accuracy of 93.33% for new HEBs predictions. In addition, we contribute 165 high quality HEBs data in total, which can promote the development of materials informatics in HEBs. Furthermore, this data-driven strategy can be expanded to accelerate the search for new HECs, paving a pathway to design novel HECs with superior properties rapidly.
This ARPA-E project developed a machine learning tool to use in formulation design of cementitious binders for concrete having 50% less embodied CO 2 and possessing twice the durability compared to concrete based on ordinary portland cement (OPC) binders. The technical focus was on limestone/calcined clay cement (LC3), the leading replacement for OPC. Here, hierarchical machine learning (HML) was used to model the flowability, set time, strength, and durability of LC3 concrete. This methodology identifies latent variables derived from domain knowledge and empirical models that develop an accurate model for a response surface from small datasets. For the flowability metric, particle packing was a dominant factor, while strength and durability were both strongly determined by the fraction of metakaolin and the water:solids ratio. Under constraints of water:binder ratio, material performance metrics, embodied CO 2 , and cost per tonne of OPC, multi-objective optimization was used to design binders parameterized by the mineral composition replacing OPC, particle size distributions, and water:solids ratio. The trained algorithm was able to predict multiple mixes met these performance criteria, and experimental testing validated the predictions. The model demonstrated here is relevant for North America, where pure kaolin deposits are found broadly. The approach is being taken forward into commercial application by Ansatz AI, a materials informatics company founded by PI Washburn and co-PI Poczos. Through collaborations with the cement and concrete industry, and funding from SBIR programs, a commercial software will be developed in future research.
Transparency and reproducibility are important aspects of validation for Machine Learning (ML) models that are not fully understood and applies independently of the application domain.We offer a case study of reproducibility that highlights the challenges encountered when attempting to reproduce analyzes obtained with Machine Learning methods in materials informatics. Our study explores prediction results obtained with ML models and issues in training data serving as input. We discuss challenges related to theory-driven and numerical errors in training data, lack of reproducibility across platforms and versions, and effects of randomness when varying hyperparameters. In addition to model accuracy, a main metric of interest in the ML community, our results show that model sensitivity may be equally important for applying ML in domain applications such a materials science.
This dataset contains submission files and raw output files from high-throughput DFT simulations to analyze the systemic errors in lattice constant, bulk moduli and formation energy predictions for a range of binary and ternary oxides using four exchange correlation functionals (LDA, PBE, PBEsol and vdW-DF-C09). This data was then used as the basis for employing materials informatics methods to predict the expected errors in the lattice constants of the studied compounds. Predicted errors were also used to better the DFT-predicted lattice parameters. Our results emphasize the link between the computed errors and the electron density and hybridization errors of a functional. In essence, these results provide “error bars” for choosing a functional for the creation of high-accuracy, high-throughput datasets as well as avenues for the development of XC functionals with enhanced performance, thereby enabling the accelerated discovery and design of new materials.