Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Materials database”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

ET-AL: Entropy-targeted active learning for bias mitigation in materials data

Growing materials data and data-driven informatics drastically promote the discovery and design of materials. While there are significant advancements in data-driven models, the quality of data resources is less studied despite its huge impact on model performance. In this work, we focus on data bias arising from uneven coverage of materials families in existing knowledge. Observing different diversities among crystal systems in common materials databases, we propose an information entropy-based metric for measuring this bias. To mitigate the bias, we develop an entropy-targeted active learning (ET-AL) framework, which guides the acquisition of new data to improve the diversity of underrepresented crystal systems. We demonstrate the capability of ET-AL for bias mitigation and the resulting improvement in downstream machine learning models. This approach is broadly applicable to data-driven materials discovery, including autonomous data acquisition and dataset trimming to reduce bias, as well as data-driven informatics in other scientific domains.

36 MATERIALS SCIENCE↗

Large language model-driven database for thermoelectric materials

Thermoelectric materials have the ability to convert waste heat into electricity, offering a valuable solution for energy harvesting. However, their widespread use is hindered by low conversion efficiency, the reliance on expensive rare earth elements, and the environmental and regulatory concerns associated with lead-based materials. A fast and cost-effective way to identify highly efficient thermoelectric materials is through data-driven methods. These approaches rely on robust and comprehensive datasets to train models. Although there are several databases on thermoelectric materials, there is still a need to collect and integrate experimental data from peer-reviewed research articles to capture diverse compositions and properties of materials. Here, in this work, we developed a comprehensive database of 7,123 thermoelectric compounds, containing key information such as chemical composition, structural detail, seebeck coefficient, electrical and thermal conductivity, power factor, and figure of merit (ZT). We used the GPTArticleExtractor workflow, powered by large language models (LLM), to extract and curate data automatically from the scientific literature published in Elsevier journals. This process enabled the creation of a structured database that addresses the challenges of manual data collection. The open access database could stimulate data-driven research and advance thermoelectric material analysis and discovery.

Database↗

High-throughput calculations of charged point defect properties with semi-local density functional theory—performance benchmarks for materials screening applications

Abstract Calculations of point defect energetics with Density Functional Theory (DFT) can provide valuable insight into several optoelectronic, thermodynamic, and kinetic properties. These calculations commonly use methods ranging from semi-local functionals with a-posteriori corrections to more computationally intensive hybrid functional approaches. For applications of DFT-based high-throughput computation for data-driven materials discovery, point defect properties are of interest, yet are currently excluded from available materials databases. This work presents a benchmark analysis of automated, semi-local point defect calculations with a-posteriori corrections, compared to 245 “gold standard” hybrid calculations previously published. We consider three different a-posteriori correction sets implemented in an automated workflow, and evaluate the qualitative and quantitative differences among four different categories of defect information: thermodynamic transition levels, formation energies, Fermi levels, and dopability limits. We highlight qualitative information that can be extracted from high-throughput calculations based on semi-local DFT methods, while also demonstrating the limits of quantitative accuracy.

36 MATERIALS SCIENCE↗

Itinerant Magnetism in Hydride-Synthesized CaCo 12 B 6

A new compound in the underexplored Ca–Co–B phase space has been discovered, validating high-throughput computations from the Open Quantum Materials Database, which predicted thermodynamic stability for CaCo 12 B 6 in the SrNi 12 B 6 structure type. The synthetic effects of different boron precursors and the advantages of using CaH 2 instead of Ca metal were demonstrated by the short synthesis duration and high purity of CaCo 12 B 6 , in contrast with traditional synthesis routes. Powder X-ray diffraction (PXRD) confirmed that CaCo 12 B 6 shares the SrNi 12 B 6 structure ( R $\bar{3}$m (#166), a = 9.469(4) Å, c = 7.468(2) Å, Z = 3) and is water- and air-stable. High-temperature in situ PXRD indicates that CaCo 12 B 6 is stable below 1050 K under vacuum in a silica capillary. CaCo 12 B 6 decomposes between 693 and 773 K during spark-plasma sintering. Density functional theory calculations indicate that CaCo 12 B 6 is metallic with a ferromagnetic ground state. X-ray absorption near-edge spectroscopy and Bader charge analysis indicate that Co atoms in CaCo 12 B 6 lack ionic character. Magnetometry reveals room-temperature paramagnetism with μ eff = 1.7(1)μ B per Co atom and a Weiss constant of +190(10)K. Ferromagnetic ordering occurs below 172(1)K, resulting in a saturation moment of 0.46 μ B per Co atom. Our findings demonstrate that the hydride route is a viable strategy for discovery of new ternary alkaline-earth-transition metal borides analogous to rare-earth-containing counterparts.

diffraction↗

Dataset describing two reference models for full-spectral lighting and daylight simulations together with implementations for two software systems

A dataset of two spectral lighting simulation reference models - one office and one factory hall - is presented. It aims to demonstrate and support full-spectral daylight and electric lighting simulations and facilitate evaluation of non-visual effects of light. The dataset includes Rhino CAD geometry, comprehensive spectral material and light source data and window system BSDF data. Example implementations in the two software tools, Radiance and OWL, enable reproducible workflows and support adoption in other software. The dataset is openly available on Zenodo. The office model reproduces Room 518 at the University of Innsbruck, including a west-facing façade and interior furnishings. The factory hall model follows the proposed geometry in the European standard 15193 for building energy performance. Interior reflectances in the office were measured in-situ using a handheld spectrometer. Exterior spectra and factory hall materials matching specified reflectances were obtained from an online spectral materials database. Glazing transmittance was derived from IGDB data using LBNL Optics/WINDOW. BSDFs for venetian blinds at various tilt angles, and for a diffusing pane adapted from the Complex Glazing Database, were generated in WINDOW. Luminaires in both models are specified with photometric files (Eulumdat/IES) and lamp spectra (Fluorescent 840, 4000 K LED). The provided example implementations (Radiance, OWL) include prepared input data and scripts to run first spectral simulations; example results are also included. The dataset is prepared to support reuse by researchers, designers and software developers for method validation, software engineering and comparison, and development of spectral metrics and controls.

Geisler-Moroder, David↗

An integrated online radioassay data storage and analytics tool for nEXO

Large-scale low-background detectors are increasingly used in rare-event searches as experimental collaborations push for enhanced sensitivity. However, building such detectors, in practice, creates an abundance of radioassay data especially during the conceptual phase of an experiment when hundreds of materials are screened for radiopurity. A tool is needed to manage and make use of the radioassay screening data to quantitatively assess detector design options. We have developed a Materials Database Application for the nEXO experiment to serve this purpose. Furthermore, this paper describes this database application, explains how it functions, and discusses how it streamlines the design of the experiment.

47 OTHER INSTRUMENTATION↗

High-Throughput Screening for Boride Superconductors

A high-throughput screening using density functional calculations is performed to search for stable boride superconductors from the existing materials database. The workflow employs the fast frozen-phonon method as the descriptor to evaluate the superconducting properties quickly. Twenty-three stable candidates were identified during the screening. The superconductivity was obtained earlier experimentally or computationally for almost all found binary compounds. Previous studies on ternary borides are very limited. Here our extensive search among ternary systems confirmed superconductivity in known systems and found several new compounds. Among these discovered superconducting ternary borides, TaMo 2 B 2 shows the highest superconducting temperature of ∼12 K. Most predicted compounds were synthesized previously; therefore, our predictions can be examined experimentally. Our work also demonstrates that the boride systems can have diverse structural motifs that lead to superconductivity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

Wide-ranging predictions of new stable compounds powered by recommendation engines

The computational search for new stable inorganic compounds is faster than ever, thanks to high-throughput density functional theory (DFT). However, stable compound searches remain highly expensive because of the enormous search space and the cost of DFT calculations. To aid these searches, recommendation engines have been developed. We conduct a systematic comparison of the performance of previously developed recommendation engines, specifically ones based on elemental substitution, data mining, and neural network prediction of formation enthalpy. After identifying ways to improve the recommendation engines, we find the neural network to be superior at recommending stable Heusler compounds. Armed with improved recommendation engines, we identify tens of thousands of compounds that are stable at zero temperature and pressure, now available in the Open Quantum Materials Database. We summarize this diverse pool of compounds, including the elusive mixed anion compounds, and two of their many applications: thermoelectricity and solar thermochemical fuel production.

Science & Technology - Other Topics↗

OPTIMADE, an API for exchanging materials data

Abstract The Open Databases Integration for Materials Design (OPTIMADE) consortium has designed a universal application programming interface (API) to make materials databases accessible and interoperable. We outline the first stable release of the specification, v1.0, which is already supported by many leading databases and several software packages. We illustrate the advantages of the OPTIMADE API through worked examples on each of the public materials databases that support the full API specification.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A representation-independent electronic charge density database for crystalline materials

Abstract In addition to being the core quantity in density-functional theory, the charge density can be used in many tertiary analyses in materials sciences from bonding to assigning charge to specific atoms. The charge density is data-rich since it contains information about all the electrons in the system. With the increasing prevalence of machine-learning tools in materials sciences, a data-rich object like the charge density can be utilized in a wide range of applications. The database presented here provides a modern and user-friendly interface for a large and continuously updated collection of charge densities as part of the Materials Project. In addition to the charge density data, we provide the theory and code for changing the representation of the charge density which should enable more advanced machine-learning studies for the broader community.

36 MATERIALS SCIENCE↗

A representation-independent electronic charge density database for crystalline materials

In addition to being the core quantity in density functional theory, the charge density can be used in many tertiary analyses in materials sciences from bonding to assigning charge to specific atoms. The charge density is data-rich since it contains information about all the electrons in the system. With increasing utilization of machine-learning tools in materials sciences, a data-rich object like the charge density can be utilized in a wide range of applications. The database presented here provides a modern and user-friendly interface for a large and continuously updated collection of charge densities as part of the Materials Project. In addition to the charge density data, we provide the theory and code for changing the representation of the charge density which should enable more advanced machine-learning studies for the broader community.

36 MATERIALS SCIENCE↗

Deep learning of electrochemical CO 2 conversion literature reveals research trends and directions

Large-scale and openly available material science databases are mainly composed of computer simulation results rather than experimental data. Some examples include the Materials Project, Open Quantum Materials Database, and Open Catalyst 2022. Unfortunately, building large-scale experimental databases remains challenging due to the difficulties in consolidating locally distributed datasets. In this work, focusing on the catalysis literature of CO 2 reduction reactions (CO 2 RRs), we present a machine learning (ML)-based protocol for selecting highly relevant papers and extracting important experimental data. First, we report a document embedding method (Doc2Vec) for collecting papers of greatest relevance to the specific target domain, which yielded 3154 CO 2 RR-related papers from six publishers. Next, we developed named entity recognition (NER) models to extract twelve entities related to material names (catalyst, electrolyte, etc.) and catalytic performance (Faradaic efficiency, current density, etc.). Further, among several tested models, the MatBERT-based approach achieved the highest accuracy, with an average F1-score of 90.4% and an F1-score of 95.2% in a boundary relaxation evaluation scheme. The accurate and accelerated NER-based data extraction from a large volume of catalysis literature enables temporal trend analyses of the CO 2 RR catalysts, products, and performances, revealing the potentially effective material space in CO 2 RRs. While this work demonstrates the effectiveness of our ML-based text mining methods for specifically CO 2 RR literature, the methods and approach are applicable to and may be used to accelerate the development of other catalytic chemical reactions.

36 MATERIALS SCIENCE↗

Path Forward: Materials Data Modernization for ASME Codes and Standards in the Artificial Intelligence Era

Development of the ASME Materials Properties Database was initiated in the early 2010s to support the ASME Codes and Standards. As information technologies advance at an accelerated pace with the artificial intelligence era on the horizon, the ASME Materials Properties Database must be further modernized from a database to a knowledgebase to ride the wave of digital information revolution and effectively support the ASME Codes and Standards in the new era. This paper is intended to provide an overview of the ASME Materials Properties Database and discuss a roadmap for its future development to facilitate understanding of and participation from different sectors of the Codes and Standards community. Further, it first reviews the basic concepts of data, information, knowledge, database, and database system as well as the pros and cons in different types of data management and then discusses the path forward for a desired evolution of the database into a self-explanatory and machine-readable knowledgebase that is consistent with human cognitive processes for the Codes and Standards development and, furthermore, provides resources for data processing and analysis to reach an eventual goal of streamlining the Codes and Standards development from the initial inquiry, throughout data submission, analysis, …, to Codes and Standards rule establishment for final publication.

36 MATERIALS SCIENCE↗

Path Forward: Materials Data Modernization for ASME Codes and Standards in the Artificial Intelligence Era

Development of the ASME Materials Properties Database was initiated in the early 2010s to support the ASME Codes and Standards. As information technologies advance at an accelerated pace with the artificial intelligence era on the horizon, the ASME Materials Properties Database must be further modernized from a database to a knowledgebase to ride the wave of digital information revolution and effectively support the ASME Codes and Standards in the new era.This paper is intended to provide an overview of the ASME Materials Properties Database and discuss a roadmap for its future development to facilitate understanding of and participation from different sectors of the Codes and Standards community. It first reviews the basic concepts of data, information, knowledge, database, and database system; as well as the pros and cons in different types of data management, and then discusses the path forward for a desired evolution of the database into a self-explanatory and machine-readable knowledgebase that is consistent with human cognitive processes for the Codes and Standards development and furthermore provides resources for data processing and analysis to reach an eventual goal of streamlining the Codes and Standards development from the initial inquiry, throughout data submission, analysis, …, to Codes and Standards rule establishment for final publication.

Ren, Weiju↗

hashin_shtrikman_mp: a package for the optimal design and discovery of multi-phase composite materials

hashin_shtrikman_mp is a tool for composites designers who have desired composite properties in mind, but who do not yet have an underlying formulation. The library utilizes the tightest theoretical bounds on the effective properties of composite materials with unspecified microstructure – the Hashin-Shtrikman bounds – to identify candidate theoretical materials, find real materials that are close to the candidates, and determine the optimal volume fractions for each of the constituents in the resulting composite. Its features include (i) leveraging of materials in the Materials Project database, (ii) integration with the Materials Project API, (iii) use of genetic machine-learning, (iv) agnosticism to underlying microstructure, and (v) ultimate engineering application, make it a tool with much broader applications than its predecessors.

97 MATHEMATICS AND COMPUTING↗

MPEX AI Digital Twins

All magnetically confined plasma fusion power plant concepts (Tokamak, Spherical Tokamak, Stellarator, Mirror, ...) must exhaust the heat and plasma from the core confinement region to the material walls. The primary channel for this exhaust is through a plasma divertor which directs plasma along open magnetic field lines to a material target. The Material Plasma Exposure eXperiment (MPEX) illustrated in Figure 1, is a high-power, steady-state linear plasma device designed to produce the plasma material interaction (PMI) conditions of the divertor of future magnetic confinement fusion power plants: energy flux 20MW/m 2 , ion fluence 1031/m 2 , pulse duration 106 sec. These goals of plasma exposure in MPEX are well beyond those achieved in magnetic fusion experimental devices. Successfully achieving these high power steady state conditions for long pulses requires operational control of the heating and particle sources and the plasma flux to the walls and target. The MPEX AI Hot Spot Controller, proposed in this project, will help achieve the operational milestones of MPEX. The MPEX device will begin commissioning at the end of FY26. A smaller proto-MPEX was operated for 14,666 plasma discharges and will resume operation in September of 2025 as proto-MPEX-lite, with reduced capability, to test a new window for the Helicon plasma source. The proto-MPEX data has undergone surrogate modeling with machine learning methods (R. Archibald, 2022 IEEE International Conference on Big Data). This proto-MPEX data will be used to begin development of the AI digital twins described in this white paper. The scientific mission of MPEX is to qualify materials of different composition for use in the high energy and plasma flux conditions of a fusion power plant. The materials exposed in MPEX will in some cases be exposed to high neutron fluxes at other ORNL facilities to measure the changes to their PMI properties. The targets exposed in MPEX will be transported under vacuum to a Surface Analysis Station (SAS). The SAS will be equipped with the following diagnostics: Focused Ion Beam (FIB) for trench milling, 100-400 angstrom resolution scanning electron microscope (SEM), surface mapping x-ray spectrometer, high resolution camera, and a future upgrade to a laser induced breakdown spectroscopy quadruple mass spectrometer (LIBS-QMS). The MPEX experiments will generate diverse pre- and post-exposure measurement data of detailed material properties down to the crystal grain level in 3D for post-exposure assessment of PMI damage (e.g. cracking, melting, erosion and redeposition of the material). Physics models for the PMI, and how the material composition and manufacturing impact its performance under high energy plasma exposure, need to be validated with MPEX data to guide the selection of new candidate materials. Our vision for the MPEX AI Digital Twins project is to supply experimental and physics model simulation data to train Artificial Intelligence (AI) models for data processing, analysis, operational control, PMI and materials simulation to maximize the scientific output of the MPEX device. Ultimately, an AI digital twin of MPEX material assessment metrics for tested and synthetic material types with simulated PMI will be trained by the AI Modeling Teams on the experimental and physics simulation data submitted to the American Science Cloud by this project. A purely empirical search for the best material is inefficient given the finite number of samples that can be tested on MPEX. In order to expand the material properties database for training the MPEX Material Assessment AI Digital Twin, and to gain physics understanding of the PMI processes, physics models of the material properties and PMI processes are required. The physics simulations provide detailed simulation data, like impact angles for plasma ions, sputtering yields, transport of the ionized sputtered target material in the plasma, and redeposition locations. This simulation data expands the measurement data for deeper physics understanding. The experimental data is essential to validate the PMI and material structure simulation models. The validated models can then be used to generate new simulation data of MPEX material assessments for synthetic material compositions that have not been exposed in MPEX. These predictive simulations, plus the whole experimental dataset, will be used to train the MPEX Material Assessment AI Digital Twin allowing a rapid generative AI search for new materials with reduced PMI damage by interpolating the domain of the training set. These new optimum materials can be simulated with the physics codes and/or tested in MPEX. The ability of AI neural networks to interpolate multi-dimensional parameter spaces and generate virtual data is exploited for a more efficient search for optimum materials. The advent of the Transformational AI Models Consortium (TAIMC) is an opportunity to engage with state of the art private and public AI developers to achieve the goals of the AI digital twins and AI accelerated physics models proposed in this project. Our partners at ORNL from the Advance Scientific Computing Research (ASCR) organization will collaborate in accelerating the integrated plasma material interaction simulation framework. This simulation framework will provide a platform for generating simulation data across a range of physical fidelities, including hybrid methods that produce multi-fidelity results. This data will be leveraged for AI model development, both for generation of surrogates and the automation of simulation campaigns. A part of the research below will include collaborative efforts with the TAIMC to (i) adapt data storage approaches to ensure AI-readiness, (ii) provide a protypical exemplar to inform and exercise constructed workflows, and (iii) generate and share data, using the TAIMC unified AI data standard, for foundational models that will be trained from multiple sources across the DOE complex. We will also collaborate with the TAIMC, as well as the planned AI modeling teams, to develop approaches for reducing the cost of data generation. These include tailored multi-fidelity approaches as well as fine-tuning strategies to augment general, large-scale foundational models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗