Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

ShenCFD

ShenCFD is a fast pseudospectral solver for fluid dynamics written to be maximally Pythonic and maximally useful for machine-learning-based turbulence model discovery.

Saenz, Juan↗

Towards a Unified theory of Fractional and Nonlocal Vector Calculus

Nonlocal and fractional-order models capture effects that classical partial differential equations cannot describe; for this reason, they are suitable for a broad class of engineering and scientific applications that feature multiscale or anomalous behavior. This has driven a desire for a vector calculus that includes nonlocal and fractional gradient, divergence and Laplacian type operators, as well as tools such as Green’s identities, to model subsurface transport, turbulence, and conservation laws. In the literature, several independent definitions and theories of nonlocal and fractional vector calculus have been put forward. Some have been studied rigorously and in depth, while others have been introduced ad-hoc for specific applications. The goal of this work is to provide foundations for a unified vector calculus by (1) consolidating fractional vector calculus as a special case of nonlocal vector calculus, (2) relating unweighted and weighted Laplacian operators by introducing an equivalence kernel, and (3) proving a form of Green’s identity to unify the corresponding variational frameworks for the resulting nonlocal volume-constrained problems. Here, the proposed framework goes beyond the analysis of nonlocal equations by supporting new model discovery, establishing theory and interpretation for a broad class of operators, and providing useful analogues of standard tools from the classical vector calculus.

97 MATHEMATICS AND COMPUTING↗

Surrogate multi-fidelity data and model fusion for scientific discovery and uncertainty quantification in Earth System Models

This whitepaper addresses the Earth and Environmental Systems Sciences Division (EESSD)’s predictability challenges in modeling the integrated water cycle and data-model integration. Specifically, it focuses on reducing and characterizing the uncertainty in the representation of process models for unresolved physics, either due to model resolution or limited by the physical under standing or computational efficiency, and the use of observational data for in-situ process parameter optimization within ESM. The described methods may also be used to determine the nature of responses (e.g. strength and direction), and hence to identify critical processes that drive the overall ESM responses to perturbation in the forcing

54 ENVIRONMENTAL SCIENCES↗

AI-based language models powering drug discovery and development

The discovery and development of new medicines is expensive, time-consuming, and often inefficient, with many failures along the way. Powered by artificial intelligence (AI), language models (LMs) have changed the landscape of natural language processing (NLP), offering possibilities to transform treatment development more effectively. Here, we summarize advances in AI-powered LMs and their potential to aid drug discovery and development. We highlight opportunities for AI-powered LMs in target identification, clinical design, regulatory decision-making, and pharmacovigilance. We specifically emphasize the potential role of AI-powered LMs for developing new treatments for Coronavirus 2019 (COVID-19) strategies, including drug repurposing, which can be extrapolated to other infectious diseases that have the potential to cause pandemics. Finally, we set out the remaining challenges and propose possible solutions for improvement.

60 APPLIED LIFE SCIENCES↗

Scale-invariant machine-learning model accelerates the discovery of quaternary chalcogenides with ultralow lattice thermal conductivity

We design an advanced machine-learning (ML) model based on crystal graph convolutional neural network that is insensitive to volumes (i.e., scale) of the input crystal structures to discover novel quaternary chalcogenides, AMM'Q 3 (A/M/M' = alkali, alkaline earth, post-transition metals, lanthanides, and Q = chalcogens). These compounds are shown to possess ultralow lattice thermal conductivity (κ l ), a desired requirement for thermal-barrier coatings and thermoelectrics. Upon screening the thermodynamic stability of ~1 million compounds using the ML model iteratively and performing density-functional theory (DFT) calculations for a small fraction of compounds, we discover 99 compounds that are validated to be stable in DFT. Taking several DFT-stable compounds, we calculate their κ l using Peierls–Boltzmann transport equation, which reveals ultralow κ l (<2 Wm -1 K -1 at room temperature) due to their soft elasticity and strong phonon anharmonicity. Our work demonstrates the high efficiency of scale-invariant ML model in predicting novel compounds and presents experimental-research opportunities with these new compounds.

36 MATERIALS SCIENCE↗

Model-Agnostic Signal Discovery with Machine Learning: Bridging the Gap Between Theory and Practice

Searches for new phenomena in complex scientific data are predominantly model-dependent, optimized for specific hypotheses, and therefore limited in their coverage of the space of possible signals. Recently, new AI-based model-agnostic search strategies, many of which have been pioneered in high-energy physics, have been proposed which provide a complementary paradigm, prioritizing broad exploration over tailored analyses. These techniques offer an opportunity to enhance the overall discovery potential of modern experiments, especially in regimes where theoretical guidance is scarce. In this document, we review the conceptual framework behind the main classes of AI-based model-agnostic strategies. We discuss the potential pitfalls of these methods, and strategies for their validation and interpretation. We aim for this document to serve as a useful reference both for practitioners and for researchers interested in learning more about these model-agnostic search strategies.

Amram, Oz [Fermilab] (ORCID:0000000237653123)↗

VLM4Bio: A Benchmark Dataset to Evaluate Pretrained Vision-Language Models for Trait Discovery from Biological Images

Images are increasingly becoming the currency for documenting biodiversity on the planet, providing novel opportunities for accelerating scientific discoveries in the field of organismal biology, especially with the advent of large vision-language models (VLMs). We ask if pre-trained VLMs can aid scientists in answering a range of biologically relevant questions without any additional fine-tuning. In this paper, we evaluate the effectiveness of 12 state-of-the-art (SOTA) VLMs in the field of organismal biology using a novel dataset, VLM4Bio, consisting of 469K question8 answer pairs involving 30K images from three groups of organisms: fishes, birds, and butterflies, covering five biologically relevant tasks. We also explore the effects of applying prompting techniques and tests for reasoning hallucination on the performance of VLMs, shedding new light on the capabilities of current SOTA VLMs in answering biologically relevant questions using images

Maruf, M [Virginia Tech, Blacksburg]↗

A Mechanism-Based Reaction–Diffusion Model for Accelerated Discovery of Thermoset Resins Frontally Polymerized by Olefin Metathesis

Frontal ring-opening metathesis polymerization (FROMP) involves a self-perpetuating exothermic reaction, which enables the rapid and energy-efficient manufacturing of thermoset polymers and composites. Current state-of-the-art reaction–diffusion FROMP models rely on a phenomenological description of the olefin metathesis kinetics, limiting their ability to model the governing thermo-chemical FROMP processes. Furthermore, the existing models are unable to predict the variations in FROMP kinetics with changes in the resin composition and as a result are of limited utility toward accelerated discovery of new resin formulations. In this work, we formulate a chemically meaningful model grounded in the established mechanism of ring-opening metathesis polymerization (ROMP). Our study aims to validate the hypothesis that the ROMP mechanism, applicable to monomer-initiator solutions below 100 °C, remains valid under the nonideal conditions encountered in FROMP, including ambient to >200 °C temperatures, sharp temperature gradients, and neat monomer environments. Through extensive simulations, we demonstrate that our mechanism-based model accurately predicts the FROMP behavior across various resin compositions, including polymerization front velocities and thermal characteristics (e.g., T max ). Additionally, we introduce a semi-inverse workflow that predicts FROMP behavior from a single experimental data point. Notably, the physiochemical parameters utilized in our model can be obtained through DFT calculations and minimal experiments, highlighting the model’s potential for rapid screening of new FROMP chemistries in pursuit of thermoset polymers with superior thermo-chemo-mechanical properties.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Language models for materials discovery and sustainability: Progress, challenges, and opportunities

Significant advancements have been made in one of the most critical branches of artificial intelligence: natural language processing (NLP). These advancements are exemplified by the remarkable success of OpenAI’s GPT-3.5/4 and the recent release of GPT-4.5, which have sparked a global surge of interest akin to an NLP gold rush. Here, in this article, we offer our perspective on the development and application of NLP and large language models (LLMs) in materials science. We begin by presenting an overview of recent advancements in NLP within the broader scientific landscape, with a particular focus on their relevance to materials science. Next, we examine how NLP can facilitate the understanding and design of novel materials and its potential integration with other methodologies. To highlight key challenges and opportunities, we delve into three specific topics: (i) the limitations of LLMs and their implications for materials science applications, (ii) the creation of a fully automated materials discovery pipeline, and (iii) the potential of GPT-like tools to synthesize existing knowledge and aid in the design of sustainable materials.

36 MATERIALS SCIENCE↗

Deep Generative Models for Materials Discovery and Machine Learning-Accelerated Innovation

Machine learning and artificial intelligence (AI/ML) methods are beginning to have significant impact in chemistry and condensed matter physics. For example, deep learning methods have demonstrated new capabilities for high-throughput virtual screening, and global optimization approaches for inverse design of materials. Recently, a relatively new branch of AI/ML, deep generative models (GMs), provide additional promise as they encode material structure and/or properties into a latent space, and through exploration and manipulation of the latent space can generate new materials. These approaches learn representations of a material structure and its corresponding chemistry or physics to accelerate materials discovery, which differs from traditional AI/ML methods that use statistical and combinatorial screening of existing materials via distinct structure-property relationships. However, application of GMs to inorganic materials has been notably harder than organic molecules because inorganic structure is often more complex to encode. In this work we review recent innovations that have enabled GMs to accelerate inorganic materials discovery. We focus on different representations of material structure, their impact on inverse design strategies using variational autoencoders or generative adversarial networks, and highlight the potential of these approaches for discovering materials with targeted properties needed for technological innovation.

36 MATERIALS SCIENCE↗

Structural constraint integration in a generative model for the discovery of quantum materials

Billions of organic molecules have been computationally generated, yet functional inorganic materials remain scarce due to limited data and structural complexity. Here, in this work, we introduce Structural Constraint Integration in a GENerative model (SCIGEN), a framework that enforces geometric constraints, such as honeycomb and kagome lattices, within diffusion-based generative models to discover stable quantum materials candidates. SCIGEN enables conditional sampling from the original distribution, preserving output validity while guiding structural motifs. This approach generates ten million inorganic compounds with Archimedean and Lieb lattices, over 10% of which pass multistage stability screening. High-throughput density functional theory calculations on 26,000 candidates shows over 95% convergence and 53% structural stability. A graph neural network classifier detects magnetic ordering in 41% of relaxed structures. Furthermore, we synthesize and characterize two predicted materials, TiPd 0.22 Bi 0.88 and Ti 0.5 Pd 1.5 Sb, which display paramagnetic and diamagnetic behaviour, respectively. Our results indicate that SCIGEN provides a scalable path for generating quantum materials guided by lattice geometry.

36 MATERIALS SCIENCE↗

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference↗

Reviving MeV-GeV indirect detection with inelastic dark matter

Thermal relic dark matter below ∼ 10 GeV is excluded by cosmic microwave background data if its annihilation to visible particles is unsuppressed near the epoch of recombination. Usual model-building measures to avoid this bound involve kinematically suppressing the annihilation rate in the low-velocity limit, thereby yielding dim prospects for indirect detection signatures at late times. In this work, we investigate a class of cosmologically viable sub-GeV thermal relics with late-time annihilation rates that are detectable with existing and proposed telescopes across a wide range of parameter space. We study a representative model of inelastic dark matter featuring a stable state χ 1 and a slightly heavier excited state χ 2 whose abundance is thermally depleted before recombination. Since the kinetic energy of dark matter in the Milky Way is much larger than it is during recombination, χ 1 χ 1 → χ 2 χ 2 upscattering can efficiently regenerate a cosmologically long-lived Galactic population of χ 2 , whose subsequent coannihilations with χ 1 give rise to observable gamma-rays in the ∼ 1 MeV − 100 MeV energy range. We find that proposed MeV gamma-ray telescopes, such as e-ASTROGAM, AMEGO, and MAST, would be sensitive to much of the thermal relic parameter space in this class of models and thereby enable both discovery and model discrimination in the event of a signal at accelerator or direct detection experiments. Published by the American Physical Society 2024

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Representations and strategies for transferable machine learning improve model performance in chemical discovery

Strategies for machine-learning (ML)-accelerated discovery that are general across material composition spaces are essential, but demonstrations of ML have been primarily limited to narrow composition variations. By addressing the scarcity of data in promising regions of chemical space for challenging targets such as open-shell transition-metal complexes, general representations and transferable ML models that leverage known relationships in existing data will accelerate discovery. Over a large set (~1000) of isovalent transition-metal complexes, we quantify evident relationships for different properties (i.e., spin-splitting and ligand dissociation) between rows of the Periodic Table (i.e., 3d/4d metals and 2p/3p ligands). We demonstrate an extension to the graph-based revised autocorrelation (RAC) representation (i.e., eRAC) that incorporates the group number alongside the nuclear charge heuristic that otherwise overestimates dissimilarity of isovalent complexes. To address the common challenge of discovery in a new space where data are limited, we introduce a transfer learning approach in which we seed models trained on a large amount of data from one row of the Periodic Table with a small number of data points from the additional row. We demonstrate the synergistic value of the eRACs alongside this transfer learning strategy to consistently improve model performance. Analysis of these models highlights how the approach succeeds by reordering the distances between complexes to be more consistent with the Periodic Table, a property we expect to be broadly useful for other material domains.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗