Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computing methodologies → machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Machine learning and shallow groundwater chemistry to identify geothermal prospects in the Great Basin, USA

This study discovers various geothermal prospects in the Great Basin, USA based on shallow groundwater chemical (geochemical) data. The geochemical data are expected to include hidden (latent) information that is a proxy for geothermal prospectivity. We processed the sparse geochemical data in the Great Basin at 14,341 locations including 18 attributes. Next, a non-negative matrix factorization with customized k-means clustering is applied to the geochemical data matrix that automatically finds three hidden geothermal signatures representing modestly, moderately, and highly confident geothermal prospects. The algorithm also evaluated the probability of occurrence of these types of resources through the studied region. There is a consistency between regional geothermal prospectivity as estimated by our ML methodology and the traditional play fairway analysis conducted over a portion of the study area. We also identify the dominant data attributes associated with each signature. Finally, our ML analyses allow us to reconstruct attributes from sparse into continuous over the study domain. The predicted continuous attributes can be used for future detailed geothermal explorations in the Great Basin.

15 GEOTHERMAL ENERGY↗

Agentic AI vs ML-Based Autotuning: A Comparative Study for Loop Reordering Optimization

High Performance Computing (HPC) applications rely heavily on code optimizations to achieve good performance on modern CPU and GPU architectures. Traditional Machine Learning auto-tuning approaches have demonstrated success in exploring high-dimensional spaces, but they often require expensive compile-run evaluations and lack adaptability for large HPC applications. The recent advances in Large Language Models (LLMs) and Agentic AI systems raise intriguing questions about the potential of these approaches to address specific optimization methodologies. This work aims to answer an essential question for the HPC community: “How Agentic AI Systems Compare to Traditional ML Autotuning Techniques?” To address this question, we present a comparative analysis between a traditional ML-based optimization approach and an Agentic AI system, evaluating their respective capabilities and limitations for loop-level optimization. In addition, we introduced a new Agentic AI system named LoopGen-AI using three different Large Language Models: GPT-4.1, Claude 4.0, and Gemini 2.5. A key finding is that LoopGen-AI achieves competitive per-formance with only a few program runs, the reasoning logs from the agents revealed that their decisions rely heavily on the combination of semantic understanding of the target kernel with dynamic feedback from the environment, highlighting a promising new dimension in performance tuning. In contrast, ML-based autotuners focus on statistical exploration, and require orders of magnitude more runs to reach peak performance. Additionally, our analysis shows that prompt engineering, particularly using Persona + Context Manager patterns, significantly impacts the effectiveness of Agentic AI. Our results indicate that while Agentic AI systems are not yet a complete replacement for ML-based autotuners, it can effectively complement traditional methods.

Rosas, Miguel Romero↗

A detailed study of interpretability of deep neural network based top taggers

Abstract Recent developments in the methods of explainable artificial intelligence (XAI) allow researchers to explore the inner workings of deep neural networks (DNNs), revealing crucial information about input–output relationships and realizing how data connects with machine learning models. In this paper we explore interpretability of DNN models designed to identify jets coming from top quark decay in high energy proton–proton collisions at the Large Hadron Collider. We review a subset of existing top tagger models and explore different quantitative methods to identify which features play the most important roles in identifying the top jets. We also investigate how and why feature importance varies across different XAI metrics, how correlations among features impact their explainability, and how latent space representations encode information as well as correlate with physically meaningful quantities. Our studies uncover some major pitfalls of existing XAI methods and illustrate how they can be overcome to obtain consistent and meaningful interpretation of these models. We additionally illustrate the activity of hidden layers as neural activation pattern diagrams and demonstrate how they can be used to understand how DNNs relay information across the layers and how this understanding can help to make such models significantly simpler by allowing effective model reoptimization and hyperparameter tuning. These studies not only facilitate a methodological approach to interpreting models but also unveil new insights about what these models learn. Incorporating these observations into augmented model design, we propose the particle flow interaction network model and demonstrate how interpretability-inspired model augmentation can improve top tagging performance.

97 MATHEMATICS AND COMPUTING↗

Algorithm to extract direction in 2D discrete distributions and a continuous Frobenius norm

In this study, we present a novel algorithm for determining directionality in 2D distributions of discrete data. We compare a reference dataset with a known direction to a measured dataset with an unknown direction by the Frobenius norm of the difference (FND) to find the unknown direction. To generalize this concept, we develop a continuous Frobenius norm of the difference (CFND) as a continuous analog of the FND and derive its analytical expression. By relating fitted and normalized 2D Gaussian distributions, we show that the CFND approximates the FND, and we validate this relationship with computer simulations. We find that a first-order approximation of the CFND between two similar Gaussian distributions takes the form of an absolute sine function, offering a simple analytical form with potential for specialized applications in segmented inverse beta decay (IBD) neutrino detectors, astronomy, machine learning, and more. Although this method may easily extend to 3D scalar fields, our focus here is on 2D real-valued fields as it directly applies to directionality. Our methodology consists of modeling a 2D Gaussian distribution, binning the data into a histogram, and encoding it as a square matrix. Rotating this matrix around its geometric center and comparing it to a measured dataset using the FND gives us rotational data that we fit with an absolute sine function. The location of the minimum of this fit is the angle closest to the true angle of the direction in the measured dataset. We present the derivation and discuss initial applications of the CFND in our novel algorithm, demonstrating its success in approximating directionality in 2D distributions.

Data Analysis, Statistics and Probability (physics↗

Machine Learning for Geothermal Resource Exploration in the Tularosa Basin, New Mexico

Geothermal energy is considered an essential renewable resource to generate flexible electricity. Geothermal resource assessments conducted by the U.S. Geological Survey showed that the southwestern basins in the U.S. have a significant geothermal potential for meeting domestic electricity demand. Within these southwestern basins, play fairway analysis (PFA), funded by the U.S. Department of Energy’s (DOE) Geothermal Technologies Office, identified that the Tularosa Basin in New Mexico has significant geothermal potential. This short communication paper presents a machine learning (ML) methodology for curating and analyzing the PFA data from the DOE’s geothermal data repository. The proposed approach to identify potential geothermal sites in the Tularosa Basin is based on an unsupervised ML method called non-negative matrix factorization with custom k-means clustering. This methodology is available in our open-source ML framework, GeoThermalCloud (GTC). Using this GTC framework, we discover prospective geothermal locations and find key parameters defining these prospects. Our ML analysis found that these prospects are consistent with the existing Tularosa Basin’s PFA studies. This instills confidence in our GTC framework to accelerate geothermal exploration and resource development, which is generally time-consuming.

15 GEOTHERMAL ENERGY↗

A framework to evaluate machine learning crystal stability predictions

The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.

Riebesell, Janosh↗

Multiscale modeling of packed-bed microwave reactors and estimation of intrinsic materials' permittivity

Modeling of packed-bed microwave reactors relies on an accurate representation of particle size, shape, and distribution within the bed, as well as the particles' dielectric properties. The measured permittivity of microwave susceptors (powders or structured materials) depends on the geometric features of the particles and the porosity of the bed, as well as the specific form factor of a structured material. These are effective properties and cannot be used to analyze other reactor configurations unless the geometric effects are removed. Therefore, we introduce a methodology for extracting the intrinsic particle permittivity from experimentally measured effective permittivity by combining cavity-based measurements with multiscale simulations and machine learning. Further, we develop the first multiscale model of packed-bed microwave reactors that incorporate particle effects (geometric features, random packing, and particle contact). This approach bridges macroscopic observables with mesoscopic physics, enabling analysis of local hotspots, arcing, and contact effects that control reactor performance. Using polymer-based spherical activated carbon (PBSAC) and silicon carbide (SiC) as examples, we demonstrate that the inferred particle permittivity is consistent with independent experimental heating profiles we collect from microwave reactors without adjustable parameters. Finally, this methodology establishes a foundation for predictive, multiscale design of microwave packed-bed reactors that explicitly accounts for particle-scale effects, enabling the estimation of intrinsic permittivity for the first time.

97 MATHEMATICS AND COMPUTING↗

Hydrogen in disordered titania: connecting local chemistry, structure, and stoichiometry through accelerated exploration

Hydrogen incorporation in native surface oxides of metal alloys often controls the onset of metal hydriding, with implications for materials corrosion and hydrogen storage. A key representative example is titania, which forms as a passivating layer on a variety of titanium alloys for structural and functional applications. These oxides tend to be structurally diverse, featuring polymorphic phases, grain boundaries, and amorphous regions that generate a disparate set of unique local environments for hydrogen. Here, we introduce a workflow that can efficiently and accurately navigate this complexity. First, a machine learning force field, trained on ab initio molecular dynamics simulations, was used to generate amorphous configurations. Density functional theory calculations were then performed on these structures to identify local oxygen environments, which were compared against experimental observations. Second, to classify subtle differences across the disordered configuration space, we employ a graph-based sampling procedure. Finally, local hydrogen binding energies and hopping kinetics are computed using exhaustive density functional theory calculations on representative configurations. Here, we leverage this methodology to show that hydrogen binding energetics are described by local oxygen coordination, which in turn is affected by stoichiometry, and form the basis of hopping kinetics and diffusion. Together these results imply that hydrogen incorporation and transport in TiO x can be tailored through compositional engineering, with implications for improving the performance and durability of titanium-derived alloys in hydrogen environments.

36 MATERIALS SCIENCE↗

An AI-Based 3D Bat Movement Tracking System at Wind Energy Facilities Using Multi-Thermal Video Cameras

The talk at the NAWEA Wind Tech 2024 conference discusses how to leverage the potential of real-time thermal-imaging methodologies in quantifying nocturnal bat activities at wind turbines, using 3D computer vision techniques within a deep learning framework. This innovation enables the automatic detection and classification of bats, birds, and insects in thermal-imaging videos captured at wind turbine sites, facilitating efficient and accurate data analysis for enhanced understanding and mitigation of bat-wind turbine interactions.

AI↗

Machine learning-based ethylene concentration estimation, real-time optimization and feedback control of an experimental electrochemical reactor

With the increase in electricity supply from clean energy sources, electrochemical reduction of carbon dioxide (CO 2 ) has received increasing attention as an alternative source of carbon-based fuels. As CO 2 reduction is becoming a stronger alternative for the clean production of chemicals, the need to model, optimize and control the electrochemical reduction of the CO 2 process becomes inevitable. However, on one hand, a first-principles model to represent the electrochemical CO 2 reduction has not been fully developed yet because of the complexity of its reaction mechanism, which makes it challenging to define a precise state-space model for the control system. On the other hand, the unavailability of efficient concentration measurement sensors continues to challenge our ability to develop feedback control systems. Gas chromatography (GC) is the most common equipment for monitoring the gas product composition, but it requires a period of time to analyze the sample, which means that GC can provide only delayed measurements during the operation. Moreover, the electrochemical CO 2 reduction process is catalyzed by a fast-deactivating copper catalyst and undergoes a selectivity shift from the product-of-interest at the later stages of experiments, which can pose a challenge for conventional control methods. To this end, machine learning (ML) techniques provide a potential approach to overcome those difficulties due to their demonstrated ability to capture the dynamic behavior of a chemical process from data. Motivated by the above considerations, we propose a machine learning-based modeling methodology that integrates support vector regression and first-principles modeling to capture the dynamic behavior of an experimental electrochemical reactor; this model, together with limited gas chromatography measurements, is employed to predict the evolution of gas-phase ethylene concentration. The model prediction is directly used in a proportional-integral (PI) controller that manipulates the applied potential to regulate the gas-phase ethylene concentration at energy-optimal set-point values computed by a real-time process optimizer (RTO). Specifically, the RTO calculates the operation set-point by solving an optimization problem to maximize the economic benefit of the reactor. Finally, suitable compensation methods are introduced to further account for the experimental uncertainties and handle catalyst deactivation. The proposed modeling, optimization, and control approaches are the first demonstration of active control for a CO 2 electrolyzer and contribute to the automation and scale-up efforts for electrified manufacturing of fuels and chemicals starting from CO 2 .

42 ENGINEERING↗

Understanding and design of interstitial oxygen conductors

Highly efficient oxygen-active materials that react with, absorb, and transport oxygen is essential for fuel cells, electrolyzers and related applications. While vacancy-mediated oxygen-ion conductors have long been the focus of research, they are limited by high migration barriers at intermediate temperatures (400–600 °C), which hinder their practical applications. In contrast, interstitial oxygen conductors exhibit significantly lower migration barriers enabling higher ionic conductivity at lower temperatures. This review systematically examines both well-established and recently identified families of interstitial oxygen-ion conductors, focusing on how their unique structural motifs such as corner-sharing polyhedral frameworks, isolated polyhedral, and cage-like architectures, facilitate low migration barriers through interstitial and/or interstitialcy diffusion mechanisms. A central discussion of this review focuses on the evolution of design strategies, from targeted donor doping, element screening, to physical-intuition descriptor material screening and machine learning approach, which leverage computational tools to explore vast chemical spaces in search for new interstitial conductors. The success of these strategies demonstrates that a significant, largely unexplored space remains for discovering high-performing interstitial oxygen conductors. Crucial features enabling high-performance interstitial oxygen diffusion include the availability of electrons for oxygen reduction and sufficient structural flexibility with accessible volume for interstitial accommodation and migration. This review concludes with a forward-looking perspective, proposing a knowledge-driven methodology that integrates current understanding with data-centric approaches to identify promising interstitial oxygen conductors outside traditional search paradigms. These approaches are expected to significantly accelerate the development of high-performance interstitial oxygen conductors for a variety of oxygen-active applications, ultimately paving the way for more efficient and sustainable energy technologies.

Interstitial oxygen conductors↗

Internet of Things Data Characterization Process: Pattern of Life Behavioral Data Study

The HoneyBee™ TARDIS LDRD team completed a data scoping study that identified the initial processes and procedures to baseline the normal and expected behaviors during operability and interoperability of Internet of Things (IoT) device networks. This research is the initial step in developing a process (or methodology) to inform a much broader information framework incorporating machine learning to determine device pattern-of-life which enables the detection of abnormal IoT behaviors on an individual device, as well as in the context of a larger network.

97 MATHEMATICS AND COMPUTING↗

Cross-domain digital twin architecture for predictive maintenance via machine learning and Large Language Models

This research introduces a comprehensive framework for creating and deploying a digital twin platform for continuous monitoring and predictive maintenance within industrial settings. Through utilizing advanced technologies, including Unreal Engine 5, Unity 3D, the Message Queue Telemetry Transport protocol, Random Forest machine learning algorithms, and Large Language Models (LLMs), we establish a platform that digitally reproduces physical equipment and translates digital controls into real-world actions. This facilitates preventive maintenance approaches and improves operational effectiveness. The digital twin platform gathers sensor data from operational equipment, analyzes it using machine learning, and delivers practical insights to prevent potential malfunctions and enhance equipment performance. Furthermore, the incorporation of a web portal enables efficient monitoring and access to historical data, educational materials, and equipment status information. Preliminary findings indicate that digital twins can transform industrial equipment management and maintenance methodologies.

97 MATHEMATICS AND COMPUTING↗

Preventing Reverse Engineering of Critical Industrial Data with DIOD

Business analytics augmented by artificial intelligence and machine learning (AI/ML) have revolutionized the role of data in the modern world. In recent years, businesses have incorporated data into their decision-making process for better prediction, risk-assessment, content creation, etc. While such businesses often seek to leverage the full use of their data through third-party AI/ML services, they are often hampered by the risks of data leaks, reverse-engineering, stolen technology, etc. that often have disastrous consequences for businesses and their stakeholders alike. Thus, there arises a need for data masking prior to its transmission that obfuscates proprietary information while preserving the information relevant for AI/ML applications. In order to meet the needs of industrial data which are significantly different from those of data warehouses, previous work proposed an efficient time and space-scalable data masking paradigm known as the deceptive infusion of data (DIOD) methodology. The present work expands upon this work by leveraging existing reverse-engineering capabilities to facilitate the decomposition of industrial data into its proprietary and AI/ML-relevant parts, referred to as fundamental and inference metadata respectively. Both sets of metadata are further obfuscated in accordance with the DIOD methodology to create the DIOD rendition of the industrial data, which is rendered immune to reverse-engineering by discarding proprietary information and only preserving AI/ML-relevant information. Additionally, constraints of the original DIOD manuscript are relaxed using mutual information by configuring the methodology to the target AI/ML application to unlock the full potential of the DIOD methodology. As an example, data from a nuclear reactor is transformed into that from a nonlinear spring-mass system with different levels of data masking as required by the generic system and the target application.

97 MATHEMATICS AND COMPUTING↗

Combining chemistry and protein engineering for new-to-nature biocatalysis

Biocatalysis, the application of enzymes to solve synthetic problems of human import, has blossomed into a powerful technology for chemical innovation. In the past decade, a threefold partnership, where nature provides blueprints for enzymatic catalysis, chemists introduce innovative activity modes with abiological substrates, and protein engineers develop new tools and algorithms to tune and improve enzymatic function, has unveiled the frontier of new-to-nature enzyme catalysis. In this Perspective, we highlight examples of interdisciplinary studies, which have helped to expand the scope of biocatalysis, including concepts of enzymatic versatility explored through the lens of biomimicry, to achieve activities and selectivities beyond those currently possible with chemocatalysis. We indicate how modern tools, such as directed evolution, computational protein design and machine learning-based protein engineering methods, have already impacted and will continue to influence enzyme engineering for new abiological transformations. As a result, a sustained collaborative effort across disciplines is anticipated to spur further advances in biocatalysis in the coming years.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cyber Threat Assessment Methodology for Autonomous and Remote Operations for Advanced Reactors (Conference Presentation)

The next generation of Advanced Reactors include planned capabilities for both Autonomous (operating without human interaction for a set period-of-time) and Remote (operating with human interaction from a separate physical location) Operations. Existing Nuclear Reactor architectures include a set of safety and security constraints tightly coupled with personnel policies and procedures. As Advanced Reactors are fielded with these new Autonomous and Remote operational capabilities, the architectures and associated infrastructure services and components will perceivably expand the overall attack surface and risk calculations with regards to safe and secure operations. This paper is part of an FY21 work program focused on ensuring Advanced Reactor designs are informed with threat-based guidance on design and operation of Secure Architectures with a specific focus on the deployment of Autonomous Systems in support of Advanced Reactor Operations. The next phase of this research program is to complement produce a methodology for assessment of the cyber threat against these architectures as well as a catalogue of Use Cases to support the Advanced Reactor community in their implementation of Autonomous and Remote Operations.

97 MATHEMATICS AND COMPUTING↗

Towards verifiable cancer digital twins: tissue level modeling protocol for precision medicine

Cancer exhibits substantial heterogeneity, manifesting as distinct morphological and molecular variations across tumors, which frequently undermines the efficacy of conventional oncological treatments. Developments in multiomics and sequencing technologies have paved the way for unraveling this heterogeneity. Nevertheless, the complexity of the data gathered from these methods cannot be fully interpreted through multimodal data analysis alone. Mathematical modeling plays a crucial role in delineating the underlying mechanisms to explain sources of heterogeneity using patient-specific data. Intra-tumoral diversity necessitates the development of precision oncology therapies utilizing multiphysics, multiscale mathematical models for cancer. This review discusses recent advancements in computational methodologies for precision oncology, highlighting the potential of cancer digital twins to enhance patient-specific decision-making in clinical settings. We review computational efforts in building patient-informed cellular and tissue-level models for cancer and propose a computational framework that utilizes agent-based modeling as an effective conduit to integrate cancer systems models that encode signaling at the cellular scale with digital twin models that predict tissue-level response in a tumor microenvironment customized to patient information. Furthermore, we discuss machine learning approaches to building surrogates for these complex mathematical models. These surrogates can potentially be used to conduct sensitivity analysis, verification, validation, and uncertainty quantification, which is especially important for tumor studies due to their dynamic nature.

60 APPLIED LIFE SCIENCES↗

Generalized representative structures for atomistic systems

A new method is presented to generate atomic structures that reproduce the essential characteristics of arbitrary material systems, phases, or ensembles. Previous methods allow one to reproduce the essential characteristics (e.g. the chemical disorder) of a large random alloy within a small crystal structure. The ability to generate small representations of random alloys, along with the restriction to crystal systems, results from using the fixed-lattice cluster correlations to describe structural characteristics. A more general description of the structural characteristics of atomic systems is obtained using complete sets of atomic environment descriptors. These are used within for generating representative atomic structures without restriction to fixed lattices. A general data-driven approach is provided here utilizing the atomic cluster expansion (ACE) basis. The N-body ACE descriptors are a complete set of atomic environment descriptors that span both chemical and spatial degrees of freedom and are used within for describing atomic structures. The generalized representative structure (GRS) method presented within generates small atomic structures that reproduce ACE descriptor distributions corresponding to arbitrary structural and chemical complexity. It is shown that systematically improvable representations of crystalline systems on fixed parent lattices, amorphous materials, liquids, and ensembles of atomic structures may be produced efficiently through optimization algorithms. With the GRS method, we highlight reduced representations of atomistic machine-learning training datasets that contain similar amounts of information and small 40–72 atom representations of liquid phases. The ability to use GRS methodology as a driver for informed novel structure generation is also demonstrated. The advantages over other data-driven methods and state-of-the-art methods restricted to high-symmetry systems are highlighted.

atomic cluster expansion↗