Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Traditional Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34

A Hierarchical Feature-Based Methodology to Perform Cervical Cancer Classification

Prevention of cervical cancer could be performed using Pap smear image analysis. This test screens pre-neoplastic changes in the cervical epithelial cells; accurate screening can reduce deaths caused by the disease. Pap smear test analysis is exhaustive and repetitive work performed visually by a cytopathologist. This article proposes a workload-reducing algorithm for cervical cancer detection based on analysis of cell nuclei features within Pap smear images. We investigate eight traditional machine learning methods to perform a hierarchical classification. We propose a hierarchical classification methodology for computer-aided screening of cell lesions, which can recommend fields of view from the microscopy image based on the nuclei detection of cervical cells. We evaluate the performance of several algorithms against the Herlev and CRIC databases, using a varying number of classes during image classification. Results indicate that the hierarchical classification performed best when using Random Forest as the key classifier, particularly when compared with decision trees, k-NN, and the Ridge methods.

60 APPLIED LIFE SCIENCES↗

AEcroscopy: A Software–Hardware Framework Empowering Microscopy Toward Automated and Autonomous Experimentation

Microscopy has been pivotal in improving the understanding of structure-function relationships at the nanoscale and is by now ubiquitous in most characterization labs. However, traditional microscopy operations are still limited largely by a human-centric click-and-go paradigm utilizing vendor-provided software, which limits the scope, utility, efficiency, effectiveness, and at times reproducibility of microscopy experiments. Here, in this work, a coupled software–hardware platform is developed that consists of a software package termed AEcroscopy (short for Automated Experiments in Microscopy), along with a field-programmable-gate-array device with LabView-built customized acquisition scripts, which overcome these limitations and provide the necessary abstractions toward full automation of microscopy platforms. The platform works across multiple vendor devices on scanning probe microscopes and electron microscopes. It enables customized scan trajectories, processing functions that can be triggered locally or remotely on processing servers, user-defined excitation waveforms, standardization of data models, and completely seamless operation through simple Python commands to enable a plethora of microscopy experiments to be performed in a reproducible, automated manner. This platform can be readily coupled with existing machine-learning libraries and simulations, to provide automated decision-making and active theory-experiment optimization to turn microscopes from characterization tools to instruments capable of autonomous model refinement and physics discovery.

47 OTHER INSTRUMENTATION↗

Feature engineering for machine learning enabled early prediction of battery lifetime

Accurate battery lifetime estimates enable accelerated design of novel battery materials and determination of optimal use protocols for longevity in deployments. Unfortunately, traditional battery testing may take years to reach thousands of cycles. Recent studies have shown that machine learning (ML) tools can predict lithium-ion battery lifetimes from 100 or fewer preliminary cycles, representing only a few weeks of cycling. Until now, conclusions about the efficacy and broad applicability of these predictions across a variety of cathode chemistries have been limited by available experimental information. In this work, we leverage a battery cycling dataset representing six cathode chemistries (NMC111, NMC532, NMC622, NMC811, HE5050, and 5Vspinel), multiple electrolyte/anode compositions, and 300 total carefully prepared pouch batteries to explore feature selection and battery chemistry's role in ML battery lifetime predictions. Here, a mean absolute error (MAE) of 78 cycles in prediction was seen for a chemistry-spanning test set from 100 preliminary cycles. Furthermore, an MAE of 103 cycles was seen when using only the first cycle. This study represents an in-depth investigation of strategies for feature selection for battery lifetime prediction, ML models' generalization across multiple battery chemistries, and predictions beyond the training set in the chemical space.

25 ENERGY STORAGE↗

Transmission-distribution long-term volt-var planning considering reactive power support capability of distributed PV

High penetration of grid-edge, inverter-based photovoltaic (PV) can cause significant voltage fluctuations not only at the distribution but also at the sub-transmission levels due to PV output intermittency. Traditional reactive power planning approaches do not consider intermittency, nor the possibility of coordinating the control of existing and future volt-ampere reactive resources. This paper proposes a reactive power planning tool for sub-transmission systems to mitigate voltage violations and fluctuations caused by high PV penetration and intermittency with a minimum investment cost. The planning tool coordinates with an optimization-based volt-var operational tool for: a) modeling the coordination of all existing var assets in both sub-transmission and distribution systems to reduce the need of new equipment, and b)selecting a set of scenarios with voltage violations, derived from PV intermittency c) testing the final investment decision. The tool obtains an investment need for each intermittency scenario with a proposed optimal power-flow framework with efficient techniques to handle a high number of discrete variables. Two options are provided for final planning decision: i) a conservative direct combination of investment need solutions and ii) a machine learning-based selection of representative investment needs at most time steps. The final investment decision options are verified using a realistic large-scale sub-transmission system and 5-minute PV and load data. The results show a significant voltage performance improvement with a lower investment cost for additional var equipment compared to conventional approaches.

14 SOLAR ENERGY↗

Review of In Situ Sensing for Directed Energy Deposition for Industrial Part Quality Assessment

As the use additive manufacturing (AM) processes continues to grow in critical industries, improved quality assurance methods are becoming increasingly sought after for qualification and certification of AM components. Traditional nondestructive evaluation of printed components is often unable to supply the required confidence in print quality to justify qualification and certification, but the layer-by-layer nature of AM provides unprecedented opportunities for in situ quality inspection. This document summarizes recent developments in process monitoring research specifically related to Directed Energy Deposition (DED). Particular attention is given to three aspects of the highlighted manuscripts: (1) the type of sensors used, (2) features extracted from each sensor modality, and (3) analysis of extracted features for AM quality assessment. Based on the review of the state-of-the-art, several observations have been made. First, none of the reviewed works have applied their trained models to real part geometries, with many of the works relying on single track experiments, thin-walled structures, and cubes. Similarly, there have not been any works demonstrating model generalizability, i.e., a model trained on data from one build allows for fruitful analysis of data from another build. Many works used machine learning techniques to distinguish different process regimes (i.e., normal, keyholing, lack-of-fusion), but very few papers have investigated stochastic variation in an already “optimized” process. Sensor fusion approaches are also limited in the DED sensing literature, but the few works that have employed such techniques have demonstrated the benefits. Finally, registration of in situ data to the build coordinate system is of paramount importance to producing industrially relevant in situ monitoring systems. Data registration allows direct correlations between process anomalies detected in the process monitoring data to localized departures in part quality, but such techniques are generally lacking in the current literature.

36 MATERIALS SCIENCE↗

Predictive Rules of Efflux Inhibition and Avoidance in Pseudomonas aeruginosa

Antibiotic-resistant bacteria rapidly spread in clinical and natural environments and challenge our modern lifestyle. A major component of defense against antibiotics in Gram-negative bacteria is a drug permeation barrier created by active efflux across the outer membrane. We identified molecular determinants defining the propensity of small peptidomimetic molecules to avoid and inhibit efflux pumps in Pseudomonas aeruginosa, a human pathogen notorious for its antibiotic resistance. Combining experimental and computational protocols, we mapped the fate of the compounds from structure-activity relationships through their dynamic behavior in solution, permeation across both the inner and outer membranes, and interaction with MexB, the major efflux transporter of P. aeruginosa. We identified predictors of efflux avoidance and inhibition and demonstrated their power by using a library of traditional antibiotics and compound series and by generating new inhibitors of MexB. The identified predictors will enable the discovery and optimization of antibacterial agents suitable for treatment of P. aeruginosa infections.

59 BASIC BIOLOGICAL SCIENCES↗

2019 Budget Request for the DOE Computational Science Graduate Fellowship (CSGF) Grant

The Department of Energy Computational Science Graduate Fellowship (DOE CSGF) is necessary to meet the continual challenging national workforce needs that arise as computational science and engineering problems continue to grow in scope and complexity. Computational science and engineering (CSE) is a multidisciplinary approach that uses scientific computing to solve practical problems methods and to supply technical tools across the scientific discovery spectrum. In particular, the DOE CSGF emphasizes high-performance computing (HPC) that enables CSE that advances science and engineering in directions important to the DOE and the economy in general. Over the past half-century, HPC has been an essential tool for DOE’s success. During this period, important missions, such as nuclear stockpile stewardship, have turned to HPC as an essential technology. Entire science disciplines, such as biology and cosmology, have been transformed through the augmentation of scientific observation via HPC. At government laboratories and in industry, DOE CSGF alumni are helping push traditional HPC boundaries while contributing to discoveries in high-energy physics, renewable energy, fusion-reactor design, additive manufacturing, nanomaterials for next-generation batteries and transistors, and turbine and advanced nuclear reactor modeling. In addition, HPC is used to address national health needs that will eventually point to cures both by helping cancer researchers manage and analyze huge troves of data, by simulating biological mechanisms, and by accelerating drug development — including continuing to rise to the challenge of pandemic-related research. A 2023 report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR office, “Can the United States Maintain Its Leadership in High-Performance Computing?” says of the Program, “The CSGF program provides a barometer for disciplines that will be of interest to future DOE computing.” An explosion in scientific and technological data has driven the need for increasingly sophisticated HPC to transform those data into scientific understanding. With access to more and more data and the proliferation of HPC, Machine Learning and Artificial Intelligence are experiencing a renaissance, complementing the now well-established use of computational simulation. Indeed, in its September 2020 subcommittee report on “AI/ML, Data Intensive Science and High-Performance Computing”, the DOE Advanced Scientific Computing Advisory Committee (ASCAC) explicitly called for a fellowship program to train computational and data scientists to tackle exascale and data-intensive computing challenges. This collaboration of empirical and theory-based modeling will increasingly inform federal policymakers whose decisions affect American society and future generations, and it requires highly skilled and intellectually agile computational scientists who can support the fast-moving DOE National Laboratory research environment. In fact, the DOE CSGF program has explicitly and consistently addressed this need.

97 MATHEMATICS AND COMPUTING↗

Novel Observing Strategies (NOS) and Earth System Digital Twins (ESDT) for Disaster Resilience

"Earth systems have now been observed continuously for more than 50 years, not only from space but also from airplanes, balloons and in-situ sensors. With the addition of commercial remote sensing providers, many more Internet-of-Things sensors and new NASA and international observatories being planned, these incredible amounts of data will soon be augmented by even larger amounts of diverse data and therefore will become more and more difficult to access, integrate, understand and utilize. At the same time, because of climate change and its impacts, in order to predict, manage, and mitigate the effects of extreme science events and disasters, the information produced by all of this data needs to be optimized, organized and analyzed in such a way that it can be utilized by traditional as well as many new non-traditional users. This calls for the development of observing systems that are agile, coordinated and responsive to events of interest, and for information systems to be dynamic, interactive and fast. With these objectives in mind, the Advanced Information Systems Technology (AIST) Program has been developing technologies that will allow for the development of Novel Observing Strategies (NOS) in a distributed, coordinated fashion (i.e., similar to an “Internet-of-Earth-Things”), as well as for the development of novel information systems or Earth System Digital Twins (ESDT) that will build a digital replica of the past and current states of the Earth systems, will derive forecasts of future states under nominal assumptions and will also offer the capability to investigate many hypothetical evolution scenarios under varying impact assumptions. Both NOS and ESDT will be facilitated by the unprecedented advances of Artificial Intelligence technologies, especially Machine Learning (ML), for extracting relevant information from these large amounts of data, for fusing and assimilating diverse data together and for running complex models faster. Together NOS and ESDT will help build the capabilities needed to anticipate, predict and mitigate the effects of future disasters."

Earth Science Remote Sensing; Information Systems↗

Dynamic Earth Energy Storage: Terawatt-year, Grid-scale Energy Storage Using Planet Earth as a Thermal Battery (GeoTES): Phase I Project (Final Report)

Grid-scale energy storage has been identified by the U.S. Department of Energy’s (DOE) Energy Storage Grand Challenge as a necessary technology to support the continued build-out of intermittent renewable energy resources required to attain a carbon-free energy future. To meet this goal, the 2018 Department of Energy Research and Innovation Act mandated the creation of a comprehensive program to accelerate the development and commercialization of next-generation energy storage technologies. One of numerous energy storage technology options is the storage of excess energy as heated geothermal brine in suitable geologic formations. This concept, known as reservoir thermal energy storage (RTES), geologic thermal energy storage (GeoTES), aquifer thermal energy storage (ATES), etc., relies on the storage of thermal energy in geologic formations for recovery and use in large-scale direct use geothermal (e.g., district heating, industrial processes, etc.) and electrical power generation applications. This thermal energy is derived from excess or waste heat from any high-temperature heat source, such as concentrated solar or from conventional thermal/nuclear generation. As such, RTES can potentially play a significant role in meeting the energy storage shortfall in the coming decades. RTES can provide energy arbitrage through both the storage and production of thermal energy stored in geologic formations for direct use applications and can serve as a source of hot fluids that can be used to generate electricity to support peak demand ramping, thus easing stress on transmission and distribution. This energy storage option has geographic benefits in that energy can be stored locally or regionally depending on the various needs/loads. RTES can also be located across an enormous geographic area, without the need for a traditional hydrothermal resource but where thermal gradients and hydrogeology allow economic exploitation of subsurface heat. The work conducted for this project includes (1) a review of lessons learned from past high-temperature RTES international projects; (2) geochemical experimental investigation and numerical simulations of potential domestic sedimentary reservoirs and (3) development of a thermo-hydrological-mechanical (THM) numerical simulation tool for optimizing formation properties and design parameters to maximize thermal energy storage performance.

15 GEOTHERMAL ENERGY↗

Gradient flow based phase-field modeling using separable neural networks

Allen–Cahn equation is a reaction–diffusion equation and is widely used for modeling phase separation. Machine learning methods for solving the Allen–Cahn equation in its strong form suffer from inaccuracies in collocation techniques, errors in computing higher-order spatial derivatives, and the large system size required by the space–time approach. To overcome these challenges, we propose solving the gradient flow of the Ginzburg–Landau free energy functional, which is equivalent to the Allen–Cahn equation, thereby avoiding the second-order spatial derivatives associated with the Allen–Cahn equation. A minimizing movement scheme is employed to solve the gradient flow problem, eliminating the complexities of a space–time approach. We utilize a separable neural network that efficiently represents the phase field through low-rank tensor decomposition. As we use the minimizing movement scheme to numerically solve the gradient flow problem, we thus, refer to the proposed method as the Separable Deep Minimizing Movement (SDMM) method. The evaluation of the functional in the minimizing movement scheme using the Gauss quadrature technique bypasses the inaccuracies associated with collocation techniques traditionally used to solve partial differential equations. A hyperbolic tangent transformation is introduced on the phase field prior to the evaluation of the functional to ensure that it remains strictly bounded within the values of the two phases. For this transformation, theoretical guarantee for energy stability of the minimizing movement scheme is established. Our results suggest that this transformation helps to improve the accuracy and efficiency significantly. The proposed method resolves the challenges faced by state-of-the-art machine learning techniques, outperforming them in both accuracy and efficiency. It is also the first machine learning method to achieve an order of magnitude speed improvement over the finite element method. In addition to its formulation and computational implementation, several case studies illustrate the applicability of the proposed method.

42 ENGINEERING↗

Quantum optimization algorithms: Energetic implications

Since the dawn of quantum computing (QC), theoretical developments like Shor's algorithm proved the conceptual superiority of QC over traditional computing. However, such quantum supremacy claims are difficult to achieve in practice because of the technical challenges of realizing noiseless qubits. In the near future, QC applications will need to rely on noisy quantum devices that offload part of their work to classical devices. One way to achieve this is by using parameterized quantum circuits in optimization or even in machine learning tasks. The energy requirements of quantum algorithms have not yet been studied extensively. Here in this article, we explore several optimization algorithms using both theoretical insights and numerical experiments to understand their impact on energy consumption. Specifically, we highlight why and how algorithms like quantum natural gradient descent, simultaneous perturbation stochastic approximations or circuit learning methods, are at least 2x to 4x more energy efficient than their classical counterparts; why feedback-based quantum optimization is energy-inefficient; and how techniques like Rosalin can improve the energy efficiency of other algorithms by a factor of ≥2 0 x. Finally, we use the NchooseK high-level programming model to run optimization problems on both gate-based quantum computers and quantum annealers. Empirical data indicate that these optimization problems run faster, have better success rates, and consume less energy on quantum annealers than on their gate-based counterparts.

97 MATHEMATICS AND COMPUTING↗

Inference offers a metric to constrain dynamical models of neutrino flavor transformation

The multimessenger astrophysics of compact objects presents a vast range of environments where neutrino flavor transformation may occur and may be important for nucleosynthesis, dynamics, and a detected neutrino signal. Development of efficient techniques for surveying flavor evolution solution spaces in these diverse environments, which augment and complement existing sophisticated computational tools, could leverage progress in this field. To this end we continue our exploration of statistical data assimilation (SDA) to identify solutions to a small-scale model of neutrino flavor transformation. SDA is a machine learning formula wherein a dynamical model is assumed to generate any measured quantities. Specifically, we use an optimization formulation of SDA wherein a cost function is extremized via the variational method. Regions of state space in which the extremization identifies the global minimum of the cost function will correspond to parameter regimes in which a model solution can exist. Our example study seeks to infer the flavor transformation histories of two monoenergetic neutrino beams coherently interacting with each other and with a matter background. We require that the solution be consistent with measured neutrino flavor fluxes at the point of detection, and with constraints placed upon the flavor content at various locations along their trajectories, such as the point of emission, and the locations of the Mikheyev-Smirnov-Wolfenstein resonances. We show how the procedure efficiently identifies solution regimes and rules out regimes where solutions are infeasible. Overall, results in this work intimate the promise of this “variational annealing” methodology to efficiently probe an array of fundamental questions that traditional numerical simulation codes render difficult to access.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

SPARC-X: Quantum simulations at extreme scale - reactive dynamics from first principles

We have developed the massively parallel electronic structure code SPARC-X: a computational framework for performing Kohn-Sham Density Functional Theory (DFT) calculations that can scale linearly with the number of atoms in the system, while being able to leverage petascale and emerging exascale parallel computers to study chemical phenomena at unprecedented length and time scales. SPARC-X exploits a recent breakthrough in electronic structure methodologies: systematically improvable, strictly local, orthonormal, discontinuous real-space bases that efficiently and systematically capture the local chemistry of the system. With further adaptation using new machine-learning techniques and the use of the massively parallel Spectral Quadrature (SQ) electronic structure method, the algorithmic complexity and prefactor associated with DFT calculations involving semilocal as well as hybrid functionals are dramatically reduced. Using petascale computational resources, SPARC-X enables quantum mechanical simulations at length and time scales previously accessible only by empirical approaches, e.g., 1,000,000 atoms for a few picoseconds using semilocal functionals or 1,000 atoms for a few picoseconds using hybrid functionals. Using exascale resources, the sizes and times targeted are two orders of magnitude larger. Such a capability has applications in a wide variety of chemical sciences, including reactive interfaces where large length- and/or long time-scales are needed and traditional force fields fail. This is particularly important in dynamic catalysis, where bond breaking and formation must be understood in detail. We developed, tested, and employed the SPARC-X framework to understand the photocatalytic properties of TiO 2 nanoparticles, revealing finite size effects that cannot be captured with standard model systems or functionals. This integrated development and application strategy ensures that SPARC-X remains a robust, efficient, and scalable software package for quantum simulations on current petascale and emerging exascale computing resources.

97 MATHEMATICS AND COMPUTING↗

A User-Friendly GUI Tool for Automated Microstructural Analysis of Fiber-Reinforced Composites and Porous Structures

Understanding and quantifying microstructural features such as fiber orientation and porosity is critical for predicting the mechanical behavior and performance of fiber-reinforced polymer composites. Traditional manual analysis is time-consuming, subjective, and unsuitable for high-throughput datasets. We present a graphical user interface (GUI) application that automates the analysis of microscopy images to extract key microstructural metrics, including fiber orientation tensors, fiber orientation distribution, porosity and pore size distribution. The app integrates multiple image segmentation techniques including global and local thresholding, clustering, and region-based approaches, offering flexibility for different types of image qualities and features. Users can load microstructural images, select regions of interest and segmentation techniques tailored to their image dataset. It also addresses a critical challenge in fiber orientation analysis: the ambiguities caused by touching, overlapping, or partially cut fibers. It supports autorun examples for standardized workflows, enabling reproducible analysis and facilitating training and benchmarking. This tool significantly reduces manual intervention, enhances consistency, and accelerates data generation for structure–property modeling, process optimization, and digital materials research. The tool is intended for use by materials scientists, engineers, and researchers engaged in composite characterization, quality control, and machine learning-based microstructural studies.

Chawla, Komal [ORNL] (ORCID:0000000190327565)↗

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES↗

ARMing the Edge: Designing Edge Computing–Capable Machine Learning Algorithms to Target ARM Doppler Lidar Processing

Abstract There is a need for long-term observations of cloud and precipitation fall speeds in validating and improving rainfall forecasts from climate models. To this end, the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) user facility Southern Great Plains (SGP) site at Lamont, Oklahoma, hosts five ARM Doppler lidars that can measure cloud and aerosol properties. In particular, the ARM Doppler lidars record Doppler spectra that contain information about the fall speeds of cloud and precipitation particles. However, due to bandwidth and storage constraints, the Doppler spectra are not routinely stored. This calls for the automation of cloud and rain detection in ARM Doppler lidar data so that the spectral data in clouds can be selectively saved and further analyzed. During the ARMing the Edge field experiment, a Waggle node capable of performing machine learning applications in situ was deployed at the ARM SGP site for this purpose. In this paper, we develop and test four algorithms for the Waggle node to automatically classify ARM Doppler lidar data. We demonstrate that supervised learning using a ResNet50-based classifier will classify 97.6% of the clear-air images and 94.7% of cloudy images correctly, outperforming traditional peak detection methods. We also show that a convolutional autoencoder paired with k -means clustering identifies 10 clusters in the ARM Doppler lidar data. Three clusters correspond to mostly clear conditions with scattered high clouds, and seven others correspond to cloudy conditions with varying cloud-base heights.

54 ENVIRONMENTAL SCIENCES↗