Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “supervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Reinforcement Learning for Intelligent Building Energy Management System Control *

A building energy management system (BEMS) is a computer-based system designed to monitor and control a building's energy needs. Modern BEMS rely on the sensing and connectivity capabilities of Internet of Things (IoT) technology to intelligently adjust the energy consumption to reduce cost while respecting the consumers' preferences. Increasingly, control decisions are made based on predictions by models trained using supervised machine learning methods, which still requires control policies to be formulated in a rule-based fashion. When using reinforcement learning (RL) instead, control policies are learned by observing the utility in terms of cost and comfort associated with actions such as a change in the heating system's setpoint. The resulting RL-based controllers can capture not only the dynamics of the building and the associated electrical devices, but also fluctuations in electricity prices and user demand, avoiding the need to combine multiple predictive models with tailored control policies. This chapter will provide an overview of RL-based approaches for BEMS. After sketching the taxonomy of general RL methods, we discuss the implications of relying on the individual methods in a BEMS context. Existing work applying RL is presented along the key devices controlled by BEMS systems. Finally, we summarize the state-of-the-art and sketch limitations and open research directions.

Kotevska, Olivera↗

Utilizing Machine Learning to Improve Neutralization Potency of an HIV-1 Antibody Targeting the gp41 N-Heptad Repeat

The N-heptad repeat (NHR) of the HIV-1 gp41 prehairpin intermediate (PHI) is an attractive potential vaccine target with high sequence conservation across diverse strains. However, despite the potency of NHR-targeting peptides and clinical efficacy of the NHR-targeting entry inhibitor enfuvirtide, no potently neutralizing NHR-directed monoclonal antibodies (mAbs) nor antisera have been identified or elicited to date. The lack of potent NHR-binding mAbs both dampens enthusiasm for vaccine development efforts at this target and presents a barrier to performing passive immunization experiments with NHR-targeting antibodies. To address this challenge, we previously developed an improved variant of the NHR-directed mAb D5, called D5_AR, which is capable of neutralizing diverse tier-2 viruses. Building on that work, here we present the 2.7Å-crystal structure of D5_AR bound to NHR mimetic peptide IQN17. We then utilize protein language models and supervised machine learning to generate small (n < 100) libraries of D5_AR variants that are subsequently screened for improved neutralization potency. We identify a variant with 5-fold improved neutralization potency, D5_FI, which is the most potent NHR-directed monoclonal antibody characterized to date and exhibits broad neutralization of tier-2 and −3 pseudoviruses as well as replicating R5 and X4 challenge strains. Additionally, our work highlights the ability of protein language models to efficiently identify improved mAb variants from relatively small libraries.

Biopolymers↗

Nowcasting Earthquakes: Imaging the Earthquake Cycle in California With Machine Learning

We propose a new machine learning-based method for nowcasting earthquakes to image the time-dependent earthquake cycle. The result is a timeseries that may correspond to the process of stress accumulation and release. The timeseries are constructed by using principal component analysis of regional seismicity. The patterns are found as eigenvectors of the cross-correlation matrix of a collection of seismicity timeseries in a coarse grained regional spatial grid (pattern recognition via unsupervised machine learning). The eigenvalues of this matrix represent the relative importance of the various eigenpatterns. Using the eigenvectors and eigenvalues, we compute the weighted correlation timeseries of the regional seismicity. This timeseries has the property that the weighted correlation generally decreases prior to major earthquakes in the region, and increases suddenly just after a major earthquake occurs. As in a previous paper, we find that this method produces a nowcasting timeseries that resembles the hypothesized regional stress accumulation and release process characterizing the earthquake cycle. We then address the problem of whether the timeseries contain information regarding future large earthquakes. For this, we compute a receiver operating characteristic and determine the decision thresholds for several future time periods of interest (optimization via supervised machine learning). We find that signals can be detected that can be used to characterize the information content of the timeseries. These signals may be useful in assessing present and near-future seismic hazards.

58 GEOSCIENCES↗

Unveiling X-ray absorption signatures of boron nitride via first-principles simulation and machine learning

Boron nitride (BN) allotropes hold great promise in many advanced applications ranging from optical and photonic devices to energy storage and battery systems to tribological components. The diverse functionalities of this material stem from BN’s highly tunable structural and electronic properties, which are governed by the versatile boron–nitrogen bonding configurations. Exploring the structural landscape of BN can unveil novel structures possessing unique properties suited for specific applications, therefore accelerating the design of next-generation advanced functional materials. In this work, we leverage boron K-edge X-ray absorption spectroscopy (XAS) as an effective probe for local structural features and chemical environments. A total of 210 BN crystal structures are generated via analogies to the extensive array of carbon allotropes, and XAS is simulated for each unique local motif within the resulting collection of structures. A mapping between structural features and spectral signatures was established by synergizing first-principle simulations with data-driven based post-analysis approaches. Specifically, we developed a neural network model that can satisfactorily predict spectra line shapes from local structural descriptors. Toward automatic spectroscopic interpretation of any new BN structures, supervised machine learning models, trained on this structure–spectrum dataset, can accurately infer local coordination environments from simulated XAS, highlighting the strength of this unique approach of combining high-fidelity first-principles simulation and machine-learning to accelerate target design of novel BN materials via rational understanding of local structure-spectrum correlations.

36 MATERIALS SCIENCE↗

Explainable machine learning for incipient anomaly detection in compact molten salt heat exchanger with overlapping feature distributions

High-temperature molten salt-cooled reactors (MSCRs) are a promising next-generation nuclear technology option, offering efficient power conversion and inherent safety features. However, the reliability of these systems depends on the robust operation of heat exchangers (HXs), which are susceptible to failure due to temperature gradients and channel plugging caused by fluid freezing. Conventional monitoring methods, relying on inlet and outlet measurements, lack the spatial resolution needed to detect early-stage faults. We propose a novel design of a compact salt-to-salt matrix-type HX design consisting of interleaved arrays of parallel tubes, with integrated synthetic fiber optic distributed temperature sensing (DTS) to enable localized detection of incipient faults. To evaluate performance of this design, we generate high-fidelity synthetic data using heat transfer computational modeling to simulate channel plugging, and introduce sensor noise for realistic modeling of measurements. The dataset comprises of 97% normal operation and 3% anomaly cases, with each anomaly class representing 1% of the data. These early anomalies result in overlapping temperature profiles between normal and faulty channels, producing a non-separable dataset that challenges traditional classification techniques. We benchmark eight supervised machine learning (ML) models and demonstrate that XGBoost achieves the highest performance. To improve transparency, we develop an explainability framework combining Shapley values and partially ordered sets (POSETs) to quantify and structurally analyze feature importance. This approach identifies both dominant predictors and ambiguous feature relationships, enhancing trust and interpretability. Our results highlight the potential of combining DTS and explainable ML with intelligent feature selection to improve predictive maintenance and ensure operational resilience in advanced nuclear systems.

Prantikos, Konstantinos [Argonne National Laborato↗

HIGH-LOW FIDELITY THERMAL HYDRAULIC COUPLING USING AI/MACHINE LEARNING ALGORITHMS

The primary goal of the US Department of Energy (DOE) office of Nuclear Energy Integrated Energy Systems (IES) program is to develop the tools and framework for coupling multi-scale and multi-physical thermal and electrical energy usage and storage systems. High- and low-fidelity (high–low) coupling is a key feature of multi-scale, multi-component systems and has been an important focus of research in the nuclear energy community for the past two decades. An essential feature of demonstrating the capability to couple high-fidelity and low-fidelity systems for real-time applications are surrogate/reduced order models (ROM). For the purposes of this study, surrogate models are essentially Blackbox models, typically developed using supervised Machine learning (ML) algorithms. The surrogate models can be used to mimic the response of high-fidelity models to represent large historical datasets and coupled with more general low-fidelity system models distributed as Functional Mock-up Interface (FMI) or Functional Mock-up Units (FMU) modules. The example is demonstrated with Spallation Neutron Source (SNS) First Target Station flow loop data. The flow loop is a liquid mercury loop with a pump, piping, heat exchange, and internal heat generation in the target window. This work elucidates some of the potential benefits and future needs of developing tools for high–low system coupling of energy systems.

Williams, Wesley↗

A neural network for determination of latent dimensionality in Nonnegative Matrix Factorization

Non-negative Matrix Factorization (NMF) has proven to be a powerful unsupervised learning method for uncovering hidden features in complex and noisy datasets with applications in data mining, text recognition, dimension reduction, face recognition, anomaly detection, blind source separation, and many other fields. An important input for NMF is the latent dimensionality of the data, that is, the number of hidden features, K, present in the explored dataset. Unfortunately, and this quantity is rarely known a priori. The existing methods for determining latent dimensionality, such as Automatic Relevance Determination (ARD), are mostly heuristic and utilize different characteristics to estimate the number of hidden features. However, all of them require human presence to make a final determination of K. Here we utilize a supervised machine learning approach in combination with a recent method for model determination, called NMFk, to determine the number of hidden features automatically. NMFk performs a set of NMF simulations on an ensemble of matrices, obtained by bootstrapping the initial dataset, and estimates which K produces stable groups of latent features that reconstruct the initial dataset well. We then train a Multi-Layter Perceptron (MLP) classifier network to determine the correct number of latent features utilizing the statistics and characteristics of the NMF solution, obtained from NMFk. In order to train the MLP classifier, a training set of 58,660 matrices with predetermined latent features were factorized with NMFk. The MLP classifier in conjunction with NMFk maintains a greater than 95% success rate when applied to a held out test set. Additionally, when applied to two well-known benchmark datasets, the swimmer and MIT face data, NMFk/MLP correctly recovers the established number of hidden features. Finally, we compare the accuracy of our method to the ARD, AIC and Stability-based methods.

97 MATHEMATICS AND COMPUTING↗

Heterogeneous Multilayer Nanopores via Chemically Tuned Dielectric Breakdown for Single‐Molecule Sensing

Solid-state nanopores are powerful platforms for single-molecule sensing, yet their performance is often constrained by fabrication complexity, noise, and limited control over surface properties. Here we report a direct method to fabricate heterogeneous multilayer nanopores using chemically tuned controlled dielectric breakdown (CT-CDB). We integrate hBN, MoS 2 , or graphene atop a silicon nitride membrane to form five distinct bilayer and tri-layer architectures, with bare SiN x nanopore as a control. CT-CDB achieves pore formation reproducibly through material-stacks with high efficiency, good pore size control, and strong yield, validated by various characterizations. Transferrin protein translocation experiments, supported by simulations, reveal that multilayer configurations modulate protein conformations, ionic current blockade and dwell time distributions, reflecting combined effects of membrane type, interfacial chemistry, and local electric field gradients. A supervised machine learning framework is implemented to assist identifying multilayer structure effects embedded in signal signatures, with over 96% accuracy. This work presents a modular and scalable framework for functional nanopore engineering with complex structural integration, thereby expanding the potential of 2D materials in single-molecule sensing applications.

2D materials↗

Extrusion parameter control optimization for DIW 3D printing using image analysis techniques

Material extrusion is a well-recognized facet of additive manufacturing that involves the fabrication of parts through the deposition of structural material from an extrusion head from a bulk supply. In the subdivision of Direct Ink Writing (DIW) additive manufacturing, challenges arise when the structural material is flowable, synchronous extrusion control and tool movement becomes critical for achieving high-quality parts with low defect populations. DIW techniques are most used in laboratory settings using expensive custom instruments and may require specialized 3D slicing software. Here, in this study, the fabrication of an inexpensive, consumer-friendly progressive cavity pump dispensing system is detailed, in which can create high-quality parts by executing G-code commands produced from a commercial slicing software. The precision and repeatability of the movement-synchronized material extrusion is demonstrated through a series of optimization schemes, entailing the alteration of various control parameters, which directly affect the extrusion properties demonstrated during a print. In situ diagnostics were implemented to evaluate the results of the established optimization experiment. Using a machine vision technique, images of the optimization prints are processed. Following this, a supervised machine learning model was trained to autonomously judge whether or not the extrusion parameters produced a passing or failing result. The machine learning scheme serves as a preliminary benchmark for future layer-by-layer evaluation of more complex DIW parts. The construction of the printer and development of in situ characterization capabilities demonstrates the ability for this printer to create high-fidelity DIW parts for a fraction of the price of other systems.

42 ENGINEERING↗

Autoencoder-Based Anomaly Detection System for Online Data Quality Monitoring of the CMS Electromagnetic Calorimeter

The CMS detector is a general-purpose apparatus that detects high-energy collisions produced at the LHC. Online data quality monitoring of the CMS electromagnetic calorimeter is a vital operational tool that allows detector experts to quickly identify, localize, and diagnose a broad range of detector issues that could affect the quality of physics data. A real-time autoencoder-based anomaly detection system using semi-supervised machine learning is presented enabling the detection of anomalies in the CMS electromagnetic calorimeter data. A novel method is introduced which maximizes the anomaly detection performance by exploiting the time-dependent evolution of anomalies as well as spatial variations in the detector response. The autoencoder-based system is able to efficiently detect anomalies, while maintaining a very low false discovery rate. The performance of the system is validated with anomalies found in 2018 and 2022 LHC collision data. In addition, the first results from deploying the autoencoder-based system in the CMS online data quality monitoring workflow during the beginning of Run 3 of the LHC are presented, showing its ability to detect issues missed by the existing system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Hydropower potential derived from streamflow extremes for Alaska, USA

Alaska is an expansive region known for its abundant natural resources, including thousands of miles of streams and rivers. These rivers represent potential opportunities for future hydropower development that could provide reliable energy supply for local communities. There is limited long-term high temporal resolution streamflow data available for the region, making data-driven estimates of potential hydropower and its variability across the state challenging. This study provides a novel data-driven approach for hydropower capacity estimation across Alaska. We use supervised machine learning to develop a relationship between the daily and peak flow duration curves in order to augment the size of our dataset from 44 sites to 67 sites. We perform a stochastic hydropower estimation across the 67 sites and identify approximately 1000 MW of total potential hydropower capacity distributed across these sites. Our study provides the first step towards more comprehensive hydropower estimation for this critical region, highlighting the need for future work integrating high-resolution spatial data, community needs, and economic constraints in estimates of potential hydropower development in Alaska.

Hydropower↗

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗

Probing Sulfur Chemical and Electronic Structure with Experimental Observation and Quantitative Theoretical Prediction of Kα and Valence-to-Core Kβ X-ray Emission Spectroscopy

An extensive experimental and theoretical study of the Kα and Kβ high-resolution X-ray emission spectroscopy (XES) of sulfur-bearing systems is presented here. This study encompasses a wide range of organic and inorganic compounds, including numerous experimental spectra from both prior published work and new measurements. Employing a linear-response time-dependent density functional theory (LR-TDDFT) approach, strong quantitative agreement is found in the calculation of energy shifts of the core-to-core Kα as well as the full range of spectral features in the valence-to-core Kβ spectrum. The ability to accurately calculate the sulfur Kα energy shift supports the use of sulfur Kα XES as a bulk-sensitive tool for assessing sulfur speciation. The fine structure of the sulfur Kβ spectrum, in conjunction with the theoretical results, is shown to be sensitive to the local electronic structure including effects of symmetry, ligand type and number, and, in the case of organosulfur compounds, to the nature of the bonded organic moiety. This agreement between theory and experiment, augmented by the potential for high-access XES measurements with the latest generation of laboratory-based spectrometers, demonstrates the possibility of broad analytical use of XES for sulfur and nearby third-row elements. The effective solution of the forward problem, i.e., successful prediction of detailed spectra from known molecular structure, also suggests future use of supervised machine learning approaches to experimental inference, as has seen recent interest for interpretation of X-ray absorption near-edge structure (XANES).

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Enhancing Nanoparticle Detection in Interferometric Scattering (iSCAT) Microscopy Using a Mask R-CNN

Interferometric scattering microscopy (iSCAT) is a label-free optical microscopy technique that enables imaging of individual nano-objects such as nanoparticles, viruses, and proteins. Essential to this technique is the suppression of background scattering and identification of signals from nano-objects. In the presence of substrates with high roughness, scattering heterogeneities in the background, when coupled with tiny stage movements, cause features in the background to be manifested in background-suppressed iSCAT images. Traditional computer vision algorithms detect these background features as particles, limiting the accuracy of object detection in iSCAT experiments. Here, in this paper, we present a pathway to improve particle detection in such situations using supervised machine learning via a mask region-based convolutional neural network (mask R-CNN). Using a model iSCAT experiment of 19.2 nm gold nanoparticles adsorbing to a rough layer-by-layer polyelectrolyte film, we develop a method to generate labeled datasets using experimental background images and simulated particle signals and train the mask R-CNN using limited computational resources via transfer learning. We then compare the performance of the mask R-CNN trained with and without inclusion of experimental backgrounds in the dataset against that of a traditional computer vision object detection algorithm, Haar-like feature detection, by analyzing data from the model experiment. Results demonstrate that including representative backgrounds in training datasets improved the mask R-CNN in differentiating between background and particle signals and elevated performance by markedly reducing false positives. The methodology for creating a labeled dataset with representative experimental backgrounds and simulated signals facilitates the application of machine learning in iSCAT experiments with strong background scattering and thus provides a useful workflow for future researchers to improve their image processing capabilities.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Maximum Entropy Theory of Multiscale Coarse-Graining via Matching Thermodynamic Forces: Application to a Molecular Crystal (TATB)

The MSCG/FM (multiscale coarse-graining via force-matching) approach is an efficient supervised machine learning method to develop microscopically informed coarse-grained (CG) models. Here we present a theory based on the principle of maximum entropy (PME) enveloping the existing MSCG/FM approaches. This theory views the MSCG/FM method as a special case of matching the thermodynamic forces from the extended ensemble described by the set of thermodynamic (relevant) system coordinates. This set may include CG coordinates, the stress tensor, applied external fields, and so forth, and may be characterized by nonequilibrium conditions. Following the presentation of the theory, we discuss the consistent matching of both bonded and nonbonded interactions. The proposed PME formulation is used as a starting point to extend the MSCG/FM method to the constant strain ensemble, which together with the explicit matching of the bonded forces is better suited for coarse-graining anisotropic media at a submolecular resolution. The theory is demonstrated by performing the fine coarse-graining of crystalline 1,3,5-triamino-2,4,6-trinitrobenzene (TATB), a well-known insensitive molecular energetic material, which exhibits highly anisotropic mechanical properties.

1,3,5-triamino-2,4,6-trinitrobenzene↗

PreMevE Update: Forecasting Ultra-Relativistic Electrons Inside Earth's Outer Radiation Belt

Energetic electrons inside Earth's Van Allen belts pose a major radiation threat to spaceborne electronics that often play vital roles in modern society. Ultra-relativistic electrons with energies greater than or equal to two megaelectron-volt (MeV) are of particular interest, and thus forecasting these ≥2 MeV electrons has a significant meaning to all space sectors. Here, we update the latest development of the predictive model for MeV electrons in the outer radiation belt. The new version, called PREdictive MEV Electron (PreMevE)-2E, forecasts ultra-relativistic electron flux distributions across the outer belt, with no need for in situ measurements of the trapped MeV electron population except at the geosynchronous orbit (GEO). Model inputs include precipitating electrons observed in low-Earthorbits by NOAA satellites, upstream solar wind speeds and densities from solar wind monitors, as well as ultra-relativistic electrons measured by one Los Alamos GEO satellite. We evaluated 32 supervised machine learning models that fall into four different classes of linear and neural network architectures, and successfully tested ensemble forecasting by using groups of top-performing models. All models are individually trained, validated, and tested by in situ electron data from NASA's Van Allen Probes mission. It is shown that the final ensemble model outperforms individual models at most L-shells, and this PreMevE-2E model can provide 25-h (~1-day) and 50-h (~2-day) forecasts with high mean performance efficiency and correlation values. Our results also suggest that this new model is dominated by nonlinear components at L-shells <~4 for ultra-relativistic electrons, different from the dominance of linear components for 1 MeV electrons as previously discovered.

79 ASTRONOMY AND ASTROPHYSICS↗

Neural Network‐Based Methods for Ocean Surface Wave Measurement Using Submarine Distributed Acoustic Sensing (DAS)

Two new data-driven models for estimating ocean surface waves from distributed acoustic sensing (DAS) submarine cable strain rate are developed using supervised machine learning on a 10-day data set collected offshore of Oliktok Point, Alaska. The new models were trained on target data from seafloor pressure moorings at three sites spaced evenly along 27.1 km of cable and were benchmarked against an empirical transfer function method previously used to estimate waves from DAS. A model which uses convolutional neural networks to transform 2-km frequency-wavenumber strain spectra to seafloor pressure spectra outperforms the benchmark in wave height prediction (RMSE of 0.15 vs. 0.41 m) and period prediction (0.29 vs. 0.37 s) when evaluated on a held-out test data set. When applied to a DAS data set collected on the same cable 2 years prior, the CNN-based model maintained similar significant wave height performance (RMSE = 0.23 m) relative to available satellite altimetry data. A two-hidden-layer, fully connected neural network which transforms 1-D strain spectra to seafloor pressure spectra also outperforms the benchmark in wave height prediction (RMSE of 0.19 vs. 0.41 m), but does not generalize as well to the prior data. Regression-based machine learning is useful for estimating waves from DAS data when the pressure-strain relationship varies temporally and spatially across different wave conditions. Models can be applied to DAS data to measure waves with higher spatial resolution and longer temporal coverage than traditional methods, which often measure waves only at a single point.

Davis, Jacob R. [Univ. of Washington, Seattle, WA ↗

Structure of chalcogen overlayers on Au(111): Density functional theory and lattice-gas modeling

Ordering of different chalcogens, S, Se, and Te, on Au(111) exhibit broad similarities but also some distinct features, which must reflect subtle differences in relative values of the long-range pair and many-body lateral interactions between adatoms. We develop lattice-gas (LG) models within a cluster expansion framework, which includes about 50 interaction parameters. These LG models are developed based on density functional theory (DFT) analysis of the energetics of key adlayer configurations in combination with the Monte Carlo (MC) simulation of the LG models to identify statistically relevant adlayer motifs, i.e., model development is based entirely on theoretical considerations. The MC simulation guides additional DFT analysis and iterative model refinement. Given their complexity, development of optimal models is also aided by strategies from supervised machine learning. The model for S successfully captures ordering motifs over a broader range of coverage than achieved by previous models, and models for Se and Te capture the features of ordering, which are distinct from those for S. More specifically, the modeling for all three chalcogens successfully explains the linear adatom rows (also subtle differences between them) observed at low coverages of ~0.1 monolayer. The model for S also leads to a new possible explanation for the experimentally observed phase with a (5 × 5)-type low energy electron diffraction (LEED) pattern at 0.28 ML and to predictions for LEED patterns that would be observed with Se and Te at this coverage.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗