Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “binary neural networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Gamma Ray Source Localization for Time Projection Chamber Telescopes Using Convolutional Neural Networks

Diverse phenomena such as positron annihilation in the Milky Way, merging binary neutron stars, and dark matter can be better understood by studying their gamma ray emission. Despite their importance, MeV gamma rays have been poorly explored at sensitivities that would allow for deeper insight into the nature of the gamma emitting objects. In response, a liquid argon time projection chamber (TPC) gamma ray instrument concept called GammaTPC has been proposed and promises exploration of the entire sky with a large field of view, large effective area, and high polarization sensitivity. Optimizing the pointing capability of this instrument is crucial and can be accomplished by leveraging convolutional neural networks to reconstruct electron recoil paths from Compton scattering events within the detector. In this investigation, we develop a machine learning model architecture to accommodate a large data set of high fidelity simulated electron tracks and reconstruct paths. We create two model architectures: one to predict the electron recoil track origin and one for the initial scattering direction. We find that these models predict the true origin and direction with extremely high accuracy, thereby optimizing the observatory’s estimates of the sky location of gamma ray sources.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Reducing the Parameter Dependency of Phase-Picking Neural Networks with Dice Loss

Training a neural network for picking seismic phase arrivals has been commonly posed as a segmentation problem. It is a highly imbalanced segmentation problem in the sense that the background vastly dominates the foreground because we are trying to pick the optimal single sample point that represents the arrival of a seismic phase in a many seconds long time window. Here, we test the Dice loss, which is a preferred loss function for highly imbalanced image segmentation problems. We show that phase-picking neural networks trained on the Dice loss behave in a binary fashion for which the prediction output is almost always either nearly 1 or nearly 0. This feature removes the strong dependence of data processing workflows on the prediction score threshold, which is an otherwise critical parameter to determine when using neural networks trained on the cross-entropy loss. When strategically used, models trained on the Dice loss can reduce the parameter dependency of machine learning-based seismic monitoring.

58 GEOSCIENCES↗

A Neural Network Approach to Predict Gibbs Free Energy of Ternary Solid Solutions

Here, we present a data-centric deep learning (DL) approach using neural networks (NNs) to predict the thermodynamics of ternary solid solutions. We explore how NNs can be trained with a dataset of Gibbs free energies computed from a CALPHAD database to predict ternary systems as a function of composition and temperature. We have chosen the energetics of the FCC solid solution phase in 226 binaries consisting of 23 elements at 11 different temperatures to demonstrate the feasibility. The number of binary data points included in the present study is 102,000. We select six ternaries to augment the binary dataset to investigate their influence on the NN prediction accuracy. We examine the sensitivity of data sampling on the prediction accuracy of NNs over selected ternary systems. It is anticipated that the current DL workflow can be further elevated by integrating advanced descriptors beyond the elemental composition and more curated training datasets to improve prediction accuracy and applicability.

42 ENGINEERING↗

Application of Convolutional and Feedforward Neural Networks for Fault Detection in Particle Accelerator Power Systems

High voltage converter modulators (HVCM) provide power to the accelerating cavities of the spallation neutron source (SNS) facility. HVCM experience catastrophic failures, which increase the downtime of the SNS and reduce beam time. The faults may occur due to different reasons including failures of the resonant capacitor, core saturation due to the magnetic flux, insulated-gate bipolar transistor (IGBT) failures, and others. We recently have setup a HVCM test stand to develop and test machine learning models for anomaly detection and fault prognostics. In this work, we propose binary classifiers and autoencoder architectures based on convolutional (CNN) and feedforward neural networks (FNN) to facilitate distinguishing normal from faulty waveforms coming from the HVCM during operation. The results indicate that the CNN binary classifier is the best model among the four showing very stable performance in the training and testing sets with impressive metrics of precision and recall reaching up to 99\% with a very small uncertainty. The FNN classifier shows the least performance with a large uncertainty in its metrics. The performances of the two autoencoders based on CNN and FNN were in between, showing very good performance nonetheless.

Radaideh, Majdi↗

Trigger Detection for the sPHENIX Experiment via Bipartite Graph Networks with Set Transformer

Trigger (interesting events) detection is crucial to high-energy and nuclear physics experiments because it improves data acquisition efficiency. It also plays a vital role in facilitating the downstream offline data analysis process. The sPHENIX detector, located at the Relativistic Heavy Ion Collider in Brookhaven National Laboratory, is one of the largest nuclear physics experiments on a world scale and is optimized to detect physics processes involving charm and beauty quarks. Furthermore, these particles are produced in collisions involving two proton beams, two gold nuclei beams, or a combination of the two and give critical insights into the formation of the early universe. This paper presents a model architecture for trigger detection with geometric information from two fast silicon detectors. Transverse momentum is introduced as an intermediate feature from physics heuristics. We also prove its importance through our training experiments. Each event consists of tracks and can be viewed as a graph. A bipartite graph neural network is integrated with the attention mechanism to design a binary classification model. Compared with the state-of-the-art algorithm for trigger detection, our model is parsimonious and increases the accuracy and the AUC score by more than 15%.

97 MATHEMATICS AND COMPUTING↗

Hybrid Approaches for Data Reduction of Spatiotemporal Scientific Applications

Scientists conduct large-scale simulations to compute derived quantities from primary data. Thus, it is crucial that data compression techniques maintain bounded errors on these derived quantities or quantities of interest (QOI). For many spatiotemporal applications, these QOIs are binary in nature and represent presence or absence of a physical phenomenon. In this work, we propose to use a hybrid approah for differential compression for such applications. We use a neural network (NN) approach to determine regions-of-interest (ROIs) where the binary QOIs are going to be prevalent. This is then used with traditional approaches that compress at a lower level (and higher accuracy) for these ROIs as compared to other regions.

Li, Xiao↗

A Comparison between Invariant and Equivariant Classical and Quantum Graph Neural Networks

Machine learning algorithms are heavily relied on to understand the vast amounts of data from high-energy particle collisions at the CERN Large Hadron Collider (LHC). The data from such collision events can naturally be represented with graph structures. Therefore, deep geometric methods, such as graph neural networks (GNNs), have been leveraged for various data analysis tasks in high-energy physics. One typical task is jet tagging, where jets are viewed as point clouds with distinct features and edge connections between their constituent particles. The increasing size and complexity of the LHC particle datasets, as well as the computational models used for their analysis, have greatly motivated the development of alternative fast and efficient computational paradigms such as quantum computation. In addition, to enhance the validity and robustness of deep networks, we can leverage the fundamental symmetries present in the data through the use of invariant inputs and equivariant layers. In this paper, we provide a fair and comprehensive comparison of classical graph neural networks (GNNs) and equivariant graph neural networks (EGNNs) and their quantum counterparts: quantum graph neural networks (QGNNs) and equivariant quantum graph neural networks (EQGNN). The four architectures were benchmarked on a binary classification task to classify the parton-level particle initiating the jet. Based on their area under the curve (AUC) scores, the quantum networks were found to outperform the classical networks. However, seeing the computational advantage of quantum networks in practice may have to wait for the further development of quantum technology and its associated application programming interfaces (APIs).

Forestano, Roy T. (ORCID:0000000203552076)↗

A Comprehensive Comparative Study of Active Learning Schemes for Nanophotonics Design

We present a benchmarking study of active learning (AL) schemes for designing planar multilayer nanophotonic metamaterials, where the design tasks are formulated as binary optimization problems. Different surrogate models, including factorization machine (FM), Gaussian process regression (GPR), and convolutional neural network (CNN), combined with different optimization methods, including exhaustive enumeration, discrete particle swarm optimization (DPSO), quantum annealing (QA), hybrid QA, and simulated annealing are studied. The benchmark cases investigated range from small problems with short binary lengths (N = 25) to large problems with N up to 100, focusing on the design of two classes of photonic structures, including antireflective coatings for the long-wavelength infrared region and transparent radiative coolers. For small problems, CNN coupled with DPSO in AL achieves the best performance. As N increases, FM with QA outperforms GPR and CNN. For FM-based AL, hybrid QA yields the best optimization results, particularly in high-dimensional cases (N = 100). These results demonstrate that the optimization method can significantly affect in AL performance as N increases, and that QA-based optimization can provide practical routes for mitigating the optimization bottleneck in high-dimensional problems.

Jung, Serang [Kyung Hee University, Korea]↗

Many-body expansion based machine learning models for octahedral transition metal complexes

Abstract Graph-based machine learning (ML) models for material properties show great potential to accelerate virtual high-throughput screening of large chemical spaces. However, in their simplest forms, graph-based models do not include any 3D information and are unable to distinguish stereoisomers such as those arising from different orderings of ligands around a metal center in coordination complexes. In this work we present a modification to revised autocorrelation descriptors, a molecular graph featurization method, for predicting spin state dependent properties of octahedral transition metal complexes (TMCs). Inspired by analytical semi-empirical models for TMCs, the new modeling strategy is based on the many-body expansion (MBE) and allows one to tune the captured stereoisomer information by changing the truncation order of the MBE. We present the necessary modifications to include this approach in two commonly used ML methods, kernel ridge regression and feed-forward neural networks. On a test set composed of all possible isomers of binary TMCs, the best MBE models achieve mean absolute errors (MAEs) of 2.75 kcal mol −1 on spin-splitting energies and 0.26 eV on frontier orbital energy gaps, a 30%–40% reduction in error compared to models based on our previous approach. We also observe improved generalization to previously unseen ligands where the best-performing models exhibit MAEs of 4.00 kcal mol −1 (i.e. a 0.73 kcal mol −1 reduction) on the spin-splitting energies and 0.53 eV (i.e. a 0.10 eV reduction) on the frontier orbital energy gaps. Because the new approach incorporates insights from electronic structure theory, such as ligand additivity relationships, these models exhibit systematic generalization from homoleptic to heteroleptic complexes, allowing for efficient screening of TMC search spaces.

Meyer, Ralf (ORCID:0000000322360261)↗

End-to-End Pipeline for Trigger Detection on Hit and Track Graphs

There has been a surge of interest in applying deep learning in particle and nuclear physics to replace labor-intensive offline data analysis with automated online machine learning tasks. This paper details a novel AI-enabled triggering solution for physics experiments in Relativistic Heavy Ion Collider and future Electron-Ion Collider. The triggering system consists of a comprehensive end-to-end pipeline based on Graph Neural Networks that classifies trigger events versus background events, makes online decisions to retain signal data, and enables efficient data acquisition. Here, the triggering system first starts with the coordinates of pixel hits lit up by passing particles in the detector, applies three stages of event processing (hits clustering, track reconstruction, and trigger detection), and labels all processed events with the binary tag of trigger versus background events. By switching among different objective functions, we train the Graph Neural Networks in the pipeline to solve multiple tasks: the edge-level track reconstruction problem, the edge-level track adjacency matrix prediction, and the graph-level trigger detection problem. We propose a novel method to treat the events as track-graphs instead of hit-graphs. This method focuses on intertrack relations and is driven by underlying physics processing. As a result, it attains a solid performance (around 72% accuracy) for trigger detection and outperforms the baseline method using hit-graphs by 2% higher accuracy.

97 MATHEMATICS AND COMPUTING↗

Machine learning for precise hit position reconstruction in Resistive Silicon Detectors

RSDs are LGAD silicon sensors with 100% fill factor, based on the principle of AC-coupled resistive read-out. Signal sharing and internal charge multiplication are the RSD key features to achieve picosecond-level time resolution and micron-level spatial resolution, thus making these sensors promising candidates as 4D-trackers for future experiments. This paper describes the use of a neural network to reconstruct the hit position of ionizing particles, an approach that can boost the performance of the RSD with respect to analytical models. The neural network has been trained in the laboratory and then validated on test beam data. The device-under-test in this work is a 450 μm-pitch matrix from the FBK RSD2 production, which achieved a resolution of about 65 μm at the DESY Test Beam Facility, a 50% improvement compared to a simple analytical reconstruction method, and a factor two better than the resolution of a standard pixel sensor of equal pitch size with binary read-out. The test beam result is compatible with the laboratory ones obtained during the neural network training, confirming the ability of the machine learning model to provide accurate predictions even in environments very different from the training one. Prospects for future improvements are also discussed.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

ℤ2 × ℤ2 Equivariant Quantum Neural Networks: Benchmarking against Classical Neural Networks

This paper presents a comparative analysis of the performance of Equivariant Quantum Neural Networks (EQNNs) and Quantum Neural Networks (QNNs), juxtaposed against their classical counterparts: Equivariant Neural Networks (ENNs) and Deep Neural Networks (DNNs). We evaluate the performance of each network with three two-dimensional toy examples for a binary classification task, focusing on model complexity (measured by the number of parameters) and the size of the training dataset. Our results show that the Z2×Z2 EQNN and the QNN provide superior performance for smaller parameter sets and modest training data samples.

Dong, Zhongtian (ORCID:0000000210003454)↗

Efficient human activity recognition with spatio-temporal spiking neural networks

In this study, we explore Human Activity Recognition (HAR), a task that aims to predict individuals' daily activities utilizing time series data obtained from wearable sensors for health-related applications. Although recent research has predominantly employed end-to-end Artificial Neural Networks (ANNs) for feature extraction and classification in HAR, these approaches impose a substantial computational load on wearable devices and exhibit limitations in temporal feature extraction due to their activation functions. To address these challenges, we propose the application of Spiking Neural Networks (SNNs), an architecture inspired by the characteristics of biological neurons, to HAR tasks. SNNs accumulate input activation as presynaptic potential charges and generate a binary spike upon surpassing a predetermined threshold. This unique property facilitates spatio-temporal feature extraction and confers the advantage of low-power computation attributable to binary spikes. We conduct rigorous experiments on three distinct HAR datasets using SNNs, demonstrating that our approach attains competitive or superior performance relative to ANNs, while concurrently reducing energy consumption by up to 94%.

60 APPLIED LIFE SCIENCES↗

In-situ sensor monitoring of multi-class gas porosity formation in laser powder bed fusion using convolutional neural network

In-situ monitoring of defect formation remains a significant challenge in the laser powder bed fusion (LPBF) process. Recent advances have enabled real-time defect detection with machine learning and in-situ sensing technologies; however, most studies focus on binary classification of keyhole pores, limiting nuanced multi-class pore differentiation and formation mechanisms. This work introduces a multi-class pore detection framework (no pore, small pores < 15 µm, and large pores > 15 µm) by leveraging photodiode sensor data alongside high-fidelity synchrotron X-ray imaging. The 15 µm threshold is selected to distinguish between two fundamentally different defect mechanisms, following the physical size-mechanism boundary established by prior high-resolution synchrotron X-ray characterization of Al6061 LPBF. Distinguishing these classes is critical because large keyhole pores are structurally detrimental, whereas small gas pores are often benign, requiring different process control strategies. Thermal emission monitoring data collected simultaneously with high-speed X-ray imaging at the Stanford Synchrotron Radiation Lightsource (SSRL), are correlated with subsurface melt pool dynamics to establish ground truth. Continuous Wavelet Transform (CWT) with optimized parameters converts the photodiode time-series signals into time–frequency images, facilitating feature extraction. Convolutional Neural Networks (CNN) are then applied for real-time multi-class pore classification in an average inference time of 1 ms per signal window. It achieves 79% accuracy and an Area Under the Receiver Operating Characteristic curve (AUC ROC) score of 0.89 with five-fold cross-validation. The results demonstrate that coupling CWT-based feature engineering with CNN architecture enables reliable multi-class pore detection in Al6061 builds using affordable in-situ sensors. This approach advances scalable and affordable quality assurance in additive manufacturing by moving beyond binary defect detection toward more nuanced classification of porosity mechanisms with in-situ sensors and machine learning.

Laser powder bed fusion, Multi-class pores, In-sit↗

Extending Power of Nature from Binary Problems to Real-Valued Graph Learning in Real World

Nature performs complex computations constantly at clearly lower cost and higher performance than digital computers. It is crucial to understand how to harness the unique computational power of nature in Machine Learning (ML). In the past decade, besides the development of Neural Networks (NNs), the community has also relentlessly explored nature-powered ML paradigms. Although most of them are still predominantly theoretical, a new practical paradigm enabled by the recent advent of CMOS-compatible room-temperature nature-based computers has emerged. By harnessing the nature's power of entropy increase, this paradigm can solve binary learning problems delivering immense speedup and energy savings compared with NNs, while maintaining comparable accuracy. Regrettably, its values to the real world are highly constrained by its binary nature. A clear pathway to its extension to real-valued problems remains elusive. This paper aims to unleash this pathway by proposing a novel end-to-end Nature-Powered Graph Learning (NP-GL) framework. Specifically, through a three-dimensional co-design, NP-GL can leverage the nature's power of entropy increase to efficiently solve real-valued graph learning problems. Experimental results across 4 real-world applications with 6 datasets demonstrate that NP-GL delivers, on average, 6970X speedup and 10^5x energy consumption reduction with comparable or even higher accuracy than Graph Neural Networks (GNNs).

artificial intelligence↗

Machine Learning Techniques for Data Reduction of Climate Applications

Scientists conduct large-scale simulations to compute derived quantities-of-interest (QoI) from primary data. Often, QoI are linked to specific features, regions, or time intervals, such that data can be adaptively reduced without compromising the integrity of QoI. For many spatiotemporal applications, these QoI are binary in nature and represent presence or absence of a physical phenomenon. We present a pipelined compression approach that first uses neural-network-based techniques to derive regions where QoI are highly likely to be present. Then, we employ a Guaranteed Autoencoder (GAE) to compress data with differential error bounds. GAE uses QoI information to apply low-error compression to only these regions. This results in overall high compression ratios while still achieving downstream goals of simulation or data collections. Experimental results are presented for climate data generated from the E3SM Simulation model for downstream quantities such as tropical cyclone and atmospheric river detection and tracking. These results show that our approach is superior to comparable methods in the literature.

Li, Xiao [University of Florida]↗

Optimisation of the Kaplan hydropower system via PID 2 and digital twin

Here, this paper proposes a proportional–integral-double–derivative (PID 2 ) optimisation method for the Kaplan hydropower system by building a digital twin. The study first uses one multilayer perceptron (MLP) to model the hydroturbine dynamic and then adopts three connected MLPs to model the generator dynamic, both in an open-loop fashion. Inspired by stochastic distribution control (SDC) theory, we regard the training of the turbine's neural network model as a process control problem, and we propose minimising entropy loss to update the network parameters. The next step is to build the digital twin by connecting the neural network models with a PID 2 controller and a lead-lag exciter and run the whole model in a closed-loop fashion. After that, a binary search approach is applied to optimise the PID 2 parameters based on the obtained digital twin model. The simulation results show that the proposed method can reduce the mean square tracking error by more than 90%. Furthermore, the method is extended to jointly optimise the PID 2 controller and excitation system gains through multiobjective optimisation, leveraging Pareto frontier analysis to balance active power and voltage tracking performance. Simulation results confirm the effectiveness of the proposed method, achieving a 83.46% reduction in relative mean square error of active power, a 47.13% reduction in terminal voltage tracking error, and an 82.78% improvement in the overall scalarized objective.

Hydropower system↗

Device-Centric Ransomware Detection using Machine Learning-Based Memory Forensics for Smart Inverters

Ransomware attacks are the fastest-growing form of cyberattacks worldwide. Recently, ransomware attacks have targeted industrial control systems (ICSs), including power grids. Lessons learned from recent incidents in ICSs show that ransomware groups can deliver ransomware into not only the organization’s control servers, but also the operational technology (OT) devices such as smart inverters and smart grid devices. This paper proposes a machine learning (ML)- based memory forensics method enabling the detection of ransomware binaries stored in the memory of a commercial smart inverter. Device firmware binary files are extracted from a Serial Peripheral Interface (SPI) flash memory, and samples of both benign and ransomware binaries are generated by a binary manipulation method and a real-world ransomware encryption, separately. A deep transfer learning (DTL) method is used to retrain a convolutional neural network (CNN)-based ransomware detection algorithm using the generated samples. The experimental result validates that the proposed ML-based memory forensics method can accurately detect ransomware files.

97 MATHEMATICS AND COMPUTING↗