Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “learning algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Improving the accessibility and transferability of machine learning algorithms for identification of animals in camera trap images: MLWIC2

Motion-activated wildlife cameras (or “camera traps”) are frequently used to remotely and noninvasively observe animals. The vast number of images collected from camera trap projects has prompted some biologists to employ machine learning algorithms to automatically recognize species in these images, or at least filter-out images that do not contain animals. These approaches are often limited by model transferability, as a model trained to recognize species from one location might not work as well for the same species in different locations. Furthermore, these methods often require advanced computational skills, making them inaccessible to many biologists. We used 3 million camera trap images from 18 studies in 10 states across the United States of America to train two deep neural networks, one that recognizes 58 species, the “species model,” and one that determines if an image is empty or if it contains an animal, the “empty-animal model.” Our species model and empty-animal model had accuracies of 96.8% and 97.3%, respectively. Furthermore, the models performed well on some out-of-sample datasets, as the species model had 91% accuracy on species from Canada (accuracy range 36%–91% across all out-of-sample datasets) and the empty-animal model achieved an accuracy of 91%–94% on out-of-sample datasets from different continents. Our software addresses some of the limitations of using machine learning to classify images from camera traps. By including many species from several locations, our species model is potentially applicable to many camera trap studies in North America. We also found that our empty-animal model can facilitate removal of images without animals globally. We provide the trained models in an R package (MLWIC2: Machine Learning for Wildlife Image Classification in R), which contains Shiny Applications that allow scientists with minimal programming experience to use trained models and train new models in six neural network architectures with varying depths.

59 BASIC BIOLOGICAL SCIENCES↗

Prediction of electric and magnetic fields from spectral data using machine learning algorithms for Doppler-free saturation spectroscopy diagnostics

The prediction of electric and magnetic field amplitudes from atomic spectral data is critical for plasma control in fusion devices such as tokamaks. Conventional approaches that rely on physics-based models are computationally expensive and unsuitable for real-time applications. In this work, we develop and benchmark three machine learning algorithms—simulation-based inference (SBI), fully connected neural networks (FCNN), and histogram-based gradient boosting regression (GBR-Hist)—to infer field intensities directly from Doppler-free saturation spectroscopy (DFSS) spectra. Synthetic datasets of spectra were generated using the EZSSS code and evaluated both with and without added Poisson noise to mimic experimental conditions. We find that SBI achieves the highest accuracy and robustness, FCNN provides a strong balance of accuracy and computational efficiency for real-time applications, and GBR-Hist offers the fastest inference but is more sensitive to noise. Furthermore, these results demonstrate the potential of machine learning to accelerate DFSS analysis and enhance its utility for plasma diagnostics and control.

Doppler-free saturation spectroscopy↗

Construction of Women’s All-Around Speed Skating Event Performance Prediction Model and Competition Strategy Analysis Based on Machine Learning Algorithms

Introduction Accurately predicting the competitive performance of elite athletes is an essential prerequisite for formulating competitive strategies. Women’s all-around speed skating event consists of four individual subevents, and the competition system is complex and challenging to make accurate predictions on their performance. Objective The present study aims to explore the feasibility and effectiveness of machine learning algorithms for predicting the performance of women’s all-around speed skating event and provide effective training and competition strategies. Methods The data, consisting of 16 seasons of world-class women’s all-around speed skating competition results, used in the present study came from the International Skating Union (ISU). According to the competition rules, distinct features are filtered using lasso regression, and a 5,000 m race model and a medal model are built using a fivefold cross-validation method. Results The results showed that the support vector machine model was the most stable among the 5,000 m race and the medal models, with the highest AUC (0.86, 0.81, respectively). Furthermore, 3,000 m points are the main characteristic factors that decide whether an athlete can qualify for the final. The 11th lap of the 5,000 m, the second lap of the 500 m, and the fourth lap of the 1,500 m are the main characteristic factors that affect the athlete’s ability to win medals. Conclusion Compared with logistic regression, random forest, K-nearest neighbor, naive Bayes, neural network, support vector machine is a more viable algorithm to establish the performance prediction model of women’s all-around speed skating event; excellent performance in the 3,000 m event can facilitate athletes to advance to the final, and athletes with outstanding performance in the 500 m event are more likely competitive for medals.

Liu, Meng↗

Landslide Likelihood Prediction using Machine Learning Algorithms

The supply of electricity via power plants is criticalto the operation of many critical infrastructure systems in mod-ern society. Natural hazards can disrupt the power supply, causepower outages that can halt economic growth, and impede emer-gency response until power is restored. The proposed work aimsto predict the landslides likelihood in these critical infrastructurelocations in the Northeastern USA using integrated databases ofexplanatory variables and machine learning algorithms. First,data related to landslides are obtained and merged, includingtopographic, soil moisture, and precipitation-related data. Fiveregression algorithms, namely: Random Forest, Extreme Gradi-ent Boosting (XGBoost), K-Nearest Neighbor regression (KNN),Linear Support Vector Regressor (SVR), and Linear regression,are utilized to predict the landslide probability and evaluatedon the dataset. The accuracy of the models is assessed by usingstatistical metrics such as mean absolute error (MAE), meansquared error (MSE), and root mean squared error (RMSE).The study results show that Random Forest outperformed othermodels with the mutual information feature selection method.It achieved an MSE of 0.0011 with mutual information-basedfeature selection and an MSE of 0.00157 without feature selection.KNN regressor outperformed the other models with an MSEof 0.00139 with correlation-based information selection. Theproposed landslide identification model with Random Forestalgorithm shows outstanding robustness and great potential intackling the landslide likelihood prediction by employing MLalgorithms.

Vasundhara Acharya↗

Application of a Machine Learning Algorithm in Generating an Evapotranspiration Data Product From Coupled Thermal Infrared and Microwave Satellite Observations

Land surface evapotranspiration (ET) is one of the main energy sources for atmospheric dynamics and a critical component of the local, regional, and global water cycles. Consequently, accurate measurement or estimation of ET is one of the most active topics in hydro-climatology research. With massive and spatially distributed observational data sets of land surface properties and environmental conditions being collected from the ground, airborne or space-borne platforms daily over the past few decades, many research teams have started to use big data science to advance the ET estimation methods. The Geostationary satellite Evapotranspiration and Drought (GET-D) product system was developed at the National Oceanic and Atmospheric Administration (NOAA) in 2016 to generate daily ET and drought maps operationally. The primary inputs of the current GET-D system are the thermal infrared (TIR) observations from NOAA GOES satellite series. Because of the cloud contamination to the TIR observations, the spatial coverage of the daily GET-D ET product has been severely impacted. Based on the most recent advances, we have tested a machine learning algorithm to estimate all-weather land surface temperature (LST) from TIR and microwave (MW) combined satellite observations. With the regression tree machine learning approach, we can combine the high accuracy and high spatial resolution of GOES TIR data with the better spatial coverage of passive microwave observations and LST simulations from a land surface model (LSM). The regression tree model combines the three LST data sources for both clear and cloudy days, which enables the GET-D system to derive an all-weather ET product. This paper reports how the all-weather LST and ET are generated in the upgraded GET-D system and provides an evaluation of these LST and ET estimates with ground measurements. The results demonstrate that the regression tree machine learning method is feasible and effective for generating daily ET under all weather conditions with satisfactory accuracy from the big volume of satellite observations.

54 ENVIRONMENTAL SCIENCES↗

Experiments on Supervised Learning Algorithms for Text Categorization

Modern information society is facing the challenge of handling massive volume of online documents, news, intelligence reports, and so on. How to use the information accurately and in a timely manner becomes a major concern in many areas. While the general information may also include images and voice, we focus on the categorization of text data in this paper. We provide a brief overview of the information processing flow for text categorization, and discuss two supervised learning algorithms, viz., support vector machines (SVM) and partial least squares (PLS), which have been successfully applied in other domains, e.g., fault diagnosis [9]. While SVM has been well explored for binary classification and was reported as an efficient algorithm for text categorization, PLS has not yet been applied to text categorization. Our experiments are conducted on three data sets: Reuter's- 21578 dataset about corporate mergers and data acquisitions (ACQ), WebKB and the 20-Newsgroups. Results show that the performance of PLS is comparable to SVM in text categorization. A major drawback of SVM for multi-class categorization is that it requires a voting scheme based on the results of pair-wise classification. PLS does not have this drawback and could be a better candidate for multi-class text categorization.

Namburu, Setu Madhavi↗

Sensitivity Analysis of Fluid–Fluid Interfacial Area, Phase Saturation and Phase Connectivity on Relative Permeability Estimation Using Machine Learning Algorithms

Recent studies have shown that relative permeability can be modeled as a state function which is independent of flow direction and dependent upon phase saturation (S), phase connectivity (X), and fluid–fluid interfacial area (A). This study evaluates the impact of each of the three state parameters (S, X, and A) in the estimation of relative permeability. The relative importance of the three state parameters in four separate quadrants of S-X-A space was evaluated using a machine learning algorithm (out-of-bag predictor importance method). The results show that relative permeability is sensitive to all the three parameters, S, X, and A, with varying magnitudes in each of the four quadrants at a constant value of wettability. We observe that the wetting-phase relative permeability is most sensitive to saturation, while the non-wetting phase is most sensitive to phase connectivity. Although the least important, fluid–fluid interfacial area is still important to make the relative permeability a more exact state function.

47 OTHER INSTRUMENTATION↗

Evaluating Combinations of Sentinel-2 Data and Machine-Learning Algorithms for Mangrove Mapping in West Africa

Creating a national baseline for natural resources, such as mangrove forests, and monitoring them regularly often requires a consistent and robust methodology. With freely available satellite data archives and cloud computing resources, it is now more accessible to conduct such large-scale monitoring and assessment. Yet, few studies examine the reproducibility of such mangrove monitoring frameworks, especially in terms of generating consistent spatial extent. Our objective was to evaluate a combination of image processing approaches to classify mangrove forests along the coast of Senegal and The Gambia. We used freely available global satellite data (Sentinel-2), and cloud computing platform (Google Earth Engine) to run two machine learning algorithms, random forest (RF), and classification and regression trees (CART). We calibrated and validated the algorithms using 800 reference points collected using high-resolution images. We further re-ran 10 iterations for each algorithm, utilizing unique subsets of the initial training data. While all iterations resulted in thematic mangrove maps with over 90% accuracy, the mangrove extent ranges between 827-2807 km2 for Senegal and 245-1271 km2 for The Gambia with one outlier for each country. We further report "Places of Agreement" (PoA) to identify areas where all iterations for both methods agree (506.6 km2 and 129.6 km2 for Senegal and The Gambia, respectively), thus have a high confidence in predicting mangrove extent. While we acknowledge the time- and cost-effectiveness of such methods for the landscape managers, we recommend utilizing them with utmost caution, as well as post-classification on-the-ground checks, especially for decision making.

Mondal, Pinki↗

Application of a machine learning algorithm (XGBoost) to offline RHIC luminosity optimization

The operation parameter optimization in 2020 RHIC low energy run is difficult. First, the RHIC luminosity is affected by many RHIC operation parameters, as well as affected by many Low Energy RHIC electron Cooling (LEReC) operation parameters. Second, the luminosity signal in this run is noisy and not sensitive to these parameter changes, especially when these parameters are very close to their optimized values. It is not easy to distinguish the effects of one parameter from all other operation parameters separately. Therefore, it is difficult to optimize the luminosity by varying these parameters one by one. To find a way for luminosity optimization, we analyze some operation parameters via a machine learning algorithm - XGBoost. After constructing a black-box surrogate model from XGBoost and plotting their partial dependency plots (PDF) and SHAP value plots for different operation parameters, we can find the effects of an individual parameter on the RHIC luminosity and optimize it accordingly.

43 PARTICLE ACCELERATORS↗

GW-PINN: A deep learning algorithm for solving groundwater flow equations

Machine learning methods provide new perspective for more convenient and efficient prediction of groundwater flow. In this study, a deep learning method “GW-PINN” without labeled data for solving groundwater flow equations with wells was proposed. GW-PINN takes the physics inform neural network (PINN) as the backbone and uses either the hard or soft constraint in the loss function for training. A locally refined sampling strategy (LRS) is adopted to generate the consistent spatial sampling points for problems with strong hydraulic head change, and then combined with an appropriate temporal sampling scheme to obtain the final spatial-temporal sampling points. A snowball-style two-stage training strategy by dividing the temporal domain into two subdomains is designed to decrease the sampling points. Five cases were designed to test the training performance of GW-PINN under different sampling strategies and two constraints. The predicted results of GW-PINN were compared with MODFLOW and the analytical solution. The results demonstrate that GW-PINN possesses strong ability in capturing the hydraulic head change for both confined and un-confined aquifers. The hard constraint owns more robust learning ability than the soft constraint. The LRS strategy can generate more accurate results with much fewer sampling points than traditional sampling strategies, and the snowball-style two-stage training strategy is significantly efficient for problems with the drastic change of hydraulic head. Additionally, the application of GW-PINN as a surrogate model for parameterized groundwater flow equations is illustrated. This study provides an option tool for efficient groundwater flow simulation, especially for those with local refinements are needed.

54 ENVIRONMENTAL SCIENCES↗

Robustness of deep learning algorithms in astronomy -- galaxy morphology studies

Deep learning models are being increasingly adopted in wide array of scientific domains, especially to handle high-dimensionality and volume of the scientific data. However, these models tend to be brittle due to their complexity and overparametrization, especially to the inadvertent adversarial perturbations that can appear due to common image processing such as compression or blurring that are often seen with real scientific data. It is crucial to understand this brittleness and develop models robust to these adversarial perturbations. To this end, we study the effect of observational noise from the exposure time, as well as the worst case scenario of a one-pixel attack as a proxy for compression or telescope errors on performance of ResNet18 trained to distinguish between galaxies of different morphologies in LSST mock data. We also explore how domain adaptation techniques can help improve model robustness in case of this type of naturally occurring attacks and help scientists build more trustworthy and stable models.

79 ASTRONOMY AND ASTROPHYSICS↗

A Comparative Study of Machine Learning Algorithms for Industry-Specific Freight Generation Model

According to Bureau of Transportation Statistics, the U.S. transportation system handled 14,329 million ton-miles of freight per day in 2020. Understanding the generation of these freight shipments is crucial for transportation researchers, planners, and policymakers to design and plan for a more efficient and connected freight transportation system. Traditionally, the freight generation modeling has been based on Ordinary Least Square (OLS) regression, although more advanced Machine Learning (ML) algorithms have been evaluated and proven to have excellent performance in various transportation applications in recent years. Furthermore, one modeling approach applied for one industry might not always be applicable for another as their freight generation logics can be quite different. The objective of this study is to apply and evaluate alternative ML algorithms in the estimation of freight generation for each of 45 industry types. Seven alternative ML algorithms, along with the base OLS regression, were evaluated and compared. In addition, the study considered different combinations of variables in both the original and logarithmic form as well as hyperparameters of those ML algorithms in the model selection for each industry type. The results showed statistically significant improvements in the root mean square error reduction by the alternative ML algorithms over the OLS for over 80% of cases. The study suggests utilizing the alternative ML algorithms can reduce the root mean square error by about 30%, depending on industry types.

97 MATHEMATICS AND COMPUTING↗

Improved machine learning algorithm for predicting ground state properties

Finding the ground state of a quantum many-body system is a fundamental problem in quantum physics. In this work, we give a classical machine learning (ML) algorithm for predicting ground state properties with an inductive bias encoding geometric locality. The proposed ML model can efficiently predict ground state properties of an n-qubit gapped local Hamiltonian after learning from only $\mathcal{O}$(log(n)) data about other Hamiltonians in the same quantum phase of matter. This improves substantially upon previous results that require $\mathcal{O}$(n c ) data for a large constant c. Furthermore, the training and prediction time of the proposed ML model scale as $\mathcal{O}$(n log n) in the number of qubits n. Numerical experiments on physical systems with up to 45 qubits confirm the favorable scaling in predicting ground state properties using a small training dataset.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Comparison of Supervised and Un-Supervised Machine Learning Algorithms for Threat Detection and Scintillator Performance for Radiation Portal Monitoring

Following the events of September 11, 2001, international border crossing have been equipped with radiation portal monitors (RPMs) to identify illicit radioactive material. Polyvinyl toluene (PVT) scintillators are commonly used due to their low cost and reasonable maintainability, however they offer low spectral resolution. Despite the fact that over twenty years has transpired since this event, radioisotopes are still typically identified by hand-crafted classification algorithms, e.g., total counts or energy windowing, and exhibit relatively poor performance in detecting threats at the low false alarm rates required to support the stream of commerce. While some improvement to performance has been realized via the use of supervised machine learning, these classification algorithms typically utilize simulations in lieu of real data due to the sparsity of data for one or more classes. Accordingly, the performance of these algorithms is somewhat less than optimal when examining experiments or simulations with model mismatch. Consequently, in this work, we examine the application of a number of unsupervised machine learning, anomaly detection based algorithms, to circumvent the inverse crime when analyzing spectroscopy data for RPMs. We also compare anomaly detection results with those obtained via the use of supervised classification detection ML algorithms when model mismatch is introduced between the simulated threat items utilized for training/testing. Finally, we compared the performance of the PVT scintillators to those obtained with higher resolution detectors using both anomaly detection and supervised classification algorithms.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

A Predictive Prescription Framework for Stochastic Unit Commitment Using Boosting Ensemble Learning Algorithms

To take unit commitment (UC) decisions under uncertain load, most existing stochastic optimization (SO) frameworks adopt a generic representation of uncertainty. While load levels that materialize on a particular day are influenced by various covariates (such as the day of the week or temperature), SO frameworks typically disregard such side observations, wasting actionable information that could significantly enhance decision quality. Here, this article proposes a contextual SO (CSO) framework for UC under uncertain load, which can effectively exploit covariate observations in conjunction with a class of machine learning (ML) algorithms to improve the out-of-sample performance of UC decisions. It shows how three ML algorithms, adaptive boosting, gradient boosted trees, and extreme gradient boosting, can be used to this end, constituting the first application of these algorithms in any CSO framework. Using real-world data harvested from the New York ISO grid, we measure the out-of-sample performance of the framework in terms of total operation cost, shed load values, locational marginal prices, and total payments by the loads, against several benchmark methods proposed in the literature. The article has an online companion (Yurdakul et al.), wherein we present additional results and lay out further mathematical formulations used in this work.

42 ENGINEERING↗

HIGH-LOW FIDELITY THERMAL HYDRAULIC COUPLING USING AI/MACHINE LEARNING ALGORITHMS

The primary goal of the US Department of Energy (DOE) office of Nuclear Energy Integrated Energy Systems (IES) program is to develop the tools and framework for coupling multi-scale and multi-physical thermal and electrical energy usage and storage systems. High- and low-fidelity (high–low) coupling is a key feature of multi-scale, multi-component systems and has been an important focus of research in the nuclear energy community for the past two decades. An essential feature of demonstrating the capability to couple high-fidelity and low-fidelity systems for real-time applications are surrogate/reduced order models (ROM). For the purposes of this study, surrogate models are essentially Blackbox models, typically developed using supervised Machine learning (ML) algorithms. The surrogate models can be used to mimic the response of high-fidelity models to represent large historical datasets and coupled with more general low-fidelity system models distributed as Functional Mock-up Interface (FMI) or Functional Mock-up Units (FMU) modules. The example is demonstrated with Spallation Neutron Source (SNS) First Target Station flow loop data. The flow loop is a liquid mercury loop with a pump, piping, heat exchange, and internal heat generation in the target window. This work elucidates some of the potential benefits and future needs of developing tools for high–low system coupling of energy systems.

Williams, Wesley↗