Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning, Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Comparison of machine learning and electrical resistivity arrays to inverse modeling for locating and characterizing subsurface targets

Here, this study evaluates the performance of multiple machine learning (ML) algorithms and electrical resistivity (ER) arrays for inversion with comparison to a conventional Gauss-Newton numerical inversion method. Four different ML models and four arrays were used for the estimation of only six variables for locating and characterizing hypothetical subsurface targets. The combination of dipole-dipole with Multilayer Perceptron Neural Network (MLP-NN) had the highest accuracy. Evaluation showed that both MLP-NN and Gauss-Newton methods performed well for estimating the matrix resistivity while target resistivity accuracy was lower, and MLP-NN produced sharper contrast at target boundaries for the field and hypothetical data. Both methods exhibited comparable target characterization performance, whereas MLP-NN had increased accuracy compared to Gauss-Newton in prediction of target width and height, which was attributed to numerical smoothing present in the Gauss-Newton approach. MLP-NN was also applied to a field dataset acquired at U.S. DOE Hanford site.

54 ENVIRONMENTAL SCIENCES↗

A Machine Learning-Based Method to Estimate Transformer Primary-Side Voltages with Limited Customer-Side AMI Measurements

Distribution control applications such as volt/var optimization, network reconfiguration, and distribution automation require accurate knowledge of the distribution system state. The lack of sufficient sensors on the primary side of distribution networks often limits the accuracy of the control decisions by these applications. The deployment of advanced metering infrastructure (AMI) provides utilities an opportunity to translate the AMI data on the secondary onto the primary so that it can be used as pseudo-measurements to augment the limited existing measurements on the primary. This paper develops a machine learning based approach for estimating service transformer primary-side voltages by using limited secondary-side AMI measurement. The machine learning model is developed by using random forest algorithm. The estimated primary-side voltages can be used by utilities as pseudo-measurements for distribution control applications. The detailed secondary model topology, which is an essential input data for many existing algorithms, is not required for the proposed method. The performance of the proposed method is validated by using AMI measurements from the field and an actual distribution feeder model of San Diego Gas & Electric Company.

advanced metering infrastructure↗

Big Data Analytics for Long-Term Meteorological Observations at Hanford Site

A growing number of physical objects with embedded sensors with typically high volume and frequently updated data sets has accentuated the need to develop methodologies to extract useful information from big data for supporting decision making. This study applies a suite of data analytics and core principles of data science to characterize near real-time meteorological data with a focus on extreme weather events. To highlight the applicability of this work and make it more accessible from a risk management perspective, a foundation for a software platform with an intuitive Graphical User Interface (GUI) was developed to access and analyze data from a decommissioned nuclear production complex operated by the U.S. Department of Energy (DOE, Richland, USA). Exploratory data analysis (EDA), involving classical non-parametric statistics, and machine learning (ML) techniques, were used to develop statistical summaries and learn characteristic features of key weather patterns and signatures. The new approach and GUI provide key insights into using big data and ML to assist site operation related to safety management strategies for extreme weather events. Specifically, this work offers a practical guide to analyzing long-term meteorological data and highlights the integration of ML and classical statistics to applied risk and decision science.

54 ENVIRONMENTAL SCIENCES↗

Machine learning results and code

Results for all the Four MLP-NN, RF, XGBR, and RF algorithms for each of geophysical array`s and machine learning python code is provided.

58 GEOSCIENCES↗

MLP-NN vs Gauss-Newton files

-Input files for randomly selected subset data (30 instances) from primary dipole-dipole forward modeling data.-Inversion results in surfer grid format as well as in .DAT format.-Scatterplots for MLP-NN vs Gauss-Newton present in the excel sheet.

58 GEOSCIENCES↗

Decayheatml

This code is designed to predict and analyze the decay heat generated in molten salt reactors (MSRs) using a hybrid approach that combines machine learning and segmented polynomial fitting. The accurate prediction of decay heat is essential for reactor safety and the optimization of spent fuel storage. The code operates through several key components: 1) Data Architecture: It incorporates a modular data architecture that handles various MSR-specific operational parameters such as power density, humidity content, and air ingress. These parameters are sampled using Sobol sequences to ensure comprehensive coverage of operational uncertainties. 2) Machine Learning Framework: The code employs a diverse set of machine learning models, including polynomial regression, decision trees, random forests, gradient boosting, support vector regression, k-nearest neighbors, multi-layer perceptrons, and symbolic regression. These models are trained to predict decay heat over a wide temporal range, from immediate shutdown up to 10,000 years. 3) Region-Optimized Training: The temporal domain is divided into multiple regions, each modeled separately to capture distinct decay heat characteristics across different time scales. This approach significantly improves the accuracy and interpretability of predictions. 4) Segmented Polynomial Interpretation (SPI): The SPI method translates machine learning predictions into piecewise polynomial equations. These equations are physically interpretable and can be directly integrated into existing engineering workflows and safety analyses. 5) Front-End Interfaces: The code includes both a Jupyter notebook interface for research development and a Streamlit web application for operational deployment. These interfaces allow users to interactively explore decay heat predictions, adjust operational parameters, and visualize results in real-time. 6) Applications: The framework supports various applications, including safety system validation and spent fuel container optimization. It enables real-time evaluation of worst-case decay heat scenarios, informing the design of passive safety systems and optimizing container designs for long-term storage. Overall, this code provides a robust, accurate, and user-friendly tool for predicting decay heat in MSRs, enhancing reactor safety, and optimizing spent fuel management.

Retamales, Mauricio Eduardo Tano [Idaho National L↗

Artificial Diversity and Defense Security (ADDSec)

Artificial Diversity and Defense Security (ADDSec) machine learning algorithms are used to classify and cluster threats so that an appropriate response can be initiated as a mitigation strategy. The package includes an ensemble of machine learning algorithms such as Support Vector Machines, naïve bayes, logistic regression, and random forest that evolve with the data to recognize anomalous behavior at the host and network levels. Inputs into the machine learning algorithms include end host system calls, system utilization, packet captures, and syslog messages. The machine learning algorithms can be retrained based on user defined intervals or on the number of packets received. ADDSEC's threat responses include Internet Protocol (IP) Address randomization, application port number randomization, and application library randomization. The IP randomization implementation is built on top of a Software Defined Networking (SDN) framework. The SDN controller installs flows on each of the SDN switches with randomized source and destination IP addresses. The application port numbers are randomized using iptables. The application library randomization is created with a LLVM compiler. All randomization schemes are transparent to the endpoints on the network. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-3379 O

Cox, RebeccaE.↗

Integrating Crack Detection and Pipe Shape Optimization for Enhanced Sewage System Durability

Crack detection in underground reinforced concrete pipes has been essential in determining the state of stormwater infrastructure. Detection models have been implemented for detecting cracks and other defects in pipes using CCTV footage for stormwater drainage systems. In addition, Finite element models have been used to determine optimum shapes and pipe thickness for different boundary conditions such as header pipes in power plants. The concept of shape optimization emerges as a crucial factor in power plant design and operation, with the potential to maximize performance while minimizing the use of materials. Shape optimization not only enhances efficiency but also contributes to reducing the environmental footprint. This paper discusses the integration of both topics by using the cracks detected in underground pipes as boundary conditions for shape optimization of the pipes. A machine learning model has been developed which uses limited data for training and outlines the location of detected cracks. A shape optimization methodology is proposed in which ANSYS modules are used to analyze fluid flow and then optimize the shape of the pipe. The crack detection model developed has been applied to a crack detected in lab setting and machine learning model used has an accuracy of 98% using a random forest algorithm.

20 FOSSIL-FUELED POWER PLANTS↗

Machine learning and deep learning for mineralogy interpretation and CO 2 saturation estimation in geological carbon Storage: A case study in the Illinois Basin

Carbon capture and storage (CCS) is a promising approach to simultaneously maintaining energy security and reducing carbon dioxide (CO 2 ) emissions under the current energy portfolio that is dominated by fossil fuel energy. Pre-injection formation characterization and post-injection CO 2 monitoring are two critical tasks to guarantee storage efficiency in CCS. The CCS projects in the Illinois Basin, the first large-scale CO 2 injection into saline aquifers in the United States, employed conventional and the latest pulsed neutron logging (PNL) tools for mineralogy interpretation and CO 2 saturation estimation, which provide valuable references for future CCS projects. Because of the inherent fuzziness of petrophysical measurements and complex subsurface heterogeneity, interpreting well-logging data is time-consuming, and its accuracy can be user-biased. In recent years, data-driven methods have been widely used to capture the non-linear patterns between input features and interpretation results. This work applied and evaluated four commonly used machine learning (ML) models, including ridge regression (RR), random forest (RF), gradient boosting regression (GBR), support vector regression (SVR), and one deep learning (DL) model, the artificial neural network (ANN). We optimized the hyperparameters of the four ML models and the DL model using the simulated annealing algorithm and the grid search strategy, respectively. The input features of the mineralogy interpretation models were eleven conventional well-logging parameters, and the label data (i.e., ground truth) were the porosity and volumetric fractions of six minerals, including quartz, feldspar, dolomite, calcite, clay, and iron minerals. The results demonstrated that the GBR and RF models were superior in predicting volumetric fractions of minerals and porosity; label data with low coefficient of variation (CV) values tended to yield better performance. For CO 2 saturation estimation, the RF was the best-performing model, followed by SVR, ANN, GBR, and RR. Furthermore, we conducted feature importance ranking using the permutation importance algorithm and found that the formation sigma and well pressure were the most important features in this study. In conclusion, the study of CCS projects in the Illinois Basin bridges the gap between the limited knowledge and understanding of geological carbon storage and the increasing demand for reliable, cost-effective, and sustainable energy solutions.

58 GEOSCIENCES↗

Predictive models of long COVID

Background: The cause and symptoms of long COVID are poorly understood. It is challenging to predict whether a given COVID-19 patient will develop long COVID in the future. Methods: We used electronic health record (EHR) data from the National COVID Cohort Collaborative to predict the incidence of long COVID. We trained two machine learning (ML) models — logistic regression (LR) and random forest (RF). Features used to train predictors included symptoms and drugs ordered during acute infection, measures of COVID-19 treatment, pre-COVID comorbidities, and demographic information. We assigned the ‘long COVID’ label to patients diagnosed with the U09.9 ICD10-CM code. The cohorts included patients with (a) EHRs reported from data partners using U09.9 ICD10-CM code and (b) at least one EHR in each feature category. We analysed three cohorts: all patients (n = 2,190,579; diagnosed with long COVID = 17,036), inpatients (149,319; 3,295), and outpatients (2,041,260; 13,741). Findings: LR and RF models yielded median AUROC of 0.76 and 0.75, respectively. Ablation study revealed that drugs had the highest influence on the prediction task. The SHAP method identified age, gender, cough, fatigue, albuterol, obesity, diabetes, and chronic lung disease as explanatory features. Models trained on data from one N3C partner and tested on data from the other partners had average AUROC of 0.75. Interpretation: ML-based classification using EHR information from the acute infection period is effective in predicting long COVID. SHAP methods identified important features for prediction. Cross-site analysis demonstrated the generalizability of the proposed methodology.

60 APPLIED LIFE SCIENCES↗

A novel probabilistic regression model for electrical peak demand estimate of commercial and manufacturing buildings

Due to the high cost of electricity in commercial and industrial sectors, demand forecast models have gained increasing attention. However, there are two unresolved issues: (1) Models are not adaptable when exposed to previously unknown data (2) The value of regression methods vs. state-of-the-art machine learning models has not been made apparent before. This study’s goal is to develop probabilistic demand estimation models. Herein, we propose a probabilistic Bayesian regression framework that can not only estimate future demands with high accuracy but also be updated once new information is available. By applying the proposed algorithm to two real-world case studies (commercial and manufacturing), we show a 40.3% and 30.8% improvement in terms of mean absolute error for the two cases. Moreover, the proposed technique outperforms powerful machine learning approaches, including support vector machine by 10.39%, random forest by 6.17%, and multilayer perceptron by 9.14% in terms of mean absolute percentage error.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Predicting solid state material platforms for quantum technologies

Semiconductor materials provide a compelling platform for quantum technologies (QT). However, identifying promising material hosts among the plethora of candidates is a major challenge. Therefore, we have developed a framework for the automated discovery of semiconductor platforms for QT using material informatics and machine learning methods. Different approaches were implemented to label data for training the supervised machine learning (ML) algorithms logistic regression, decision trees, random forests and gradient boosting. We find that an empirical approach relying exclusively on findings from the literature yields a clear separation between predicted suitable and unsuitable candidates. In contrast to expectations from the literature focusing on band gap and ionic character as important properties for QT compatibility, the ML methods highlight features related to symmetry and crystal structure, including bond length, orientation and radial distribution, as influential when predicting a material as suitable for QT.

36 MATERIALS SCIENCE↗

Estimating Subhourly Inverter Clipping Loss From Satellite-Derived Irradiance Data

Photovoltaic system production simulations are conventionally run using hourly weather datasets. Hourly simulations are sufficiently accurate to predict the majority of long-term system behavior but cannot resolve high-frequency effects like inverter clipping caused by short-duration irradiance variability. Direct modeling of this subhourly clipping error is only possible for the few locations with high-resolution irradiance datasets. This paper describes a method of predicting the magnitude of this error using a machine learning regressor ensemble model, comprised of a random forest and an XGBoost model, and 30-minute satellite irradiance data. The method predicts a correction for each 30-minute interval with the potential to roll up into 60-minute corrections to match an hourly energy model. The model is trained and validated at locations where the error can be directly simulated from 1-minute ground data. The validation shows low bias at most ground station locations. The model is also applied to gridded satellite irradiance to produce a heatmap of the estimated clipping error across the United States. Finally, the relative importance of each predictor satellite variable is retrieved from the model and discussed.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Technical note: Uncertainties in eddy covariance CO 2 fluxes in a semiarid sagebrush ecosystem caused by gap-filling approaches

Abstract. Gap-filling eddy covariance CO2 fluxes is challenging at dryland sites due to small CO2 fluxes. Here, four machine learning (ML) algorithms including artificial neural network (ANN), k-nearest neighbors (KNNs), random forest (RF), and support vector machine (SVM) are employed and evaluated for gap-filling CO2 fluxes over a semiarid sagebrush ecosystem with different lengths of artificial gaps. The ANN and RF algorithms outperform the KNN and SVM in filling gaps ranging from hours to days, with the RF being more time efficient than the ANN. Performances of the ANN and RF are largely degraded for extremely long gaps of 2 months. In addition, our results suggest that there is no need to fill the daytime and nighttime net ecosystem exchange (NEE) gaps separately when using the ANN and RF. With the ANN and RF, the gap-filling-induced uncertainties in the annual NEE at this site are estimated to be within 16 g C m−2, whereas the uncertainties by the KNN and SVM can be as large as 27 g C m−2. To better fill extremely long gaps of a few months, we test a two-layer gap-filling framework based on the RF. With this framework, the model performance is improved significantly, especially for the nighttime data. Therefore, this approach provides an alternative in filling extremely long gaps to characterize annual carbon budgets and interannual variability in dryland ecosystems.

Yao, Jingyu↗