Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Predictive links between microbial communities and biological oxygen utilization in the Arctic Ocean

Microbial metabolism influences rates of net community production (NCP), exerting a direct biological control on marine oxygen and carbon fluxes. In the Arctic, it is increasingly important to understand and quantify this process, as ecological and oceanographic conditions shift due to changing climate. Here, we describe potential ecological links between pelagic microbial diversity and an NCP precursor, biological oxygen utilization, using machine learning and paired observations of community structure and metabolic activity from a seasonally and spatially variable transect of the Arctic Ocean (2019–2020 MOSAiC Expedition). Community structure was determined using 16S (prokaryotic) and 18S (eukaryotic) rRNA gene amplicon sequencing, and metabolic activity was derived from ΔO 2 /Ar. Using self-organizing maps, we identified clear successional patterns in observed microbial community structure that were seasonally driven in the upper ocean and vertically stratified with depth. Metabolic activity was also stratified, with a primarily net heterotrophic water column (median −1.5% biological oxygen saturation), excepting periodic oxygen supersaturation (maximum: 13.6%) within the mixed layer. Using DNA sequences as predictor variables, we then constructed a random forest regression model that reliably reconstructed biological oxygen concentrations (root mean squared error = 4.14 μmol kg −1 ). Top predictors from this model were from heterotrophic (bacteria) or potentially mixotrophic (dinoflagellate) taxa. These analyses highlight biologically driven diagnostic tools that can be used to expand biogeochemical datasets and improve the microbial perspectives and metabolisms represented in ecological models of net productivity and carbon flux in a changing Arctic Ocean.

Chamberlain, Emelia J. [Univ. of San Diego, San Di↗

Detection and classification capabilities of two multibeam sonars

Multibeam sonars can be used to detect and classify marine fauna under conditions when optical sensors are ineffective. In this work, we compare the detection and classification capabilities of two multibeam sonars with different operating frequencies, the Tritech Gemini 720is (720 kHz) and the Teledyne BlueView M900-2250 (2250 kHz). The two sonars were deployed with overlapping swaths in a narrow tidal channel with peak currents of approximately 2 m/s where seals, schools of fish, and diving birds were intermittently present. In comparing data concurrently acquired by both sonars, we observe differences in the appearance of detected targets and detection capabilities. The detected target tracks are classified using a random forest model as either individual biological targets (flora and fauna, including fish, diving birds, seals, and plant matter), fish schools, or nonbiological targets (including entrained air and sonar artifacts). Despite the observed differences in detection capabilities for the two sonars, automatic classification distinguishes between target classes with precision and recall above 0.7 and 0.9, respectively. Furthermore, these results suggest that while similar methodologies can be used for data analysis, some outcomes from environmental studies using multibeam sonars may be instrument-specific.

16 TIDAL AND WAVE POWER↗

Formation of Amorphous Carbon Multi‐Walled Nanotubes from Random Initial Configurations

Amorphous carbon nanotubes (a‐CNT) with up to four walls and sizes ranging from 200 to 3200 atoms have been simulated, starting from initial random configurations and using the Gaussian Approximation Potential. The important variables (like density, height, and diameter) required to successfully simulate a‐CNTs were predicted with the machine learning random forest technique. The width of the a‐CNT models ranged between 0.55–2 nm with an average inter‐wall spacing of 0.31 nm. The topological defects in a‐CNTs were analyzed and new defect configurations were observed. The electronic density of states and localization in these phases were discussed and delocalized electrons in the π subspace were identified as an important factor for inter‐layer cohesion. Spatial projection of the electronic conductivity favors axial transport along connecting hexagons, while non‐hexagonal parts of the network either hinder or bifurcate the electronic transport. A vibrational density of states was calculated and is potentially an experimentally comparable fingerprint of the material. The appearance of a low‐frequency radial breathing mode was discussed and the thermal conductivity at 300 K was estimated using the Green‐Kubo formula.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Monitoring covariance in multivariate time series: Comparing machine learning and statistical approaches

Abstract In complex systems with multiple variables monitored at high‐frequency, variables are not only temporally autocorrelated, but they may also be nonlinearly related or exhibit nonstationarity as the inputs or operation changes. One approach to handling such variables is to detrend them prior to monitoring and then apply control charts that assume independence and stationarity to the residuals. Monitoring controlled systems is even more challenging because the control strategy seeks to maintain variables at prespecified mean levels, and to compensate, correlations among variables may change, making monitoring the covariance essential. In this paper, a vector autoregressive model (VAR) is compared with a multivariate random forest (MRF) and a neural network (NN) for detrending multivariate time series prior to monitoring the covariance of the residuals using a multivariate exponentially weighted moving average (MEWMA) control chart. Machine learning models have an advantage when the data's structure is unknown or may change. We design a novel simulation study with nonlinear, nonstationary, and autocorrelated data to compare the different detrending models and subsequent covariance monitoring. The machine learning models have superior performance for nonlinear and strongly autocorrelated data and similar performance for linear data. An illustration with data from a reverse osmosis process is given.

Weix, Derek↗

Network Modeling of Complex Data Sets

We demonstrate a selection of network and machine learning techniques useful in the analysis of complex datasets, including 2-way similarity networks, Markov clustering, enrichment statistical networks, FCROS differential analysis, and random forests. We demonstrate each of these techniques on the Populus trichocarpa gene expression atlas.

Jones, Piet C.↗

Use of Longitudinal Serum Analysis and Machine Learning to Develop a Classifier for Cancer Early Detection

Early detection of solid tumors through a simple screening process, such as the proteomic analysis of biofluids, has the potential to significantly alter the management and outcomes of cancers. The application of advanced targeted proteomics measurements and data analysis strategies to uniformly collected serum or plasma samples would enable longitudinal studies of cancer risk, progression, and response to therapy that have the potential to significantly reduce cancer burden in general. In this article, we describe a generalizable workflow combining robust, multiplexed targeted proteomics measurements applied to longitudinal samples from the Department of Defense Serum Repository with a Random Forest machine learning method for developing and initially evaluating the performance of candidate biomarker panels for early detection of cancers. The effectiveness of this approach was demonstrated in a cohort of 175 head and neck squamous cell carcinoma patients. The outlined protocols include methods for sample preparation, instrument analysis, and data analysis and interpretation using this workflow.

Longitudinal analysis, machine learning, cancer, e↗

Adrastea: An Efficient FPGA Design Environment for Heterogeneous Scientific Computing and Machine Learning

We present Adrastea, an efficient FPGA design environment for developing scientific machine learning applications. FPGA development is challenging, from deployment, proper toolchain setup, programming methods, interfacing FPGA kernels, and more importantly, the need to explore design space choices to get the best performance and area usage from the FPGA kernel design. Adrastea provides an automated and scalable design flow to parameterize, implement, and optimize complex FPGA kernels and associated interfaces. We show how virtualization of the development environment via virtual machines is leveraged to simplify the setup of the FPGA toolchain while deploying the FPGA boards and while scaling up the automated design space exploration to leverage multiple machines concurrently. Adrastea provides an automated build and test environment of FPGA kernels. By exposing design space hyper-parameters, Adrastea can automatically search the design space in parallel to optimize the FPGA design for a given metric, usually performance or area. Adrastea simplifies the task of interfacing with the FPGA kernels with a simplified interface API. To demonstrate the capabilities of Adrastea, we implement a complex random forest machine learning kernel with 10,000 input features while achieving extremely low computing latency without loss of prediction accuracy, which is required by a scientific edge application at SNS. We also demonstrate Adrastea using an FFT kernel and show that for both applications Adrastea is able to systematically and efficiently evaluate different design options, which reduced the time and effort required to develop the kernel from months of manual work to days of automatic builds.

Young, Aaron↗

A data-driven framework for predicting machining stability: employing simulated data, operational modal analysis, and enhanced transfer learning

Chatter, a self-excited vibration phenomenon, presents a significant challenge in machining operations, particularly in high-speed milling, where it can degrade tool life, reduce material removal efficiency, and compromise workpiece quality. Addressing this challenge requires a reliable predictive model that can accommodate the complex dynamics of various machining scenarios. This study introduces a novel, data-driven approach to predicting machining stability, leveraging over 140,000 simulated datasets and employing advanced techniques such as operational modal analysis (OMA), enhanced transfer learning (TL), and receptance coupling substructure analysis (RCSA). By integrating these methodologies, the framework effectively classifies and predicts chatter across diverse operational modes, achieving robust and accurate outcomes. Our model utilizes a Random Forest (RF) classifier trained with the comprehensive dataset, which demonstrates substantial improvements in both predictive accuracy and robustness. Specifically, the RF model achieved an accuracy rate of 85%, an area under the curve (AUC) of 0.90, and an F1 score of 0.88, underscoring its capability to adapt to varying machining configurations. These results highlight the framework’s potential to enhance operational efficiency and machining quality by providing reliable chatter predictions across a broad range of machining parameters. In conclusion, this research thus offers a significant advancement in predictive maintenance for machining processes, enabling more stable and efficient manufacturing operations.

42 ENGINEERING↗

Machine Learning Analysis of Impact of Western US Fires on Central US Hailstorms

Fires, including wildfires, harm air quality and essential public services like transportation, communication, and utilities. These fires can also influence atmospheric conditions, including temperature and aerosols, potentially affecting severe convective storms. Here, we investigate the remote impacts of fires in the western United States (WUS) on the occurrence of large hail (size: $\geqslant$ 2.54 cm) in the central US (CUS) over the 20-year period of 2001–20 using the machine learning (ML), Random Forest (RF), and Extreme Gradient Boosting (XGB) methods. The developed RF and XGB models demonstrate high accuracy (> 90%) and F1 scores of up to 0.78 in predicting large hail occurrences when WUS fires and CUS hailstorms coincide, particularly in four states (Wyoming, South Dakota, Nebraska, and Kansas). The key contributing variables identified from both ML models include the meteorological variables in the fire region (temperature and moisture), the westerly wind over the plume transport path, and the fire features (i.e., the maximum fire power and burned area). Importantly, the results confirm a linkage between WUS fires and severe weather in the CUS, corroborating the findings of our previous modeling study conducted on case simulations with a detailed physics model.

54 ENVIRONMENTAL SCIENCES↗

Accelerated screening of functional atomic impurities in halide perovskites using high-throughput computations and machine learning

The pressing need for novel materials that can serve rising demands in solar cell and optoelectronic technologies makes the nexus of halide perovskites, high-throughput computations, and machine learning, very promising. Ever increasing amounts of data on the structure, fundamental properties, and device performance of halide perovskites provide opportunities for learning chemical rules and design principles that make these materials attractive, and applying them across wide chemical spaces. In this work, we show that impurity properties of halide perovskites computed using density functional theory (DFT) can be combined with machine learning (ML) to deliver predictive models and quick identification of optoelectronically active impurity atoms. Our computation lead to the largest reported dataset of the formation energies and charge transition levels of Pb-site impurities in methylammonium lead halide (MAPbX 3 ) perovskites. Descriptors are defined to uniquely represent any impurity atom in any MAPbX 3 compound and mapped to the computed impurity properties using regression techniques such as Gaussian process regression, neural networks, and random forests. We use the best optimized predictive models to make predictions for hundreds of impurities across 9 MAPbX 3 compounds and create lists of dominating impurities, that is, impurities that can shift the equilibrium Fermi level in the perovskite as determined by native point defects. Finally, this accelerated screening powered by computations and machine learning can guide the identification of problematic impurities that may cause undesired recombination of charge carriers, as well as impurities that can be deliberately introduced to tune the perovskite conductivity and resulting photovoltaic absorption.

36 MATERIALS SCIENCE↗

Risk-Aware Framework Development for Disruption Prediction: Alcator C-Mod and DIII-D Survival Analysis

Abstract Survival regression models can achieve longer warning times at similar receiver operating characteristic performance than previously investigated models. Survival regression models are also shown to predict the time until a disruption will occur with lower error than other predictors. Time-to-event predictions from time-series data can be obtained with a survival analysis statistical framework, and there have been many tools developed for this task which we aim to apply to disruption prediction. Using the open-source Auton-Survival package we have implemented disruption predictors with the survival regression models Cox Proportional Hazards, Deep Cox Proportional Hazards, and Deep Survival Machines. To compare with previous work, we also include predictors using a Random Forest binary classifier, and a conditional Kaplan-Meier formalism. We benchmarked the performance of these five predictors using experimental data from the Alcator C-Mod and DIII-D tokamaks by simulating alarms on each individual shot. We find that developing machine-relevant metrics to evaluate models is an important area for future work. While this study finds cases where disruptive conditions are not predicted, there are instances where the desired outcome is produced. Giving the plasma control system the expected time-to-disruption will allow it to determine the optimal actuator response in real time to minimize risk of damage to the device.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Evaluating ecosystem water use efficiency under drought stress: a case study of the Helan Mountain region, northwest China

Context Water use efficiency (WUE) is a fundamental ecological indicator links carbon assimilation and water loss in terrestrial ecosystems. Understanding its responses to drought stress is essential for adaptive ecosystem management, particularly in climate-sensitive mountain landscapes. Objectives This study aimed to investigate drought-driven variations in WUE across major vegetation types in the Helan Mountain region of Northwest China. Specifically, we sought to identify dominant ecological drivers of WUE variability and to disentangle their relative importance and causal pathways. Methods We quantified WUE using the Moderate Resolution Imaging Spectroradiometer (MODIS) products and the Drought Severity Index (DSI) data from 2001 to 2020. To examine WUE – drought relationships across contrasting vegetation types, we employed a spatially explicit analytical framework integrating Random Forest (RF) modeling, partial correlation analysis, and structural equation modeling (SEM). Results Regional WUE exhibited relatively stable interannual dynamics, yet pronounced spatial heterogeneity that was strongly modulated by drought conditions. Vegetation properties, particularly Leaf Area Index (LAI) and Normalized Difference Vegetation Index (NDVI), emerged as the dominant determinants of WUE, with NDVI alone explaining over 20% of its spatial variance in forest and grassland during non-drought periods. SEM analyses revealed that climate forcing influenced WUE mainly through indirect pathways mediated by soil moisture availability and vegetation structural dynamics, rather than through direct climatic controls. Among all regulating factors, LAI acted as the central control node governing ecosystem carbon–water coupling. In contrast, short-term climatic stress, especially atmospheric demand and drought duration, exerted weak or negative direct effects on WUE. Ecosystem-specific responses were observed, with croplands mainly regulated by soil water availability, whereas forests and grasslands showed more sensitive to atmospheric drought stress. Together, these results reveal a hierarchical control framework where soil–vegetation interactions mediate climate impacts on WUE, driving strong spatial heterogeneity in drought responses across mountain landscapes. Conclusions Our findings highlight the pivotal role of indirect drought effects mediated by vegetation and soil processes in shaping ecosystem WUE. The identified soil–vegetation–climate regulatory hierarchy provides mechanistic insight into landscape–scale drought sensitivity and supports integrated modeling approaches for evaluating ecosystem resilience and sustainable management in arid mountain regions.

China↗

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Classification of orthostatic intolerance through data analytics

Imbalance in the autonomic nervous system can lead to orthostatic intolerance manifested by dizziness, lightheadedness, and a sudden loss of consciousness (syncope); these are common conditions, but they are challenging to diagnose correctly. Uncertainties about the triggering mechanisms and the underlying pathophysiology have led to variations in their classification. This study uses machine learning to categorize patients with orthostatic intolerance. Here we use random forest classification trees to identify a small number of markers in blood pressure, and heart rate time-series data measured during head-up tilt to (a) distinguish patients with a single pathology and (b) examine data from patients with a mixed pathophysiology. Next, we use Kmeans to cluster the markers representing the time-series data. We apply the proposed method analyzing clinical data from 186 subjects identified as control or suffering from one of four conditions: postural orthostatic tachycardia (POTS), cardioinhibition, vasodepression, and mixed cardioinhibition and vasodepression. Classification results confirm the use of supervised machine learning. We were able to categorize more than 95% of patients with a single condition and were able to subgroup all patients with mixed cardioinhibitory and vasodepressor syncope. Clustering results confirm the disease groups and identify two distinct subgroups within the control and mixed groups. The proposed study demonstrates how to use machine learning to discover structure in blood pressure and heart rate time-series data. The methodology is used in classification of patients with orthostatic intolerance. Diagnosing orthostatic intolerance is challenging, and full characterization of the pathophysiological mechanisms remains a topic of ongoing research. This study provides a step toward leveraging machine learning to assist clinicians and researchers in addressing these challenges.

60 APPLIED LIFE SCIENCES↗

A general spatial-temporal framework for short-term building temperature forecasting at arbitrary locations with crowdsourcing weather data

Weather forecasting has been a critical component to predict and control building energy consumption for better building energy management. Without accessibility to other data sources, the onsite observed temperatures or the airport temperatures are used in forecast models. In this paper, we present a novel approach by utilizing the crowdsourcing weather data from neighboring personal weather stations (PWS) to improve the weather forecast accuracy around buildings using a general spatial-temporal modeling framework. The final forecast is based on the ensemble of local forecasts for the target location using neighboring PWSs. Our approach is distinguished from existing literature in various aspects. First, we leverage the crowdsourcing weather data from PWS in addition to public data sources. In this way, the data is at much finer time resolution (e.g., at 5-minute frequency) and spatial resolution (e.g., arbitrary location vs grid). Second, our proposed model incorporates spatial-temporal correlation information of weather variables between the target building and a set of neighboring PWSs so that underlying correlations can be effectively captured to improve forecasting performance. Here, we demonstrate the performance of the proposed framework by comparing to the benchmark models on temperature forecasting for a building located at an arbitrary location at San Antonio, Texas, USA. In general, the proposed model framework equipped with machine learning technique such as Random Forest can improve forecasting by 50% compares with persistent model and has 90% chance to outperform airport forecast in short-term forecasting. In a real-time setting, the proposed model framework can provide more accurate temperature forecasting results compared with using airport temperature forecast for most forecast horizon. Moreover, we analyze the sensitivity of model parameters to gain insights on how crowdsourcing data from the neighboring personal weather stations impacts forecasting performance. Finally, we implement our model in other cities such as Syracuse and Chicago to test the model's performance in different landforms and climate types.

54 ENVIRONMENTAL SCIENCES↗

Macroscopic Traffic Modeling Using Probe Vehicle Data: A Machine Learning Approach

Abstract The macroscopic fundamental diagram (MFD) captures an orderly relationship among traffic flow, density, and speed at the network level. It is a simple yet powerful tool for modeling traffic dynamics in large urban networks with broad application in traffic control and management. However, empirically derived MFDs in urban regions require high-resolution traffic data from the network. Having the network flow and vehicular density estimated at the (granular) census tract level using vehicle probe data, we apply machine learning methods to predict the MFDs across U.S. urban areas and capture the impacts of location-specific input features on the network flow–density relationships at a large scale. The results show that, among the four tested machine learning approaches (Random Forest, XGBoost, Support Vector Machine, and Neural Network), XGBoost delivers the best performance in predicting network traffic flow based on vehicular density and location attributes. Using interaction Shapley Additive explanation (SHAP) values and partial correlation analysis, we examine the factors influencing MFD shapes across different locations. Our empirical findings reveal that across U.S. urban areas, network topology, transportation infrastructure, and land use are primary factors shaping MFD curves, while demand and trip-related factors play a lesser role. Specifically, higher ranking roads, centrality, and development levels correlate positively with network capacity and critical density, whereas negative associations are observed for network connectivity, mixed-use development, and road roughness levels.

Jin, Ling↗

A data-centric weak supervised learning for highway traffic incident detection

Using the data from loop detector sensors for near-real-time detection of traffic incidents on highways is crucial to averting major traffic congestion. While recent supervised machine learning methods offer solutions to incident detection by leveraging human-labeled incident data, the false alarm rate is often too high to be used in practice. Specifically, the inconsistency in the human labeling of the incidents significantly affects the performance of supervised learning models. To that end, we focus on a data-centric approach to improve the accuracy and reduce the false alarm rate of traffic incident detection on highways. We develop a weak supervised learning workflow to generate high-quality training labels for the incident data without the ground truth labels, and we use those generated labels in the supervised learning setup for final detection. This approach comprises three stages. First, we introduce a data preprocessing and curation pipeline that processes traffic sensor data to generate high-quality training data through leveraging labeling functions, which can be domain knowledge-related or simple heuristic rules. Second, we evaluate the training data generated by weak supervision using three supervised learning models-random forest, k-nearest neighbors, and a support vector machine ensemble-and long short-term memory classifiers. The results show that the accuracy of all of the models improves significantly after using the training data generated by weak supervision. Third, we develop an online real-time incident detection approach that leverages the model ensemble and the uncertainty quantification while detecting incidents. Finally, we show that our proposed weak supervised learning workflow achieves a high incident detection rate (0.90) and low false alarm rate (0.08).

97 MATHEMATICS AND COMPUTING↗

Predicting defects in laser powder bed fusion using in-situ thermal imaging data and machine learning

Variation in the local thermal history during the Laser Powder Bed Fusion (LPBF) process in Additive Manufacturing (AM) can cause micropore defects, which add to the uncertainty of the mechanical properties (e.g., fatigue life, tensile strength) of the built materials. In-situ sensing has been proposed for monitoring the AM process to minimize defects, but successful minimization requires establishing a quantitative relationship between the sensing data and the porosity, which is particularly challenging with a large number of variables (e.g., laser speed, power, scan path, powder property). Physics-based modeling can simulate such an in-situ sensing-porosity relationship, but it is computationally costly. In this work, we develop Machine Learning (ML) models that can use in-situ thermographic data to predict the micropore of LPBF stainless steel materials. This work considers two identified key features from the thermal histories: the time above the apparent melting threshold ($\tau$) and the maximum radiance ( T max ). These features are computed, stored for each voxel in the built material, and then used as inputs. The binary state of each voxel, either defective or normal, is the output. Different ML models are trained and tested for the binary classification task. In addition to using the thermal features of each voxel to predict its own state, the thermal features of neighboring voxels are also included as inputs. This is shown to improve the prediction accuracy, which is consistent with thermal transport physics around each voxel contributing to its final state. Among the models trained, the F1 scores on test sets reach above 0.96 for Random Forests. Feature importance analysis based on the ML models shows that T max is more important to the voxel state than $\tau$ . The analysis also finds that the thermal history of the voxels above the present voxel is more influential than those beneath it. Our study significantly extends the capability of using in-situ thermographic data to predict porosity in LPBF materials. Finally, since ML models are fast, they may play integral roles in the optimization and control of such AM technologies.

36 MATERIALS SCIENCE↗