Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “machine learning, Random Forest”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Pyrocumulonimbus Events Over British Columbia in 2017: An Ensemble Model Study of Parameter Sensitivities and Climate Impacts of Wildfire Smoke in the Stratosphere

Abstract The Pyrocumulonimbus (pyroCb) events over British Columbia in 2017 were observed in the lower stratosphere for about 8–10 months after the smoke injections. Several previous studies used global climate models to investigate the physical parameters for the 2017 pyroCb events, but the conclusions show strong model dependency. In this study, we use the Energy Exascale Earth System Model (E3SM) and complete an ensemble of runs exploring three injection parameters: smoke aerosol mass, the percentage of black carbon within the smoke aerosols, and plume injection height. Additionally, we consider the heterogeneous reaction of ozone and primary organic matter. According to the satellite daily observed aerosol optical depth (AOD), we find that the best ensemble member is the simulation with 0.4 Tg of smoke, 3% of which is black carbon, a 13.5 km smoke injection height, and a 10 −5 probability factor of the heterogeneous reaction. Besides AOD, we examine the ensemble score based on the metrics used in previous studies: the metrics of the extinction coefficient at 18 km altitude and the maximum plume height. The conclusion of the best estimate of the injection parameters for 2017 pyroCb events shows strong not only model but also evaluation metric dependency. We use the Random Forest machine learning technique to quantify the relative importance of each parameter in accurately simulating the 2017 pyroCb events and find that the injection height is the most critical feature, no matter which metric is used to score the ensemble members.

54 ENVIRONMENTAL SCIENCES↗

Star formation rate and stellar mass calibrations based on infrared photometry and their dependence on stellar population age and extinction

The stellar mass (M $\star$ ) and the star formation rate (SFR) are among the most important features that characterize galaxies. Measuring these fundamental properties accurately is critical for understanding the present state of galaxies, their history, and future evolution. Infrared (IR) photometry is widely used to measure the M $\star$ and SFR of galaxies because the near-IR traces the continuum emission of the majority of their stellar populations (SPs), and the mid/far-IR traces the dust emission powered by star-forming activity. This work explores the dependence of the IR emission of galaxies on their extinction, and the age of their SPs. It aims to provide accurate and precise IR-photometry SFR and M $\star$ calibrations that account for SP age and extinction while providing quantification of their scatter. We used the CIGALE spectral energy distribution (SED) fitting code to create model SEDs of galaxies with a wide range of star formation histories, dust content, and interstellar medium properties. We fit the relations between M $\star$ and SFR with IR and optical photometry of the model-galaxy SEDs with the Markov chain Monte Carlo (MCMC) method. As an independent confirmation of the MCMC fitting method, we performed a machine-learning random forest (RF) analysis on the same data set. The RF model yields similar results to the MCMC fits, thus validating the latter. This work provides calibrations for the SFR using a combination of the WISE bands 1 and 3, or the JWST NIR-F200W and MIRI-F2100W. It also provides mass-to-light ratio calibrations based on the WISE band-1, the JWST NIR-F200W, and the optical u - r or g - r colors. These calibrations account for the biases attributed to the SP age, while they are given in the form of extinction-dependent and extinction-independent relations. The proposed calibrations show robust estimations while minimizing the scatter and biases throughout a wide range of SFRs and stellar masses. The SFR calibration offers better results, especially in dust-free or passive galaxies where the contributions of old SPs or biases from the lack of dust are significant. Similarly, the M $\star$ calibration yields significantly better results for dusty and high-SFR galaxies where dust emission can otherwise bias the estimations.

79 ASTRONOMY AND ASTROPHYSICS↗

A unified understanding of minimum lattice thermal conductivity

Here, we propose a first-principles model of minimum lattice thermal conductivity ($κ^{min}_L$) based on a unified theoretical treatment of thermal transport in crystals and glasses. We apply this model to thousands of inorganic compounds and find a universal behavior of $κ^{min}_L$ in crystals in the high-temperature limit: The isotropically averaged $κ^{min}_L$ is independent of structural complexity and bounded within a range from ~0.1 to ~2.6 W/(m K), in striking contrast to the conventional phonon gas model which predicts no lower bound. We unveil the underlying physics by showing that for a given parent compound, $κ^{min}_L$ is bounded from below by a value that is approximately insensitive to disorder, but the relative importance of different heat transport channels (phonon gas versus diffuson) depends strongly on the degree of disorder. Moreover, we propose that the diffuson-dominated $κ^{min}_L$ in complex and disordered compounds might be effectively approximated by the phonon gas model for an ordered compound by averaging out disorder and applying phonon unfolding. With these insights, we further bridge the knowledge gap between our model and the well-known Cahill–Watson–Pohl (CWP) model, rationalizing the successes and limitations of the CWP model in the absence of heat transfer mediated by diffusons. Finally, we construct graph network and random forest machine learning models to extend our predictions to all compounds within the Inorganic Crystal Structure Database (ICSD), which were validated against thermoelectric materials possessing experimentally measured ultralow κ L . Our work offers a unified understanding of $κ^{min}_L$, which can guide the rational engineering of materials to achieve .

42 ENGINEERING↗

Moisture availability mediates the relationship between terrestrial gross primary production and solar-induced chlorophyll fluorescence: Insights from global-scale variations

Effective use of solar-induced chlorophyll fluorescence (SIF) to estimate and monitor gross primary production (GPP) in terrestrial ecosystems requires a comprehensive understanding and quantification of the relationship between SIF and GPP. To date, this understanding is incomplete and somewhat controversial in the literature. Here we derived the GPP/SIF ratio from multiple data sources as a diagnostic metric to explore its global-scale patterns of spatial variation and potential climatic dependence. We found that the growing season GPP/SIF ratio varied substantially across global land surfaces, with the highest ratios consistently found in boreal regions. Spatial variation in GPP/SIF was strongly modulated by climate variables. The most striking pattern was a consistent decrease in GPP/SIF from cold-and-wet climates to hot-and-dry climates. We propose that the reduction in GPP/SIF with decreasing moisture availability may be related to stomatal responses to aridity. Furthermore, we show that GPP/SIF can be empirically modeled from climate variables using a machine learning (random forest) framework, which can improve the modeling of ecosystem production and quantify its uncertainty in global terrestrial biosphere models. Finally, our results point to the need for targeted field and experimental studies to better understand the patterns observed and to improve the modeling of the relationship between SIF and GPP over broad scales.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning of factors for improving oyster hatchery production

Oyster aquaculture and restoration in the Chesapeake Bay are vital, yet hatcheries frequently struggle with inconsistent larval growth and sudden mass mortality events. Unpredictable disruptions in larval production cause large economic losses, represent a perceived risk to growers, and impede industry expansion. To better understand associations between production yield and its potential predictors, we applied machine learning (random forest, and neural network) and statistical (generalized additive model) models to a comprehensive dataset of environmental, water quality, and operational parameters from a Maryland oyster hatchery, aiming to identify key yield predictors and develop a robust forecasting tool. We used recursive Boruta algorithm for variable selection, pinpointing critical predictors, and employed cross-validation to fine-tune model settings. Shapley value analysis offered crucial insights into model interpretations, highlighting week number, Normalized Difference Vegetation Index, salinity, turbidity, and fecundity as primary drivers of yield variability. For low-yield cases, salinity-related variables were particularly important. Our findings provide an early warning system for potential production downturns, empowering hatchery operators to make data-driven decisions for optimizing water conditions, feeding schedules, and broodstock management. By boosting predictability and efficiency, this research directly supports economic stability of the oyster industry and ecological health of the Chesapeake Bay.

Vishwakarma, Srishti [Oak Ridge National Laborator↗

Improving Effective Mass Estimations in Pu-Metal Annuli

Determining the effective mass of 240 Pu in a plutonium metal item can be achieved through a number of destructive and non-destructive assay techniques. However, these techniques have one or more shortcomings. These include the need for large quantities of plutonium, long measurement time, or lack of sufficient accuracy. While efforts have been made to mitigate these issues by estimating 240 Pu quantities through neutron coincidence counting techniques, these estimates are subject to systematic bias, and their estimates are not well characterized when other factors of the annulus’ physical form and composition are accounted for. In this work, we expand upon these non-destructive assay techniques via the implementation of random forest machine learning models, which produce correction functions that augment and improve the effective mass estimates derived from classical leakage multiplication, singles, doubles, and triples multiplicity counting equations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Securing Grid-interactive Efficient Buildings (GEB) through Cyber Defense and Resilient System (CYDRES)

The DOE CYDRES project is driven by the urgent need to address critical research gaps in the domain of cyber-physical security of smart buildings, including Grid-interactive Efficient Buildings (GEBs). CYDRES, a real-time advanced building resilient platform, aims to enhance the cyber-attack-immune capabilities of buildings through multi-layered prevention, detection, and adaptation mechanisms. CYDRES consists of five key modules: a multi-layer network analyzer, an Automatic Fault Detection, Diagnosis, and Prognosis (AFDDP) framework, an intelligent mode selector, a cyber-resilient control framework, and a situation awareness platform. The Network Analyzer employs a data-driven framework that includes a protocol state learning tool and a CRF (Conditional Random Field) command validator. In Hardware-In-the-Loop (HIL) testbeds, it achieved 100% detection accuracy with a false alarm rate of 3%, validating its efficacy in identifying selected cyber-attacks. The AFDDP framework leverages pattern matching, PCA (Principal Component Analysis)-based strategies, and a DBN (Dynamic Bayesian Network)-based fault diagnosis approach to pinpoint the causes of physical system abnormalities using Building Automation System (BAS) data. In HIL experiments, the AFDDP module attained a detection accuracy of over 95% with a false alarm rate below 7%. Additionally, the fault detector utilized machine learning (Random Forest) and deep learning (Multi-Layer Perceptron) methods with acoustic sensor data to achieve a 100% fault detection accuracy in Heating, Ventilation, and Air-Conditioning (HVAC) equipment. The Mode Selector offered real-time impact analysis, allowing immediate actions to protect BASs in the face of emerging threats. The cyber-resilient control framework included an adaptive Model Predictive Control (MPC) and a measurement compensator, reducing temperature violations by up to 94% and improving the total demand flexibility by up to 70% in HIL experiments. Such HIL experiments covered a cyber-attack case and a physical fault case, showcasing CYDRES’ efficiency in maintaining operational continuity during threats. The situation awareness platform in Grafana enhanced real-time threat detection and response visualization, augmenting the operational awareness for building operators. CYDRES demonstrated high technical effectiveness in various test scenarios, particularly in HIL environments. The project's phased development approach ensured efficient use of resources, highlighting its practical feasibility and readiness for commercialization. By enhancing the security and resilience of building operations, CYDRES represents a significant advance in mitigating risks associated with cyber-physical systems, thereby enhancing public confidence in the safety of modern building infrastructure. Future directions for the project include expanding testing protocols, refining AFDDP methodologies, exploring more comprehensive resilient control strategies, and testing in real commercial buildings.

42 ENGINEERING↗

Resource Occurrence and Productivity in Existing and Proposed Wind Energy Lease Areas on the Northeast US Shelf

States in the Northeast United States have the ambitious goal of producing more than 22 GW of offshore wind energy in the coming decades. The infrastructure associated with offshore wind energy development is expected to modify marine habitats and potentially alter the ecosystem services. Species distribution models were constructed for a group of fish and macroinvertebrate taxa resident in the Northeast US Continental Shelf marine ecosystem. These models were analyzed to provide baseline context for impact assessment of lease areas in the Middle Atlantic Bight designated for renewable wind energy installations. Using random forest machine learning, models based on occurrence and biomass were constructed for 93 species providing seasonal depictions of their habitat distributions. We developed a scoring index to characterize lease area habitat use for each species. Subsequently, groups of species were identified that reflect varying levels of lease area habitat use ranging across high, moderate, low, and no reliance on the lease area habitats. Among the species with high to moderate reliance were black sea bass ( Centropristis striata ), summer flounder ( Paralichthys dentatus ), and Atlantic menhaden ( Brevoortia tyrannus ), which are important fisheries species in the region. Potential for impact was characterized by the number of species with habitat dependencies associated with lease areas and these varied with a number of continuous gradients. Habitats that support high biomass were distributed more to the northeast, while high occupancy habitats appeared to be further from the coast. There was no obvious effect of the size of the lease area on the importance of associated habitats. Model results indicated that physical drivers and lower trophic level indicators might strongly control the habitat distribution of ecologically and commercially important species in the wind lease areas. Therefore, physical and biological oceanography on the continental shelf proximate to wind energy infrastructure development should be monitored for changes in water column structure and the productivity of phytoplankton and zooplankton and the effects of these changes on the trophic system.

17 WIND ENERGY↗

Brief communication: Monitoring snow depth using small, cheap, and easy-to-deploy snow–ground interface temperature sensors

Abstract. Temporally continuous snow depth estimates are vital for understanding changing snow patterns and impacts on permafrost in the Arctic. We trained a random forest machine learning model to predict snow depth from variability in snow–ground interface temperature. The model performed well on Alaska's Seward Peninsula where it was trained and at Arctic evaluation sites (RMSE ≤ 0.15 m). It performed poorly at temperate sites with deeper snowpacks, partially due to training data limitations. Small temperature sensors are cheap and easy to deploy, so this technique enables spatially distributed and temporally continuous snowpack monitoring at high latitudes to an extent previously infeasible.

54 ENVIRONMENTAL SCIENCES↗

Analysis of Correlation between Cold Weather Meteorological Variables and Electricity Outages

The significance of the impact of weather on the electric grid has grown as climate change continues to increase the frequency and intensity of extreme weather events. In recent years (2021-2022) in particular, extreme winter weather has affected the grid in locations in the US rarely exposed to extreme low temperatures, snow and icing conditions. Here we analyze the correlation between cold weather meteorological variables and electricity outages during two large winter storm events, Uri (February 2021) and Landon (February 2022) using Random Forest machine learning and Pearson’s correlation coefficient. Our geographical focus across the two storms is the state of Texas. Extrapolation of the method to winter weather impacts over other years and additional locations is proposed.

Dumas, Melissa↗

Automatic Classification of Biological Targets in a Tidal Channel Using a Multibeam Sonar

Multibeam sonars are widely used for environmental monitoring of fauna at marine renewable energy sites. However, they can rapidly accrue vast volumes of data, which poses a challenge for data processing. Here, using data from a deployment in a tidal channel with peak currents of 1–2 m s –1 , we demonstrate the data-reduction benefits of real-time automatic classification of targets detected and tracked in multibeam sonar data. First, we evaluate classification capabilities for three machine learning algorithms: random forests, support vector machines, and k-nearest neighbors. For each algorithm, a hill-climbing search optimizes a set of hand-engineered attributes that describe tracked targets. Here, the random forest algorithm is found to be most effective—in postprocessing, discriminating between biological and nonbiological targets with a recall rate of 0.97 and a precision of 0.60. In addition, 89% of biological targets are correctly classified as either seals, diving birds, fish schools, or small targets. Model dependence on the volume of training data is evaluated. Second, a real-time implementation of the model is shown to distinguish between biological targets and nonbiological targets with nearly the same performance as in postprocessing. From this, we make general recommendations for implementing real-time classification of biological targets in multibeam sonar data and the transferability of trained models.

16 TIDAL AND WAVE POWER↗

Predicting Antimicrobial Resistance Using Partial Genome Alignments

Antimicrobial resistance (AMR) is an important global health threat that impacts millions of people worldwide each year. Developing methods that can detect and predict AMR phenotypes can help to mitigate the spread of AMR by informing clinical decision making and appropriate mitigation strategies. Many bioinformatic methods have been developed for predicting AMR phenotypes from whole-genome sequences and AMR genes, but recent studies have indicated that predictions can be made from incomplete genome sequence data. In order to more systematically understand this, we built random forest-based machine learning classifiers for predicting susceptible and resistant phenotypes for Klebsiella pneumoniae (1,640 strains), Mycobacterium tuberculosis (2,497 strains), and Salmonella enterica (1,981 strains). We started by building models from alignments that were based on a reference chromosome for each species. We then subsampled each chromosomal alignment and built models for the resulting subalignments, finding that very small regions, representing approximately 0.1 to 0.2% of the chromosome, are predictive. In K. pneumoniae, M. tuberculosis, and S. enterica, the subalignments are able to predict multiple AMR phenotypes with at least 70% accuracy, even though most do not encode an AMR-related function. We used these models to identify regions of the chromosome with high and low predictive signals. Finally, subalignments that retain high accuracy across larger phylogenetic distances were examined in greater detail, revealing genes and intergenic regions with potential links to AMR, virulence, transport, and survival under stress conditions. IMPORTANCE Antimicrobial resistance causes thousands of deaths annually worldwide. Understanding the regions of the genome that are involved in antimicrobial resistance is important for developing mitigation strategies and preventing transmission. Machine learning models are capable of predicting antimicrobial resistance phenotypes from bacterial genome sequence data by identifying resistance genes, mutations, and other correlated features. They are also capable of implicating regions of the genome that have not been previously characterized as being involved in resistance. In this study, we generated global chromosomal alignments for Klebsiella pneumoniae, Mycobacterium tuberculosis, and Salmonella enterica and systematically searched them for small conserved regions of the genome that enable the prediction of antimicrobial resistance phenotypes. In addition to known antimicrobial resistance genes, this analysis identified genes involved in virulence and transport functions, as well as many genes with no previous implication in antimicrobial resistance.

59 BASIC BIOLOGICAL SCIENCES↗

A machine learning approach to galaxy properties: joint redshift–stellar mass probability distributions with Random Forest

We demonstrate that highly accurate joint redshift–stellar mass probability distribution functions (PDFs) can be obtained using the Random Forest (RF) machine learning (ML) algorithm, even with few photometric bands available. As an example, we use the Dark Energy Survey (DES), combined with the COSMOS2015 catalogue for redshifts and stellar masses. We build two ML models: one containing deep photometry in the griz bands, and the second reflecting the photometric scatter present in the main DES survey, with carefully constructed representative training data in each case. We validate our joint PDFs for 10 699 test galaxies by utilizing the copula probability integral transform and the Kendall distribution function, and their univariate counterparts to validate the marginals. Benchmarked against a basic set-up of the template-fitting code bagpipes, our ML-based method outperforms template fitting on all of our predefined performance metrics. In addition to accuracy, the RF is extremely fast, able to compute joint PDFs for a million galaxies in just under 6 min with consumer computer hardware. Such speed enables PDFs to be derived in real time within analysis codes, solving potential storage issues. As part of this work we have developed galpro 1, a highly intuitive and efficient python package to rapidly generate multivariate PDFs on-the-fly. galpro is documented and available for researchers to use in their cosmology and galaxy evolution studies.

79 ASTRONOMY AND ASTROPHYSICS↗

Predicting chatter using machine learning and acoustic signals from low-cost microphones

Machining chatter is a phenomenon resulting from self-oscillation between a machining tool and workpiece. This self-oscillation results in variation on the machined product that reduces the ability to meet desired specifications. Chatter is a widely studied topic as it directly relates to the quality of machined products. Here, this study details the application of a Random Forest (RF) classifier with Recursive Feature Elimination (RFE) to machining audio collected by a single microphone during down-milling operations. This approach allows straightforward feature elimination that results in an easily understood set of analyzed dimensions. Stability is predicted solely based on the classification output of the RF classifier. Our approach proves highly predictive with consistent machining setup and a small sample set. We also review transferability between machining setups and present key findings. Our RF approach demonstrates the ability to analyze and classify chatter through a low-cost approach with limited training data required. The motivation for using a single microphone is to enable detection on machines without other sensors, such as accelerometers, present in the machining setup. The value of the in-process sensor and chatter classifier is highlighted because the machining setup included asymmetric dynamics that reduced the accuracy of the traditional analytical stability solution. We see a natural progression to deploying this audio-only methodology with real-time processing and classification using either a laptop or smartphone. This progression will allow visual indicators during the machining process that can alert machinists of progression into unstable machining processes.

42 ENGINEERING↗

Macroscopic Traffic Modeling Using Probe Vehicle Data: A Machine Learning Approach

Abstract The macroscopic fundamental diagram (MFD) captures an orderly relationship among traffic flow, density, and speed at the network level. It is a simple yet powerful tool for modeling traffic dynamics in large urban networks with broad application in traffic control and management. However, empirically derived MFDs in urban regions require high-resolution traffic data from the network. Having the network flow and vehicular density estimated at the (granular) census tract level using vehicle probe data, we apply machine learning methods to predict the MFDs across U.S. urban areas and capture the impacts of location-specific input features on the network flow–density relationships at a large scale. The results show that, among the four tested machine learning approaches (Random Forest, XGBoost, Support Vector Machine, and Neural Network), XGBoost delivers the best performance in predicting network traffic flow based on vehicular density and location attributes. Using interaction Shapley Additive explanation (SHAP) values and partial correlation analysis, we examine the factors influencing MFD shapes across different locations. Our empirical findings reveal that across U.S. urban areas, network topology, transportation infrastructure, and land use are primary factors shaping MFD curves, while demand and trip-related factors play a lesser role. Specifically, higher ranking roads, centrality, and development levels correlate positively with network capacity and critical density, whereas negative associations are observed for network connectivity, mixed-use development, and road roughness levels.

Jin, Ling↗

Physics-Infused AI/ML Based Digital-Twin Framework for Flow-Induced-Vibration Damage Prediction in a Nuclear Reactor Heat Exchanger

This report summarizes some of the ongoing work related to the development of an expert-elicitation-digital-twin framework for real time damage state prediction in heat exchanger components of a nuclear reactor. The framework is targeted towards predicting damage associated with coupled low cycle fatigue (associated with regular heat-up, cool-down and power operation transients) and high cycle fatigue (associated with flow induced vibration transients). The overall framework will be based on a NoSQL based database, physics-infused-geometry-dependent virtual-sensor data, different AI/ML techniques-based data-driven-predictive-model applications (Apps) and real-time plant sensor measurements available through few existing sensors. Towards this overall goal, this report updates some of the ongoing work, such as on implementation of a NoSQL Database (such as MongoDB), FE based heat transfer analysis of a heat exchanger (e.g. of a PWR steam generator) for generating geometry-dependent virtual sensor data and evaluation of various AI/ML models such as based on multivariate linear regression, ensembled decision-tree based Random-Forest and Gradient-Boosting regression and high-dimensional-kernel-function-transformation based Support-Vector-Machine regression models. The AI/ML models were evaluated for predicting multi-time-series thermal states at thousands of 3D point-clouds

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Analysis of Random Forest Modeling Strategies for Multi-Step Wind Speed Forecasting

Although the random forest (RF) model is a powerful machine learning tool that has been utilized in many wind speed/power forecasting studies, there has been no consensus on optimal RF modeling strategies. This study investigates three basic questions which aim to assist in the discernment and quantification of the effects of individual model properties, namely: (1) using a standalone RF model versus using RF as a correction mechanism for the persistence approach, (2) utilizing a recursive versus direct multi-step forecasting strategy, and (3) training data availability on model forecasting accuracy from one to six hours ahead. These questions are investigated utilizing data from the FINO1 offshore platform and Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) C1 site, and testing results are compared to the persistence method. At FINO1, due to the presence of multiple wind farms and high inter-annual variability, RF is more effective as an error-correction mechanism for the persistence approach. The direct forecasting strategy is seen to slightly outperform the recursive strategy, specifically for forecasts three or more steps ahead. Finally, increased data availability (up to ~8 equivalent years of hourly training data) appears to continually improve forecasting accuracy, although changing environmental flow patterns have the potential to negate such improvement. We hope that the findings of this study will assist future researchers and industry professionals to construct accurate, reliable RF models for wind speed forecasting.

54 ENVIRONMENTAL SCIENCES↗