Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Brief communication: Monitoring snow depth using small, cheap, and easy-to-deploy snow–ground interface temperature sensors

Abstract. Temporally continuous snow depth estimates are vital for understanding changing snow patterns and impacts on permafrost in the Arctic. We trained a random forest machine learning model to predict snow depth from variability in snow–ground interface temperature. The model performed well on Alaska's Seward Peninsula where it was trained and at Arctic evaluation sites (RMSE ≤ 0.15 m). It performed poorly at temperate sites with deeper snowpacks, partially due to training data limitations. Small temperature sensors are cheap and easy to deploy, so this technique enables spatially distributed and temporally continuous snowpack monitoring at high latitudes to an extent previously infeasible.

54 ENVIRONMENTAL SCIENCES↗

A semi-agnostic ansatz with variable structure for variational quantum algorithms

Quantum machine learning—and specifically Variational Quantum Algorithms (VQAs)—offers a powerful, flexible paradigm for programming near-term quantum computers, with applications in chemistry, metrology, materials science, data science, and mathematics. Here, one trains an ansatz, in the form of a parameterized quantum circuit, to accomplish a task of interest. However, challenges have recently emerged suggesting that deep ansatzes are difficult to train, due to flat training landscapes caused by randomness or by hardware noise. This motivates our work, where we present a variable structure approach to build ansatzes for VQAs. Our approach, called VAns (Variable Ansatz), applies a set of rules to both grow and (crucially) remove quantum gates in an informed manner during the optimization. Consequently, VAns is ideally suited to mitigate trainability and noise-related issues by keeping the ansatz shallow. We employ VAns in the variational quantum eigensolver for condensed matter and quantum chemistry applications, in the quantum autoencoder for data compression and in unitary compilation problems showing successful results in all cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Effects of forest structural and compositional change on forest microclimates across a gradient of disturbance severity

Forest structural diversity and community composition are key in regulating forest microclimates. When disturbance affects structural diversity or composition, forest microclimates may be altered due to changes in soil temperature, soil water content, and light availability. It is unclear however which structural or compositional components, when changed or to what extent, result in microclimatic change. To address this question, we used data from a large scale, manipulative stem-girdling experiment in northern, lower Michigan—the Forest Resilience and Threshold Experiment (FoRTE). FoRTE follows a factorial design with multiple levels of disturbance severity (0, 45, 65, 85%) based on targeted reductions in gross leaf area index via stem-girdling induced mortality. These disturbance severity treatments are applied in two ways: either as top-down (largest trees are killed) or bottom-up (small to medium trees killed) treatments. We examined how multiple components of structural diversity and community composition changed as a product of disturbance severity and type, and then tested for resulting effects on forest microclimates (light availability, soil temperature, and soil water), using a multivariate, Random Forest framework. We found that measures of community composition (species richness, species evenness, and Shannon-Wiener Diversity Index) and stand structure (basal area, standard deviation of DBH, tree size diversity) declined more following disturbance than did measures of canopy cover, heterogeneity, arrangement, or height. However, when changes in each variable from pre- to post-disturbance, measured as log change, were employed in a multivariate, Random Forest regression framework, structural diversity measures of heterogeneity (rugosity, top rugosity), cover (canopy cover), and arrangement (porosity) were the most influential variables, but with differences among bottom-up and top-down treatments We found that the death of large trees from disturbance impacts soil temperature, water, and light environments more substantially and uniformly across disturbance gradients than does the death of smaller trees. Furthermore, our results have implications for both statistical and process-based modeling of forest disturbance.

54 ENVIRONMENTAL SCIENCES↗

Ensemble Estimation of Historical Evapotranspiration for the Conterminous U.S.

Abstract Evapotranspiration (ET) is the largest component of the water budget, accounting for the majority of the water available from precipitation. ET is challenging to quantify because of the uncertainties associated with the many ET equations currently in use, and because observations of ET are uncertain and sparse. In this study, we combine information provided by available ET data and equations to produce a new monthly data set for ET for the conterminous U.S. (CONUS). These maps are produced from 1895 to 2018 at an 800 m spatial scale, marking a finer resolution than currently available products over this time period. In our approach, the relative performance of a suite of ET equations is assessed using water balance, flux tower, and remotely sensed ET estimates. At the observation locations, we use error distributions to quantify relative weights for the equations and use these in a modified Bayesian model averaging weighted ensemble approach. The relative weights are spatially generalized using a random forest regression, which is applied to wall‐to‐wall explanatory variable maps to generate CONUS‐wide relative weight maps and ensemble estimates. We assess the performance of the ensemble using a reserved subset of the observations and compare this performance against other national‐scale map products for historical to modern ET. The ensemble ET maps are shown to provide an improved accuracy over the alternative comparison products. These ET maps could be useful for a variety of hydrologic modeling and assessment applications that benefit from a long record, such as the study of periods of water scarcity through time.

Environmental Sciences & Ecology↗

Classification Analysis of Southwest Pacific Tropical Cyclone Intensity Changes Prior to Landfall

This study evaluates the ability of a random forest classifier to identify tropical cyclone (TC) intensification or weakening prior to landfall over the western region of the Southwest Pacific Ocean (SWPO) basin. For both Australia mainland and SWPO island cases, when a TC first crosses land after spending ≥24 h over the ocean, the closest hour prior to the intersection is considered as the landfall hour. If the maximum wind speed (V max ) at the landfall hour increased or remained the same from the 24-h mark prior to landfall, the TC is labeled as intensifying and if the V max at the landfall hour decreases, the TC is labeled as weakening. Geophysical and aerosol variables closest to the 24 h before landfall hour were collected for each sample. The random forest model with leave-one-out cross validation and the random oversampling example technique was identified as the best-performing classifier for both mainland and island cases. The model identified longitude, initial intensity, and sea skin temperature as the most important variables for the mainland and island landfall classification decisions. Incorrectly classified cases from the test data were analyzed by sorting the cases by their initial intensity hour, landfall hour, monthly distribution, and 24-h intensity changes. TC intensity changes near land strongly impact coastal preparations such as wind damage and flood damage mitigations; hence, this study will contribute to improve identifying and prioritizing prediction of important variables contributing to TC intensity change before landfall.

54 ENVIRONMENTAL SCIENCES↗

Latent Stochastic Differential Equations for Modeling Quasar Variability and Inferring Black Hole Properties

Quasars are bright and unobscured active galactic nuclei (AGN) thought to be powered by the accretion of matter around supermassive black holes at the centers of galaxies. The temporal variability of a quasar’s brightness contains valuable information about its physical properties. The UV/optical variability is thought to be a stochastic process, often represented as a damped random walk described by a stochastic differential equation (SDE). Upcoming wide-field telescopes such as the Rubin Observatory Legacy Survey of Space and Time (LSST) are expected to observe tens of millions of AGN in multiple filters over a ten year period, so there is a need for efficient and automated modeling techniques that can handle the large volume of data. Latent SDEs are machine learning models well suited for modeling quasar variability, as they can explicitly capture the underlying stochastic dynamics. In this work, we adapt latent SDEs to jointly reconstruct multivariate quasar light curves and infer their physical properties such as the black hole mass, inclination angle, and temperature slope. Our model is trained on realistic simulations of LSST ten year quasar light curves, and we demonstrate its ability to reconstruct quasar light curves even in the presence of long seasonal gaps and irregular sampling across different bands, outperforming a multioutput Gaussian process regression baseline. Our method has the potential to provide a deeper understanding of the physical properties of quasars and is applicable to a wide range of other multivariate time series with missing data and irregular sampling.

79 ASTRONOMY AND ASTROPHYSICS↗

Utilizing physics-based input features within a machine learning model to predict wind speed forecasting error

Machine learning is quickly becoming a commonly used technique for wind speed and power forecasting. Many machine learning methods utilize exogenous variables as input features, but there remains the question of which atmospheric variables are most beneficial for forecasting, especially in handling non-linearities that lead to forecasting error. This question is addressed via creation of a hybrid model that utilizes an autoregressive integrated moving-average (ARIMA) model to make an initial wind speed forecast followed by a random forest model that attempts to predict the ARIMA forecasting error using knowledge of exogenous atmospheric variables. Variables conveying information about atmospheric stability and turbulence as well as inertial forcing are found to be useful in dealing with non-linear error prediction. Streamwise wind speed, time of day, turbulence intensity, turbulent heat flux, vertical velocity, and wind direction are found to be particularly useful when used in unison for hourly and 3 h timescales. The prediction accuracy of the developed ARIMA–random forest hybrid model is compared to that of the persistence and bias-corrected ARIMA models. The ARIMA–random forest model is shown to improve upon the latter commonly employed modeling methods, reducing hourly forecasting error by up to 5 % below that of the bias-corrected ARIMA model and achieving an R 2 value of 0.84 with true wind speed.

17 WIND ENERGY↗

Dual origin of ferropericlase inclusions within super-deep diamonds

Ferropericlase [(Mg,Fe)O] is one of the major constituents of Earth's lower mantle and the most abundant mineral inclusion in sub-lithospheric diamonds. Although a lower mantle origin for ferropericlase inclusions has often been suggested, some studies have proposed that many of these inclusions may instead form at much shallower depths, in the deep upper mantle or transition zone. No straightforward method exists to discriminate ferropericlase of lower-mantle origin without characteristic mineral associations, such as co-existing former bridgmanite. To explore ferropericlase-diamond growth relationships, we have investigated the crystallographic orientation relationships (CORs), determined by single-crystal X-ray diffraction, between 57 ferropericlase inclusions and 37 diamonds from Juina (Brazil) and Kankan (Guinea). We show that ferropericlase inclusions can develop specific (16 inclusions in 12 diamonds), rotational statistical (9 inclusions in 7 diamonds) and random (32 inclusions in 25 diamond) CORs with respect to their diamond hosts. All measured inclusions showing a specific COR were found to be Fe-rich (X FeO > 0.30). Coexistence of non-randomly and randomly oriented ferropericlase inclusions within the same diamond indicates that their CORs may be variably affected by local growth conditions. However, the occurrence of specific CORs only for Fe-rich inclusions indicates that Fe-rich ferropericlases have a distinct genesis and are syngenetic with their host diamonds. In conclusion, this result provides strong support for a dual origin for ferropericlase in Earth's mantle, with Fe-rich compositions likely indicating redox growth in the upper mantle, while more Mg-rich compositions with random COR mostly representing ambient lower mantle trapped as protogenetic inclusions.

58 GEOSCIENCES↗

Ensemble cure kinetics network (ECK-Net): A method to derive cure kinetics of thermosetting resin

This paper introduces an Ensemble Cure Kinetics Network (ECK-Net), a neural network (NN)–based framework for modeling the cure kinetics of thermosetting resins within a phenomenological context. ECK-Net replaces traditional analytic models, which require extensive chemical insight and multiple isothermal/non-isothermal experiments, with a data-driven surrogate that maps nonlinear relationships between temperature, degree of cure, and reaction rate from differential scanning calorimetry data. The proposed approach predicts input-dependent kinetic coefficients of a generalized nth-order reaction equation rather than reaction rates directly, enabling a single unified model to represent various epoxy systems without relying on iso-conversional analysis or predefined functional forms. To ensure robustness, multiple independently trained networks under different random initializations are blended through an ensemble strategy, effectively mitigating the stochastic variability inherent to neural networks. The framework is validated using experimental datasets from multiple resin systems, including aerospace-grade materials (Toray 3900-2, Cycom 5320-1, and Hexcel 8552) and a windmill-grade resin (RIMR 035c). The model accurately reproduces the temporal evolution of the degree of cure under manufacturers’ recommended cure cycles across all tested resins systems, yielding Pearson’s correlation coefficients of 0.992, 0.994, 0.993, 0.997, respectively. To demonstrate process-level applicability, the trained network was implemented within the Abaqus environment to simulate out-of-autoclave (OOA) curing process of the CFRP panel composed of Toray T830H-6K/3900-2D prepreg. The simulation results showed excellent agreement with experimental temperature response (maximum peak temperature, simulation: 189.6 °C, experiment: 188.5 °C) and the final degree of cure (simulation: 0.948, experiment: 0.960 ± 0.013), confirming ECK-Net’s capability as a reliable alternative to conventional cure kinetics modeling methods.

Composite curing↗

Metal Foam Morphology Affects the Run to Run Reproducibility of OER Using Nickel Catalysts

Nickel (Ni) foam-based electrodes are excellent catalysts for the oxygen evolution reaction; however, we found that the random pore-size distribution of Ni foams contributes to a significant variability in electrochemically active surface area, compromising experimental reproducibility. We provide insights into quantifying this critical material property, verified by four electrochemical laboratories.

25 ENERGY STORAGE↗

Probabilistic measures for biological adaptation and resilience

This paper introduces an approach to quantifying ecological resilience in biological systems, particularly focusing on noisy systems responding to episodic disturbances with sudden adaptations. Incorporating concepts from nonequilibrium statistical mechanics, we propose a measure termed “ecological resilience through adaptation,” specifically tailored to noisy, forced systems that undergo physiological adaptation in the face of stressful environmental changes. Randomness plays a key role, accounting for model uncertainty and the inherent variability in the dynamical response among components of biological systems. Our measure of resilience is rooted in the probabilistic description of states within these systems and is defined in terms of the dynamics of the ensemble average of a model-specific observable quantifying success or well-being. Our approach utilizes stochastic linear response theory to compute how the expected success of a system, originally in statistical equilibrium, dynamically changes in response to a environmental perturbation and a subsequent adaptation. Importantly, the resulting mathematical derivations allow for the estimation of resilience in terms of ensemble averages of simulated or experimental data. Finally, through a simple but clear conceptual example, we illustrate how our resilience measure can be interpreted and compared to other existing frameworks in the literature. The methodology is general but inspired by applications in plant systems, with the potential for broader application to complex biological processes.

60 APPLIED LIFE SCIENCES↗

Equatorial waves triggering extreme rainfall and floods in southwest Sulawesi, Indonesia

On the basis of detailed analysis of a case study and long-term climatology, it is shown that equatorial waves and their interactions serve as precursors for extreme rain and flood events in the central Maritime Continent region of southwest Sulawesi, Indonesia. Meteorological conditions on January 22, 2019, leading to heavy rainfall and devastating flooding in this area were studied. It is shown that a convectively coupled Kelvin wave (CCKW) and a convectively coupled Rossby wave (CCERW) embedded within the larger-scale envelope of the Madden-Julian Oscillation (MJO) enhanced convective phase, contributed to the onset of a mesoscale convective system which developed over the Java Sea. Low-Level convergence from the CCKW forced mesoscale convective organization and orographic ascent of moist air over the slopes of southwest Sulawesi. Climatological analysis shows that 92 % of December-January-February floods and 76% of extreme rain events in this region were immediately preceded by positive low-level westerly wind anomalies. It is estimated that both CCKWs and CCERWs propagating over Sulawesi double the chance of floods and extreme rain event development, which the probability of such hazardous events occurring during their combined activity is eight times greater than on a random day. While the MJO is a key component shaping tropical atmospheric variability, it is shown that its usefulness as a single factor for extreme weather-driven hazard prediction is limited.

Latos, Beata↗

Comparison of linear regression, k-nearest neighbour and random forest methods in airborne laser-scanning-based prediction of growing stock

Abstract In this study, for five sites around the world, we look at the effects of different model types and variable selection approaches on forest yield modelling performances in an area-based approach (ABA). We compared ordinary least squares regression (OLS), k-nearest neighbours (kNN) and random forest (RF). Our objective was to test if there are systematic differences in accuracy between OLS, kNN and RF in ABA predictions of growing stock volume. The analyses are based on a 5-fold cross-validation at five study sites: an eucalyptus plantation, a temperate forest and three different boreal forests. Two completely independent validation datasets were also available for two of the boreal sites. For the kNN, we evaluated multiple measures of distance including Euclidean, Mahalanobis, most similar neighbour (MSN) and an RF-based distance metric. The variable selection approaches we examined included a heuristic approach (for OLS, kNN and RF), exhaustive search among all combinations (OLS only) and all variables together (RF only). Performances varied by model type and variable selection approaches among sites. OLS and RF had similar accuracies and were more efficient than any of the kNN variants. Variable selection did not affect RF performance. Heuristic and exhaustive variable selection performed similarly for OLS. kNN fared the poorest amongst model types, and kNN with RF distance was prone to overfitting when compared with a validation dataset. Additional caution is therefore required when building kNN models for volume prediction though ABA, being preferable instead to opt for models based on OLS with some variable selection, or RF with all variables together.

Cosenza, Diogo N.↗

Modeling freight mode choice using machine learning classifiers: a comparative study using Commodity Flow Survey (CFS) data

This study explores the usefulness of machine learning classifiers for modeling freight mode choice. We investigate eight commonly used machine learning classifiers, namely Naïve Bayes, Support Vector Machine, Artificial Neural Network, K-Nearest Neighbors, Classification and Regression Tree, Random Forest, Boosting and Bagging, along with the classical Multinomial Logit model. US 2012 Commodity Flow Survey data are used as the primary data source; we augment it with spatial attributes from secondary data sources. The performance of the classifiers is compared based on prediction accuracy results. The current research also examines the role of sample size and training-testing data split ratios on the predictive ability of the various approaches. In addition, the importance of variables is estimated to determine how the variables influence freight mode choice. The results show that the tree-based ensemble classifiers perform the best. Specifically, Random Forest produces the most accurate predictions, closely followed by Boosting and Bagging. With regard to variable importance, shipment characteristics, such as shipment distance, industry classification of the shipper and shipment size, are the most significant factors for freight mode choice decisions.

42 ENGINEERING↗

Peak Rain Rate Sensitivity to Observed Cloud Condensation Nuclei and Turbulence in Continental Warm Shallow Clouds During CACTI

Abstract Warm clouds strongly affect Earth's energy budget but remain imperfectly represented in climate models, partly due to the complexity and covariability of relevant processes influencing warm rain. This work presents a detailed analysis of different factors affecting rain rate peak intensity (RR) in continental warm clouds. Clouds were identified with vertically pointing radar and lidar observations and categorized via a temperature‐based cloud type classification algorithm from which warm clouds were isolated. Observations and retrievals of liquid water path (LWP), cloud condensation nuclei concentration (N CCN ), cloud depth, and cloud duration of more than 3,000 separate warm clouds sampled during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign are analyzed in this work. Multiple linear regression (MLR) and random forest (RF) models are applied to assess the relative impact of these variables on RR. Overall, RR tends to increase as cloud depth, LWP, and cloud duration increase, or N CCN decreases. Cloud depth affects RR the most while N CCN impacts it the least. When considering over 170 warm clouds observed at least 1 hr in which in‐cloud turbulence is retrieved, the effect of N CCN on RR remains most likely suppressive, but it is not significant at a 75% level for MLR and is highly uncertain for RF. The impact of in‐cloud turbulence depends on the moment and location it is sampled. Cloud base turbulence around the time of RR suppresses RR, while cloud top turbulence effects are inconclusive. Possible difficulties in isolating robust CCN and turbulence effects on RR are discussed.

54 ENVIRONMENTAL SCIENCES↗

The active CGCG 077-102 NED02 galaxy within the Abell 2063 galaxy cluster

Context.Within the framework of investigating the link between the central super massive black holes in the cores of galaxies and the galaxies themselves, we detected a variable X-ray source in the center of CGCG 077-102 NED02, which is a member of the CGCG 077-102 galaxy pair within the Abell 2063 cluster of galaxies. Aims.Our goal is to combine X-ray and optical data to demonstrate that this object harbors an active super massive black hole in its core, and to relate this to the dynamical status of the galaxy pair within the Abell 2063 cluster. Methods.We usedChandraandXMM-Newtonarchival data to derive the X-ray spectral shape and variability. We also obtained optical spectroscopy to detect the expected emission lines that are typically found in active galactic nuclei. Finally, we used public ZTF imaging data to investigate the optical variability. Results.There is no evidence of multiple X-ray sources or extended components within CGCG 077-102 NED02. Single X-ray spectral models fit the source well. We detect significant, nonrandom inter-observation 0.5–10 keV X-ray flux variabilities, for observations separated by ∼4 days for short-term variations and by up to ∼700 days for long-term variations. Optical spectroscopy points toward a passive galaxy for CGCG 077-102 NED01 and a Seyfert for CGCG 077-102 NED02. The classification of CGCG 077-102 NED02 is also consistent with its X-ray luminosity of over 10 42 erg s −1 . We do not detect short-term variability in the optical ZTF light curves. However, we find a significant long-term stochastic variability in theg-band that can be well described by the damped random walk model with a best-fit characteristic damping timescale ofτ DRW = 30 −12 +28 days. Finally, the CGCG 077-102 galaxy pair is deeply embedded within the Abell 2063 potential, with a long enough history within this massive structure to have been affected by the influence of this cluster for a long time. Conclusions.Our observations point toward a moderately massive black hole in the center of CGCG 077-102 NED02 of ∼10 6 M ⊙ . As compared to another similar pair in the literature, CGCG 077-102 NED02 is not heavily obscured, perhaps because of the surrounding intracluster medium ram-pressure stripping.

Astronomy & Astrophysics↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗

How Well do Earth System Models Capture Apparent Relationships Between Phytoplankton Biomass and Environmental Variables?

Abstract As phytoplankton form the base of the marine food web, understanding the controls on their abundance is fundamental to understanding marine ecology and its sensitivity to global climate change. While many Earth System Models (ESMs) predict phytoplankton biomass, it is unclear whether they properly capture the mechanistic relationships that control this quantity in the real ocean. We used Random Forest analysis to analyze the output of 13 ESMs as well as two observational data sets. The target variable was phytoplankton carbon and the predictors included environmental parameters known to influence phytoplankton, including nutrients, light, mixed layer depth, salinity, temperature, and upwelling. We examined the following: (a) What fractions of variability in ESMs and observations can be linked to the large‐scale environmental variables simulated by ESMs? (b) What are the dominant predictors and relationships affecting phytoplankton biomass? (c) How well do ESMs simulate phytoplankton carbon and do they simulate the relationships we see in observations? About 88%–96% of the variability in observational data sets and greater than 98% in the ESMs was accounted for by environmental variables known to influence phytoplankton biomass. The dominant predictors in the observational data sets were shortwave radiation and dissolved iron, with temperature and ammonium also relatively important. All the ESMs show that shortwave radiation is the most important variable and most of them predict the right sign of sensitivity to most variables. However, the models predict that biomass reaches maximum levels at unrealistically low levels of iron and unrealistically high levels of light.

Environmental Sciences & Ecology↗