Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Cross-national analysis of food security drivers: comparing results based on the Food Insecurity Experience Scale and Global Food Security Index

Abstract The second UN Sustainable Development Goal establishes food security as a priority for governments, multilateral organizations, and NGOs. These institutions track national-level food security performance with an array of metrics and weigh intervention options considering the leverage of many possible drivers. We studied the relationships between several candidate drivers and two response variables based on prominent measures of national food security: the 2019 Global Food Security Index (GFSI) and the Food Insecurity Experience Scale’s (FIES) estimate of the percentage of a nation’s population experiencing food security or mild food insecurity (FI ). We compared the contributions of explanatory variables in regressions predicting both response variables, and we further tested the stability of our results to changes in explanatory variable selection and in the countries included in regression model training and testing. At the cross-national level, the quantity and quality of a nation’s agricultural land were not predictive of either food security metric. We found mixed evidence that per-capita cereal production, per-hectare cereal yield, an aggregate governance metric, logistics performance, and extent of paid employment work were predictive of national food security. Household spending as measured by per-capita final consumption expenditure (HFCE) was consistently the strongest driver among those studied, alone explaining a median of 92% and 70% of variation (based on out-of-sample R 2 ) in GFSI and FI , respectively. The relative strength of HFCE as a predictor was observed for both response variables and was independent of the countries used for model training, the transformations applied to the explanatory variables prior to model training, and the variable selection technique used to specify multivariate regressions. The results of this cross-national analysis reinforce previous research supportive of a causal mechanism where, in the absence of exceptional local factors, an increase in income drives increase in food security. However, the strength of this effect varies depending on the countries included in regression model fitting. We demonstrate that using multiple response metrics, repeated random sampling of input data, and iterative variable selection facilitates a convergence of evidence approach to analyzing food security drivers.

42 ENGINEERING↗

Formation of Amorphous Carbon Multi‐Walled Nanotubes from Random Initial Configurations

Amorphous carbon nanotubes (a‐CNT) with up to four walls and sizes ranging from 200 to 3200 atoms have been simulated, starting from initial random configurations and using the Gaussian Approximation Potential. The important variables (like density, height, and diameter) required to successfully simulate a‐CNTs were predicted with the machine learning random forest technique. The width of the a‐CNT models ranged between 0.55–2 nm with an average inter‐wall spacing of 0.31 nm. The topological defects in a‐CNTs were analyzed and new defect configurations were observed. The electronic density of states and localization in these phases were discussed and delocalized electrons in the π subspace were identified as an important factor for inter‐layer cohesion. Spatial projection of the electronic conductivity favors axial transport along connecting hexagons, while non‐hexagonal parts of the network either hinder or bifurcate the electronic transport. A vibrational density of states was calculated and is potentially an experimentally comparable fingerprint of the material. The appearance of a low‐frequency radial breathing mode was discussed and the thermal conductivity at 300 K was estimated using the Green‐Kubo formula.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Predicting nutrition and environmental factors associated with female reproductive disorders using a knowledge graph and random forests

Female reproductive disorders (FRDs) are common health conditions that may present with significant symptoms. Diet and environment are potential areas for FRD interventions. We utilized a knowledge graph (KG) method to predict factors associated with common FRDs (for example, endometriosis, ovarian cyst, and uterine fibroids). We harmonized survey data from the Personalized Environment and Genes Study (PEGS) on internal and external environmental exposures and health conditions with biomedical ontology content. We merged the harmonized data and ontologies with supplemental nutrient and agricultural chemical data to create a KG. We analyzed the KG by embedding edges and applying a random forest for edge prediction to identify variables potentially associated with FRDs. We also conducted logistic regression analysis for comparison. Across 9765 PEGS respondents, the KG analysis resulted in 8535 significant or suggestive predicted links between FRDs and chemicals, phenotypes, and diseases. Amongst these links, 32 were exact matches when compared with the logistic regression results, including comorbidities, medications, foods, and occupational exposures. Mechanistic underpinnings of predicted links documented in the literature may support some of our findings. Our KG methods are useful for predicting possible associations in large, survey-based datasets with added information on directionality and magnitude of effect from logistic regression. These results should not be construed as causal but can support hypothesis generation. This investigation enabled the generation of hypotheses on a variety of potential links between FRDs and exposures. Future investigations should prospectively evaluate the variables hypothesized to impact FRDs.

60 APPLIED LIFE SCIENCES↗

Circulating levels of micronutrients and risk of osteomyelitis: a Mendelian randomization study

Background Few observational studies have investigated the effect of micronutrients on osteomyelitis, and these findings are limited by confounding and conflicting results. Therefore, we conducted Mendelian randomization (MR) analyses to evaluate the association between blood levels of eight micronutrients (copper, selenium, zinc, vitamin B12, vitamin C, and vitamin D, vitamin B6, vitamin E) and the risk of osteomyelitis. Methods We performed the two-sample and multivariable Mendelian randomization (MVMR) to investigate causation, where instrument variables for the predictor (micronutrients) were derived from the summary data of micronutrients from independent cohorts of European ancestry. The outcome instrumental variables were used from the summary data of European-ancestry individuals ( n = 486,484). The threshold of statistical significance was set at p < 0.00625. Results We found a significant causal association that elevated zinc heightens the risk of developing osteomyelitis in European ancestry individuals OR = 1.23 [95% confidence interval (CI) [1.07, 1.43]; p = 4.26E-03]. Similarly, vitamin B6 showed a similar significant causal effect on osteomyelitis as a risk factor OR = 2.78 (95% CI [1.34, 5.76]; p = 6.04E-03; in the secondary analysis). Post-hoc analysis suggested this result (vitamin B6). However, the multivariable Mendelian randomization (MVMR) provides evidence against the causal association between zinc and osteomyelitis OR = 0.98(95% CI [−0.11, 0.07]; p = 7.20E-1). After searching in PhenoScanner, no SNP with confounding factors was found in the analysis of vitamin B6. There was no evidence of a reverse causal impact of osteomyelitis on zinc and vitamin B6. Conclusion This study supported a strong causal association between vitamin B6 and osteomyelitis while reporting a dubious causal association between zinc and osteomyelitis.

Zhang, Xu↗

The variability structure function of the highest luminosity quasars on short time-scales

ABSTRACT The stochastic photometric variability of quasars is known to follow a random-walk phenomenology on emission time-scales of months to years. Some high-cadence rest-frame optical monitoring in the past has hinted at a suppression of variability amplitudes on shorter time-scales of a few days or weeks, opening the question of what drives the suppression and how it might scale with quasar properties. Here, we study a few thousand of the highest luminosity quasars in the sky, mostly in the luminosity range of $L_{\rm bol}$$=[46.4, 47.3]$ and redshift range of $z=[0.7, 2.4]$. We use a data set from the NASA/Asteroid Terrestrial-impact Last Alert System facility with nightly cadence, weather permitting, which has been used before to quantify strong regularity in longer term rest-frame-UV variability. As we focus on a careful treatment of short time-scales across the sample, we find that a linear function is sufficient to describe the UV variability structure function. Although the result can not rule out the existence of breaks in some groups completely, a simpler model is usually favoured under this circumstance. In conclusion, the data are consistent with a single-slope random walk across rest-frame time-scales of $\Delta t=[10, 250]$ d.

Tang, Ji-Jia (ORCID:0000000218600886)↗

A data-driven global soil heterotrophic respiration dataset and the drivers of its inter-annual variability

Soil heterotrophic respiration (SHR), one of the primary carbon fluxes from terrestrial ecosystems to the atmosphere, is important for carbon-climate feedbacks because of its sensitivity to available litter and soil carbon, climatic conditions, and nutrient availability. However, until recently limited SHR data were available, and most published global SHR estimates have either a short time span, coarse spatial resolution, or reply on overly-simple model formulations. To better understand and quantify the global distribution of SHR and its sensitivity to climate variability, we produced a new global SHR dataset using Random Forest algorithms, up-scaling 455 point data from the Global Soil Respiration Database (SRDB 4.0) with gridded fields of climatic, edaphic and productivity as explanatory variables. We estimated a global total SHR of 46.8 Pg C yr-1 over 1985-2013 (95% confidence interval: 38.6-56.3 Pg C yr-1), with a significant increasing trend of 0.03 Pg C yr-2 during this period. We found that the choice of soil moisture datasets contributes more to the difference among these data-driven SHR members rather than that of productivity, temperature and precipitation data sources. We also analyzed the influence of climatic variables on the inter-annual variability (IAV) of our SHR product. Water availability was the dominant driver of IAV at global scales, although the inferred sensitivity depends on the choice of the soil moisture gridded dataset. At the ecosystem scale, temperature strongly controls the IAV of SHR in tropical forests, while water availability dominates in extra-tropical forest and semi-arid regions. Our machine-learning gridded SHR dataset and outputs from process-based land surface models (TRENDYv6) show agreement for a strong association between water variability and SHR IAV at the global scale, but the two approaches lead to different temporal trend globally and different controlling variables for IAV at the ecosystem scale. Our study provides evidence for the pervasive and important role of water availability in driving SHR, indicating both a direct effect limiting decomposition rates and an indirect effect through the amount of fresh organic matter made available to SHR from productivity. In consideration of potential limitations and uncertainties remaining in our data-driven SHR datasets, we call for a more scientifically designed observation network for SHR, more observation data compilation, and increased use of deep learning methods making maximum use of observation data in hand. This will benefit process-based models, and improve our understanding of SHR response to future anomalous environmental conditions.

Yao, Yitong↗

Approximating a linear multiplicative objective in watershed management optimization

Implementing management practices in a cost-efficient manner is critical for regional efforts to reduce the amount of pollutants entering the Chesapeake Bay. We study the problem of selecting a subset of practices that minimizes pollutant load—subject to budgetary and environmental constraints—as simulated in a widely used regulatory watershed model. Mimicking the computation of pollutant load in the regulatory model, we formulate this problem as a continuous optimization model with a linear multiplicative objective function and linear constraints. To lay the groundwork for incorporating additional stakeholder requirements in the future, especially those that would require integer variables, we present and study a continuous linear optimization model that approximates the nonlinear model. The linear model, which requires an exponential number of variables, arises naturally as an alternative model for the same underlying physical process. We examine the theoretical behavior of these optimization models and investigate restrictions of the linear model to handle its large number of variables. Through extensive computational tests on real and randomly generated instances, we demonstrate that the linear model and its restrictions provide optimal solutions close to those of the nonlinear model in practice, despite poor approximation properties in the worst case. We conclude that the linear model—together with our approach to handling its large number of variables—provides a viable framework from which to extend the optimization model to better meet the needs of the Chesapeake Bay watershed management stakeholders.

54 ENVIRONMENTAL SCIENCES↗

Uncertainty-aware mixed-variable machine learning for materials design

Abstract Data-driven design shows the promise of accelerating materials discovery but is challenging due to the prohibitive cost of searching the vast design space of chemistry, structure, and synthesis methods. Bayesian optimization (BO) employs uncertainty-aware machine learning models to select promising designs to evaluate, hence reducing the cost. However, BO with mixed numerical and categorical variables, which is of particular interest in materials design, has not been well studied. In this work, we survey frequentist and Bayesian approaches to uncertainty quantification of machine learning with mixed variables. We then conduct a systematic comparative study of their performances in BO using a popular representative model from each group, the random forest-based Lolo model (frequentist) and the latent variable Gaussian process model (Bayesian). We examine the efficacy of the two models in the optimization of mathematical functions, as well as properties of structural and functional materials, where we observe performance differences as related to problem dimensionality and complexity. By investigating the machine learning models’ predictive and uncertainty estimation capabilities, we provide interpretations of the observed performance differences. Our results provide practical guidance on choosing between frequentist and Bayesian uncertainty-aware machine learning models for mixed-variable BO in materials design.

36 MATERIALS SCIENCE↗

Observation of spatter-induced stochastic lack-of-fusion in laser powder bed fusion using in situ process monitoring

Material produced via additive manufacturing (AM) continues to exhibit variable mechanical properties despite apparent optimization of processing parameters, inhibiting qualification efforts and limiting use in critical applications. Stochastic lack-of-fusion flaws may help explain this variability, but the origin of these seemingly random defects has to this point remained unclear. In this work, we show that spatter particles, material ejected from the laser melt pool, are directly responsible for generating stochastic lack-of-fusion in laser-based powder bed fusion components through the application of spatial statistics. Herein a statistically significant, causal relationship between spatter particles and stochastic lack-of-fusion is established, and the spatial and morphological relationships between spatter and internal flaws are investigated. The occurrence of spatter-induced lack-of-fusion in relation to the inert gas flow and laser trajectory direction is also investigated, and recommendations for mitigating the occurrence of spatter are evaluated.

36 MATERIALS SCIENCE↗

Limits on the evolutionary rates of biological traits

Abstract This paper focuses on the maximum speed at which biological evolution can occur. I derive inequalities that limit the rate of evolutionary processes driven by natural selection, mutations, or genetic drift. These rate limits link the variability in a population to evolutionary rates. In particular, high variances in the fitness of a population and of a quantitative trait allow for fast changes in the trait’s average. In contrast, low variability makes a trait less susceptible to random changes due to genetic drift. The results in this article generalize Fisher’s fundamental theorem of natural selection to dynamics that allow for mutations and genetic drift, via trade-off relations that constrain the evolutionary rates of arbitrary traits. The rate limits can be used to probe questions in various evolutionary biology and ecology settings. They apply, for instance, to trait dynamics within or across species or to the evolution of bacteria strains. They apply to any quantitative trait, e.g., from species’ weights to the lengths of DNA strands.

59 BASIC BIOLOGICAL SCIENCES↗

Limits on the Evolutionary Rates of Biological Traits

This paper focuses on the maximum speed at which biological evolution can occur. I derive inequalities that limit the rate of evolutionary processes driven by natural selection, mutations, or genetic drift. These rate limits link the variability in a population to evolutionary rates. In particular, high variances in the fitness of a population and of a quantitative trait allow for fast changes in the trait’s average. In contrast, low variability makes a trait less susceptible to random changes due to genetic drift. The results in this article generalize Fisher’s fundamental theorem of natural selection to dynamics that allow for mutations and genetic drift, via trade-of relations that constrain the evolutionary rates of arbitrary traits. The rate limits can be used to probe questions in various evolutionary biology and ecology settings. They apply, for instance, to trait dynamics within or across species or to the evolution of bacteria strains. They apply to any quantitative trait, e.g., from species’ weights to the lengths of DNA strands.

59 BASIC BIOLOGICAL SCIENCES↗

Multifrequency Models of Black Hole Photon Rings from Low-luminosity Accretion Disks

Images of black holes encode both astrophysical and gravitational properties. Detecting highly lensed features in images can differentiate between these two effects. We present an accretion disk emission model coupled to the Adaptive Analytical Ray Tracing (AART) code that allows a fast parameter space exploration of black hole photon ring images produced from synchrotron emission from 10 to 670 GHz. As an application, we systematically study several disk models and compute their total flux density, average radii, and optical depth. The model parameters are chosen around fiducial values calibrated to general relativistic magnetohydrodynamic (GRMHD) simulations and observations of M87*. For the parameter space studied, we characterize the transition between optically thin and thick regimes and the frequency at which the first photon ring is observable. Our results highlight the need for careful definitions of photon ring radius in the image domain, as in certain models the highly lensed photon ring is dimmer than the direct emission at certain angles. We find that at low frequencies the ring radii are set by the electron temperature, while at higher frequencies the magnetic field strength plays a more significant role, demonstrating how multifrequency analysis can also be used to infer plasma parameters. Lastly, we show how our implementation can qualitatively reproduce multifrequency black hole images from GRMHD simulations when adding time variability to our disk model through Gaussian random fields. This approach provides a new method for simulating observations from the Event Horizon Telescope and the proposed Black Hole Explorer space mission.

79 ASTRONOMY AND ASTROPHYSICS↗

Probabilistic Context Neighborhood model for lattices

Here we present the Probabilistic Context Neighborhood model designed for two-dimensional lattices as a variation of a Markov random field assuming discrete values. In this model, the neighborhood structure has a fixed geometry but a variable order, depending on the neighbors’ values. Our model extends the Probabilistic Context Tree model, originally applicable to one-dimensional space. It retains advantageous properties, such as representing the dependence neighborhood structure as a graph in a tree format, facilitating an understanding of model complexity. Furthermore, we adapt the algorithm used to estimate the Probabilistic Context Tree to estimate the parameters of the proposed model. We illustrate the accuracy of our estimation methodology through simulation studies. Additionally, we apply the Probabilistic Context Neighborhood model to spatial real-world data, showcasing its practical utility.

97 MATHEMATICS AND COMPUTING↗

A Predictor-Corrector Strategy for Adaptivity in Dynamical Low-Rank Approximations

Here, in this paper, we present a predictor-corrector strategy for constructing rank-adaptive, dynamical low-rank approximations (DLRAs) of matrix-valued ODE systems. The strategy is a compromise between (i) low-rank step-truncation approaches that alternately evolve and compress solutions and (ii) strict DLRA approaches that augment the low-rank manifold using subspaces generated locally in time by the DLRA integrator. The strategy is based on an analysis of the error between a forward temporal update into the ambient full-rank space, which is typically computed in a step-truncation approach before recompressing, and the standard DLRA update, which is forced to live in a low-rank manifold. We use this error, without requiring its full-rank representation, to correct the DLRA solution. A key ingredient for maintaining a low-rank representation of the error is a randomized SVD, which introduces some degree of stochastic variability into the implementation. The strategy is formulated and implemented in the context of discontinuous Galerkin spatial discretizations of PDEs and applied to several versions of DLRA methods found in the literature as well as a new variant. Numerical experiments comparing the predictor-corrector strategy to other methods demonstrate robustness to overcome shortcomings of step truncation or strict DLRA approaches: The former may require more memory than is strictly needed, while the latter may miss transients solution features that cannot be recovered. The effect of randomization, tolerances, and other implementation parameters is also explored.

97 MATHEMATICS AND COMPUTING↗

Operator Relaxation and the Optimal Depth of Classical Shadows

Classical shadows are a powerful method for learning many properties of quantum states in a sample-efficient manner, by making use of randomized measurements. Here we study the sample complexity of learning the expectation value of Pauli operators via “shallow shadows,” a recently proposed version of classical shadows in which the randomization step is effected by a local unitary circuit of variable depth t. Here we show that the shadow norm (the quantity controlling the sample complexity) is expressed in terms of properties of the Heisenberg time evolution of operators under the randomizing (“twirling”) circuit—namely the evolution of the weight distribution characterizing the number of sites on which an operator acts nontrivially. For spatially contiguous Pauli operators of weight k, this entails a competition between two processes: operator spreading (whereby the support of an operator grows over time, increasing its weight) and operator relaxation (whereby the bulk of the operator develops an equilibrium density of identity operators, decreasing its weight). From this simple picture we derive (i) an upper bound on the shadow norm which, for depth t~log⁡(k), guarantees an exponential gain in sample complexity over the t=0 protocol in any spatial dimension, and (ii) quantitative results in one dimension within a mean-field approximation, including a universal subleading correction to the optimal depth, found to be in excellent agreement with infinite matrix product state numerical simulations. Our Letter connects fundamental ideas in quantum many-body dynamics to applications in quantum information science, and paves the way to highly optimized protocols for learning different properties of quantum states.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Machine Learning of Key Variables Impacting Extreme Precipitation in Various Regions of the Contiguous United States

Abstract Amplification in extreme precipitation intensity and frequency can cause severe flooding and impose significant social and economic consequences. Variations in extreme precipitation intensity, frequencies, and return periods can be attributed to many physical variables across spatial and temporal scales. Here we employ ensemble machine learning (ML) methods, namely random forest (RF), eXtreme Gradient Boosting (XGB), and artificial neural networks (ANN), to explore key contributing variables to monthly extreme precipitation intensity and frequency in six regions over the United States. We further establish emulators for return periods. Results show that the ML models for intensity perform better in regions with obvious seasonality (i.e., Northern Great Plains, Southern Great Plains, and West Coast) than the other three regions (Northeast, Southwest, and Rocky Mountains), while for frequency the models perform well for most regions. The Shapley additive explanation is used to help explain the relationships between extreme precipitation characteristics and identify top variables for RF and XGB. We find that latent heat flux, relative humidity, soil moisture, and large‐scale subsidence are key common variables across the regions for both monthly intensity and frequency, and their compound effects are non‐negligible. The developed ML models capture the probability and return period of extreme precipitation well for all regions and may be used for decision making (e.g., infrastructure planning and design).

54 ENVIRONMENTAL SCIENCES↗

Land surface dynamics and meteorological forcings modulate land surface temperature characteristics

This study examines the effect of land cover, vegetation health, climatic forcings, elevation heat loads, and terrain characteristics (LVCET) on land surface temperature (LST) distribution over West Africa (WA). We employ fourteen machine-learning models, which preserve nonlinear relationships, to downscale LST and other predictands while preserving the geographical variability of WA. Our results showed that the random forest model performs best in downscaling predictands. This is important for the sub-region since it has limited access to mainframes to power multiplex machine-learning algorithms. In contrast to the northern regions, the southern regions consistently exhibit healthy vegetation. Also, areas with unhealthy vegetation coincide with hot LST clusters. The positive Normalized Difference Vegetation Index (NDVI) trends in the Sahel underscore rainfall recovery and subsequent Sahelian greening. The southwesterly winds cause the upwelling of cold waters, lowering LST in southern WA and highlighting the cooling influence of water bodies on LST. Identifying regions with elevated LST is paramount for prioritizing greening initiatives, and our study underscores the importance of considering LVCET factors in urban planning. Topographic slope-facing angles, heat loads, and diurnal anisotropic heat all contribute to variations in LST, emphasizing the need for a holistic approach when designing resilient and sustainable landscapes.

54 ENVIRONMENTAL SCIENCES↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗