Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

The frequency of extreme X-ray variability for radio-quiet quasars

ABSTRACT We analyse 1598 serendipitous Chandra X-ray observations of 462 radio-quiet quasars to constrain the frequency of extreme amplitude X-ray variability that is intrinsic to the quasar corona and innermost accretion flow. The quasars in this investigation are all spectroscopically confirmed, optically bright (mi ≤ 20.2), and contain no identifiable broad absorption lines in their optical/ultraviolet spectra. This sample includes quasars spanning z ≈ 0.1–4 and probes X-ray variability on time-scales of up to ≈12 rest-frame years. Variability amplitudes are computed between every epoch of observation for each quasar and are analysed as a function of time-scale and luminosity. The tail-heavy distributions of variability amplitudes at all time-scales indicate that extreme X-ray variations are driven by an additional physical mechanism and not just typical random fluctuations of the coronal emission. Similarly, extreme X-ray variations of low-luminosity quasars seem to be driven by an additional physical mechanism, whereas high-luminosity quasars seem more consistent with random fluctuations. The amplitude at which an X-ray variability event can be considered extreme is quantified for different time-scales and luminosities. Extreme X-ray variations occur more frequently at long time-scales (Δt ≳ 300 d) than at shorter time-scales and in low-luminosity quasars compared to high-luminosity quasars over a similar time-scale. A binomial analysis indicates that extreme intrinsic X-ray variations are rare, with a maximum occurrence rate of $\lt 2.4{{\ \rm per\ cent}}$ of observations. Finally, we present X-ray variability and basic optical emission-line properties of three archival quasars that have been newly discovered to exhibit extreme X-ray variability.

Timlin, III, John D.↗

Probabilistic Context Neighborhood model for lattices

Here we present the Probabilistic Context Neighborhood model designed for two-dimensional lattices as a variation of a Markov random field assuming discrete values. In this model, the neighborhood structure has a fixed geometry but a variable order, depending on the neighbors’ values. Our model extends the Probabilistic Context Tree model, originally applicable to one-dimensional space. It retains advantageous properties, such as representing the dependence neighborhood structure as a graph in a tree format, facilitating an understanding of model complexity. Furthermore, we adapt the algorithm used to estimate the Probabilistic Context Tree to estimate the parameters of the proposed model. We illustrate the accuracy of our estimation methodology through simulation studies. Additionally, we apply the Probabilistic Context Neighborhood model to spatial real-world data, showcasing its practical utility.

97 MATHEMATICS AND COMPUTING↗

Evaluation of residual gas fraction estimation methods for cycle-to-cycle combustion variability analysis and modeling

Cycle-to-cycle combustion variability in spark-ignition engines during normal operation is mainly caused by random perturbations of the in-cylinder conditions such as the flow velocity field, homogeneity of the air-fuel distribution, spark energy discharge, and turbulence intensity of the flame front. Such perturbations translate into the variability of the energy released observed at the end of the combustion process. During normal operating conditions, the cycle-to-cycle variability (CCV) of the energy release behaves as random uncorrelated noise. However, during diluted combustion, in either the form of exhaust gas recirculation (EGR) or excess air (lean operation), the CCV tends to increase as dilution increases. Moreover, when the ignition limit is reached at high dilution levels, the combustion CCV is exacerbated by sporadic occurrences of incomplete combustion events, and the uncorrelation assumption no longer holds. The low or null energy released by partial burns and misfires has an impact on the following combustion event due to the residual gas that carries burned and unburned gases, which contributes to the deterministic coupling between engine cycles. Many residual gas fraction estimation methods, however, only address the nominal case where complete combustion occurs and combustion events are uncorrelated. Here we evaluate the efficacy of such methods on capturing the effects of partial burns and misfires on the residual gas estimate for high-EGR operation. The advantages and disadvantages of each method are discussed based on their ability to generate cycle-to-cycle estimates. Finally, a comparison between the different estimation techniques is presented based on their usefulness for control-oriented modeling.

42 ENGINEERING↗

A Predictor-Corrector Strategy for Adaptivity in Dynamical Low-Rank Approximations

Here, in this paper, we present a predictor-corrector strategy for constructing rank-adaptive, dynamical low-rank approximations (DLRAs) of matrix-valued ODE systems. The strategy is a compromise between (i) low-rank step-truncation approaches that alternately evolve and compress solutions and (ii) strict DLRA approaches that augment the low-rank manifold using subspaces generated locally in time by the DLRA integrator. The strategy is based on an analysis of the error between a forward temporal update into the ambient full-rank space, which is typically computed in a step-truncation approach before recompressing, and the standard DLRA update, which is forced to live in a low-rank manifold. We use this error, without requiring its full-rank representation, to correct the DLRA solution. A key ingredient for maintaining a low-rank representation of the error is a randomized SVD, which introduces some degree of stochastic variability into the implementation. The strategy is formulated and implemented in the context of discontinuous Galerkin spatial discretizations of PDEs and applied to several versions of DLRA methods found in the literature as well as a new variant. Numerical experiments comparing the predictor-corrector strategy to other methods demonstrate robustness to overcome shortcomings of step truncation or strict DLRA approaches: The former may require more memory than is strictly needed, while the latter may miss transients solution features that cannot be recovered. The effect of randomization, tolerances, and other implementation parameters is also explored.

97 MATHEMATICS AND COMPUTING↗

Operator Relaxation and the Optimal Depth of Classical Shadows

Classical shadows are a powerful method for learning many properties of quantum states in a sample-efficient manner, by making use of randomized measurements. Here we study the sample complexity of learning the expectation value of Pauli operators via “shallow shadows,” a recently proposed version of classical shadows in which the randomization step is effected by a local unitary circuit of variable depth t. Here we show that the shadow norm (the quantity controlling the sample complexity) is expressed in terms of properties of the Heisenberg time evolution of operators under the randomizing (“twirling”) circuit—namely the evolution of the weight distribution characterizing the number of sites on which an operator acts nontrivially. For spatially contiguous Pauli operators of weight k, this entails a competition between two processes: operator spreading (whereby the support of an operator grows over time, increasing its weight) and operator relaxation (whereby the bulk of the operator develops an equilibrium density of identity operators, decreasing its weight). From this simple picture we derive (i) an upper bound on the shadow norm which, for depth t~log⁡(k), guarantees an exponential gain in sample complexity over the t=0 protocol in any spatial dimension, and (ii) quantitative results in one dimension within a mean-field approximation, including a universal subleading correction to the optimal depth, found to be in excellent agreement with infinite matrix product state numerical simulations. Our Letter connects fundamental ideas in quantum many-body dynamics to applications in quantum information science, and paves the way to highly optimized protocols for learning different properties of quantum states.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Machine Learning of Key Variables Impacting Extreme Precipitation in Various Regions of the Contiguous United States

Abstract Amplification in extreme precipitation intensity and frequency can cause severe flooding and impose significant social and economic consequences. Variations in extreme precipitation intensity, frequencies, and return periods can be attributed to many physical variables across spatial and temporal scales. Here we employ ensemble machine learning (ML) methods, namely random forest (RF), eXtreme Gradient Boosting (XGB), and artificial neural networks (ANN), to explore key contributing variables to monthly extreme precipitation intensity and frequency in six regions over the United States. We further establish emulators for return periods. Results show that the ML models for intensity perform better in regions with obvious seasonality (i.e., Northern Great Plains, Southern Great Plains, and West Coast) than the other three regions (Northeast, Southwest, and Rocky Mountains), while for frequency the models perform well for most regions. The Shapley additive explanation is used to help explain the relationships between extreme precipitation characteristics and identify top variables for RF and XGB. We find that latent heat flux, relative humidity, soil moisture, and large‐scale subsidence are key common variables across the regions for both monthly intensity and frequency, and their compound effects are non‐negligible. The developed ML models capture the probability and return period of extreme precipitation well for all regions and may be used for decision making (e.g., infrastructure planning and design).

54 ENVIRONMENTAL SCIENCES↗

Land surface dynamics and meteorological forcings modulate land surface temperature characteristics

This study examines the effect of land cover, vegetation health, climatic forcings, elevation heat loads, and terrain characteristics (LVCET) on land surface temperature (LST) distribution over West Africa (WA). We employ fourteen machine-learning models, which preserve nonlinear relationships, to downscale LST and other predictands while preserving the geographical variability of WA. Our results showed that the random forest model performs best in downscaling predictands. This is important for the sub-region since it has limited access to mainframes to power multiplex machine-learning algorithms. In contrast to the northern regions, the southern regions consistently exhibit healthy vegetation. Also, areas with unhealthy vegetation coincide with hot LST clusters. The positive Normalized Difference Vegetation Index (NDVI) trends in the Sahel underscore rainfall recovery and subsequent Sahelian greening. The southwesterly winds cause the upwelling of cold waters, lowering LST in southern WA and highlighting the cooling influence of water bodies on LST. Identifying regions with elevated LST is paramount for prioritizing greening initiatives, and our study underscores the importance of considering LVCET factors in urban planning. Topographic slope-facing angles, heat loads, and diurnal anisotropic heat all contribute to variations in LST, emphasizing the need for a holistic approach when designing resilient and sustainable landscapes.

54 ENVIRONMENTAL SCIENCES↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗

The relation between quasars’ optical spectra and variability

Abstract Brightness variation is an essential feature of quasars, but its mechanism and relationship to other physical quantities are not understood well. We aimed to find the relationship between the optical variability and spectral features to reveal the regularity behind the random variation. It is known that a quasar’s Fe ii/Hβ flux ratio and equivalent width of [O iii]5007 are negatively correlated; this is called Eigenvector 1. In this work, we visualized the relationship between the position on this Eigenvector 1 (EV1) plane and how the brightness of the quasars had changed after ∼10 yr. We conducted three analyses, using a different quasar sample in each. The first analysis showed the relation between the quasars’ distributions on the EV1 plane and how much they had changed brightness, using 13438 Sloan Digital Sky Survey quasars. This result shows how brightness changes later are clearly related to the position on the EV1 plane. In the second analysis, we plotted the sources reported as “changing-look quasars” (or “changing-state quasars”) on the EV1 plane. This result shows that the position on the EV1 plane corresponds to the activity level of each source, and the bright or dim states of them are distributed on the opposite sides divided by the typical quasar distribution. In the third analysis, we examined the transition vectors on the EV1 plane using sources with multiple-epoch spectra. This result shows that the brightening and dimming sources move on a similar path and they reach a position corresponding to the opposite activity level. We also found this trend is opposite to the empirical rule that $R_{\rm {Fe\, \small {II}}}$ positively correlated with the Eddington ratio, which has been proposed based on the trends of a large number of quasars. From all these analyses, it is indicated that quasars tend to oscillate between both sides of the distribution ridge on the EV1 plane; each of them corresponds to a dim state and a bright state. This trend in optical variation suggests that significant brightness changes, such as changing-look quasars, are expected to repeat.

Astronomy & Astrophysics↗

Analysis of human performance differences between students and operators when using the Rancor Microworld simulator

Here, from within the umbrella of the Simplified Human Error Experimental Program (SHEEP) framework, this paper analyzes human performance differences between professional and student operators when using a simplified simulator (i.e., Rancor Microworld). This paper represents a crucial step in understanding the fidelity of the simplified simulators and student operators within the SHEEP study. This paper explores a randomized factorial experimental design that features two independent variables: participant type and event class. Six human performance measurements are considered in the experiment. The experiment is conducted using 20 professional reactor operators employed at actual nuclear power plants (NPPs), along with 20 trained students. The experimental data are analyzed via statistical analysis methods. Finally, this paper examines the differences in human performance between actual operators and students when using Rancor Microworld.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Quantifying nitrogen loss hotspots and mitigation potential for individual fields in the US Corn Belt with a metamodeling approach

The high productivity in the US Corn Belt is largely enabled by the consumption of millions of tons of manufactured fertilizer. Excessive application of nitrogen (N) fertilizer has been pervasive in this region, and the unrecovered N eventually escaped from croplands in forms of nitrous oxide (N 2 O) emission and N leaching. Mitigating these negative impacts is hindered by a lack of practical information on where to focus and how much mitigation potential to expect. At a large scale, process-based crop models are the primary tools for predicting variables required by decision making, but their applications are prohibited by expensive computational and data storage costs. To overcome these challenges, we built a series of metamodels to learn the key mechanisms regarding the carbon (C) and N cycle from a well-validated process-based biogeochemical model, ecosys. The trained metamodel captures over 98% of the variability of the ecosys simulated outputs for 99 randomly selected counties in Iowa, Illinois, and Indiana. To identify hotspots with high mitigation potential, we introduce net societal benefit (NSB) as an indicator for synthesizing the loss in yield and social benefits through emissions and pollutants avoided. Our results show that reducing N fertilizer by 10% leads to 9.8% less N 2 O emissions and 9.6% less N leaching at the cost of 4.9% more SOC depletion and 0.6% yield reduction over the study region. The estimated total annual NSB is $\$395$ M (uncertainty ranges from $\$114$ M to $\$1271$ M), including $\$334$ from social benefits (uncertainty ranges from $\$46$ M to $\$1076$ M), $\$100$ M from saving fertilizer (uncertainty ranges from $\$13$ M to $\$455$ M), and –$\$40$ M due to yield changes (uncertainty ranges from –$\$261$ M to $\$69$ M). For the median scenario, we noted that 20% of the study area accounts for nearly 50% of the NSB, and thus represent hotspot locations for targeted mitigation. Although the uncertainty range suggests that developing such a high-resolution framework is not yet settled and the scenario based estimations are not appropriate to inform the management practices for individual farmers, our efforts shed light on the new generation of analytical tools for life cycle assessment.

54 ENVIRONMENTAL SCIENCES↗

Polarization-agnostic continuous-variable quantum key distribution

Here, we introduce a polarization-agnostic method for Gaussian-modulated coherent-state (GCMS) continuous-variable quantum key distribution (CVQKD). Due to the random and continuous nature of the GCMS protocol, Alice, the transmitter, can encode two distinct quadratures in each of two orthogonal polarization modes, such that Bob, the receiver, measures valid GCMS quadratures in a single polarization mode even when polarization changes occur during transmission. This method does not require polarization correction in the optical domain, does not require monitoring both polarization modes, reduces loss by eliminating optical components, and avoids the noise injected by polarization correction algorithms.

Williams, Brian P. [Oak Ridge National Laboratory ↗

Evolving Metrics for Resource Adequacy Assessment

Resource adequacy analysis quantifies the likelihood of capacity shortfall on a power system in a probabilistic manner. Using a combination of statistical techniques and power system fundamentals, the analysis typically evaluates hundreds or thousands of stochastic random samples (replications) of varying load, generator outages, variable renewable energy availability, and other aspects of power system uncertainty. In this range of uncertainty, there are - at times - periods where the power system's available resources are insufficient to meet system demand, referred to as a shortfall event. Today's power systems' rapidly evolving generation mix is changing the types of data needed by system planners and regulators, which can often render traditional resource adequacy metrics insufficient for ensuring resource adequacy for tomorrow's grid. In this paper we provide a critical assessment of traditional measures of shortfall risk in power systems, discussing their shortcomings and how they compare to metrics used in other domains. From this analysis we propose four steps forward for improving power system resource adequacy risk metrics in the future.

ENERGY PLANNING, POLICY, AND ECONOMY,POWER TRANSMI↗

Beyond Expected Values Evolving Metrics for Resource Adequacy Assessment

Resource adequacy analysis quantifies the likelihood of capacity shortfall on a power system in a probabilistic manner. Using a combination of statistical techniques and power system fundamentals, the analysis typically evaluates hundreds or thousands of stochastic random samples (replications) of varying load, generator outages, variable renewable energy availability, and other aspects of power system uncertainty. In this range of uncertainty, there are - at times - periods where the power system's available resources are insufficient to meet system demand, referred to as a shortfall event. Today's power systems' rapidly evolving generation mix is changing the types of data needed by system planners and regulators, which can often render traditional resource adequacy metrics insufficient for ensuring resource adequacy for tomorrow's grid. In this paper we provide a critical assessment of traditional measures of shortfall risk in power systems, discussing their shortcomings and how they compare to metrics used in other domains. From this analysis we propose four steps forward for improving power system resource adequacy risk metrics in the future.

ENERGY PLANNING, POLICY, AND ECONOMY↗

Data-assisted combustion simulations with dynamic submodel assignment using random forests

This investigation outlines a data-assisted approach that employs random forest classifiers for local and dynamic submodel assignment in turbulent-combustion simulations. This method is demonstrated in simulations of a single-element GOX/GCH4 rocket combustor; a priori as well as a posteriori assessments are conducted to (i) evaluate the accuracy and adjustability of the classifier for targeting different quantities of interest (QoIs), and (ii) assess improvements, resulting from the data-assisted combustion model assignment, in predicting target QoIs during simulation runtime. Results from the a priori study show that random forests, trained with local flow properties as input variables and combustion model errors as training labels, assign three different combustion models – finite-rate chemistry (FRC), flamelet progress variable (FPV) model, and inert mixing (IM) – with reasonable classification performance even when targeting multiple QoIs. Applications in a posteriori studies demonstrate improved predictions from data-assisted simulations, in temperature and CO mass fraction, when compared with monolithic FPV calculations. An additional a posteriori data-assisted simulation of a modified configuration demonstrates that the present approach can be successfully applied to different configurations, as long as thermophysical behavior can be represented by the training data. Furthermore, these results demonstrate that this data-driven framework holds promise for dynamic combustion submodel assignments in reacting flow simulations.

42 ENGINEERING↗

Probabilistic Forecasting of Generators Startups and Shutdowns in the MISO System Based on Random Forest

Solving security constrained unit commitment (SCUC) problems to plan an economical generation schedule for day-head electricity market has been an important research topic in recent years. Mixed integer programming method (MIP), the-state-of-art approach for solving SCUC problem, is known computationally hard when the number of binary status variables is large. In this paper, a machine learning-based algorithm - random forest (RF), was applied to forecast the startups (SU) and shutdowns (SD) hours of generators, based on historical hourly system condition observations in the Midcontinent Independent System Operator (MISO) system. The main purpose is to reduce the number of binary status variables, by fixing the SU/SD hours to a narrow range of high confidence. This would significantly reduce the size of the decision space, and therefore speed up SCUC solutions with reduced uncertainty.

Lin, Xinming↗

Model Inputs, Outputs, and Scripts associated with: “Combined effects of stream hydrology and land use on basin-scale hyporheic zone denitrification in the Columbia River Basin”

This data package is associated with the publication “Combined effects of stream hydrology and land use on basin‐scale hyporheic zone denitrification in the Columbia River Basin”, published in Water Resource Research (Son et al.2022) available at https://doi.org/10.1029/2021WR031131. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) denitrification at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions, including two-step denitrification and aerobic respiration, are considered within the HZ. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values. This study uses the RCM to explore the spatial patterns of HZ denitrification across reaches with different sizes and land use in the CRB. Our main objective is to use the RCM as a virtual reality model, and the machine-learning models as surrogates that encapsulate the complexities of the physics-based model while identifying the importance of different variables that are not evident in the model conceptualization. We do not include a direct comparison of the modeled HZ denitrification and measurements; however, the RCM can capture the overall spatial patterns of the HZ denitrification because the model inputs and its reaction networks are based on well-established theory and a physical-based model. The combination of the model-based predictions and a machine-learning approach (e.g., random forest) is used to improve our understanding of what variables of the model are associated with spatial patterns of the modeled denitrification across reaches with different sizes and land uses, and to develop a proxy model using measurable variables to reproduce the simulated patterns.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Machine learning predictions for local electronic properties of disordered correlated electron systems

We present a scalable machine learning (ML) model to predict local electronic properties such as on-site electron number and double occupation for disordered correlated electron systems. Our approach is based on the locality principle, or the nearsightedness nature, of many-electron systems, which means local electronic properties depend mainly on the immediate environment. A ML model is developed to encode this complex dependence of local quantities on the neighborhood. We demonstrate our approach using the square-lattice Anderson-Hubbard model, which is a paradigmatic system for studying the interplay between Mott transition and Anderson localization. We develop a lattice descriptor based on the group-theoretical method to represent the on-site random potentials within a finite region. The resultant feature variables are used as input to a multilayer fully connected neural network, which is trained from data sets of variational Monte Carlo (VMC) simulations on small systems. We show that the ML predictions agree reasonably well with the VMC data. Our work underscores the promising potential of ML methods for multiscale modeling of correlated electron systems.

36 MATERIALS SCIENCE↗