Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random variables”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Importance of Depth and Artificial Structure as Predictors of Female Red Snapper Reproductive Parameters

Abstract The Red Snapper Lutjanus campechanus is a structure‐associated species occurring across a wide depth range in the northern Gulf of Mexico. We used the random forest machine learning algorithm to understand which habitat and individual fish characteristics could predict reproductive parameters of female Red Snapper. We evaluated fish captured from 2016 to 2018 on three artificial structure types with various structure heights at depths of 100 m or less. Overall, we found that depth and month were important predictors for most reproductive parameters, but the type of structure (artificial reefs, oil platforms, and rigs‐to‐reefs structures) was not important. Maturity was correctly classified in 88.9% of the cases when using the random forest ensemble model, with important predictors including FL, depth, structure height, and month of collection. Spawning seasonality (measured as gonadosomatic index [GSI]) was correctly classified in 59.5% of the cases when using histology reproductive phase, FL, month, and depth variables. Reproductively active or inactive females were correctly classified in 89.3% of the cases using GSI, month, FL, and depth, while females in the developing versus spawning capable phases were correctly classified in 82.2% of the cases using GSI, FL, month, and depth. Histological indicators that show potential spawning within a 36‐h period were correctly classified 61.5% of the time, with the best predictors being depth, FL, GSI, and month. Stepwise regression indicated that month was the only factor that significantly predicted contrasts in relative batch fecundity, with significantly greater values in August compared to all other months. Our findings suggest that female Red Snapper reproductive effort is not consistently or well predicted by artificial structure type or height but that a combination of fish FL, month, and depth can predict reproductive characteristics of female Red Snapper.

Brown‐Peterson, Nancy J.↗

An Observationally Trained Markov Model for MJO Propagation

A Markovian stochastic model is developed for studying the propagation of the Madden-Julian Oscillation (MJO). This model represents the daily changes in real time multivariate MJO (RMM) indices as random functions of their current state and background conditions. The probability distribution function of the RMM changes is obtained using a machine learning algorithm trained to maximize MJO forecast skills using observed daily indices of RMM and different modes of variability. Skillful forecasts are obtained for lead times between 8 and 27 days. Large ensemble simulations by the stochastic model show that with monsoonal changes in the background state, MJO propagation across the Maritime Continent (MC) is most likely to be disrupted in boreal spring and summer when MJO events propagate from favorable conditions over the Indian Ocean to unfavorable ones over the MC, and predictability is higher during spring and summer when MJO activity is away from the MC region.

54 ENVIRONMENTAL SCIENCES↗

Curvature perturbations from stochastic particle production during inflation

We calculate the curvature power spectrum sourced by spectator fields that are excited repeatedly and non-adiabatically during inflation. In the absence of detailed information of the nature of spectator field interactions, we consider an ensemble of models with intervals between the repeated interactions and interaction strengths drawn from simple probabilistic distributions. We show that the curvature power spectra of each member of the ensemble shows rich structure with many features, and there is a large variability between different realizations of the same ensemble. Such features can be probed by the cosmic microwave background (CMB) and large scale structure observations. They can also have implications for primordial black hole formation and CMB spectral distortions. The geometric random walk behavior of the spectator field allows us to calculate the ensembleaveraged power spectrum of curvature perturbations semi-analytically. For sufficiently large stochastic sourcing, the ensemble averaged power spectrum shows a scale dependence arising from the time spent by modes outside the horizon during the period of particle production, in spite of there being no preferred scale in the underlying model. We find that the magnitude of the ensemble-averaged power spectrum overestimates the typical power spectra in the ensemble because the ensemble distribution of the power spectra is highly non-Gaussian with fat tails.

79 ASTRONOMY AND ASTROPHYSICS↗

Using machine learning to derive cloud condensation nuclei number concentrations from commonly available measurements

Cloud condensation nuclei (CCN) number concentrations are an important aspect of aerosol–cloud interactions and the subsequent climate effects; however, their measurements are very limited. We use a machine learning tool, random decision forests, to develop a random forest regression model (RFRM) to derive CCN at 0.4 % supersaturation ([CCN0.4]) from commonly available measurements. The RFRM is trained on the long-term simulations in a global size-resolved particle microphysics model. Using atmospheric state and composition variables as predictors, through associations of their variabilities, the RFRM is able to learn the underlying dependence of [CCN0.4] on these predictors, which are as follows: eight fractions of PM 2.5 (NH 4 , SO 4 , NO 3, secondary organic aerosol (SOA), black carbon (BC), primary organic carbon (POC), dust, and salt), seven gaseous species (NO x , NH 3 , O 3 , SO 2 , OH, isoprene, and monoterpene), and four meteorological variables (temperature (T), relative humidity (RH), precipitation, and solar radiation). The RFRM is highly robust: it has a median mean fractional bias (MFB) of 4.4 % with ≈96.33 % of the derived [CCN0.4] within a good agreement range of -60% 2.5 speciation (NH 4 , SO 4 , NO 3 , and organic carbon (OC)), NO x , O 3 , SO 2 , T, and RH, as well as [CCN0.4] are available. We modify, optimize, and retrain the developed RFRM to make predictions from 19 to 9 of these available predictors. This retrained RFRM (RFRM-ShortVars) shows a reduction in performance due to the unavailability and sparsity of measurements (predictors); it captures the [CCN0.4] variability and magnitude at SGP with ≈67.02 % of the derived values in the good agreement range. This work shows the potential of using the more commonly available measurements of PM 2.5 speciation to alleviate the sparsity of CCN number concentrations' measurements.

54 ENVIRONMENTAL SCIENCES↗

Generation of random geological models using multi-randomization for machine learning

Generating high-fidelity geological models is essential for advancing machine learning (ML) methods in automated seismic interpretation. For instance, seismic images paired with corresponding fault labels are foundational for ML-based fault detection from seismic migration sections. While several open-access datasets of random geological models exist, open-source tools specifically designed to produce large volumes of such models for ML applications remain scarce. To address this gap, we present RGM (Random Geological Model), an open-source software package for efficiently generating 2D and 3D synthetic geological models tailored for ML workflows. RGM supports the creation of diverse model components, including medium property distributions (P-/S-wave velocities and density), seismic reflectivity images (i.e., synthetic migration sections), relative geological time, and discrete fault attributes such as probability, dip, strike, rake, and displacement. It also accommodates the creation of complex geological features such as salt bodies and unconformities. The model generation algorithm employs a multi-randomization strategy, yielding an effectively infinite-dimensional model space that encompasses a wide range of geological scenarios and associated seismic features. Furthermore, RGM incorporates a method to generate synthetic elastic migration images using analytical elastic reflection coefficients combined with frequency-dependent scaling. This functionality enables the creation of training datasets for ML models that leverage elastic seismic images. RGM is implemented in modern object-oriented Fortran, allowing users to flexibly control statistical parameters governing model variability. We demonstrate the capability, performance, and geological realism of the package through comprehensive 2D and 3D examples.

58 GEOSCIENCES↗

Using electronic health record metadata to predict housing instability amongst veterans

Housing instability is considered a significant life stressor and preemptive screening should be applied to identify those at risk for homelessness as early as possible so that they can be targeted for specialized care. We developed models to classify patient outcomes for an established VA Homelessness Screening Clinical Reminder (HSCR), which identifies housing instability, in the two months prior to its administration. Logistic Regression and Random Forest models were fit to classify responses using the last 18 months of document activity. We measure concentration of risk across stratifications of predicted probability and observe an enriched likelihood of finding confirmed false negative responses from veterans with diagnosed housing instability. Positive responses were 34 times more likely to be detected within the top 1 % of patients predicted at risk than from those randomly selected. There is a 1 in 4 chance of detecting false negatives within the top 1 % of predicted risk. Machine learning methods can classify between episodes of housing instability using a data-driven approach that does not rely on variables curated from domain experts. This method has the potential to improve clinicians’ ability to identify veterans who are experiencing housing instability but are not captured by HSCR.

60 APPLIED LIFE SCIENCES↗

De novo design of high-affinity binders of bioactive helical peptides

Many peptide hormones form an α-helix on binding their receptors, and sensitive methods for their detection could contribute to better clinical management of disease. De novo protein design can now generate binders with high affinity and specificity to structured proteins. However, the design of interactions between proteins and short peptides with helical propensity is an unmet challenge. Here we describe parametric generation and deep learning-based methods for designing proteins to address this challenge. We show that by extending RFdiffusion to enable binder design to flexible targets, and to refining input structure models by successive noising and denoising (partial diffusion), picomolar-affinity binders can be generated to helical peptide targets by either refining designs generated with other methods, or completely de novo starting from random noise distributions without any subsequent experimental optimization. The RFdiffusion designs enable the enrichment and subsequent detection of parathyroid hormone and glucagon by mass spectrometry, and the construction of bioluminescence-based protein biosensors. The ability to design binders to conformationally variable targets, and to optimize by partial diffusion both natural and designed proteins, should be broadly useful.

59 BASIC BIOLOGICAL SCIENCES↗

Tools for Water Ingress Testing

The Safety Storage and Engineering Team, as part of the Production Support Services division (PSS-2), is tasked with ensuring the safety of containers used for handling and storage of nuclear materials. As part of this work, water ingress tests are conducted to evaluate the water-tightness of containers intended for in-glovebox use. In collaboration, the statistics group of the Computer and Computational Sciences Division (CCS-6) provided support in developing a statistically defensible approach for determining appropriate sample sizes for water ingress testing. Water ingress testing involves multiple measurements on multiple containers. Our approach uses a simple random effects model to analyze a pilot data set, implementing prediction limits to evaluate the efficacy of collecting additional data. Although this study capitalizes on available data, our approach can be used with estimates of the ratio of between and within variability and average values, often available from past testing or expert knowledge. An interactive Shiny tool was developed as a final user-friendly product for future testing. The Shiny interface is an open-source package providing a framework for building web applications. Raw data exploration and prediction interval-based sample size assessments can quickly be conducted by the engineering team without needing to interact with the underlying code.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

What matters the most? Understanding individual tornado preparedness using machine learning

Scholars from various disciplines have long attempted to identify the variables most closely associated with individual preparedness. Therefore, we now have much more knowledge regarding these factors and their association with individual preparedness behaviors. However, it has not been sufficiently discussed how decisive many of these factors are in encouraging preparedness. In this article, we seek to examine what factors, among the many examined in previous studies, are most central to engendering emergency preparedness in individuals particularly for tornadoes by utilizing a relatively uncommon machine learning technique in disaster management literature. Using unique survey data, we find that in the case of tornado preparedness the most decisive variables are related to personal experiences and economic circumstances rather than basic demographics. Our findings contribute to scholarly endeavors to understand and promote individual tornado preparedness behaviors by highlighting the variables most likely to shape tornado preparedness at an individual level.

54 ENVIRONMENTAL SCIENCES↗

A probabilistic creep model incorporating test condition, initial damage, and material property uncertainty

Uncertainty is prevalent in the creep resistance of alloys, where at elevated temperature and low pressure, rupture can range across logarithmic decades. In this study, a probabilistic continuum-damage-mechanics (CDM)-based model is derived to capture the uncertainty of creep resistance. To meet this objective, creep data for alloy 304 Stainless Steel is gathered. A constitutive model, “Sinh”, is calibrated deterministically to determine the statistical variability of the material properties. Three sources of uncertainty are injected into the model: test condition (stress and temperature), initial damage, and material properties. Probabilistic simulations are carried out by (a) calibrating probability distribution functions (pdfs) for each source of uncertainty (b) randomly sampling the pdfs using Monte Carlo methods and (c) executing simulations to replicate the uncertain creep behavior. A sensitivity analysis is performed to evaluate the relative effect of each source of uncertainty. In full probabilistic simulations, the cumulative uncertainty of creep behavior is evaluated. The probabilistic model accurately predicts the creep deformation and rupture of the available experiments. The probabilistic model is validated for interpolation but lacks extrapolation ability. Several future works are proposed to further improve the model.

36 MATERIALS SCIENCE↗

Nanometer-thick ultraflat cantilever resonators

Nanomechanical devices made from ultrathin materials are transforming diverse fields, including sensing, signal processing, and quantum technologies. However, as these materials become thinner, their low bending rigidity poses significant fabrication challenges, and achieving nanometer-thick flat cantilevers with consistent and predictable mechanical responses has remained elusive despite decades of research. Here we present nanometer-thick, ultraflat cantilever resonators fabricated using atomic layer deposition. By effectively mitigating the effects of uncontrollable built-in strain and geometric disorder, the ultraflat nanocantilevers exhibit resonance frequencies closely aligned with thin-plate theory predictions and display low sample-to-sample variability. These cantilevers maintain mechanical stability in both vacuum and air environments, even at large length-to-thickness ratios of up to 3000. The ultraflat nanocantilevers are approaching the thickness limit, beyond which thermal fluctuations at room temperature can spontaneously induce random ripples in otherwise flat films.

nanofabrication↗

Fabrication-conscious neural network based inverse design of single-material variable-index multilayer films

Multilayer films with continuously varying indices for each layer have attracted great deal of attention due to their superior optical, mechanical, and thermal properties. However, difficulties in fabrication have limited their application and study in scientific literature compared to multilayer films with fixed index layers. In this work we propose a neural network based inverse design technique enabled by a differentiable analytical solver for realistic design and fabrication of single material variable-index multilayer films. This approach generates multilayer films with excellent performance under ideal conditions. We furthermore address the issue of how to translate these ideal designs into practical useful devices which will naturally suffer from growth imperfections. By integrating simulated systematic and random errors just as a deposition tool would into the optimization process, we demonstrated that the same neural network that produced the ideal device can be retrained to produce designs compensating for systematic deposition errors. Furthermore, the proposed approach corrects for systematic errors even in the presence of random fabrication imperfections. The results outlined in this paper provide a practical and experimentally viable approach for the design of single material multilayer film stacks for an extremely wide variety of practical applications with high performance.

36 MATERIALS SCIENCE↗

Evaluation of Two-Dimensional to One-Dimensional Site Response for Idaho National Laboratory

We perform two-dimensional (2D) site response analyses accounting for spatial variability of soil properties and subsurface geometry of the Eastern Snake River Plane (ESRP), and quantify their effects on ground surface motion relative to one-dimensional (1D) site response analyses at the Idaho National Laboratory (INL). We first present the development of random field idealizations of the repeated basalt lava flows, heterogeneously inter-layered with sediments, from seismic velocity data collected over four decades in the ESRP. Using realizations of the stochastic fields mapped on 2D deterministic finite element models, we perform 2D viscoelastic and equivalent-linear wave propagation simulations, and quantify the mean and variance of site response aggravation factors, defined as the response spectral ratio of 2D to 1D analyses on the ground surface. Results are shown to be insensitive to the constitutive material behavior considered here, for strains induced by rock outcrop peak ground acceleration (PGA) as high as 0.7g: viscoelastic and equivalent-linear analyses predict peak mean 2D/1D aggravation factor 1.05 at period T=0.075 sec (i.e. the 2D response spectrum is 5% higher than the corresponding 1D at that period, on average), which corresponds to the wavelength of the horizontal correlation length of the random field (50m). For periods longer than the fundamental period of the site (here, T 1 =0.3125 sec), the propagating wavelengths are too long to be affected by the 1D site response and the aggravation factor becomes equal to 1. The standard deviation of the natural logarithms of the 2D/1D aggravation factors is ~0.15 for periods shorter than the fundamental period of the site, and decays thereafter at a steady rate.

58 GEOSCIENCES↗

Redefining Resource Adequacy for Modern Power Systems: A Report of the Redefining Resource Adequacy Task Force

Today's rapidly increasing levels of wind, solar, storage, and load flexibility require the industry to rethink reliability planning and resource adequacy methods for modern power systems. Periods with a risk of shortfall often no longer coincide with peak demand - reliability risks are less about peak load and more about the daily setting of the sun, extended cloud cover, wind speeds, cold snaps, and heat waves. In addition, demand is increasingly flexible. Key resources are time-sensitive, as batteries need time to recharge and electricity customers can only be asked to provide demand response for just so long. And reliability failures are often correlated - with one another and with the weather. Two driving factors require the industry to reconsider its analytical approach for resource adequacy: (1) Chronological grid operations: The increasing importance of variable renewable resources (such as wind and solar) and of energy-limited resources (storage and demand response) make it essential to understand the full year of chronological operation of the grid. Specific attention must be paid to hourly, seasonal, and inter-annual resource variability. The sequence of the variability is key, as energy-limited resources such as batteries or demand response require either a preceding period or subsequent period of high production to be useful for grid reliability. (2) Correlated events: Historically, resource adequacy analysis focused on shortfalls caused by random, discrete mechanical failures of large generating units. In contrast, shortfalls today are often caused by multiple, correlated events caused by common weather patterns. Resource adequacy analysis must increasingly shift its focus to these correlated events. The redesign of resource adequacy methods will benefit from a set of guiding principles to better allow for sharing of insights and best practices, interregional resource coordination, and a smoother regulatory process for resource procurement. The objective of this report is to move this redesign forward. It provides an overview of key drivers changing the way resource adequacy needs to be evaluated, identifies shortcomings of conventional approaches, and outlines first principles for practitioners to consider as they adapt their approaches. The central message is: what got us here won't get us there.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Development of Short-Term Forecasting Models Using Plant Asset Data and Feature Selection

Nuclear power plants collect and store large volumes of heterogeneous data from various components and systems. With recent advances in machine learning (ML) techniques, these data can be leveraged to develop diagnostic and short-term forecasting models to better predict future equipment condition. Maintenance operations can then be planned in advance whenever degraded performance is predicted, thus resulting in fewer unplanned outages and the optimization of maintenance activities. This enables lower maintenance costs and improves the overall economics of nuclear power. This paper focuses on developing a short-term forecasting process that leverages a feature selection process to distill large volumes of heterogeneous data and predict specific equipment parameters. A variety of feature selection methods, including Shapley Additive Explanations (SHAP) and variance inflation factor (VIF), were used to select the optimal features as inputs for three ML methods: long short-term memory (LSTM) networks, support vector regression (SVR), and random forest (RF). Each combination of model and input features was used to predict a pump bearing temperature both 1 and 24 hours in advance, based on actual plant system data. The optimal inputs for the LSTM and SVR were selected using the SHAP values, while the optimal input for the RF consisted solely of the response variable itself. Each model produced similar 1-hour-ahead predictions, with root mean square errors (RMSEs) of roughly 0.006. For the 24-hour-ahead predictions, differences could be seen between LSTM, SVR, and RF, as reflected by model performances of 0.036 +- 0.014, 0.0026 +- 0, and 0.063 +- 0.004 RMSE, respectively. As big data and continuous online monitoring become more widely available, the proposed feature selection process can be used for many applications beyond the prediction of process parameters within nuclear infrastructure.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A flexible class of priors for orthonormal matrices with basis function-specific structure

Statistical modeling of high-dimensional matrix-valued data motivates the use of a low-rank representation that simultaneously summarizes key characteristics of the data and enables dimension reduction. Low-rank representations commonly factor the original data into the product of orthonormal basis functions and weights, where each basis function represents an independent feature of the data. However, the basis functions in these factorizations are typically computed using algorithmic methods that cannot quantify uncertainty or account for basis function correlation structure a priori. While there exist Bayesian methods that allow for a common correlation structure across basis functions, empirical examples motivate the need for basis function-specific dependence structure. We propose a prior distribution for orthonormal matrices that can explicitly model basis function-specific structure. The prior is used within a general probabilistic model for singular value decomposition to conduct posterior inference on the basis functions while accounting for measurement error and fixed effects. We discuss how the prior specification can be used for various scenarios and demonstrate favorable model properties through synthetic data examples. Finally, we apply our method to two-meter air temperature data from the Pacific Northwest, enhancing our understanding of the Earth system’s internal variability.

97 MATHEMATICS AND COMPUTING↗

Modeling household-level party composition behavior for multiparty activities: a random parameter nested logit modeling approach

This study presents findings of a household-level party composition model for multiparty activities. It exploits data from a comprehensive Household Travel Survey conducted by Chicago Metropolitan Agency of Planning. The study estimates a random parameter nested logit model to capture households’ unobserved preference heterogeneity and non-proportional substitution patterns in terms of activity party composition for multiparty activities. A wide variety of household demographics, activity attributes and residential neighborhood characteristics are examined in this paper. The magnitude of the impacts of the determinants are tested in this study by analyzing the elasticity of the variables, which suggests that household demographics and attributes of the multiparty activities have significant effects on the household-level activity party composition. Residential neighborhood characteristics, although somewhat less impactful, still play a meaningful role. This model will be implemented within the POLARIS transportation systems simulator to improve the activity generation modeling workflow, and the prediction accuracy of various activity-travel components.

activity party composition↗

Estimates of Lake Nitrogen, Phosphorus, and Chlorophyll‐ a Concentrations to Characterize Harmful Algal Bloom Risk Across the United States

Abstract Excess nutrient pollution contributes to the formation of harmful algal blooms (HABs) that compromise fisheries and recreation and that can directly endanger human and animal health via cyanotoxins. Efforts to quantify the occurrence, drivers, and severity of HABs across large areas is difficult due to the resource intensive nature of field monitoring of lake nutrient and chlorophyll‐aconcentrations. To better characterize how nutrients interact with other environmental factors to produce algal blooms in freshwater systems, we used spatially explicit and temporally matched climate, landscape, in‐lake characteristic, and nutrient inventory data sets to predict nutrients and chlorophyll‐aacross the conterminous US (CONUS). Using a nested modeling approach, three random forest (RF) models were trained to explain the spatiotemporal variation in total nitrogen (TN), total phosphorus (TP), and chlorophyll‐aconcentrations across US EPA's National Lakes Assessment (n = 2,062). Concentrations of TN and TP were the most important predictors and, with other variables, the RF model accounted for 68% of variation in chlorophyll‐a. We then used these RF models to extrapolate lake TN and TP predictions to lakes without nutrient observations and predict chlorophyll‐afor ∼112,000 lakes across the CONUS. Risk for high chlorophyll‐aconcentrations is highest in the agriculturally dominated Midwest, but other areas of risk emerge in nutrient pollution hot spots across the country. These catchment and lake‐specific results can help managers identify potential nutrient pollution and chlorophyll‐ahot spots that may fuel blooms, prioritize at‐risk lakes for additional monitoring, and optimize management to protect human health and other environmental end goals.

Environmental Sciences & Ecology↗