Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Normal Distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Investigating the ecological fallacy through sampling distributions constructed from finite populations

Correlation coefficients and linear regression values computed from group averages can differ from correlation coefficients and linear regression values computed using individual scores. This observation known as the ecological fallacy often assumes that all the individual scores are available from a population. In many situations, one must use a sample from the larger population. In such cases, the computed correlation coefficient and linear regression values will depend on the sample that is chosen and the underlying sampling distribution. The sampling distribution of correlation coefficients and linear regression values for group averages will be identical to the sampling distribution for individuals for normally distributed variables for random samples drawn from infinitely large continuous distributions. However, data that is acquired in practice is often acquired when sampling without replacement from a finite population. Our objective is to demonstrate through Monte Carlo simulations that the sampling distributions for correlation and linear regression will also be similar for individuals and group averages when sampling without replacement from normally distributed variables. These simulations suggest that when a random sample from a population is selected, the correlation coefficients and linear regression values computed from individual scores will not be more accurate in estimating the entire population values compared to samples when group averages are used as long as the sample size is the same.

97 MATHEMATICS AND COMPUTING

Variance-Reduced Accelerated First-Order Methods: Central Limit Theorems and Confidence Statements

In this paper, we consider a strongly convex stochastic optimization problem and propose three classes of variable sample-size stochastic first-order methods: (i) the standard stochastic gradient descent method, (ii) its accelerated variant, and (iii) the stochastic heavy-ball method. In each scheme, the exact gradients are approximated by averaging across an increasing batch size of sampled gradients. We prove that when the sample size increases at a geometric rate, the generated estimates converge in mean to the optimal solution at an analogous geometric rate for schemes (i)–(iii). Based on this result, we provide central limit statements, whereby it is shown that the rescaled estimation errors converge in distribution to a normal distribution with the associated covariance matrix dependent on the Hessian matrix, the covariance of the gradient noise, and the step length. If the sample size increases at a polynomial rate, we show that the estimation errors decay at a corresponding polynomial rate and establish the associated central limit theorems (CLTs). Under certain conditions, we discuss how both the algorithms and the associated limit theorems may be extended to constrained and nonsmooth regimes. As a result, we provide an avenue to construct confidence regions for the optimal solution based on the established CLTs and test the theoretical findings on a stochastic parameter estimation problem.

Lei, Jinlong

Data for "Quantifying the Propagation of Parametric Uncertainty on Flux Balance Analysis"

In the repository are example scripts that perform uncertainty injection and propagation to flux balance analysis with outputs for a small sample size (for demonstration purpose only). For proper analysis, user should download the scripts and run for a large sample size (e.g., 10,000 samples). If you use the scripts, please cite the following Metabolic Engineering article: “Quantifying the propagation of parametric uncertainty on flux balance analysis” (https://doi.org/10.1016/j.ymben.2021.10.012) There are two subdirectories: /uncFBA/uncBiom: injection of normally distributed noise to biomass precursor coeffcients and ATP maintenance (growth-associated ATP maintenance (GAM) and non-growth associated ATP maintenance (NGAM)) /uncFBA/uncRHS: departure from steady-state by adding noise drawn from normal distribution to the RHS terms of mass balance constraints

Metabolomics

Determining reference standard strength for neutron-irradiated reduced activation ferritic/martensitic steel F82H by Bayesian method

The deterministic approach widely adopted in the design of structural components relies on systematically defined design limits using empirically determined safety factors. However, this approach is not always appropriate because structures are subjected to a variety of loads in the practical environment, which may result in excessively conservative design limits. In recent years, a more rigorous probabilistic approach that incorporates material strength distributions has become an important solution. In the probabilistic approach, the probability density functions of material strength properties underpin the design criteria. Here, the objective of this study is to identify the density distribution functions that best describe tensile properties of irradiated F82H to define a reference strength for DEMO design. Due to the limited number of existing data, this study specifically employs a Bayesian prediction method based on Monte Carlo simulations to determine a material reference value with statistical reliability and to investigate its effectiveness. For example, the dependence of tensile properties of 300 °C irradiated materials on irradiation damage and the range predicted by 95% Bayesian estimation was evaluated. As a statistical model for the dose dependence of statistical parameters, the normal distribution exhibited a better fit for 0.2% proof strength and tensile strength, whereas the distribution of total elongation data gave comparable reference values for both the normal and Weibull distribution models. Both models gave comparable criteria for the distribution of total elongation data. The Weibull model also gave better results for uniform elongation. The function best describing the model was a logarithmic law for both 0.2% proof strength and tensile strength, while a power law for both total and uniform elongation, which allowed for more comprehensive data prediction of irradiation data with statistical accuracy for DEMO reactor design.

36 MATERIALS SCIENCE

Yet Another Discriminant Analysis (YADA): A Probabilistic Model for Machine Learning Applications

This paper presents a probabilistic model for various machine learning (ML) applications. While deep learning (DL) has produced state-of-the-art results in many domains, DL models are complex and over-parameterized, which leads to high uncertainty about what the model has learned, as well as its decision process. Further, DL models are not probabilistic, making reasoning about their output challenging. In contrast, the proposed model, referred to as Yet Another Discriminate Analysis(YADA), is less complex than other methods, is based on a mathematically rigorous foundation, and can be utilized for a wide variety of ML tasks including classification, explainability, and uncertainty quantification. YADA is thus competitive in most cases with many state-of-the-art DL models. Ideally, a probabilistic model would represent the full joint probability distribution of its features, but doing so is often computationally expensive and intractable. Hence, many probabilistic models assume that the features are either normally distributed, mutually independent, or both, which can severely limit their performance. YADA is an intermediate model that (1) captures the marginal distributions of each variable and the pairwise correlations between variables and (2) explicitly maps features to the space of multivariate Gaussian variables. Numerous mathematical properties of the YADA model can be derived, thereby improving the theoretic underpinnings of ML. Validation of the model can be statistically verified on new or held-out data using native properties of YADA. However, there are some engineering and practical challenges that we enumerate to make YADA more useful.

97 MATHEMATICS AND COMPUTING

Evaluating the impact of anatomical and physiological variability on human equivalent doses using PBPK models

Abstract Addressing human anatomical and physiological variability is a crucial component of human health risk assessment of chemicals. Experts have recommended probabilistic chemical risk assessment paradigms in which distributional adjustment factors are used to account for various sources of uncertainty and variability, including variability in the pharmacokinetic behavior of a given substance in different humans. In practice, convenient assumptions about the distribution forms of adjustment factors and human equivalent doses (HEDs) are often used. Parameters such as tissue volumes and blood flows are likewise often assumed to be lognormally or normally distributed without evaluating empirical data for consistency with these forms. In this work, we performed dosimetric extrapolations using physiologically based pharmacokinetic (PBPK) models for dichloromethane (DCM) and chloroform that incorporate uncertainty and variability to determine if the HEDs associated with such extrapolations are approximately lognormal and how they depend on the underlying distribution shapes chosen to represent model parameters. We accounted for uncertainty and variability in PBPK model parameters by randomly drawing their values from a variety of distribution types. We then performed reverse dosimetry to calculate HEDs based on animal points of departure for each set of sampled parameters. Corresponding samples of HEDs were tested to determine the impact of input parameter distributions on their central tendencies, extreme percentiles, and degree of conformance to lognormality. This work demonstrates that the measurable attributes of human variability should be considered more carefully and that generalized assumptions about parameter distribution shapes may lead to inaccurate estimates of extreme percentiles of HEDs.

Toxicology

Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO

Abstract Streambed grain sizes control river hydro‐biogeochemical (HBGC) processes and functions. However, measuring their quantities, distributions, and uncertainties is challenging due to the diversity and heterogeneity of natural streams. This work presents a photo‐driven, artificial intelligence (AI)‐enabled, and theory‐based workflow for extracting the quantities, distributions, and uncertainties of streambed grain sizes from photos. Specifically, we first trained You Only Look Once, an object detection AI, using 11,977 grain labels from 36 photos collected from nine different stream environments. We demonstrated its accuracy with a coefficient of determination of 0.98, a Nash–Sutcliffe efficiency of 0.98, and a mean absolute relative error of 6.65% in predicting the median grain size of 20 ground‐truth photos representing nine typical stream environments. The AI is then used to extract the grain size distributions and determine their characteristic grain sizes, including the 10th, 50th, 60th, and 84th percentiles, for 1,999 photos taken at 66 sites within a watershed in the Northwest US. The results indicate that the 10th, median, 60th, and 84th percentiles of the grain sizes follow log‐normal distributions, with most likely values of 2.49, 6.62, 7.68, and 10.78 cm, respectively. The average uncertainties associated with these values are 9.70%, 7.33%, 9.27%, and 11.11%, respectively. These data allow for the computation of the quantities, distributions, and uncertainties of streambed HBGC parameters, including Manning's coefficient, Darcy‐Weisbach friction factor, top layer interstitial velocity magnitude, and nitrate uptake velocity. Additionally, major sources of uncertainty in grain sizes and their impact on HBGC parameters are examined.

58 GEOSCIENCES

Generative learning of densities on manifolds

A generative modeling framework is proposed that combines diffusion models and manifold learning to efficiently sample data densities on manifolds. The approach utilizes Diffusion Maps to uncover possible low-dimensional underlying (latent) spaces in the high-dimensional data (ambient) space. Two approaches for sampling from the latent data density are described. The first is a score-based diffusion model, which is trained to map a standard normal distribution to the latent data distribution using a neural network. The second one involves solving an Itô stochastic differential equation in the latent space. Additional realizations of the data are generated by lifting the samples back to the ambient space using Double Diffusion Maps , a recently introduced technique typically employed in studying dynamical system reduction; here the focus lies in sampling densities rather than system dynamics. The proposed approaches enable sampling high dimensional data densities restricted to low-dimensional, a priori unknown manifolds. The efficacy of the proposed framework is demonstrated through a benchmark problem and a material with multiscale structure.

Double diffusion maps

Consistent and reproducible computation of the glass transition temperature from molecular dynamics simulations

In many fields, from semiconductors for opto-electronic applications to ionic liquids (ILs) for separations, the glass transition temperature (Tg) of a material is a useful gauge for its potential use in practical settings. As a result, there is a great deal of interest in predicting Tg using molecular simulations. However, the uncertainty and variation in the trend shift method, a common approach in simulations to predict Tg, can be high. This is due to the need for human intervention in defining a fitting range for linear fits of density with temperature assumed for the liquid and glass phases across the simulated cooling. The definition of such fitting ranges then defines the estimate for the Tg as the intersection of linear fits. We eliminate this need for human intervention by leveraging the Shapiro–Wilk normality test and proposing an algorithm to define the fitting ranges and, consequently, Tg. Through this integration, we incorporate into our automated methodology that residuals must be normally distributed around zero for any fit, a requirement that must be met for any regression problem. Consequently, fitting ranges for realizing linear fits for each phase are statistically defined rather than visually inferred, obtaining an estimate for Tg without any human intervention. The method is also capable of finding multiple linear regimes across density vs temperature curves. We compare the predictions of our proposed method across multiple IL and semiconductor molecular dynamics simulation results from the literature and compare other proposed methods for automatically detecting Tg from density–temperature data. We believe that our proposed method would allow for more consistent predictions of Tg. We make this methodology available and open source through GitHub.

Chemistry

Welch Method and Bootstrapping Applied to Subcritical Gamma Noise

We measured the prompt neutron decay constant 𝛼 of the CROCUS zero-power reactor at the Swiss Federal Institute of Technology Lausanne using cross-power spectral density (CPSD) analysis of gamma-gamma correlations from two trans-stilbene organic scintillators positioned near the reactor core. We measured critical and subcritical states, with water levels ranging from 960 mm (critical) to 800 mm (𝜌=−1.4 $ subcritical). Our analysis used the Welch method, dividing signal segments for fast Fourier transform (FFT) frequency analysis and applying bootstrapping uncertainty quantification that uses Welch-defined segments. Results demonstrated a clear increase in the measured 𝛼 as reactor reactivity decreased, distinguishing critical from subcritical conditions. At the 960-mm critical level, 𝛼 was estimated at 155.9 ± 0.7 s −1 , and for the 800-mm subcritical level, 𝛼 increased significantly to 367.3 ± 6.9 s –1 . A linear regression of subcritical states yielded a critical estimate of 154.0 ± 3.1 s –1 , aligning with the static 𝛼 estimate at critical. The bootstrapping method produced normally distributed 𝛼 estimates, confirming data consistency. The gamma CPSD 𝛼 estimates clearly distinguish reactor states and improve monitoring of zero-power reactors. The future deployment of modular and microreactors as potential candidates for noise analysis is demonstrated in CROCUS, particularly zero-power mock-ups of new designs. The improvement of noise analysis in the subcritical domain from this work will support experimental data for reactor deployment and procedure.

CROCUS

Safe Reinforcement Learning-Based Transient Stability Control for Islanded Microgrids With Topology Reconfiguration

This paper proposes a safe reinforcement learning (RL)-based transient stability emergency control (TSEC) method for islanded microgrids. RL requires extensive interaction with the environment to learn control strategies, hence, a data-driven approach is used as a substitute for time-consuming time-domain simulation calculations. Deep sigma point processes (DSPP), which is a Gaussian process model, is utilized to predict the normal distribution of transient stability of microgrids and to construct a transient stability chance constraint. Reward-constrained policy optimization (RCPO) can simultaneously achieve objective prediction, policy learning, and constraint cost coefficient update across multiple timescales. RCPO interacts with the DSPP-based microgrid environment through a multi-process parallel manner, greatly increasing the training speed. Case studies on a real islanded microgrid demonstrate that the proposed method can efficiently and quickly obtain the optimal emergency control strategy while adhering to all hard constraints.

14 SOLAR ENERGY

Double Bootstrapping

This code performs a simulation experiment that involves (i) drawing random numbers from the normal distribution, (ii) resampling elements from arrays with replacement, and (iii) computing various quantities like mean, standard deviation, etc. Further information is available in section 4 of FERMILAB-FN-1273-ETD [https://inspirehep.net/literature/2925453].

Shyamsundar, Prasanth [Fermi National Accelerator

Toward the validation of crowdsourced experiments for lightness perception

Crowdsource platforms have been used to study a range of perceptual stimuli such as the graphical perception of scatterplots and various aspects of human color perception. Given the lack of control over a crowdsourced participant’s experimental setup, there are valid concerns on the use of crowdsourcing for color studies as the perception of the stimuli is highly dependent on the stimulus presentation. Here, we propose that the error due to a crowdsourced experimental design can be effectively averaged out because the crowdsourced experiment can be accommodated by the Thurstonian model as the convolution of two normal distributions, one that is perceptual in nature and one that captures the error due to variability in stimulus presentation. Based on this, we provide a mathematical estimate for the sample size needed to produce a crowdsourced experiment with the same power as the corresponding in-person study. We tested this claim by replicating a large-scale, crowdsourced study of human lightness perception with a diverse sample with a highly controlled, in-person study with a sample taken from psychology undergraduates. Our claim was supported by the replication of the results from the latter. These findings suggest that, with sufficient sample size, color vision studies may be completed online, giving access to a larger and more representative sample. With this framework at hand, experimentalists have the validation that choosing either many online participants or few in person participants will not sacrifice the impact of their results.

97 MATHEMATICS AND COMPUTING

Building ControlScore: General Service Administration Office Building Deployment

Improvements to building control systems can lead to energy savings and increased occupant comfort. In an optimized system, process variables such as air temperature will closely follow their desired setpoints and avoid excess energy use. Typically, experts must manually inspect individual control loops to identify poor performance and opportunities for improvement. However, this approach is difficult in modern buildings that have a prohibitively large number of controllers. To address this issue, Pacific Northwest National Laboratory (PNNL) created the ControlScore concept, which takes operating data from the many controllers within a building and generates standardized scores for each loop on a scale of 0 to 10 (a score of 0 indicates poor control, a score of 10 indicates good control). PNNL applied the Building ControlScore application to all available data from a General Services Administration office building within the period of January 1, 2023, to March 9, 2023. The building scored a 4.7 overall, with all 74 of the building’s loops fitting a roughly normal distribution centered around 5. These results indicate that the analyzed systems have below-average performance with room for improvement, especially in the poorly scored systems. Airflow loops tended to have much lower scores than zone temperature loops. The lowest and highest performing systems in the building section were identified, as were all loops with a score less than 1. While the ControlScore identifies loops and systems that aren’t meeting their designated setpoints, it does not indicate the cause of those issues. For example, consider a supply air terminal unit’s airflow loop that received a low score due to it delivering less air than specified by the setpoint. The lower-than-desired airflow could be due to equipment limitations (e.g., the terminal unit or duct serving is too small to accommodate that airflow), malfunctioning equipment (e.g., a stuck damper or bad sensor), or something else entirely. The ControlScore does not diagnose problems it simply identifies the symptoms that can be explored and addressed by building operators.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Working with Bezier Curves as bases for Functional Expansion Tallies

Functional expansion tallies (FETs) are a powerful tool for getting more information per history from Monte Carlo simulations, but in the past they have been constrained to orthogonal bases. Bezier curves, from computer aided design (CAD), could be well suited for FET due to their ability to assume many arbitrary shapes, but are non-orthogonal. Recent development has made non-orthongal FET possible. The convergence of B ´ezier curve FETs in both polynomial order and number of samples is explored. These bases are well suited for representing normal distributions, and opens the door to possible other CAD derived FET bases.

97 MATHEMATICS AND COMPUTING

Working with Bézier Curves as bases for Functional Expansion Tallies

Functional expansion tallies (FETs) are powerful tools for getting more information per history from Monte Carlo simulations, but in the past they have been constrained to orthogonal bases. Bézier curves are used widely in computer aided design (CAD) geometry kernels and could be well-suited for FETs due to their ability to assume many arbitrary shapes, but they use nonorthogonal bases. Recent developments in 2021 have made nonorthogonal FETs possible. The convergence of Bézier curve FETs in both polynomial order and number of samples is explored in this work. It is shown that these bases are well-suited for representing normal distributions and this opens the door to the possibility of other CAD-derived FET bases.

97 - MATHEMATICS AND COMPUTING

The Multi-Phase Circumgalactic Medium of DESI Emission-Line Galaxies at z~1.5

We study the multi-phase circumgalactic medium (CGM) of emission line galaxies (ELGs) at $z\sim1.5$, traced by MgII$\lambda2796$, $\lambda2803$ and CIV$\lambda1548$, $\lambda1550$ absorption lines, using approximately 7,000 ELG-quasar pairs from the Dark Energy Spectroscopic Instrument. Our results show that both the mean rest equivalent width ($W_{0}$) profiles and covering fractions of MgII and CIV increase with ELG stellar mass at similar impact parameters, but show similar distributions when normalized by the virial radius. Moreover, warm CIV gas has a more extended distribution than cool MgII gas. The dispersion of MgII and CIV gas velocity offsets relative to the galaxy redshifts rises from $\sim100 \, \rm km \, s^{-1}$ within halos to $\sim 200 \, \rm km \, s^{-1}$ beyond. We explore the relationships between MgII and CIV $W_{0}$ and show that the two are not tightly coupled: at a fixed absorption strength of one species, the other varies by several-fold, indicating distinct kinematics between the gas phases traced by each. We measure the line ratios, FeII/MgII and CIV/MgII, of strong MgII absorbers and find that at $<0.2$ virial radius, the FeII/MgII ratio is elevated, while the CIV/MgII ratio is suppressed compared with the measurements on larger scales, both with $\sim4-5\, σ$ significance. We argue that multiphase gas that is not co-spatial is required to explain the observational results. Finally, by combining with measurements from the literature, we investigate the redshift evolution of CGM properties and estimate the neutral hydrogen, metal, and dust masses in the CGM of DESI ELGs -- found to be comparable to those in the ISM.

Lan, Ting-Wen [Taiwan, Natl. Taiwan U.; Academia S

A Pseudoreversible Normalizing Flow for Stochastic Dynamical Systems with Various Initial Distributions

Here, we present a pseudoreversible normalizing flow method for efficiently generating samples of the state of a stochastic differential equation (SDE) with various initial distributions. The primary objective is to construct an accurate and efficient sampler that can be used as a surrogate model for computationally expensive numerical integration of SDEs, such as those employed in particle simulation. After training, the normalizing flow model can directly generate samples of the SDE’s final state without simulating trajectories. The existing normalizing flow model for SDEs depends on the initial distribution, meaning the model needs to be retrained when the initial distribution changes. The main novelty of our normalizing flow model is that it can learn the conditional distribution of the state, i.e., the distribution of the final state conditional on any initial state, such that the model only needs to be trained once and the trained model can be used to handle various initial distributions. This feature can provide a significant computational saving in studies of how the final state varies with the initial distribution. Additionally, we propose to use a pseudoreversible network architecture to define the normalizing flow model, which has sufficient expressive power and training efficiency for a variety of SDEs in science and engineering, e.g., in particle physics. We provide a rigorous convergence analysis of the pseudoreversible normalizing flow model to the target probability density function in the Kullback–Leibler divergence metric. Numerical experiments are provided to demonstrate the effectiveness of the proposed normalizing flow model.

97 MATHEMATICS AND COMPUTING