Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Conditional distribution”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Conditional distribution estimation of building characteristics with diffusion models for urban energy modeling

Understanding current energy consumption behavior in communities is critical for informing future energy use decisions and enabling efficient energy management. Urban energy models, which are used to simulate these energy use patterns, require large datasets with detailed building characteristics for accurate outcomes. However, such detailed characteristics at the individual building level are often unknown and costly to acquire, or unavailable. Through this work, we propose using a generative modeling approach to generate realistic building attributes to fill in the data gaps and finally provide complete characteristics as inputs to energy models. Our model learns complex, building-level patterns from training on a large-scale residential building stock model containing 2.2 million buildings. We employ a tabular diffusion-based framework that is designed to handle heterogeneous (discrete and continuous) features in tabular building data, such as occupancy, floor area, heating, cooling, and other equipment details. We develop a capability for conditional diffusion, enabling the imputation of missing building characteristics conditioned on known attributes. We conduct a comprehensive validation of our conditional diffusion model, firstly by comparing the generated conditional distributions against the underlying data distribution, and secondly, by performing a case study for a Baltimore residential region, showing the practical utility of our approach. Our work is one of the first to demonstrate the potential of generative modeling to accelerate building energy modeling workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

SIDDA: SInkhorn Dynamic Domain Adaptation for image classification with equivariant neural networks

Modern neural networks (NNs) often do not generalize well in the presence of a ‘covariate shift’; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels given the data remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more robust, domain-invariant features. Domain adaptation (DA) methods include a broad range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SInkhorn Dynamic Domain Adaptation (SIDDA), an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, real astronomical observations, and remote sensing data. These datasets exhibit covariate shifts due to noise, blurring, differences between telescopes, and variations in imaging wavelengths. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with symmetry-aware equivariant NNs (ENNs). We find that SIDDA consistently enhances the generalization capabilities of NNs, achieving up to a ${\approx}40\%$ improvement in classification accuracy on unlabeled target data, while also providing a more modest performance gain of $\lesssim 1\%$ on labeled source data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, if SIDDA achieves proper domain alignment, it also enhances model calibration on both source and target data, with the most significant gains in the unlabeled target domain—achieving over an order of magnitude improvement in the expected calibration error and Brier score. SIDDA’s versatility across various NN models and datasets, combined with its automated approach to domain alignment, has the potential to significantly advance multi-dataset studies by enabling the development of highly generalizable models.

79 ASTRONOMY AND ASTROPHYSICS

SIDDA: SInkhorn Dynamic Domain Adaptation

Modern neural networks (NNs) often do not generalize well in the presence of a "covariate shift"; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more domain-invariant features. Domain adaptation (DA) methods include a range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SIDDA, an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, and real astronomical observations. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with equivariant neural networks (ENNs). We find that SIDDA enhances the generalization capabilities of NNs, achieving up to a ≈40% improvement in classification accuracy on unlabeled target data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, we find that SIDDA enhances model calibration on both source and target data--achieving over an order of magnitude improvement in the ECE and Brier score. SIDDA's versatility, combined with its automated approach to domain alignment, has the potential to advance multi-dataset studies by enabling the development of highly generalizable models.

Pandya, Sneh [Northeastern U.]

The specification of distributed boundary conditions in numerical simulation of the ionosphere

The approach described makes it possible to use boundary conditions which do not have to be associated with any particular point of the solution interval. The approach was applied in a program which solves simultaneously the coupled differential equations describing the concentrations of the O(+) ions, their temperature, the electron temperature and the concentrations of O2(+), NO(+) and N2(+) ions. Several versions of the proposed scheme are described for various kinds of boundary conditions. Attention is given to a given columnar ionic content, a given incoming or outgoing flux, or a given height of the peak concentration.

Waldman, H.

Cluster infall for mass calibration in the stage-IV era

The outskirts of galaxy clusters present a promising avenue for constraining cluster masses in a way that is robust to the impact of baryonic physics. We assess the accuracy to which the cluster infall regions can be used for cluster mass calibration. Building on previous work, we parametrize the velocity distribution 𝑃⁡(𝑣r,𝑣tan|𝑟,𝑀) of dark matter halos on scales 𝑟 ≥ 5⁢ℎ −1 Mpc as the product of the marginalized distribution 𝑃⁡(𝑣 r |𝑟,𝑀) and the conditional distribution 𝑃⁡(𝑣 tan |𝑣 r ,𝑟,𝑀), calibrating the radial and mass dependence of these distributions in numerical simulations. We then project our model along the line of sight to obtain accurate predictions for the distributions of line-of-sight velocities at a given projected radius and cluster mass 𝑃⁡(𝑣 LOS |𝑅,𝑀), which we can observe with spectroscopic survey data. Furthermore, with our model, we forecast that spectra from the Dark Energy Spectroscopic Instrument can constrain cluster masses with subpercent-level precision, comparable to that of stage-IV weak lensing surveys.

79 ASTRONOMY AND ASTROPHYSICS

Elliptically-Contoured Tensor-variate Distributions with Application to Image Learning

Statistical analysis of tensor-valued data has largely used the tensor-variate normal (TVN) distribution that may be inadequate for data arising from distributions with heavier or lighter tails. We study a general family of elliptically contoured (EC) TV distributions and derive its characterizations, moments, marginal, and conditional distributions. We describe procedures for maximum likelihood estimation from data that are (1) uncorrelated draws from an EC distribution, (2) from a scale mixture of the TVN distribution, and (3) from an underlying but unknown EC distribution, for which we extend Tyler’s robust estimator. A detailed simulation study highlights the benefits of choosing an EC distribution over the TVN for heavier-tailed data. We develop TV classification rules using discriminant analysis and EC errors and show that they better predict cats and dogs from images in the Animal Faces-HQ dataset than the TVN-based rules. A novel tensor-on-tensor regression and TV analysis of variance (TANOVA) framework under EC errors is also demonstrated to better characterize gender, age, and ethnic origin than the usual TVN-based TANOVA in the celebrated labeled faces of the wild dataset.

97 MATHEMATICS AND COMPUTING

Carbon monoxide over the Amazon Basin during the wet season

Measurements of the distribution of carbon monoxide in the lower atmosphere (150-5000 m) over the Amazon region of Brazil during the wet season, taken in conjunction with the 1987 NASA Global Tropospheric Experiment/Amazon Boundary Layer Experiment (ABLE 2B), are analyzed. About 100 hr of airborne, in situ CO measurements were obtained using a tunable diode laser system, providing insights into factors influencing the basin-scale distribution of CO in the Amazonian troposphere during wet season conditions. Distribution of CO over the altitudes 0.15-4.5 km was influenced by such factors as surface emissions from biological sources and long-range transport of pollutants from Northern Hemisphere sources. It is noted that the disruption of mixed layer growth and decay processes has a particularly important influence on CO concentration in the daytime lower troposphere and that the correlation of CO with O3 was positive under conditions influenced by Northern Hemisphere air and negative under all other conditions observed.

Harriss, R. C.

Range reference atmosphere models

A description is given of the methods used to establish the statistical parameters and models for wind and various thermodynamic quantities at an altitude of 0-70 km for nine geographical locations. It is noted that wind is modeled as a vector quantity using the bivariate normal probability function. With the five parameters of the bivariate normal distribution, the distribution for wind speed is derived as a generalized Rayleigh distribution. In addition, the frequency of wind direction is derived, and the conditional distribution of wind speed given the wind direction is derived. It is pointed out that these and other wind models are consistent with the rigorous mathematical properties of the bivariate normal probability theory. The thermodynamic quantities are consistent with the hydrostatic equation and the equation of state for the mean values. With these methods, many statistical relationships can be derived.

Smith, O. E.

Calculations of heavy ion charge state distributions for nonequilibrium conditions

Numerical calculations of the charge state distributions of test ions in a hot plasma under nonequilibrium conditions are presented. The mean ionic charges of heavy ions for finite residence times in an instantaneously heated plasma and for a non-Maxwellian electron distribution function are derived. The results are compared with measurements of the charge states of solar energetic particles, and it is found that neither of the two simple cases considered can explain the observations.

Luhn, A.

NASA GPM GV Science Requirements

An important scientific objective of the NASA portion of the GPM Mission is to generate quantitatively-based error characterization information along with the rainrate retrievals emanating from the GPM constellation of satellites. These data must serve four main purposes: (1) they must be of sufficient quality, uniformity, and timeliness to govern the observation weighting schemes used in the data assimilation modules of numerical weather prediction models; (2) they must extend over that portion of the globe accessible by the GPM core satellite to which the NASA GV program is focused - (approx.65 degree inclination); (3) they must have sufficient specificity to enable detection of physically-formulated microphysical and meteorological weaknesses in the standard physical level 2 rainrate algorithms to be used in the GPM Precipitation Processing System (PPS), i.e., algorithms which will have evolved from the TRMM standard physical level 2 algorithms; and (4) they must support the use of physical error modeling as a primary validation tool and as the eventual replacement of the conventional GV approach of statistically intercomparing surface rainrates fiom ground and satellite measurements. This approach to ground validation research represents a paradigm shift vis-&-vis the program developed for the TRMM mission, which conducted ground validation largely as a statistical intercomparison process between raingauge-derived or radar-derived rainrates and the TRMM satellite rainrate retrievals -- long after the original satellite retrievals were archived. This approach has been able to quantify averaged rainrate differences between the satellite algorithms and the ground instruments, but has not been able to explain causes of algorithm failures or produce error information directly compatible with the cost functions of data assimilation schemes. These schemes require periodic and near-realtime bias uncertainty (i.e., global space-time distributed conditional accuracy of the retrieved rainrates) and local error covariance structure (i.e., global space-time distributed error correlation information for the local 4-dimensional space-time domain -- or in simpler terms, the matrix form of precision error). This can only be accomplished by establishing a network of high quality-heavily instrumented supersites selectively distributed at a few oceanic, continental, and coastal sites. Economics and pragmatics dictate that the network must be made up of a relatively small number of sites (6-8) created through international cooperation. This presentation will address some of the details of the methodology behind the error characterization approach, some proposed solutions for expanding site-developed error properties to regional scales, a data processing and communications concept that would enable rapid implementation of algorithm improvement by the algorithm developers, and the likely available options for developing the supersite network.

Smith, E.

A Simple Stochastic Model for Generating Broken Cloud Optical Depth and Top Height Fields

A simple and fast algorithm for generating two correlated stochastic twodimensional (2D) cloud fields is described. The algorithm is illustrated with two broken cumulus cloud fields: cloud optical depth and cloud top height retrieved from Moderate Resolution Imaging Spectrometer (MODIS). Only two 2D fields are required as an input. The algorithm output is statistical realizations of these two fields with approximately the same correlation and joint distribution functions as the original ones. The major assumption of the algorithm is statistical isotropy of the fields. In contrast to fractals and the Fourier filtering methods frequently used for stochastic cloud modeling, the proposed method is based on spectral models of homogeneous random fields. For keeping the same probability density function as the (first) original field, the method of inverse distribution function is used. When the spatial distribution of the first field has been generated, a realization of the correlated second field is simulated using a conditional distribution matrix. This paper is served as a theoretical justification to the publicly available software that has been recently released by the authors and can be freely downloaded from http://i3rc.gsfc.nasa.gov/Public codes clouds.htm. Though 2D rather than full 3D, stochastic realizations of two correlated cloud fields that mimic statistics of given fields have proved to be very useful to study 3D radiative transfer features of broken cumulus clouds for better understanding of shortwave radiation and interpretation of the remote sensing retrievals.

Prigarin, Sergei M.

A generative modeling approach to reconstructing 21 cm tomographic data

Abstract Analyses of the cosmic 21 cm signal are hampered by astrophysical foregrounds that are far stronger than the signal itself. These foregrounds, typically confined to a wedge-shaped region in Fourier space, often necessitate the removal of a vast majority of modes, thereby degrading the quality of the data anisotropically. To address this challenge, we introduce a novel deep generative model based on stochastic interpolants to reconstruct the 21 cm data lost to wedge filtering. Our method leverages the non-Gaussian nature of the 21 cm signal to effectively map wedge-filtered 3D lightcones to samples from the conditional distribution of wedge-recovered lightcones. We demonstrate how our method is able to restore spatial information effectively, considering both varying cosmological initial conditions and astrophysics. Furthermore, we discuss a number of future avenues where this approach could be applied in analyses of the 21 cm signal, potentially offering new opportunities to improve our understanding of the Universe during the epochs of cosmic dawn and reionization. Code, pre-trained models, and scripts for making plots in this paper can be found here .

Sabti, Nashwan (ORCID:000000027924546X)

Power conditioning equipment for a thermoelectric outer planet spacecraft, volume 1, book 1

Equipment was designed to receive power from a radioisotope thermoelectric generator source, condition, distribute, and control this power for the spacecraft loads. The TOPS mission, aimed at a representative tour of the outer planets, would operate for an estimated 12 year period. Unique design characteristics required for the power conditioning equipment results from the long mission time and the need for autonomous on-board operations due to large communications distances and the associated time delays of ground initiated actions. The salient features of the selected power subsystem configuration are: (1) The PCE regulates the power from the radioisotope thermoelectric generator power source at 30 vdc by means of a quad-redundant shunt regulator; (2) 30 vdc power is used by certain loads, but is more generally inverted and distributed as square-wave ac power; (3) a protected bus is used to assure that power is always available to the control computer subsystem to permit corrective action to be initiated in response to fault conditions; and (4) various levels of redundancy are employed to provide high subsystem reliability.

Andrews, R. E.

Continuing Development of the Nuclear Data Processing Code AMPX [Poster]

The ENDF/B-VIII.1 evaluation library has seen a great growth in the thermal neutron scattering sub-library. The SCALE code system has traditionally approached CE transport by assuming that the CE library on disk represented the fully expanded cumulative probability distributions, conditional on exiting angle and marginal on exiting energy. While this is a complete description of the data, it comes at the potential cost of large amounts of on-disk storage. This approach was strained by several TSL files in ENDF/B-VIII.1, such as graphite, which contained data for a large number of Bragg edges. In the fully expanded probability distributions, this was found to be a disproportionately large fraction of the SCALE CE library.

GNDS

Application of Statistical Methods of Rain Rate Estimation to Data From The TRMM Precipitation Radar

The TRMM Precipitation Radar is well suited to statistical methods in that the measurements over any given region are sparsely sampled in time. Moreover, the instantaneous rain rate estimates are often of limited accuracy at high rain rates because of attenuation effects and at light rain rates because of receiver sensitivity. For the estimation of the time-averaged rain characteristics over an area both errors are relevant. By enlarging the space-time region over which the data are collected, the sampling error can be reduced. However. the bias and distortion of the estimated rain distribution generally will remain if estimates at the high and low rain rates are not corrected. In this paper we use the TRMM PR data to investigate the behavior of 2 statistical methods the purpose of which is to estimate the rain rate over large space-time domains. Examination of large-scale rain characteristics provides a useful starting point. The high correlation between the mean and standard deviation of rain rate implies that the conditional distribution of this quantity can be approximated by a one-parameter distribution. This property is used to explore the behavior of the area-time-integral (ATI) methods where fractional area above a threshold is related to the mean rain rate. In the usual application of the ATI method a correlation is established between these quantities. However, if a particular form of the rain rate distribution is assumed and if the ratio of the mean to standard deviation is known, then not only the mean but the full distribution can be extracted from a measurement of fractional area above a threshold. The second method is an extension of this idea where the distribution is estimated from data over a range of rain rates chosen in an intermediate range where the effects of attenuation and poor sensitivity can be neglected. The advantage of estimating the distribution itself rather than the mean value is that it yields the fraction of rain contributed by the light and heavy rain rates. This is useful in estimating the fraction of rainfall contributed by the rain rates that go undetected by the radar. The results at high rain rates provide a cross-check on the usual attenuation correction methods that are applied at the highest resolution of the instrument.

Meneghini, R.