Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Data-Efficient Generation of Protein Conformational Ensembles with Backbone-to-Side-Chain Transformers

Excitement at the prospect of using data-driven generative models to sample configurational ensembles of biomolecular systems stems from the extraordinary success of these models on a diverse set of high-dimensional sampling tasks. Unlike image generation or even the closely related problem of protein structure prediction, there are currently no data sources with sufficient breadth to parametrize generative models for conformational ensembles. To enable discovery, a fundamentally different approach to building generative models is required: models should be able to propose rare, albeit physical, conformations that may not arise in even the largest data sets. Here, in this work, we introduce a modular strategy to generate conformations based on “backmapping” from a fixed protein backbone that (1) maintains conformational diversity of the side chains and (2) couples the side-chain fluctuations using global information about the protein conformation. Our model combines simple statistical models of side-chain conformations based on rotamer libraries with the now ubiquitous transformer architecture to sample with atomistic accuracy. Together, these ingredients provide a strategy for rapid data acquisition and hence a crucial ingredient for scalable physical simulation with generative neural networks.

36 MATERIALS SCIENCE↗

Misclassification in Workers’ Telecommuting Frequency Choices Using a Generalized Extreme Value Model

Telecommuting frequency is a response variable collected in travel surveys and is, therefore, prone to errors leading to mismeasurements or misclassification. Misclassification of explanatory variables is a common risk when using statistical modeling techniques. We define “misclassification” as a response reported or recorded in the wrong category; for example, a variable is recorded as a 1 when it should be 0. Here, in this context, this study aims to develop a statistical model to analyze telecommuting data which accounts for potential misclassification errors by building on existing literature in econometrics. The empirical analysis was undertaken using the 2017 National Household Travel Survey (NHTS) and the general extreme value (GEV) models available in the literature. Specifically, the frequency of telecommuting days was analyzed using the negative binomial (NB) model recast as the multinomial logit (MNL) model. By nature—and consistent with other studies—NHTS data are prone to errors that can be classified as intentional or unintentional misinformation provided by the person being interviewed. Ignoring these errors while modeling telecommuting frequencies using standard discrete count models can result in biased parameter estimates. The misclassification parameter was calculated for both over-reporting and under-reporting scenarios. The misclassification errors can be as high as 14% over-reported and 10% under-reported, particularly for the neighboring values. Statistical fit comparison between the models shows that models that ignore misclassification have worse data fit and biased parameter estimates with significant policy implications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Sustainability Trait Modeling of Field-Grown Switchgrass ( Panicum virgatum ) Using UAV-Based Imagery

Unmanned aerial vehicles (UAVs) provide an intermediate scale of spatial and spectral data collection that yields increased accuracy and consistency in data collection for morphological and physiological traits than satellites and expanded flexibility and high-throughput compared to ground-based data collection. In this study, we used UAV-based remote sensing for automated phenotyping of field-grown switchgrass (Panicum virgatum), a leading bioenergy feedstock. Using vegetation indices calculated from a UAV-based multispectral camera, statistical models were developed for rust disease caused by Puccinia novopanici, leaf chlorophyll, nitrogen, and lignin contents. For the first time, UAV remote sensing technology was used to explore the potentials for multiple traits associated with sustainable production of switchgrass, and one statistical model was developed for each individual trait based on the statistical correlation between vegetation indices and the corresponding trait. Also, for the first time, lignin content was estimated in switchgrass shoots via UAV-based multispectral image analysis and statistical analysis. The UAV-based models were verified by ground-truthing via correlation analysis between the traits measured manually on the ground-based with UAV-based data. The normalized difference red edge (NDRE) vegetation index outperformed the normalized difference vegetation index (NDVI) for rust disease and nitrogen content, while NDVI performed better than NDRE for chlorophyll and lignin content. Overall, linear models were sufficient for rust disease and chlorophyll analysis, but for nitrogen and lignin contents, nonlinear models achieved better results. As the first comprehensive study to model switchgrass sustainability traits from UAV-based remote sensing, these results suggest that this methodology can be utilized for switchgrass high-throughput phenotyping in the field.

60 APPLIED LIFE SCIENCES↗

The Data Synergy Effects of Time-Series Deep Learning Models in Hydrology

When fitting statistical models to variables in geoscientific disciplines such as hydrology, it is a customary practice to stratify a large domain into multiple regions (or regimes) and study each region separately. Traditional wisdom suggests that models built for each region separately will have higher performance because of homogeneity within each region. However, each stratified model has access to fewer and less diverse data points. Here, through two hydrologic examples (soil moisture and streamflow), we show that conventional wisdom may no longer hold in the era of big data and deep learning (DL). We systematically examined an effect we call data synergy, where the results of the DL models improved when data were pooled together from characteristically different regions. The performance of the DL models benefited from modest diversity in the training data compared to a homogeneous training set, even with similar data quantity. Moreover, allowing heterogeneous training data makes eligible much larger training datasets, which is an inherent advantage of DL. A large, diverse data set is advantageous in terms of representing extreme events and future scenarios, which has strong implications for climate change impact assessment. The results here suggest the research community should place greater emphasis on data sharing.

54 ENVIRONMENTAL SCIENCES↗

Application of Linear Additive Conditions for Near-Infrared Diffuse Reflectance Absorption Spectroscopy

Determining the homogeneity of material mixing in real time during product processing is critical for quality control. According to the Kubelka–Munk (K-M) function of diffuse reflectance absorption spectrum, absorbance (A) is approximately linear with the content of the components when the sample scattering coefficient (S) is in a certain range. The S is determined by the particle size of powder samples. Therefore, this study determined particle size ranges that satisfy linear additivity in near-infrared diffuse reflectance spectroscopy (NIRDRS). Thus, the proposed NIRDRS analysis technique can be used to determine the homogeneity of material mixes or analyze the percentages of the components in the mixture. In this study, vitamin B3 and vitamin C were used for preparing mixed samples with varying percentages. The experimental results revealed that linear additivity is satisfied when the powder particle size is in the range of less than 280, 280–450, and 450–900 μm. When the confidence level is 0.01, the actual mixed spectra are not significantly different from the “simulated mixed spectra” constructed by linear addition, with their relative deviations less than 1.08%. The absolute errors of the actual and analytic percentages were within 2.98% for each component in the mixtures. The above conclusions also hold for sorghum, which has a complex material composition. Statistical models cannot analyze the percentages of components in the mixture. In contrast, linear addition and direct calibration approach avoids the use of a large number of samples for statistical modeling and analyze the percentages of mixed samples. Meanwhile, it can be used to discriminate and analyze the material mixing uniformity by building a mechanistic model.

Feng, Zhiyue↗

Grain size estimation in fluvial gravel bars using uncrewed aerial vehicles: A comparison between methods based on imagery and topography

Abstract Grain size assessments are necessary for understanding the various geomorphological, hydrological and ecological processes that occur within rivers. Recent research has shown that the application of Structure‐from‐Motion (SfM) photogrammetry to imagery from uncrewed aerial vehicles (UAVs) shows promise for rapidly characterising grain sizes along rivers in comparison to traditional field‐based methods. Here, we evaluated the applicability of different methods for estimating grain sizes in gravel bars along a study reach in the Olentangy River in Columbus, Ohio. We collected imagery of these gravel bars with a UAV and processed those images with SfM photogrammetry software to produce three‐dimensional point clouds and orthomosaics. Our evaluation compared statistical models calibrated on topographic roughness, which was computed from the point clouds, and to those based on image texture, which was computed from the orthomosaics. Our results showed that statistical models calibrated on image texture were more accurate than those based on topographic roughness. This might be because of site‐specific patterns of grain size, shape and imbrication. Such patterns would have complicated the detection of topographic signatures associated with individual grains. Our work illustrates that UAV‐SfM approaches show potential to be used as an accessible method for characterising surface grain sizes along rivers at higher spatial and temporal resolutions than those provided by traditional methods.

Wong, Tyler↗

How Bayesian methods can improve R -matrix analyses of data: The example of the d t reaction

The 3 H(d, n) 4 He reaction is of significant interest in nuclear astrophysics and nuclear applications. It is an important, early step in big-bang nucleosynthesis and a key process in nuclear fusion reactors. We use one- and two-level R-matrix approximations to analyze data on the cross section for this reaction at center-of-mass energies below 215 keV. We critically examine the data sets using a Bayesian statistical model that allows for both common-mode and additional point-to-point un- certainties. We use Markov Chain Monte Carlo sampling to evaluate this R-matrix-plus-statistical model and find two-level R-matrix results that are stable with respect to variations in the channel radii. The S factor at 40 keV evaluates to 25.36(19) MeV b (68% credibility interval). We discuss our Bayesian analysis in detail and provide guidance for future applications of Bayesian methods to R-matrix analyses. We also discuss possible paths to further reduction of the S-factor uncertainty.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Sustainability Trait Modeling of Field-Grown Switchgrass (Panicum virgatum) Using UAV-Based Imagery

Unmanned aerial vehicles (UAVs) provide an intermediate scale of spatial and spectral data collection that yields increased accuracy and consistency in data collection for morphological and physiological traits than satellites and expanded flexibility and high-throughput compared to ground-based data collection. In this project, we used UAV-based multispectral imagery collected from MicaSense RedEdge-M on a DJI Matrice 600 Pro for automated phenotyping of field-grown switchgrass (Panicum virgatum), a leading bioenergy feedstock. The raw images were processed with Pix4D Mapper to create the reflectance data, and vegetation indices were calculated from a UAV-based multispectral camera. Statistical models were developed for rust disease caused by Puccinia novopanici, leaf chlorophyll, nitrogen, and lignin contents. For the first time, UAV remote sensing technology was used to explore the potential for multiple traits associated with sustainable production of switchgrass. One statistical model was developed for each individual trait based on the statistical correlation between vegetation indices and the corresponding trait.

CBI↗

First Measurement of 87 Rb( α , xn ) Cross Sections at Weak r -process Energies in Supernova ν -driven Ejecta to Investigate Elemental Abundances in Low-metallicity Stars

Observed abundances of Z ∼ 40 elements in metal-poor stars vary from star to star, indicating that the rapid and slow neutron capture processes may not contribute alone to the synthesis of elements beyond iron. The weak r-process was proposed to produce Z ∼ 40 elements in a subset of old stars. Thought to occur in the ν-driven ejecta of a core-collapse supernova, ( α, xn ) reactions would drive the nuclear flow toward heavier masses at T = 2−5 GK. However, current comparisons between modeled and observed yields do not bring satisfactory insights into the stellar environment, mainly due to the uncertainties of the nuclear physics inputs where the dispersion in a given reaction rate often exceeds 1 order of magnitude. Involved rates are calculated with the statistical model where the choice of an α -optical-model potential ( α OMP) leads to such a poor precision. The first experiment on 87 Rb( α, xn ) reactions at weak r -process energies is reported here. Total inclusive cross sections were assessed at E c.m. = 8.1−13 MeV (3.7−7.6 GK) with the active target MUlti-Sampling Ionization Chamber. With an N = 50 seed nucleus, the measured values agree with statistical model estimates using the α OMP Atomki-V2. A reevaluated reaction rate was incorporated into new nucleosynthesis calculations, focusing on ν-driven ejecta conditions known to be sensitive to this specific rate. These conditions were found to fail to reproduce the lighter heavy element abundances in metal-poor stars.

79 ASTRONOMY AND ASTROPHYSICS↗

Model-independent determination of the dipole response of 66 Zn using quasimonoenergetic and linearly polarized photon beams

Background: Photon strength functions are an important ingredient in calculations relevant for the nucleosynthesis of heavy elements. The relation to the photoabsorption cross section allows to experimentally constrain photon strength functions by investigating the photoresponse of atomic nuclei. Purpose: Here, we determine the photoresponse of 66 Zn in the energy region of 5.6 MeV to 9.9 MeV and analyze the contribution of the 'elastic' decay channel back to the ground state. In addition, for the elastic channel electric and magnetic dipole transitions were separated. Methods: Nuclear resonance fluorescence experiments were performed using a linearly polarized quasi-monoenergetic photon beam at the High Intensity γ-ray Source. Photon beam energies from 5.6 to 9.9 MeV with an energy spread of about 3% were selected in steps of 200–300 keV. Two high purity germanium detectors were used for the subsequent γ-ray spectroscopy. Results: Full photoabsorption cross sections are extracted from the data making use of the monoenergetic character of the photon beam. For the ground-state decay channel, the average contribution of electric and magnetic dipole strengths is disentangled. The average branching ratio back to the ground state is determined as well. Conclusions: The new results indicate lower cross sections when compared to the values extracted from a former experiment using bremsstrahlung on 66 Zn. In the latter, the average branching ratio to the ground state is estimated from statistical-model calculations in order to analyze the data. Corresponding estimates from statistical-model calculations underestimate this branching ratio compared to the values extracted from the present analysis, which would partly explain the high cross sections determined from the bremsstrahlung data.

59 ≤ A ≤ 89↗

Alternating Conditional Expectations: Introducing a Non‐Parametric Statistical Method to Interpret Long‐Term Greenhouse Gas Flux Measurements Over Semi‐Arid and Wetland Ecosystems

Abstract We explore the potential of using a non‐parametric statistical method called Alternating Conditional Expectations, ACE, to quantify functional relationships in biogeosciences. Here, ACE is used to quantify the non‐linear and multi‐faceted responses of greenhouse gas fluxes to a set of biophysical forcings, when the shapes of those response surfaces are unknown. We evaluated the statistical method over two contrasting ecosystems and two contrasting time steps. One case involved quantifying the biophysical controls of water vapor and carbon dioxide (CO 2 ) fluxes over a semi‐arid oak savanna using daily integrated fluxes. The other case evaluated the responses of CO 2 and methane (CH 4 ) flux measurements to a set of biophysical forcings at a restored tidal wetland using thirty‐minute averages. The statistical model, based on 4 independent variables, explained up over 90% of the variation in daily integrated flux densities of water vapor and net carbon dioxide exchange at the savanna site. This fit was defined by distinct non‐linear responses to such drivers as gross primary production, photosynthetically active radiation, air temperature, vapor pressure deficit and soil moisture. At the tidal wetland site, we evaluated net carbon dioxide and methane fluxes with short‐term measurements to capture the influence of rising and falling tides and seasonality in biological activity. The statistical model defined the shape of the forcing of fluxes due to the roles of carbon exudates, water table depth, oxygen level in the water column, temperature and vegetation status. The statistical fits of the greenhouse gas fluxes were less precise than the savanna case. The fetch varies on a run‐to‐run basis as it is comprised of a heterogeneous mosaic of open water and vegetation. Furthermore, it is difficult to monitor the environmental conditions of the archaea and bacteria in the sediments that produce methane and carbon dioxide.

Environmental Sciences & Ecology↗

Neural networks for parameter estimation in intractable models

The goal is to use deep learning models to estimate parameters in statistical models when standard likelihood estimation methods are computationally infeasible. For instance, inference for max-stable processes is exceptionally challenging even with small datasets, but simulation is straightforward. Data from model simulations are used to train deep neural networks and learn statistical parameters from max-stable models. The proposed neural network-based method provides a competitive alternative to current approaches, as demonstrated by considerable accuracy and computational time improvements. Finally, it serves as a proof of concept for deep learning in statistical parameter estimation and can be extended to other estimation problems.

97 MATHEMATICS AND COMPUTING↗

Data assimilation for combustion ignition delay time simulation using schlieren image velocimetry

This study sought to improve the accuracy of simulating spray penetration and combustion ignition delay by means of data assimilation (DA). The simulations were conducted using the Reynolds-averaged Navier–Stokes (RANS) equations and assimilating the schlieren image data. In DA, an ensemble square root filter (EnSRF) was used to build the statistical model, making the simulation results more accurate without any change in the governing equations. Recognizing that the spray-cone injection angle has a large effect on penetration, we created ensemble members with different injection angles. And we applied the two-component velocity distribution calculated via SIV and updated both velocity and temperature by using a DA statistical model derived from RANS ensemble simulations. The ignition delay time is generally known to vary even under the same experimental conditions because it is influenced by many factors. In this study, we attempted the transient DA-assisted RANS simulation to predict the ignition delay time even when the temporal resolution and accuracy of the observation data ware insufficient. Our trials offer an example of how a combination of techniques can be effectively used to assimilate experimental data obtained under restricted conditions.

Combustion simulation↗

Predicting September Arctic Sea Ice: A Multimodel Seasonal Skill Comparison

This study quantifies the state of the art in the rapidly growing field of seasonal Arctic sea ice prediction. A novel multimodel dataset of retrospective seasonal predictions of September Arctic sea ice is created and analyzed, consisting of community contributions from 17 statistical models and 17 dynamical models. Prediction skill is compared over the period 2001–20 for predictions of pan-Arctic sea ice extent (SIE), regional SIE, and local sea ice concentration (SIC) initialized on 1 June, 1 July, 1 August, and 1 September. This diverse set of statistical and dynamical models can individually predict linearly detrended pan-Arctic SIE anomalies with skill, and a multimodel median prediction has correlation coefficients of 0.79, 0.86, 0.92, and 0.99 at these respective initialization times. Regional SIE predictions have similar skill to pan-Arctic predictions in the Alaskan and Siberian regions, whereas regional skill is lower in the Canadian, Atlantic, and central Arctic sectors. The skill of dynamical and statistical models is generally comparable for pan-Arctic SIE, whereas dynamical models outperform their statistical counterparts for regional and local predictions. The prediction systems are found to provide the most value added relative to basic reference forecasts in the extreme SIE years of 1996, 2007, and 2012. SIE prediction errors do not show clear trends over time, suggesting that there has been minimal change in inherent sea ice predictability over the satellite era. Overall, this study demonstrates that there are bright prospects for skillful operational predictions of September sea ice at least 3 months in advance.

54 ENVIRONMENTAL SCIENCES↗

Evaluation of Light-Element Reactions in the Resolved Resonance Region

Light-element reactions at low energies in the resolved resonance region are important for a range of applications in basic and applied sciences including nuclear reactors, nonproliferation, cultural heritage, forensics and environmental control, rare event investigations and nuclear astrophysics. In this paper, we report on an effort to evaluate charged-particle cross sections in the resolved resonance region and produce evaluated nuclear data files for further processing and inclusion in evaluated data libraries. We discuss the open issues in R-matrix calculations as we extend to higher energies, such as dealing with the rapidly growing number of open channels and merging with the regime of smooth cross sections described by the statistical model, and present attempts to address these issues in neutron-induced reactions relevant to nuclear reactor applications.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An integrated statistical-thermodynamic model for fission gas release and swelling in nuclear fuels

Here, we propose a new model for burst fission gas release induced by microcracking in ceramic nuclear fuels such as uranium dioxide. The model stipulates that the densities of defects in the fuel material, such as microcracks and fission gas bubbles on grain boundaries, evolve in accordance with the second law of thermodynamics. Central to the model is the notion of an effective temperature, conjugate to the configurational entropy of the fuel material, and directly linked to the burnup. The model predicts that microcracking, driven by the internal stress state of the fuel material, reduces the bubble storage capacity of grain boundaries, and accounts for burst fission gas release during rapid temperature transients that simulate power transients, reactor startup, and loss-of-coolant accident conditions.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Projecting Future Energy Production from Operating Wind Farms in North America. Part III: Variability

Abstract Daily expected wind power production from operating wind farms across North America are used to evaluate capacity factors (CF) computed using simulation output from the Weather Research and Forecasting (WRF) Model and to condition statistical models linking atmospheric conditions to electricity production. In Parts I and II of this work, we focus on making projections of annual energy production and the occurrence of electrical production drought. Here, we extend evaluation of the CF projections for sites in the Northeast, Midwest, southern Great Plains (SGP), and southwest U.S. coast (SWC) using statewide wind-generated electricity supply to the grid. We then quantify changes in the time scales of CF variability and the seasonality. Currently, wind-generated electricity is lowest in summer in each region except SWC, which causes a substantial mismatch with electricity demand. While electricity of residential heating may shift demand, research presented here suggests that summertime CF are likely to decline, potentially exacerbating the offset between seasonal peak power production and current load. The reduction in summertime CF is manifest for all regions except the SGP and appears to be linked to a reduction in synoptic-scale variability. Using fulfillment of 50% and 90% of annual energy production to quantify interannual variability, it is shown that wind power production exhibits higher (earlier fulfillment) or lower (later fulfillment) production for periods of over 10–30 years as a result of the action of internal climate modes. Significance Statement Electrical power system reassessment and redesign may be needed to aid efficient increased use of variable renewables in the generation of electricity. Currently wind-generated electricity in many regions of North America exhibits a minimum in summertime and hence is not well synchronized with electricity demand, which tends to be maximized in summer. Future projections indicate evidence of reductions in wind power during summer that would amplify this offset. However, electrification of heating may lead to increased wintertime demand, which would lead to greater synchronization.

Coburn, Jacob↗