Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Power grid frequency prediction using spatiotemporal modeling

Understanding power system dynamics is essential for interarea oscillation analysis and the detection of grid instabilities. The FNET/GridEye is a GPS-synchronized wide-area frequency measurement network that provides an accurate picture of the normal real-time operational condition of the power system-dynamics, giving rise to new and intricate spatiotemporal patterns of power loads. We propose to model FNET/GridEye grid frequency data from the U.S. Eastern Interconnection with a spatiotemporal statistical model. We predict the frequency data at locations without observations, a critical need during disruption events where measurement data are inaccessible. Spatial information is accounted for either as neighboring measurements in the form of covariates or with a spatiotemporal correlation model captured by a latent Gaussian field. Finally, the proposed method is useful in estimating power system dynamic response from limited phasor measurements and holds promise for predicting instability that may lead to undesirable effects such as cascading outages.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Permafrost Carbon: Progress on Understanding Stocks and Fluxes Across Northern Terrestrial Ecosystems

Significant progress in permafrost carbon science made over the past decades include the identification of vast permafrost carbon stocks, the development of new pan-Arctic permafrost maps, an increase in terrestrial measurement sites for CO 2 and methane fluxes, and important factors affecting carbon cycling, including vegetation changes, periods of soil freezing and thawing, wildfire, and other disturbance events. Process-based modeling studies now include key elements of permafrost carbon cycling and advances in statistical modeling and inverse modeling enhance understanding of permafrost region C budgets. By combining existing data syntheses and model outputs, the permafrost region is likely a wetland methane source and small terrestrial ecosystem CO 2 sink with lower net CO 2 uptake toward higher latitudes, excluding wildfire emissions. For 2002–2014, the strongest CO2 sink was located in western Canada (median: -52 g C m -2 y -1 ) and smallest sinks in Alaska, Canadian tundra, and Siberian tundra (medians: -5 to -9 g C m -2 y -1 ). Eurasian regions had the largest median wetland methane fluxes (16–18 g CH4 m -2 y -1 ). Quantifying the regional scale carbon balance remains challenging because of high spatial and temporal variability and relatively low density of observations. More accurate permafrost region carbon fluxes require: (a) the development of better maps characterizing wetlands and dynamics of vegetation and disturbances, including abrupt permafrost thaw; (b) the establishment of new year-round CO 2 and methane flux sites in underrepresented areas; and (c) improved models that better represent important permafrost carbon cycle dynamics, including non-growing season emissions and disturbance effects.

54 ENVIRONMENTAL SCIENCES↗

Exploring scenarios for enhanced fuel compression and performance on the National Ignition Facility with machine-learning-aided design techniques

Recent fusion experiments on the National Ignition Facility (NIF) have achieved ignition, producing multi-MJ fusion yields for input laser energies of roughly 2 MJ [Abu-Shawareb et al., Phys. Rev. Lett. 132, 065102 (2024)]. Building on the success of the target designs that have achieved ignition, we explore new implosion scenarios predicted to generate significantly more compression of the dense DT ice layer and correspondingly higher yields while preserving many of the key physics characteristics of present-day ignition designs. Our main result is a novel 3-shock implosion scheme that effectively minimizes the shock-induced entropy in the dense, accelerating DT shell and maximizes the resulting fuel compression subject to a fixed leading shock strength consistent with present-day ignition experiments, which is necessary to melt the crystalline high-density carbon ablator. Compared to the first NIF experiment to fulfill Lawson's ignition criterion, shot N210808 [Abu-Shawareb et al., Phys. Rev. Lett. 129, 075001 (2022)], our design exhibits a 40% increase in simulated peak areal density (ρR) and a 5× increase in 1D fusion yield using a 4% lighter ablator and identical DT payloads. We also present a complete integrated 2D hohlraum design and laser pulse specifications capable of generating the desired 3-shock drive and maintaining control of the low-mode capsule implosion symmetry, where the increase in simulated 2D yield relative to N210808 is > 10×. This new implosion regime was discovered with help from a machine-learning-enabled capsule design optimization framework. We outline the workflow this automated tool uses to identify improved design candidates by running several rounds of capsule simulations, constructing a surrogate model mapping input variations to key physics output quantities, and querying the resulting statistical model to propose adjustments to the x-ray drive and capsule to reach a set of physics objectives prescribed by the designer.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Warming of the lower Columbia River , 1853 to 2018

Abstract Water temperature is a critical ecological indicator; however, few studies have statistically modeled century‐scale trends in riverine or estuarine water temperature, or their cause. Here, we recover, digitize, and analyze archival temperature measurements from the 1850s onward to investigate how and why water temperatures in the lower Columbia River are changing. To infill data gaps and explore changes, we develop regression models of daily historical Columbia River water temperature using time‐lagged river flow and air temperature as the independent variables. Models were developed for three time periods (mid‐19th, mid‐20th, and early 21st century), using archival and modern measurements (1854–1876; 1938–present). Daily and monthly averaged root‐mean‐square errors overall are 0.89°C and 0.77°C, respectively for the 1938–2018 period. Results suggest that annual averaged water temperature increased by 2.2°C ± 0.2°C since the 1850s, a rate of 1.3°C ± 0.1°C/century. Increased water temperatures are seasonally dependent. An increase of approximately 2.0°C ± 0.2°C/century occurs in the July–Dec time‐frame, while springtime trends are statistically insignificant. Rising temperatures change the probability of exceeding ecologically important thresholds; since the 1850s, the number of days with water temperatures over 20°C increased from ~5 to 60 per year, while the number below 2°C decreased from ~10 to 0 days/per year. Overall, the modern system is warmer, but exhibits less temperature variability. The reservoir system reduces sensitivity to short‐term atmospheric forcing. Statistical experiments within our modeling framework suggest that increased water temperature is driven by warming air temperatures (~29%), altered river flow (~14%), and water resources management (~57%).

54 ENVIRONMENTAL SCIENCES↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

Enabling energy‐efficient manufacturing of pharmaceutical solid oral dosage forms via integrated techno‐economic analysis and advanced process modeling

Abstract The global pharmaceutical industry is a trillion‐dollar market. However, the pharmaceutical sector often lags in manufacturing innovation and automation which limits its potential to maximize energy efficiency. The integration of techno‐economic analysis (TEA) with advanced process models as part of an overarching smart manufacturing platform, can help industries create business models, which can be adapted for manufacturing to reduce energy consumption and operating costs while ensuring product quality which can further enable a more sustainable process operation. In this study, a rational design of experiment on three unit‐operations (wet granulation, drying, and milling) was performed on a batch (case 1) and continuous (case 2) pharmaceutical process to obtain experimental data. Process models for predicting product quality and energy efficiency of each of the three‐unit operations were developed. The experimental data were used to validate the models and good agreement was observed. The energy consumption of each unit operation was calculated using statistical models relating the power consumption and the process parameters. The developed process models and energy models were further integrated into a TEA framework, which quantified the energy and monetary cost of manufacturing for both batch and continuous manufacturing cases. With this integrated framework, energy costs savings of ~33% was obtained in the continuous manufacturing process (case 2) over the batch process (case 1).

Sampat, Chaitanya↗

Model-based economic analysis under uncertainty for PFAS treatment by granular activated carbon and ion exchange technologies

Recent drinking water regulations have imposed the need for per- and polyfluoroalkyl substances (PFAS) remediation. In response, treatment facilities may be required to retrofit existing treatment schemes to treat PFAS below maximum contaminant levels (MCLs). Adsorption technologies such as granular activated carbon (GAC) and ion exchange (IX) have been demonstrated to be effective; however, there are limited techno-economic metrics available which provide guidance on technology selection and design for diverse PFAS-containing source water conditions. Process systems engineering (PSE) tools which can traditionally perform these analyses are hindered by the data availability, model validity, and understanding of treatment phenomena for emerging contaminants. This work employs published data regressions, statistical models, process models, techno-economic analyses, and other process systems tools in a model-based uncertainty framework to consider the limitations of emerging contaminant research. Through this analysis framework, economic results are provided as probabilistic distributions based on the uncertainty of the models and diverse conditions that treatment facilities experience.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analytical methods for superresolution dislocation identification in dark-field X-ray microscopy

In this work, we develop several inference methods to estimate the position of dislocations from images generated using dark-field X-ray microscopy (DFXM)—achieving superresolution accuracy and principled uncertainty quantification. Using the framework of Bayesian inference, we incorporate models of the DFXM contrast mechanism and detector measurement noise, along with initial position estimates, into a statistical model coupling DFXM images with the dislocation position of interest. We motivate several position estimation and uncertainty quantification algorithms based on this model. We then demonstrate the accuracy of our primary estimation algorithm on synthetic realistic DFXM images of edge dislocations in single-crystal aluminum. We conclude with a discussion of our methods’ impact on future dislocation studies and possible future research avenues.

36 MATERIALS SCIENCE↗

The Need for a System Science Approach to Global Magnetospheric Models

This perspective advocates for the need of a combined system science approach to global magnetospheric models and to spacecraft magnetospheric data to answer the question “Do simulations behave in the same manner as the magnetosphere does?” (instead of the standard validation question “How well do simulations reproduce spacecraft data?”). This approach will 1) validate global magnetospheric models statistically, without the need for a direct comparison against spacecraft data, 2) expose the deficiencies of the models, and 3) provide physics support to the system analysis performed on the magnetospheric system.

79 ASTRONOMY AND ASTROPHYSICS↗

Reactor Network Analysis with Various Reaction Mechanisms to Investigate Hydrogen vs. Methane Fuel at Varying Flame Temperatures with Experimental Data

Emissions data were evaluated for a set of test hardware that adapts Collins Aerospace’s aeroengine liquid fueled injector technology to a ground-based turbine. The hardware was specifically developed to operate on 100% hydrogen; however it has the ability to operate on both pure methane and hydrogen and mixtures in between. 16 injector configurations and seven factors (air split, fuel and air swirl, pressure drop, preheat temperature, fuel composition, and flame temperature) were investigated based on a Box Behnken statistical model. Of the 16, configuration 3 was selected for further discussion being in the middle of the statistical design space. A chemical reactor network was developed based on the experimental data obtained and was assessed with three different mechanisms: GRI Mech 3.0, UCSD, and Galway. All factors were held constant except the fuel composition to directly compare chemical pathways for methane and hydrogen. This study was conducted across different flame temperatures of 1500K, 1675K, and 1850K for both hydrogen and methane. Additional lower flame temperatures of 1100K and 1300K were evaluated for hydrogen. Moreover, results for fully premixed and non-premixed results were compared to obtain information on the implications of mixedness on emissions. These results were compared with the Leonard and Stegmaier plot for methane and a similar plot was constructed for hydrogen.

08 HYDROGEN↗

Leveraging Large Language Models for Understanding Fundamental Principles of Catalysis

Heterogeneous catalysis presents a distinct challenge for artificial intelligence (AI). Data sets are often small and inconsistently reported, catalyst representations are not standardized, and extracting fundamental knowledge requires integrating performance data, spectroscopic characterizations, and mechanistic models across multiple scales. Language offers a unifying representation across these modalities, making catalysis well suited for leveraging large language models (LLMs). By standardizing how catalytic data is represented, LLMs make dispersed experimental results more accessible to downstream statistical modeling. In this perspective, we focus our discussion around three opportunities where LLMs can significantly contribute to catalysis: (1) text to properties; (2) text to structure; and (3) text to mechanistic models. The discussion is followed by a perspective section on LLM-readiness of data, aligning LLM outputs with scientific correctness, and bridging lab-scale discovery to industrial deployment. Across each area, the most productive applications couple dispersed chemical knowledge with physics-grounded validation to produce verifiable hypotheses and actionable representations.

Catalysts↗

Assessing correlated truncation errors in modern nucleon-nucleon potentials

We test the BUQEYE model of correlated effective field theory (EFT) truncation errors on Reinert, Krebs, and Epelbaum's semilocal momentum-space implementation of the chiral EFT (𝜒⁢EFT ) expansion of the nucleon-nucleon (NN) potential. This Bayesian model hypothesizes that dimensionless coefficient functions extracted from the order-by-order corrections to NN observables can be treated as draws from a Gaussian process (GP). We combine a variety of graphical and statistical diagnostics to assess when predicted observables have a 𝜒⁢EFT convergence pattern consistent with the hypothesized GP statistical model. Our conclusions are that, first, the BUQEYE model is generally applicable to the potential investigated here, which enables statistically principled estimates of the impact of higher EFT orders on observables. Second, parameters defining the extracted coefficients such as the expansion parameter 𝑄 must be well chosen for the coefficients to exhibit a regular convergence pattern—a property we exploit to obtain posterior distributions for such quantities. Third, the assumption of GP stationarity across lab energy and scattering angle is not generally met; this necessitates adjustments in future work. We provide a workflow and interpretive guide for our analysis framework, and show what can be inferred about probability distributions for 𝑄, the EFT breakdown scale Λ 𝑏 , the scale associated with soft physics in the 𝜒⁢EFT potential 𝑚 eff , and the GP hyperparameters. All our results can be reproduced using a publicly available Jupyter notebook, which can be straightforwardly modified to analyze other 𝜒⁢EFT NN potentials.

Bayesian methods↗

Discovering the Unknowns: A First Step

This article aims at discovering the unknown variables in the system through data analysis. The main idea is to use the time of data collection as a surrogate variable and try to identify the unknown variables by modeling gradual and sudden changes in the data. We use Gaussian process modeling and a sparse representation of the sudden changes to efficiently estimate the large number of parameters in the proposed statistical model. The method is tested on a realistic dataset generated using a one-dimensional implementation of a Magnetized Liner Inertial Fusion (MagLIF) simulation model, and encouraging results are obtained.

42 ENGINEERING↗

Experimentally Inferred Fusion Yield Dependencies of OMEGA Inertial Confinement Fusion Implosions

Statistical modeling of experimental and simulation databases has enabled the development of an accurate predictive capability for deuterium-tritium layered cryogenic implosions at the OMEGA laser. Here, a physics-based statistical mapping framework is described and used to uncover the dependencies of the fusion yield. This model is used to identify and quantify the degradation mechanisms of the fusion yield in direct-drive implosions on OMEGA. The yield is found to be reduced by the ratio of laser beam to target radius, the asymmetry in inferred ion temperatures from the $l$ = 1 mode, the time span over which tritium fuel has decayed, and parameters related to the implosion hydrodynamic stability. When adjusted for tritium decay and $l$ = 1 mode, the highest yield in OMEGA cryogenic implosions is predicted to exceed 2 × 10 14 fusion reactions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Identifying microbial drivers in biological phenotypes with a Bayesian network regression model

Abstract In Bayesian Network Regression models, networks are considered the predictors of continuous responses. These models have been successfully used in brain research to identify regions in the brain that are associated with specific human traits, yet their potential to elucidate microbial drivers in biological phenotypes for microbiome research remains unknown. In particular, microbial networks are challenging due to their high dimension and high sparsity compared to brain networks. Furthermore, unlike in brain connectome research, in microbiome research, it is usually expected that the presence of microbes has an effect on the response (main effects), not just the interactions. Here, we develop the first thorough investigation of whether Bayesian Network Regression models are suitable for microbial datasets on a variety of synthetic and real data under diverse biological scenarios. We test whether the Bayesian Network Regression model that accounts only for interaction effects (edges in the network) is able to identify key drivers (microbes) in phenotypic variability. We show that this model is indeed able to identify influential nodes and edges in the microbial networks that drive changes in the phenotype for most biological settings, but we also identify scenarios where this method performs poorly which allows us to provide practical advice for domain scientists aiming to apply these tools to their datasets. BNR models provide a framework for microbiome researchers to identify connections between microbes and measured phenotypes. We allow the use of this statistical model by providing an easy‐to‐use implementation which is publicly available Julia package at https://github.com/solislemuslab/BayesianNetworkRegression.jl .

59 BASIC BIOLOGICAL SCIENCES↗

Adaptation Strategies Strongly Reduce the Future Impacts of Climate Change on Simulated Crop Yields

Abstract Simulations of crop yield due to climate change vary widely between models, locations, species, management strategies, and Representative Concentration Pathways (RCPs). To understand how climate and adaptation affects yield change, we developed a meta‐model based on 8703 site‐level process‐model simulations of yield with different future adaptation strategies and climate scenarios for maize, rice, wheat and soybean. We tested 10 statistical models, including some machine learning models, to predict the percentage change in projected future yield relative to the baseline period (2000–2010) as a function of explanatory variables related to adaptation strategy and climate change. We used the best model to produce global maps of yield change for the RCP4.5 scenario and identify the most influential variables affecting yield change using Shapley additive explanations. For most locations, adaptation was the most influential factor determining the projected yield change for maize, rice and wheat. Without adaptation under RCP4.5, all crops are expected to experience average global yield losses of 6%–21%. Adaptation alleviates this average projected loss by 1–13 percentage points. Maize was most responsive to adaptive practices with a projected mean yield loss of −21% [range across locations: −63%, +3.7%] without adaptation and −7.5% [range: −46%, +13%] with adaptation. For maize and rice, irrigation method and cultivar choice were the adaptation types predicted to most prevent large yield losses, respectively. When adaptation practices are applied, some areas are predicted to experience yield gains, especially at northern high latitudes. These results reveal the critical importance of implementing adequate adaptation strategies to mitigate the impact of climate change on crop yields.

54 ENVIRONMENTAL SCIENCES↗

Bayesian characterization of uncertainties surrounding fluvial flood hazard estimates

Fluvial floods drive severe risk to riverine communities. There is strong evidence of increasing flood hazards in many regions around the world. The choice of methods and assumptions used in flood hazard estimates can impact the design of risk management strategies. In this study, we characterize the expected flood hazards conditioned on the uncertain model structures, model parameters, and prior distributions of the parameters. We construct a Bayesian framework for river stage return level estimation using a nonstationary statistical model that relies exclusively on the Indian Ocean Dipole Index. We show that ignoring uncertainties can lead to biased estimation of expected flood hazards. We find that the considered model parametric uncertainty is more influential than model structures and model priors. Our results highlight the importance of incorporating uncertainty in extreme flood stage estimates, and are of practical use for informing water infrastructure designs in a changing climate.

54 ENVIRONMENTAL SCIENCES↗

Bayesian model-data comparison incorporating theoretical uncertainties

Accurate comparisons between theoretical models and experimental data are critical for scientific progress. However, inferred physical model parameters can vary significantly with the chosen physics model, highlighting the importance of properly accounting for theoretical uncertainties. In this Letter, we present a Bayesian framework that explicitly quantifies these uncertainties by statistically modeling theory errors, guided by qualitative knowledge of a theory’s varying reliability across the input domain. We demonstrate the effectiveness of this approach using two systems: a simple ball drop experiment and multi-stage heavy-ion simulations. In both cases incorporating model discrepancy leads to improved parameter estimates, with systematic improvements observed as additional experimental observables are integrated.

Bayesian methods↗