Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

State densities of heavy nuclei in the static-path plus random-phase approximation

We report that nuclear state densities are important inputs to statistical models of compound-nucleus reactions. State densities are often calculated with self-consistent mean-field approximations that do not include important correlations and must be augmented with empirical collective enhancement factors. Here, we benchmark the static-path plus random-phase approximation (SPA + RPA) to the state density in a chain of samarium isotopes 148–155 Sm against exact results (up to statistical errors) obtained with the shell-model Monte Carlo (SMMC) method. The SPA + RPA method incorporates all static fluctuations beyond the mean field together with small-amplitude quantal fluctuations around each static fluctuation. Using a pairing plus quadrupole interaction, we show that the SPA + RPA state densities agree well with the exact SMMC densities for both the even- and odd-mass isotopes. For the even-mass isotopes, we also compare our results with mean-field state densities calculated with the finite-temperature Hartree-Fock-Bogoliubov (HFB) approximation. We find that the SPA + RPA repairs the deficiencies of the mean-field approximation associated with broken rotational symmetry in deformed nuclei and with the violation of particle-number conservation in the pairing condensate. In particular, in deformed nuclei the SPA + RPA reproduces the rotational enhancement of the state density relative to the mean-field state density.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Bayesian model averaging for analysis of lattice field theory results

Statistical modeling is a key component in the extraction of physical results from lattice field theory calculations. Although the general models used are often strongly motivated by physics, many model variations can frequently be considered for the same lattice data. Model averaging, which amounts to a probability-weighted average over all model variations, can incorporate systematic errors associated with model choice without being overly conservative. We discuss the framework of model averaging from the perspective of Bayesian statistics, and give useful formulae and approximations for the particular case of least-squares fitting, commonly used in modeling lattice results. In addition, we frame the common problem of data subset selection (e.g. choice of minimum and maximum time separation for fitting a two-point correlation function) as a model selection problem and study model averaging as a straightforward alternative to manual selection of fit ranges. Numerical examples involving both mock and real lattice data are given.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Statistical Downscaling of Climate Models for Solar Resource Assessment

This study presents the development of statistical models to efficiently downscale future projections of solar irradiance for solar energy applications. A climate data set simulated from a Regional Climate Model (RCM) obtained from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) is selected as input to the statistical models to create high-resolution global horizontal irradiance (GHI) over the contiguous United States (CONUS). Our approach builds statistical downscaling models that (1) regrid RCM data (0.22 degree and daily spatiotemporal resolution), (2) correct bias of GHI projections, (3) downscale the future GHI project from daily-scale to hourly-scale, and (4) spatially downscale to generate GHI at 8-km resolution. To calibrate and validate the statistical models, we adapt and use the National Solar Radiation Database (NSRDB). Preliminary results show that the statistical downscaling approach downscales future projections of GHI under two climate scenarios (RCP4.5 and RCP8.5) with a nBIAS of 3%, nMAE of 34% and nRMSE of 46% estimated against NSRDB for the contiguous United State. This presentation will summarize the implemented methodology and validation results as well as future extension of this research.

climate data↗

Benchmark of numerical modeling approaches on the systematic performance evaluation of wave energy converters

Different numerical modeling methods have been developed and applied to evaluate a variety of performance indicators of wave energy converters (WECs), including the power performance, structural loads, levelized cost of energy, etc. Based on the modeling fidelity, the commonly used numerical modeling approaches can be classified as linear modeling, weakly nonlinear modeling and fully nonlinear modeling approaches. Each method differs in accuracy and computational efficiency, making them suitable for different stages of WEC design. However, the selection of modeling approach could significantly impact evaluation outcomes. For instance, simplified linear models may underestimate structural loads or overestimate energy production in some operational conditions, potentially leading to less cost-effective designs. Given the widespread utilization of these models, it is essential to understand the uncertainties brought by them in performance evaluations. This work is dedicated to benchmarking different linear-potential-flow-based numerical models for evaluating the systematic performance of WECs. Three representative numerical modeling approaches are considered in this work, including linear frequency-domain modeling, statistically linearized spectral-domain modeling and Cummins equation-based nonlinear time-domain modeling. A generic point absorber WEC is considered as the research reference in this work, and different sea sites are taken into account. The numerical models are utilized to predict critical performance indicators, including power performance, the annual energy production, the capacity factor, the levelized cost of energy and the PTO fatigue loads. By comparing the results, this work identifies the uncertainties associated with different modeling approaches in evaluating WEC performance.

Fatigue↗

Bridging the gap between experiments and simulations using machine learning

The physics of inertial confinement fusion is rich and complex. Simulation codes that are used to design experiments are computationally expensive and lack the predictive capability required for extensive parameter exploration in search of a high-performing design for laser direct drive. In this work we use deep learning to build a fast emulator of experiments. To facilitate the development of the deep-learning model, an autoencoder is used to reduce the dimensionality of the input space. Two deep learning models are developed. One model is trained on a vast array of simulation data and is subsequently calibrated to expensive and limited experimental data using a technique known as “transfer learning.” The other model is trained on a statistical model and is subsequently calibrated using experimental data. A comparative study of the two predictive models is carried out. The models potentially reproduce key experimental observables with high accuracy and unprecedented inference times relative to those achieved with simulation codes. These models facilitate rapid exploration of a high dimensional input parameter space.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Determining reference standard strength for neutron-irradiated reduced activation ferritic/martensitic steel F82H by Bayesian method

The deterministic approach widely adopted in the design of structural components relies on systematically defined design limits using empirically determined safety factors. However, this approach is not always appropriate because structures are subjected to a variety of loads in the practical environment, which may result in excessively conservative design limits. In recent years, a more rigorous probabilistic approach that incorporates material strength distributions has become an important solution. In the probabilistic approach, the probability density functions of material strength properties underpin the design criteria. Here, the objective of this study is to identify the density distribution functions that best describe tensile properties of irradiated F82H to define a reference strength for DEMO design. Due to the limited number of existing data, this study specifically employs a Bayesian prediction method based on Monte Carlo simulations to determine a material reference value with statistical reliability and to investigate its effectiveness. For example, the dependence of tensile properties of 300 °C irradiated materials on irradiation damage and the range predicted by 95% Bayesian estimation was evaluated. As a statistical model for the dose dependence of statistical parameters, the normal distribution exhibited a better fit for 0.2% proof strength and tensile strength, whereas the distribution of total elongation data gave comparable reference values for both the normal and Weibull distribution models. Both models gave comparable criteria for the distribution of total elongation data. The Weibull model also gave better results for uniform elongation. The function best describing the model was a logarithmic law for both 0.2% proof strength and tensile strength, while a power law for both total and uniform elongation, which allowed for more comprehensive data prediction of irradiation data with statistical accuracy for DEMO reactor design.

36 MATERIALS SCIENCE↗

Real-time image denoising of mixed Poisson–Gaussian noise in fluorescence microscopy images using ImageJ

Fluorescence microscopy imaging speed is fundamentally limited by the measurement signal-to-noise ratio (SNR). To improve image SNR for a given image acquisition rate, computational denoising techniques can be used to suppress noise. However, common techniques to estimate a denoised image from a single frame either are computationally expensive or rely on simple noise statistical models. These models assume Poisson or Gaussian noise statistics, which are not appropriate for many fluorescence microscopy applications that contain quantum shot noise and electronic Johnson–Nyquist noise, therefore a mixture of Poisson and Gaussian noise. In this paper, we show convolutional neural networks (CNNs) trained on mixed Poisson and Gaussian noise images to overcome the limitations of existing image denoising methods. The trained CNN is presented as an open-source ImageJ plugin that performs real-time image denoising (within tens of milliseconds) with superior performance (SNR improvement) compared to conventional fluorescence microscopy denoising methods. The method is validated on external datasets with out-of-distribution noise, contrast, structure, and imaging modalities from the training data and consistently achieves high-performance ( > <#comment/> 8 d B ) denoising in less time than other fluorescence microscopy denoising methods.

47 OTHER INSTRUMENTATION↗

Diagnostics, Prognostics, and Optimization for Lithium-Ion Battery Systems

Health management of lithium-ion battery systems presents a host of challenges due to their complex physics, large numbers of components, and a wide variety of degradation behaviors across different battery types. Dr. Paul Gasper will present on research from the Electrochemical Energy Storage Group on Lithium-ion battery diagnostics, prognostics, and optimization. Diagnostics research, including state-estimation via machine-learning from electrochemical impedance spectroscopy and DC pulses as well as continuous state-estimation via Kalman filters, will highlight the ongoing challenges for accurately measuring the state of batteries without performing time-consuming characterization tests. NLR's industry-recognized battery prognostics work, which predicts real-world battery degradation by identifying degradation rate models from accelerated aging data using statistical modeling and machine-learning, will be used to demonstrate the critical impact of battery controls, thermal management, and operating strategy on durability and lifetime. Finally, the use of prognostic models for financial or lifetime optimization will be discussed.

25 ENERGY STORAGE↗

Self-diagnosis of model suitability for continuous measurements of stream-dissolved organic carbon derived from in situ UV–visible spectroscopy

Application of high-frequency monitoring of dissolved organic carbon (DOC) is difficult in instances where training datasets are challenging to develop (e.g., remote locations) and the relationship between optical features and DOC concentration changes due to environmental or landscape shifts (e.g., climate or land-use change). We developed and compared three partial least squares (PLS) models using in situ water level measurements, conductivity, and UV–Vis spectral attenuation to predict DOC. Two site-specific models were developed using data from a hillslope-dominated forest or a low-relief wetland-pond-dominated stream catchment. The third model, using data from both sites, exhibited the best performance (DOC range = 4–15.5 mg C L −1 , mean = 8.38 mg C L −1 , training RMSE = 0.34 mg C L −1 , internal validation RMSE = 0.50 mg C L −1 , external validation RMSE = 2.43 mg C L −1 ). We further demonstrate using PLS model statistics to monitor performance and elucidate when and how models should be updated. These statistics, Hotelling's T 2 and squared prediction errors, are useful consistency checks for the predictions made and detect underlying inconsistencies that, if undetected, can reduce the robustness of DOC prediction. For example, via the T 2 statistic, we identified the summer–autumn transition as a period when DOC composition differed from what was represented in the training dataset. We also determined that elevated SUVA 254 values contributed to the overall bias observed in predictions made during the subsequent year as part of the external validation. This enabled the application of a bias correction that reduced the RMSE from 2.43 to 0.89 mg C L −1 . The method presented here could be applied to future monitoring programs enabling model updates to monitor DOC fluxes accurately from optical datasets (e.g., attenuance or fluorescence) in the face of developing datasets in remote locations or environmental change. In conclusion, implementation of this approach may also identify possible regime shifts or landscape and hydrologic change associated with climate and other environmental changes relevant to terrestrial to aquatic fluxes.

UV-VIS spectroscopy↗

Hybrid geological modeling: Combining machine learning and multiple-point statistics

Accurately modeling and constructing a geologically realistic subsurface model remains an outstanding problem as the morphology controls the flow behaviors. Particularly, one of the pattern-based methods, namely cross-correlation based simulation, has been proved to be an effective way to reconstruct a realistic model, at both small and large scales. However, conditioning to point data in the large-scale problems is still a crucial issue in these algorithms, since there is always a trade-off between the quality of the realizations and the degree of point data reproduction. Specifically, it is not practical to build a training image (TI) which includes all the possibilities and variabilities. Therefore, finding a pattern that can represent the point data and, at the same time, preserving the connectivities is difficult. This leads to producing highly-connected realizations with a significant mismatch or poor models with a reasonable degree of point data reproduction. To accurately reproduce the densely distributed hard data, pixel-based methods can also produce some unrealistic artifacts around the hard data. In this paper, to overcome this challenge, however, we use pattern-based methods as they often produce more disconnected geobodies when dealing with dense hard data, and proposed a hybrid algorithm using the pattern-based methods and convolutional neural network (CNN). The trained CNN model is utilized to improve the quality of conditioning to point data for the original realizations generated by the pattern-based algorithm. As such, the mismatch locations are identified, and the same regions are used in the training of CNN to mimic the procedure through which a missing region can be filled. To evaluate the performance of the proposed hybrid algorithm, it is tested on cases with different dimensions and different numbers of facies. Then, the newly improved realizations are compared with the initial realizations generated by the pattern-based algorithm. The comparison is also conducted by the flow simulation test. And it indicates that the proposed hybrid algorithm can better reproduce the point data, while the connectivities are better preserved.

58 GEOSCIENCES↗

A Comparison of Infectious Disease Forecasting Methods across Locations, Diseases, and Time

Accurate infectious disease forecasting can inform efforts to prevent outbreaks and mitigate adverse impacts. This study compares the performance of statistical, machine learning (ML), and deep learning (DL) approaches in forecasting infectious disease incidences across different countries and time intervals. We forecasted three diverse diseases: campylobacteriosis, typhoid, and Q-fever, using a wide variety of features (n = 46) from public datasets, e.g., landscape, climate, and socioeconomic factors. We compared autoregressive statistical models to two tree-based ML models (extreme gradient boosted trees [XGB] and random forest [RF]) and two DL models (multi-layer perceptron and encoder–decoder model). The disease models were trained on data from seven different countries at the region-level between 2009–2017. Forecasting performance of all models was assessed using mean absolute error, root mean square error, and Poisson deviance across Australia, Israel, and the United States for the months of January through August of 2018. The overall model results were compared across diseases as well as various data splits, including country, regions with highest and lowest cases, and the forecasted months out (i.e., nowcasting, short-term, and long-term forecasting). Overall, the XGB models performed the best for all diseases and, in general, tree-based ML models performed the best when looking at data splits. There were a few instances where the statistical or DL models had minutely smaller error metrics for specific subsets of typhoid, which is a disease with very low case counts. Feature importance per disease was measured by using four tree-based ML models (i.e., XGB and RF with and without region name as a feature). The most important feature groups included previous case counts, region name, population counts and density, mortality causes of neonatal to under 5 years of age, sanitation factors, and elevation. This study demonstrates the power of ML approaches to incorporate a wide range of factors to forecast various diseases, regardless of location, more accurately than traditional statistical approaches.

59 BASIC BIOLOGICAL SCIENCES↗

In-process monitoring and prediction of droplet quality in droplet-on-demand liquid metal jetting additive manufacturing using machine learning

Abstract In droplet-on-demand liquid metal jetting (DoD-LMJ) additive manufacturing, complex physical interactions govern the droplet characteristics, such as size, velocity, and shape. These droplet characteristics, in turn, determine the functional quality of the printed parts. Hence, to ensure repeatable and reliable part quality it is necessary to monitor and control the droplet characteristics. Existing approaches for in-situ monitoring of droplet behavior in DoD-LMJ rely on high-speed imaging sensors. The resulting high volume of droplet images acquired is computationally demanding to analyze and hinders real-time control of the process. To overcome this challenge, the objective of this work is to use time series data acquired from an in-process millimeter-wave sensor for predicting the size, velocity, and shape characteristics of droplets in DoD-LMJ process. As opposed to high-speed imaging, this sensor produces data-efficient time series signatures that allows rapid, real-time process monitoring. We devise machine learning models that use the millimeter-wave sensor data to predict the droplet characteristics. Specifically, we developed multilayer perceptron-based non-linear autoregressive models to predict the size and velocity of droplets. Likewise, a supervised machine learning model was trained to classify the droplet shape using the frequency spectrum information contained in the millimeter-wave sensor signatures. High-speed imaging data served as ground truth for model training and validation. These models captured the droplet characteristics with a statistical fidelity exceeding 90%, and vastly outperformed conventional statistical modeling approaches. Thus, this work achieves a practically viable sensing approach for real-time quality monitoring of the DoD-LMJ process, in lieu of the existing data-intensive image-based techniques.

Gaikwad, Aniruddha (ORCID:0000000285642621)↗

Deep learning-based predictive models for laser direct drive at the Omega Laser Facility

The rich and complex physics of inertial confinement fusion provides a unique and challenging space for high-fidelity first-principles modeling. Consequently, simulation codes that are used to design experiments are computationally expensive and lack the predictive capability required for extensive parameter exploration in search of a high-performing design for laser direct drive. In this article, we present two deep-learning-based predictive models intended to address these difficulties. The first model (TL DNN) acts as a fast emulator of simulations as well as experiments at the Omega Laser Facility. This model is trained on a simulation database and subsequently calibrated on experimental data using transfer learning. To facilitate the development of this model, an autoencoder is developed to reduce the dimensionality of the input space by compressing the laser pulse input. The model predicts key experimental scalar observables of Omega experiments with high accuracy and minimal computational cost. This deep neural net enables rapid exploration of a high-dimensional input parameter space for an optimal implosion design. The second model (DNN SM+) aims to extend the statistical modeling work of Lees et al. [Phys. Rev. Lett. 127, 105001 (2021)], by increasing the complexity of the model space and allowing for coupling between degradation terms. Since the model capacity of DNN SM+ is higher than the model of Lees et al., DNN SM+ can potentially provide an improvement in predictive capability, and we use this model to provide insight into complicated degradation dependencies.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

InClass nets: independent classifier networks for nonparametric estimation of conditional independence mixture models and unsupervised classification

Abstract Conditional independence mixture models (CIMMs) are an important class of statistical models used in many fields of science. We introduce a novel unsupervised machine learning technique called the independent classifier networks (InClass nets) technique for the nonparameteric estimation of CIMMs. InClass nets consist of multiple independent classifier neural networks (NNs), which are trained simultaneously using suitable cost functions. Leveraging the ability of NNs to handle high-dimensional data, the conditionally independent variates of the model are allowed to be individually high-dimensional, which is the main advantage of the proposed technique over existing non-machine-learning-based approaches. Two new theorems on the nonparametric identifiability of bivariate CIMMs are derived in the form of a necessary and a (different) sufficient condition for a bivariate CIMM to be identifiable. We use the InClass nets technique to perform CIMM estimation successfully for several examples. We provide a public implementation as a Python package called RainDancesVI.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Validation of MCNP Critical Benchmark Models of PU-MET-FAST-016 [Slides]

The validation of PU-MET-FAST-016 contributed to the centralized LANL benchmark repository currently under development. The revisions made the MCNP models statistically, significantly more similar to the benchmark models. The overall impact on the USL is negligible. The revisions to the PU-MET-FAST-016 models provide value to the Los Alamos Benchmark Suite without invalidating past Whisper results.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Predictive Indicators of the Performance of Large Language Models

In several mission contexts, it is desirable to estimate the performance of large language models (LLMs) on tasks that we cannot run directly. In light of published “scaling laws” our hypothesis is that some tasks should be consistently more challenging than others based on characteristics of the task. The goal of this project was to begin quantifying how much information about LLM performance can be gained from the features of a model and a task. Two of our statistical models struggled to converge. Pass/fail test results may provide limited information for inference beyond model quality and task difficulty, but we see no evidence at this time for significant feature interaction effect sizes, arguing for simple models. Future work extending the models to capitalize on perplexity of ground truth answers is suggested. This project also introduces “Depth of Knowledge Variant Testing” as a strategy for more finely assessing language models on open domain question and answer tasks. We developed sets of questions that ask a language model to produce similar information while demonstrating increasing depth of knowledge, and also relabeled existing Q&A test questions with their depth of knowledge. Our results suggest further consideration of Bloom’s taxonomy and further refinement of prompts to properly elicit information at varying depths. In the course of this work, we set up a basic infrastructure for standardizing tasks and testing many language models on these tasks. In addition to testing the predictive quality of model features and performance across test suites, with this project we have introduced two new task features to contextualize each test question: the Dewey Classification main category of information covered, and the Bloom’s taxonomy level that corresponds to the depth of knowledge probed by the question. Splits across these and other features produced over five hundred task subtypes with distinct feature vectors, which we tested on half a dozen models.

97 MATHEMATICS AND COMPUTING↗

Bayesian chain graph models to characterize microbe-environment dynamics

Microbiome data require statistical models that can simultaneously decode microbes' reaction to the environment and interactions among microbes. While a multiresponse linear regression model seems like a straight-forward solution, we argue that treating it as a graphical model is problematic given that the regression coefficient matrix does not encode the conditional dependence structure between response and predictor nodes. This observation is especially important in biological settings when we have prior knowledge on the edges from specific experimental interventions that can only be properly encoded under a conditional dependence model. Here, we propose a chain graph model with two sets of nodes (predictors and responses) whose solution yields a graph with edges that indeed represent conditional dependence, thus agreeing with the experimenter's intuition on the average behavior of nodes under treatment. The solution to our model is sparse via the Bayesian linear regression (LASSO). In addition, we propose an adaptive extension so that different shrinkages can be applied to different edges to incorporate edge-specific prior knowledge. Our model is computationally inexpensive through an efficient Gibbs sampling algorithm and can account for binary, counting, and compositional responses via an appropriate hierarchical structure. We test the performance of our model in a variety of simulated datasets, thereby showing superior performance to state-of-the-art approaches. We further apply our model to human gut and soil microbial compositional datasets, and we highlight that CG-LASSO can estimate biologically meaningful network structures in the data.

compositional data↗

Lithium-ion battery physics and statistics-based state of health model

A pseudo-2d model using COMSOL Multiphysics® software is developed to simulate performance and performance degradation of Li-ion batteries consisting of layered and olivine cathodes with graphite anode when subjected to peak shaving grid service. Multiple degradation pathways are considered, including solid electrolyte interphase (SEI) formation and breakdown at the anode, cathode dissolution and its synergistic effect on SEI formation at the anode. The model is validated by simulating commercial cylindrical cell performance. A global model is developed to simulate performance across all chemistries, along with individual chemistry models using global model parameters as initial values. There is good agreement between these models for various optimization parameters such as SEI equilibrium potential, cathode dissolution exchange current density, solvent diffusivity in the SEI and SEI ionic conductivity. To circumvent time constraints related to the COMSOL model, a 0d global model is developed which fits data well and provides more clarity on differences in cathode dissolution exchange current density. Again, good agreement for various optimization parameters is obtained among the COMSOL global & individual chemistry models and the 0-d model. The lessons learned from the physics-based model is used to develop a top down statistics-based model using current, voltage and anode volumetric change per mole lithium intercalated, along with their interactions as degradation predictors. This model predicts out of sample degradation for multiple grid services and electric vehicle drive cycle with high accuracy and provides the pathway to develop an efficient battery management system combining machine learning and findings from physics-based computationally intensive algorithms.

Crawford, Aladsair J.↗