Criteria for and extrapolation in overstress models
Accelerated life test models, criteria for model selection, and extrapolation in overstress models
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Accelerated life test models, criteria for model selection, and extrapolation in overstress models
Estimating the probability of failure for complex real-world systems using high-fidelity computational models is often prohibitively expensive, especially when the probability is small. Exploiting low-fidelity models can make this process more feasible, but merging information from multiple low-fidelity and high-fidelity models poses several challenges. Here, this paper presents a robust multi-fidelity surrogate modeling strategy in which the multi-fidelity surrogate is assembled using an active learning strategy using an on-the-fly model adequacy assessment set within a subset simulation framework for efficient reliability analysis. The multi-fidelity surrogate is assembled by first applying a Gaussian process correction to each low-fidelity model and assigning a model probability based on the model's local predictive accuracy and cost. Three strategies are proposed to fuse these individual surrogates into an overall surrogate model based on model averaging and deterministic/stochastic model selection. The strategies also dictate which model evaluations are necessary. No assumptions are made about the relationships between low-fidelity models, while the high-fidelity model is assumed to be the most accurate and most computationally expensive model. Through two analytical and two numerical case studies, including a case study evaluating the failure probability of Tristructural isotropic-coated (TRISO) nuclear fuels, the algorithm is shown to be highly accurate while drastically reducing the number of high-fidelity model calls (and hence computational cost).
Abstract Acoustic telemetry studies often rely on the assumption that premature tag failure does not affect the validity of inferences. However, in some cases this assumption is possibly or likely invalid and it is necessary to apply a correction to estimation procedures. The question of which approaches and specific models are best suited to modeling acoustic tag failures has received little research attention. In this short communication, we present a meta-analysis of 42 acoustic tag-life studies, originally used to correct survival studies involving outmigrating juvenile salmonids in the Columbia/Snake river basin. We compare the performance of nine alternative parametric models including common failure–time/survival models and the vitality models of Li and Anderson Theor Popul Biol 76:118–131, (2009) and Demogr Res 28:341–372, (2013). The tag-life studies used acoustic tags from three different tag manufacturers, had expected lifetimes between 12 and 61 days, and had dry weights ranging from 0.22 to 1.65 g. In 57% of the cases, the vitality models of Li and Anderson Theor Popul Biol 76:118–131, (2009) and Demogr Res 28:341–372, (2013) fit the tag-failure times best. The vitality models were also the second-best choices in 17% of the cases. Together, the vitality models, log-logistic, (19%), and gamma models (14%) accounted for 90% of the models selected. Unlike more traditional failure–time models (e.g., Weibull, Gompertz, gamma, and log-logistic), the vitality models are capable of characterizing both the early onset of tag failure due to manufacturing errors and the anticipated battery life. We provide further guidance on appropriate sample sizes (50–100 tags) and procedures to be considered when applying precise tag-life corrections in release–recapture survival studies.
This presentations provides a high-level overview of the NETL-developed techno-economic models associated with the CO2 transport and CO2 storage components of the carbon capture and storage (CCS)/carbon capture, utilization, and storage (CCUS) value chain. It also discusses the models’ capabilities through the discussion of select modeling applications (both internal and external) and highlights current model modifications and future work. It was presented at the CCUS 2023 conference, organized and presented by the Society of Petroleum Engineers (SPE), American Association of Petroleum Geologists (AAPG), and Society of Exploration Geophysicists (SEG) and held in Houston, Texas, April 25-27, 2023.
Subsurface flow research is essential for the sustainable management of natural resources and the environment. Deep learning (DL) has significantly advanced this field by developing efficient and accurate surrogate models to replace computationally expensive physics‐based simulations. These surrogate models are commonly used to predict the spatiotemporal evolution of state variables, such as gas saturation and reservoir pressure, in heterogeneous geological formations. Despite the various DL models applied to this task, there is a lack of studies systematically comparing their performance. This absence of comparative analysis leads to somewhat arbitrary DL model selection in subsurface flow research, resulting in suboptimal performance and potentially inaccurate predictions. To bridge this gap, we conduct a systematic comparison study of three popular DL architectures—U‐Net, Fourier Neural Operators (FNO), and Segmentation Transformer (SETR)—in surrogate modeling of underground hydrogen storage (UHS). We focus on UHS due to its promise of enhancing clean energy resilience and its cyclic operational conditions that represent common scenarios in various subsurface applications. We evaluate the models based on accuracy, training cost, and inference speed. The comparison shows that U‐Net achieves the highest accuracy, followed by SETR and FNO. Despite its lower accuracy, FNO has the highest inference speed. SETR offers competitive accuracy with the least training memory usage, demonstrating the potential of transformers in learning subsurface flow. Our results provide guidance for selecting DL models for surrogate modeling in a wide range of subsurface flow problems.
The sensitivity of photosynthesis to environmental changes is essential for understanding carbon cycle responses to global climate change and for the development of modeling approaches that explains its spatial and temporal variability. We collected a large variety of published sensitivity functions of gross primary productivity (GPP) to different forcing variables to assess the response of GPP to environmental factors. These include the responses of GPP to temperature; vapor pressure deficit, some of which include the response to atmospheric CO 2 concentrations; soil water availability (W); light intensity; and cloudiness. These functions were combined in a full factorial light use efficiency (LUE) model structure, leading to a collection of 5600 distinct LUE models. Each model was optimized against daily GPP and evapotranspiration fluxes from 196 FLUXNET sites and ranked across sites based on a bootstrap approach. The GPP sensitivity to each environmental factor, including CO 2 fertilization, was shown to be significant, and that none of the previously published model structures performed as well as the best model selected. From daily and weekly to monthly scales, the best model's median Nash-Sutcliffe model efficiency across sites was 0.73, 0.79 and 0.82, respectively, but poorer at annual scales (0.23), emphasizing the common limitation of current models in describing the interannual variability of GPP. Although the best global model did not match the local best model at each site, the selection was robust across ecosystem types. The contribution of light saturation and cloudiness to GPP was observed across all biomes (from 23% to 43%). Temperature and W dominates GPP and LUE but responses of GPP to temperature and W are lagged in cold and arid ecosystems, respectively. The findings of this study provide a foundation towards more robust LUE-based estimates of global GPP and may provide a benchmark for other empirical GPP products.
Machine-readable chemical structure representations are foundational in all attempts to harness machine learning for the prediction of reactivities, selectivities, and chemical properties directly from molecular structure. The featurization of discrete chemical structures into a continuous vector space is a critical phase undertaken before model selection, and the development of new ways to quantitatively encode molecules is an active area of research. Here, we highlight the application and suitability of different representations, from expert-guided “engineered” descriptors to automatically “learned” features, in different prediction tasks relevant to organic and organometallic chemistry, where differing amounts of training data are available. These tasks include statistical models of stereo- and enantioselectivity, thermochemistry, and kinetics developed using experimental and quantum chemical data. The use of expert-guided molecular descriptors provides an opportunity to incorporate chemical knowledge, domain expertise, and physical constraints into statistical modeling. In applications to stereoselective organic and organometallic catalysis, where data sets may be relatively small and 3D-geometries and conformations play an important role, mechanistically informed features can be used successfully to obtain predictive statistical models that are also chemically interpretable. We provide an overview of several recent applications of this approach to obtain quantitative models for reactivity and selectivity, where topological descriptors, quantum mechanical calculations of electronic and steric properties, along with conformational ensembles, all feature as essential ingredients of the molecular representations used. Alternatively, more flexible, general-purpose molecular representations such as attributed molecular graphs can be used with machine learning approaches to learn the complex relationship between a structure and prediction target. This approach has the potential to out-perform more traditional representation methods such as “hand-crafted” molecular descriptors, particularly as data set sizes grow. One area where this is particularly relevant is in the use of large sets of quantum mechanical data to train quantitative structure–property relationships. A general approach toward curating useful data sets and training highly accurate graph neural network models is discussed in the context of organic bond dissociation enthalpies, where this strategy outperforms regression using precomputed descriptors. Finally, we describe how graph neural network predictions can be incorporated into mechanistically informed statistical models of chemical reactivity and selectivity. Once trained, this approach avoids the expensive computational overhead associated with quantum mechanical calculations, while maintaining chemical interpretability. We illustrate examples for which fast predictions of bond dissociation enthalpy and of the identities of radicals formed through cleavage of a molecule’s weakest bond are used in simple physical models of site-selectivity and reactivity.
Classical problems in computational physics such as data-driven forecasting and signal reconstruction from sparse sensors have recently seen an explosion in deep neural network (DNN) based algorithmic approaches. However, most DNN models do not provide uncertainty estimates, which are crucial for establishing the trustworthiness of these techniques in downstream decision making tasks and scenarios. In recent years, ensemble-based methods have achieved significant success for the uncertainty quantification in DNNs on a number of benchmark problems. However, their performance on real-world applications remains under-explored. In this work, we present an automated approach to DNN discovery and demonstrate how this may also be utilized for ensemble-based uncertainty quantification. Specifically, we propose the use of a scalable neural and hyperparameter architecture search for discovering an ensemble of DNN models for complex dynamical systems. We highlight how the proposed method not only discovers high-performing neural network ensembles for our tasks, but also quantifies uncertainty seamlessly. This is achieved by using genetic algorithms and Bayesian optimization for sampling the search space of neural network architectures and hyperparameters. Subsequently, a model selection approach is used to identify candidate models for an ensemble set construction. Afterwards, a variance decomposition approach is used to estimate the uncertainty of the predictions from the ensemble. We demonstrate the feasibility of this framework for two tasks — forecasting from historical data and flow reconstruction from sparse sensors for the sea-surface temperature. In conclusion, we demonstrate superior performance from the ensemble in contrast with individual high-performing models and other benchmarks.
The objectives are the following: (1) to develop and demonstrate in-space performance of both passive and active damping systems for suppression of micro-amplitude vibration on an actual application structure and operate despite uncertain dynamics and uncertain disturbance characteristics; and (2) to correlate ground and in-space performance - the performance metric is vibration attenuation. The goals are to achieve vibration suppression equivalent to 5 percent passive damping in selected models and 15 percent active damping in selected modes. Various aspects of this experiment are presented in viewgraph form.
Abstract The global decline of water quality in rivers and streams has resulted in a pressing need to design new watershed management strategies. Water quality can be affected by multiple stressors including population growth, land use change, global warming, and extreme events, with repercussions on human and ecosystem health. A scientific understanding of factors affecting riverine water quality and predictions at local to regional scales, and at sub‐daily to decadal timescales are needed for optimal management of watersheds and river basins. Here, we discuss how machine learning (ML) can enable development of more accurate, computationally tractable, and scalable models for analysis and predictions of river water quality. We review relevant state‐of‐the art applications of ML for water quality models and discuss opportunities to improve the use of ML with emerging computational and mathematical methods for model selection, hyperparameter optimization, incorporating process knowledge into ML models, improving explainablity, uncertainty quantification, and model‐data integration. We then present considerations for using ML to address water quality problems given their scale and complexity, available data and computational resources, and stakeholder needs. When combined with decades of process understanding, interdisciplinary advances in knowledge‐guided ML, information theory, data integration, and analytics can help address fundamental science questions and enable decision‐relevant predictions of riverine water quality.
Traffic emissions significantly impact near-road air quality and public health. This research applies a Bayesian modeling framework to investigate these impacts using high-resolution traffic and air pollutant data from an urban corridor in Columbia, South Carolina. Despite a data collection period truncated by the COVID-19 lockdown, the Bayesian approach successfully identified significant predictors and quantified model uncertainty. Employing Bayesian Model Selection and Averaging enhanced prediction accuracy and evaluated model uncertainty. Findings indicate that higher temperatures and increased moisture levels elevate particulate matter (PM 1.0 , PM 2.5 , PM 10 ) concentrations, while traffic speed significantly affects nitrogen dioxide (NO 2 ) levels. Specifically, higher average traffic speeds (indicative of smoother flow) correspond to lower NO 2 concentrations, suggesting that less congested conditions reduce NO 2 emissions. This study highlights the robustness of Bayesian methods for generating reliable air quality insights even under data-constrained conditions. The findings underscore the importance of traffic flow management (e.g., reducing congestion) for mitigating near-road NO 2 exposure and provide a basis for developing targeted public health strategies.
Contrast-variation small-angle neutron scattering (CV-SANS) is a widely used technique for quantifying hydration water in soft matter systems, but it is predominantly applied in the dilute regime or for systems with a well-defined structure factor. Here, CV-SANS was used to quantify the number of hydration water molecules associating with three water-soluble polymers with different critical solution temperatures and types of water–solute interactions in dilute, semidilute, and concentrated solution through the exploration of novel methods of data fitting and analysis. Multiple SANS fitting workflows with varying levels of model assumptions were evaluated and compared to give insight into SANS model selection. These fitting pathways ranged from general, model-free algorithms to more standard form and structure factor fitting. In addition, Monte Carlo bootstrapping was evaluated as a method to estimate parameter uncertainty through simulation of technical replicates. The most robust fitting workflow for dilute solutions was found to be form factor fitting without CV-SANS ( i.e. polymer in 100% D 2 O). For semidilute and concentrated solutions, while the model-free approach can be mathematically defined for CV-SANS data, the addition of a structure factor imposes physical constraints on the optimization problem, suggesting that the optimal fitting pathway should include appropriate form and structure factor models. The measured hydration numbers were consistent with the number of tightly bound water molecules associated with each monomer unit, and the concentration dependence of the hydration number was largely governed by the chemistry-specific interactions between water and polymer. Polymers with weaker water–polymer interactions ( i.e. those with fewer hydration water molecules) were found to have more bound water at higher concentrations than those with stronger water–polymer interactions due to the increase in the number of forced water–polymer contacts in the concentrated system. This SANS-based method to count hydration water molecules can be applied to polymers in any concentration regime, which will lead to improved understanding of water–polymer interactions and their impact on materials design.
Summary Evolutionary history plays a key role driving patterns of trait variation across plant species. For scaling and modeling purposes, grass species are typically organized into C 3 vs C 4 plant functional types (PFTs). Plant functional type groupings may obscure important functional differences among species. Rather, grouping grasses by evolutionary lineage may better represent grass functional diversity. We measured 11 structural and physiological traits in situ from 75 grass species within the North American tallgrass prairie. We tested whether traits differed significantly among photosynthetic pathways or lineages (tribe) in annual and perennial grass species. Critically, we found evidence that grass traits varied among lineages, including independent origins of C 4 photosynthesis. Using a rigorous model selection approach, tribe was included in the top models for five of nine traits for perennial species. Tribes were separable in a multivariate and phylogenetically controlled analysis of traits, owing to coordination of important structural and ecophysiological characteristics. Our findings suggest grouping grass species by photosynthetic pathway overlooks variation in several functional traits, particularly for C 4 species. These results indicate that further assessment of lineage‐based differences at other sites and across other grass species distributions may improve representation of C 4 species in trait comparison analyses and modeling investigations.
Flows through a transonic diffuser were investigated with the PARC code using five turbulence models to determine the effects of turbulence model selection on flow prediction. Three of the turbulence models were algebraic models: Thomas (the standard algebraic turbulence model in PARC), Baldwin-Lomax, and Modified Mixing Length-Thomas (MMLT). The other two models were the low Reynolds number k-epsilon models of Chien and Speziale. Three diffuser flows, referred to as the no-shock, weak-shock, and strong-shock cases, were calculated with each model to conduct the evaluation. Pressure distributions, velocity profiles, locations of shocks, and maximum Mach numbers in the duct were the flow quantities compared. Overall, the Chien k-epsilon model was the most accurate of the five models when considering results obtained for all three cases. However, the MMLT model provided solutions as accurate as the Chien model for the no-shock and the weak-shock cases, at a substantially lower computational cost (measured in CPU time required to obtain converged solutions). The strong shock flow, which included a region of shock-induced flow separation, was only predicted well by the two k-epsilon models.
This paper reports on an on-going Project to investigate techniques to diagnose complex dynamical systems that are modeled as hybrid systems. In particular, we examine continuous systems with embedded supervisory controllers that experience abrupt, partial or full failure of component devices. We cast the diagnosis problem as a model selection problem. To reduce the space of potential models under consideration, we exploit techniques from qualitative reasoning to conjecture an initial set of qualitative candidate diagnoses, which induce a smaller set of models. We refine these diagnoses using parameter estimation and model fitting techniques. As a motivating case study, we have examined the problem of diagnosing NASA's Sprint AERCam, a small spherical robotic camera unit with 12 thrusters that enable both linear and rotational motion.
The TEAMS model analyzer is a supporting tool developed to work with models created with TEAMS (Testability, Engineering, and Maintenance System), which was developed by QSI. In an effort to reduce the time spent in the manual process that each TEAMS modeler must perform in the preparation of reporting for model reviews, a new tool has been developed as an aid to models developed in TEAMS. The software allows for the viewing, reporting, and checking of TEAMS models that are checked into the TEAMS model database. The software allows the user to selectively model in a hierarchical tree outline view that displays the components, failure modes, and ports. The reporting features allow the user to quickly gather statistics about the model, and generate an input/output report pertaining to all of the components. Rules can be automatically validated against the model, with a report generated containing resulting inconsistencies. In addition to reducing manual effort, this software also provides an automated process framework for the Verification and Validation (V&V) effort that will follow development of these models. The aid of such an automated tool would have a significant impact on the V&V process.
Reliable and accurate estimation of an electric bus’s instantaneous energy consumption is critical in evaluating energy impacts of planning and control of electric bus operations. In this study, we developed machine learning-based long short-term memory (LSTM) and artificial neural network (ANN) models to estimate 1 Hz energy consumption of electric buses based on continuous monitoring data of electric buses in Chattanooga, Tennessee, in 2019 and 2020. We propose a data-partitioning algorithm to separate energy charging and discharging modes before applying data-driven estimation models. Here, a K-fold cross-validation-based model selection process was conducted to identify the optimal model structure and input variables in terms of prediction accuracy. The estimation results show the predicted mean absolute percentage error rates of LSTM and ANN models were 3% and 5%, respectively. We compared the proposed models with existing models in the literature based on the same testing data to demonstrate the predictability of our models.
Recent works exploring deep learning application to dynamical systems modeling have demonstrated that embedding physical priors into neural networks can yield more effective, physically-realistic, and data-efficient models. However, in the absence of complete prior knowledge of a dynamical system's physical characteristics, determining the optimal structure and optimization strategy for these models can be difficult. In this work, we explore methods for discovering neural state space dynamics models for system identification. Starting with a design space of block-oriented state space models and structured linear maps with strong physical priors, we encode these components into a model genome alongside network structure, penalty constraints, and optimization hyperparameters. Demonstrating the overall utility of the design space, we employ an asynchronous genetic search algorithm that alternates between model selection and optimization and obtains accurate physically consistent models of three physical systems: an aerodynamics body, a continuous stirred tank reactor, and a two tank interacting system.