Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Intercomparison of flood inundation models across land use types and hydrological flood stages

Flood Inundation Mapping (FIM) model selection is a key operational decision because accurate, rapid mapping underpins early warning and resource allocation. FIM performance is context-dependent and can vary with hydrograph phase, land-use/land-cover (LULC), and the evaluation benchmark. Intercomparison studies typically assess a single near-peak snapshot against one reference dataset. Here, we provide a context-stratified intercomparison across (i) multiple hydrograph phases, (ii) LULC classes, and (iii) benchmark types, for five FIM approaches spanning a wide range of physical complexity and operational cost (TRITON, LISFLOOD-FP, HEC-RAS 2D, ARC-Curve2Flood, and OWP HAND-FIM). We use the Hurricane Matthew flood (2016) in the Neuse River Basin, North Carolina, USA, as a case study. Using high-resolution remote sensing-derived flood inundation maps, hand-labeled points, and building footprints, we assess model skill across two rising and two falling hydrograph limbs and across major LULC types. Results show that model rankings shift systematically across contexts: LISFLOOD-FP ranks highest in three of four flood phases, while TRITON leads during one rising limb phase; LISFLOOD-FP performs best in vegetated areas, whereas HEC-RAS improves relative performance in agricultural and urban areas; and benchmark choice influences conclusions, with LISFLOOD-FP performing best for flooded-building detection in the late falling limb, while TRITON ranks highest against hand-labeled points. We also report representative wall-clock runtimes for each workflow to provide use-case context for operational feasibility. Together, these results offer transferable guidance for model selection and for designing large-scale, benchmark-aware FIM intercomparison studies.

Nikrou, Parvaneh [University of Alabama]↗

Hierarchical, rotation‐equivariant neural networks to select structural models of protein complexes

Abstract Predicting the structure of multi‐protein complexes is a grand challenge in biochemistry, with major implications for basic science and drug discovery. Computational structure prediction methods generally leverage predefined structural features to distinguish accurate structural models from less accurate ones. This raises the question of whether it is possible to learn characteristics of accurate models directly from atomic coordinates of protein complexes, with no prior assumptions. Here we introduce a machine learning method that learns directly from the 3D positions of all atoms to identify accurate models of protein complexes, without using any precomputed physics‐inspired or statistical terms. Our neural network architecture combines multiple ingredients that together enable end‐to‐end learning from molecular structures containing tens of thousands of atoms: a point‐based representation of atoms, equivariance with respect to rotation and translation, local convolutions, and hierarchical subsampling operations. When used in combination with previously developed scoring functions, our network substantially improves the identification of accurate structural models among a large set of possible models. Our network can also be used to predict the accuracy of a given structural model in absolute terms. The architecture we present is readily applicable to other tasks involving learning on 3D structures of large atomic systems.

Eismann, Stephan↗

Selecting Appropriate Model Complexity: An Example of Tracer Inversion for Thermal Prediction in Enhanced Geothermal Systems

Abstract A major challenge in the inversion of subsurface parameters is the ill‐posedness issue caused by the inherent subsurface complexities and the generally spatially sparse data. Appropriate simplifications of inversion models are thus necessary to make the inversion process tractable and meanwhile preserve the predictive ability of the inversion results. In this study, we investigate the effect of model complexity on fracture aperture inversion and thermal performance prediction in a field‐scale EGS model. Principal component analysis was used to map the aperture field to a low‐dimensional latent space. The complexity of the inversion model was quantitatively represented by the percentage of total variance in the original aperture fields preserved by the latent space. Tracer, pressure and flow rate data were used to invert for fracture aperture through an ensemble‐based inversion method, and the inferred aperture field was used to predict thermal performance. With an over‐simplified aperture model, ensemble collapse occurred. The inverted aperture models failed to resolve necessary flow and transport features, leading to a biased thermal performance prediction. A complex aperture model involved excessive features and was prone to overinterpreting the inversion data. Both the tracer/pressure/flow rate data reproduction and thermal prediction showed significant uncertainties, making it difficult to properly estimate long‐term thermal performance. Fortunately, our results indicate that there exists an appropriate model complexity which can simultaneously match inversion data and predict thermal performance with an acceptable uncertainty. The quality of the fit of tracer data appears to be a useful indicator of such an appropriate model complexity.

15 GEOTHERMAL ENERGY↗

Electron collisional excitation cross-section measurements and modeling for select Ni-like to Ge-like gold transitions

We have experimentally determined the electron collisional excitation cross-sections for several 3d→4f and 3d→5f excitations in Ni- to Ge-like Au at energies of ~ 0.4, 1, 2, and 3 keV above threshold energy, E T , for the 3d→4f excitations ( E T ~ 2.5 keV) and ~ 0.2, 1, and 2 keV above threshold energy for the 3d→5f excitations ( E T ~ 3.3 keV). The cross-section measurements are possible by using the GSFC micro-calorimeter to record emission spectra from beam plasmas created in the Livermore EBIT-I electron beam ion trap. The cross-sections are experimentally determined from the ratio of the measured intensities of the collisionally excited lines to the intensities of the radiative recombination lines in monoenergetic electron distribution EBIT-I plasmas. The effects of polarization and Auger processes in the beam plasmas are accounted for in the cross-section determination. Experimentally determined cross-sections are compared with those from HULLAC, DWS, and FAC calculations. Finally, the measurements exhibit significant differences with the calculations of these excitation cross-sections.

74 ATOMIC AND MOLECULAR PHYSICS↗

False Indications of Dose-Response Nonlinearity in Large Epidemiologic Cancer Radiation Cohort Studies; A Simulation Exercise

This study explores the likely prevalence of false indications of dose-response nonlinearity in large epidemiologic cancer radiation cohort studies (A-bomb survivors, INWORKS, Techa River). Reasons: Increasing numbers of tests of nonlinearity are being made in studies. Hypothesized nonlinear dose-response models have been justified to policy makers by analyses that rely in part on isolated findings that could be statistical fluctuations. After removing dose nonlinearity (linearization) by adjusting person-years of observation at each dose category, indications of nonlinearity, necessarily false, were counted in 5,000 randomized replications of six datasets. The average frequency of any false positive for five indicators of nonlinearity tested against a linear null was roughly 25% in Monte Carlo simulations per study, consistent with binomial calculations, increasing to ~50% within 6 studies assessed. Comparable frequencies were found using Akaike's information criterion (AIC) for model selection or multi-model averaging. False above-zero threshold doses were found more than 50% of the time, averaging to 0.05 Gy, consistent with findings in the 6 studies. Such bias, uncorrected, could distort meta-analyses of multiple studies, because meta-analyses can incorporate high P value findings. AIC-based correction for the extra threshold parameter lowered these false occurrences to 8 to 19%. Given the simulation rates, the possibility of false positives might be noted when isolated findings of nonlinearity are discussed in a regulatory context. When reporting a threshold dose with a P value > 0.05, it would be informative to note the expected high false prevalence rate due to bias.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Support interference of wind tunnel models: A selective annotated bibliography

This bibliography, with abstracts, consists of 143 citations arranged in chronological order by dates of publication. Selection of the citations was made for their relevance to the problems involved in understanding or avoiding support interference in wind tunnel testing throughout the Mach number range. An author index is included.

Tuttle, M. H.↗

Support interference of wind tunnel models: A selective annotated bibliography

This bibliography, with abstracts, consists of 143 citations arranged in chronological order by dates of publication. Selection of the citations was made for their relevance to the problems involved in understanding or avoiding support interference in wind tunnel testing throughout the Mach number range. An author index is included.

Tuttle, M. H.↗

Low-frequency variability and wavenumber selection in models with zonally symmetric forcing

The authors consider a two-layer quasigeostrophic model with linear surface drag and forcing that relaxes to a zonal baroclinically unstable equilibrium state consisting of a meridionally confined temperature gradient. It is observed that the most energetic wave in the time-mean climate has near zero frequency and is not driven by upscale nonlinear energy transfers. This wave has a zonal scale near the long-wave cutoff of the equilibrium state, and its energy balance is mainly between baroclinic generation and dissipation. This maintenance mechanism is different from that suggested by beta-plane, two-dimensional, and quasigeostrophic turbulence arguments and may be relevant to the dynamics of zonally asymmetric low-frequency variability in the atmosphere, particularly in the Southern Hemisphere.

Whitaker, Jeffrey S.↗

Generative diffusion model surrogates for mechanistic agent-based biological models

Mechanistic, multicellular, agent-based models are commonly used to investigate tissue, organ, and organism-scale biology at single-cell resolution. The Cellular-Potts Model (CPM) is a powerful and popular framework for developing and interrogating these models. CPMs become computationally expensive at large space- and time- scales making application and investigation of developed models difficult. Surrogate models may allow for the accelerated evaluation of CPMs of complex biological systems. However, the stochastic nature of these models means each set of parameters may give rise to different model configurations, complicating surrogate model development. In this work, we leverage denoising diffusion probabilistic models (DDPMs) to train a generative AI surrogate of a CPM used to investigate in vitro vasculogenesis. We describe the use of an image classifier to learn the characteristics that define unique areas of a 2-dimensional parameter space. We then apply this classifier to aid in surrogate model selection and verification. Our CPM model surrogate generates model configurations 20,000 timesteps ahead of a reference configuration and demonstrates approximately a 22x reduction in computational time as compared to native code execution. Our work represents a step towards the implementation of DDPMs to develop digital twins of stochastic biological systems.

97 MATHEMATICS AND COMPUTING↗

Evaluating Climate Models with the CLIVAR 2020 ENSO Metrics Package

El Niño–Southern Oscillation (ENSO) is the dominant mode of interannual climate variability on the planet, with far-reaching global impacts. It is therefore key to evaluate ENSO simulations in state-of-the-art numerical models used to study past, present, and future climate. Recently, the Pacific Region Panel of the International Climate and Ocean: Variability, Predictability and Change (CLIVAR) Project, as a part of the World Climate Research Programme (WCRP), led a community-wide effort to evaluate the simulation of ENSO variability, teleconnections, and processes in climate models. The new CLIVAR 2020 ENSO metrics package enables model diagnosis, comparison, and evaluation to 1) highlight aspects that need improvement; 2) monitor progress across model generations; 3) help in selecting models that are well suited for particular analyses; 4) reveal links between various model biases, illuminating the impacts of those biases on ENSO and its sensitivity to climate change; and to 5) advance ENSO literacy. By interfacing with existing model evaluation tools, the ENSO metrics package enables rapid analysis of multipetabyte databases of simulations, such as those generated by the Coupled Model Intercomparison Project phases 5 (CMIP5) and 6 (CMIP6). The CMIP6 models are found to significantly outperform those from CMIP5 for 8 out of 24 ENSO-relevant metrics, with most CMIP6 models showing improved tropical Pacific seasonality and ENSO teleconnections. Only one ENSO metric is significantly degraded in CMIP6, namely, the coupling between the ocean surface and subsurface temperature anomalies, while the majority of metrics remain unchanged.

54 ENVIRONMENTAL SCIENCES↗

Accuracy Enhancement of Nuclear Power Plant Simulators Utilizing High Accuracy Simulation Predictions

More recently, reactor core simulators for core designs associated with commercial nuclear power plants that utilize what is believed to be higher fidelity models have been developed. Features such as neutronics models that utilize transport equation solvers with fine spatial meshes and many energy-groups, thermal-hydraulic models that utilize sub-channel solvers with fine spatial mesh and capable of treating a wide range of fluid conditions, and fuel-coolant chemistry interaction models capable of treating CRUD deposition are to be found in these higher fidelity core simulators. These reactor core simulators require access to higher performance computers, characterized by many processors, cores and large memory. So associated with utilization of these simulators is access to high performance computers and ability to accommodate in one’s workflow longer execution times. By contrast, currently used core simulators by the nuclear industry can execute on engineering workstations and have execution times of seconds to minutes. The desirability for having short execution times is not only desired for support of time critical tasks but supports the mental process of decision making by engineers. The goal of the work reported upon here has the objective of retaining the fidelity of higher fidelity models while retaining the ability to utilize engineering workstations. Beyond the core simulator goal, additional goals of this work include incorporating the just described core simulator capability into a Nuclear Steam Supply System (NSSS) simulator, and to incorporate the resulting capability into an environment supportive of design and operational decision making associated with nuclear power stations. The model selected for the core neutronics model is the NESTLE code, for the core thermal-hydraulic model is the CTF code utilizing coarse mesh, and for the NSSS model is the RELAP5-3D code. WSC’s proprietary 3KEYMASTERTM platform is being used to provide software coupling, user interface, visualization, and reporting. The NESTLE core neutronics simulator was first integrated with the CTF core thermal-hydraulic simulator using CTF developed communication commands which are also used for CTF to communicate with RELAP5-3D under WSC’s proprietary 3KEYMASTERTM platform. To assure NESTLE prediction consistency with higher fidelity core neutronic simulators, buffer codes have been created to automatically generate from output files written by the VERA core simulator the NESTLE nodal neutronic parameter’ library, geometry, and pin-power reconstruction input files, thereby avoiding a number of challenges associated with utilizing lattice physics codes and providing consistency with VERA predictions. To treat absorber rod effects a multi-set library is utilized, where a set refers to a specific absorber rod fully inserted pattern. A coarse spatial mesh CTF model was developed with features added that support using CTF as envisioned in the engineering quality simulator. A hybrid meshing approach was implemented to allow for automated construction of models with mixed levels of refinement. Specifically, a core model could resolve some assemblies at a nodal level (4 subchannels per assembly) and others at a pin-resolution (one subchannel per coolant subchannel in the assembly). The intention is that this will allow for better resolution of limiting conditions such as DNBR and PCT, which are based on local rod and subchannel conditions. Further development was done of features that enhance the capabilities for the envisioned engineering quality simulator that has been developed, but now for RELAP-3D. The RELAP5-3D code development includes ability to model more than 999 components and the addition of the cross-channels turbulence mixing model and the void drift model that are implemented in CTF, aiming to achieve closer prediction agreement of the two codes for transient simulations, specifically, more accurate matches of the overall mass, momentum, and energy exchanges of both the liquid and gas phases between the neighboring core assemblies. Graphics were also developed for the Instructor Station for this project under WSC’s proprietary 3KEYMASTERTM platform to facilitate design and operational decision making.

42 ENGINEERING↗

Using automated machine learning for the upscaling of gross primary productivity

Estimating gross primary productivity (GPP) over space and time is fundamental for understanding the response of the terrestrial biosphere to climate change. Eddy covariance flux towers provide in situ estimates of GPP at the ecosystem scale, but their sparse geographical distribution limits larger-scale inference. Machine learning (ML) techniques have been used to address this problem by extrapolating local GPP measurements over space using satellite remote sensing data. However, the accuracy of the regression model can be affected by uncertainties introduced by model selection, parameterization, and choice of explanatory features, among others. Recent advances in automated ML (AutoML) provide a novel automated way to select and synthesize different ML models. In this work, we explore the potential of AutoML by training three major AutoML frameworks on eddy covariance measurements of GPP at 243 globally distributed sites. We compared their ability to predict GPP and its spatial and temporal variability based on different sets of remote sensing explanatory variables. Explanatory variables from only Moderate Resolution Imaging Spectroradiometer (MODIS) surface reflectance data and photosynthetically active radiation explained over 70 % of the monthly variability in GPP, while satellite-derived proxies for canopy structure, photosynthetic activity, environmental stressors, and meteorological variables from reanalysis (ERA5-Land) further improved the frameworks' predictive ability. We found that the AutoML framework Auto-sklearn consistently outperformed other AutoML frameworks as well as a classical random forest regressor in predicting GPP but with small performance differences, reaching an r 2 of up to 0.75. We deployed the best-performing framework to generate global wall-to-wall maps highlighting GPP patterns in good agreement with satellite-derived reference data. This research benchmarks the application of AutoML in GPP estimation and assesses its potential and limitations in quantifying global photosynthetic activity.

54 ENVIRONMENTAL SCIENCES↗

General Multifidelity Surrogate Models: Framework and Active-Learning Strategies for Efficient Rare Event Simulation

Estimating the probability of failure for complex real-world systems using high-fidelity computational models is often prohibitively expensive, especially when the probability is small. Exploiting low-fidelity models can make this process more feasible, but merging information from multiple low-fidelity and high-fidelity models poses several challenges. Here, this paper presents a robust multi-fidelity surrogate modeling strategy in which the multi-fidelity surrogate is assembled using an active learning strategy using an on-the-fly model adequacy assessment set within a subset simulation framework for efficient reliability analysis. The multi-fidelity surrogate is assembled by first applying a Gaussian process correction to each low-fidelity model and assigning a model probability based on the model's local predictive accuracy and cost. Three strategies are proposed to fuse these individual surrogates into an overall surrogate model based on model averaging and deterministic/stochastic model selection. The strategies also dictate which model evaluations are necessary. No assumptions are made about the relationships between low-fidelity models, while the high-fidelity model is assumed to be the most accurate and most computationally expensive model. Through two analytical and two numerical case studies, including a case study evaluating the failure probability of Tristructural isotropic-coated (TRISO) nuclear fuels, the algorithm is shown to be highly accurate while drastically reducing the number of high-fidelity model calls (and hence computational cost).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗