Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

CFD Simulations of Lower Plenum Mixing

Review of model development and validation performed in the Advanced Reactor Technologies (ART) program for thermal mixing at the outlet of High Temperature Gas Reactors (HTGRs). Understanding the mixing that occurs in the lower plenum in an HTGR is necessary to facilitate design improvements and to perform reactor safety analysis. Numerical models are one possible approach to gain a better understanding of mixing in the lower plenum. Given the complexity of the geometry and the intense mixing present, it is important to perform validation of numerical models. Three models have been developed during FY2025: a porous media with Pronghorn, a Reynolds Averaged Navier Stokes (RANS) with STAR-CCM+, and a Large Eddy Simulation (LES) with NekRS. The reference facility is a scaled-down version of the lower plenum of the High Temperature Gas-Cooled Reactor - Pebble-bed Module (HTR-PM) demonstration reactor. Preliminary results of the porous media and the RANS shows general good agreement against experimental benchmark data. Future work will leverage high-fidelity results obtained through LES to guide model selection and improvements to the lower-fidelity models, with particular attention to the Pronghorn porous media.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Sensitivity of the Assimilated Ozone in the UTLS to Model and Data Selection Changes

This presentation will discuss the sensitivity of assimilated ozone fields in the upper troposphere and lower stratosphere (UTLS) to a number of factors, focusing mainly on aspects of data selection and the prediction model. This is important, because assimilation represents an attempt to construct our best estimates of the true ozone field; however, inaccuracies in the UTLS ozone distribution translate into an uncertainty in factors such as the calculated radiative forcing of climate or the inferred stratosphere-troposphere exchange (STE) of ozone. The 3D ozone data assimilation system, from NASA's Global Modeling and Assimilation Office (GMAO), combines observations of total ozone column and stratospheric profiles with predictions from an off-line, parameterized chemistry and transport model (pCTM) to produce six-hourly, global analyses. The first experiments discussed assimilate ozone retrievals from the Earth-Probe Total Ozone Mapping Spectrometer (EPTOMS) and stratospheric profiles from the Solar Backscatter UltraViolet/2 (SBUV/2) instrument. The SBUV/2 ozone data have a coarse vertical resolution, with increased uncertainty below the ozone maximum, and TOMS provides only total ozone columns. Thus, the assimilated ozone profiles in the UTLS region are only weakly constrained by the incoming SBUV and TOMS data. Consequently, the assimilated ozone distribution should be sensitive to changes in inputs to the statistical analysis scheme. Sensitivity studies have been conducted to examine the responses to TOMS and SBUV/2 data selection, modifications of the forecast and observation error covariance models, and the model formulation (turning off chemistry or using different wind analyses in the pCTM). The second set of experiments includes an additional data type: ozone retrieved from infrared limb-emission by MIPAS on Envisat. These data offer not only improved vertical resolution in the stratosphere, but also give measurements in the polar night. Comparisons of the assimilated ozone fields from both sets of experiments with independent observations, primarily ozone sondes, are used to determine the impact of each of these changes. It is shown that many of the changes have a significant impact on the UTLS ozone estimates. Implications for interpretation of STE and radiative forcing of climate are discussed.

Pawson, Steven↗

Development of Building Design Optimization Methodology: Residential Building Applications

Building design optimization is a highly complex problem, requiring long computational running processes because of the many options that exist when a building is being designed. This paper introduces an integrated approach through which to perform this optimization within an acceptable time frame. The approach includes the methods of variable selection, model simplification, and a sequential optimization process. Using singular value decomposition, a large number of design variables is reduced to a smaller subset that can be solved more quickly through the optimization algorithm. To expedite the variable selection process, a modeling approach that quickly simulates annual energy consumption was developed to replace full annual energy simulations. The developed methodology was applied to two residential buildings in the US, and the results are discussed herein. To assess the accuracy of the integrated optimization methodology, the optimized life cycle costs are compaa variables demonstrating the strongest contributions in the optimization study were identified. The proposed methodology significantly shortened the time requirements for the optimization processes of the two case studies by 74% and 84%; the optimized life cycle costs were within 0.05% and 0.06%, respectively, of the optimum point.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Improvement of the $\mathrm{BISON U_3Si_2}$ modeling capabilities based on multiscale developments to modeling fission gas behavior

Uranium silicide (U 3 Si 2 ) is a concept explored as a potential alternative to UO 2 fuel used in light water reactors (LWRs) since it may improve accident tolerance and economics due to its higher thermal conductivity and increased uranium density. U 3 Si 2 has been previously used in research reactors in the form of dispersion fuel, but operated at lower temperatures than commercial LWRs. The research reactor data illustrated that significant gaseous swelling occurs as the fuel burnup increases. Therefore, it is imperative to understand the fission gas behavior of U 3 Si 2 under higher temperature LWR operating conditions. In this work, molecular dynamics and phase-field modeling techniques are used to reduce the uncertainty in select modeling assumptions made in developing the fission gas behavior model for U 3 Si 2 in the BISON fuel performance code. These lower length scale informed models are then utilized in the validation of BISON U 3 Si 2 modeling capabilities to simulate the ATF-1 experiments irradiated in the Advanced Test Reactor (ATR). Sensitivity analysis (SA) and uncertainty quantification (UQ) are included as part of the validation process to identify where further experiments and lower length scale modeling would be beneficial. Here, the multiscale modeling approach utilized in this work can be applied to new fuel concepts being explored for both LWRs and advanced reactors (e.g., uranium nitride, uranium carbide).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Full-sky Models of Galactic Microwave Emission and Polarization at Subarcminute Scales for the Python Sky Model

Polarized foreground emission from the Galaxy is one of the biggest challenges facing current and upcoming cosmic microwave background (CMB) polarization experiments. We develop new models of polarized Galactic dust and synchrotron emission at CMB frequencies that draw on the latest observational constraints; that employ the “polarization fraction tensor” framework to couple intensity and polarization in a physically motivated way; and that allow for stochastic realizations of small-scale structure at subarcminute angular scales currently unconstrained by full-sky data. We implement these models into the publicly available Python Sky Model (PySM) software and additionally provide PySM interfaces to select models of dust and CO emission from the literature. We characterize the behavior of each model by quantitatively comparing it to observational constraints in both maps and power spectra, demonstrating an overall improvement over previous PySM models. Finally, we synthesize models of the various Galactic foreground components into a coherent suite of three plausible microwave skies that span a range of astrophysical complexity allowed by current data. Author contributions to this paper can be found at the end of this work.

Group, The Pan-Experiment Galactic Science↗

Computational modeling of multispectral remote sensing systems: Background investigations

A computational model of the deterministic and stochastic process of remote sensing has been developed based upon the results of the investigations presented. The model is used in studying concepts for improving worldwide environment and resource monitoring. A review of various atmospheric radiative transfer models is presented as well as details of the selected model. Functional forms for spectral diffuse reflectance with variability introduced are also presented. A cloud detection algorithm and the stochastic nature of remote sensing data with its implications are considered.

Aherron, R. M.↗

Project trades model for complex space missions

A Project Trades Model (PTM) is a collection of tools/simulations linked together to rapidly perform integrated system trade studies of performance, cost, risk, and mission effectiveness. An operating PTM captures the interactions between various targeted systems and subsystems through an exchange of computed variables of the constituent models. Selection and implementation of the order, method of interaction, model type, and envisioned operation of the ensemble of tools rpresents the key system engineering challenge of the approach. This paper describes an approach to building a PTM and using it to perform top-level system trades for a complex space mission. In particular, the PTM discussed here is for a future Mars mission involving a large rover.

web services↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Parameter Estimation for Compact Binary Coalescence Signals with the First Generation Gravitational-Wave Detector Network

Compact binary systems with neutron stars or black holes are one of the most promising sources for ground-based gravitational-wave detectors. Gravitational radiation encodes rich information about source physics; thus parameter estimation and model selection are crucial analysis steps for any detection candidate events. Detailed models of the anticipated waveforms enable inference on several parameters, such as component masses, spins, sky location and distance, that are essential for new astrophysical studies of these sources. However, accurate measurements of these parameters and discrimination of models describing the underlying physics are complicated by artifacts in the data, uncertainties in the waveform models and in the calibration of the detectors. Here we report such measurements on a selection of simulated signals added either in hardware or software to the data collected by the two LIGO instruments and the Virgo detector during their most recent joint science run, including a blind injection where the signal was not initially revealed to the collaboration. We exemplify the ability to extract information about the source physics on signals that cover the neutron-star and black-hole binary parameter space over the component mass range 1M25M and the full range of spin parameters. The cases reported in this study provide a snapshot of the status of parameter estimation in preparation for the operation of advanced detectors.

Aasi, J.↗

Physics vs structure: A systematic benchmark of learning strategies for multi-zone building thermal dynamics

Recent advances in physics-informed and data-driven machine learning promise improved thermal models for advanced building control, yet there is limited quantitative evidence on when added physics structure and architectural complexity are beneficial. Here, this work presents a systematic benchmark of five representative system identification methods for modeling multi-zone building thermal dynamics: linear state-space models, multi-layer perceptrons, neural state-space models, neural ordinary differential equations, and physically-consistent neural networks. The methods are evaluated across multiple data regimes and zone coupling strategies. Using a high-fidelity multi-zone commercial building emulator, we examine short-term and long-term prediction accuracy, computational efficiency, and ease of development. Our results reveal critical trade-offs between prediction performance, model complexity, and physical consistency. We demonstrate that decoupled, nonlinear black-box models consistently outperform coupled physics-constrained architectures in both predictive accuracy and out-of-distribution robustness in majority of the test cases for the building type considered in the study. Our findings quantify the cost of complexity in building thermal modeling and provide concrete, actionable, scenario-based guidelines for selecting model classes for control-oriented applications.

Building thermal modeling↗

Singling out modified gravity parameters and data sets reveals a dichotomy between Planck and lensing

ABSTRACT An important route to testing general relativity (GR) at cosmological scales is usually done by constraining modified gravity (MG) parameters added to the Einstein perturbed equations. Most studies have analysed so far constraints on pairs of MG parameters, but here, we explore constraints on one parameter at a time while fixing the other at its GR value. This allows us to analyse various models while benefiting from a stronger constraining power from the data. We also explore which specific data sets are in tension with GR. We find that models with (μ = 1, η) and (μ, η = 1) exhibit a 3.9σ and 3.8σ departure from GR when using Planck18 + Supernovae type Ia (SNe) + Baryon acoustic oscillation (BAO), while (μ, η) shows a tension of 3.4σ. We find no tension with GR for models with the MG parameter Σ fixed to its GR value. Using a Bayesian model selection analysis, we find that some one-parameter MG models are moderately favoured over ΛCDM when using all data set combinations except Planck cosmic microwave background lensing and dark energy survey data. Namely, Planck18 shows a moderate tension with GR that only increases when adding any combination of redshift space distortion, SNe, or BAO. However, adding lensing diminishes or removes these tensions, which can be attributed to the ability of lensing in constraining the MG parameter Σ. The two overall groups of data sets are found to have a dichotomy when performing consistency tests with GR, which may be due to systematic effects, lack of constraining power, or modelling. These findings warrant further investigation using more precise data from ongoing and future surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Multi‐Model Ensembles in Ecosystem Modeling: Challenges and Best Practices for Decision‐Making

Ecosystem models are increasingly central to the decision-making for environmental policy, conservation planning, and climate-related investments. Yet, the growing reliance on Multi-Model Ensembles (MMEs) of ecosystem models by practitioners and policymakers, sometimes under tight timelines and imperfect information, has frequently outpaced the scientific rigor required to ensure ensemble reliability. Here, MMEs refer to approaches that combine targeted predictions from multiple models with the expectation of improving robustness and quantifying predictive uncertainty. Poorly designed MMEs may create a false sense of confidence and lead to suboptimal policy and market decisions. This perspective argues that robust decision-making-relevant MMEs must be grounded on two pillars: (1) rigorous Model Intercomparison Projects (MIPs), which identify inter-model agreement and disagreement, characterize model uncertainties, and evaluate robustness with observationally based benchmarks—MIPs' diagnostic evaluation is so critical that it must be needed to drive MME's decision in model selection and weighting, especially when only a limited number of models available; and (2) co-design by both stakeholders and scientists to ensure that scenarios, metrics and uncertainty requirements provide decision-relevant information. Building upon the past success and lessons from the existing MIPs-MMEs efforts (e.g., climate/Earth system/crop), we derived the theoretical basis for MMEs, addressed their specific challenges in ecosystem modeling, and highlighted proper consideration of model numbers and diversity, risk of model inter-dependence, effective calibration of model parameters, possible overdue of some ecosystem model development, critical roles of open benchmark data across a wide range of conditions, and suggested use of Artificial Intelligence to support MIPs-MMEs. We highlighted the under-recognized opportunity for MIPs and MMEs to drive scientific progress and innovation through identifying better performing models, systematic benchmarking, feedback loops, and targeted model improvement. By following actionable best practice guidelines, MMEs can evolve from ad hoc aggregation of models into a trusted backbone of environmental policy and decision-making.

ecosystem modeling↗

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa↗

Models, data, and scripts associated with “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning”

This data package is associated with the publication “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning’’ submitted to the Journal of Geophysical Research: Machine Learning and Computation (Scheibe et al. 2024). River sediment respiration observations are expensive and labor intensive to obtain and there is no physical model for predicting this quantity. The Worldwide Hydrobiogeochemisty Observation Network for Dynamic River Systems (WHONDRS) observational data set (Goldman et al.; 2020) is used to train machine learning (ML) models to predict respiration rates at unsampled sites. This repository archives training data, ML models, predictions, and model evaluation results for the purposes of reproducibility of the results in the associated manuscript and community reuse of the ML models trained in this project. One of the key challenges in this work was to find an optimum configuration for machine learning models to work with this feature-rich (i.e. 100+ possible input variables) data set. Here, we used a two-tiered approach to managing the analysis of this complex data set: 1) a stacked ensemble of ML models that can automatically optimize hyperparameters to accelerate the process of model selection and tuning and 2) feature permutation importance to iteratively select the most important features (i.e. inputs) to the ML models. The major elements of this ML workflow are modular, portable, open, and cloud-based, thus making this implementation a potential template for other applications. This data package is associated with the GitHub repository found at Please see the file level metadata (flmd; “sl-archive-whondrs_flmd.csv”) for a list of all files contained in this data package and descriptions for each. Please see the data dictionary (dd; “sl-archive-whondrs_dd.csv”) for a list of all column headers contained within comma separated value (csv) files in this data package and descriptions for each. The GitHub repository is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning models trained on the data in “input_data”; (3) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; (4) “examples” contains the visualization of the results in this repository including plotting scripts for the manuscript (e.g., model evaluation, FPI results) and scripts for running predictions with the ML models (i.e., reusing the trained ML models); (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. Furthermore, depending on the number of features used to train the ML models, the preprocessing and postprocessing scripts, and their intermediate results, can also be different branch-to-branch. The “main-*” branches are meant to be starting points (i.e. trunks) for each model branch (i.e. sprouts). Please see the Branch Navigation section in the top-level README.md in the GitHub repository for more details. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please the top-level README.md in the GitHub repository for more details on the automation.

13C↗

Energy Storage Valuation: A Review of Use Cases and Modeling Tools

An enticing prospect that drives adoption of energy storage systems (ESS) is its ability to be used in a diverse set of use cases and the potential to take advantage of multiple unique value streams. The Energy Storage Grand Challenge (ESGC) technology development pathways for storage technologies draw from a set of use cases in the electrical power system, each with their own specific cost and performance needs. In addition to the need for cost and performance improvements for storage technologies, there a need for robust valuation methods to enable effective policy, investment, business models, and resource planning. There are numerous storage valuation tools available to the public, many of which can analyze the value of an ESS project with inputs and characteristics that reflect a specific storage use case. To effectively reach ESS stakeholders that may be interested in learning about valuation models, this report will draw from publicly available tools developed by the Department of Energy (DOE) and frame their functionalities and capabilities within the context of three distinct use case families. This report examines three of the ESGC use case families in depth and provides a methodology in which interested stakeholders can determine which DOE modeling tool is best suited to value ESS for their specific case. The high-level objectives for this report include: (1) Provide specific sub use-cases for each use case family for further characterization; (2) Provide technical parameters and relevant data for three example use cases that could be used in a valuation tool; (3) Identify a list of publicly available DOE tools that can provide energy storage valuation insights for ESS use case stakeholders; (4) Provide information on the capabilities and different options in each modeling tool; (5) Make conclusions on which are best suited for valuing certain functional/performance requirements and which tools might be applicable to other use cases; and (6) Show the methodology that informs a Model Selection Platform (MSP) framework that educates stakeholders on different DOE models and provides a streamlined way to choose the right model that most closely matches their needs.

25 ENERGY STORAGE↗

Quantification of Type I Interferon Inhibition by Viral Proteins: Ebola Virus as a Case Study

Type I interferons (IFNs) are cytokines with both antiviral properties and protective roles in innate immune responses to viral infection. They induce an antiviral cellular state and link innate and adaptive immune responses. Yet, viruses have evolved different strategies to inhibit such host responses. One of them is the existence of viral proteins which subvert type I IFN responses to allow quick and successful viral replication, thus, sustaining the infection within a host. We propose mathematical models to characterise the intra-cellular mechanisms involved in viral protein antagonism of type I IFN responses, and compare three different molecular inhibition strategies. We study the Ebola viral protein, VP35, with this mathematical approach. Approximate Bayesian computation sequential Monte Carlo, together with experimental data and the mathematical models proposed, are used to perform model calibration, as well as model selection of the different hypotheses considered. Finally, we assess if model parameters are identifiable and discuss how such identifiability can be improved with new experimental data.

59 BASIC BIOLOGICAL SCIENCES↗

Calculation of atmospheric loss from microwave radiometric noise temperature measurements

Microwave propagation loss in the atmosphere can be inferred from microwave radiometric noise temperature measurements. The relevant equations are given and a derivation and calculation is made assuming various physical models. Comparison is made with the commonly used lumped element atmospheric model (isothermal and uniform loss) and the model with linear temperature and exponential loss distributions. The results are useful for estimating the integral inversion differences due to the model selection. This indicates that the commonly used lumped element atmospheric model is a very good approximation with judicious choice of the effective physical temperature. For the worst case comparison, the lumped element model agrees with the variable parameter model within 0.2 dB up to a propagation loss of 3 dB.

Stelzried, C.↗

Implementation of and Ada real-time executive: A case study

Current Ada language implementations and runtime environments are immature, unproven and are a key risk area for real-time embedded computer system (ECS). A test-case environment is provided in which the concerns of the real-time, ECS community are addressed. A priority driven executive is selected to be implemented in the Ada programming language. The model selected is representative of real-time executives tailored for embedded systems used missile, spacecraft, and avionics applications. An Ada-based design methodology is utilized, and two designs are considered. The first of these designs requires the use of vendor supplied runtime and tasking support. An alternative high-level design is also considered for an implementation requiring no vendor supplied runtime or tasking support. The former approach is carried through to implementation.

Laird, James D.↗