Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Combined selection of the dynamic model and modeling error in nonlinear aeroelastic systems using Bayesian Inference

Here, we report a Bayesian framework for concurrent selection of physics-based models and (modeling) error models. We investigate the use of colored noise to capture the mismatch between the predictions of calibrated models and observational data that cannot be explained by measurement error alone within the context of Bayesian estimation for stochastic ordinary differential equations. Proposed models are characterized by the average data-fit, a measure of how well a model fits the measurements, and the model complexity measured using the Kullback–Leibler divergence. The use of a more complex error models increases the average data-fit but also increases the complexity of the combined model, possibly over-fitting the data. Bayesian model selection is used to find the optimal physical model as well as the optimal error model. The optimal model is defined using the evidence, where the average data-fit is balanced by the complexity of the model. The effect of colored noise process is illustrated using a nonlinear aeroelastic oscillator representing a rigid NACA0012 airfoil undergoing limit cycle oscillations due to complex fluid–structure interactions. Several quasi-steady and unsteady aerodynamic models are proposed with colored noise or white noise for the model error. The use of colored noise improves the predictive capabilities of simpler models.

42 ENGINEERING↗

A comparison of probabilistic generative frameworks for molecular simulations

Generative artificial intelligence is now a widely used tool in molecular science. Despite the popularity of probabilistic generative models, numerical experiments benchmarking their performance on molecular data are lacking. Here, in this work, we introduce and explain several classes of generative models, broadly sorted into two categories: flow-based models and diffusion models. We select three representative models: neural spline flows, conditional flow matching, and denoising diffusion probabilistic models, and examine their accuracy, computational cost, and generation speed across datasets with tunable dimensionality, complexity, and modal asymmetry. Our findings are varied, with no one framework being the best for all purposes. In a nutshell, (i) neural spline flows do best at capturing mode asymmetry present in low-dimensional data, (ii) conditional flow matching outperforms other models for high-dimensional data with low complexity, and (iii) denoising diffusion probabilistic models appear the best for low-dimensional data with high complexity. Our datasets include a Gaussian mixture model and the dihedral torsion angle distribution of the Aib9 peptide, generated via a molecular dynamics simulation. We hope our taxonomy of probabilistic generative frameworks and numerical results may guide model selection for a wide range of molecular tasks.

Artificial intelligence↗

Numerical Analysis of Novel Plate Type Heat Exchanger with Oval-Twisted Channels

The novel heat exchanger (HX) designs implementing geometrical and surface modification combinations are expected to perform better than traditional HX technologies. Applied enhancement techniques seek to achieve (1) higher overall heat transfer performance, (2) increased compactness, and (3) simplified or comparable manufacturability. One innovative enhancement technique is to generate swirling flow vortices induced by channel or tube twisting. The turbulence generated by the twisted cross-section greatly enhances the heat transfer rate with minimal increases in pressure drop. Plate-type HXs and Printed Circuit Heat Exchangers (PCHXs) are compact designs that achieve high heat transfer rates per unit volume by utilizing several small channels, which maximizes the heat transfer surface area between the hot and cold fluids. Currently, advanced manufacturing technologies enable the design and fabrication of compact-type units with complicated channel geometries to achieve the highest performance and meet the compactness criteria of innovative HX technology. The proposed HX design concept combines the compactness of plate-type HXs and twisted channels, which provide additional turbulence and flow swirl enhancement. The plate-type oval-twisted HX (PTOTHX) is a crossflow configuration, with 16 short channels on one side (for hot fluid) and 8 long channels on the perpendicular side (for cold fluid). The inlet plenums have flow guide vanes to redirect flow and produce uniformity across the various flow paths. The compact size and purportedly improved heat transfer performance of the PTOTHX investigated herein prove its viability in various applications. Some notable potential nuclear applications of the PTOTHX include reactor core, spent fuel cooling, and residual heat dissipation. To establish a reference case, circular channels (PTCHX) are also considered in the present study for comparison with (PTOTHX). This paper aims to outline the numerical analysis procedures for determining the viability of the PTOTHX by comparing its heat transfer performance with the PTCHX units. The computational study used STAR-CCM+, a commercial computational fluid dynamics (CFD) code. Sensitivity analysis and model selection studies are conducted to determine the appropriate mesh density and turbulence model to provide the reported results. Numerical analyses comparing the Nusselt number (Nu) of the PTOTHX design with a comparable HX unit, including a circular cross-section and no twisting (PTCHX), show an overall heat transfer performance increase of 29-55% for balanced flow and 29-59% for imbalanced flow. Oval-cross-sectional twisted channels induce swirling flow vortices, enhancing the working fluid's convective heat transfer capabilities.

25 ENERGY STORAGE↗

Field-scale dynamics of planting dates in the US Corn Belt from 2000 to 2020

Crop planting dates are a dynamic feature of agricultural systems that respond to short- and long-term climate signals, crop and cultivar selection, and technology changes. Planting date records are essential for yield gap analyses, accurate crop modeling, and tracking farmer adaptations to weather and climate change. Although planting dates have high variation at local scales due to heterogeneity in farm resources and decision-making, available long-term data on planting dates is largely restricted to aggregated regional statistics or, at best, satellite-derived datasets with limited spatiotemporal extent and at resolutions unable to distinguish individual fields (> 250 m). Here, we generated retrospective annual field-scale (30 m) planting date maps for both maize and soybeans spanning 2000-2020 across a 12 state region in the United States Corn Belt based on Landsat satellite data and a large ground sample of over 28,000 maize and soybean fields. Using training data from 2015-2020 for model selection, we found that planting date predictions improved with harmonic regression of Landsat data and additional annual weather covariates. The preferred random forests model approximately doubled performance compared to a null model based on state median planting dates, capturing 47% of field-level variation for maize (mean absolute error, MAE = 7.4 days) and 44% for soybeans (MAE = 7.5 days) against held-out ground truth test data for 2008-2014. We also evaluated the full 2000-2020 dataset with state agricultural statistics, finding strong agreement with median planting dates for maize (R 2 = 0.76, MAE = 4.4 days) and slightly lower agreement for soybeans (R 2 = 0.65, MAE = 5.4 days) when aggregated to the state level. We then used this new dataset to analyze environmental determinants of planting dates at a finer-scale than previously possible, controlling for unobserved variation at the sub-state district level. We found that during 2000-2020, each standard deviation increase in rainfall delayed planting by ~ 2.5 days, and fields with higher soil productivity ratings tended to be planted earlier. We did not find meaningful trends over the last two decades in planting dates for maize or soybeans, in contrast to trends towards earlier planting dates late last century and predicted for this period in climate adaptation studies. We hypothesize increases in early season rainfall may have inhibited these shifts towards earlier planting. Remotely sensed planting dates will be a useful tool for yield gap analyses, crop simulation modeling, and ongoing assessment of climate adaptation.

54 ENVIRONMENTAL SCIENCES↗

Data-driven analysis and prediction of stable phases for high-entropy alloy design

High-entropy alloys (HEAs) represent a promising class of materials with exceptional structural and functional properties. However, their design and optimization pose challenges due to the large composition-phase space coupled with the complex and diverse nature of the phase formation dynamics. In this study, a data-driven approach that utilizes machine learning (ML) techniques to predict HEA phases and their composition-dependent phases is proposed. By employing a comprehensive dataset comprising 5692 experimental records encompassing 50 elements and 11 phase categories, we compare the performance of various ML models. Our analysis identifies the most influential features for accurate phase prediction. Furthermore, the class imbalance is addressed by employing data augmentation methods, raising the number of records to 1500 in each category, and ensuring a balanced representation of phase categories. The results show that XGBoost and Random Forest consistently outperform the other models, achieving 86% accuracy in predicting all phases. Additionally, this work provides an extensive analysis of HEA phase formers, showing the contributions of elements and features to the presence of specific phases. We also examine the impact of including different phases on ML model accuracy and feature significance. Notably, the findings underscore the need for ML model selection based on specific applications and desired predictions, as feature importance varies across models and phases. This study significantly advances the understanding of HEA phase formation, enabling targeted alloy design and fostering progress in the field of materials science.

36 MATERIALS SCIENCE↗

Assimilation of Multiscale Data into Multifidelity Biogeochemical Models (Final Report)

Quantitative predictions of subsurface processes rely on computational models that capture, with different degrees of fidelity, complex interactions between hydrologic and biogeochemical processes. Molecular- and pore-scale models provide a high-fidelity representation of these processes but are impractical at the field scale. Reduced complexity (e.g., field-scale or data-driven) models sacrifice some degree of fidelity in favor of computational efficiency. Uncertainty and assimilation of data into model predictions, pose a question of model selection: Given a significant difference in computational cost, can a lower-fidelity model be preferable to its higher-fidelity counterpart? Since a predictive model must be computable in reasonable time it has to operate at the field scale, with molecular- and pore-scale data (obtained either from simulations or measurements) determining both the model's structure and parameters. Availability of such multi-resolution data raises a question hitherto undressed in hydrogeology: Can a coarse model dynamically ``learn'' its own structure (e.g., adjust the reaction pathways in its transport module) as more fine-scale data become available during simulations? Our results in multifidelity simulations, data assimilation, and machine learning led us to hypothesize that the answer to these questions is ``yes''. The overarching goal of this project was to confirm this hypothesis by developing scalable, computationally efficient tools for assimilation of data into multiscale models, in which fine-scale data and simulations dynamically inform and autonomously modify coarse-scale models. Our research led to nine manuscripts, three of which have been published and the other six are currently under review.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Verification and validation of developed short-term forecasting models

Recent advancements in machine learning (ML) and artificial intelligence (AI) technologies provide an opportunity for leveraging data-driven algorithms to predict future nuclear power plant (NPP) operating conditions by using recorded plant process data. Successfully implementing these models can lead to cost-reducing, conditioned-based predictive maintenance through optimized maintenance schedules and a reduction of unnecessary maintenance activities. This report discusses the verification and validation of short-term forecasting processes (i.e., data cleaning, feature selection, model optimization, and forecasting) developed in previous reports. The verification and validation (V&V) process demonstrates the expected precision and accuracy when the ML model encounters new datasets from different systems. Shapley additive explanations were used as the primary means of feature selection across these different data set. Individual models were trained for each data set, then validated through a cross-validation procedure. In this report, two different ML models were tasked to predict variables from three different plant process data sets with varying prediction horizons. The results indicate that support vector regression (SVR) outperformed long short-term memory (LSTM) neural networks in regard to each data set and each prediction horizon in this study, but further tuning and optimization could improve long short-term memory results. However, each forecasting model showed reduced performance as the prediction horizon was extended from 1 hour to 1 day ahead. Research is ongoing to evaluate the optimal input variable space, which is based on a given set of process parameters, to further improve forecasting accuracy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Surrogate model evaluation and building energy benchmarking for commercial buildings

Building energy consumption benchmarking involves challenges associated with various energy patterns for different building types; heating, ventilating, and air-conditioning (HVAC) system types; and climates. Given significant variation in energy use patterns, accurate prediction of long-term energy use using surrogate models remains challenging. Multiple linear regression (MLR) is commonly used for building energy benchmarking because of its simple structure; however, it lacks accuracy compared to other black-box models. Although many studies have compared surrogate models and offer guidance on model selection based on metrics, they do not provide detailed analysis on improving the surrogate model accuracy. In this paper, we implement a surrogate model using polynomial ridge regression (i.e., MLR with interaction terms combined with ridge regularization) for small office and retail strip mall buildings across six HVAC system types and all climate zones, for electricity and natural gas in baseline and proposed scenarios. A simulation workflow is developed using OpenStudio TM /EnergyPlus TM to generate simulation data using measures over a wide range of efficiency inputs. Enhancements based on statistical insights are used for improving the model accuracy using filters, input transformations, and change points. Surrogate models achieved average coefficient of variation of the root mean squared error (CVRMSE) values of 2.17, 1.06, 2.05, and 3.26 for proposed electricity, proposed natural gas, baseline electricity, and baseline natural gas, respectively, with enhancements reducing CVRMSE by an average of 14.9% across all combinations. We provide model interpretation via Shapley additive explanations to determine which input variables most influence energy consumption and provide supportive arguments for enhancements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Global centroid moment tensor solutions in a heterogeneous earth: the CMT3D catalogue

SUMMARY For over 40 yr, the global centroid-moment tensor (GCMT) project has determined location and source parameters for globally recorded earthquakes larger than magnitude 5.0. The GCMT database remains a trusted staple for the geophysical community. Its point-source moment-tensor solutions are the result of inversions that model long-period observed seismic waveforms via normal-mode summation for a 1-D reference earth model, augmented by path corrections to capture 3-D variations in surface wave phase speeds, and to account for crustal structure. While this methodology remains essentially unchanged for the ongoing GCMT catalogue, source inversions based on waveform modelling in low-resolution 3-D earth models have revealed small but persistent biases in the standard modelling approach. Keeping pace with the increased capacity and demands of global tomography requires a revised catalogue of centroid-moment tensors (CMT), automatically and reproducibly computed using Green's functions from a state-of-the-art 3-D earth model. In this paper, we modify the current procedure for the full-waveform inversion of seismic traces for the six moment-tensor parameters, centroid latitude, longitude, depth and centroid time of global earthquakes. We take the GCMT solutions as a point of departure but update them to account for the effects of a heterogeneous earth, using the global 3-D wave speed model GLAD-M25. We generate synthetic seismograms from Green's functions computed by the spectral-element method in the 3-D model, select observed seismic data and remove their instrument response, process synthetic and observed data, select segments of observed and synthetic data based on similarity, and invert for new model parameters of the earthquake’s centroid location, time and moment tensor. The events in our new, preliminary database containing 9382 global event solutions, called CMT3D for ‘3-D centroid-moment tensors’, are on average 4 km shallower, about 1 s earlier, about 5 per cent larger in scalar moment, and more double-couple in nature than in the GCMT catalogue. We discuss in detail the geographical and statistical distributions of the updated solutions, and place them in the context of earlier work. We plan to disseminate our CMT3D solutions via the online ShakeMovie platform.

58 GEOSCIENCES↗

A new database website for nuclear level densities

We introduce a new open-access, web-based database (http://nld.ascsn.net), Current Archive of Nuclear Density of Levels (CANDL), that hosts experimental nuclear level density (NLD) datasets from a variety of techniques and energy ranges. Built using the Dash framework in Python, the database is designed to be interactive and user-friendly, allowing researchers to search, visualize, fit, and export NLD data with minimal effort. This resource includes data extracted from evaporation spectra, Oslo method variants, and other experimental techniques that cover excitation energies beyond the neutron resonance region. The database supports on-the-fly fitting with two widely-used phenomenological models—the Constant Temperature (CT) model and the Back-Shifted Fermi Gas (BSFG) model—selected for their simplicity and computational efficiency. Future versions aim to include additional datasets and model types, as well as easy-to-use interfaces to data science techniques. Here, this platform offers a vital tool for the nuclear physics, astrophysics, medicine, and reactor design communities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Improvement of two-phase closure models in CTF using Bayesian inference

Under the Consortium for Advanced Simulation of Light Water Reactors (CASL) program, extensive capabilities have been developed in CTF to analyze light-water reactors (LWRs) for normal operating conditions, departure from nucleate boiling (DNB), and system transients. However, further improvements are required in the modeling and simulation of boiling water reactors (BWRs), which is a focus of the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program. In this work, CTF validation results were used to optimize selected modeling coefficients by calibrating to experimental data using a Bayesian inference approach. Here, calibration studies were conducted to improve (vapor) void fraction prediction without worsening the two-phase pressure drop prediction, as well as to improve the two-phase pressure drop prediction. Calibration was performed for interfacial drag and wall shear models. Surrogates were developed to alleviate the computational expense required for sampling the parameter space using Markov chain Monte Carlo (MCMC). An assessment performed with calibrated models demonstrated an improvement of CTF in its prediction of key parameters such as void fraction and two-phase pressure drop.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Recent progress in microscale modeling of RF sheaths

The microscale properties of RF sheaths in the ion cyclotron range of frequencies (ICRF) are investigated by means of analytical theory, nonlinear fluid and particle-in-cell (PIC) code modeling. Previous work that parametrized RF sheath properties, specifically the RF sheath impedance and the rectified (DC) sheath potential, is generalized to include the effect of net DC current flow through the sheath. Analytical results are presented in the low frequency limit where the displacement current is negligible, and tested against results from a fluid numerical model. Here, it is shown that when the sheath draws DC electron current, the voltage rectification is reduced from the zero current case, and the electron admittance is increased. In separate but related work on the microscale model, selected cases have been simulated with PIC codes to validate, further illuminate and extend fluid model results and their parametrizations. Quantitative agreement in trends for voltage rectification and sheath admittance vs. RF driving voltage is found.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

“Which Projections Do I Use?” Strategies for Climate Model Ensemble Subset Selection Based on Regional Stakeholder Needs

Climate model (or earth system model) projections are increasingly used for climate adaptation planning and impact assessments. As part of this process, many end‐users evaluate a subset of downscaled climate projections without being aware of the implications of downscaling methodology for statistics or event outcomes. Approaches for determining a subset of global climate models to use often focus on values from the raw models, rather than from their downscaled counterparts, in other words assuming that the statistical distribution of the multi‐model ensemble does not change post downscaling. This study demonstrates that a downscaled ensemble will typically retain the change distribution as a raw ensemble, but individual models can differ dramatically post‐downscaling. We recommend that subset‐selection methods account for this possibility and that decision‐relevant downscaled climate projections provide proper descriptions of fitness‐for‐purpose and essential caveats, so that non‐specialists can interpret the results with an appropriate level of confidence.

54 ENVIRONMENTAL SCIENCES↗

Open‐source photovoltaic model pipeline validation against well‐characterized system data

Abstract All freely available plane‐of‐array (POA) transposition models and photovoltaic (PV) temperature and performance models in pvlib‐python and pvpltools‐python were examined against multiyear field data from Albuquerque, New Mexico. The data include different PV systems composed of crystalline silicon modules that vary in cell type, module construction, and materials. These systems have been characterized via IEC 61853‐1 and 61853‐2 testing, and the input data for each model were sourced from these system‐specific test results, rather than considering any generic input data (e.g., manufacturer's specification [spec] sheets or generic Panneau Solaire [PAN] files). Six POA transposition models, 7 temperature models, and 12 performance models are included in this comparative analysis. These freely available models were proven effective across many different types of technologies. The POA transposition models exhibited average normalized mean bias errors (NMBEs) within ±3%. Most PV temperature models underestimated temperature exhibiting mean and median residuals ranging from −6.5°C to 2.7°C; all temperature models saw a reduction in root mean square error when using transient assumptions over steady state. The performance models demonstrated similar behavior with a first and third interquartile NMBEs within ±4.2% and an overall average NMBE within ±2.3%. Although differences among models were observed at different times of the day/year, this study shows that the availability of system‐specific input data is more important than model selection. For example, using spec sheet or generic PAN file data with a complex PV performance model does not guarantee a better accuracy than a simpler PV performance model that uses system‐specific data.

14 SOLAR ENERGY↗

Calibrating Microscopic Car-Following Models for Adaptive Cruise Control Vehicles: Multiobjective Approach

Adaptive cruise control (ACC) vehicles are the first step toward comprehensive vehicle automation. However, the impacts of such vehicles on the underlying traffic flow are not yet clear. Therefore, it is of interest to accurately model vehicle-level dynamics of commercially available ACC vehicles so that they may be used in further modeling efforts to quantify the impact of commercially available ACC vehicles on traffic flow. Importantly, not only model selection but also the calibration approach and error metric used for calibration are critical to accurately model ACC vehicle behavior. In this work, we explore the question of how to calibrate car following models to describe ACC vehicle dynamics. Specifically, we apply a multi-objective calibration approach to understand the tradeoff between calibrating model parameters to minimize speed error vs. spacing error. Three different car-following models are calibrated for data from six vehicles. The results are in line with recent literature and verify that targeting a low spacing error does not compromise the speed accuracy whether the opposite is not true for modeling ACC vehicle dynamics.

33 ADVANCED PROPULSION SYSTEMS↗

Grid Topology Discovery Algorithm Evaluation of Suitability for Utility Deployment (CRADA 606 Final Report)

This work presents the results of a field-informed demonstration aimed at evaluating the practical suitability of a topology discovery algorithm for utility environments. We demonstrated an algorithm that uses a graph-theory-informed state estimation approach for model selection. In collaboration with Survalent and Peninsula Light Co., the algorithm was applied to real feeder models and field measurements from supervisory control and data acquisition (SCADA) and advanced metering infrastructure (AMI) systems to identify the operational topology of a power distribution system. The demonstration assessed the algorithm’s performance under realistic data conditions, including sparse and noisy measurements, and examined its ability to identify the most likely network configurations. The results confirmed that the approach can effectively narrow down feasible topologies, providing operators with improved situational awareness of network status. Key lessons learned emphasize the need for systematic data validation and strategic sensor placement to enhance observability. These insights inform future deployment strategies and guide refinements for broader adoption in utility operations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine learning-enhanced hybrid modeling approach for better identification of a building thermal network model and improved prediction

The gray-box modeling approach, which uses a semi-physical thermal network model, has been widely used in building prediction applications, such as model predictive control (MPC). However, unmeasured disturbances, such as occupants, lighting, and in/exfiltration loads, make it challenging to apply this approach to practical buildings. In this word, we propose a hybrid modeling approach that integrates the gray-box model with a model for unmeasured disturbance. After reviewing several system identification approaches, we systematically designed the unmeasured disturbance model with a model selection process based on statistical tests to make it robust. We generated data based on the building model calibrated by real operational data and then trained the hybrid model for two different weather conditions. The hybrid model approach demonstrates an RMSE reduction of approximately 0.2–0.9 °C and 0.3–2 °C on 1-day ahead temperature prediction compared to the Conventional approach for mild (Berkeley, CA) and cold (Chicago, IL) climates, respectively. In addition, this approach was applied to experimental data obtained from the laboratory building to be used for the MPC application, showing superior prediction performances.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗