Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data base for model validation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Epidural anesthesia needle guidance by forward-view endoscopic optical coherence tomography and deep learning

Epidural anesthesia requires injection of anesthetic into the epidural space in the spine. Accurate placement of the epidural needle is a major challenge. To address this, we developed a forward-view endoscopic optical coherence tomography (OCT) system for real-time imaging of the tissue in front of the needle tip during the puncture. We tested this OCT system in porcine backbones and developed a set of deep learning models to automatically process the imaging data for needle localization. A series of binary classification models were developed to recognize the five layers of the backbone, including fat, interspinous ligament, ligamentum flavum, epidural space, and spinal cord. The classification models provided an average classification accuracy of 96.65%. During puncture, it is important to maintain a safe distance between the needle tip and the dura mater. Regression models were developed to estimate that distance based on the OCT imaging data. Based on the Inception architecture, our models achieved a mean absolute percentage error of 3.05% ± 0.55%. Overall, our results validated the technical feasibility of using this novel imaging strategy to automatically recognize different tissue structures and measure the distances ahead of the needle tip during the epidural needle placement.

60 APPLIED LIFE SCIENCES↗

Inter-well connectivity detection in CO 2 WAG projects using statistical recurrent unit models

Routine well-wise injection and production measurements contain significant information on subsurface structure and properties. Data-driven technology that interprets surface data into subsurface structure or properties can assist operators in making informed decisions by providing a better understanding of field assets. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO 2 EOR projects utilizing the water-alternating-gas (WAG) process. SRU is a special type of recurrent neural network (RNN) that allows for better characterization of temporal trends, by learning various statistics of the input at different time scales. In our application, the complete states (injection rate, pressure and cumulative injection) at injectors and pressure states at producers are fed to SRU as the input and the phase rates at producers are treated as the output. Once the SRU is trained and validated, it is then used to assess the connectivity of each injector to any producer using permutation variable importance method, wherein inputs corresponding to an injector are shuffled and the increase in prediction error at a given producer is recorded as the importance (connectivity metric) of the injector to the producer. This method is tested in both synthetic and field-scale cases. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. This significantly improves confidence in our data-driven procedure. The novelty of this work is that it is purely data-driven method and can directly interpret routine surface measurements to intuitive subsurface knowledge. Furthermore, the streamline-based validation procedure provides physics-based backing to the results obtained from data analytics. This study results in a reliable and efficient data analytics framework that is well-suited for large field applications.

42 ENGINEERING↗

Gridded daily weather data for North America with comprehensive uncertainty quantification

Abstract Access to daily high-resolution gridded surface weather data based on direct observations and over long time periods is essential for many studies and applications including vegetation, wildlife, soil health, hydrological modelling, and as driver data in Earth system models. We present Daymet V4, a 40-year daily meteorological dataset on a 1 km grid for North America, Hawaii, and Puerto Rico, providing temperature, precipitation, shortwave radiation, vapor pressure, snow water equivalent, and day length. The dataset includes an objective quantification of uncertainty based on strict cross-validation analysis for temperature and precipitation results. The dataset represents several improvements from a previous version, and this data descriptor provides complete documentation for updated methods. Improvements include: reductions in the timing bias of input reporting weather station measurements; improvement to the three-dimensional regression model techniques in the core algorithm; and a novel approach to handling high elevation temperature measurement biases. We show cross-validation analyses with the underlying weather station data to demonstrate the technical validity of new dataset generation methods, and to quantify improved accuracy.

54 ENVIRONMENTAL SCIENCES↗

Hybrid Cyber-attack Detection in Photovoltaic Farms

Here, to address the cyber-physical security in PV farms, a hybrid cyber-attack detection is proposed in this manuscript. To secure PV farms, the proposed method integrates model-based and data-driven methods by fusing the detection score at the device and system levels. First, a model-based cyber-attack detection method is developed for each PV inverter. A residual between the estimation of the Kalman filter and measurement is calculated. By leveraging the calculated residual from all inverters, a squared Mahalanobis distance is developed for device detection score generation. At the system level, a convolutional neural network (CNN) is proposed to detect cyber-attack using the waveform data at the point of common coupling (PCC) in PV farms. To improve the CNN detection accuracy, a set of well-designed features are extracted from the raw waveform data. Finally, a weighted detection score fusion method is proposed to combine device and system detection scores by using their complementary strength. The feasibility and robustness of the proposed method are validated by testing cases and a comparative experiment.

14 SOLAR ENERGY↗

ANN-based ground motion model for Turkey using stochastic simulation of earthquakes

SUMMARY Turkey is characterized by a high level of seismic activity attributed to its complex tectonic structure. The country has a dense network to record earthquake ground motions; however, to study previous earthquakes and to account for potential future ones, ground motion simulations are required. Ground motion simulation techniques offer an alternative means of generating region-specific time-series data for locations with limited seismic networks or regions with seismic data gaps, facilitating the study of potential catastrophic earthquakes. In this research, a local ground motion model (GMM) for Turkey is developed using region-specific simulated records, thus constructing a homogeneous data set. The simulations employ the stochastic finite-fault approach and utilize validated input-model parameters in distinct regions, namely Afyon, Erzincan, Duzce, Istanbul and Van. To overcome the limitations of linear regression-based models, artificial neural network is used to establish the form of equations and coefficients. The predictive input parameters encompass fault mechanism (FM), focal depth (FD), moment magnitude (Mw), Joyner and Boore distance (RJB) and average shear wave velocity in the top 30 m (Vs30). The data set comprises 7359 records with Mw ranging between 5.0 and 7.5 and RJB ranging from 0 to 272 km. The results are presented in terms of spectral ordinates within the period range of 0.03–2.0 s, as well as peak ground acceleration and peak ground velocity. The quantification of the GMM uncertainty is achieved through the analysis of residuals, enabling insights into inter- and intra-event uncertainties. The simulation results and the effectiveness of the model are verified by comparing the predicted values of ground motion parameters with the observed values recorded during previous events in the region. The results demonstrate the efficacy of the proposed model in simulating physical phenomena.

Karimzadeh, Shaghayegh (ORCID:0000000337531676)↗

An Open-Source Virtual Testbed for a Real Net-Zero Energy Community

Net zero energy communities (NZECs) are critical to ensure sustainability and resilience of modernized power systems. System modeling helps overcome technical challenges in designing and operating NZECs. In this paper, we present the modeling work based on a real NZEC. Two sets of models are developed: higher-fidelity physics-based models considering the interaction between subsystems of the studied NZEC and capturing fast-dynamics; and lower-fidelity data-driven models requiring less resource to establish and/or run. All models are validated against measurements from this real NZEC. In addition, we create a simulation framework which streamlines the processes for simulation and thus allows using developed models to form a virtual testbed. To demonstrate the usage of the virtual testbed, a case study is conducted where a building-to-grid integration control is evaluated via simulation. The evaluation results suggest that the tested control significantly smooths the power draw of the studied community and doesn’t sacrifice the thermal comfort to a great extent.

Huang, Sen↗

A Multiscale Deep Learning Model for Soil Moisture Integrating Satellite and In Situ Data

Deep learning (DL) models trained on hydrologic observations can perform extraordinarily well, but they can inherit deficiencies of the training data, such as limited coverage of in situ data or low resolution/accuracy of satellite data. In this work, we propose a novel multiscale DL scheme learning simultaneously from satellite and in situ data to predict 9 km daily soil moisture (5 cm depth). Based on spatial cross-validation over sites in the conterminous United States, the multiscale scheme obtained a median correlation of 0.901 and root-mean-square error of 0.034 m 3 /m 3 . It outperformed the Soil Moisture Active Passive satellite mission's 9 km product, DL models trained on in situ data alone, and land surface models. Our 9 km product showed better accuracy than previous 1 km satellite downscaling products, highlighting limited impacts of improving resolution. Not only is our product useful for planning against floods, droughts, and pests, our scheme is generically applicable to geoscientific domains with data on multiple scales, breaking the confines of individual data sets.

54 ENVIRONMENTAL SCIENCES↗

Predictive Data-driven Platform for Subsurface Energy Production

Subsurface energy activities such as unconventional resource recovery, enhanced geothermal energy systems, and geologic carbon storage require fast and reliable methods to account for complex, multiphysical processes in heterogeneous fractured and porous media. Although reservoir simulation is considered the industry standard for simulating these subsurface systems with injection and/or extraction operations, reservoir simulation requires spatio-temporal “Big Data” into the simulation model, which is typically a major challenge during model development and computational phase. In this work, we developed and applied various deep neural network-based approaches to (1) process multiscale image segmentation, (2) generate ensemble members of drainage networks, flow channels, and porous media using deep convolutional generative adversarial network, (3) construct multiple hybrid neural networks such as convolutional LSTM and convolutional neural network-LSTM to develop fast and accurate reduced order models for shale gas extraction, and (4) physics-informed neural network and deep Q-learning for flow and energy production. We hypothesized that physicsbased machine learning/deep learning can overcome the shortcomings of traditional machine learning methods where data-driven models have faltered beyond the data and physical conditions used for training and validation. We improved and developed novel approaches to demonstrate that physics-based ML can allow us to incorporate physical constraints (e.g., scientific domain knowledge) into ML framework. Outcomes of this project will be readily applicable for many energy and national security problems that are particularly defined by multiscale features and network systems.

58 GEOSCIENCES↗

Customizable adaptive regularization techniques for B-spline modeling

B-spline models are a powerful way to represent scientific data sets with a functional approximation. However, these models can suffer from spurious oscillations when the data to be approximated are not uniformly distributed. Model regularization (i.e., smoothing) has traditionally been used to minimize these oscillations; unfortunately, it is sometimes impossible to sufficiently remove unwanted artifacts without smoothing away key features of the data set. In this article, we present a method of model regularization that preserves significant features of a data set while minimizing artificial oscillations. Our method varies the strength of a smoothing parameter throughout the domain automatically, removing artifacts in poorly-constrained regions while leaving other regions unchanged. Further, the proposed method selectively incorporates regularization terms based on first and second derivatives to maintain model accuracy while minimizing numerical artifacts. The behavior of our method is validated on a collection of two- and three-dimensional data sets produced by scientific simulations. In addition, a key tuning parameter is highlighted and the effects of this parameter are presented in detail. This paper is an extension of our previous conference paper at the 2022 International Conference on Computational Science (ICCS) (Lenz et al., 2022).

97 MATHEMATICS AND COMPUTING↗

A Process-based Model with Temperature, Water, and Lab-derived Data Improves Predictions of Daily Culex pipiens/restuans Mosquito Density

While the number of human cases of mosquito-borne diseases has increased in North America in the last decade, accurate modeling of mosquito population density has remained a challenge. Longitudinal mosquito trap data over the many years needed for model calibration, and validation is relatively rare. In particular, capturing the relative changes in mosquito abundance across seasons is necessary for predicting the risk of disease spread as it varies from year to year. We developed a discrete, semi-stochastic, mechanistic process-based mosquito population model that captures life-cycle egg, larva, pupa, adult stages, and diapause for Culex pipiens (Diptera, Culicidae) and Culex restuans (Diptera, Culicidae) mosquito populations. This model combines known models for development and survival into a fully connected age-structured model that can reproduce mosquito population dynamics. Mosquito development through these stages is a function of time, temperature, daylight hours, and aquatic habitat availability. The time-dependent parameters are informed by both laboratory studies and mosquito trap data from the Greater Toronto Area. The model incorporates city-wide water-body gauge and precipitation data as a proxy for aquatic habitat. This approach accounts for the nonlinear interaction of temperature and aquatic habitat variability on the mosquito life stages. We demonstrate that the full model predicts the yearly variations in mosquito populations better than a statistical model using the same data sources. This improvement in modeling mosquito abundance can help guide interventions for reducing mosquito abundance in mitigating mosquito-borne diseases like West Nile virus.

59 BASIC BIOLOGICAL SCIENCES↗

Structured Neural Network Modeling for Developing Digital Twins Models of Hydropower Generation Units

Dynamic modeling is a key part in the development of digital twin (DT) for dynamic systems. This is true for hydropower systems, where whole system modeling including penstock, turbine and generators, etc is important in realizing actuate modeling for the real systems. On the other hand, in response to the large variations of the power demand due to increased penetration of renewables such as wind and solar, hydropower systems are now required to operate in a large power generation range. This situation triggers the nonlinear characteristics of the generation unit with respect to its models. As such, it is imperative to use data driven modeling such as neural networks to learn the nonlinear dynamics of the hydropower generation unit. To achieve this objective, this study constructs a modeling and learning algorithm integrated with multiple structured neural network models for the modeling of turbine shaft speed, penstock pressure, and generator power output based on the generator power control setpoint, field current, and field voltage. In addition, the study uses the hydropower data from Tacoma Public Utilities to train and validate the proposed neural network algorithm. The results have shown that this structured neural network modeling approach can learn the system dynamics effectively by using the real-time data collected from the hydropower system with the desired modeling results.

Wang, Hong↗

Estimation of the Thermal-Fluids and Thermal-Structural Performance of Helium-Cooled Modular Finger-Type Divertors

Over the past decade, our group has investigated the thermal-fluid performance of the helium-cooled modular divertor with multiple jets (HEMJ) and a simplified “flat” design of the HEMJ for long-pulse magnetic fusion energy (MFE) reactors. Experimental studies were performed in a helium (He) loop at the prototypical pressure of 10 MPa, nearly prototypical He temperatures and incident heat fluxes using test sections made from stainless steel and tungsten alloys. Correlations for average Nusselt numbers and pressure loss coefficients were developed from the data and are used to validate computational fluid dynamics (CFD) models. Here, this work presents updated thermal-fluids performance curves based on these correlations that estimate the maximum heat flux that can be accommodated by the plasma-facing surface and coolant pumping power requirements at prototypical operating conditions. Thermal-structural performance curves developed from ITER structural design criteria are introduced, which include protection against ductile and non-ductile failure, ratcheting fatigue, and creep fatigue. The performance design curves for these finger-type divertors demonstrate that the “flat” design, with a significantly less complicated geometry than the HEMJ, has thermal-fluid and thermal-structural performance comparable to the original HEMJ concept, and may be superior in terms of non-ductile failure criteria.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Detection of Marine Borne Radioactive Sources Report

This report describes a preliminary analysis on the utility of modern capabilities to localize radiological sources in marine environment. This study analyzed expected background sources and identified possible approaches based on modeling and simulation analyses. The effort culminated in the design and fabrication of validation experiments which provided favorable supporting data. The conclusions of this effort are not conclusive as to the viability of the approach nor is the underlining hypothesis discounted. Continuation of the experimental approach is recommended with a focus on open-water experimentation.

54 ENVIRONMENTAL SCIENCES↗

Blueprints for Training Information Bottlenecks for Collider Analyses

Dimensionality reduction is a crucial aspect of data analysis in high energy physics, even if accompanied by information loss. Several methods, including histogram- and kernel-based analyses, are only computationally feasible for low-dimensional data. Furthermore, simulation models used in HEP can often only be validated for low-dimensional data. We provide several blueprints for using machine learning to create low-dimensional data representations (continuous event variables and discrete classification labels) for use in signal discovery and parameter estimation tasks. We also describe how to design the learned representation to facilitate a) searches with unknown model parameters and b) validation of simulation models in data control regions.

43 PARTICLE ACCELERATORS↗

A graph neural network (GNN) approach to basin-scale river network learning: the role of physics-based connectivity and data fusion

Abstract. Rivers and river habitats around the world are under sustained pressure from human activities and the changing global environment. Our ability to quantify and manage the river states in a timely manner is critical for protecting the public safety and natural resources. In recent years, vector-based river network models have enabled modeling of large river basins at increasingly fine resolutions, but are computationally demanding. This work presents a multistage, physics-guided, graph neural network (GNN) approach for basin-scale river network learning and streamflow forecasting. During training, we train a GNN model to approximate outputs of a high-resolution vector-based river network model; we then fine-tune the pretrained GNN model with streamflow observations. We further apply a graph-based, data-fusion step to correct prediction biases. The GNN-based framework is first demonstrated over a snow-dominated watershed in the western United States. A series of experiments are performed to test different training and imputation strategies. Results show that the trained GNN model can effectively serve as a surrogate of the process-based model with high accuracy, with median Kling–Gupta efficiency (KGE) greater than 0.97. Application of the graph-based data fusion further reduces mismatch between the GNN model and observations, with as much as 50 % KGE improvement over some cross-validation gages. To improve scalability, a graph-coarsening procedure is introduced and is demonstrated over a much larger basin. Results show that graph coarsening achieves comparable prediction skills at only a fraction of training cost, thus providing important insights into the degree of physical realism needed for developing large-scale GNN-based river network models.

54 ENVIRONMENTAL SCIENCES↗

Availability of Critical Benchmark Experiments for the Pebble Tanker Transportation Model for Nuclear Criticality Safety Validation of TRISO Pebbles

This study addresses the need for comprehensive investigations into TRi-structural ISOtropic (TRISO) fuel pebble transportation validation. In this work, an exploratory model, the pebble tanker(PT), was developed with the aim of facilitating the validation of nuclear criticality safety calculations in the context of industrial-scale transportation of TRISO fuel. The PT model was designed to investigate the availability and applicability of critical benchmark experiments crucial for assessing the transportation of these pebbles. This work incorporated sensitivity/uncertainty (S/U) similarity studies to quantify the applicability of critical benchmark experiments and to address nuclear data uncertainties in the context of TRISO transportation. Two container models were investigated: one for the Hermes-type pebble and one for the Pebble Bed Modular Reactor (PBMR)–type pebble. The models were simplified, considering fuel, containment, and either water or air, to enable a focus on the underlying physics of applications involving TRISO fuel pebbles using the PT model. A crucial aspect under consideration was the capacity of the transport package to hold pebbles while ensuring subcriticality in the flooded state. An approach in the criticality validation process involves assessing the similarity between systems through an integral index parameter evaluation. This involves calculating a correlation coefficient (referred to as c k ) based on shared nuclear data–induced uncertainty between a benchmark experiment and the application of the PT model. To facilitate this analysis, the SCALE tools, particularly the CSAS6-Shift, TSUNAMI-3D-Shift, and TSUNAMI-IP sequences, were employed for comprehensive studies in neutronics and S/U analysis. Our findings showed that there are sufficient critical experimental benchmarks to perform this validation of the PT model in the most reactive state, i.e. when the tanker is flooded. This paper provides valuable insights into validating a transport package for Generation IV TRISO fuel pebbles.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

MOOSE-based Tritium Migration Analysis Program, Version 8 (TMAP8) for advanced open-source tritium transport and fuel cycle modeling

Tritium management is critical for the safety, sustainability, and economics of fusion energy systems, and advanced and reliable modeling tools help accelerate the development of tritium technologies. This paper presents the Tritium Migration Analysis Program, Version 8 (TMAP8), an open-source, MOOSE-based application developed to provide state-of-the-art tritium transport and fuel cycle modeling capabilities. TMAP8 aims to expand the capabilities of previous versions (i.e., TMAP4 and TMAP7) by leveraging modern computational techniques, ensuring high software quality assurance standards (key to building trust), and enabling multispecies, multiscale, and multiphysics simulations for integrated tritium transport modeling in complex geometries. This paper outlines TMAP8’s scope and rigorous development practices, emphasizing its transparency, accessibility, modularity, and reliability. We present the current suite of verification and validation cases based on those from TMAP4, demonstrating TMAP8’s accuracy and reliability against analytical solutions and experimental data. Additionally, the paper showcases TMAP8’s integrated fuel cycle modeling capabilities, highlighting its applicability at various scales and levels. The TMAP8 code and documentation are openly available, promoting collaborative development and widespread adoption within the fusion community. Future work will soon expand TMAP8’s verification and validation suite to include those from TMAP7 and other recent experimental studies for validation.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗