Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Towards fast, accurate predictions of RF simulations via data-driven modeling: Forward and lateral models

Three machine learning techniques (multilayer perceptron, random forest, and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. A single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time to complete, which is acceptable for many purposes, but too slow for integrated modeling and real-time control applications. More accurate simulations with fast electron diffusion are even slower, requiring multiple hours of run time with parallel processing. The machine learning models use a database of 16,000+ GEN-RAY/CQL3D simulations for training, validation, and testing. Latin hypercube sampling methods implemented in πScope ensure that the database covers the range of 9 input parameters (n e0 , T e0 , I p , B t , R 0 , n ∥︀ , Z e f f , V loop , P LHCD ) with sufficient density in all regions of parameter space. The surrogate models reduce the computation time from minutes-hours to ms with high accuracy across the input parameter space. Data-driven surrogate models also allow for solving inverse and “lateral” problems. A surrogate model for the inverse problem maps from a desired current drive or power deposition profile to a set of input parameters that would result in such a profile, while a surrogate model for the lateral problem maps from a measured experimental quantity such as hard x-ray emission to a current drive or power deposition profile. In conclusion, the πScope database creation workflow is flexible and applicable to other RF simulation codes such as TORIC.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Data-Driven Modeling of a High Capacity Cryogenic System for Control Optimization

The Cryogenic Moderator System (CMS) is responsible for maintaining a steady flow of cold neutrons for numerous physics experiments at the Spallation Neutron Source (SNS) in Oak Ridge National Laboratory (ORNL). Sudden losses in beam power, known as beam trips, cause a major disturbance to the CMS due to large step changes in cooling demands. Ongoing efforts on upgrading the neutron beam power from 1.4 to 2.0MW are expected to generate larger transients that can further strain the CMS subsystems if they are not properly controlled. To manage such disturbances, four flow valves and one electric heater are adjusted by five decentralized proportional-integral-derivative (PID) controllers. However, the original PID gains were calibrated empirically based only on tracking performance and not based on disturbance rejection. To address this issue without compromising current CMS operations, a control-oriented model was developed to recalibrate the PID controllers offline. The zero-dimensional (0-D) model was based on simple physics-based principles and data-driven system identification techniques. The CMS was broken into several subsystems for analysis, each of which corresponds to a parametric model tied to the thermodynamic states of the working fluid. The model parameters were identified using the nonlinear least squares method where the residuals were calculated from available sensor data. Simulation results show that the proposed model can capture the dynamics of the CMS at steady state and during beam trips.

Maldonado Puente, Bryan↗

An ensemble data assimilation modeling system for operational outdoor microalgae growth forecasting

Microalgae have received increasing attention as a potential feedstock for biofuel or biobased products. Forecasting the microalgae growth is beneficial for managers in planning pond operations and harvesting decisions. This study proposed a biomass forecasting system comprised of the Huesemann Algae Biomass Growth Model (BGM), the Modular Aquatic Simulation System in Two Dimensions (MASS2), ensemble data assimilation (DA), and numerical weather prediction Global Ensemble Forecast System (GEFS) ensemble meteorological forecasts. The novelty of this study is to seek the use of ensemble DA to improve both BGM and MASS2 model initial conditions with the assimilation of biomass and water temperature measurements and consequently improve short-term biomass forecasting skills. This study introduces the theory behind the proposed integrated biomass forecasting system, with an application undertaken in pseudo-real-time in three outdoor ponds cultured with Chlorella sorokiniana in Delhi, California, United States. Results from all three case studies demonstrate that the biomass forecasting system improved the short-term (i.e., 7-day) biomass forecasting skills by about 60% on average, comparing to forecasts without using the ensemble DA method. Given the satisfactory performances achieved in this study, it is probable that the integrated BGM-MASS2-DA forecasting system can be used operationally to inform managers in making pond operation and harvesting planning decisions.

59 BASIC BIOLOGICAL SCIENCES↗

Large Destabilization of (TiVNb)-Based Hydrides via (Al, Mo) Addition: Insights from Experiments and Data-Driven Models

High-entropy alloys (HEAs) represent an interesting alloying strategy that can yield exceptional performance properties needed across a variety of technology applications, including hydrogen storage. Examples include ultrahigh volumetric capacity materials (BCC alloys → FCC dihydrides) with improved thermodynamics relative to conventional high-capacity metal hydrides (like MgH 2 ), but still further destabilization is needed to reduce operating temperature and increase system-level capacity. Here, in this work, we demonstrate efficient hydride destabilization strategies by synthesizing two new Al 0.05 (TiVNb) 0.95–x Mo x (x = 0.05, 0.10) compositions. We specifically evaluate the effect of molybdenum (Mo) addition on the phase structure, microstructure, hydrogen absorption, and desorption properties. Both alloys crystallize in a bcc structure with decreasing lattice parameters as the Mo content increases. The alloys can rapidly absorb hydrogen at 25 °C with capacities of 1.78 H/M (2.79 wt %) and 1.79 H/M (2.75 wt %) with increasing Mo content. Pressure-composition isotherms suggest a two-step reaction for hydrogen absorption to a final fcc dihydride phase. The experiments demonstrate that increasing Mo content results in a significant hydride destabilization, which is consistent with predictions from a gradient boosting tree data-driven model for metal hydride thermodynamics. Furthermore, improved desorption properties with increasing Mo content and reversibility were observed by in situ synchrotron X-ray diffraction, in situ neutron diffraction, and thermal desorption spectroscopy.

36 MATERIALS SCIENCE↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Crystallization↗

Dynamic in-context learning with conversational models for data extraction and materials property prediction

The advent of natural language processing and large language models (LLMs) has revolutionized the extraction of data from unstructured scholarly papers. However, ensuring data trustworthiness remains a significant challenge. In this paper, we introduce PropertyExtractor, an open-source tool that leverages advanced conversational LLMs such as Google gemini-pro and OpenAI gpt-4, blends zero-shot with few-shot in-context learning, and employs engineered prompts for the dynamic refinement of structured information hierarchies—enabling autonomous, efficient, scalable, and accurate identification, extraction, and verification of material property data. Our tests on material data demonstrate precision and recall that exceed 95% with an error rate of ∼9%, highlighting the effectiveness and versatility of the toolkit. Finally, databases for 2D material thicknesses, a critical parameter for device integration, and energy bandgap values are developed using PropertyExtractor. In particular, for the thickness database, the rapid evolution of the field has outpaced both experimental measurements and computational methods, creating a significant data gap. Our work addresses this gap and showcases the potential of PropertyExtractor as a reliable and efficient tool for the autonomous generation of various material property databases, advancing the field.

Ekuma, Chinedu E. (ORCID:0000000258527556)↗

Effect of vaccination on the case fatality rate for COVID-19 infections 2020–2021: multivariate modelling of data from the US Department of Veterans Affairs

Objectives: To evaluate the benefits of vaccination on the case fatality rate (CFR) for COVID-19 infections. Design, setting and participants: The US Department of Veterans Affairs has 130 medical centres. We created multivariate models from these data—339 772 patients with COVID-19—as of 30 September 2021. Outcome measures: The primary outcome for all models was death within 60 days of the diagnosis. Logistic regression was used to derive adjusted ORs for vaccination and infection with Delta versus earlier variants. Models were adjusted for confounding factors, including demographics, comorbidity indices and novel parameters representing prior diagnoses, vital signs/baseline laboratory tests and outpatient treatments. Patients with a Delta infection were divided into eight cohorts based on the time from vaccination to diagnosis. A common model was used to estimate the odds of death associated with vaccination for each cohort relative to that of unvaccinated patients. Results: 9.1% of subjects were vaccinated. 21.5% had the Delta variant. 18 120 patients (5.33%) died within 60 days of their diagnoses. The adjusted OR for a Delta infection was 1.87±0.05, which corresponds to a relative risk (RR) of 1.78. The overall adjusted OR for prior vaccination was 0.280±0.011 corresponding to an RR of 0.291. Raw CFR rose steadily after 10–14 weeks. The OR for vaccination remained stable for 10–34 weeks. Conclusions: Our CFR model controls for the severity of confounding factors and priority of vaccination, rather than solely using the presence of comorbidities. Our results confirm that Delta was more lethal than earlier variants and that vaccination is an effective means of preventing death. After adjusting for major selection biases, we found no evidence that the benefits of vaccination on CFR declined over 34 weeks. We suggest that this model can be used to evaluate vaccines designed for emerging variants.

59 BASIC BIOLOGICAL SCIENCES↗

Data-driven modeling of coarse mesh turbulence for reactor transient analysis using convolutional recurrent neural networks

Advanced nuclear reactors often exhibit complex thermal-fluid phenomena during transients. To accurately capture such phenomena, a coarse-mesh three-dimensional (3-D) modeling capability is desired for modern nuclear-system code. In the coarse-mesh 3-D modeling of advanced-reactor transients that involve flow and heat transfer, accurately predicting the turbulent viscosity is a challenging task that requires an accurate and computationally efficient model to capture the unresolved fine-scale turbulence. In this work, we propose a data-driven coarse-mesh turbulence model based on local flow features for the transient analysis of thermal mixing and stratification in a sodium-cooled fast reactor. The model has a coarse mesh setup to ensure computational efficiency, while it is trained by fine-mesh computational fluid dynamics (CFD) data to ensure accuracy. A novel neural network architecture, combining a densely connected convolutional network and a long-short-term-memory network, is developed that can efficiently learn from the spatial temporal CFD transient simulation results. The neural network model was trained and optimized on a loss-of flow transient and demonstrated high accuracy in predicting the turbulent viscosity field during the whole transient. The trained model's generalization capability was also investigated on two other transients with different inlet conditions. The study demonstrates the potential of applying the proposed data-driven approach to support the coarse-mesh multi-dimensional modeling of advanced reactors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

Comprehensive framework for data-driven model form discovery of the closure laws in thermal-hydraulics codes

The two-phase two-fluid model is a basis of many thermal-hydraulics codes used in design, licensing, and safety considerations of nuclear power plants. Thermal-hydraulics codes rely on the closure laws to close the system of conservation equations and describe the interactions between phases. These laws, derived from years of experimental investigations, are semi-empirical correlations that lack generality and have a limited range of applicability. Increase of computational power, availability of new experiments, and development of high-fidelity simulations has increased the number of validation data. The discrepancies between the code predictions and the validation data are a great source of knowledge. Missing physics that are not included in the model but are important for the considered phenomena can be discovered by propagating the information from the experimental results through the model. Furthermore, physics-discovered data-driven model form (P3DM) methodology integrates available integral effect tests and separate effects tests to determine the necessary corrections to the model form of the closure laws. In contrast to existing calibration techniques, the methodology modifies the functional form of the closure laws. Based on the functional form of the correction, the missing physics that were not included in the original model can be discovered. The methodology provides the alternative to the machine learning approach, in which the model is discovered in the form of the intractable black-box relation. In this work, the methodology was applied to the CTF subchannel code to improve the prediction of the two-phase flow phenomena.

42 ENGINEERING↗

Data-driven Stellar Models

We developed a data-driven model to map stellar parameters ( T eff , log g , and [Fe/H]) accurately and precisely to broadband stellar photometry. This model must, and does, simultaneously constrain the passband-specific dust reddening vector in the Milky Way, R. The model uses a neural network to learn the (de-reddened) absolute magnitude in one band and colors across many bands, given stellar parameters from spectroscopic surveys and parallax constraints from Gaia. To demonstrate the effectiveness of this approach, we train our model on a data set with spectroscopic parameters from LAMOST, APOGEE, and GALAH, Gaia parallaxes, and optical and near-infrared photometry from Gaia, Pan-STARRS 1, Two Micron All Sky Survey and Wide-field Infrared Survey Explorer. Testing the model on these data sets leads to an excellent fit and a precise—and by construction—accurate prediction of the color–magnitude diagrams in many bands. This flexible approach rigorously links spectroscopic and photometric surveys, and also results in an improved, T eff -dependent R. As such, it provides a simple and accurate method for predicting photometry in stellar evolutionary models. Our model will form a basis to infer stellar properties, distances, and dust extinction from photometric data, which should be of great use in 3D mapping of the Milky Way. Our trained model can be obtained at doi:10.5281/zenodo.3902382.

79 ASTRONOMY AND ASTROPHYSICS↗

Poisson Log-Normal Process for Count Data Prediction

Modeling count data is important in physics and other scientific disciplines, where measurements often involve discrete, non-negative quantities such as photon or neutrino detection events. Traditional parametric approaches can be trained to generate integer-count predictions but may struggle with capturing complex, non-linear dependencies often observed in the data. Gaussian process (GP) regression provides a robust non-parametric alternative to modeling continuous data; however, it cannot generate integer outputs. We propose the Poisson Log-Normal (PoLoN) process, a framework that employs GP to model Poisson log-rates. As in GP regression, our approach relies on the correlations between data points captured via GP kernel structure rather than explicit functional parameterizations. We demonstrate that the PoLoN predictive distribution is Poisson-LogNormal and provide an algorithm for optimizing kernel hyperparameters. Furthermore, we adapt the PoLoN approach to the problem of detecting weak localized signals superimposed on a smoothly varying background - a task of considerable interest in many areas of science and engineering. Our framework allows us to predict the strength, location and width of the detected signals. We evaluate PoLoN's performance using both synthetic and real-world datasets, including the open dataset from CERN which was used to detect the Higgs boson at the Large Hadron Collider. Our results indicate that the PoLoN process can be used as a non-parametric alternative for analyzing, predicting, and extracting signals from integer-valued data.

Saha, Anushka [Rutgers U., Piscataway]↗

A Short-Term Solar Forecasting Platform Using a Physics-Based Smart Persistence Model and Data Imputation Method

Electrical energy plays vital role in our socio-economic activity and therefore ensuring the reliability of the electric grid, from the generation, transmission and distribution level is critical. In order to maintain the power system parameter viz., frequency, voltage, etc., optimally, balancing of generation and consumption is very much essential. However, solar energy is infirm power by nature this is due to cloud cover / other local phenomena. Hence, Photovoltaic (PV) power generation brings a significant challenge to the grid operator due to the variability of the solar energy. The complexity of this challenge in terms of planning and dispatch ability of PV resources, aggravates with the high penetration of solar energy into the electric grid. In this setting, reliable solar radiation forecasting models based on accurate and quality input data become essential. In order to develop a suitable model for predicting solar radiation, quality historical / real time measurement is also needed. Under this study NIWE and NREL jointly developed / tested short-term solar forecasting frameworks using a smart persistence and physics-based smart persistence models for intra-hour forecasting of solar radiation (PSPI) and benchmarked 9 different data imputation techniques in 15 Solar Radiation Resource Assessment (SRRA) stations, located at different parts of India. During any measurement campaign, due to various technical reasons, we may miss few observations. However, the missing observation often reduce the performance of any forecasting model. Therefore, suitable data imputation method would assist us to obtain continuous observation of solar radiation. A station-by-station and method-by-method analysis was carried out to understand the performance of each model. Based on our analysis, among all the data imputation methods, the Kalman data imputation method is better for Indian Weather condition. In addition, Kalman StructTS, Linear, Stine and Arima methods yield slightly inferior accuracy compared to Kalman, but outperform the other methods. The extended solar radiation data are used by solar forecasting models to provide the prediction of solar radiation at 15 SRRA stations. As far as short term forecasting model is concerned, the PSPI model outperforms the Smart Persistence model. However, the forecast error is increases with the forecasting horizon.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Modeling Grid Data Flows for Transmission and Distribution Operations: Review, Design, Next Steps

Operational scenarios of the power grids grow multifold to accommodate the diverse needs of both the utilities and end consumers, and the various other stakeholders in-between. To comprehensively model and apply analytics to support objectives and business functions of grid sectors, a reliable approach to characterize and design data flows is crucial. The flows bridge business functions with communications protocols, stakeholders such as the grid actors, and data interfaces comprising different data objects. Additionally, constraints applied to the flow such as cybersecurity, trust, privacy, and ownership among others intersect these entities, requiring the delineation of their interactions under different scenarios. This paper aims to not only highlight relevant research in the space of grid data flows, but also proposes, for the transmission-distribution sector, a novel modeling approach that marries the aforementioned entities: objectives, business functions, data interfaces, communication protocols, data stakeholders, and flow constraints. It elaborates on the design philosophy and the significance of each entity within the model and applies it to an example function of fault location, isolation and service restoration (FLISR). Finally, the next steps to extend the application of this data flow model for other practical operational scenarios are discussed.

Sundararajan, Aditya [ORNL] (ORCID:000000033577854↗

Data-driven modeling of dynamic occupant thermostat override behavior for demand response applications

Buildings consume nearly 40% of global energy and produce similar emissions. Whiletechnological advances address efficiency, occupant behavior causes energy use variations up to 300% between identical buildings. This gap between predicted and actual building performance impacts building design, operations, and grid demand management programs. Through analyses of smart thermostat data from 1,400 single-occupant homes, the researchdemonstrates that occupants respond to 8°F thermostat setpoint changes within a median of 15 minutes, while 2°F changes trigger responses within a median of 30 minutes. This highlights an understudied temporal relationship between thermostat setbacks and response time of occupant behaviors. Models of such behavior dynamics are required to incorporate occupant impacts into building performance simulation. A key contribution of this dissertation is the Thermal Frustration Theory (TFT), which positsthat thermal discomfort driven behaviors are caused by the time-accumulation of discomfort, not simply a temperature deviation threshold or a delay from an initiating event. Using a dataset of 634 thermostats, each with 25+ manual setpoint changes, a comparative analysis of TFT and comfort zone and a delayed response theories demonstrated that personalized TFT models better predict when manual setpoint change occur. This was measured by the area under the curve statistical measure (AUC); all three models perform similarly by a Matthews Correlation Coefficient measure. Higher AUC performance is especially important for modeling occupant behavior in demand response programs where false negatives of rare occupant interactions could adversely affect grid stability. EnergyPlus based simulations were conducted with TFT-derived occupant models, demonstrating the ability to identify parameters of known TFT models from only data observable with smart thermostats, even under the presence of noise from routine overrides. Overall, the dissertation highlights that thermostat interactions are neither static,instantaneous, nor driven solely by the environment. Instead, temporal accumulation of discomfort and routine-based behavior play important roles. The methodology and results offer a pathway towards more accurate modeling of human-building interactions for policy assessment, building design, and demand response programs.

Sharma, Kunind [Northeastern University] (ORCID:00↗

Assimilation of citizen science data in snowpack modeling using a new snow data set: Community Snow Observations

A physically based snowpack evolution and redistribution model was used to test the effectiveness of assimilating crowd-sourced snow depth measurements collected by citizen scientists. The Community Snow Observations project gathers, stores, and distributes measurements of snow depth recorded by recreational users and snow professionals in high mountain environments. These citizen science measurements are valuable since they come from terrain that is relatively undersampled and can offer in situ snow information in locations where snow information is sparse or nonexistent. The present study investigates (1) the improvements to model performance when citizen science measurements are assimilated, and (2) the number of measurements necessary to obtain those improvements. Model performance is assessed by comparing time series of observed (snow pillow) and modeled snow water equivalent values, by comparing spatially distributed maps of observed (remotely sensed) and modeled snow depth, and by comparing fieldwork results from within the study area. The results demonstrate that few citizen science measurements are needed to obtain improvements in model performance, and these improvements are found in 62 % to 78 % of the ensemble simulations, depending on the model year. Model estimations of total water volume from a subregion of the study area also demonstrate improvements in accuracy after CSO measurements have been assimilated. These results suggest that even modest measurement efforts by citizen scientists have the potential to improve efforts to model snowpack processes in high mountain environments, with implications for water resource management and process-based snow modeling.

54 ENVIRONMENTAL SCIENCES↗

Efficient data-driven models for prediction and optimization of geothermal power plant operations

Increasing the capacity of geothermal energy as a renewable resource calls for development and deployment of efficient control and optimization technologies for geothermal power plants. A data-driven prediction and optimization model is presented as a cost-effective and efficient alternative to physics-based approach. The model predicts power output and operational cost by propagating the influence of control and disturbance variables within an artificial neural network (ANN). Numerical experiments with simulated and field data from a real geothermal power plant are first used to demonstrate the prediction performance of the ANN model. The model is then adopted to maximize the net predicted power production by automatically adjusting the working fluid circulation rate. The optimization performance of the model in evaluated using a thermodynamic flowsheet simulation model. The workflow is applied to model and control the effect of ambient temperature on an air-cooled binary cycle power plant, which is complex and costly to perform using a physics-based predictive model. As a result, the performance of the method is demonstrated by applying it to both simulated and field datasets from a binary cycle geothermal power plant.

15 GEOTHERMAL ENERGY↗