Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

The Data Mine model for accessible partnerships in data science

Abstract The Data Mine at Purdue University is a pioneering experiential learning community for undergraduate and graduate students of any background to learn data science. The first data‐intensive experience embedded in a large learning community, The Data Mine had nearly 1300 students in academic year (AY) 2022–2023 and nearly 1700 students for AY 2023–2024. The Data Mine embodies data‐infused education, research, and collaboration. Students learn Python, R, SQL, and shell‐scripting, while working on weekly projects within a high‐performance computing (HPC) cluster. In the Corporate Partners cohort, students work on teams of 5–15 students, led by a paid student team leader. Each cohort follows an Agile approach, working on data‐intensive projects provided by industry partners and mentored by company employees. Students develop professional and data skills throughout the academic year, from August through April. Many students return in subsequent years to the program, increasing their tenure with a Corporate Partner. Student teams are inherently interdisciplinary; students from 133 different majors are involved in the program, ranging from new incoming students through PhD level students. These interdisciplinary teams of students bring new perspectives to challenging problems in which data science is a key part of the solution. The interdisciplinary teams foster an environment of synthesis with ideas and solutions. Students come together with different life experiences, different levels of technical skill, but also varying ways they navigate paths to solutions because of the variety of majors represented, resulting in a more creative and robust solution than a traditional data science program. This article is categorized under: Applications of Computational Statistics > Education in Computational Statistics

Betz, Margaret A.↗

Plant-specific Model and Data Analysis using Dynamic Security Modeling and Simulation

The requirements for U.S. nuclear power plants to maintain a large on-site physical security force contribute to their high operational costs. The cost of maintaining the current physical security posture is approximately 10% of the overall operation and maintenance budget for commercial nuclear power plants. The goal of the Light Water Reactor Sustainability (LWRS) program’s physical security pathway is to develop tools, methods, and technologies and provide the technical basis for an optimized physical security posture. The conservatisms built into current security postures may be analyzed and minimized in order to reduce security costs while still ensuring adequate security and operational safety. The research performed at Idaho National Laboratory within LWRS program’s physical security pathway has successfully developed a dynamic force-on-force modeling framework using various computer simulation tools and integrating them with the dynamic assessment Event Modeling Risk Assessment using Linked Diagrams (EMRALD) tool. This document provides an update on the progress in applying a dynamic computational framework that links results from a commercially available force-on-force simulation tool, a commercially available thermal-hydraulic tool, and EMRALD to an operating commercial nuclear power plant. This report is only a summary of the progress and does not contain specific modeling results as those contain sensitive security information. This process of including plant procedures and multiple analysis results is being called Modeling and Analysis for Safety Security using Dynamic EMRALD Framework or MASS-DEF. Previous reports described how a user could integrate their plant-specific force-on-force models with the dynamic simulation tool EMRALD, model operator actions, integrate with probabilistic risk assessment tools, such as CAFTA (Computer Aided Fault Tree Analysis System) or SAPHIRE (Systems Analysis Programs for Hands-on Integrated Reliability Evaluations), and with thermal-hydraulic tools, such as RELAP-5. Previous reports applied various combinations of available simulations codes with EMRALD using generic plant models to demonstrate how to perform the analysis. This report documents the results of applying the dynamic computational framework to an actual nuclear facility using their security scenarios and timelines. This report does not contain any plant's sensitive information and/or Safeguards Information. The purpose of this study was to verify that results achieved using generic models are similar to actual plant results and to refine our guidance on the use of the framework. This assessment enables further analysis, such as what-if scenarios and staff-reduction evaluation, thereby optimizing physical security at plants.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Improving Text Classification with Large Language Model-Based Data Augmentation

Large Language Models (LLMs) such as ChatGPT possess advanced capabilities in understanding and generating text. These capabilities enable ChatGPT to create text based on specific instructions, which can serve as augmented data for text classification tasks. Previous studies have approached data augmentation (DA) by either rewriting the existing dataset with ChatGPT or generating entirely new data from scratch. However, it is unclear which method is better without comparing their effectiveness. This study investigates the application of both methods to two datasets: a general-topic dataset (Reuters news data) and a domain-specific dataset (Mitigation dataset). Our findings indicate that: 1. ChatGPT generated new data consistently enhanced model’s classification results for both datasets. 2. Generating new data generally outperforms rewriting existing data, though crafting the prompts carefully is crucial to extract the most valuable information from ChatGPT, particularly for domain-specific data. 3. The augmentation data size affects the effectiveness of DA; however, we observed a plateau after incorporating 10 samples. 4. Combining the rewritten sample with new generated sample can potentially further improve the model’s performance.

97 MATHEMATICS AND COMPUTING↗

Transfer-Learnt Energy Models for Predicting Electricity Consumption in Buildings with Limited and Sparse Field Data

Modeling energy consumption is critical for energy-efficient utilization of the electric appliances in a building, smart grid programs (like demand-response), and many other smart home applications. State-of-the-art energy modeling techniques either rely on theoretical models, or extensive instrumentation of the building envelope to gather ``big" data to train a deep neural network. While theoretical models are often limited by their estimation accuracy, it is not always feasible to gather a significant amount of field data. In this paper, we explore transfer learning-based strategies to train much more accurate model for energy estimation when using a sparse field data. We transferred knowledge, in the form of data and parameters, from the simulation framework to the field data. We evaluated the efficacy of our approach on field data collected from six commercial buildings and our results indicate that transfer learning-based models trained over one month data can perform comparative (and in some cases better) than the state-of-the-art machine learning and deep learning solutions.

Jain, Milan↗

Convergence in simulating global soil organic carbon by structurally different models after data assimilation

Abstract Current biogeochemical models produce carbon–climate feedback projections with large uncertainties, often attributed to their structural differences when simulating soil organic carbon (SOC) dynamics worldwide. However, choices of model parameter values that quantify the strength and represent properties of different soil carbon cycle processes could also contribute to model simulation uncertainties. Here, we demonstrate the critical role of using common observational data in reducing model uncertainty in estimates of global SOC storage. Two structurally different models featuring distinctive carbon pools, decomposition kinetics, and carbon transfer pathways simulate opposite global SOC distributions with their customary parameter values yet converge to similar results after being informed by the same global SOC database using a data assimilation approach. The converged spatial SOC simulations result from similar simulations in key model components such as carbon transfer efficiency, baseline decomposition rate, and environmental effects on carbon fluxes by these two models after data assimilation. Moreover, data assimilation results suggest equally effective simulations of SOC using models following either first‐order or Michaelis–Menten kinetics at the global scale. Nevertheless, a wider range of data with high‐quality control and assurance are needed to further constrain SOC dynamics simulations and reduce unconstrained parameters. New sets of data, such as microbial genomics‐function relationships, may also suggest novel structures to account for in future model development. Overall, our results highlight the importance of observational data in informing model development and constraining model predictions.

54 ENVIRONMENTAL SCIENCES↗

Multiscale modeling high-order methods and data-driven modeling

Projection-based reduced-order models (ROMs) comprise a promising set of data-driven approaches for accelerating the simulation of high-fidelity numerical simulations. Standard projection-based ROM approaches, however, suffer from several drawbacks when applied to the complex nonlinear dynamical systems commonly encountered in science and engineering. These limitations include a lack of stability, accuracy, and sharp a posteriori error estimators. This work addresses these limitations by leveraging multiscale modeling, least-squares principles, and machine learning to develop novel reduced-order modeling approaches, along with data-driven a posteriori error estimators, for dynamical systems. Theoretical and numerical results demonstrate that the two ROM approaches developed in this work - namely the windowed least-squares method and the Adjoint Petrov - Galerkin method - yield substantial improvements over state-of-the-art approaches. Additionally, numerical results demonstrate the capability of the a posteriori error models developed in this work.

97 MATHEMATICS AND COMPUTING↗

Bayesian Optimization Framework for Imperfect Data or Models

Conventional Bayesian optimization methods implicitly assume that the data and model being optimized are “perfect.” This assumption leads to inaccurate posterior probability distribution functions (PDFs) when applied to “imperfect” data or models. The new Bayesian optimization framework presented in this report provides a way to parameterize the effect of imperfections usually encountered in a prior PDF of generalized data or a model on the posterior PDF. The effects of imperfections are parameterized by a set of constraints imposed on the posterior expectation values of deviations between the data and the model and on their covariance matrix elements. A particular set of values for these constraints conveys an evaluator’s best estimate of the effect of imperfections on the corresponding posterior expectation values. When a prior PDF of generalized data is assumed to be normal, an expression for a posterior PDF satisfying an arbitrary set of constraints is derived analytically for linear models. An analogous iterative algorithm is given for nonlinear models. The corresponding posterior PDF should be used to estimate any posterior expectation values in the presence of imperfections parameterized by that set of constraints. A posterior PDF of a conventional Bayesian optimization method is recovered analytically when all evaluator-specified constraints are set to zero (i.e., in the absence of any imperfections). The analytical expressions derived in this report for normal PDFs and linear models were verified numerically by a Metropolis–Hastings Monte Carlo method. The methods presented herein could be applied to any kind of data or models, including differential cross-section data or integral benchmark experiments.

97 MATHEMATICS AND COMPUTING↗

Surrogate multi-fidelity data and model fusion for scientific discovery and uncertainty quantification in Earth System Models

This whitepaper addresses the Earth and Environmental Systems Sciences Division (EESSD)’s predictability challenges in modeling the integrated water cycle and data-model integration. Specifically, it focuses on reducing and characterizing the uncertainty in the representation of process models for unresolved physics, either due to model resolution or limited by the physical under standing or computational efficiency, and the use of observational data for in-situ process parameter optimization within ESM. The described methods may also be used to determine the nature of responses (e.g. strength and direction), and hence to identify critical processes that drive the overall ESM responses to perturbation in the forcing

54 ENVIRONMENTAL SCIENCES↗

WaterTAP3 Model Input Data for NAWI's Eight Source Water Baseline Analyses

This folder contains the input data for the WaterTAP3 model that was used for the eight NAWI (National Alliance for Water Innovation) source water baselines studies published in the Environmental Science and Technology special issue: Technology Baselines and Innovation Priorities for Water Treatment and Supply. There are also eight other separate DAMS submissions, one per source water, that include the model results for the published studies. In this data submission, all model inputs across the eight baselines are included. The data structure and content are described in a README.txt file. For more details on how to use the data in WaterTAP3 please refer to the model documentation and GitHub site found at "WaterTAP3 Github" linked in the submission resources.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Prototype energy models for data centers

Data centers in the United States consume about two percent of the nation’s electricity. Because heat gains from IT equipment drive cooling demand, data centers offer unique opportunities for energy savings. However, no prototype energy model for data centers is available in the suite of existing U.S. Department of Energy’s Commercial Prototype Building Models. Here we present the development of two new data center prototype models and their implementation in OpenStudio and EnergyPlus. The small-size data center model represents a computer room in a building served by computer room air conditioners (CRACs); while the large-sized model represents stand-alone data centers served by computer room air handlers (CRAHs) with a central chiller plant. For each data center model, two levels of IT equipment (ITE) load density were considered, to cover the wide range of IT power density of data centers: 40 and 100 W/ft 2 (430 and 1076 W/m 2 ) for the computer room, and 100 and 500 W/ft 2 (1076 and 5382 W/m 2 ) for the stand-alone data center. All other assumptions, such as building envelope, lighting, HVAC efficiencies and schedules, were based on the minimal requirements of ASHRAE Standard 90.1 at various vintages. We introduced a novel concept of supply and return air approach temperatures to capture the essential effects of non-uniform airflow and temperature distribution in data centers. The approach temperatures were pre-computed by computational fluid dynamics (CFD) simulations for various configurations of ITE loads and airflow containment management in data centers. A new feature was developed in EnergyPlus to implement the approach temperature method. A case study was conducted to demonstrate the use of the data center models. The two data center models cover all U.S. climate zones and can be used to evaluate energy saving measures for data centers, as well as to support development of data center energy efficiency codes and standards.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Mechanical separations of corn stover anatomical fractions in an integrated feedstock preprocessing system: An experimental and data-driven modeling study

High variabilities of material attributes in lignocellulosic biomass present risks for biofuel and biochemical productions and must be mitigated via preprocessing. Since almost no mechanical device is originally designed for processing biomass, how to operate existing apparatuses with efficient performance has not been investigated extensively. This work presents a study on an integrated screening and air classification to separate cobs and stalks from husks and leaves in corn stover. Prototype machine learning models were developed to assess the feasibility of predicting the process outcome based on the measurable parameters. The models trained upon limited experimental data rendered decent predictive accuracy of yield and purity. The experimental data and modeling results collectively suggest decreasing throughput leads to a higher purity. To the contrary, if throughput increases, a lower purity is likely. A possible trade-off between yield and purity of the separated streams indicates the need for optimal combinations of feedstock size, moisture, and throughput to achieve optimized separations. The results of this study also suggest the need to further improve model predictability by developing more accurate formulations for physics governing the integrated unit operations. To accomplish this, additional experimental data needs to be generated for model training.

09 - BIOMASS FUELS↗

Northern Hemisphere Snow Drought in Earth System Model Simulations and ERA5‐Land Data in 1980–2014

Abstract Low snow levels over the past few decades and predictions of a low‐to‐no snow future have spurred research into snow droughts, which pose a threat to water security and management. Systematic data‐model comparisons of snow drought have been lacking, hindering our understanding of the drivers of snow drought in the past. To address this gap, we analyzed snow drought events using standardized snow water equivalent index derived from monthly results of four numerical experiments using the E3SM Land Model (ELM) and ERA5‐Land data during the period of 1980–2014. Additionally, we compared snow drought duration calculated from models with those from the ERA5‐Land data during selected El Niño‐Southern Oscillation (ENSO) years. The numerical experiments were conducted with ELM driven by two prescribed atmospheric forcings, and with the coupled land‐atmosphere configuration of E3SM with and without plant hydraulics scheme feedback. Analysis reveals that 20%–30% of snow droughts occur due to factors other than above‐normal temperature and low snowfall, such as low soil moisture, warm soil temperature, and low relative humidity, etc., especially in high latitudes (50° North). Furthermore, our study highlights the exacerbating effect of ENSO events on snow drought conditions in various regions, despite some discrepancies between model and ERA5‐Land results. We also identified limitations of the coupled land‐atmosphere models in our current configuration in capturing the spatial patterns of snow droughts. This study underscores the challenge of predicting and mitigating snow drought and the need for a comprehensive understanding of the factors contributing to snow drought.

54 ENVIRONMENTAL SCIENCES↗

A model-independent data assimilation (MIDA) module and its applications in ecology

Abstract. Models are an important tool to predict Earth system dynamics. An accurate prediction of future states of ecosystems depends on not only model structures but also parameterizations. Model parameters can be constrained by data assimilation. However, applications of data assimilation to ecology are restricted by highly technical requirements such as model-dependent coding. To alleviate this technical burden, we developed a model-independent data assimilation (MIDA) module. MIDA works in three steps including data preparation, execution of data assimilation, and visualization. The first step prepares prior ranges of parameter values, a defined number of iterations, and directory paths to access files of observations and models. The execution step calibrates parameter values to best fit the observations and estimates the parameter posterior distributions. The final step automatically visualizes the calibration performance and posterior distributions. MIDA is model independent, and modelers can use MIDA for an accurate and efficient data assimilation in a simple and interactive way without modification of their original models. We applied MIDA to four types of ecological models: the data assimilation linked ecosystem carbon (DALEC) model, a surrogate-based energy exascale earth system model: the land component (ELM), nine phenological models and a stand-alone biome ecological strategy simulator (BiomeE). The applications indicate that MIDA can effectively solve data assimilation problems for different ecological models. Additionally, the easy implementation and model-independent feature of MIDA breaks the technical barrier of applications of data–model fusion in ecology. MIDA facilitates the assimilation of various observations into models for uncertainty reduction in ecological modeling and forecasting.

58 GEOSCIENCES↗

Data-driven model for divertor plasma detachment prediction

We present a fast and accurate data-driven surrogate model for divertor plasma detachment prediction leveraging the latent feature space concept in machine learning research. Our approach involves constructing and training two neural networks: an autoencoder that finds a proper latent space representation (LSR) of plasma state by compressing the multi-modal diagnostic measurements and a forward model using multi-layer perception (MLP) that projects a set of plasma control parameters to its corresponding LSR. By combining the forward model and the decoder network from autoencoder, this new data-driven surrogate model is able to predict a consistent set of diagnostic measurements based on a few plasma control parameters. In order to ensure that the crucial detachment physics is correctly captured, highly efficient 1D UEDGE model is used to generate training and validation data in this study. The benchmark between the data-driven surrogate model and UEDGE simulations shows that our surrogate model is capable of providing accurate detachment prediction (usually within a few per cent relative error margin) but with at least four orders of magnitude speed-up, indicating that performance-wise, it has the potential to facilitate integrated tokamak design and plasma control. Comparing with the widely used two-point model and/or two-point model formatting, the new data-driven model features additional detachment front prediction and can be easily extended to incorporate richer physics. This study demonstrates that the complicated divertor and scrape-off-layer plasma state has a low-dimensional representation in latent space. Understanding plasma dynamics in latent space and utilising this knowledge could open a new path for plasma control in magnetic fusion energy research.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An open source fast fluid dynamics model for data center thermal management

Although computational fluid dynamics (CFD) has been widely adopted to improve data center thermal management, the high computational demand limits its applications, such as multivariate optimal design and operation. Fast fluid dynamics (FFD), which has been applied for fast airflow simulation, shows great potential. However, few research applied FFD for optimal design and operation of data center thermal management. This research improves the FFD model for data centers and conducts a comprehensive evaluation and demonstration. First, the FFD model is improved by solving the advection and diffusion equations together using an upwind scheme instead of a semi-Lagrangian advection solver in the conventional FFD model. Second, new features for data centers are added, such as a pressure correction method to simulate plenum airflow and dynamic boundary conditions for IT racks. The new FFD model is first validated with two indoor environment cases and the results show that the new FFD model has slightly better overall prediction accuracy and faster speed compared to the conventional FFD model. It is also observed that both FFD models achieve acceptable accuracy, except for a few localized disparities with experimental data, which might be due to simplified handling of turbulence viscosity near the boundaries. Furthermore, validation with a real data center shows that the FFD model achieves a similar level of accuracy as CFD when compared to the experimental measurements with some level of uncertainties. It is then demonstrated for data center optimal design and operation, which saves 53.4–58.8% of annual energy while still meeting the thermal requirements. In conclusion, with a much faster speed and comparable accuracy compared to CFD, the FFD model parallelized on a graphics processing unit is promising for practical model-based data center early design and operation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗