Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A Processing and Analytics System for Microscopy Data Workflows: The Pycroscopy Ecosystem of Packages

Major advancements in fields as diverse as biology and quantum computing have relied on a multitude of microscopy techniques. Despite the considerable proliferation of these instruments, significant bottlenecks remain in terms of processing, analysis, storage, and retrieval of the acquired datasets. Aside from lack of file standards, individual domain-specific analysis packages are often disjoint from the underlying datasets, and thus keeping track of analysis and processing steps remains tedious for the end-user, hampering reproducibility. Here, in this study, the pycroscopy ecosystem of packages is introduced, an open-source python-based ecosystem underpinned by a common data model. The data model, termed the N-dimensional spectral imaging data format, is realized in pycroscopy's sidpy package. This package is built on top of dask arrays, thus leveraging dask array attributes, but expanding them to accelerate microscopy relevant analysis and visualization. Several examples of the use of the pycroscopy ecosystem to create workflows for data ingestion and analysis of scanning transmission electron microscopy (STEM) and scanning probe microscopy data are shown. Adoption of such standardized routines will be critical to usher in the next generation of autonomous instruments where processing, computation, and meta-data storage will be critical to overall experimental operations.

97 MATHEMATICS AND COMPUTING↗

The Data Mine model for accessible partnerships in data science

Abstract The Data Mine at Purdue University is a pioneering experiential learning community for undergraduate and graduate students of any background to learn data science. The first data‐intensive experience embedded in a large learning community, The Data Mine had nearly 1300 students in academic year (AY) 2022–2023 and nearly 1700 students for AY 2023–2024. The Data Mine embodies data‐infused education, research, and collaboration. Students learn Python, R, SQL, and shell‐scripting, while working on weekly projects within a high‐performance computing (HPC) cluster. In the Corporate Partners cohort, students work on teams of 5–15 students, led by a paid student team leader. Each cohort follows an Agile approach, working on data‐intensive projects provided by industry partners and mentored by company employees. Students develop professional and data skills throughout the academic year, from August through April. Many students return in subsequent years to the program, increasing their tenure with a Corporate Partner. Student teams are inherently interdisciplinary; students from 133 different majors are involved in the program, ranging from new incoming students through PhD level students. These interdisciplinary teams of students bring new perspectives to challenging problems in which data science is a key part of the solution. The interdisciplinary teams foster an environment of synthesis with ideas and solutions. Students come together with different life experiences, different levels of technical skill, but also varying ways they navigate paths to solutions because of the variety of majors represented, resulting in a more creative and robust solution than a traditional data science program. This article is categorized under: Applications of Computational Statistics > Education in Computational Statistics

Betz, Margaret A.↗

Plant-specific Model and Data Analysis using Dynamic Security Modeling and Simulation

The requirements for U.S. nuclear power plants to maintain a large on-site physical security force contribute to their high operational costs. The cost of maintaining the current physical security posture is approximately 10% of the overall operation and maintenance budget for commercial nuclear power plants. The goal of the Light Water Reactor Sustainability (LWRS) program’s physical security pathway is to develop tools, methods, and technologies and provide the technical basis for an optimized physical security posture. The conservatisms built into current security postures may be analyzed and minimized in order to reduce security costs while still ensuring adequate security and operational safety. The research performed at Idaho National Laboratory within LWRS program’s physical security pathway has successfully developed a dynamic force-on-force modeling framework using various computer simulation tools and integrating them with the dynamic assessment Event Modeling Risk Assessment using Linked Diagrams (EMRALD) tool. This document provides an update on the progress in applying a dynamic computational framework that links results from a commercially available force-on-force simulation tool, a commercially available thermal-hydraulic tool, and EMRALD to an operating commercial nuclear power plant. This report is only a summary of the progress and does not contain specific modeling results as those contain sensitive security information. This process of including plant procedures and multiple analysis results is being called Modeling and Analysis for Safety Security using Dynamic EMRALD Framework or MASS-DEF. Previous reports described how a user could integrate their plant-specific force-on-force models with the dynamic simulation tool EMRALD, model operator actions, integrate with probabilistic risk assessment tools, such as CAFTA (Computer Aided Fault Tree Analysis System) or SAPHIRE (Systems Analysis Programs for Hands-on Integrated Reliability Evaluations), and with thermal-hydraulic tools, such as RELAP-5. Previous reports applied various combinations of available simulations codes with EMRALD using generic plant models to demonstrate how to perform the analysis. This report documents the results of applying the dynamic computational framework to an actual nuclear facility using their security scenarios and timelines. This report does not contain any plant's sensitive information and/or Safeguards Information. The purpose of this study was to verify that results achieved using generic models are similar to actual plant results and to refine our guidance on the use of the framework. This assessment enables further analysis, such as what-if scenarios and staff-reduction evaluation, thereby optimizing physical security at plants.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Improving Text Classification with Large Language Model-Based Data Augmentation

Large Language Models (LLMs) such as ChatGPT possess advanced capabilities in understanding and generating text. These capabilities enable ChatGPT to create text based on specific instructions, which can serve as augmented data for text classification tasks. Previous studies have approached data augmentation (DA) by either rewriting the existing dataset with ChatGPT or generating entirely new data from scratch. However, it is unclear which method is better without comparing their effectiveness. This study investigates the application of both methods to two datasets: a general-topic dataset (Reuters news data) and a domain-specific dataset (Mitigation dataset). Our findings indicate that: 1. ChatGPT generated new data consistently enhanced model’s classification results for both datasets. 2. Generating new data generally outperforms rewriting existing data, though crafting the prompts carefully is crucial to extract the most valuable information from ChatGPT, particularly for domain-specific data. 3. The augmentation data size affects the effectiveness of DA; however, we observed a plateau after incorporating 10 samples. 4. Combining the rewritten sample with new generated sample can potentially further improve the model’s performance.

97 MATHEMATICS AND COMPUTING↗

Transfer-Learnt Energy Models for Predicting Electricity Consumption in Buildings with Limited and Sparse Field Data

Modeling energy consumption is critical for energy-efficient utilization of the electric appliances in a building, smart grid programs (like demand-response), and many other smart home applications. State-of-the-art energy modeling techniques either rely on theoretical models, or extensive instrumentation of the building envelope to gather ``big" data to train a deep neural network. While theoretical models are often limited by their estimation accuracy, it is not always feasible to gather a significant amount of field data. In this paper, we explore transfer learning-based strategies to train much more accurate model for energy estimation when using a sparse field data. We transferred knowledge, in the form of data and parameters, from the simulation framework to the field data. We evaluated the efficacy of our approach on field data collected from six commercial buildings and our results indicate that transfer learning-based models trained over one month data can perform comparative (and in some cases better) than the state-of-the-art machine learning and deep learning solutions.

Jain, Milan↗

Convergence in simulating global soil organic carbon by structurally different models after data assimilation

Abstract Current biogeochemical models produce carbon–climate feedback projections with large uncertainties, often attributed to their structural differences when simulating soil organic carbon (SOC) dynamics worldwide. However, choices of model parameter values that quantify the strength and represent properties of different soil carbon cycle processes could also contribute to model simulation uncertainties. Here, we demonstrate the critical role of using common observational data in reducing model uncertainty in estimates of global SOC storage. Two structurally different models featuring distinctive carbon pools, decomposition kinetics, and carbon transfer pathways simulate opposite global SOC distributions with their customary parameter values yet converge to similar results after being informed by the same global SOC database using a data assimilation approach. The converged spatial SOC simulations result from similar simulations in key model components such as carbon transfer efficiency, baseline decomposition rate, and environmental effects on carbon fluxes by these two models after data assimilation. Moreover, data assimilation results suggest equally effective simulations of SOC using models following either first‐order or Michaelis–Menten kinetics at the global scale. Nevertheless, a wider range of data with high‐quality control and assurance are needed to further constrain SOC dynamics simulations and reduce unconstrained parameters. New sets of data, such as microbial genomics‐function relationships, may also suggest novel structures to account for in future model development. Overall, our results highlight the importance of observational data in informing model development and constraining model predictions.

54 ENVIRONMENTAL SCIENCES↗

Multiscale modeling high-order methods and data-driven modeling

Projection-based reduced-order models (ROMs) comprise a promising set of data-driven approaches for accelerating the simulation of high-fidelity numerical simulations. Standard projection-based ROM approaches, however, suffer from several drawbacks when applied to the complex nonlinear dynamical systems commonly encountered in science and engineering. These limitations include a lack of stability, accuracy, and sharp a posteriori error estimators. This work addresses these limitations by leveraging multiscale modeling, least-squares principles, and machine learning to develop novel reduced-order modeling approaches, along with data-driven a posteriori error estimators, for dynamical systems. Theoretical and numerical results demonstrate that the two ROM approaches developed in this work - namely the windowed least-squares method and the Adjoint Petrov - Galerkin method - yield substantial improvements over state-of-the-art approaches. Additionally, numerical results demonstrate the capability of the a posteriori error models developed in this work.

97 MATHEMATICS AND COMPUTING↗

Bayesian Optimization Framework for Imperfect Data or Models

Conventional Bayesian optimization methods implicitly assume that the data and model being optimized are “perfect.” This assumption leads to inaccurate posterior probability distribution functions (PDFs) when applied to “imperfect” data or models. The new Bayesian optimization framework presented in this report provides a way to parameterize the effect of imperfections usually encountered in a prior PDF of generalized data or a model on the posterior PDF. The effects of imperfections are parameterized by a set of constraints imposed on the posterior expectation values of deviations between the data and the model and on their covariance matrix elements. A particular set of values for these constraints conveys an evaluator’s best estimate of the effect of imperfections on the corresponding posterior expectation values. When a prior PDF of generalized data is assumed to be normal, an expression for a posterior PDF satisfying an arbitrary set of constraints is derived analytically for linear models. An analogous iterative algorithm is given for nonlinear models. The corresponding posterior PDF should be used to estimate any posterior expectation values in the presence of imperfections parameterized by that set of constraints. A posterior PDF of a conventional Bayesian optimization method is recovered analytically when all evaluator-specified constraints are set to zero (i.e., in the absence of any imperfections). The analytical expressions derived in this report for normal PDFs and linear models were verified numerically by a Metropolis–Hastings Monte Carlo method. The methods presented herein could be applied to any kind of data or models, including differential cross-section data or integral benchmark experiments.

97 MATHEMATICS AND COMPUTING↗

Surrogate multi-fidelity data and model fusion for scientific discovery and uncertainty quantification in Earth System Models

This whitepaper addresses the Earth and Environmental Systems Sciences Division (EESSD)’s predictability challenges in modeling the integrated water cycle and data-model integration. Specifically, it focuses on reducing and characterizing the uncertainty in the representation of process models for unresolved physics, either due to model resolution or limited by the physical under standing or computational efficiency, and the use of observational data for in-situ process parameter optimization within ESM. The described methods may also be used to determine the nature of responses (e.g. strength and direction), and hence to identify critical processes that drive the overall ESM responses to perturbation in the forcing

54 ENVIRONMENTAL SCIENCES↗

WaterTAP3 Model Input Data for NAWI's Eight Source Water Baseline Analyses

This folder contains the input data for the WaterTAP3 model that was used for the eight NAWI (National Alliance for Water Innovation) source water baselines studies published in the Environmental Science and Technology special issue: Technology Baselines and Innovation Priorities for Water Treatment and Supply. There are also eight other separate DAMS submissions, one per source water, that include the model results for the published studies. In this data submission, all model inputs across the eight baselines are included. The data structure and content are described in a README.txt file. For more details on how to use the data in WaterTAP3 please refer to the model documentation and GitHub site found at "WaterTAP3 Github" linked in the submission resources.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

City and County Commercial Building Inventories

The Commercial Building Inventories provide modeled data on commercial building type, vintage, and area for each U.S. city and county. Please note this data is modeled and more precise data may be available through county assessors or other sources. Commercial building stock data is estimated using CoStar Realty Information, Inc. building stock data. This data is part of a suite of state and local energy profile data available at the "State and Local Energy Profile Data Suite" link below and builds on Cities-LEAP energy modeling, available at the "EERE Cities-LEAP Page" link below. Examples of how to use the data to inform energy planning can be found at the "Example Uses" link below.

Array↗

Prototype energy models for data centers

Data centers in the United States consume about two percent of the nation’s electricity. Because heat gains from IT equipment drive cooling demand, data centers offer unique opportunities for energy savings. However, no prototype energy model for data centers is available in the suite of existing U.S. Department of Energy’s Commercial Prototype Building Models. Here we present the development of two new data center prototype models and their implementation in OpenStudio and EnergyPlus. The small-size data center model represents a computer room in a building served by computer room air conditioners (CRACs); while the large-sized model represents stand-alone data centers served by computer room air handlers (CRAHs) with a central chiller plant. For each data center model, two levels of IT equipment (ITE) load density were considered, to cover the wide range of IT power density of data centers: 40 and 100 W/ft 2 (430 and 1076 W/m 2 ) for the computer room, and 100 and 500 W/ft 2 (1076 and 5382 W/m 2 ) for the stand-alone data center. All other assumptions, such as building envelope, lighting, HVAC efficiencies and schedules, were based on the minimal requirements of ASHRAE Standard 90.1 at various vintages. We introduced a novel concept of supply and return air approach temperatures to capture the essential effects of non-uniform airflow and temperature distribution in data centers. The approach temperatures were pre-computed by computational fluid dynamics (CFD) simulations for various configurations of ITE loads and airflow containment management in data centers. A new feature was developed in EnergyPlus to implement the approach temperature method. A case study was conducted to demonstrate the use of the data center models. The two data center models cover all U.S. climate zones and can be used to evaluate energy saving measures for data centers, as well as to support development of data center energy efficiency codes and standards.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Mechanical separations of corn stover anatomical fractions in an integrated feedstock preprocessing system: An experimental and data-driven modeling study

High variabilities of material attributes in lignocellulosic biomass present risks for biofuel and biochemical productions and must be mitigated via preprocessing. Since almost no mechanical device is originally designed for processing biomass, how to operate existing apparatuses with efficient performance has not been investigated extensively. This work presents a study on an integrated screening and air classification to separate cobs and stalks from husks and leaves in corn stover. Prototype machine learning models were developed to assess the feasibility of predicting the process outcome based on the measurable parameters. The models trained upon limited experimental data rendered decent predictive accuracy of yield and purity. The experimental data and modeling results collectively suggest decreasing throughput leads to a higher purity. To the contrary, if throughput increases, a lower purity is likely. A possible trade-off between yield and purity of the separated streams indicates the need for optimal combinations of feedstock size, moisture, and throughput to achieve optimized separations. The results of this study also suggest the need to further improve model predictability by developing more accurate formulations for physics governing the integrated unit operations. To accomplish this, additional experimental data needs to be generated for model training.

09 - BIOMASS FUELS↗