Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Scalable Volume Visualization for Big Scientific Data Modeled by Functional Approximation

Considering the challenges posed by the space and time complexities in handling extensive scientific volumetric data, various data representations have been developed for the analysis of large-scale scientific data. Multivariate functional approximation (MFA) is an innovative data model designed to tackle substantial challenges in scientific data analysis. It computes values and derivatives with high-order accuracy throughout the spatial domain, mitigating artifacts associated with zero- or first-order interpolation. However, the slow query time through MFA makes it less suitable for interactively visualizing a large MFA model. In this work, we develop the first scalable interactive volume visualization pipeline, MFA-DVV, for the MFA model encoded from large-scale datasets. Our method achieves low input latency through distributed architecture, and its performance can be further enhanced by utilizing a compressed MFA model while still maintaining a high-quality rendering result for scientific datasets. We conduct comprehensive experiments to show that MFA-DVV can decrease the input latency and achieve superior visualization results for big scientific data compared with existing approaches.

big scientific dataset↗

PV Performance Modeling - Data and Resources

The Photovoltaic (PV) Performance Modeling Collaborative (PVPMC) organized a blind PV performance modeling intercomparison to allow PV modelers to blindly test their models and modeling ability against real system data. Measured weather and irradiance data were provided along with detailed descriptions of PV systems from two locations (Albuquerque, New Mexico, USA and Roskilde, Denmark). Participants were asked to simulate the plane-of-array irradiance, module temperature, and DC power output from six systems and submit their results to Sandia for processing. This dataset includes seven MS-Excel sheets with instructions, notes and all necessary data (weather, irradiance, temperature, power) used for the data analysis of the blind modeling comparison. The hourly data represent six different systems from Albuquerque, NM and Roskilde, Denmark over a period of one year. These data are useful for PV performance model validation studies.

14 SOLAR ENERGY↗

Guidelines for Publicly Archiving Terrestrial Model Data to Enhance Usability, Intercomparison, and Synthesis

Scientific communities are increasingly publishing data to evaluate, accredit, and build on published research. However, guidelines for curating data for publication are sparse for model-related research, limiting the usability of archived simulation data. In particular, there are no established guidelines for archiving data related to terrestrial models that simulate land processes and their coupled interactions with climate. Terrestrial modelers have a unique set of challenges when publishing data due to the diversity of scientific domains, research questions, and the types and scales of simulations. Researchers in the U.S. Department of Energy’s (DOE) projects use a variety of multiscale models to advance robust predictions of terrestrial and subsurface ecosystem processes. Here, we synthesize archiving needs for data associated with different DOE models, and provide guidelines for publishing terrestrial model data components following FAIR (Findable, Accessible, Interoperable, Reusable) principles. The guidelines recommend archiving model inputs and testing data used in final simulation runs along with associated codes, workflow scripts, and metadata in public repositories. Researchers should consider archiving model outputs if they are within the storage limits of the repository. We also provide considerations for how to bundle files into different data publications with citable digital object identifiers. Finally, we identify repository features and tools that would enable storage and reuse of model data. Given the diversity of DOE terrestrial models, these guidelines are transferable to other model types and will enable efficient reuse of simulation data for purposes such as model intercomparisons, initialization, benchmarking, synthesis, and comparisons with field observations.

58 GEOSCIENCES↗

The Role of Snowmelt and Subsurface Heterogeneity in Headwater Hydrology of a Mountainous Catchment in Colorado: A Model‐Data Integration Approach

Mountainous headwater streams are sustained by both snowmelt‐driven streamflow and groundwater discharge in the Upper Colorado River Basin. However, predicting headwater stream discharge magnitude and peak flow timing is challenging in mountainous terrains, where snowmelt rates vary with vegetation type and elevation, and heterogeneous subsurface physical properties influence groundwater storage and its release. We used a model‐data integration approach to investigate the roles of snowmelt and subsurface structure in stream discharge and groundwater level. We ran an ensemble of 100 integrated surface‐subsurface hydrologic models for a mountainous headwater catchment near Crested Butte, Colorado, USA. We also evaluated and calibrated these models against observed data sets, including snow depth measurements using distributed temperature probes, stream discharge, and groundwater levels. Calibration with multiple data sources using neural density estimators has further constrained uncertainty in subsurface properties and snowmelt rates. Results indicated that observed slower snowmelt rates in evergreen forests delayed the peak flow and baseflow onset. In upstream areas with lower subsurface permeability, water was stored within the subsurface but was not released as interflow or shallow groundwater flow, and thereby not contributing to downstream streamflow during recession limb periods. Double peaks in groundwater occurred in areas with spatial subsurface heterogeneity, in our case due to the contrast between granodiorite and Mancos shale. These process‐based insights into groundwater and snowmelt dynamics in mountainous headwaters will help improve predictions of headwater hydrology.

Wang, Lijing [University of Connecticut, Storrs, C↗

Evolution of the ATLAS event data model for the HL-LHC

The upcoming high-luminosity run of the CERN Large Hadron Collider (HL-LHC) will yield an unprecedented volume of data. In order to process this data, the ATLAS collaboration is evolving its offline software to be able to use heterogeneous resources such as graphical processing units (GPUs) and field-programmable gate arrays (FPGAs). To reduce conversion overheads, the event data model (EDM) should be compatible with the requirements of these resources. While the ATLAS EDM has long allowed representing data as a structure of arrays, further evolution of the EDM can enable more efficient sharing of data between CPU and GPU resources. Some of this work will be summarized here, including extensions to allow controlling how memory for event data is allocated and the implementation of jagged vectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data-model files associated with the manuscript titled "The Importance of Explicitly Representing the Streambed in Watershed Models" (Shuai et al., 2023 HP)

This data package contains the model inputs and outputs used in the manuscript titled "The Importance of Explicitly Representing the Streambed in Watershed Models" (Shuai et al., 2023 HP). The data.zip file contains the data used to drive the Advanced Terrestrial Simulator (ATS) model simulations. The model.zip file contains the XML input file for ATS. The notebook.zip file contains the Jupyter notebooks for pre- and post- processing model results. The figures.zip file contains the raw figures associated with the manuscript. We generated this data package in support of the manuscript and research reproducibility. The background of this study is that the streambed itself has not been explicitly represented in watershed models, although the streambed characteristics are significantly different from those of its surrounding soil. We aim to answer the following questions: 1) How do streambed properties including hydraulic conductivity, thickness and resolution impact the groundwater-surface water exchange fluxes across the streambed? 2) Is an explicit representation of streambed important in watershed modeling?; 3) Is a high-resolution streambed needed for watershed simulations?

54 ENVIRONMENTAL SCIENCES↗

BIM Interoperability Tool for Improving IFC-based Model Data Exchange Between Architectural Design and Structural Analysis for Linear Piping Components

There is a growing need in the AEC industry for the digitalization of model-based data exchange in BIM workflows. However, users continue to face difficulties exchanging data between BIM-based computer-aided design (CAD) and computer-aided engineering (CAE) software, even with the open, non-proprietary data exchange format called Industry Foundation Classes (IFC). Proposed solutions in academic research focus primarily on building systems, with comparatively little attention to piping models. Therefore, this paper introduces an interoperability tool for enabling piping model data exchange between architectural and structural analysis domains. The tool is tested using a piping model created in Autodesk Revit and shows marked improvement for IFC-based CAD-to-CAE interoperability over existing practice.

97 MATHEMATICS AND COMPUTING↗

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage↗

PFLOTRAN modeling data and scripts associated with “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics” submitted to Water Resources Research (Terry et al. 2025). The data package contains the groundwater modeling dataset from PFLOTRAN software. It includes the python script for mesh generation, boundary condition setting, PFLOTRAN input deck formation and postprocessing. It couples groundwater flow and species transport for Hanford Reach river corridor and pipelines the model generation and processing. This model can be used to easily generate the model and analysis for Hanford site. It can also be adjusted to other hydrologic area with ease. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package consists of 6 folders: (1) “data” contains all necessary data as input and intermediate data for processing; (2) “mesh” contains all mesh related files to generate mesh in Hanford Reach river corridor; (3) “model_run” contains the generated script for PFLOTRAN modeling; (4) “notebooks” contains all the Python script to generate the model; (5) “output” contains all the output from the computation; (6) “postprocessing” contains the Python script to generate scientific figure for manuscript. All files are .csv (comma-separated values), .h5 (HDF5 format), .in (input files), .ipynb (Jupyter notebooks), .p (Python pickle), .png (images), .PNG (images), .py (Python scripts), .pyc (Python bytecode), .r (R scripts), .sh (shell scripts), .txt (text files), .vtu (3D mesh/visualization format), .xz (compressed archive), or .zip (compressed archive).

54 ENVIRONMENTAL SCIENCES↗

Machine learning-enabled model-data integration for predicting subsurface water storage

Subsurface water storage (SWS) is a key variable of the climate system and a storage component for precipitation and radiation anomalies, inducing persistence in the climate system. It plays a critical role in climate-change projections and can mitigate the impacts of climate change on ecosystems. However, because of the difficult accessibility of the underground, hydrologic properties and dynamics of SWS are poorly known. Direct observations of SWS are limited, and accurate incorporation of SWS dynamics into Earth system land models remains challenging. We propose a machine learning-enabled model-data integration framework to improve the SWS prediction at local to conus scales in a changing climate by leveraging all the available observation and simulation resources, as well as to inform the model development and guide the observation collection. The accurate prediction will enable an optimal decision of water management and land use and improve the ecosystem's resilience to the climate change.

Lu, Dan↗

Model Data Archive for Manuscript Titled "Evaluation of a Coupled Surface–Subsurface Hydrologic Model Using Dense Water‑Level Sensors in a Mixed Urban–Rural Watershed"

This archive provides scripts, input files, and datasets used for the implementation and evaluation of a fully coupled surface–subsurface hydrologic model in the Neches River Basin, southeast Texas. The study uses the Advanced Terrestrial Simulator (ATS) to simulate coupled surface–subsurface hydrologic processes over a mixed urban–rural watershed and evaluates model performance using a dense network of 136 in situ water-level sensors, nine U.S. Geological Survey (USGS) stream gauges, and SSEBop-derived evapotranspiration estimates during the period October 2014–June 2024. The workflow is implemented primarily in Python 3 using the Watershed Workflow package. The Jupyter notebooks can be executed using open-source software such as Anaconda JupyterLab or Visual Studio Code. Other data files include TXT, CSV, XML, SHP, TIF, NetCDF, HDF5, and ExodusII files, which can be processed using the provided Python scripts. ATS input files are provided in XML format and can be edited using any commonly used text editor. This archive contains: *Scripts and input files used to generate the ATS model setup, including watershed discretization, mesh generation, parameter mapping, and model configuration. *Jupyter notebooks used for preprocessing observational data, evaluating streamflow, water levels, and evapotranspiration, computing performance metrics, and generating the figures presented in the manuscript. *ATS simulation outputs and processed observational datasets, including OneRain and DD6 water-level sensors, USGS streamflow observations, GIS data, and supporting spatial datasets used throughout the study.

Dense water-level sensor network↗

Bayesian model-data comparison incorporating theoretical uncertainties

Accurate comparisons between theoretical models and experimental data are critical for scientific progress. However, inferred physical model parameters can vary significantly with the chosen physics model, highlighting the importance of properly accounting for theoretical uncertainties. In this Letter, we present a Bayesian framework that explicitly quantifies these uncertainties by statistically modeling theory errors, guided by qualitative knowledge of a theory’s varying reliability across the input domain. We demonstrate the effectiveness of this approach using two systems: a simple ball drop experiment and multi-stage heavy-ion simulations. In both cases incorporating model discrepancy leads to improved parameter estimates, with systematic improvements observed as additional experimental observables are integrated.

Bayesian methods↗

Utah FORGE: 2024 Discrete Fracture Network Model Data

The Utah FORGE 2024 Discrete Fracture Network (DFN) Model dataset provides a set of files representing discrete fracture network modeling for the FORGE site near Milford, Utah. The dataset includes four distinct DFN model file sets, each corresponding to different time frames and modeling approaches in 2024. These models characterize both natural and induced fractures in the geothermal reservoir, which consists of crystalline granitic and metamorphic rock approximately 8,000 feet below the ground surface. The dataset includes a reference DFN model from February 2024 that incorporates planar fractures and well trajectories, as well as upscaled permeability, porosity, compressibility, and storage values on specified grids. Additionally, there are models based on new microseismic (MEQ) data from May and July 2024, including fracture planes fitted to the latest MEQ catalog datasets, tensile fractures from hydraulic stimulation, and an alternative connected DFN for modeling purposes. Coordinate data is provided in both global and local frames, with detailed instructions on the transformations used to align with principal stress orientations. The dataset also includes notes and calculation files for estimating fracture sizes and differences between various fracture sets. There are subfolders for Global Coordinates and Local Coordinates. To move from the global to the local coordinate frame, fractures and wells were a) rotated 20 degrees counterclockwise looking down about the global point (335376.400482041, 4263189.99998761, 250.093546450195) to better align with the principal stresses; and b) translated by (-335408.68, -4263010.9, 1150). Upscaled permeability values using the _XYZ suffix show directions with respect to the global XYZ coordinate frame, while those using the _IJK suffix are aligned with local coordinate frame.

15 GEOTHERMAL ENERGY↗

Block Island Acoustic Propagation Modeling Data

This dataset contains acoustic propagation model outputs, computational subroutines, and analysis tools for underwater sound propagation in the Block Island region, including parabolic equation (PE) model results, visualization products, and comprehensive modeling software tools.

17 WIND ENERGY↗

Hydrologic Model Data for the East Fork Poplar Creek Watershed Simulated with the Advanced Terrestrial Simulator (ATS): Streamflow and Network Expansion–Contraction Dynamics

This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.

54 ENVIRONMENTAL SCIENCES↗

Northern Pacific Turbulence Intensity Model Data in Observational Space

The dataset archives model-simulated turbulence intensity and meteorological profiles and timeseries at the lidar buoys sites off the coast of California (Humboldt and Morro Bay). The simulated data are interpolated in time and/or space according to observed quantities. The simulations were carried out for the north Pacific region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing.

17 WIND ENERGY↗

Gulf of Mexico Turbulence Intensity Model Data in Observational Space

The dataset archives model-simulated turbulence intensity and meteorological profiles and timeseries at the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars. The simulated data are interpolated in time and/or space according to observed quantities. The simulations were carried out for the Gulf of Mexico region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing.

17 WIND ENERGY↗

Data-model files associated with the manuscript "Modeling the Effects of Wetland Restoration on Coastal Hydrology: A Case Study of Elkhorn Slough Watershed, California"

This package contains the data, simulation setups, notebooks and figures used in “Modeling the Effects of Wetland Restoration on Coastal Hydrology: A Case Study of Elkhorn Slough Watershed, California” (Xu et al., 2025). In this study, we selected Elkhorn Slough, a tidal estuary, in California, to investigate the impact of wetland restoration and sea level rise on coastal hydrology using the process-based coastal hydrologic model, Advanced Terrestrial Simulator (ATS), informed by site-specific data. We designed a novel modeling workflow for incorporating wetland restoration features into land cover and soil properties for the model parameterization. The validation results demonstrate a strong agreement between modeled and observed data. We studied the characteristics of coastal watershed hydrology, then focused on the surface water dynamics at two wetland sites within Elkhorn Slough, a reference site and a restored site. Our simulation results indicate that the restored site successfully maintains surface elevation, resulting in reduced surface inundation. We also examined the impact of wetland restoration under expected sea level rise over the next few decades. The low-lying Yampah Marsh, the reference site, is likely to be inundated due to future sea level rise when highest tides arrive; while a higher percentage of Hester Marsh, the restored site, would retain marsh vegetation in coming decades, regardless of tidal conditions. Our study provides important information for examining the outcome of restoration practices that include surface elevation in tidal wetlands under climate changes.Several files can be found from this data package.1. README.md: This file describes the title, journal, co-authors, abstract, repository structure and model version.2. Simulation_Setups.zip: The file contains the model configuration files (XML format) for ATS. 3. Notebooks.zip: The file contains the Jupyter notebooks for generating the pre- and post-restoration meshes and the meshes of future scenarios. 4. Figures.zip: The file contains the figures used in the manuscript.5. Data.zip: The file contains the data used to drive the model simulations, including watershed and wetlands boundaries, mesh files and references to additional datasets (e.g., meteorological forcing, tidal dataset, DEMs, land cover, soil properties). Also, it contains water level observations at the restored wetland.

54 ENVIRONMENTAL SCIENCES↗