Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Improving and Automating Building Model Data Exchange

There are many instances throughout a project’s lifecycle where there arises a need for quick and accurate risk assessment of building designs. For example, an unexpected design change during construction may necessitate structural engineers to perform a seismic risk assessment on analytical models of the updated building design using high fidelity structural analysis software, such as ANSYS or Abaqus. However, the efficiency of such workflows often depends upon the interoperability of architectural design software and structural analysis software. When the quality of this interoperability is lacking or even non-existent, the efficiency of virtual engineering workflows is hampered, which increases project costs. A McGraw Hill industry survey of professional users of Building Information Modeling (BIM) technologies found that there is high demand for BIM interoperability for structural analysis, but that the value/difficulty ratio is currently too low for practical use. There have been efforts by the academic community to facilitate model data exchange between the architectural design and structural analysis domains, but such solutions have not been widely adopted by industry, face technical challenges, and oftentimes are limited in applicability for users of various BIM software. Therefore, INL is developing capabilities to improve, automate, and generalize model data exchange between architectural BIM software (e.g., Revit) and structural analysis software (e.g., SAP2000, ANSYS). The goal is to help expedite and automate as much of the pre-processing step for creating analytical models in finite element analysis software as reasonably as possible. Such a "BIM-to-FEA" conversion tool should provide direct benefit to end-users through accuracy, automation, quick turn-around, and wide applicability. To generalize the application of this BIM-to-FEA conversion tool and increase its useability among the many different commercial BIM software currently used by industry, the program is being developed with the concept of openBIM. OpenBIM is the application of non-proprietary, open data standards that allow for BIM model data exchange in a format that is accessible, retainable, and useable for all users. The most widely used open, non-proprietary data exchange format for BIM is the Industry Foundation Classes (IFC) schema. IFC is developed by buildingSMART international and is ISO certified (ISO 16739-1:2018). The BIM-to-FEA conversion tool is being developed for compatibility with typical commercial building designs of steel framed structures. The tool is currently capable of importing architectural BIM data of framed building structures, recognizing and extracting the aspects of the model that are required for structural analysis, adjusting the connectivity of frame members, and finally exporting to an analytical model stored in the IFC format. The exported IFC analytical model can then be imported into various openBIM compliant software, such as SAP2000. Such capabilities have already been tested on commercial software, as shown above, and continue to be improved. Work is underway to test the conversion on various commercial BIM software, develop a user-friendly interface, incorporate the program into the broader DeepLynx data warehouse project being developed by INL, and to eventually open-source the tool for the benefit of the community. Future development of the tool envisions the ability for efficient iterative risk assessment of generative building designs, all within a workflow utilizing open-source tools. One such open-source tool will be MOOSE, an advanced finite element analysis tool developed at INL. The conversion tool will also branch out from typical commercial building designs and will aim to incorporate nuclear construction. The aim will be to convert both structural and non-structural components of nuclear facilities, such as curved concrete containment structures and piping systems, respectively.

97 MATHEMATICS AND COMPUTING↗

Evaluating alternative ebullition models for predicting peatland methane emission and its pathways via data–model fusion

Abstract. Understanding the dynamics of peatland methane (CH4) emissions and quantifying sources of uncertainty in estimating peatland CH4 emissions are critical for mitigating climate change. The relative contributions of CH4 emission pathways through ebullition, plant-mediated transport, and diffusion, together with their different transport rates and vulnerability to oxidation, determine the quantity of CH4 to be oxidized before leaving the soil. Notwithstanding their importance, the relative contributions of the emission pathways are highly uncertain. In particular, the ebullition process is more uncertain and can lead to large uncertainties in modeled CH4 emissions. To improve model simulations of CH4 emission and its pathways, we evaluated two model structures: (1) the ebullition bubble growth volume threshold approach (EBG) and (2) the modified ebullition concentration threshold approach (ECT) using CH4 flux and concentration data collected in a peatland in northern Minnesota, USA. When model parameters were constrained using observed CH4 fluxes, the CH4 emissions simulated by the EBG approach (RMSE = 0.53) had a better agreement with observations than the ECT approach (RMSE = 0.61). Further, the EBG approach simulated a smaller contribution from ebullition but more frequent ebullition events than the ECT approach. The EBG approach yielded greatly improved simulations of pore water CH4 concentrations, especially in the deep soil layers, compared to the ECT approach. When constraining the EBG model with both CH4 flux and concentration data in model–data fusion, uncertainty of the modeled CH4 concentration profiles was reduced by 78 % to 86 % in comparison to constraints based on CH4 flux data alone. The improved model capability was attributed to the well-constrained parameters regulating the CH4 production and emission pathways. Our results suggest that the EBG modeling approach better characterizes CH4 emission and underlying mechanisms. Moreover, to achieve the best model results both CH4 flux and concentration data are required to constrain model parameterization.

59 BASIC BIOLOGICAL SCIENCES↗

Model data for a watershed-scale study in the Portage River Basin (OH) examining the effects of subsurface drainage on the hydrologic response of an agricultural watershed.

This study builds on Rathore et al. (2024, WRR) and investigates the role of artificial tile-drainage on various aspects of watershed hydrological response, with a particular focus on peakflow. The model-data for the original modeling-focused paper (Rathore et al., 2024, WRR) is archived at Rathore et al. (2024, ESS-DIVE). Hence, this model-data archive provides scripts that are specific to this study that includes model updates, processing and analysis scripts. For details and models files of original model, readers are referred to Rathore et al. (2024, ESS-DIVE). The key difference between the model configuration in this study and Rathore et al. (2024, WRR) is that the tile drains are applied to the entire domain, to study the impact of tile-drains on different aspects of hydrological response. Additional scenario considering intensified precipitation after a dry period was also simulated. The Watershed Workflow package is implemented in Python3. The Jupyter notebooks can be executed through multiple open-source tools, for example, Anaconda Jupyter Lab, VS Studio Code, etc. Other data files include CSV and HDF5 files, which can be read through Python scripts.

54 ENVIRONMENTAL SCIENCES↗

Model data for numerical evaluation of photosensitive tracers as a strategy for separating surface and subsurface transient storage in streams

This model-data archive pertains to a study aimed at the numerical evaluation of the photosensitive tracers as a potential strategy for separating the effects of surface and hyporheic storage zones (SSZs and HSZs, respectively). Separating effects of SSZ and HSZ are important for accurately representing stream function as HSZs and SSZs expose solutes to significantly different biogeochemical conditions. We perform numerical experiments using a multiscale reactive transport model for stream corridors implemented in ATS code, which allows for representing multiple storage zones with their respective travel time distributions and biogeochemistry. For each of the numerical experiment, we provide in this model-data archive python wrapper to drive ATS (forward_model.py), synthetic observation (synthetic_btc.py, BTC_observed.csv), forward model files including ATS input (multiscale_transport.tpl), PFLOTRAN inputs (reactions_channel.tpl, reactions_hz.tpl, reactions_sz.tpl), MCMC (mcmc_run.py), predictive uncertainty (pred_uncert.py, BTCs_simulated_day.csv, BTCs_simulated_night.csv), MCMC outputs (tracer_test_GR.npy, tracer_test_logps.npy, tracer_test_parameters.npy) and Jupyter notebook for post-processing and visualization (post-processing.ipynb). For the denitrification application, ATS input file (denitrification_multisubgrid.tpl), PFLOTRAN input files (denitrification_channel.in, denitrification_hz.in), predictive uncertainty (pred_uncert_denitrification.py, BTCs_simulated_DO.csv, BTCs_simulated_DOC.csv, BTCs_simulated_Nitrate.csv).

54 ENVIRONMENTAL SCIENCES↗

Model-Data for Joint Estimation of Biogeochemical Model Parameters from Multiple Experiments: A Bayesian Approach Applied to Mercury Methylation

This modeling archive supports the manuscript submitted for publication in the Environmental Modeling and Software. This study is supported by ORNL-SFA and IDEAS-Watershed. This study aims to improve calibration of complex biogeochemical models using datasets from multiple experiments targeting specific subprocesses. The proposed Bayesian joint-fitting scheme calibrates the entire biogeochemical model in one go using all the available datasets and estimate parameter uncertainties using Markov Chain Monte Carlo (MCMC). This allows for complete propagation of uncertainties and utilization of the information shared between different datasets. Mapping joint distribution of parameters guides model improvement by identifying null spaces in the parameter space. This archive contains files used to perform MCMC, post-process outputs and visualize results.

East Fork Poplar Creek↗

A hybrid data–model approach to map soil thickness in mountain hillslopes

Abstract. Soil thickness plays a central role in the interactions between vegetation, soils, and topography, where it controls the retention and release of water, carbon, nitrogen, and metals. However, mapping soil thickness, here defined as the mobile regolith layer, at high spatial resolution remains challenging. Here, we develop a hybrid model that combines a process-based model and empirical relationships to estimate the spatial heterogeneity of soil thickness with fine spatial resolution (0.5 m). We apply this model to two aspects of hillslopes (southwest- and northeast-facing, respectively) in the East River watershed in Colorado. Two independent measurement methods – auger and cone penetrometer – are used to sample soil thickness at 78 locations to calibrate the local value of unconstrained parameters within the hybrid model. Sensitivity analysis using the hybrid model reveals that the diffusion coefficient used in hillslope diffusion modeling has the largest sensitivity among all input parameters. In addition, our results from both sampling and modeling show that, in general, the northeast-facing hillslope has a deeper soil layer than the southwest-facing hillslope. By comparing the soil thickness estimated between a machine-learning approach and this hybrid model, the hybrid model provides higher accuracy and requires less sampling data. Modeling results further reveal that the southwest-facing hillslope has a slightly faster surface soil erosion rate and soil production rate than the northeast-facing hillslope, which suggests that the relatively less dense vegetation cover and drier surface soils on the southwest-facing slopes influence soil properties. With seven parameters in total for calibration, this hybrid model can provide a realistic soil thickness map with a relatively small amount of sampling dataset comparing to machine-learning approach. Integrating process-based modeling and statistical analysis not only provides a thorough understanding of the fundamental mechanisms for soil thickness prediction but also integrates the strengths of both statistical approaches and process-based modeling approaches.

58 GEOSCIENCES↗

Model Data for the Mesh Convergence Study Demonstrating Benefits of Mixed-polyhedral Mesh in Integrated Hydrology Simulations

This archived model data is related to a study introducing a unique method that employs a stream-aligned mixed-polyhedral mesh to effectively and accurately represent river valleys, stream corridors, and narrow engineered channels in integrated hydrology simulations. The study finds that utilizing stream-aligned mixed-polyhedral meshes in integrated hydrology simulations achieves accuracy on par with a finely refined TIN-based mesh while markedly diminishing computational costs. This archive contains scripts and data files needed to generate the ATS model input, including mesh and ATS input files, for all mesh scenarios using the Watershed Workflow package. Additionally, this archive also provides key outputs from the model simulations that are used in the analysis and post-processing scripts to reproduce figures in the manuscript. The Watershed Workflow package is implemented in Python3. The Jupyter notebooks can be executed through multiple open-source tools, for example, Anaconda Jupyter Lab, VS Studio Code, etc. Other data files include CSV and HDF5 files, which can be read through Python scripts. The input files for the ATS model, open-source integrated hydrology, and transport model are in XML format and can be edited in any commonly used text editors.

54 ENVIRONMENTAL SCIENCES↗

Scalable Volume Visualization for Big Scientific Data Modeled by Functional Approximation

Considering the challenges posed by the space and time complexities in handling extensive scientific volumetric data, various data representations have been developed for the analysis of large-scale scientific data. Multivariate functional approximation (MFA) is an innovative data model designed to tackle substantial challenges in scientific data analysis. It computes values and derivatives with high-order accuracy throughout the spatial domain, mitigating artifacts associated with zero- or first-order interpolation. However, the slow query time through MFA makes it less suitable for interactively visualizing a large MFA model. In this work, we develop the first scalable interactive volume visualization pipeline, MFA-DVV, for the MFA model encoded from large-scale datasets. Our method achieves low input latency through distributed architecture, and its performance can be further enhanced by utilizing a compressed MFA model while still maintaining a high-quality rendering result for scientific datasets. We conduct comprehensive experiments to show that MFA-DVV can decrease the input latency and achieve superior visualization results for big scientific data compared with existing approaches.

big scientific dataset↗

Robust Online Sequential RVFLNs for Data Modeling of Dynamic Time-Varying Systems with Application of an Ironmaking Blast Furnace

In a world where the increasing complexity of modern industrial processes brings difficulties for accurate mathematical modeling, taking advantage of data has become an efficient solution to complex dynamic process modeling issue. In this paper, we develop a novel robust online sequential version of random vector functional-link networks (RVFLNs) for data-driven modeling of dynamic time-varying system and applied it in a blast furnace (BF) ironmaking process. First, to overcome the time-varying dynamics of process and to enable the RVFLNs to learn online with avoiding data saturation, an improved online sequential version of RVFLNs (OS-RFVLNs) is first presented by online sequential learning with forgetting factor. This improved OS-RVFLNs algorithm is not only suitable for the real-time and large data transfer situation, but also can adjust the sensitivity of the algorithm to different samples with the help of the introduced forgetting factor. Second, since the output weights of the improved OS-RVFLNs as well as other RVFLNs algorithms are obtained by the least squares approach, a robustness problem may occur when the training dataset is contaminated with various outliers. To solve this problem, a Cauchy distribution weighted M-estimator is introduced to improve the robustness of the improved OS- RVFLNs. For this proposed robust OS-RVFLNs (R-OS- RVFLNs), since the weights of different outlier data are properly determined by the Cauchy distribution function, their corresponding contribution on modeling can be properly distinguished. Thus robust and better modeling results can be achieved. Experiments using actual industrial data of BF ironmaking process and comparative studies have demonstrated that the proposed method produces a better estimation accuracy and stronger robustness than other methods.

Blast furnace (BF), Modelling, Dynamic systems↗

Transforming the study of organisms: Phenomic data models and knowledge bases

The rapidly decreasing cost of gene sequencing has resulted in a deluge of genomic data from across the tree of life; however, outside a few model organism databases, genomic data are limited in their scientific impact because they are not accompanied by computable phenomic data. The majority of phenomic data are contained in countless small, heterogeneous phenotypic data sets that are very difficult or impossible to integrate at scale because of variable formats, lack of digitization, and linguistic problems. One powerful solution is to represent phenotypic data using data models with precise, computable semantics, but adoption of semantic standards for representing phenotypic data has been slow, especially in biodiversity and ecology. Some phenotypic and trait data are available in a semantic language from knowledge bases, but these are often not interoperable. In this review, we will compare and contrast existing ontology and data models, focusing on nonhuman phenotypes and traits. We discuss barriers to integration of phenotypic data and make recommendations for developing an operationally useful, semantically interoperable phenotypic data ecosystem.

59 BASIC BIOLOGICAL SCIENCES↗

PV Performance Modeling - Data and Resources

The Photovoltaic (PV) Performance Modeling Collaborative (PVPMC) organized a blind PV performance modeling intercomparison to allow PV modelers to blindly test their models and modeling ability against real system data. Measured weather and irradiance data were provided along with detailed descriptions of PV systems from two locations (Albuquerque, New Mexico, USA and Roskilde, Denmark). Participants were asked to simulate the plane-of-array irradiance, module temperature, and DC power output from six systems and submit their results to Sandia for processing. This dataset includes seven MS-Excel sheets with instructions, notes and all necessary data (weather, irradiance, temperature, power) used for the data analysis of the blind modeling comparison. The hourly data represent six different systems from Albuquerque, NM and Roskilde, Denmark over a period of one year. These data are useful for PV performance model validation studies.

14 SOLAR ENERGY↗

Guidelines for Publicly Archiving Terrestrial Model Data to Enhance Usability, Intercomparison, and Synthesis

Scientific communities are increasingly publishing data to evaluate, accredit, and build on published research. However, guidelines for curating data for publication are sparse for model-related research, limiting the usability of archived simulation data. In particular, there are no established guidelines for archiving data related to terrestrial models that simulate land processes and their coupled interactions with climate. Terrestrial modelers have a unique set of challenges when publishing data due to the diversity of scientific domains, research questions, and the types and scales of simulations. Researchers in the U.S. Department of Energy’s (DOE) projects use a variety of multiscale models to advance robust predictions of terrestrial and subsurface ecosystem processes. Here, we synthesize archiving needs for data associated with different DOE models, and provide guidelines for publishing terrestrial model data components following FAIR (Findable, Accessible, Interoperable, Reusable) principles. The guidelines recommend archiving model inputs and testing data used in final simulation runs along with associated codes, workflow scripts, and metadata in public repositories. Researchers should consider archiving model outputs if they are within the storage limits of the repository. We also provide considerations for how to bundle files into different data publications with citable digital object identifiers. Finally, we identify repository features and tools that would enable storage and reuse of model data. Given the diversity of DOE terrestrial models, these guidelines are transferable to other model types and will enable efficient reuse of simulation data for purposes such as model intercomparisons, initialization, benchmarking, synthesis, and comparisons with field observations.

58 GEOSCIENCES↗

The Role of Snowmelt and Subsurface Heterogeneity in Headwater Hydrology of a Mountainous Catchment in Colorado: A Model‐Data Integration Approach

Mountainous headwater streams are sustained by both snowmelt‐driven streamflow and groundwater discharge in the Upper Colorado River Basin. However, predicting headwater stream discharge magnitude and peak flow timing is challenging in mountainous terrains, where snowmelt rates vary with vegetation type and elevation, and heterogeneous subsurface physical properties influence groundwater storage and its release. We used a model‐data integration approach to investigate the roles of snowmelt and subsurface structure in stream discharge and groundwater level. We ran an ensemble of 100 integrated surface‐subsurface hydrologic models for a mountainous headwater catchment near Crested Butte, Colorado, USA. We also evaluated and calibrated these models against observed data sets, including snow depth measurements using distributed temperature probes, stream discharge, and groundwater levels. Calibration with multiple data sources using neural density estimators has further constrained uncertainty in subsurface properties and snowmelt rates. Results indicated that observed slower snowmelt rates in evergreen forests delayed the peak flow and baseflow onset. In upstream areas with lower subsurface permeability, water was stored within the subsurface but was not released as interflow or shallow groundwater flow, and thereby not contributing to downstream streamflow during recession limb periods. Double peaks in groundwater occurred in areas with spatial subsurface heterogeneity, in our case due to the contrast between granodiorite and Mancos shale. These process‐based insights into groundwater and snowmelt dynamics in mountainous headwaters will help improve predictions of headwater hydrology.

Wang, Lijing [University of Connecticut, Storrs, C↗

Evolution of the ATLAS event data model for the HL-LHC

The upcoming high-luminosity run of the CERN Large Hadron Collider (HL-LHC) will yield an unprecedented volume of data. In order to process this data, the ATLAS collaboration is evolving its offline software to be able to use heterogeneous resources such as graphical processing units (GPUs) and field-programmable gate arrays (FPGAs). To reduce conversion overheads, the event data model (EDM) should be compatible with the requirements of these resources. While the ATLAS EDM has long allowed representing data as a structure of arrays, further evolution of the EDM can enable more efficient sharing of data between CPU and GPU resources. Some of this work will be summarized here, including extensions to allow controlling how memory for event data is allocated and the implementation of jagged vectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data-model files associated with the manuscript titled "The Importance of Explicitly Representing the Streambed in Watershed Models" (Shuai et al., 2023 HP)

This data package contains the model inputs and outputs used in the manuscript titled "The Importance of Explicitly Representing the Streambed in Watershed Models" (Shuai et al., 2023 HP). The data.zip file contains the data used to drive the Advanced Terrestrial Simulator (ATS) model simulations. The model.zip file contains the XML input file for ATS. The notebook.zip file contains the Jupyter notebooks for pre- and post- processing model results. The figures.zip file contains the raw figures associated with the manuscript. We generated this data package in support of the manuscript and research reproducibility. The background of this study is that the streambed itself has not been explicitly represented in watershed models, although the streambed characteristics are significantly different from those of its surrounding soil. We aim to answer the following questions: 1) How do streambed properties including hydraulic conductivity, thickness and resolution impact the groundwater-surface water exchange fluxes across the streambed? 2) Is an explicit representation of streambed important in watershed modeling?; 3) Is a high-resolution streambed needed for watershed simulations?

54 ENVIRONMENTAL SCIENCES↗

BIM Interoperability Tool for Improving IFC-based Model Data Exchange Between Architectural Design and Structural Analysis for Linear Piping Components

There is a growing need in the AEC industry for the digitalization of model-based data exchange in BIM workflows. However, users continue to face difficulties exchanging data between BIM-based computer-aided design (CAD) and computer-aided engineering (CAE) software, even with the open, non-proprietary data exchange format called Industry Foundation Classes (IFC). Proposed solutions in academic research focus primarily on building systems, with comparatively little attention to piping models. Therefore, this paper introduces an interoperability tool for enabling piping model data exchange between architectural and structural analysis domains. The tool is tested using a piping model created in Autodesk Revit and shows marked improvement for IFC-based CAD-to-CAE interoperability over existing practice.

97 MATHEMATICS AND COMPUTING↗

Machine-learning interatomic potentials for interfaces in all-solid-state batteries: Perspectives on training data, model selection, and validation

Interfaces play a pivotal role in dictating the performance and reliability of all-solid-state batteries (ASSBs), where complex electro-chemo-mechanical phenomena at grain boundaries (GBs) and interfaces can lead to degradation and failure. Traditional atomistic simulation methods, such as first-principles calculations and classical molecular dynamics, face limitations in modeling these interfaces due to either high computational cost or insufficient transferability to the diverse atomic environments evolving at interfaces. Machine-learning interatomic potentials (MLIPs) have emerged as a transformative approach, enabling large-scale, high-accuracy simulations of disordered and chemically complex systems by leveraging the predictability of machine learning models trained on first-principles data. Recent applications of MLIPs have demonstrated their ability to capture intricate behaviors at ASSB interfaces, including ion transport, interfacial evolution, and degradation mechanisms, with accuracy and efficiency unattainable by conventional methods. This prospective paper presents comprehensive analysis and practical guidance for MLIP development for GBs and interfaces in ASSBs, with a focus on three key pillars: data generation, model selection, and validation. Here, we review the current state of MLIP applications for GBs and interfaces in both general and ASSB-specific materials, highlighting best practices and challenges in constructing diverse and representative datasets, choosing appropriate machine learning architectures, and rigorously validating model performance. We also discuss emerging strategies and opportunities for improved reliability and efficiency of MLIPs to simulate realistic interfaces in ASSBs.

Energy - Storage↗

PFLOTRAN modeling data and scripts associated with “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics” submitted to Water Resources Research (Terry et al. 2025). The data package contains the groundwater modeling dataset from PFLOTRAN software. It includes the python script for mesh generation, boundary condition setting, PFLOTRAN input deck formation and postprocessing. It couples groundwater flow and species transport for Hanford Reach river corridor and pipelines the model generation and processing. This model can be used to easily generate the model and analysis for Hanford site. It can also be adjusted to other hydrologic area with ease. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package consists of 6 folders: (1) “data” contains all necessary data as input and intermediate data for processing; (2) “mesh” contains all mesh related files to generate mesh in Hanford Reach river corridor; (3) “model_run” contains the generated script for PFLOTRAN modeling; (4) “notebooks” contains all the Python script to generate the model; (5) “output” contains all the output from the computation; (6) “postprocessing” contains the Python script to generate scientific figure for manuscript. All files are .csv (comma-separated values), .h5 (HDF5 format), .in (input files), .ipynb (Jupyter notebooks), .p (Python pickle), .png (images), .PNG (images), .py (Python scripts), .pyc (Python bytecode), .r (R scripts), .sh (shell scripts), .txt (text files), .vtu (3D mesh/visualization format), .xz (compressed archive), or .zip (compressed archive).

54 ENVIRONMENTAL SCIENCES↗