Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “standardized data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

The Marine and Hydrokinetic ToolKit (MHKiT) for Data Quality Control and Analysis [Slides]

The ability to collect, ingest, condition, reduce, quality control, process, visualize, and store data in a standardized way is critical at all stages of Marine Energy (ME) research and technology/project development. MHKiT is an open-source, standardized suite of ME data processing functions that provides the ability to ingest, condition, reduce, quality control, process, visualize and store ME data. MHKiT is developed in both Python and Matlab.

16 TIDAL AND WAVE POWER↗

Fine Temperature Grid Continuous Energy Cross Section Generation for Monte Carlo Analysis of Xe-100 Design

The standard “A Compact ENDF (ACE)” data libraries used by Monte Carlo based reactor physics codes calculations are provided by Los Alamos National Laboratory (LANL) with a temperature interval mostly of 300 K (e.g. 300 K, 600 K, 900 K) for the cross sections and between 100 K and 200 K for the thermal scattering libraries (TSL). However, some codes such as MCNP lack capability to perform on-the-fly temperature interpolation during simulation both for neutron and TSL cross-sections. To evaluate the impact related to Doppler broadening and spectrum shift associated with TSL changes, this paper explores the potential of adopting a temperature grid finer than the ones contained in the standard data libraries. A 50 K temperature grid was employed to quantify the error in neutronics calculations due to temperature grid resolution. This was achieved by comparing the results of this study (50 K temperature interval) against the results obtained with standard data libraries (>100 K temperature interval). While the adopted grid primarily relies on the ENDF/B-VII.1 library, for neutron cross-sections, it utilizes ENDF/B-VIII.0 library for TSL. The analyses confirmed that the accuracy of neutronics calculations is satisfactory when using a 50 K temperature grid. Notably, adopting a 50 K temperature grid, as opposed to standard libraries or coarser temperature grids, could lead to a difference of no more than a few hundred pcm in dk for both fresh fuel and burnt fuel. The most sensitive reaction type to the temperature grid was as expected identified as the capture cross-section of U-238.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

pvlib iotools—Open-source Python functions for seamless access to solar irradiance data

Access to accurate solar resource data is critical for numerous applications, including estimating the yield of solar energy systems, developing radiation models, and validating irradiance datasets. However, lack of standardization in data formats and access interfaces across providers constitutes a major barrier to entry for new users. pvlib python’s iotools subpackage aims to solve this issue by providing standardized Python functions for reading local files and retrieving data from external providers. All functions follow a uniform pattern and return convenient data outputs, allowing users to seamlessly switch between data providers and explore alternative datasets. The pvlib package is community-developed on GitHub: https://github.com/pvlib/pvlib-python. As of pvlib python version 0.9.5, the iotools subpackage supports 12 different datasets, including ground measurement, reanalysis, and satellite-derived irradiance data. The supported ground measurement networks include the Baseline Surface Radiation Network (BSRN), NREL MIDC, SRML, SOLRAD, SURFRAD, and the US Climate Reference Network (CRN). Additionally, satellite-derived and reanalysis irradiance data from the following sources are supported: PVGIS (SARAH & ERA5), NSRDB PSM3, and CAMS Radiation Service (including McClear clear-sky irradiance).

14 SOLAR ENERGY↗

The need for standardization and improved open (meta)data practices in metaproteomics

Metaproteomics enables functional insight into microbial communities by identifying and quantifying proteins in complex samples. Yet, heterogeneous analytical workflows and the lack of standardization across experimental and bioinformatics stages hinder reproducibility and comparability, limiting integration with other omics data. We here present a community-developed reporting checklist tailored to the specific needs of metaproteomics. We also outline current efforts to enable structured and interoperable metadata capture, drawing on standards from proteomics and microbiome research wherever possible. By promoting transparent reporting and advancing metadata practices, our recommendations aim to align metaproteomics more closely with FAIR principles and support reproducible and interoperable research practices.

Armengaud, Jean [Universite Paris-Saclay, France]↗

ndi

The Nuclear Data Interface (NDI) is an application programming interface (API) that allows access to standard nuclear data parameters while hiding the underlying details of the data libraries and their storage. It allows access to multigroup transport data (neutron and gamma), thermonuclear burn data, dosimetry data, production/depletion chain data, radiochemistry data, and secondary neutron multiplicity data. The name NDI refers to both the code and data formats supported by the code.

Saller, Thomas↗

RNTuple: A CMS Perspective

ROOT is planning to move from TTree to RNTuple as the data storage format for HL-LHC in order to, for example, speed up the IO, make the files smaller, and have a modern C++ API. Initially, RNTuple was not planned to support the same set of C++ data structures as TTree supports. CMS has explored the necessary transformations in its standard persistent data types to switch to RNTuple. Many challenges were encountered as alternative data structures were explored. This contribution will discuss the challenges uncovered and how collaboration with the ROOT team allowed them to be overcome. The solution to the challenges allows progressive changes to the CMS data types rather than requiring a sudden change to all data types to be stored in RNTuple. Once the solution was achieved, storage performance comparisons using the CMS data types were possible between RNTuple and TTree. This contribution will also present the results of those comparisons.

Smith, Nick↗

Advances in building data management for building performance standards using the SEED platform

Reducing energy consumption and greenhouse gas emissions in the built environment is a critical step in achieving emission goals to mitigate climate change impacts. Local, federal, and international jurisdictions are deploying several methods to reduce energy and emissions such as voluntary and mandatory benchmarking and building performance standards, requiring building owners to reach energy and emission targets. Jurisdictions leveraging benchmarking and building performance standards require knowledge of the buildings covered; which is a large task due to staffing constraints, limited information on building characteristics and tax parcel data, and the need for advanced data management techniques to align datasets. This paper describes an open-source platform's recent advances to create consistent taxonomies, identify erroneous data, enable auditability, and track building performance. The paper concludes with two use cases on how the platform has been used by jurisdictions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Deep learning to estimate permeability using geophysical data

Time-lapse electrical resistivity tomography (ERT) is a popular geophysical method to estimate three-dimensional (3D) permeability fields from electrical potential difference measurements. Traditional inversion and data assimilation methods are used to ingest this ERT data into hydrogeophysical models to estimate permeability. Due to ill-posedness and the curse of dimensionality, existing inversion strategies provide poor estimates and low resolution of the 3D permeability field. Recent advances in deep learning provide us with powerful algorithms to overcome this challenge. This paper presents a deep learning (DL) framework to estimate the 3D subsurface permeability from time-lapse ERT data. To test the feasibility of the proposed framework, we train DL-enabled inverse models on simulation data. Each measurement in both synthetic and field data is standardized by removing the mean and scaling the time-series to unit variance. This pre-processing step is necessary to bring simulation data closer to field observations. Subsurface process models based on hydrogeophysics are used to generate this synthetic data. Training performed on limited simulation data resulted in the DL model over-fitting. An advanced data augmentation based on mixup is implemented to generate additional training samples to overcome this issue. This mixup technique creates weakly labeled (low-fidelity) samples from strongly labeled (high-fidelity) data. The weakly labeled training data is then used to develop DL-enabled inverse models and reduce over-fitting. As both time-lapse ERT (1133048 features/realization) and 3D permeability (585453 features/realization) data samples are from a high-dimensional space, principal component analysis (PCA) is employed to reduce dimensionality. Encoded ERT and encoded permeability are generated using the trained PCA estimators. A deep neural network is then trained to map the encoded ERT to encoded permeability. This mixup training and unsupervised learning allowed us to build a fast and reasonably accurate DL-based inverse model under limited simulation data. Results show that proposed weak supervised learning can capture salient spatial features in the 3D permeability field. Quantitatively, the average mean squared error (in terms of the natural log) on the strongly labeled training, validation, and test datasets is less than 0.5. The R 2 -score (global metric) is greater than 0.75, and the percent error in each cell (local metric) is less than 10%. Finally, an added benefit in terms of computational cost is that the proposed DL-based inverse model is at least O(10 4 ) times faster than running a forward model once it is trained. Data generation, DL model training, and hyperparameter tuning to identify optimal neural network architectures utilized high-performance computing resources while the DL inference is performed on a standard laptop. Approximately, O(10 5 ) processor hours are used for generating data and DL tuning and training. We acknowledge that the data generation and DL model development are expensive. But once a DL model is trained, it can be re-used for inversion rapidly for the given system, with set physics and domain. Note that traditional inversion may require multiple forward model simulations (e.g., in the order of 10 to 1000), which are very expensive. This computational savings ≈ O(10 5 ) – O(10 7 )) makes the proposed DL-based inverse model attractive for subsurface imaging and real-time ERT monitoring applications due to fast and yet reasonably accurate estimations of permeability field.

58 GEOSCIENCES↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

Analytic continuation of noisy data using Adams Bashforth residual neural network

We propose a data-driven learning framework for the analytic continuation problem in numerical quantum many-body physics. Designing an accurate and efficient framework for the analytic continuation of imaginary time using computational data is a grand challenge that has hindered meaningful links with experimental data. The standard Maximum Entropy (MaxEnt)-based method is limited by the quality of the computational data and the availability of prior information. Also, the MaxEnt is not able to solve the inversion problem under high level of noise in the data. Here we introduce a novel learning model for the analytic continuation problem using a Adams-Bashforth residual neural network (AB-ResNet). Additionally, the advantage of this deep learning network is that it is model independent and, therefore, does not require prior information concerning the quantity of interest given by the spectral function. More importantly, the ResNet-based model achieves higher accuracy than MaxEnt for data with higher level of noise. Finally, numerical examples show that the developed AB-ResNet is able to recover the spectral function with accuracy comparable to MaxEnt where the noise level is relatively small.

97 MATHEMATICS AND COMPUTING↗

Verification of the 3-Region Advanced Test Reactor MCNP Model

The verification of the 3-region homogenized fuel Advanced Test Reactor MCNP model. The 3-region model was compared to the 19-plate model found in the 94-CIC report. Flux tallies, energy deposition tallies, and quarter core mesh tallies were used to compare the two models. The 3-region model needed updating in order to make good comparisons between the models. The percent error from the flux and energy deposition tallies data shows that experiment positions inside the flux trap have higher errors than positions outside the fuel ring. The standard deviation data obtained from the mesh tallies shows that the two models agree within two standard deviations throughout the reactor. It is concluded that the model works adequately for what it is used for.

99 GENERAL AND MISCELLANEOUS↗

Verification of the 3-Region Advanced Test Reactor MCNP Model

The verification of the 3-region homogenized fuel Advanced Test Reactor MCNP model. The 3-region model was compared to the 19-plate model found in the 94-CIC report. Flux tallies, energy deposition tallies, and quarter core mesh tallies were used to compare the two models. The 3-region model needed updating in order to make good comparisons between the models. The percent error from the flux and energy deposition tallies data shows that experiment positions inside the flux trap have higher errors than positions outside the fuel ring. The standard deviation data obtained from the mesh tallies shows that the two models agree within two standard deviations throughout the reactor. It is concluded that the model works adequately for what it is used for.

99 GENERAL AND MISCELLANEOUS↗

A Case Study of Multimodal, Multi-institutional Data Management for the Combinatorial Materials Science Community

Although the convergence of high-performance computing, automation, and machine learning has significantly altered the materials design timeline, transformative advances in functional materials and acceleration of their design will require addressing the deficiencies that currently exist in materials informatics, particularly a lack of standardized experimental data management. The challenges associated with experimental data management are especially true for combinatorial materials science, where advancements in automation of experimental workflows have produced datasets that are often too large and too complex for human reasoning. The data management challenge is further compounded by the multimodal and multi-institutional nature of these datasets, as they tend to be distributed across multiple institutions and can vary substantially in format, size, and content. Furthermore, modern materials engineering requires the tuning of not only composition but also of phase and microstructure to elucidate processing–structure–property–performance relationships. To adequately map a materials design space from such datasets, an ideal materials data infrastructure would contain data and metadata describing (i) synthesis and processing conditions, (ii) characterization results, and (iii) property and performance measurements. In this work, we present a case study for the low-barrier development of such a dashboard that enables standardized organization, analysis, and visualization of a large data lake consisting of combinatorial datasets of synthesis and processing conditions, X-ray diffraction patterns, and materials property measurements generated at several different institutions. While this dashboard was developed specifically for data-driven thermoelectric materials discovery, we envision the adaptation of this prototype to other materials applications, and, more ambitiously, future integration into an all-encompassing materials data management infrastructure.

36 MATERIALS SCIENCE↗

LungMAP Portal Ecosystem: Systems-level Exploration of the Lung

An improved understanding of the human lung necessitates advanced systems models informed by an ever-increasing repertoire of molecular omics, cellular, imaging, and pathological datasets. To centralize and standardize information across broad lung research efforts we expanded the LungMAP.net website into a new gateway portal. This portal connects a broad spectrum of research networks, bulk and single-cell multi-omics data and a diverse collection of image data that span mammalian lung development, and disease. The data are standardized across species and technologies using harmonized data and metadata models that leverage recent advances including those from the Human Cell Atlas, diverse ontologies, and the LungMAP CellCards initiative. To cultivate future discoveries, we have aggregated a diverse collection of single-cell atlases for multiple species (human, rhesus, mouse), to enable consistent queries across technologies, cohorts, age, disease, and drug treatment. These atlases are provided as independent and integrated queryable datasets, with an emphasis on dynamic visualization, figure generation, re-analysis, cell-type curation, and automated reference-based classification of user-provided single-cell genomics datasets (Azimuth). Finally, as this resource grows, we intend to increase the breadth of available interactive interfaces, supported data types, data portals and datasets from LungMAP and external research efforts.

59 BASIC BIOLOGICAL SCIENCES↗

Data Interfaces for Automated Vehicle Services - A Municipality Perspective

As Automated Vehicle (AV) services proliferate, data sharing between AV operators and municipal agents is assuming greater importance. Information on the dynamic nature of the road system such as incidents to avoid, weather hazards (such as flooding), construction and detours, as well as active safety concerns (e.g. - riots) is important for AV operators. Such information cannot be directly sensed from a vehicle's sensor array, but instead must be communicated in a timely and trustworthy channel. Municipalities are interested in pushing this information to AV operators to support emergency response efforts, reduce traffic in construction zones, and generally improve operation of the system. Similarly, information on vehicle safety such as disengagements, as well as critical information on the use of roadway system (trips, origin and destination patterns) are important performance factors for municipalities to understand utilization and plan for appropriate infrastructure. As mobility shifts to on-demand options, the need for safe and coordinated pick-up and drop-off zones will increase (potentially reducing parking needs). For all of these reasons, communication flows between AV operators and municipalities are becoming increasingly important. This paper investigates the functions, emerging practices and protocols for sharing of such critical data, and identifies gaps in and challenges in existing practices. Additionally, case studies are used to highlight the impacts of data sharing between AV operators and municipalities.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Robust Method to Identify Groundwater Affected By Redox Conditions at Los Alamos National Laboratory Legacy Cleanup Site - 20497

To facilitate the mission of the Department of Energy's (DoE's) Environmental Management Program at Los Alamos National Laboratory (LANL), a significant number of groundwater monitoring wells with depths ranging from 152-396 meter were completed for characterization and monitoring. Groundwater at Los Alamos tends to be oxygen saturated with very low concentrations of organic matter. Drilling fluids can introduce residual carbon causing an increase in microbial activity and local reducing conditions around the well, potentially impacting representativeness of groundwater data quality for redox sensitive constituents. We assert that the redox state of the wells can robustly be identified by aqueous solution concentrations of iron and manganese. An automated review process has been built into a computer-based data management system enabling an efficient screening process for identifying reducing conditions that are not caught using standard data-validation protocol, and ensures that important data quality issues are thoroughly identified and addressed in an efficient and timely manner. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Development of Methane Emissions Model to Assess Fuel Recovery Potential at Gas Well Sites Using On Site Compression

The U.S. natural gas production and consumption has increased 85.5% since 2005 primarily due to the unconventional production methods of horizontal drilling and hydraulic fracturing. Natural gas used as a fuel has a lower greenhouse gas (GHG) footprint than coal and petroleum due to lower Carbon Dioxide (CO2) emissions when combusted. However, the “greener” benefit to natural gas may be negated by leaks in production and transmission systems. Methane (CH4), the primary hydrocarbon in natural gas, has an estimated Global Warming Potential (GWP) of 28-36 over 100 years, meaning it can absorb 28-36 more energy than CO2 which has a GWP of 1.0. Natural gas well sites are prone to methane emissions, or leaks and irregular gas releases, vented to atmosphere throughout production and transmission. The U.S. Department of Energy (DOE) and the National Energy Technology Laboratory (NETL) has recently granted West Virginia University (WVU) funding under agreement DE-FOA-0002005, to “Advance technologies to mitigate methane emissions and increase the efficiency of the natural gas transportation infrastructure”. As part of this funding WVU was tasked with identifying and quantifying sources of methane emissions at unconventional well sites, processing this data, and developing a system to recapture these emissions. A 0-D Simulink model was developed, utilizing standardized methodologies, data from previously conducted studies, as well as collected data from well sites in the Marcellus shale play region. The model was developed to predict emission rates from various components at natural gas well sites as well as the potential to utilize these emissions as fuel for the natural gas powered compressor engines on-site. This model was utilized to run high, medium, and low cases for four identified emission sources, engine size, pneumatic controller count, liquid level production which dictates tank emissions, and compressor packing vent emissions. Due to discrepancies in transient tank emission data, a high and low emission factor for tanks was used, resulting in two sets of 81 executed cases, and 162 unique cases of total site emissions and potential for fuel consumption. Each of the cases were run over 86,400 seconds at a 1 Hz, representative of a full 24 hour day of operation. The fuel consumption offset an average of 557% of fuel consumption on an energy density basis across all 81 cases with the high tank emission factor with a maximum offset of 2334%. The fuel consumption offset was an average of 82.9% for all 81 cases with the low tank emission factor with a maximum offset of 337%. This study highlights flaws in the use of publicly available methane number calculations to determine natural gas’s suitability as an engine fuel as well as the lack of public data for transient liquid storage tank emissions.

03 NATURAL GAS↗

Optimization-Based Model Reduction Scheme for Renewable Energy Power Plants Using Standardized Testing Scenarios

This paper presents an optimization-based model reduction scheme for renewable energy (RE) power plants consisting of inverter-based resources (IBRs) operating in grid-following (GFL) or grid-forming (GFM) modes. More importantly, the datasets feeding the optimization-based model reduction scheme are generated and re-used through the standardized grid-interactive testing scenarios. Particularly, the proposed scheme makes use of the power plant point of common coupling (PCC) measurements of various quantities specified by standardized tests (e.g., voltage and frequency ride through) as per IEEE 2800, to estimate the parameters of the reduced-order model such that its dynamic performance aligns with the original detailed power plant model. The proposed model reduction approach does not require the parameters of individual IBRs and using standardized test data as input to the formulated optimization problem simplifies the reduced-order modelling scheme. Extensive case studies following standardized test scenarios verified the remarkable accuracy of the proposed approach.

Yallamilli, Ram S. [Purdue University]↗