Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “standardized data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Standard experimental paradigm designs and data exclusion practices in cognitive psychology can inadvertently introduce systematic “shadow” biases in participant samples

Abstract Standard cognitive psychology research practices can introduce inadvertent sampling biases that reduce the reliability and generalizability of the findings. Researchers commonly acknowledge and understand that any given study sample is not perfectly generalizable, especially when implementing typical experimental constraints (e.g., limiting recruitment to specific age ranges or to individuals with normal color vision). However, less obvious systematic sampling constraints, referred to here as “shadow” biases, can be unintentionally introduced and can easily go unnoticed. For example, many standard cognitive psychology study designs involve lengthy and tedious experiments with simple, repetitive stimuli. Such testing environments may 1) be aversive to some would-be participants (e.g., those high in certain neurodivergent symptoms) who may self-select not to enroll in such studies, or 2) contribute to participant attrition, both of which reduce the sample’s representativeness. Likewise, standard performance-based data exclusion efforts (e.g., minimum accuracy or response time) or attention checks can systematically remove data from participants from subsets of the population (e.g., those low in conscientiousness). This commentary focuses on the theoretical and practical issues behind these non-obvious and often unacknowledged “shadow” biases, offers a simple illustration with real data as a proof of concept of how applying attention checks can systematically skew latent/hidden variables in the included population, and then discusses the broader implications with suggestions for how to manage and reduce, or at a minimum acknowledge, the problem.

Siritzky, Emma M.↗

Creating the First Interactive Mobility Data Landscape

As the world continues to be driven increasingly driven by data, the ways this data is sorted and collected is critical for researchers and specification creators. In the world of mobility data, there are few if any de jure standards connecting data, and there is little knowledge on the gaps that exist in the data. In this project we have created an interactive landscape where mobility data and data specifications can be categorized and organized in an easy-to-use living document.

ADVANCED PROPULSION SYSTEMS↗

EXSCLAIM!

Due to recent improvements in image resolution and acquisition speed, materials microscopy is experiencing an explosion of published imaging data. The standard publication format, while sufficient for traditional data ingestion scenarios where a select number of images can be critically examined and curated manually, is not conducive tolarge-scale data aggregation or analysis, hindering data sharing and reuse. Most images in publications are presented as components of a larger figure with their explicit context buried in the main body or caption text, so even if aggregated, collections of images with weak or no digitized contextual labels have limited value. To solve the problem of curating labeled microscopy data from literature, we introduce the EXSCLAIM! Python toolkit for the automatic EXtraction, Separation, and Caption-based natural Language Annotation of IMages from scientific literature. The software is implemented through a three part pipeline: the JournalScraper, which searches the web and downloads figures and captions based on a user provided query, the CaptionDistributor, which separates caption text based on the subfigure each portion of the caption refers to, and the FigueSeparator, which separates figures into component subfigures and extracts other visual information. Also included is a Django user interface for exploring the resulting dataset.

CHAN, MARIA↗

DOE EV Data Collection - Vehicle Data

Vehicle data consist of electric vehicle performance data collected directly from the vehicle during standard operations. Data were collected using onboard data loggers that were either installed by the project team or preinstalled by the original equipment manufacturer. Data recorded by the data loggers were made accessible via an online web portal or an application programming interface. Different data loggers were used (HEM, ViriCiti, and Geotab), and the method for each vehicle is defined in the vehicle attributes file. Some systems collected data on a “trip-level” basis, in which each row of a table represents a single trip (the period between a key-on and key-off event), whereas other data were collected on a per-day basis, in which each row represents a single day of operation. Data were collected over a range of data collection periods, depending on the project. Data have been anonymized by removing information or decreasing information resolution as necessary so that fleets are not identifiable. Due to the wide range of vehicle types represented and variation in data collection, data parameters and frequencies differ between vehicles and fleets The **Performance Data Daily/Trip Data Dictionaries** contain definitions for each available parameter associated with a vehicle’s operations, aggregated at either a daily or trip level. The parameters available will vary from vehicle to vehicle, but every possible parameter will be defined. The **Vehicle Attributes Data Dictionary** contains definitions for each available parameter associated with a vehicle’s physical and functional attributes and fleet context. The **Vehicle Attributes** table contains specific vehicle characteristics, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. The **Vehicle Data** tables contain the data from each vehicle’s operations, aggregated at either a daily or trip level, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Physics-Informed Gaussian Process Regression for States Estimation and Forecasting in Power Grids

Real-time state estimation and forecasting are critical for the efficient operation of power grids. In this paper, a physics-informed Gaussian process regression (PhI-GPR) method is presented and used for forecasting and estimating the phase angle, angular speed, and wind mechanical power of a three-generator power grid system using sparse measurements. In standard data-driven Gaussian process regression (GPR), parameterized models for the prior statistics are fit by maximizing the marginal likelihood of observed data. In the PhI-GPR method, we propose to compute the prior statistics offline by solving stochastic differential equations (SDEs) governing the power grid dynamics. The short-term forecast of a power grid system dominated by wind generation is complicated by the stochastic nature of the wind and the resulting uncertainty in wind mechanical power. Here, we assume that the power grid dynamics are governed by swing equations, with the wind mechanical power fluctuating randomly in time. We solve these equations for the mean and covariances of the power grid states using the Monte Carlo simulation method. We demonstrate that the proposed PhI-GPR method can accurately forecast and estimate observed and unobserved states. For the considered problem, PhI-GPR has computational advantages over the ensemble Kalman filter (EnKF) method: In PhI-GPR, ensembles are computed offline and independently of the data acquisition process, whereas for EnFK, ensembles are computed online with data acquisition, rendering real-time forecast more challenging. We also demonstrate that the PhI-GPR forecast is more accurate than the EnKF forecast when the random mechanical wind power is non-Markovian. In contrast, the two methods produce similar forecasts for the Markovian mechanical wind power. For observed states, we show that PhI-GPR provides a forecast comparable to the standard data-driven GPR; both forecasts are significantly more accurate than the autoregressive integrated moving average (ARIMA) forecast. We also show that the ARIMA forecast is more sensitive to observation frequency and measurement errors than the PhI-GPR forecast.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Design and requirements of a hydrogen component reliability database (HyCReD)

Hydrogen technologies are expected to play a key role in the decarbonization of several sectors including energy storage and transportation. Rigorous investigation and quantification of the risk and reliability issues associated with hydrogen technologies will be critical to ensuring both their wider adoption and safe, economical operations. Quantitative risk assessment (QRA) is an important tool that has been used to enable the safe deployment of many engineering systems, including hydrogen fueling stations and hydrogen storage systems. However, QRA studies require reliability data which is currently lacking for expanding applications of hydrogen systems. Here, to address this gap, we present a new structure for a hydrogen component reliability database (HyCReD) that can be used to generate reliability data to be used in QRA, reliability, safety studies, maintenance planning, and more. Building on our previous work examining four major hydrogen safety data collection tools (West et al., 2022) [1], our approach in this work was to consult scientific literature on reliability data collection as well as a number of existing reliability engineering databases in the oil & gas, chemical processing, and nuclear power plant sectors. The evaluation of these databases led to identifying best practices to be implemented in a data collection framework for a hydrogen component reliability database. Based on these best practices, a set of 24 requirements for the proposed database are presented, covering its characteristics and the types of data to be collected. We define the structure of the HyCReD database and 25 data elements to be collected, spanning system description, failure, shutdown, or near-miss events, and maintenance events. The data elements are then defined according to international standards used in the safety and reliability practice and potential choice lists are provided for each field. Since this database is being piloted for hydrogen fueling stations, a generic station component hierarchy developed by West (2021) [2] is used to standardize system data. Finally, we demonstrate populating the database with information extracted from five narrative reports on hydrogen fueling station incidents.

08 HYDROGEN↗

Science Area 1: Standard Award: Model-Data Fusion to Examine Multiscale Dynamical Controls on Snow Cover and Critical Zone Moisture Inputs (Final Report)

In many mountain watersheds of the world, seasonal snowpacks play an important role as natural reservoirs of water. Seasonal snowpacks accumulate water during cold, wet winter months that subsequently melts. Downstream communities depend on water from melting seasonal snowpacks to support agricultural, industrial, and municipal water needs. Rapidly melting snowpacks can also present a flooding hazard, particularly if snowpacks melt at rates faster than anticipated and where adequate reservoir capacity is unavailable to buffer river flows associated with melt. The spatial and temporal dynamics of snow accumulation and melt also play an important role in supporting upland ecosystems in mountain landscapes. Snowmelt provides soil moisture that enable terrestrial ecosystem productivity and exert control on soil microorganisms that play important roles in global carbon cycles. Climate warming is gradually decreasing the amount of precipitation in mountain watersheds arriving as snow, presenting potentially profound disruptions to mountain ecosystems, as well as downstream delivery of water. The overarching goal of this project was to understand how interactions between the near-surface atmosphere and surface topography control the input, accumulation, retention, and release of water from mountain snowpacks. Over a 5-year period, this project pursued an approach combining high-resolution regional climate modeling, satellite and airborne remote sensing data, and ground-based observations to develop and analyze benchmark datasets to address overarching science questions and hypotheses. Key products include a continuous, long-term, high spatiotemporal resolution (1 km/1 hr) dataset characterizing key climate variables in the Upper Colorado River Basin. The dataset included historical estimates of precipitation, temperature, humidity, solar and longwave radiation, and wind speeds. Additionally, the project developed a 20+ year long, 30 m spatial, daily temporal multi-sensor dataset characterizing snow presence/absence in the East/Taylor River watersheds in the Central Rocky Mountains of Colorado. The project supported training of 1 postdoctoral scholar, 1 Ph.D. student, and 1 M.S. student.

54 ENVIRONMENTAL SCIENCES↗

EDX Spatial: Leveraging cloud and hybrid data management resources for spatial data

In the last few years, the National Energy Technology Laboratory has started leveraging cloud-hosted services for hosting spatial data collections published on the Energy Data eXchange (EDX). Using cloud-hosted storage and compute options offers multiple benefits for visualization and tool development through utilization of spatial data resources. The transition from on-premises services to cloud-hosted services has enabled a few key features: increased accessibility of large derivative datasets, dynamic integration of external authoritative data resources directly from outside entities through representational state transfer application programming interfaces (REST API), and the ability to produce complex mapping applications, dashboards, and online maps. As a result of the shift, NETL has launched EDX Spatial, a platform that leverages on-premises resources combined with cloud compute capabilities to enable enhanced online mapping interfaces and optimize data access. In addition, a unified workflow for handling the public release of spatial data products through EDX has been developed, including standardization of data hosting practices, metadata, symbology, and application elements. This poster reviews the opportunity of leveraging cloud-hosted and hybrid data management solutions for spatial data, and discusses the benefits and lessons learned while leveraging these services through the data repository EDX.

Morkner, Paige↗

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

The Geothermal Data Repository: Ten Years of Supporting the Geothermal Industry with Open Access to Geothermal Data: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) is celebrating its tenth anniversary! Over the last decade it has grown from the simple idea of storing public data in a centralized location to a valuable tool at the center of the US geothermal scientific community and an integral part of the DOE Geothermal Technologies Office (DOE GTO) project management strategy. Researchers funded by the DOE GTO have contributed over 1,300 data submissions to the GDR. These data have been used to further advancements in geothermal science, economic analysis, exploration, research, development, and operational efficiency. The adoption of open data methodologies and a data management strategy that prioritizes universal open access and standardized, interoperable data have further increased the value of GDR data, making them available across a distributed network of data sharing partners and improving their utility to other industries and related fields, including material science and space exploration. Incorporating feedback from users has been critical to the GDRs success, allowing it to grow over the years to meet the evolving needs of the geothermal community. This paper will explore some of many changes that occurred throughout the GDRs tenure and the lessons learned along the way, as well as highlight some of the new features and recent improvements that been implemented to support innovation, reduce duplication of effort, and advance the geothermal industry as a whole.

accessibility↗

Commercial, industrial, and institutional discount rate estimation for efficiency standards analysis: Sector-level data 1998–2021

Underlying each of the Department of Energy’s (DOE’s) federal appliance and equipment energy conservation standards are a set of complex analyses of the projected costs and benefits of regulation. Any new or amended standard must be designed to achieve significant additional energy conservation, provided that it is technologically feasible and economically justified (42 U.S.C. 6295(o)(2)(A)). DOE determines economic justification based on whether the benefits exceed the burdens, considering a variety of factors, including the economic impact of the standard on consumers of the product and the savings in lifetime operating cost compared to any increase in price or maintenance expenses (42 U.S.C. 6295(o)(2)(B)). As part of this determination, DOE conducts a Life-Cycle Cost (LCC) analysis, which models the combined impact of appliance first cost and operating cost changes on a representative commercial building sample in order to identify the fraction of customers achieving LCC savings or incurring net cost at the considered efficiency levels. Thus, the commercial discount rate value(s) used to calculate the present value of energy cost savings within the LCC model implicitly plays a role in estimating the economic impact of potential standard levels. This report provides an in-depth discussion of the commercial discount rate estimation process. It is an update to previous reports on estimating commercial discount rates from firm-level financial data (Fujita, 2016). Major topics covered in this report include: Discount rate estimation methods and rationale; -Data sources used and data limitations; -Discount rate distributions for use in standards analysis; -Discount rate estimation methods and distributions specific to the small business subgroup analysis. Going forward, this report will be updated as data allow and analyses necessitate.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The Geothermal Data Repository: Ten Years of Supporting the Geothermal Industry with Open Access to Geothermal Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) is celebrating its tenth anniversary! Over the last decade it has grown from the simple idea of storing public data in a centralized location to a valuable tool at the center of the US geothermal scientific community and an integral part of the DOE Geothermal Technologies Office (DOE GTO) project management strategy. Researchers funded by the DOE GTO have contributed over 1,300 data submissions to the GDR. These data have been used to further advancements in geothermal science, economic analysis, exploration, research, development, and operational efficiency. The adoption of open data methodologies and a data management strategy that prioritizes universal open access and standardized, interoperable data have further increased the value of GDR data, making them available across a distributed network of data sharing partners and improving their utility to other industries and related fields, including material science and space exploration. Incorporating feedback from users has been critical to the GDR's success, allowing it to grow over the years to meet the evolving needs of the geothermal community. This paper will explore some of many changes that occurred throughout the GDRs tenure and the lessons learned along the way, as well as highlight some of the new features and recent improvements that been implemented to support innovation, reduce duplication of effort, and advance the geothermal industry as a whole.

access↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

Commercial, industrial, and institutional discount rate estimation for efficiency standards analysis: Sector-level data 1998–2022

Underlying each of the U.S. Department of Energy’s (DOE’s) federal appliance and equipment energy conservation standards are a set of complex analyses of the projected costs and benefits of regulation. Any new or amended standard must be designed to achieve significant additional energy conservation, provided that it is technologically feasible and economically justified (42 U.S.C. 6295(o)(2)(A)). DOE determines economic justification based on whether the benefits exceed the burdens, considering a variety of factors, including the economic impact of the standard on consumers of the product and the savings in lifetime operating cost compared to any increase in price or maintenance expenses (42 U.S.C. 6295(o)(2)(B)). As part of this determination, DOE conducts a life-cycle cost (LCC) analysis, which models the combined impact of appliance first cost and operating cost changes on a representative commercial building sample to identify the fraction of customers achieving LCC savings or incurring net cost at the considered efficiency levels. Thus, the commercial discount rate value(s) used to calculate the present value of energy cost savings within the LCC model implicitly plays a role in estimating the economic impact of potential standard levels. This report provides an in-depth discussion of the commercial discount rate estimation process. It is an update to previous reports on estimating commercial discount rates from firm-level and sector-level financial data (e.g., Fujita, 2021, 2016). Major topics covered in this report include the following: -Discount rate estimation methods and rationale -Data sources used and data limitations -Discount rate distributions for use in standards analysis -Discount rate estimation methods and distributions specific to the small business subgroup analysis A version of this analysis was most recently released in 2022. Going forward, this report will be updated as data allow and analyses necessitate.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

The Coastal Carbon Library and Atlas: Open source soil data and tools supporting blue carbon research and policy

Abstract Quantifying carbon fluxes into and out of coastal soils is critical to meeting greenhouse gas reduction and coastal resiliency goals. Numerous ‘blue carbon’ studies have generated, or benefitted from, synthetic datasets. However, the community those efforts inspired does not have a centralized, standardized database of disaggregated data used to estimate carbon stocks and fluxes. In this paper, we describe a data structure designed to standardize data reporting, maximize reuse, and maintain a chain of credit from synthesis to original source. We introduce version 1.0.0. of the Coastal Carbon Library, a global database of 6723 soil profiles representing blue carbon‐storing systems including marshes, mangroves, tidal freshwater forests, and seagrasses. We also present the Coastal Carbon Atlas, an R‐shiny application that can be used to visualize, query, and download portions of the Coastal Carbon Library. The majority (4815) of entries in the database can be used for carbon stock assessments without the need for interpolating missing soil variables, 533 are available for estimating carbon burial rate, and 326 are useful for fitting dynamic soil formation models. Organic matter density significantly varied by habitat with tidal freshwater forests having the highest density, and seagrasses having the lowest. Future work could involve expansion of the synthesis to include more deep stock assessments, increasing the representation of data outside of the U.S., and increasing the amount of data available for mangroves and seagrasses, especially carbon burial rate data. We present proposed best practices for blue carbon data including an emphasis on disaggregation, data publication, dataset documentation, and use of standardized vocabulary and templates whenever appropriate. To conclude, the Coastal Carbon Library and Atlas serve as a general example of a grassroots F.A.I.R. (Findable, Accessible, Interoperable, and Reusable) data effort demonstrating how data producers can coordinate to develop tools relevant to policy and decision‐making.

Holmquist, James R.↗