Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data integrity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

CORAL: A framework for rigorous self-validated data modeling and integrative, reproducible data analysis

Abstract Background Many organizations face challenges in managing and analyzing data, especially when relevant datasets arise from multiple sources and methods. Analyzing heterogeneous datasets and additional derived data requires rigorous tracking of their interrelationships and provenance. This task has long been a Grand Challenge of data science and has more recently been formalized in the FAIR principles: that all data objects be Findable, Accessible, Interoperable, and Reusable, both for machines and for people. Adherence to these principles is necessary for proper stewardship of information, for testing regulatory compliance, for measuring the efficiency of processes, and for facilitating reuse of data-analytical frameworks. Findings We present the Contextual Ontology-based Repository Analysis Library (CORAL), a platform that greatly facilitates adherence to all 4 of the FAIR principles, including the especially difficult challenge of making heterogeneous datasets Interoperable and Reusable across all parts of a large, long-lasting organization. To achieve this, CORAL's data model requires that data generators extensively document the context for all data, and our tools maintain that context throughout the entire analysis pipeline. CORAL also features a web interface for data generators to upload and explore data, as well as a Jupyter notebook interface for data analysts, both backed by a common API. Conclusions CORAL enables organizations to build FAIR data types on the fly as they are needed, avoiding the expense of bespoke data modeling. CORAL provides a uniquely powerful platform to enable integrative cross-dataset analyses, generating deeper insights than are possible using traditional analysis tools.

97 MATHEMATICS AND COMPUTING↗

Demonstration and Evaluation of a Non-Invasive, Low-Cost, Strap-On Sensor for Natural Gas Meters

The U.S. General Services Administration (GSA) is interested in installing internet-connected, gas submeters to better understand gas consumption in its portfolio of buildings. The GSA in partnership with the National Renewable Energy Laboratory conducted a demonstration to assess a specific submeter technology. This technology was implemented at two GSA separate facilities located in Dallas, Texas. This demonstration evaluated hardware and software installations and integrations, data integrity and accuracy, and included an economic analysis. This demonstration evaluated a gas submeter technology provided by the vendor Vata Verks, who produces a non-invasive strap-on submeter. The company was founded with a mission to conserve water (and eventually gas) cheaply and simply, as explained on its website. The product intends to streamline submeter deployments for gas and water by eliminating most hardware costs while allowing for easy integration of submetered data into other systems such as the Building Automation System and removing tenant and building disruption This demonstration evaluated the product features when used on a gas utility meter.

03 NATURAL GAS↗

CO 2 Storage Site Screening Platform Development and CO 2 Storage Resource Analysis in SECARB Offshore Reservoirs Using SAS Viya

A major goal of the SECARB Offshore Partnership (DE-FE0031557) is to screen deep saline aquifers and hydrocarbon reservoirs in the central Gulf of Mexico for CO 2 sequestration and CO 2 -enhanced oil and gas recovery (EOR/EGR) and estimate the corresponding CO 2 storage resources for select reservoirs. CO 2 storage potential associated with offshore CO 2 -EOR is considerable and likely represents “low hanging fruit” for near-term CO 2 storage given the in-place infrastructure in the region. It is for these reasons that this assessment focuses on oil and gas fields. To this end, three major objectives have been completed and include (1) managing geological data derived from different sources, (2) building a reservoir screening platform for CO 2 storage, and (3) ranking the reservoirs based on the estimated CO 2 storage resources. The SAS ® Viya platform was used for data management and analytics. The Viya platform is a cloud service platform that provides data integration, data management, quick analytics, data visualization, machine learning functions, and application programming interfaces (APIs) for multi-programming languages. Different sources of data containing geologic information, reservoir properties, and EOR/EGR information were collected, cleaned, formatted, and loaded into the SAS ® Viya platform for evaluation. The major geological characteristics of both shelf and deep-water areas of the central Gulf were examined and compared to define the appropriate reservoir screening criteria. Next, a CO 2 storage site screening system was built in the SAS ® Viya platform with the pre-defined criteria. Finally, the CO 2 storage resources of the screened reservoirs were calculated and reported at the BOEM field level to identify fields with the highest estimated CO 2 storage resource. The fields with the largest total estimated CO 2 storage resource are located in the Mississippi Canyon protraction area. Due to proximity to the Mississippi Delta (indicative of less infrastructure) and large estimated CO 2 storage resources, future development activities may wish to focus efforts in the Mississippi Canyon protraction area.

02 PETROLEUM↗

Integrating multimodal data through interpretable heterogeneous ensembles

Motivation: Integrating multimodal data represents an effective approach to predicting biomedical characteristics, such as protein functions and disease outcomes. However, existing data integration approaches do not sufficiently address the heterogeneous semantics of multimodal data. In particular, early and intermediate approaches that rely on a uniform integrated representation reinforce the consensus among the modalities but may lose exclusive local information. The alternative late integration approach that can address this challenge has not been systematically studied for biomedical problems. Results: We propose Ensemble Integration (EI) as a novel systematic implementation of the late integration approach. EI infers local predictive models from the individual data modalities using appropriate algorithms and uses heterogeneous ensemble algorithms to integrate these local models into a global predictive model. We also propose a novel interpretation method for EI models. We tested EI on the problems of predicting protein function from multimodal STRING data and mortality due to coronavirus disease 2019 (COVID-19) from multimodal data in electronic health records. We found that EI accomplished its goal of producing significantly more accurate predictions than each individual modality. It also performed better than several established early integration methods for each of these problems. The interpretation of a representative EI model for COVID-19 mortality prediction identified several disease-relevant features, such as laboratory test (blood urea nitrogen and calcium) and vital sign measurements (minimum oxygen saturation) and demographics (age). These results demonstrated the effectiveness of the EI framework for biomedical data integration and predictive modeling.

59 BASIC BIOLOGICAL SCIENCES↗

Merged Observatory Data Files (MODFs): an integrated observational data product supporting process-oriented investigations and diagnostics

A large and ever-growing body of geophysical information is measured in campaigns and at specialized observatories as a part of scientific expeditions and experiments. These collections of observed data include many essential climate variables (as defined by the Global Climate Observing System) but are often distinguished by a wide range of additional non-routine measurements that are designed to not only document the state of the environment but also the drivers that contribute to that state. These field data are used not only to further understand environmental processes through observation-based studies but also to provide baseline data to test model performance and to codify understanding to improve predictive capabilities. To address the considerable barriers and difficulty in utilizing these diverse and complex data for observation–model research, the Merged Observatory Data File (MODF) concept has been developed. A MODF combines measurements from multiple instruments into a single file that complies with well-established data format and metadata practices and has been designed to parallel the development of corresponding Merged Model Data Files (MMDFs). Using the MODF and MMDF protocols will facilitate the evolution of model intercomparison projects into model intercomparison and improvement projects by putting observation and model data “on the same page” in a timely manner. The MODF concept was developed especially for weather forecast model studies in the Arctic. The surprisingly complex process of implementing MODFs in that context refined the concept itself. Thus, this article explains the concept of MODFs by providing details on the issues that were revealed and resolved during that first specific implementation. Detailed instructions are provided on how to make MODFs, and this article can be considered a MODF creation manual.

54 ENVIRONMENTAL SCIENCES↗

ARM Aerial Facility (AAF) Merged Value-Added Product Report for Historical G-1 Field Campaigns

For 30 years, the U.S. Department of Energy (DOE) Office of Science supported an instrumented Grumman Gulfstream-1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) user facility Data Center (ADC) and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated data set was recently developed covering the final six years of G-1 operations (2013 to 2018). The integrated data set includes data collected from 236 flights (766.4 hours). Four of the seven field campaigns were based in the U.S. One campaign collected data from the wildfires in the U.S. Pacific Northwest and agricultural burns in the lower Mississippi River valley as part of the Biomass Burning Observation Project (BBOP) in 2013. In 2015, the ARM Cloud Aerosol Precipitation Experiment provided data on atmospheric rivers and associated aerosol-cloud interactions that produce heavy precipitation on the U.S. west coast during the early spring. Research data from Airborne Carbon Measurements-V (ACME-V), collected during the summer of 2015, gave scientists insight into trends and variability of trace gases in the atmosphere over the North Slope of Alaska to improve arctic climate models. In the early summer and autumn of 2016, the Holistic Interactions of Shallow Clouds, Aerosols, and Land-Ecosystems (HI-SCALE) campaign provided an extensive data set geared toward coupled processes that affect the life cycle of shallow clouds through the interaction among aerosol, cloud, land surface, and ecosystems. In 2014 (March and October), the airborne sampling moved outside of the U.S. to the city of Manaus in central Amazonia, Brazil, where residential and industrial emissions were extensively characterized by flights of the G-1. The GoAmazon2014/15 aircraft campaign data are being integrated with aquatic and terrestrial ecosystem measurements to quantify anthropogenic perturbations to a usually pristine tropical environment. Another international airborne mission was carried out in the Eastern North Atlantic region. The Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA) campaign saw the G-1 aircraft fly from Terceira Island in the Azores during the summer of 2017 and the winter of 2018. The campaign studied both seasons to measure key aerosol and cloud processes under various meteorological and cloud conditions with different aerosol sources. Then the G-1 deployed to the Sierras de Córdoba range in central Argentina from October to November 2018 for the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) campaign to study orographic convective cloud interactions with their surrounding environment. These comprehensive datastreams provide much-needed insight into spatiotemporal variability of thermodynamic quantities, aerosol and cloud states, and properties for addressing essential science questions in Earth system process studies.

54 ENVIRONMENTAL SCIENCES↗

ARM Aerial Facility (AAF) Merged Value-Added Product Report for Historical G-1 Field Campaigns

For 30 years, the U.S. Department of Energy (DOE) Office of Science supported an instrumented Grumman Gulfstream-1 (G-1) aircraft for atmospheric field campaigns. Data from the final decade of G-1 operations were archived by the Atmospheric Radiation Measurement (ARM) user facility Data Center (ADC) and made publicly available at no cost to all registered users. To ensure a consistent data format and to improve the accessibility of the ARM airborne data, an integrated data set was recently developed covering the final six years of G-1 operations (2013 to 2018). The integrated data set includes data collected from 236 flights (766.4 hours). Four of the seven field campaigns were based in the U.S. One campaign collected data from the wildfires in the U.S. Pacific Northwest and agricultural burns in the lower Mississippi River valley as part of the Biomass Burning Observation Project (BBOP) in 2013. In 2015, the ARM Cloud Aerosol Precipitation Experiment provided data on atmospheric rivers and associated aerosol-cloud interactions that produce heavy precipitation on the U.S. west coast during the early spring. Research data from Airborne Carbon Measurements-V (ACME-V), collected during the summer of 2015, gave scientists insight into trends and variability of trace gases in the atmosphere over the North Slope of Alaska to improve arctic climate models. In the early summer and autumn of 2016, the Holistic Interactions of Shallow Clouds, Aerosols, and Land-Ecosystems (HI-SCALE) campaign provided an extensive data set geared toward coupled processes that affect the life cycle of shallow clouds through the interaction among aerosol, cloud, land surface, and ecosystems. In 2014 (March and October), the airborne sampling moved outside of the U.S. to the city of Manaus in central Amazonia, Brazil, where residential and industrial emissions were extensively characterized by flights of the G-1. The GoAmazon2014/15 aircraft campaign data are being integrated with aquatic and terrestrial ecosystem measurements to quantify anthropogenic perturbations to a usually pristine tropical environment. Another international airborne mission was carried out in the Eastern North Atlantic region. The Aerosol and Cloud Experiments in the Eastern North Atlantic (ACE-ENA) campaign saw the G-1 aircraft fly from Terceira Island in the Azores during the summer of 2017 and the winter of 2018. The campaign studied both seasons to measure key aerosol and cloud processes under various meteorological and cloud conditions with different aerosol sources. Then the G-1 deployed to the Sierras de Córdoba range in central Argentina from October to November 2018 for the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) campaign to study orographic convective cloud interactions with their surrounding environment. These comprehensive datastreams provide much-needed insight into spatiotemporal variability of thermodynamic quantities, aerosol and cloud states, and properties for addressing essential science questions in Earth system process studies.

54 ENVIRONMENTAL SCIENCES↗

WIND Toolkit Raw Data Reformatted

Wind Integration National Dataset Toolkit The Wind Integration National Dataset (WIND) Toolkit is an update and expansion of the Eastern Wind Integration Data Set and Western Wind Integration Data Set. It supports the next generation of wind integration studies.

17 WIND ENERGY↗

Integral Nuclear Data and Benchmarking Needs for Fusion Energy Systems

Fusion energy systems are currently being designed and optimized using radiation transport codes. To deal with the unique environment inside a fusion-based system, many of these designs incorporate novel materials able to withstand the high radiation fields, ensure adequate cooling and thermal protection, and produce tritium. Validation plays a vital role in building trust in the predictive power of these models and computational methods. Validation of a code consists of modeling documented real-world experiments and comparing the code-predicted response to the measured response. Adequate validation requires measured responses from real-world experiments, also known as integral data, that mimic the system being designed, including materials, impinging radiation, and temperature, among other variables. The most trusted integral data are experimental responses that have been through a rigorous benchmarking process that develops a recommended computational model and evaluates all experimental uncertainties. Finally, there are a few research groups around the world that have been producing integral data for fusion applications, but a substantial investment is needed to address the unique validation needs of the fusion community.

Fusion↗

Enhancing Biopreparedness through a Model System to Understand the Molecular Mechanisms that Lead to Pathogenesis and Disease Transmission: NW-BRaVE

The science of biopreparedness to counter biological threats hinges on understanding the fundamental principles and molecular mechanisms that lead to pathogenesis and disease transmission. Our vision to address this challenge is to create a powerful and user-friendly platform to elucidate the fundamental principles of how molecular interactions drive pathogen-host relationships and host shifts. We will enable groundbreaking discoveries by integrating a wide range of structural, genomics, proteomics, and other advanced omics measurements, along with evolutionary and artificial intelligence predictions. To make sure the system is applicable to real-world problems, we will develop it in the context of a tractable model system, the small, abundant, and accessible photosynthetic cyanobacteria and their constantly co-adapting viral pathogens, cyanophages. This model will maintain the system’s applicability to real-world problems and techniques, but the overall focus will be on elucidating general principles of detecting, assessing, and surveilling molecular interaction, adaptation, and coevolution that are system agnostic and therefore extensible to other viral-host interactions. Our overall objectives are to (1) identify the molecular complexes that comprise the cyanobacteria redox macromolecular subsystem and how they dynamically change with bacteriophage infection in situ, using cryo-electron tomography; (2) profile regulatory changes during infection using proteomics, multiomics, and experimental validation, and integrate the data with in situ structures; (3) use genomics and metagenomics to determine environmental and population factors across time scales that impact the interactions between marine cyanobacteria and their cyanophage parasites, predicting the evolutionary origins of in situ structural and functional interactions, convergence and coevolution; and (4) develop a data integration and transformation platform that facilitates the integration of in situ, proteomic, and evolutionary measurements of molecular interactions to surveil diverse hosts and parasites in various environmental contexts. These objectives address Focus Area 2 Reveal Molecular Interactions Across Biological Scales for Design of Targeted Interventions. Our powerful and user-friendly platform will enhance connections between the often-siloed fields of structure, molecular phenotype, and evolutionary genomics that are key to biopreparedness, but in need of integration (Figure 1). We will build an integrated navigation tool to facilitate the effective use of globally distributed experimental data for integrated analysis and predictive modeling. The project will develop, implement, and test a platform to assess host-pathogen molecular interactions, adaptation to hosts and host shifts, and coevolution between hosts and pathogens, successfully impacting the research community by revolutionizing abilities to study any host-pathogen interaction, encourage diverse community contributions, and gain fundamental insights into how proteins adapt to new contexts. This ability will be critical for designing early interventions to address future threats. We will build surveillance training capability, aiming for a fair and equitable response to future pandemics and biothreats.

59 BASIC BIOLOGICAL SCIENCES↗

Guided construction of single cell reference for human and mouse lung

Accurate cell type identification is a key and rate-limiting step in single-cell data analysis. Single-cell references with comprehensive cell types, reproducible and functionally validated cell identities, and common nomenclatures are much needed by the research community for automated cell type annotation, data integration, and data sharing. Here, we develop a computational pipeline utilizing the LungMAP CellCards as a dictionary to consolidate single-cell transcriptomic datasets of 104 human lungs and 17 mouse lung samples to construct LungMAP single-cell reference (CellRef) for both normal human and mouse lungs. CellRefs define 48 human and 40 mouse lung cell types catalogued from diverse anatomic locations and developmental time points. We demonstrate the accuracy and stability of LungMAP CellRefs and their utility for automated cell type annotation of both normal and diseased lungs using multiple independent methods and testing data. We develop user-friendly web interfaces for easy access and maximal utilization of the LungMAP CellRefs.

59 BASIC BIOLOGICAL SCIENCES↗

IS BLOCKCHAIN A SUITABLE TECHNOLOGY FOR ENSURING THE INTEGRITY OF DATA SHARED BY LIGHTING AND OTHER BUILDING SYSTEMS?

Increasing amounts of data are available from lighting and other building systems. In commercial buildings, these data can be used to improve system and energy performance, detect and diagnose faults, and facilitate maintenance. Building data are not only of interest to owners and operators of the building systems, however. Owners and operators of similar buildings, manufacturers of building systems, utilities, and city agencies also have interesting use cases. Energy data can, for example, be used to verify the performance of energy-conservation measures, issue renewableenergy certificates, and financially settle grid services. Increased data sharing, however, significantly expands the cyber-attack surface, creating new challenges. In this work, blockchain is explored as an option for ensuring the integrity of data that are shared by lighting and other building systems. Blockchain fundamentals and variants are briefly reviewed, and value propositions relevant to building systems are discussed. A recently developed blockchain applicability framework (BAF) that builds upon and addresses the limitations of previous applicability models is also briefly reviewed. The BAF is used to assess the suitability of blockchain over other technologies or approaches for building-data applications, using emerging connected lighting systems as an example use case.

blockchain, connected lighting system, data integr↗

Rate Equations for Reversible Disproportionation Reactions and Fitting to Time-Course Data

Integrated rate equations are straightforward to fit to experimental data to verify a proposed mechanism and to extract kinetic parameters. Such equations are derived for reversible disproportionation/comproportionation reactions with any set of initial concentrations. Extraction of forward and reverse rate constants from experimental data by fitting the rate law to the data is demonstrated for the disproportionation of 2,2,6,6-tetramethyl-1-piperidinyl-N-oxyl (TEMPO) under acidic conditions where the approach to equilibrium is observed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ensemble Federated Machine Learning‐Based Cybersecurity Situational Awareness in Microgrid Network

Cyber-physical microgrids are vulnerable to stealthy cybersecurity threats that disguise their actions through the exploitation of system knowledge. Such actions can severely impacts microgrids deployed in defense bases, slowing the response time of military forces during national emergencies. Several machine-learning algorithms have been proposed to detect intrusions in the grid networks; however, these traditional machine-learning algorithms lack data privacy and are subject to several adversarial machine-learning threats. This paper proposes a novel federated machine learning (FML)-based three-model framework to detect and identify stealthy data-integrity attacks while ensuring data privacy in microgrid networks. The proposed architecture uses a variational mode decomposition technique to extract derived features from incoming measurement and control datasets. The extraction of these derived features allows FML models to learn minute variations in data patterns that allow them to perform significantly better than the models trained with generic datasets consisting of raw features. Our experimental results show the efficient performance of the proposed methodology against different types of data integrity attacks while considering primary and secondary controllers in microgrids. Further, the applied FML-integrated random forest ensemble algorithm outperforms the existing generic FML algorithms during noisy and noise-free datasets with prediction latencies of only 91–134 µs per sample within the 0.1 s sampling interval and requires communication bandwidth of around ∼8.25 KB/s at the control center and ∼2.7 KB/s per edge client for communication.

24 POWER TRANSMISSION AND DISTRIBUTION↗