Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Digital Information Platform (DIP) Request for Information (RFI) Informational Session

DIP published an RFI on March 23, 2021. An information session is being hosted to go over the concept of DIP, review the content of the RFI and clarify what is being asked. Important information is emerging in digital form and additional effort is required to utilize the information. DIP will provide the capability to integrate key data sources to data processing services that are used to build aviation services for the traditional/commercial community and new emergents. DIP is currently planning and will take a community-involved, collaborative approach to define and develop the framework. DIP is seeking community and stakeholder input to operational needs, metrics of interest, and demonstration contribution. Participants will have an opportunity to ask questions at the end.

Mirna Johnson↗

RadLab: A Comprehensive Database and Analysis Toolkit for Space Radiation Measurements Relevant to Space Radiation Biology

RadLab, a new addition to the NASA Open Science Data Repository (OSDR), is a public platform for space radiation data relevant to human space exploration. RadLab consists of a database, a submission portal, and user-friendly visualization and data analysis tools, including a graphical user interface (GUI) and an application programming interface (API). Investigators from ISS partners including Germany, Italy, Canada, Hungary, the Czech Republic, Russia, Japan have committed to providing data from their instruments. RadLab will also include data from other spacecraft in LEO: the Space Shuttle, the Mir space station, biosatellites; and beyond LEO: the lunar and the Martian surface, the heliocentric orbit at 1 AU, Mars orbit, and Earth-Mars space. Once fully operational, RadLab will provide open, centralized access to space radiation physics data relevant to human space exploration; a platform for submission of data by agencies and research institutions responsible for radiation detectors deployed in space; analysis tools to facilitate detector and dataset intercomparison to better understand space habitat radiation environments; capabilities for space biology investigators to determine the radiation environment to which samples were exposed. A RadLab Working Group (RLWG) has been formed, modeled on the GeneLab Analysis Working Groups and comprised of data contributors and users. RLWG tasks include identifying data sources, normalizing data from diverse detectors, expanding the analysis toolkit and, perhaps most importantly, sharing ideas for research exploiting capabilities of RadLab. We will provide an overview of RadLab data and capabilities and discuss examples of its potential as a resource for open science.

radiation↗

NEWS Integrated Analysis (NEWS-IA) Dataset: Budgets and Input Data

This documentation provides a summary of the underlying physical and mathematical descriptions of the global energy and water cycle budgets used to produce the NASA Energy and Water Cycle Study (NEWS) Integrated Analysis. In addition to discussion of the budget equations, a summary is provided concerning NEWS regions, input data sources, and data fields found in the NEWS-IA input data file.

J Brent Roberts↗

Data shuffling with hierarchical tuple spaces

Methods and systems for shuffling data to generate a dataset are described. A first map module may generate first pair data, and a second map module may generate second pair data, from source data. The first map module may insert the first pair data into a first local tuple space accessible to the first map module. The second map module may insert the second pair data into a second local tuple space accessible to the second map module. A shuffle module may request pair data that includes a particular key. The first and second pair data may be inserted into a global tuple space accessible by the first and second map modules. The shuffle module may identify the requested pair data in the global tuple space, and may fetch the identified pair data from a memory. The shuffle module may shuffle the fetched pair data to generate the dataset.

Kayi, Abdullah↗

Data shuffling with hierarchical tuple spaces

Methods and systems for shuffling data are described. A processor may generate pair data from source data. The processor may insert the pair data into local tuple spaces. In response to a request for a particular key, the processor may determine a presence of the requested key in a global tuple space. The processor may, in response to a presence of the requested key in the global tuple space, update the global tuple space. The update may be based on the pair data among the local tuple spaces including the existing key. The processor may, in response to an absence of the requested key in the global tuple space, insert pair data including the missing key from the local tuple spaces into the global tuple space. The processor may fetch the requested pair data, and may shuffle the fetched data to generate a dataset.

Andrade Costa, Carlos Henrique↗

GRIP: Constraint-based Explanation of Missing Answers for Graph Queries

Abstract: A useful feature in graph query engines is to clarify “Why certain entities (nodes, attribute values or edges) are missing” in query answers. This task is even more challenging when the relevant data is already missing in the underlying data source. Missing data, on the other hand, can be inferred by enforcing data constraints for graphs. We demonstrate GRIP, a system that exploits data constraints to clarify missing answers for graph queries. (1) Constraint-based ex- planation. Given a desired yet missing entity in the query answer, GRIP ensures to generate finite and minimal sequences of data con- strains (an “explanation”) that should be consecutively enforced to ?? to ensure its occurrence for the same query. (2) Answering “why” and“how” questions. Users can query GRIP with both“Why”(“Why” the element is missing) and “How” questions (“How” to refine the graph to include the missing answer). GRIP engine supports run- time generation of explanations by incrementally maintaining a set of bi-directional search trees. (3) Interactive exploration. GRIP provides a user-friendly GUI to support interactive ad visual exploration of explanations, including both automated generation and step-by-step inspection of graph manipulations.

graphs↗

One-way transfer device with secure reverse channel

A data diode provides a flexible device for collecting data from a data source and transmitting the data to a data destination using one-way data transmission across a main channel. On-board processing elements allow the data diode to identify automatically the type of connectivity provided to the data diode and configure the data diode to handle the identified type of connectivity. Either or both of the inbound and outbound side of the data diode may comprise one or both of wired and wireless communication interfaces. A secure reverse channel, separate from the main channel, allows carefully predetermined communications from the data destination to the data source.

Lee, Sang Cheon↗

One-way transfer device with secure reverse channel

A data diode provides a flexible device for collecting data from a data source and transmitting the data to a data destination using one-way data transmission across a main channel. On-board processing elements allow the data diode to identify automatically the type of connectivity provided to the data diode and configure the data diode to handle the identified type of connectivity. Either or both of the inbound and outbound side of the data diode may comprise one or both of wired and wireless communication interfaces. A secure reverse channel, separate from the main channel, allows carefully predetermined communications from the data destination to the data source.

Lee, Sang Cheon↗

Kosh

Kosh allows codes to store, query, share data via an easy-to-use Python API. Kosh lies on top of Sina and as a result can use any database backend supported by Sina. In adition Kosh aims to make data access and sharing as simple as possible. Via "loaders" kosh can open files associated with datasets in a seemless fashion independently of the actual file format. Kosh's loader can also load data in different format, although numpy is the most usual output type. Once loaded data from sources, data can be further processed via "transformers"

Doutriaux, Charles↗

Towards a distributed information architecture for avionics data

Avionics data at the National Aeronautics and Space Administration's (NASA) Jet Propulsion Laboratory (JPL consists of distributed, unmanaged, and heterogeneous information that is hard for flight system design engineers to find and use on new NASA/JPL missions. The development of a systematic approach for capturing, accessing and sharing avionics data critical to the support of NASA/JPL missions and projects is required. We propose a general information architecture for managing the existing distributed avionics data sources and a method for querying and retrieving avionics data using the Object Oriented Data Technology (OODT) framework. OODT uses XML messaging infrastructure that profiles data products and their locations using the ISO-11179 data model for describing data products. Queries against a common data dictionary (which implements the ISO model) are translated to domain dependent source data models, and distributed data products are returned asynchronously through the OODT middleware. Further work will include the ability to 'plug and play' new manufacturer data sources, which are distributed at avionics component manufacturer locations throughout the United States.

Information architecture↗

Sub-City-Scale Air Quality Forecasts Combining Models, Satellites, and Surface Measures

Poor air quality is a global major concern, especially in cities with their higher emissions and large numbers of exposed people. Air quality monitoring has traditionally relied on ground-based measurements from a few accurate but expensive regulatory-grade monitors, leading to limited spatial data coverage. More recently, these have been supplemented with other data sources, including satellite observations of pollutants, atmospheric chemistry model simulations, and low-cost monitors allowing for denser spatial data collection at the expense of lower accuracy compared to regulatory-grade monitors. Each of these air quality data sources have their own benefits and drawbacks, and so there is an opportunity to combine these data together while respecting their relative strengths and weaknesses in order to generate a more comprehensive and detailed picture of local air quality. I will present our current work towards such a combination, with a focus on producing higher spatial resolution estimates and near-term forecasts of Nitrogen Dioxide in urban areas in the United States. Our method combines data from the GEOS Composition Forecasting (GEOS-CF) model system, TROPOMI tropospheric NO2 satellite data products, and ground measurements from the EPA regulatory monitoring network using a combination of simple intuitive relationships and machine learning techniques. I will show the performance of this proposed method in several urban areas in the United States, comparing it with baseline approaches using single data sources separately. Overall, we find that combining these disparate datasets together leads to more accurate air quality forecasts in the target areas than is possible using each data source separately.

Air quality↗

Application of remote sensing to estimating soil erosion potential

A variety of remote sensing data sources and interpretation techniques has been tested in a 6136 hectare watershed with agricultural, forest and urban land cover to determine the relative utility of alternative aerial photographic data sources for gathering the desired land use/land cover data. The principal photographic data sources are high altitude 9 x 9 inch color infrared photos at 1:120,000 and 1:60,000 and multi-date medium altitude color and color infrared photos at 1:60,000. Principal data for estimating soil erosion potential include precipitation, soil, slope, crop, crop practice, and land use/land cover data derived from topographic maps, soil maps, and remote sensing. A computer-based geographic information system organized on a one-hectare grid cell basis is used to store and quantify the information collected using different data sources and interpretation techniques. Research results are compared with traditional Universal Soil Loss Equation field survey methods.

Morris-Jones, D. R.↗

HarDWR - Raw Water Rights Records

A dataset within the Harmonized Database of Western U.S. Water Rights (HarDWR). For a detailed description of the database, please see the meta-record v2.0. Changelog v2.0 - Switched source data from collecting records from each state independently to using the WestDAAT dataset v1.0 - Initial public release Description In order to hold a water right in the western United States, an entity, (e.g., an individual, corporation, municipality, sovereign government, or non-profit) must register a physical document with the state's water regulatory agency. State water agencies each maintain their own database containing all registered water right documents within the state, along with relevant metadata such as the point of diversion and place of use of the water. All western U.S. states have digitized their individual water rights databases, as well as geospatial data defining the areas in which water rights are managed. Each state maintains and provides their own water rights data in accordance with individual state regulations and standards. In addition, while all states make their water rights publicly available, each provides their records in unique formats, meaning that file types, field availability, and terms vary from state to state. This leads to additional challenges to managing resources which cross state lines, or conducting consistent multi-state water analyses. For the first version of HarDWR, we collected the water rights databases from 11 Western States of the United States. In order to preform regional analyses with the collected data, the raw records had to be harmonized into one single format. The Water Data Exchange (WaDE) is a program dedicated to the sharing of water-related data for the Western U.S. in a singular consistent format. Created by the Western States Water Council (WSWC) to facilitate the collection and dissemination of water data among WSWC's member states and the public, WaDE provides an important service for those interested in water resource planning and management in their focus region. Of the services which WaDE provides, the one of the most interesting is the WestDAAT dataset, which is a collection of water rights data provided by the 18 WSWC member states that have been standardized into a single format, much like we had done on a more limited scale with HarDWR v1. For this version of HarDWR we decided to use WestDAAT, specifically a snapshot created in Feburary 2024, as our water rights source data. A full explanation of the benefits gained from this switch can be found in the description of the updated Harmonized Water Rights Records v2.0, but in short it has allowed us to focus more of our efforts on answering research questions and gaining a more realistic understanding of how water rights are allocated. For more information on how the data for WestDAAT was collected, please see the WaDE data summary. Terms of Use While WaDE works directly with the state agencies to collect and standardize the water rights records, the ultimate authority for the water rights data remains the individual states. Each state, and their respective water right authorities, have made their water right records available for non-commercial reference uses. In addition, the states make no guarantees as to the completeness, accuracy, or timeliness of their respective databases, let alone the modifications which we, the authors of this paper, have made to the collected records. None of the states should be held liable for using this data outside of its intended use. As several of the states update their water rights databases daily, the information provided here is not the latest possible, and should not be used for legal purposes. WestDAAT itself has irregular updates. Additional questions about the data the source states provided should be directed to the respective state agencies (see methods.csv and organization.csv files described below). In addition, although data was presented here was not collected directly from the states, several states requested specifically worked disclaimers when sharing their data. These disclaimers are included here as an acknowledgement from where the water rights data is primarily sourced. Colorado: "The data made available here has been modified for use from its original source, which is the State of Colorado. THE STATE OF COLORADO MAKES NO REPRESENTATIONS OR WARRANTY AS TO THE COMPLETENESS, ACCURACY, TIMELINESS, OR CONTENT OF ANY DATA MADE AVAILABLE THROUGH THIS SITE. THE STATE OF COLORADO EXPRESSLY DISCLAIMS ALL WARRANTIES, WHETHER EXPRESS OR IMPLIED, INCLUDING ANY IMPLIED WARRANTIES OF MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE. The data is subject to change as modifications and updates are complete. It is understood that the information contained in the Web feed is being used at one's own risk." Montana: "The Montana State Library provides this product/service for informational purposes only. The Library did not produce it for, nor is it suitable for legal, engineering, or surveying purposes. Consumers of this information should review or consult the primary data and information sources to ascertain the viability of the information for their purposes. The Library provides these data in good faith but does not represent or warrant its accuracy, adequacy, or completeness. In no event shall the Library be liable for any incorrect results or analysis; any direct, indirect, special, or consequential damages to any party; or any lost profits arising out of or in connection with the use or the inability to use the data or the services provided. The Library makes these data and services available as a convenience to the public, and for no other purpose. The Library reserves the right to change or revise published data and/or services at any time." Oregon: "This product is for informational purposes and may not have been prepared for, or be suitable for legal, engineering, or surveying purposes. Users of this information should review or consult the primary data and information sources to ascertain the usability of the information." File Descriptions The unmodified February, 2024 WestDAAT snapshot is composed of nine files. Below is a brief description of each file, as well as how they were utilized for HarDWR. WaDEDataDictionaryTerms.xlsx: As the file's name implies, this is a data dictionary for all of the below named files. This file describes the column names for each of the following files, with the exception of citation.txt which does not have any columns. The descriptions for each file are divided by tab,with the same name as their associated file, within this document. allocationamount.csv: The "main" file of the group, it contains the water right records for each state. Of particular note, each water right is broken down into one or more water allocations. Allocations may be withdrawn from one or more locations, or even multiple allocations associated with a particular location. This is a more subtle and realistic representation of how water is used than what was available in the first version of HarDWR. For the records from some states, this can mean that multiple allocations listed under a single right will appear as rows within this file. citation.txt: A combination of contact information for WaDE personnel, disclaimer about how the data should be used, and guidelines for citing WestDAAT. methods.csv: A file describing the source and method by which WaDE collected water rights data from each state. organization.csv: A file listing the water rights authoritative agencies for each state. sites.csv: This file provides the geographic, and other descriptors, of the physical location of allocations, called 'sites'. To reiterate, it is possible for one allocation to be associated with multiple sites, as well as one site to be associated with multiple allocations. The two descriptors which we were most interested in where the site's coordinates, as well as whether the site was classified as a Point of Diversion (POD) or a Place of Use (POU). As a general rule, PODs are geographic points, while POUs are areas typically represented as property boundaries or irregularly shaped polygons. sites_pouGeometry.csv: For those allocations with a POU site, this file contains the defining points for the associated polygons. variables.csv: A file describing the units in which an allocation's water amount is reported within WestDAAT. This information is essentially a repeat of the 'AllocationFlow_CFS' and 'AllocationVolume_AF' columns within allocationamount.csv, at least for our purposes. watersources: This file describes the source of water from which each site extracts from. For our purposes, this table was used to determine whether the water came from Surface Water, Groundwater, or Unspecified Water.

Lisk, Matthew↗

A search for hard X-rays from five strong extragalactic radio sources

Data from the University of California at San Diego (UCSD) hard X-ray instrument on OSO-7 have been used to search the regions containing 3C 390.3, M87, 3C 273, 3C 111, and 3C 129. We have a weak detection of 3C 390.3 and a possible detection of 3C 111. Sensitive upper limits on hard X-rays have been obtained for the other objects. The available observational data indicate a correlation between hard X-ray emission, self-absorbed gigahertz radio emission, and strong optical emission lines in extragalactic sources. Self-Compton X-ray models have been calculated for the sources observed in this paper, and the angular sizes, lifetimes, and B fields in these sources have been estimated. These model calculations agree with existing VLBI and time variability data.

Mushotzky, R. F.↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗

A Prototype Software to Demonstrate a Data Catalog for Hanford Environmental Datasets

Ensuring that data on long-term environmental remediation at the Hanford Site is high-quality, traceable, and easily accessible is an ongoing challenge, complicated by decades of data collection, multiple contractors maintaining data sources, and the wide range of data types. A centralized data catalog, known as the Hanford Environmental Information and Data Index (HEIDI), has been under development as part of the Hanford Environmental Data Management (HEDM) program to address these challenges. HEIDI fulfills a critical need to bring together a wide range of data types and sizes from multiple authoritative data sources, while documenting the data pedigree and quality information (i.e., traceable to the data source/originator). This document describes additional development and maturation of the HEIDI prototype. Key accomplishments included deploying the catalog software, Esri Geoportal Server, on a server accessible to Hanford Local Area Network users, conducting cybersecurity evaluations, investigating integrated authentication solutions, and conducting functional testing of the catalog prototype. The server-based deployment enabled targeted feedback, leading to enhancements including improved accessibility features and an expanded metadata schema. Specifications for the server-based deployment of the prototype catalog and the HEIDI metadata schema are provided in this document to support subsequent HEIDI deployment by the U.S. Department of Energy Richland Operations Office.

54 ENVIRONMENTAL SCIENCES↗

Modelling above-ground biomass stock over Norway using national forest inventory data with ArcticDEM and Sentinel-2 data

Boreal forests constitute a large portion of the global forest area, yet they are undersampled through field surveys, and only a few remotely sensed data sources provide structural information wall-to-wall throughout the boreal domain. ArcticDEM is a collection of high-resolution (2 m) space-borne stereogrammetric digital surface models (DSM) covering the entire land area north of 60° of latitude. The free-availability of ArcticDEM data offers new possibilities for aboveground biomass mapping (AGB) across boreal forests, and thus it is necessary to evaluate the potential for these data to map AGB over alternative open-data sources (i.e., Sentinel-2). This study was performed over the entire land area of Norway north of 60° of latitude, and the Norwegian national forest inventory (NFI) was used as a source of field data composed of accurately geolocated field plots (n=7710) systematically distributed across the study area. Separate random forest models were fitted using NFI data, and corresponding remotely sensed data consisting of either: i) a canopy height model (ArcticCHM) obtained by subtracting a high-quality digital terrain model (DTM) from the ArcticDEM DSM height values, ii) Sentinel-2 (S2), or iii) a combination of the two (ArcticCHM+S2). Furthermore, we assessed the effect of the forest- and terrain-specific factors on the models’ predictive accuracy. The best model (,i.e., ArcticCHM+S2) explained nearly 60% of the variance of the training set, which translated in the largest accuracy in terms of root mean square error (RMSE=41.4 t/ha). This result highlights the synergy between 3D and multispectral data in AGB modelling. Furthermore, this study showed that despite the importance of ArcticCHM variables, the S2 model performed slightly better than ArcticCHM model. This finding highlights some of the limitations of ArcticDEM, which, despite the unprecedented spatial resolution, is highly heterogeneous due to the blending of multiple acquisitions across different years and seasons. We found that both forest- and terrain-specific characteristics affected the uncertainty of the ArcticCHM+S2 model and concluded that the combined use of ArcticCHM and Sentinel-2 represents a viable solution for AGB mapping across boreal forests. The synergy between the two data sources allowed for a reduction of the saturation effects typical of multispectral data while ensuring the spatial consistency in the output predictions due to the removal of artifacts and data voids present in ArcticCHM data. While the main contribution of this study is to provide the first evidence of the best-case-scenario (i.e., availability of accurate terrain models) that ArcticDEM data can provide for large-scale AGB modelling, it remains critically important for other studies to investigate how ArcticDEM may be used in areas where no DTMs are available as is the case for large portions of the boreal zone.

space-borne imagery↗