Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Open-Source Data for MAC-POSTS: Mobility Data Analytics Center - Prediction, Optimization, and Simulation Toolkit for Transportation Systems

MAC-POSTS (Mobility Data Analytics Center - Prediction, Optimization, and Simulation toolkit for Transportation Systems) is a toolkit for dynamic transportation network modeling. Developed by the Mobility Data Analytics Center (MAC) at Carnegie Mellon University, this package implements many classic dynamic transportation network models, as well as new models proposed by MAC members. It has served as one building block for many other models and research projects. As such, this package used to be treated as an internal research project of the MAC lab, and admittedly, the code base is messy, and the interface is hard to use. However, we are working hard to make it a generally usable and useful toolkit for dynamic transportation network modeling. We would really appreciate any feedback, comments, suggestions, or criticisms.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

FEST: Facility Energy Saving and Securing Technology Using Multi-Source Data: (Milestone 2 Report)

This project aims to demonstrate three technologies developed in-house at LLNL and University of Michigan-Dearborn (UMD), at a military site, including i) Grid Data Crossing (called GriD-Xing) for increasing smart meter data usability, ii) Facility Energy Optimization (called Facility E-GO) for improving facility energy efficiency in both operation and planning perspectives, and iii) Co-simulation tool (called Co-Sim) for enhancing smart meter and energy facility networks resilience and security. In this report, Millstone 2 - Integration A: Integrate GriD-Xing and facility optimization tools is documented.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Facility Energy Saving and Securing Technology Using Multi-Source Data (FEST) (Final Report)

Facility Energy Saving and Securing Technology (FEST) project focuses on developing algorithmic tools and techniques for analyzing cyber-physical security of DoD’s military site facilities, and optimal scheduling, operation and planning of their Distributed Energy Resources (DER). The project is led by LLNL with the team including University of Michigan-Dearborn and XENDEE. The military site partner providing the energy metering data is White Sands Missile Range (WSMR). This report summarizes the work performed during the project and future directions for follow-on research.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Foundational Industrial Energy Dataset (FIED): Open-Source Data on Industrial Facilities

The state of data on industrial energy use has co-evolved over several decades with the demands of industrial energy analysis. The most recent development - analysis in support of decarbonizing the industrial sector - has changed the characteristics of industrial data that are useful for analysts and model developers. Although data and its collection processes may be cast from a conventional viewpoint as objective and free from the influence of social dynamics, this provides an incomplete picture of not only the processes by which information is generated, but also the limitations and opportunities of data to be useful for analysis. The foundational industry energy data set (FIED) is a result of the confluence of trends in open data and the demand for higher resolution industrial energy analysis. The general approach to compiling the FIED involves accessing, filtering, and formatting data published by federal organizations on the Internet for public use. Unlike most industrial energy datasets, which are published by the U.S. Energy Information Administration (EIA), the FIED relies on core datasets from the U.S. Environmental Protection Agency (EPA). The FIED addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization - including estimates of energy use, greenhouse gas emissions, and design capacities - for facilities that are identified by latitude and longitude. This enables local-level analysis of existing combustion equipment, as well as regional comparisons with traditional industrial energy data estimates. The report summarizes the general logic behind compiling the FIED. The FIED itself and its Python code are available from OpenEI and GitHub, respectively.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Biolink Model: A universal schema for knowledge graphs in clinical, biomedical, and translational science

Abstract Within clinical, biomedical, and translational science, an increasing number of projects are adopting graphs for knowledge representation. Graph‐based data models elucidate the interconnectedness among core biomedical concepts, enable data structures to be easily updated, and support intuitive queries, visualizations, and inference algorithms. However, knowledge discovery across these “knowledge graphs” (KGs) has remained difficult. Data set heterogeneity and complexity; the proliferation of ad hoc data formats; poor compliance with guidelines on findability, accessibility, interoperability, and reusability; and, in particular, the lack of a universally accepted, open‐access model for standardization across biomedical KGs has left the task of reconciling data sources to downstream consumers. Biolink Model is an open‐source data model that can be used to formalize the relationships between data structures in translational science. It incorporates object‐oriented classification and graph‐oriented features. The core of the model is a set of hierarchical, interconnected classes (or categories) and relationships between them (or predicates) representing biomedical entities such as gene, disease, chemical, anatomic structure, and phenotype. The model provides class and edge attributes and associations that guide how entities should relate to one another. Here, we highlight the need for a standardized data model for KGs, describe Biolink Model, and compare it with other models. We demonstrate the utility of Biolink Model in various initiatives, including the Biomedical Data Translator Consortium and the Monarch Initiative, and show how it has supported easier integration and interoperability of biomedical KGs, bringing together knowledge from multiple sources and helping to realize the goals of translational science.

60 APPLIED LIFE SCIENCES↗

Newly reconstructed Arctic surface air temperatures for 1979–2021 with deep learning method

A precise Arctic surface air temperature (SAT) dataset, that is regularly updated, has more complete spatial and temporal coverage, and is based on instrumental observations, is critically important for timely monitoring and improving understanding of the rapid change in the Arctic climate. In this study, a new monthly gridded Arctic SAT dataset dated back to 1979 was reconstructed with a deep learning method by combining surface air temperatures from multiple data sources. The source data include the observations from land station of GHCN (Global Historical Climatology Network), ICOADS (International Comprehensive Ocean-Atmosphere Data Set) over the oceans, drifting ice station of Russian NP (North Pole), and buoys of IABP (International Arctic Buoy Programme). The last two are crucial for improving the representation of the in-situ observed temperatures within the Arctic. The newly reconstructed dataset includes monthly Arctic SAT beginning in 1979 and daily Arctic SAT beginning in 2011. This dataset would represent a new improvement in developing observational temperature datasets and can be used for a variety of applications.

54 ENVIRONMENTAL SCIENCES↗

Urban-scale Energy Modeling: Scaling Beyond Tax Assessor Data

In an attempt to attain building-specific characteristics for urban-scale building energy models, county-specific tax assessors’ data is often an initial data source. This data source can contain valuable information such as year built, area, height, HVAC type, and roof/wall descriptions.We will show examples of 2,000 fields from Hamilton County in Tennessee with examples of many fields which are not relevant to urban-scale building energy modeling, are incorrect compared to other data sources, and highlight some lessons learned working with such a data source.There are currently 3,142 counties in the United States, each with their own data format, field definitions, and data access policy. As urban-scale involves city-scale analysis potentially covering multiple counties and matures toward state- or nation-scale analysis, county-by-county approaches are not scalable. While there are efforts to unify these datasets, there is an increasing proliferation of data and algorithms that can cover wider areas and provide more accurate inputs for urban-scale models. This paper summarizes computer vision of imagery, cartographic layers, building type assessment, and model generation used to achieve scalable detection and analysis of buildings.

New, Joshua↗

Linking resource availability to pantropical forest canopy resistance and resilience to cyclone disturbance

Statement of purpose: Tropical cyclones are intensifying and occurring at higher latitudes in recent decades, but the mechanisms underpinning the resistance (ability to withstand disturbance-induced change) and resilience (pace of return to pre-disturbance reference values) of tropical forests to cyclones remains largely unexplored at the pantropical scale. We conducted a meta-analysis to investigate the role of soil resource availability (i.e., total soil phosphorus concentration) in mediating site-level forest canopy resistance and resilience to cyclones pan-tropically. We evaluated cyclone-induced and post-cyclone litterfall mass (g/m2/day), phosphorus (P) and nitrogen (N) fluxes (mg/m2/day), as well as concentrations (mg/g) across 73 case studies in Australia, Guadeloupe, Hawaii, Mexico, Puerto Rico, and Taiwan. The dataset zip file includes three data and two metadata files: - The compiled Litterfall Mass Flux data from tropical forests across the globe prior to and after varying tropical cyclone disturbances are provided in Litterfall_Mass.csv. This data file also includes site location, geographical characteristics, elevation, soil phosphorus concentration, geology, and several variables related to each tropical cyclone disturbance. - The compiled Litterfall Nitrogen and Phosphorus Flux data from tropical forests across the globe prior to and after varying tropical cyclone disturbances are provided in Litterfall_Nutrients.csv. This data file also includes site location, geographical characteristics, elevation, soil phosphorus concentration, geology, and several variables related to each tropical cyclone disturbance. - Tropical cyclone track data compiled from HURDAT2 and IBTrACS databases and used as input in the HURRECON model (https://github.com/hurrecon-model/HurreconR) to generate wind data is provided in hurdat2-1851-2019-052520.txt. - The metadata file (Metadata_Meta-analysis_Litterfall-Mass.pdf) has the complete information on each variable included in the Litterfall_Mass.csv dataset, the data sources, and data processing information. - The metadata file (Metadata_Meta-analysis_Litterfall-Nutrients.pdf) has the complete information on each variable included in the Litterfall_Nutrients.csv dataset, the data sources, and data processing information.

54 ENVIRONMENTAL SCIENCES↗

Riverine Plastic Pollution: Sampling and Analysis Methods

Riverine plastic pollution has been found in all major U.S. rivers, but the exact amount of plastic being released to the oceans has not been quantified. Field studies conducted in U.S. rivers have used a range of sampling and analysis techniques and rarely measured the mass of the plastic collected. Measurements of riverine plastic pollution are needed to calibrate and validate models used to estimate the U.S. riverine plastic emissions to the oceans. This report surveys measurement methods used to quantify riverine pollution and current estimates of U.S. riverine plastic pollution from measurements and models. Measurement methods include field sampling and laboratory analysis. Field sampling methods are described for large (macro) and small (micro) plastic particles. Laboratory analysis methods are described for macro and microplastic with an emphasis on the detailed characterization processes of microplastics. Waterborne leachate analysis is also briefly described. Three models are described that estimate plastic pollution based on mismanaged plastic waste in the river catchment basins. The models were validated and calibrated with global data sources. The data sources were predominantly outside of the U.S., where the magnitude and composition of plastic pollution is different than what is found in U.S. rivers. Comprehensive measurements of riverine plastics are needed not only to characterize the riverine plastic pollution, but also parameterize and validate models of plastic fate and transport. This report also describes five key U.S. rivers that span a range of sizes and environmental conditions that could be sampled to obtain data to support characterization and model development of plastic pollution from rivers to oceans. Sampling and analysis protocol recommendations are made to ensure the highest quality of data are collected in the five rivers.

54 ENVIRONMENTAL SCIENCES↗

Assessing United States County-Level Exposure for Research on Tropical Cyclones and Human Health

Tropical cyclone epidemiology can be advanced through exposure assessment methods that are comprehensive and consistent across space and time, as these facilitate multiyear, multistorm studies. Further, an understanding of patterns in and between exposure metrics that are based on specific hazards of the storm can help in designing tropical cyclone epidemiological research. a) Provide an open-source data set for tropical cyclone exposure assessment for epidemiological research; and b) investigate patterns and agreement between county-level assessments of tropical cyclone exposure based on different storm hazards. We created an open-source data set with data at the county level on exposure to four tropical cyclone hazards: peak sustained wind, rainfall, flooding, and tornadoes. The data cover all eastern U.S. counties for all land-falling or near-land Atlantic basin storms, covering 1996–2011 for all metrics and up to 1988–2018 for specific metrics. We validated measurements against other data sources and investigated patterns and agreement among binary exposure classifications based on these metrics, as well as compared them to use of distance from the storm’s track, which has been used as a proxy for exposure in some epidemiological studies. Our open-source data set was typically consistent with data from other sources, and we present and discuss areas of disagreement and other caveats. Over the study period and area, tropical cyclones typically brought different hazards to different counties. Therefore, when comparing exposure assessment between different hazard-specific metrics, agreement was usually low, as it also was when comparing exposure assessment based on a distance-based proxy measurement and any of the hazard-specific metrics. Our results provide a multihazard data set that can be leveraged for epidemiological research on tropical cyclones, as well as insights that can inform the design and analysis for tropical cyclone epidemiological research.

60 APPLIED LIFE SCIENCES↗

BAMCensus (The Behavior and Advanced Mobility Census Dataset Aggregator) [SWR-25-120]

This software is a high-performance tool developed in Rust for downloading and processing large-scale geospatial datasets, specifically focusing on US Census data. It is designed to address scaling limitations found in existing tools, such as R's [tidycensus](https://walker-data.com/tidycensus/), by providing performant streaming dataset JOIN operations between various US Census datasets (like ACS and LEHD) and their corresponding geometries stored on the TIGER/Lines web server. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. The tool automates the process of joining these data sources, returning aggregated data to the user based on a specified census GEOID type. Its primary motivation stems from the need for a high-performance solution to combine spatial datasets with graph traversals within the context of mobility analysis tooling being developed at NREL's Behavior and Advanced Mobility (BAM) group.

Fitzgerald, Robert [National Renewable Energy Labo↗

High-Fidelity, Large-Scale, Realistic Dataset Development

The final report summarizes the work performed for supporting the ARPA-E Grid Optimization Competition (Challenge 2 and Challenge 3) within the stated period. Challenge 2 For the challenge period, the main responsibility of the team is to investigate, gen- erate, and deliver parts of the data sets for the competition, based on the competition model for Challenge 2, existing data sets from Challenge 1, and data source supplied by other data set teams. Challenge 3 For the challenge period, the main responsibility of the team is to propose, create, deliver, and maintain the data format during the competition period. The data format will specify how the benchmark data will be represented and communicated to competitors. It will also specify how competitors should report back the solutions. The data format will be closely aligned with the problem formulation (maintained by the formulation team) and the solution validation process (maintained by the validation team). Our team is also responsible in investigating, generating, and delivering parts of the data sets for the competition. The data sets will be created based on the competition model for Challenge 3, existing data sets from Challenge 1 and Challenge 2, and data source supplied by other data set teams.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗