Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Enabling pan-repository reanalysis for big data science of public metabolomics data

Public untargeted metabolomics data is a growing resource for metabolite and phenotype discovery; however, accessing and utilizing these data across repositories pose significant challenges. Therefore, here we develop pan-repository universal identifiers and harmonized cross-repository metadata. This ecosystem facilitates discovery by integrating diverse data sources from public repositories including MetaboLights, Metabolomics Workbench, and GNPS/MassIVE. Our approach simplified data handling and unlocks previously inaccessible reanalysis workflows, fostering unmatched research opportunities.

El Abiead, Yasin↗

Data-Driven Buy Clean: Decarbonization and Beyond

This report was compiled to provide recommendations on the availability of public background data from the U.S. Federal life cycle assessment (LCA) Data Commons to be conformant with the Association for Life Cycle Assessment (ACLCA) 2022 Product Category Rule (PCR) Open Standard to build technical tools that can assist industry in creating more comparable Type II Environmental Product Declarations (EPDs) for Federal Buy Clean and sustainability initiatives. The Federal LCA Commons is not only a public data source but also a consistently structured, self-referencing mega-repository for data developed by federal agency experts (in agency repositories) and by academia, nonprofit organizations, and industry (via the US Life Cycle Inventory Database). The Federal LCA Commons Technical Working Group is continuously improving the standardization of data documentation, formatting, and nomenclature to ensure lossless data loading and accurate data representation. This report and appendixes include the following: 1) An introduction to data-driven Buy Clean and decarbonization initiatives at the federal level; 2) The current status and associated challenges with LCA data and EPD standards and comparability; 3) Opportunities for the Federal LCA Commons to support conformance with the ACLCA 2022 PCR Open Standard and provide resources to implement the Federal Sustainability Plan, Buy Clean Program, and Inflation Reduction Act (IRA) sustainability goals and objectives. To date, the Federal LCA Commons is the result of coordinated work by National Renewable Energy Laboratory (NREL), the U.S. Department of Agriculture (USDA), the Environmental Protection Agency (EPA), the National Energy Technology Laboratory (NETL), the Argonne National Laboratory (ANL), the U.S. Army Corps of Engineers (USACE), the Federal Highway Administration (FHWA), the U.S. Forest Service (USFS), the Federal Aviation Administration (FAA), the Department of Defense (DoD) and the National Institute of Standards and Technologies (NIST). The Federal LCA Commons will continue to combine databases from the collaborating agencies while remaining a public resource. There are several initiatives among the collaborating agencies to expand the Federal LCA Commons and dedicated federal funding and resources could accelerate and strengthen these initiatives.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

CO2-Locate: A Dynamic Database and Tool for Accessing National Oil and Gas Well Data to Inform Carbon Storage Projects

The CO2-Locate Database is a growing compilation of publicly available wellbore resources that have been merged based on common attributes across data sources with an attribute schema developed to be consistent across disparate resources, reduce data gaps, and eliminate record redundancy. The first version of CO2-Locate has been published to Energy Data eXchange (EDX) and includes the integrated public wells dataset as well as additional geospatial summary layers of key wellbore characteristics to protect proprietary resources. Additionally, the CO2-Locate database has been deployed into a web application, enabling easy access, data filtering capabilities, and visualization of U.S. wellbore infrastructure by stakeholders to inform injection site selection and risk assessments.

Dyer, Alec S. [NETL Site Support Contractor, Natio↗

Data analytics for intermodal freight transportation applications

With the growth of intermodal freight transportation, it is important that transportation planners and decision-makers are knowledgeable about freight flow data to make informed decisions. This is particularly true with Intelligent Transportation Systems (ITS) offering new capabilities for intermodal freight transportation. Specifically, ITS enables access to multiple different data sources, but they have different formats, resolutions, and time scales. Thus, knowledge of data science is essential to be successful in future ITS-enabled intermodal freight transportation systems. This chapter discusses the commonly used descriptive and predictive data analytic techniques in intermodal freight transportation applications. These techniques cover the entire spectrum of univariate, bivariate, and multivariate analyses. In addition to illustrating how to apply these techniques manually, this chapter will also show how to apply them using the statistical software R. Additional exercises are provided for those who wish to apply the described techniques to more complex problems.

Huynh, Nathan↗

U.S. Geothermal District Heating Systems and Funding Data

This dataset contains detailed information on geothermal district heating (GDH) systems across the United States, including technical specifications, costs, energy usage, and funding sources. Data includes system details such as state, site, location, operational years, temperature, flow rate, capacity, energy use, load factor, costs (original, updated, and operational), funding sources, employment impact, and project status. Additionally included are contributions to GDH systems, including federal, state, and local grants, as well as loans from agencies such as USDA and USDOE. This dataset serves as a resource for analyzing the economic and technical landscape of geothermal district heating in the United States.

15 GEOTHERMAL ENERGY↗

Bayesian merging of numerical modeling and remote sensing for saltwater intrusion quantification in the Vietnamese Mekong Delta

Saltwater intrusion has become one of the most concerning issues in the Vietnamese Mekong Delta (VMD) due to its increasing impacts on agriculture and food security of Vietnam. Reliable estimation of salinity plays a crucial role to mitigate the impacts of saltwater intrusion. Here, this study developed a hybrid technique that merges satellite imagery with numerical simulations to improve the estimation of salinity in the VMD. The salinity derived from Landsat images and by numerical simulations was fused using the Bayesian inference technique. The results indicate that our technique significantly reduces the uncertainties and improves the accuracy of salinity estimates. The Nash–Sutcliffe coefficient is 0.74, which is much higher than that of numerical simulation (0.63) and Landsat estimation (0.6). The correlation coefficient between the ensemble and measured salinity is relatively high (0.88). The variance of the ensemble salinity errors (5.0 ppt 2 ) is lower than that of Landsat estimation (10.4 ppt 2 ) and numerical simulations (9.6 ppt 2 ). The proposed approach shows a great potential to combine multiple data sources of a variable of interest to improve its accuracy and reliability wherever these data are available.

54 ENVIRONMENTAL SCIENCES↗

A Comprehensive Northern Hemisphere Particle Microphysics Data Set From the Precipitation Imaging Package

Microphysical observations of precipitating particles are critical data sources for numerical weather prediction models and remote sensing retrieval algorithms. However, obtaining coherent data sets of particle microphysics is challenging as they are often unindexed, distributed across disparate institutions, and have not undergone a uniform quality control process. This work introduces a unified, comprehensive Northern Hemisphere particle microphysical data set from the National Aeronautics and Space Administration precipitation imaging package (PIP), accessible in a standardized data format and stored in a centralized, public repository. Data is collected from 10 measurement sites spanning 34° latitude (37°N–71°N) over 10 years (2014–2023), which comprise a set of 1,070,000 precipitating minutes. The provided data set includes measurements of a suite of microphysical attributes for both rain and snow, including distributions of particle size, vertical velocity, and effective density, along with higher-order products including an approximation of volume-weighted equivalent particle densities, liquid equivalent snowfall, and rainfall rate estimates. The data underwent a rigorous standardization and quality assurance process to filter out erroneous observations to produce a self-describing, scalable, and achievable data set. Case study analyses demonstrate the capabilities of the data set in identifying physical processes like precipitation phase-changes at high temporal resolution. Bulk precipitation characteristics from a multi-site intercomparison also highlight distinct microphysical properties unique to each location. This curated PIP data set is a robust database of high-quality particle microphysical observations for constraining future precipitation retrieval algorithms, and offers new insights toward better understanding regional and seasonal differences in bulk precipitation characteristics.

54 ENVIRONMENTAL SCIENCES↗

Real-Twin

Real-Twin is a unified, model-agnostic scenario generation tool designed to streamline and standardize the evaluation of emerging mobility technologies. It provides an end-to-end framework that includes robust workflows, integrated tools, and comprehensive metrics to generate, calibrate, and benchmark microscopic traffic simulation scenarios across multiple platforms. Key Features of Real-Twin include: - Unified Scenario Generation: generate transferable, simulation-ready scenarios from heterogeneous data sources using a consistent workflow. - Automated Calibration Workflow: bridges simulation and real-world data, minimizing manual effort and making traffic simulation more accessible to researchers and engineers. - Model-Agnostic Compatibility: supports SUMO, VISSIM, and AIMSUN for cross-platform scenario generation and benchmarking. Enables reliable comparisons and reproducibility across different simulation tools. - Consistent Scenarios across Different Simulators: generate comparable simulation scenarios across different microscopic traffic simulators, providing users the ability to conduct benchmarking and cross-validation that are crucial for ensuring the reliability and reproducibility of simulation results. - Emerging Technology Support: includes a scenario database and pipeline for studying autonomous vehicles (AVs), with planned extensions to CAVs, EVs, and other advanced technologies.

Wang, Chieh (Ross) [Oak Ridge National Laboratory ↗

Transit Rider/Travel Behavior Inventory Survey - Minneapolis-St. Paul Metro - 2010

The Metropolitan Council administered a system-wide transit on-board survey between September and November 2010 as part of the 2010 Travel Behavior Inventory to provide detailed transit usage patterns and rider information to support modeling and planning efforts. The study’s focus was to capture only the most current and reliable data necessary to determine future public transportation needs in the Minneapolis region. The study was structured to collect detailed transit ridership data for different routes during different times of day to develop a disaggregate transit trip table that will support advanced travel demand modeling. The study was designed to leverage existing data sources, such as the 2005 on-board survey, to provide the best quality data to update the Travel Behavior Inventory in the Minneapolis-Saint Paul metropolitan area.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Hestia SW-IFL Onroad Fossil Fuel Carbon Dioxide (FFCO2) product: Road segment-level annual FFCO2 emissions across Arizona (2017-2022), version 1.1

The SW-IFL onroad fossil fuel carbon dioxide (FFCO2) emissions data product represents CO2 emissions from the combustion of fossil fuels by motor vehicles (e.g., passenger cars, trucks, buses, motorcycles) traveling on designated roadways. The emissions are represented geographically on each road segment within the state of Arizona spanning the 2017 to 2022 time period. This data product was developed as part of the Southwest Urban Corridor Integrated Field Laboratory (SW-IFL) project, which aims to provide new knowledge and tools that address urban environmental issues by integrating high-resolution observations, modeling, and civic engagement. The emissions data are provided in CSV (input data, ONR_FFCO2_AZ_county.csv) and GeoPackage form (output polyline objects - about 786,000 road segments, XXXX_AZ_v1.1.gpkg) designated by road class (interstates, arterials, collectors, local). The metadata file (Metadata_SW-IFL_Onroad_annualFFCO2_v1.1.docx) provides details about attributes and data formats. The GeoPackage emissions data are provided separately for local roads and nonlocal roads (interstates, arterials, collectors). The method file (Methods_SW-IFL_Onroad_annualFFCO2_v1.1.docx) describes the data processing flow and data sources. Update on 2024-04-17: Updates were made to both the input emission data file (.csv) and output segment-level emission file (.gpkg). There was an update in county-level emission input data (ONR_FFCO2_AZ_county.csv) and the entire road segments were reprocessed to reflect this update.Update on 2024-04-29: Update was made to one output segment-level emission file (Nonlocal_AZ_v1.0.gpkg). There was an error in the AADT values and the data were reprocessed to reflect this update.Update on 2024-10-22: Temporal coverage was extended to include 2022. VMT values were recalculated using new AADT data and the entire road segments were reprocessed to reflect these updates.

54 ENVIRONMENTAL SCIENCES↗

Monitoring L2 Milestone Summary: Converged HPC Center-Wide Data Analytics

This document is the summary for the ASC 2025 L2 milestone Converged HPC Center-Wide Data Analytics. It describes the convergence of the center-wide monitoring at LC, the data sources involved, and different ways this has enabled analyzing and gaining insights from the data.

97 MATHEMATICS AND COMPUTING↗

NCAR-RAL Surface Hydrometeorological Observation Network Data for LASSO-CACTI Overview Paper

This data set contains the 15 minute resolution surface meteorology and soils data from the 15 NCAR/RAL weather stations that were operated around central Argentina during the RELAMPAGO (Remote sensing of Electrification, Lightning, And Meso-scale/micro-scale Processes with Adaptive Ground Observations) Extended Observing Period (EOP). Data providence, citation, and acknowledgement This ARM data set is a copy of v1.0 of the NCAR data set obtained in June 2024 from https://doi.org/10.26023/KW8Z-F2WX-H0Y. The citation for the original data source is: Gochis, D., et al. 2019. NCAR-RAL Surface Hydrometeorological Observation Network Data. Version 1.0. UCAR/NCAR - Earth Observing Laboratory. https://doi.org/10.26023/KW8Z-F2WX-H0Y Accessed June 2024. In addition to the citation reference and any other acknowledgements, please acknowledge NCAR/EOL in your publications with text such as: “Data provided by NCAR/EOL under the sponsorship of the National Science Foundation. https://data.eol.ucar.edu/”

air temperature↗

USEEIO v2.0, The US Environmentally-Extended Input-Output Model v2.0

USEEIO v2.0 is an environmental-economic model of US goods and services that can be used for life cycle assessment, footprinting, national prioritization, and related applications. This paper describes the development of the model and accompanies the release of a full model dataset as well as various supporting datasets of national environmental totals by US industry. Novel methodological elements since USEEIO v1 models include waste sector disaggregation, final demand vectors for US consumption and production, a domestic form of the model that can be used to separate domestic and foreign impacts, and price adjustment matrices for converting outputs to purchaser price and in various US dollar years. Improvements in modeling national totals of industry and environmental flows are described. The model is validated through reproduction of national totals from input data sources and through analysis of changes from the most recent complete USEEIO model that can be explained based on data updates or method changes. The model datasets can all be reproduced with open source software packages.

54 ENVIRONMENTAL SCIENCES↗

Temporal Variability in Reservoir Surface Area Is an Important Source of Uncertainty in GHG Emission Estimates

Ebullitive methane (CH 4 ) emissions in lentic ecosystems tend to concentrate at river-lake interfaces and within shallow littoral zones. However, inconsistent definitions of the littoral zone and static representations of the lake or reservoir surface area contribute to major uncertainties in greenhouse gas (GHG) emissions estimates, particularly in reservoirs with large water-level fluctuations. This study examines temporal variation in littoral and total surface areas of US reservoirs and demonstrates how different methods and data sources lead to discrepencies in reservoir GHG emissions at large scales and over time. We also explore variability in remotely sensed water occurrence according to maximum surface area, reservoir purposes, and hydrologic regions. Notably, the largest relative variability in surface area is exhibited by small reservoirs with a maximum surface area <1 km 2 and non-hydroelectric reservoirs. Additionally, we use a case study of measured CH 4 emissions from the southeastern United States (Douglas Reservoir) to illustrate the effects of varying surface area on reservoir-wide GHG estimates. Upscaled CH 4 emissions in Douglas Reservoir differed by nearly two-fold depending on the source of total surface area data and whether estimates accounted for seasonal fluctuations in surface area. During seasonal drawdown in Douglas Reservoir, relative littoral area varies non-linearly; periods of lower pool elevation (and thus larger relative littoral area) likely contribute disproportionately high CH 4 emission rates compared to the commonly sampled summer season when water levels are at full-pool elevation. Improved GHG monitoring and upscaling techniques require accounting for temporal variability in reservoir surface extent and littoral area.

54 ENVIRONMENTAL SCIENCES↗

An Open-Access Repository of Synchrophasor Data Quality Examples: Curation and Example Applications

Synchrophasor measurements are critical in providing wide-area situational awareness to power system operators. However, data artifacts may be introduced due to various issues such as loss of communication, loss of GPS signal, internal clock error, and vendor-specific implementation of phasor estimation algorithms. Tools designed to provide actionable insights from synchrophasor data, hence, must be designed to be robust to these data quality issues. In this work, two years of synchrophasor data sourced from multiple electric utilities in the United States were analyzed to identify examples of data quality problems. These examples were then labeled and published in the Grid Event Signature Library, a publicly available repository of power system measurements hosted by the Oak Ridge National Laboratory. This paper describes the data curation process, and illustrates two application use cases where the dataset can be valuable to the research community. In the first use case, a random forest classifier is trained to distinguish power system disturbance signatures from data anomalies introduced in synchrophasor measurements due to clock errors. The second use case studies the impact of data quality issues on an example synchrophasor application (specifically, event start time determination). The choice of data quality problems investigated is informed by the examples in the repository curated in this work.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Multifidelity_Timeseries

SAND2025-03305O Multifidelity Timeseries is a user-friendly tool designed to create advanced models for analyzing time-series data. It offers three modeling options, allowing users to choose the best fit for their specific needs. The software efficiently processes multiple data sources without the need for complex sampling methods. It helps uncover patterns and insights using data. The result is it is easier to make informed decisions for projects. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Katona, Ryan [Sandia National Lab. (SNL-CA), Liver↗

Google Health Trends performance reflecting dengue incidence for the Brazilian states

Abstract Background Dengue fever is a mosquito-borne infection transmitted by Aedes aegypti and mainly found in tropical and subtropical regions worldwide. Since its re-introduction in 1986, Brazil has become a hotspot for dengue and has experienced yearly epidemics. As a notifiable infectious disease, Brazil uses a passive epidemiological surveillance system to collect and report cases; however, dengue burden is underestimated. Thus, Internet data streams may complement surveillance activities by providing real-time information in the face of reporting lags. Methods We analyzed 19 terms related to dengue using Google Health Trends (GHT), a free-Internet data-source, and compared it with weekly dengue incidence between 2011 to 2016. We correlated GHT data with dengue incidence at the national and state-level for Brazil while using the adjusted R squared statistic as primary outcome measure (0/1). We used survey data on Internet access and variables from the official census of 2010 to identify where GHT could be useful in tracking dengue dynamics. Finally, we used a standardized volatility index on dengue incidence and developed models with different variables with the same objective. Results From the 19 terms explored with GHT, only seven were able to consistently track dengue. From the 27 states, only 12 reported an adjusted R squared higher than 0.8; these states were distributed mainly in the Northeast, Southeast, and South of Brazil. The usefulness of GHT was explained by the logarithm of the number of Internet users in the last 3 months, the total population per state, and the standardized volatility index. Conclusions The potential contribution of GHT in complementing traditional established surveillance strategies should be analyzed in the context of geographical resolutions smaller than countries. For Brazil, GHT implementation should be analyzed in a case-by-case basis. State variables including total population, Internet usage in the last 3 months, and the standardized volatility index could serve as indicators determining when GHT could complement dengue state level surveillance in other countries.

59 BASIC BIOLOGICAL SCIENCES↗

Driver Identification Dataset

The ORNL Driver Identification Dataset was created to collect and analyze driving behavior data from 50 different drivers. Each driver operated a 2014 Kenworth T270 Class 6 truck around Fort Collins, Colorado while various data sources recorded their driving behavior and vehicle performance. The dataset includes CANbus (Controller Area Network) data, GPS data, inertial measurement data, and biometric data from a heart rate monitor. A cyberattack was executed during each drive, which caused multiple dashboard warning lights to illuminate and set the tachometer and speedometer to zero, regardless of actual speed. The attack was stopped either after one minute or if the driver pulled over. By downloading the dataset, you agree to the following: 1) I will not use or disclose the data for any purpose other than Research as that term is defined in 10 CFR 745.102. 2) I will not, under any circumstances, request or accept private or linking identifiers for the data used. 3) I will not attempt to determine the identity of the individuals associated with the data. 4) I will use appropriate safeguards to prevent the use or disclose of the data for any purpose other than Research.

99 GENERAL AND MISCELLANEOUS↗