Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Sources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Transfer learning for probabilistic localization of hidden cracks in concrete structures

Abstract The utility of discriminative supervised learning models built using multiple training-data sources is investigated for hidden crack localization in concrete. Feed-forward neural network (FFNN) is chosen as the model architecture, and transfer learning is used to assimilate the information obtained from different sources (computational physics simulations and laboratory experiments). The labeled training data consists of values of a damage index and the known locations of hidden cracks. The classification models need to learn how the presence of damage (hidden cracks) affects the damage index at different sensors for different test conditions. To this end, diagnostic FFNN models are built by sequentially adding and training new hidden layers to assimilate labeled information from computer models (different model geometries, test conditions, crack lengths, crack locations) and laboratory experiments on a plain cement slab. These transfer learning-based models are then used to localize damage in concrete specimens that reflect real-world conditions (i.e., specimens with steel reinforcement and randomly distributed aggregate). The actual damage state in these specimens is determined by extracting cores and performing petrographic studies on the extracted cores. The damage probability estimated by transfer learning-based models is compared with the petrographic damage rating index (DRI) to identify the most suitable approach to train the diagnostic models. The transfer learning-based diagnostic methodology shows promise and could be used in various structural health monitoring applications, where sufficient labeled data are typically not available from a single data source.

Miele, S.↗

Congruity of genomic and epidemiological data in modelling of local cholera outbreaks

Cholera continues to be a global health threat. Understanding how cholera spreads between locations is fundamental to the rational, evidence-based design of intervention and control efforts. Traditionally, cholera transmission models have used cholera case-count data. More recently, whole-genome sequence data have qualitatively described cholera transmission. Integrating these data streams may provide much more accurate models of cholera spread; however, no systematic analyses have been performed so far to compare traditional case-count models to the phylodynamic models from genomic data for cholera transmission. Here, we use high-fidelity case-count and whole-genome sequencing data from the 1991 to 1998 cholera epidemic in Argentina to directly compare the epidemiological model parameters estimated from these two data sources. We find that phylodynamic methods applied to cholera genomics data provide comparable estimates that are in line with established methods. Our methodology represents a critical step in building a framework for integrating case-count and genomic data sources for cholera epidemiology and other bacterial pathogens.

59 BASIC BIOLOGICAL SCIENCES↗

Travel Patterns and Characteristics of Low-Income Population in New York State: 2017 Update

This study examines the key characteristic of low-income people, focusing on New York State populations and households and their comparison with the rest of the United States. The characteristics includes their demographics, trip activities, accessibility, travel attitudes, and equity. The major data source used is 2017 National Households Travel Survey (NHTS). Supplemental data sources are also used such as American Community Survey and Census Transportation Planning Products for a more comprehensive analysis.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Automatic Traffic Queue-End Identification using Location-Based Waze User Reports

Traffic queues, especially queues caused by non-recurrent events such as incidents, are unexpected to high-speed drivers approaching the end of queue (EOQ) and become safety concerns. Though the topic has been extensively studied, the identification of EOQ has been limited by the spatial-temporal resolution of traditional data sources. This study explores the potential of location-based crowdsourced data, specifically Waze user reports. It presents a dynamic clustering algorithm that can group the location-based reports in real time and identify the spatial-temporal extent of congestion as well as the EOQ. The algorithm is a spatial-temporal extension of the density-based spatial clustering of applications with noise (DBSCAN) algorithm for real-time streaming data with an adaptive threshold selection procedure. Here, the proposed method was tested with 34 traffic congestion cases in the Knoxville, Tennessee area of the United States. It is demonstrated that the algorithm can effectively detect spatial-temporal extent of congestion based on Waze report clusters and identify EOQ in real-time. The Waze report-based detection are compared to the detection based on roadside sensor data. The results are promising: The EOQ identification time of Waze is similar to the EOQ detection time of traffic sensor data, with only 1.1 min difference on average. In addition, Waze generates 1.9 EOQ detection points every mile, compared to 1.8 detection points generated by traffic sensor data, suggesting the two data sources are comparable in respect of reporting frequency. The results indicate that Waze is a valuable complementary source for EOQ detection where no traffic sensors are installed.

99 GENERAL AND MISCELLANEOUS↗

Multisource Data Fusion Outage Location in Distribution Systems via Probabilistic Graphical Models

Efficient outage location is critical to enhancing the resilience of power distribution systems. However, accurate outage location requires combining massive evidence received from diverse data sources, including smart meter (SM) last gasp signals, customer trouble calls, social media messages, weather data, vegetation information, and physical parameters of the network. This is a computationally complex task due to the high dimensionality of data in distribution grids. In this paper, we propose a multi-source data fusion approach to locate outage events in partially observable distribution systems using Bayesian networks (BNs). A novel aspect of the proposed approach is that it takes multi-source evidence and the complex structure of distribution systems into account using a probabilistic graphical method. Our method can radically reduce the computational complexity of outage location inference in high-dimensional spaces. The graphical structure of the proposed BN is established based on the network’s topology and the causal relationship between random variables, such as the states of branches/customers and evidence. Utilizing this graphical model, accurate outage locations are obtained by leveraging a Gibbs sampling (GS) method, to infer the probabilities of de-energization for all branches. Compared with commonly-used exact inference methods that have exponential complexity in the size of the BN, GS quantifies the target conditional probability distributions in a timely manner. As a result, a case study of several real-world distribution systems is presented to validate the proposed method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A general spatial-temporal framework for short-term building temperature forecasting at arbitrary locations with crowdsourcing weather data

Weather forecasting has been a critical component to predict and control building energy consumption for better building energy management. Without accessibility to other data sources, the onsite observed temperatures or the airport temperatures are used in forecast models. In this paper, we present a novel approach by utilizing the crowdsourcing weather data from neighboring personal weather stations (PWS) to improve the weather forecast accuracy around buildings using a general spatial-temporal modeling framework. The final forecast is based on the ensemble of local forecasts for the target location using neighboring PWSs. Our approach is distinguished from existing literature in various aspects. First, we leverage the crowdsourcing weather data from PWS in addition to public data sources. In this way, the data is at much finer time resolution (e.g., at 5-minute frequency) and spatial resolution (e.g., arbitrary location vs grid). Second, our proposed model incorporates spatial-temporal correlation information of weather variables between the target building and a set of neighboring PWSs so that underlying correlations can be effectively captured to improve forecasting performance. Here, we demonstrate the performance of the proposed framework by comparing to the benchmark models on temperature forecasting for a building located at an arbitrary location at San Antonio, Texas, USA. In general, the proposed model framework equipped with machine learning technique such as Random Forest can improve forecasting by 50% compares with persistent model and has 90% chance to outperform airport forecast in short-term forecasting. In a real-time setting, the proposed model framework can provide more accurate temperature forecasting results compared with using airport temperature forecast for most forecast horizon. Moreover, we analyze the sensitivity of model parameters to gain insights on how crowdsourcing data from the neighboring personal weather stations impacts forecasting performance. Finally, we implement our model in other cities such as Syracuse and Chicago to test the model's performance in different landforms and climate types.

54 ENVIRONMENTAL SCIENCES↗

Learning From User Behavior: A Survey-Assist Algorithm for Longitudinal Mobility Data Collection

GPS-based travel surveys are widely used in mobility studies to gather crucial qualitative data, like purpose, transportation mode and replaced mode. However, survey response still poses a burden to users, especially in long-term mobility studies, leading to response fatigue. We explore a survey-assist strategy to ease this burden by a novel, user-level modeling approach that leverages past responses from each user to predict responses for new trips, without relying on external data sources like GIS data. We investigate three main algorithms for predicting responses: (i) clustering trips and extrapolating responses for similar trips, (ii) using random forest classification, and (iii) clustering that uses a hybrid algorithm to determine spatial structure, which is then fed as input to a classic random forest classifier. The clustering approach can flexibly predict responses for even complex qualitative survey questions; it achieved F-scores of 65%. The random forest pipeline uses architecture that restricts it to predicting three predetermined survey questions: trip purpose, mode, and replaced mode. However, it achieved F-scores of 78%. While the survey-assist approach has been implemented by several proprietary systems, to our knowledge, this is the first exploration in the academic literature. It follows that this is also the first rigorous evaluation of multiple algorithms that can implement the approach. The evaluation uses a large scale, publicly available, longitudinal dataset consisting of ~ 92k trips from 235 users over a period of roughly one and a half years. With this approach, travel surveys can be pre-filled with the predicted responses for each trip, thus streamlining the survey process for users. Combined with an active learning system that requests user input on low-confidence predictions, models can be updated and improved over time to better support the long-term collection of longitudinal qualitative data.

clustering↗

The case for data science in experimental chemistry: examples and recommendations

The physical sciences community is increasingly taking advantage of the possibilities offered by modern data science to solve problems in experimental chemistry and potentially to change the way we design, conduct and understand results from experiments. Successfully exploiting these opportunities involves considerable challenges. In this Expert Recommendation, we focus on experimental co-design and its importance to experimental chemistry. We provide examples of how data science is changing the way we conduct experiments, and we outline opportunities for further integration of data science and experimental chemistry to advance these fields. Our recommendations include establishing stronger links between chemists and data scientists; developing chemistry-specific data science methods; integrating algorithms, software and hardware to ‘co-design’ chemistry experiments from inception; and combining diverse and disparate data sources into a data network for chemistry research.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Real-time Multi-granular Analytics Framework for HIT Systems

Streaming analytics is the process of ingesting and digesting live data from multiple data sources. In the healthcare domain, as the importance of extracting immediate insights while data are streaming into the system grows, the focus is shifting from batch processing to streaming analytics. With data increasing dramatically at high speeds, many informatics designs have been proposed to adapt healthcare domain into this new environment. In our previous work, we introduced a prototype of health informatics technology (HIT) framework that aims to address challenges in adopting state-of-the-art technologies to enable advanced healthcare analytic tasks in new streaming environments. We recently made major updates to the framework so that anomaly from multiple streaming data sources at different granularity levels can be detected in near real-time. In this paper, we detail the implementation and deployment of the framework in Kubernetes clusters and report its performances when tested on electronic health record (EHR) data of Veterans Affairs.

Park, Byung↗

BEPAM Model Code and CABBI Simulation Results for "Assessing the Additional Carbon Savings with Biofuel"

This dataset consists of various input data that are used in the GAMS model. All the data are in the format of .inc which can be read within GAMS or Notepad. Main data sources include: acreage data (acre), crop budget data ($/acre), crop yield data (e.g. bushel/acre), Soil carbon sequestration data (KgCO2/ha/yr). Model details can be found in the "Assessing the Additional Carbon Savings with Biofuel" and GAMS model package. ## File Description (1) GAMS Model.zip: This includes all the input files and scripts for running the model (2) Table*.csv: These files include the data from the tables in the manuscript (3) Figure2_3_4.csv: This contains the data used to create the figures in the manuscript (4) BaselineResults.csv: This includes a summary of the model results. (5) SensitivityResults_*.csv: Model results from the various sensitivity analyses performed (6) LUC_emission.csv: land use change emissions by crop reporting district for changes of pasturelands to annual crops.

Anticipated baseline approach↗

Commercial, industrial, and institutional discount rate estimation for efficiency standards analysis: Sector-level data 1998–2021

Underlying each of the Department of Energy’s (DOE’s) federal appliance and equipment energy conservation standards are a set of complex analyses of the projected costs and benefits of regulation. Any new or amended standard must be designed to achieve significant additional energy conservation, provided that it is technologically feasible and economically justified (42 U.S.C. 6295(o)(2)(A)). DOE determines economic justification based on whether the benefits exceed the burdens, considering a variety of factors, including the economic impact of the standard on consumers of the product and the savings in lifetime operating cost compared to any increase in price or maintenance expenses (42 U.S.C. 6295(o)(2)(B)). As part of this determination, DOE conducts a Life-Cycle Cost (LCC) analysis, which models the combined impact of appliance first cost and operating cost changes on a representative commercial building sample in order to identify the fraction of customers achieving LCC savings or incurring net cost at the considered efficiency levels. Thus, the commercial discount rate value(s) used to calculate the present value of energy cost savings within the LCC model implicitly plays a role in estimating the economic impact of potential standard levels. This report provides an in-depth discussion of the commercial discount rate estimation process. It is an update to previous reports on estimating commercial discount rates from firm-level financial data (Fujita, 2016). Major topics covered in this report include: Discount rate estimation methods and rationale; -Data sources used and data limitations; -Discount rate distributions for use in standards analysis; -Discount rate estimation methods and distributions specific to the small business subgroup analysis. Going forward, this report will be updated as data allow and analyses necessitate.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Commercial, industrial, and institutional discount rate estimation for efficiency standards analysis: Sector-level data 1998–2022

Underlying each of the U.S. Department of Energy’s (DOE’s) federal appliance and equipment energy conservation standards are a set of complex analyses of the projected costs and benefits of regulation. Any new or amended standard must be designed to achieve significant additional energy conservation, provided that it is technologically feasible and economically justified (42 U.S.C. 6295(o)(2)(A)). DOE determines economic justification based on whether the benefits exceed the burdens, considering a variety of factors, including the economic impact of the standard on consumers of the product and the savings in lifetime operating cost compared to any increase in price or maintenance expenses (42 U.S.C. 6295(o)(2)(B)). As part of this determination, DOE conducts a life-cycle cost (LCC) analysis, which models the combined impact of appliance first cost and operating cost changes on a representative commercial building sample to identify the fraction of customers achieving LCC savings or incurring net cost at the considered efficiency levels. Thus, the commercial discount rate value(s) used to calculate the present value of energy cost savings within the LCC model implicitly plays a role in estimating the economic impact of potential standard levels. This report provides an in-depth discussion of the commercial discount rate estimation process. It is an update to previous reports on estimating commercial discount rates from firm-level and sector-level financial data (e.g., Fujita, 2021, 2016). Major topics covered in this report include the following: -Discount rate estimation methods and rationale -Data sources used and data limitations -Discount rate distributions for use in standards analysis -Discount rate estimation methods and distributions specific to the small business subgroup analysis A version of this analysis was most recently released in 2022. Going forward, this report will be updated as data allow and analyses necessitate.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Commercial, industrial, and institutional discount rate estimation for efficiency standards analysis: Sector-level data 1998–2023

Underlying each of the U.S. Department of Energy’s (DOE’s) federal appliance and equipment energy conservation standards are a set of complex analyses of the projected costs and benefits of regulation. Any new or amended standard must be designed to achieve significant additional energy conservation, provided that it is technologically feasible and economically justified (42 U.S.C. 6295(o)(2)(A)). DOE determines economic justification based on whether the benefits exceed the burdens, considering a variety of factors, including the economic impact of the standard on consumers of the product and the savings in lifetime operating cost compared to any increase in price or maintenance expenses (42 U.S.C. 6295(o)(2)(B)). As part of this determination, DOE conducts a life-cycle cost (LCC) analysis, which models the combined impact of appliance first cost and operating cost changes on a representative commercial building sample to identify the fraction of customers achieving LCC savings or incurring net cost at the considered efficiency levels. Thus, the commercial discount rate value(s) used to calculate the present value of energy cost savings within the LCC model implicitly plays a role in estimating the economic impact of potential standard levels. This report provides an in-depth discussion of the commercial discount rate estimation process relying on the Capital Asset Pricing Model (CAPM) to estimate a business’ cost of equity, and by adding a risk adjustment factor to the risk-free rate associated with long-term U.S. Treasury bonds to estimate their cost of debt. It is an update to previous reports on estimating commercial discount rates from firm-level and sector-level financial data (e.g., Fujita, 2021, 2016). Major topics covered in this report include the following: • Discount rate estimation methods and rationale • Data sources used and data limitations • Discount rate distributions for use in standards analysis • Discount rate estimation methods and distributions specific to the small business subgroup analysis.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART): Data Format and Preparation Guidance (V.5.0)

The Joint Office of Energy and Transportation maintains the Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART), which provides a centralized hub for submitting electric vehicle (EV) charging infrastructure data directed by the Federal Highway Administration (23 CFR 680.112(a)-(c)). EV-ChART provides a streamlined data submission process and an integrated set of analytic tools, connects to other data sources, and empowers data sharing and access across stakeholders, including the public. Any data shared publicly will be aggregated and anonymized to stay in accordance with 23 CFR 680. This EV-ChART Data Format and Preparation Guidance provides a comprehensive overview of the data reporting requirements as authorized under 23 CFR 680.112(a)-(c)). The guidance is intended to be used alongside the EV-ChART Data Input Template, which defines the tabular data structure that these data submissions must follow. Per 23 CFR 680.112(a)-(c), the annual and quarterly data submissions are required of all National Electric Vehicle Infrastructure (NEVI) Formula Program projects, as well as projects for the construction of publicly accessible EV chargers that are funded with funds made available under Title 23, United States Code, including any EV charging infrastructure project funded with federal funds that is treated as a project on a federal-aid highway. One-time data submissions are required of both the NEVI Formula Program projects and grants awarded under 23 U.S.C. 151(f) for projects that are for EV charging stations located along and designed to serve the users of designated Alternative Fuel Corridors (AFCs). Other information and data required in 23 CFR 680, such as 23 CFR 680.112(d), 23 CFR 680.116(c), and 23 CFR 680.106(a), are not discussed in this guidance.

33 ADVANCED PROPULSION SYSTEMS↗

STAIR 2.0: A Generic and Automatic Algorithm to Fuse Modis, Landsat, and Sentinel-2 to Generate 10 m, Daily, and Cloud-/Gap-Free Surface Reflectance Product

Remote sensing datasets with both high spatial and high temporal resolution are critical for monitoring and modeling the dynamics of land surfaces. However, no current satellite sensor could simultaneously achieve both high spatial resolution and high revisiting frequency. Therefore, the integration of different sources of satellite data to produce a fusion product has become a popular solution to address this challenge. Many methods have been proposed to generate synthetic images with rich spatial details and high temporal frequency by combining two types of satellite datasets—usually frequent coarse-resolution images (e.g., MODIS) and sparse fine-resolution images (e.g., Landsat). In this paper, we introduce STAIR 2.0, a new fusion method that extends the previous STAIR fusion framework, to fuse three types of satellite datasets, including MODIS, Landsat, and Sentinel-2. In STAIR 2.0, input images are first processed to impute missing-value pixels that are due to clouds or sensor mechanical issues using a gap-filling algorithm. The multiple refined time series are then integrated stepwisely, from coarse- to fine- and high-resolution, ultimately providing a synthetic daily, high-resolution surface reflectance observations. We applied STAIR 2.0 to generate a 10-m, daily, cloud-/gap-free time series that covers the 2017 growing season of Saunders County, Nebraska. Moreover, the framework is generic and can be extended to integrate more types of satellite data sources, further improving the quality of the fusion product. View Full-Text

47 OTHER INSTRUMENTATION↗

Times (and the Grid) are A-Changin'

<p style="text-align: left;"><span style="background-color: rgb(255, 255, 255); font-family: &quot;Neue Plak&quot;, -apple-system, BlinkMacSystemFont, Roboto, &quot;Helvetica Neue&quot;, Helvetica, Tahoma, Arial, sans-serif; font-size: 16px; color: rgb(111, 114, 135);">This webinar highlighted two datasets that provide forward-looking projections of electric grid GHG emissions for the United States. These data sets were created to help present-day analyses and decision-support tools incorporate what we know about the ongoing evolution of the U.S. electric grid.</span><p style="text-align: left;"><br><p style="text-align: left;"><span style="background-color: rgb(255, 255, 255); font-family: &quot;Neue Plak&quot;, -apple-system, BlinkMacSystemFont, Roboto, &quot;Helvetica Neue&quot;, Helvetica, Tahoma, Arial, sans-serif; font-size: 16px; color: rgb(111, 114, 135);">Additionally, the webinar compared these two databases to other data sources to summarize differences (i.e., attributional versus consequential emissions rates, age and source of underlying data, regionalization approach, and boundary conditions) and highlighted the value and need for expert guidance when an analyst or developer is choosing how to represent electric-sector GHG emissions.</span><p style="text-align: left;"><br><p style="text-align: left;"><span style="background-color: rgb(255, 255, 255); font-family: &quot;Neue Plak&quot;, -apple-system, BlinkMacSystemFont, Roboto, &quot;Helvetica Neue&quot;, Helvetica, Tahoma, Arial, sans-serif; font-size: 16px; color: rgb(111, 114, 135);">This webinar was intended to set the stage for ACLCA roundtable events to further discuss these issues.</span>

Jamieson, Matthew↗

WBS 1.2.3.405 - Life Cycle Assessment of Storage Technologies

Recent commitments by the Biden administration have established targets to achieve a net-zero energy system by 2050. Meeting these targets will spur a rapid transition to clean energy technologies and a commensurate need to develop and deploy energy storage technologies at scale. Pumped Storage Hydro (PSH) is expected to be part of this solution because its ability to provide grid flexibility and stability and enable the dispatching of disparate variable renewable energy technologies. Despite PSH being a mature technology with a history of deployment dating back several decades, there is very little information on the greenhouse gas (GHG) implications of PSH as compared to other storage technologies. The objective of this project is to perform a full lifecycle assessment (LCA) of new PSH projects in the U.S. This LCA includes all project phases (resource extraction, construction, operation, maintenance, end-of-life). The functional unit for this study is 1 kWh electricity delivered by system to grid substation connection point and the estimated lifetime for our base case is 80 years. Data used in this study are based on over 30 potential PSH projects that are in preliminary planning phases and are represent a wide range of potential closed-loop PSH systems in terms of location, technology, and capacity. The project approach, data sources, and modeling assumptions have been informed by a technical review committee of stakeholders that include experts from academia, national and international government, industry, and utilities. The GHGs and energy return on investment (EROI) from PSH will be compared to other storage technologies (e.g., stationary battery storage). Results from this project will improve the PSH community's understanding of the environmental impacts and sustainability of new PSH projects and how PSH compares to other storage technologies. The approach used in this project relies on open-source programming. The analysis framework (source code and data) and will be made publicly available at the end of the project. In addition to reporting results for the base case, we will perform rigorous sensitivity analysis to identify the major drivers, understand impacts of different configurations, and future energy markets. Results from this project will be published in a suitable journal.

ENERGY PLANNING, POLICY, AND ECONOMY,HYDRO ENERGY↗

Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART): Data Format and Preparation Guidance, Version 2.0

The Joint Office of Energy and Transportation maintains the Electric Vehicle Charging Analytics and Reporting Tool (EV-ChART), which provides a centralized hub for submitting electric vehicle (EV) charging infrastructure data directed by the Federal Highway Administration (23 CFR 680.112) EV-ChART will provide a streamlined data submission process and an integrated set of analytic tools, connect to other data sources, and empower data sharing and access across stakeholders, including the public. Any data shared publicly will be aggregated and anonymized to stay in accordance with 23 CFR 680. This EV-ChART Data Format and Preparation Guidance provides a comprehensive overview of the data reporting requirements as authorized under 23 CFR 680.112. The guidance is intended to be used alongside the EV-ChART Data Input Template, which defines the tabular data structure that these data submissions must follow.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗