Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DATA BASE MANAGEMENT”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Inception Based Deep Convolutional Neural Network for Remaining Useful Life Estimation of Turbofan Engines

Accurate estimation of the remaining useful life (RUL) is a key component of condition based maintenance (CBM) and prognosis and health management (PHM). Data-based models for the estimation of RUL are of particular interest because expert knowledge of systems is not always available and physical modeling is often not feasible. In this paper, a deep convolutional neural network (CNN) architecture is investigated for its ability to estimate the RUL of turbofan engines. The input to the model is a window of time series data collected from the engine under test. Inputting raw sensor data allows features to be learned instead of manually determined. To incorporate the ability to detect features of differing lengths, inception modules are used in the neural network architecture. The model is trained and tested using the new Commercial Modular Aero-Propulsion System Simulation (N-CMAPSS) data set and high prognosis accuracy was achieved. The developed model was used in the 2021 PHM Society Data Challenge and received second place, further validating its ability to accurately estimate RUL.

DeVol, Nathaniel↗

Making a Water Data System Responsive to Information Needs of Decision Makers

Evidence-based environmental management requires data that are sufficient, accessible, useful and used. A mismatch between data, data systems, and data needs for decision making can result in inefficient and inequitable capital investments, resource allocations, environmental protection, hazard mitigation, and quality of life. In this paper, we examine the relationship between data and decision making in environmental management, with a focus on water management. We focus on the concept of decision-driven data systems —data systems that incorporate an assessment of decision-makers' data needs into their design. The aim of the research was to examine the process of translating data into effective decision making by engaging stakeholders in the development of a water data system. Using California's legislative mandate for state agencies to integrate existing water and other environmental data as a case study, we developed and applied a participatory approach to inform data-system design and identify unmet data needs. Using workshops and focused stakeholder meetings, we developed 20 diverse use cases to assess data sources, availability, characteristics, gaps, and other attributes of data used for representative decisions. Federal and state agencies made up about 90% of the data sources, and could readily adapt to a federated data system, our recommended model for the state. The remaining 10% of more-specialized data, central to important decisions across multiple use cases, would require additional investment or incentives to achieve data consistency, interoperability, and compatibility with a federated system. Based on this assessment, we propose a typology of different types of data limitations and gaps described by stakeholders. We also propose technical, governance, and stakeholder engagement evaluation criteria to guide planning and building environmental data systems. Data-system governance involving both producers and users of data was seen as essential to achieving workable standards, stable funding, convenient data availability, resilience to institutional change, and long-term buy-in by stakeholders. Our work provides a replicable lesson for using decision-maker and stakeholder engagement to shape the design of an environmental data system, and inform a technical design that addresses both user and producer needs.

Cantor, Alida↗

Method for automatic correction of offset drift in online sensors

Abstract Successful operation and optimization of water treatment systems hinge on the availability of high-quality online sensor measurements. Ideally, the available measurements should be simultaneously accurate (i.e., unbiased and precise), representative, voluminous, and timely. This remains a pain-point in current water infrastructures, forming a barrier to a wider adoption of advanced and autonomous control systems. While short-lived symptoms, such as outliers and spikes, can be detected or corrected with state-of-the-art tools for fault detection and identification, it is much more difficult to detect, diagnose, and correct the symptoms of slow faults, such as changes in offset or sensitivity due to drift. The time scale of drift is often longer than the time scales of the system dynamics of interest. Moreover, sensor drift has been shown to occur at the same time and with similar rates when sensors are exposed to the same conditions. This challenges data quality management strategies based on redundancy. In this contribution, we develop a new method, including both a hands-off sensor calibration mechanism and an information-seeking control architecture that can handle the unique challenge of simultaneous and similar drift in online sensors.

Chowdhury, Dhrubajit↗

Developing Novel Performance Measures for Traffic Congestion Management and Operational Planning Based on Connected Vehicle Data

In this study, the authors present their efforts in exploring a new type of traffic data, referred to as internet-connected vehicle (ICV) data, for traffic congestion management and operational planning. Most currently manufactured vehicles contain onboard GPS and cellular modules, and they constantly connect to automobile manufacturers' clouds via cellular networks and upload their status. Some automobile manufacturers have recently redistributed the nonpersonal part of such data, such as geolocation, to third-party organizations for innovative applications. Compared with the traditional vehicle GPS data, the ICV data contain high-resolution GPS waypoints accompanied with the vehicles' abnormal moving events (e.g., hard braking). The ICV data also have huge potential in congestion management and operational planning. They explore to identify and analyze traffic congestion on both freeways and arterials using the ICV data. The ICV data adopted for this research are redistributed by Wejo Data Service, representing 10%-15% of all moving vehicles in the Dallas-Fort Worth (DFW) area in Texas. Through one case study for a freeway segment and one for an arterial segment, new traffic performance metrics based on the characteristics of ICV data have been presented. The highlights of these efforts are as follows: (I) queue length and propagation at freeway bottlenecks can be directly measured based on where and when most internet-connected vehicles slow down and join the queue; (II) an internet-connected vehicle's actual delay time on arterials can be directly measured according to its slow movement percentage, without assuming the nondelay travel speed; and (III) the ICV data set are also combined with the high-resolution traffic signal events to generate a ground-truth time-space diagram (TSD) on arterials - a common visualization of arterial signal performance for transportation planning and operations.

33 ADVANCED PROPULSION SYSTEMS↗

The global spectrum of plant form and function: enhanced species-level trait dataset

Here we provide the ‘Global Spectrum of Plant Form and Function Dataset’, containing species mean values for six vascular plant traits. Together, these traits –plant height, stem specific density, leaf area, leaf mass per area, leaf nitrogen content per dry mass, and diaspore (seed or spore) mass – define the primary axes of variation in plant form and function. The dataset is based on ca. 1 million trait records received via the TRY database (representing ca. 2,500 original publications) and additional unpublished data. It provides 92,159 species mean values for the six traits, covering 46,047 species. The data are complemented by higher-level taxonomic classification and six categorical traits (woodiness, growth form, succulence, adaptation to terrestrial or aquatic habitats, nutrition type and leaf type). Data quality management is based on a probabilistic approach combined with comprehensive validation against expert knowledge and external information. Intense data acquisition and thorough quality control produced the largest and, to our knowledge, most accurate compilation of empirically observed vascular plant species mean traits to date.

54 ENVIRONMENTAL SCIENCES↗

Blockchain Empowered Provenance Framework for Sensor Identity Management and Data Flow Security in Fossil-based Power Plants

The overall goals of this project are divided into 3 important phases. First, a peer-to-peer SCADA network is to be established where identity profile of sensors from fossil-power plants will be stored in an immutable manner in order to verify data flow integrity in real time. Second, it is aimed to develop a Blockchain-based provenance platform to audit equipment operations along with data and process integrity violations. In this phase, various operational rules will be incorporated in the ledger on behalf of the operating sensor so its abnormal use can be easily detected. In addition, it is aimed to achieve provenance when the rate of transactions is exceedingly high, thus enabling to design a scalable provenance scheme that can work in real time when number of data collecting points or sensor units increase in fossil-based power plants. Finally, a working testbed will be developed to integrate the Blockchain platform with sensor-based fossil power plant comprising of Remote Terminal Units (RTU) and SCADA devices. Using multiple Raspberry-PIs mimicking the role of RTUs that are interfaced with multiple sensors devices will be deployed to serve as a critical component of fossil power plant and the performance metrics such as throughput, latency, bootstrapping time, etc., will be evaluated on the testbed.

20 FOSSIL-FUELED POWER PLANTS↗

Evaluation of Data Catalog Software for Hanford Site Environmental Datasets

Environmental information and data underpin achievement of the U.S. Department of Energy (DOE) Office of Environmental Management (EM) mission at the Hanford Site. The Hanford Environmental Data Management (HEDM) Program is the DOE Richland Operations Office (RL) approach to develop and implement a formal program for managing environmental data and the associated records, materials, and systems at the Hanford Site. The current project, contract, organization, and contractor-specific efforts at managing environmental data sets are insufficient to provide orderly, long-term, site-wide access. A vital element to be created within the HEDM program plan is a catalog of data sources, called the Hanford Environmental Information and Data Index (HEIDI), that will enable long-term access and retrievability for the multiple independent sources of data that might otherwise be difficult to discover. This report compares leading open source and commercial data catalog platforms using criteria to assess the functionality needed to develop the HEIDI catalog of Hanford data sources that connects and exchanges data with established Hanford Local Area Network (HLAN) enterprise information technology systems. Proprietary platforms evaluated included ArcGIS Enterprise Sites, Junar, OpenDataSoft, and Socrata, and non-proprietary platforms included Energy Data eXchange (EDX), Comprehensive Knowledge Archive Network (CKAN), and DKAN (a Drupal-based open data portal based on CKAN). Capabilities supporting data discoverability, retrieval, and archival, as well as metadata standard requirements and integration into the HLAN were rated as either failing to meet requirements (F), meeting requirements (M), or exceeding requirements by delivering additional desired features (E). The lowest rating for any capability area was assigned as the overall rating for the platform. These findings enable DOE-RL and the contractors implementing the HEDM plan to focus on candidate tools likely to meet the requirements for implementing HEIDI. All of the platforms receiving an overall rating of ‘F’ were unable to be deployed on Hanford infrastructure or within dedicated cloud resources. A propriety software-as-a-service (SaaS) model of delivering a data catalog (e.g., found in software such as Junar and OpenDataSoft) favors consistency across customers at the expense of customization and configurable roles that are needed for Hanford work. Hosting data on a shared commercial platform places limits on dataset size (maximum of 240 Mb for OpenDataSoft), a significant limitation for HEIDI implementation. EDX, a government data catalog based on CKAN, received the ‘F’ rating due to an inability to incorporate authentication from HLAN into the system. Among platforms rated ‘M’ or ‘E’, only the Socrata platform had a SaaS delivery model. In contrast to other SaaS platforms, Socrata provided custom roles and gateways that allow local datasets to be incorporated into an online catalog. Socrata also complies with the Federal Risk and Authorization Management Program, a significant benefit for cloud-based management of Hanford data. The other platforms rated ‘M’ or ‘E’, ArcGIS Enterprise Sites, CKAN, and DKAN, provide fully self-hosted options, allowing for greater control and flexibility with the HEIDI catalog. These widely used tools have supportive communities of practice, extensive customization options, and demonstrated deployments that provide evidence that they can meet requirements, often deliver additional desired features, and work well with federal government systems. Completely customized alternatives built on a collection of applications were not evaluated because achieving similar performance to CKAN or DKAN requires substantial resources, especially in the absence of the active communities that have grown to support these tools. ArcGIS Enterprise Sites, Socrata, CKAN, and DKAN were evaluated as strong candidates for successful implementation with HEIDI.

54 ENVIRONMENTAL SCIENCES↗

Hydrologic applicability of satellite-based precipitation estimates for irrigation water management in the data-scarce region

Reliable precipitation estimates are crucial for planning and managing water resources, monitoring hydrologic extremes, and fulfilling irrigation water requirements. Accurate precipitation estimates are particularly challenging in complex mountain terrains, where monitoring gauges are often sparsely distributed due to their remote locations, and high installation and long-term operation costs. Recent advances in satellite-based precipitation estimates offer promising opportunities to improve our understanding of hydrologic processes and their applications for irrigation water management. Several datasets are available varying considerably in terms of their data sources, quality control methods, estimation procedure, and spatiotemporal resolutions. Choosing the most suitable dataset for a particular application is a complex task. In this study, we (1) evaluate the performance of six satellite-based precipitation estimates (SPEs): i) CHIRPS v2.0, ii) CMORPH v1.0, iii) ERA5, iv) IMERG v6, v) MSWEP v2.8, and vi) PERSIANN-CDR against the gauge precipitation using continuous statistical and categorical indices, (2) integrate SPEs with a calibrated semi-distributed hydrologic model to predict streamflow, and (3) demonstrate practical implications of improved streamflow prediction for irrigation water management in the central Himalayan region, Nepal. Our results illustrate that satellite-based precipitation estimates have competitive performance in capturing a wide range of rainfall characteristics, with demonstrated variability across river basins and time scales. Further, there are no significant discrepancies observed in satellite-based precipitation estimates for estimating irrigation water requirements for the three major crops (maize, wheat, and paddy) during the cropping period across the selected river basins, showing a greater promise for irrigation water management planning and decision making.

54 ENVIRONMENTAL SCIENCES↗

Cross-facility science with the Superfacility Project at LBNL

As data sets from DOE user science facilities grow in both size and complexity there is an urgent need for new capabilities to transfer, analyze and manage the data underlying scientific discoveries. LBNL's Superfacility project brings together experimental and observational research instruments with computational and network facilities at the National Energy Research Scientific Computing Center (NERSC) and the Energy Sciences Network (ESnet) with the goal of enabling user science. Here, we report on recent innovations in the Superfacility project, including advanced data management, API-based automation, real-Time interactive user interfaces, and supported infrastructure for 'edge' services.

Enders, B↗

Management and Storage of Scientific Data

Scientific discoveries rely heavily on efficient access, search, and management of massive data sets. Data management technologies have, for decades, provided foundational capabilities for scientific computing. Just as storage, input/output (I/O), and data management have been fundamental to simulation-based science for many years, so too are capable data-management technologies key to the success of today’s scientific workflows utilizing data intensive and machine learning (ML) techniques. The Department of Energy, Office of Science, Advanced Scientific Computing Research (ASCR) program has invested broadly in data-management research focused on high-performance computing (HPC) systems, from parallel file systems that store data to application software that makes these systems more productive. Still, advances in technology combined with growing diversity of supported science strongly motivate continued investment in this area. In January 2022, ASCR convened a workshop to identify priority research directions in the area of data management for high-performance and scientific computing. Attendees were challenged to identify promising approaches that would support the breadth of the DOE mission, including the explosion of artificial intelligence (AI) uses and the growing needs of experimental and observational science. Technological and science drivers were identified and considered as they relate to key aspects of data management such as interfaces, architectural design, and FAIR principles (Findable, Accessible, Interoperable, and Reusable). The thoughts of the workshop participants were distilled into a set of four priority research directions with the potential for high impact on DOE science. These research directions are summarized in the following pages.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

AmeriFlux BASE Flux/Met Data QA/QC and Processing (AMF-BASE-QAQC) v1.0.0

The AmeriFlux BASE Flux/Met Data QA/QC and Processing (AMF-BASE-QAQC) code provides tools to review and prepare continuous flux/met data submitted to the AmeriFlux Management Project for publication as the AmeriFlux BASE data product. The code provides 3 core functionalities: Format QA/QC assesses submitted data files for compliance with the required submission format; Data QA/QC assesses the data quality; BASE Publish prepares the data for publication.

Christianson, Danielle↗

Bridging Cloud and Edge Computing at NREL Using CONNECT: Cloud Optimized Networking for Next-Gen Edge Computing Technologies [Slides]

CONNECT is an innovative on-premise hardware and software solution that integrates edge and cloud computing infrastructure at NREL. Built on the AWS Greengrass middleware and leveraging the MQTT protocol, CONNECT enables real-time data streaming from IoT devices and gateways to both cloud and local services, empowering researchers to rapidly capture, analyze, and act upon edge-generated data while leveraging cloud capabilities. The platform addresses research infrastructure challenges by providing a pre-approved platform which is already configured with the correct networking and cybersecurity baselines thus eliminating procurement delays and enabling on-demand availability. CONNECT's hybrid architecture efficiently manages burstable workloads, allowing research teams to dynamically scale computational capacity, handle peak data loads, and reduce operational bottlenecks. Advanced capabilities include built-in GPU support for executing machine learning models which enables low-latency inference at the edge from models trained in the cloud. This architecture supports real-time analytics and filtering, providing a mechanism to allow only transmitting and processing high-value data. Cloud-based configuration management permits engineers to manage on-premise systems remotely, optimizing operational efficiency. By bridging edge and cloud computing, CONNECT provides NREL researchers with a flexible, scalable platform that accelerates scientific discovery while maintaining robust security and performance standards.

97 MATHEMATICS AND COMPUTING↗

Field surveying data of low-cost networked flood sensors in southeast Texas

Floods are common natural disasters worldwide and pose substantial risks to life, property, food production, and natural resources. Effective measures for flood mitigation and warning are essential. Southeast Texas is still at significant risk of flooding, and Lamar University is assisting the region with asset management of a flood sensor network for flooding events. This network provides real-time water stage information. Lamar University developed a survey program to measure elevation and coordinates at each sensor site location to make this data more useful for flood monitoring and mapping. This paper overviews the measurement of the elevation and coordinates of 74 networked flood sensors and various flood stage thresholds at critical points that flood decision-makers can use for reference at each site. In the first phase of this program, these sensors were deployed throughout a 7-county region spanning nearly 6,000 square miles in Southeast Texas. The latitude and longitude of the sensors and their elevations were determined using survey-grade Global Navigation Satellite System (GNSS) technology. Various Continually Operating Reference Stations (CORS) were utilized for post-processing to achieve sub-inch resolution. The flood stage thresholds, water level sensors elevation, and the elevations and positions of other critical surrounding points are viewable to the public through two online repositories and a web-based sensor management dashboard. The data is used to aid with decisions related to road closures or modeling efforts by mitigation decision-makers, emergency managers, and the public, including the Texas Department of Transportation, Houston Transtar, the National Weather Service, and the Sabine River Authority of Texas (SRA).

54 ENVIRONMENTAL SCIENCES↗

NuScale Systems Engineering Program Overview and Status

NuScale systems engineering program overview, including technical data management and model-based engineering, product structures, architecture views, and model-based systems engineering. Includes a discussion on technical risk and opportunity management and technical readiness levels, as well as risk-informed performance-based principles and methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data for Identifying the best high-biomass sorghum hybrids based on biomass yield potential and feedstock quality affected by nitrogen fertility management under various environments

Data were collected from agronomy fields in Urbana and Ewing, IL, during the 2022 and 2023 growing seasons. The dataset includes dry biomass yield, nitrogen, phosphorus, and potassium concentrations and removals, and chemical composition elements (cellulose, hemicellulose, lignin, and soluble fractions) for 13 high-biomass sorghum hybrids. data_sharing.xlsx contains 20 columns and 104 rows. Below is the explanation of all variables in the file: Year: 2022; 2023 Location: Urbana, IL; Ewing, IL N rate (kg-N/ha): 0; 112 Hybrid #: H1-H13 Pedigree: Pedigree for 13 hybrids Dry biomass yield (Mg/ha): Aboveground dry biomass yield N (g/kg): Nitrogen concentration in plant tissue P (g/kg): Phosphorus concentration in plant tissue K (g/kg): Potassium concentration in plant tissue N (kg/ha): Nitrogen removal by aboveground biomass P (kg/ha): Phosphorus removal by aboveground biomass K (kg/ha): Potassium removal by aboveground biomass Cellulose (g/kg): Cellulose concentration in plant tissue Hemicellulose (g/kg): Hemicellulose concentration in plant tissue Lignin (g/kg): Lignin concentration in plant tissue Soluble (g/kg): Soluble concentration in plant tissue Cellulose (Mg/ha): Cellulose content in aboveground biomass Hemicellulose (Mg/ha): Hemicellulose content in aboveground biomass Lignin (Mg/ha): Lignin content in aboveground biomass Soluble (Mg/ha): Soluble content in aboveground biomass

environmental adaptability↗