Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert↗

Adaptable Standards for Discovery, Access, and Usability of Oak Ridge National Laboratory’s Data Portals and Catalogs

Oak Ridge National Laboratory (ORNL) is leveraging its established capabilities and subject matter expertise in data curation, governance, management, national security, and risk assessment and mitigation to support the US Department of Energy (DOE) Grid Modernization Initiative. Using standards modeled by the National Institute of Standards and Technology (NIST), the Data Curation Network (DCN), the Oak Ridge Leadership Computing Facility (OLCF), and other leading organizations in the fields of energy research, high-performance computing, and national and homeland security, ORNL seeks to provide a federated approach to research data discovery, use, and interoperability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Multimodal Event Catalog and Waveform Data Set That Supports Explosion Monitoring from Nevada, U.S.A.

Multimodal, curated data sets and nuisance event catalogs remain rare in the explosion monitoring community relative to curated seismic data sets. The source of this relative absence is the difficultly in deploying multimodal receivers that sense the seismic, acoustic, and other modalities from multiphysics sources. We provide such a data set in this study that delivers seismic, infrasound, and electromagnetic (magnetometer) sensor records collected over a two–week period, within 255 km of a 10 ton buried chemical explosion called DAG–4 that was located at 37.1146°, –116.0693° on 22 June 2019 21:06:19.88 UTC. This catalog includes 485 seismic, seismoacoustic, and infrasound–only events that an expert analyst manually built by reviewing waveforms from 29 seismic and infrasound sensors. Our data release includes waveforms from these 29 seismic, infrasound, and seismoacoustic stations and two magnetometer stations and their station metadata. We deliver these waveforms in NNSA KB Core CSS.w format (i4) with a corresponding wfdisc table that provides the header information. Here, we expect that this data set will provide a valuable, benchmark resource to develop signal processing algorithms and explosion monitoring methods against manual, human observations.

58 GEOSCIENCES↗

An Update on the Geothermal Data Repository's Data Standards and Pipelines: Geospatial Data and Distributed Acoustic Sensing Data: Preprint

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented data standards and automated data pipelines for the following data types: 1) drilling data, 2) geospatial datasets, and 3) DAS data. An additional data pipeline is proposed for stimulation data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how we can improve this process.

cloud-optimized↗

An Update on the Geothermal Data Repository's Data Standards and Pipelines: Geospatial Data and Distributed Acoustic Sensing Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented data standards and automated data pipelines for the following data types: 1) drilling data, 2) geospatial datasets, and 3) DAS data. An additional data pipeline is proposed for stimulation data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how we can improve this process.

cloud-optimized↗

An Update on the Geothermal Data Repository's Data Standards and Pipelines: Geospatial Data and Distributed Acoustic Sensing Data

The Department of Energy's (DOE) Geothermal Data Repository (GDR) team has implemented or is currently implementing data standards and automated data pipelines for the following geothermal data types: 1) drilling data, 2) geospatial datasets, and 3) Distributed Acoustic Sensing (DAS) data. These data standards and pipelines are intended to improve the real-world applicability of geothermal machine learning outputs through improving the quality of data. More specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, allowing more time to be spent on actual research. By automating this process, the burden of standardization is taken off of the user, overall increasing the availability of standardized data. This paper provides an update on the GDR's transition toward data standardization through automated data pipelines and calls for feedback from the community on how the GDR team can improve this process.

cloud-optimized↗

Electronic structure simulations in the cloud computing environment

The transformative impact of modern computational paradigms and technologies, such as high-performance computing, quantum computing, and cloud computing, has opened up profound new opportunities for scientific simulations. Scalable computational chemistry is one beneficiary of this technological progress. The main focus of this paper is on the performance of various quantum chemical formulations, ranging from low-order methods to high-accuracy approaches, implemented in different computational chemistry packages, such as NWChem, NWChemEx, SPEC, ExaChem, and FLOSIC codes on the Azure Quantum Element (AQE) Microsoft cloud services. We pay particular attention to the intricate workflows for performing composite chemistry simulations, associated data curation, and mechanisms for accuracy assessment, as defined by the enabling cloud Computational Chemistry as a Service (CCaaS). Our focus also extends to Arrows' automated workflow for high throughput simulations. Finally, we provide a perspective on the role of cloud computing in supporting the mission of leadership computational facilities (LCFs).

computational chemistry, electronic structure, Clo↗

Data-Centric AI and the Open Energy Data Initiative (OEDI)

This presentation emphasizes the critical importance of data-centric AI. The limitations of model-centric AI when dealing with poor or insufficient data are highlighted, and it is illustrated how training models on inaccurate or noisy data leads to suboptimal results. This talk advocates for a hybrid approach that combines a focus on data quality and model parameters to achieve optimal results. The Open Energy Data Initiative (OEDI) is introduced as a valuable resource for obtaining high-quality energy-related datasets, hosting nearly 2,000 publicly accessible datasets, including 99 solar-related datasets, totaling over 2.7 petabytes of data. OEDI's data lakes enable users to query and work with data without extensive transfers. In conclusion, the significance of data-centric AI and adherence to data curation best practices is emphasized, positioning OEDI as a prime source of high-quality data for AI and machine learning in the renewable energy sector.

AI↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine

03 NATURAL GAS↗

OEDI—Solar Grid Integration Data and Analytics Library

As a part of the Open Energy Data Initiative, this effort aims to develop and demonstrate novel distribution state estimation, control optimization, and transient analysis as well as provide access to data, data integration, and mapping information. More specifically, the focus of the effort will be on physics-based distribution system state estimation, hybrid (physics-based and machine learning) distribution optimal power flow, and event detection/analysis for solar integration and analytics. This work will enable reproducible, robust, replicable, and generalizable R&D in simulation and emulation of solar system integration. These test models and datasets will provide an integrated library for developing and testing power system operation technologies. To make the library user-friendly, this project will provide data curation tools such as data translators, mapping scripts and APIs, database schemas and metadata, interfaces and user dashboard, source code for the reference algorithms, description of the use-cases/scenarios, and comprehensive information on all the assumptions.

14 SOLAR ENERGY↗

GHRSST-14 DAS-TAG Report

The DAS-TAG provides the informatics and data management expertise in emerging information technologies for the GHRSST community. It provides expertise in data and metadata formats and standards, fosters improvements for GHRSST data curation, experiments with new data processing paradigms, and evaluates services and tools for data usage. It provides a forum for producer and distributor data management issues and coordination.

data processing↗

Curating Virtual Data Collections

NASAs Earth Observing System Data and Information System (EOSDIS) contains a rich set of datasets and related services throughout its many elements. As a result, locating all the EOSDIS data and related resources relevant to particular science theme can be daunting. This is largely because EOSDIS data's organizing principle is affected more by the way they are produced than around the expected end use. Virtual collections oriented around science themes can overcome this by presenting collections of data and related resources that are organized around the user's interest, not around the way the data were produced. Virtual collections consist of annotated web addresses (URLs) that point to data and related resource addresses, thus avoiding the need to copy all of the relevant data to a single place. These URL addresses can be consumed by a variety of clients, ranging from basic URL downloaders (wget, curl) and web browsers to sophisticated data analysis programs such as the Integrated Data Viewer.

data integration↗

Evolution of NASA's Earth Science Digital Object Identifier Registration System

NASA's Earth Science Data and Information System (ESDIS) Project has implemented a fully automated system for assigning Digital Object Identifiers (DOIs) to Earth Science data products being managed by its network of 12 distributed active archive centers (DAACs). A key factor in the successful evolution of the DOI registration system over last 7 years has been the incorporation of community input from three focus groups under the NASA's Earth Science Data System Working Group (ESDSWG). These groups were largely composed of DOI submitters and data curators from the 12 data centers serving the user communities of various science disciplines. The suggestions from these groups were formulated into recommendations for ESDIS consideration and implementation. The ESDIS DOI registration system has evolved to be fully functional with over 5,000 publicly accessible DOIs and over 200 DOIs being held in reserve status until the information required for registration is obtained. The goal is to assign DOIs to the entire 8000+ data collections under ESDIS management via its network of discipline-oriented data centers. DOIs make it easier for researchers to discover and use earth science data and they enable users to provide valid citations for the data they use in research. Also for the researcher wishing to reproduce the results presented in science publications, the DOI can be used to locate the exact data or data products being cited.

DOI; digital object identifier; mappin↗

Real-Time Health Monitoring for Gas Turbine Components Using Online Learning and High-Dimensional Data (Final Report)

Capital-intensive turbomachinery, such as gas turbines and combined cycle plants, are constantly being monitored for performance anomalies, faults, and physical degradation. Although these power-generating assets are equipped with hundreds of sensors, existing monitoring tools can only handle moderate-sized data. As a result, only a handful of aggregate metrics are used to monitor machine health. At the same time, developing advanced tools suitable for large datasets have been restricted by the lack of appropriate data. The objective of this proposal was to demonstrate a Big Data analytics framework for fault detection and diagnosis in gas turbine applications. We develop a predictive analytics framework methodology guided by these experimental data, industrial data from our collaborators, and physics-based models with engineering domain knowledge. Our analytics framework consists of four key components: (1) a data curation process that addresses data storage, data quality assessments, and integrity checks, (2) a feature engineering component that utilizes statistical methods and transformation algorithms guided by physics-based models to extract high-fidelity fault features that can be leveraged for fault detection and classifying fault severities, (3) a Machine Learning-based fault detection and diagnostics algorithms for detecting operational and hardware faults in the combustion and the turbines section. We utilize two industry-class gas turbine component test rigs to generate first of its kind data for critical gas turbine faults with varying severity levels. Advanced gas turbine test facilities will be interrogated using state-of-the-art instrumentation techniques to build fault signatures and data trends for key combustor and turbine faults. Data generated from a combustor test rig (Georgia Tech) and a turbine test rig (Penn State) during both normal operation and with seeded faults serve as the basis for the Big Data sets. The test conditions in the two test facilities include common, critical events that occur in the operation. Utilizing the combustor test rig, we examine two common combustor faults: lean blowout and centerbody degradation. For the turbine section we develop analytic models for monitoring cooling faults in the gas turbine.

20 FOSSIL-FUELED POWER PLANTS↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

Second Report of the Nuclear Data Subcommittee of the Nuclear Science Advisory Committee

The central importance of the nuclear data curated by the US Nuclear Data Program (USNDP) for clean energy generation, national security, nonproliferation, medical applications, and space exploration as well as basic science was described in a prior report issued by the DOE/NSF Nuclear Science Advisory Committee subcommittee on Nuclear Data (NSAC-ND) in September 2022. In this report, we present a set of fourteen (14) recommendations that will enhance and advance DOE-NP's stewardship of nuclear data. The first three recommendations focus on the existing core USNDP capabilities, namely: 1) Support the nuclear structure evaluation workforce to improve the currency, consistency, and accessibility of the Evaluated Nuclear Structure Data File (ENSDF); 2) Enhance nuclear reaction evaluation within the USNDP in support of the Evaluated Nuclear Data File (ENDF) through expansion of the workforce and integration of high-performance computing, automation, and machine learning and; 3) Continue atomic mass evaluation in support AME and NUBASE databases. This is followed by eight (8) recommendations representing new cross-cutting initiatives involving both measurement and evaluation to address outstanding nuclear data needs. These new initiatives require a highly trained, diverse workforce that includes personnel with expertise from both inside and outside the nuclear physics community from which evaluators have traditionally been recruited. As such, many of these initiatives are accomplished via a Topical Nuclear Data Collaborations (TNDC). A TNDC is made up of domestic and international stakeholders, subject matter and nuclear data experts, and nuclear data evaluators and features a workforce development plan to ensure that nuclear data evaluators maintain currency in the relevant applications and are seen as equity partners in the endeavor. These include: 1) Establish a coordinated effort to improve evaluation and modeling in nuclear astrophysics for stellar dynamics, multi-messenger astronomy and nucleosynthesis; 2) Initiate a TNDC to develop and maintain nuclear structure evaluation beyond discrete states, including nuclear level densities, photon strength functions and photonuclear data for improved reaction modeling, and exploring nuclear structure at finite temperature; 3) Create a TNDC to perform correlated fission data evaluation, including cross sections, fragment yields, v(A), v(E n ) for nuclear energy, national security, nonproliferation and basic science; 4) From a panel of subject matter experts to establish and annually update a roster of key decay data to nurture its accelerated dissemination including both measurement and evaluation for targeted high-value nuclides for national security, nonproliferation and medical applications; 5) Comprehensive, consistent neutron-induced structure and reaction data for nuclear energy, national security, nonproliferation and planetary nuclear spectroscopy; 6) Charged-particle stopping powers for detector design, space effects and ion beam therapy; 7) High-energy reactions for space exploration and medical nuclide production, and; 8) The creation of an infrastructure for open data and data preservation for use by the entire nuclear physics community. All told, these initiatives require approximately $6.5M increase in NP support of the USNDP in fiscal year 2023 dollars and would require at least 3-5 years to carry out due to the length of time needed to recruit and train new nuclear data researchers. This relatively modest investment would help ensure that the fruits of the nuclear data research carried out by DOE-NP and its collaborators would be brought to bear to address some of the most important needs of our nation and the world. To ensure effective execution of this plan, we present an overview of recruitment, training, and retention goals for the USNDP, the centerpiece of which is a mutually agreed upon code of conduct. Finally, we identify the facility and instrumentation needed to perform the recommended experimental activities. This includes a short review of target fabrication capabilities, reactors, neutron beam, light- and heavy-stable ion, gamma-ray, high-energy and radioactive ion beam facilities. Lastly, a more complete appendix of experimental facilities previously compiled is included with new input provided for 6 facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗