Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

VA EDH Data Curation Documentation – FY2021, Rev. 2

The health and well-being of the Nation’s men and women who have served in uniform is the highest priority for the U.S. Department of Veterans Affairs (VA). VA is committed to providing timely access to high-quality, recovery-oriented, evidence-based mental health care that anticipates and responds to Veterans’ needs and supports the reintegration of returning Service members into their communities. VA is working to eliminate suicide among all Veterans by developing and implementing innovative suicide prevention approaches and resources.

60 APPLIED LIFE SCIENCES↗

VA EDH Data Curation Documentation FY23-Q1

The health and well-being of the Nation’s men and women who have served in uniform is the highest priority for the U.S. Department of Veterans Affairs (VA). VA is committed to providing timely access to high-quality, recovery-oriented, evidence-based mental health care that anticipates and responds to Veterans’ needs and supports the reintegration of returning Service members into their communities. VA is working to eliminate suicide among all Veterans by developing and implementing innovative suicide prevention approaches and resources.

99 GENERAL AND MISCELLANEOUS↗

VA Community Determinants of Health Data Curation Documentation FY26-Q3

The U.S. Department of Veterans Affairs (VA) places the health and well-being of our nation’s veterans as its top priority. VA is dedicated to offering timely access to high-quality, evidence-based mental health care that meets the needs of veterans and supports their reintegration into society. One of our core missions is to prevent suicide among veterans through innovative approaches and resources.

99 GENERAL AND MISCELLANEOUS↗

Curating Carbon Storage Data for Reuse: Enabling Research and Modeling from Earth’s Surface to Subsurface

The volume of public geologic carbon storage (GCS) data resources has continued to increase in recent years as the result of an increase in funding from government, industry, and academia towards national, basin, regional and field scale studies to ensure carbon capture and storage becomes a commercially viable operation. Despite the increasing volume of data, GCS data applied towards analyses such as geologic, cost, and risk modeling continues to be multi-sourced and often disparate in nature, published across government agencies, websites, data repositories and buried in derivative reports and documents. Much of the time preparing for an analysis and derivative product development is spent collecting, aggregating, transforming and preparing input data. There have been significant efforts within the DOE National Energy Technology Laboratory’s Carbon Storage Program to optimize multi-source, multi-scale subsurface geologic data curation and aggregation to support data discovery, interoperability, and reuse. Methods include the use of artificial intelligence, machine learning, and data science techniques. This talk will discuss the workflows, best practices, and processes developed to support the aggregation and curation of data through the whole system – surface to subsurface data - that support multi-scale, multi-purpose analysis for carbon storage research.

Morkner, Paige↗

A Relevancy Algorithm for Curating Earth Science Data Around Phenomenon

Earth science data are being collected for various science needs and applications, processed using different algorithms at multiple resolutions and coverages, and then archived at different archiving centers for distribution and stewardship causing difficulty in data discovery. Curation, which typically occurs in museums, art galleries, and libraries, is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest. Curating data sets around topics or areas of interest addresses some of the data discovery needs in the field of Earth science, especially for unanticipated users of data. This paper describes a methodology to automate search and selection of data around specific phenomena. Different components of the methodology including the assumptions, the process, and the relevancy ranking algorithm are described. The paper makes two unique contributions to improving data search and discovery capabilities. First, the paper describes a novel methodology developed for automatically curating data around a topic using Earthscience metadata records. Second, the methodology has been implemented as a standalone web service that is utilized to augment search and usability of data in a variety of tools.

earth science phenomena↗

Rucio at LSST/Rubin

In this presentation, we will explore the Rucio experience with the Rubin Observatory experiment. Our discussion will cover several key areas: Scalability Tests: Insights into the performance and scalability evaluations of Rucio in the context of Rubin's data needs and what we have learned, especially with many small files. Role in Rubin's Data Curation: Rubin's Data Butler: An overview of how Rucio, along with with Rubin's Data Butler using Hermes-K, which involves message passing through Kafka, is integrated in the Rubin's data curation system. Monitoring and Support: Current status of Rucio and PostgreSQL monitoring and Rucio deployment and support within the Rubin environment. Tape RSE Implementation: Deal with the order of magnitude more files going to tape than HEP. Future Needs: An examination of Rubin's evolving requirements for Rucio services and how we plan to address them.

Lee, Dennis [Fermilab]↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

Data Mining for Anomaly Detection

The Vehicle Integrated Prognostics Reasoner (VIPR) program describes methods for enhanced diagnostics as well as a prognostic extension to current state of art Aircraft Diagnostic and Maintenance System (ADMS). VIPR introduced a new anomaly detection function for discovering previously undetected and undocumented situations, where there are clear deviations from nominal behavior. Once a baseline (nominal model of operations) is established, the detection and analysis is split between on-aircraft outlier generation and off-aircraft expert analysis to characterize and classify events that may not have been anticipated by individual system providers. Offline expert analysis is supported by data curation and data mining algorithms that can be applied in the contexts of supervised learning methods and unsupervised learning. In this report, we discuss efficient methods to implement the Kolmogorov complexity measure using compression algorithms, and run a systematic empirical analysis to determine the best compression measure. Our experiments established that the combination of the DZIP compression algorithm and CiDM distance measure provides the best results for capturing relevant properties of time series data encountered in aircraft operations. This combination was used as the basis for developing an unsupervised learning algorithm to define "nominal" flight segments using historical flight segments.

Biswas, Gautam↗

New developments in space radiation research at NASA: Annotating data using a novel radiation biology ontology

Like many interdisciplinary sciences, data producers and consumers in the field of radiation biology often use a wide variety of terminology to describe their experiments and data. Furthermore, space systems and technologies are rapidly evolving, and a shared understanding and common terminology for these is also lacking. The efficiency of research organizations can be enhanced by standardizing metadata through the use of knowledge resources like ontologies. Employing a sophisticated model such as a formal ontology to standardize metadata enables automated data acquisition processes and supports more complete, accurate meta-analysis through more efficient and complete data discovery and retrieval, particularly when using multiple data sources. Thus, we developed the Radiation Biology Ontology (RBO) in order to improved radiation biology metadata uniformity and transparency. We used open-source software (the Ontology Development Kit, Protégé and WebProtégé) and worked within the OBO Foundry framework, which includes a set of ontology development principles and practices for ontology consistency, uniformity, and accountability. The RBO has now been incorporated into two radiation research data repositories, NASA’s GeneLab omics database (https://genelab.nasa.gov), and the European Commission STORE database (https://www.storedb.org/). Continuous build integration tools allowed our international RBO collaboration to be more efficient and focus its efforts on semantic model design. Currently, the RBO contains over 300 annotated classes and individuals specific to the study of radiation on biological systems, as well as imports of many additional classes from other OBO Foundry ontologies that relate to and/or provide context for these RBO entities. We publish the RBO through the OBO Foundry, so that it is available for browsing, download, and querying through NCBI Bioportal web site and application programming interface. The NASA Ames Life Science Data Archive (ALSDA) is also in the process of adopting use of the RBO, taking NASA one step closer to a knowledge-based system for space biology data. It is our hope that the global communities of radiation research Investigators, data curators and data analysts can similarly leverage the RBO and will contribute to its further development.

radiation↗

Jackson, L., Johnson, M.B., Latrach, A., Grimes, D., Martinez, C., and Mclaughlin, J.F., 2024, Multidisciplinary geotechnical data collection, curation, and analysis for conformity with the regulatory framework for geologic carbon storage in Wyoming, USA: Geological Society of America Abstracts with Programs. Vol. 56, No. 5, 2024, doi: 10.1130/abs/2024AM-405024

Title: Multidisciplinary Geotechnical Data Collection, Curation, and Analysis for Conformity with the Regulatory Framework for Geologic Carbon Storage in Wyoming, USA. Text: Construction and operation of wells for geologic sequestration of carbon dioxide necessitate that they are permitted under the Environmental Protection Agency’s Underground Injection Control Class VI requirements. Class VI wells conform to stringent requirements to ensure long-term safety and integrity of the storage site and the protection of Underground Sources of Drinking Water. Entities pursuing Class VI permitting must provide comprehensive geologic site characterization, including regional geologic structure and stratigraphy, aquifer information, reservoir and confining unit geomechanical properties, geochemical analyses, assessment of trapping capacity and mechanisms, and a variety of other of multidisciplinary geotechnical data. The Wyoming Class VI Site Characterization Database Project is focused on developing a geologic site characterization database of geotechnical information, which has been compiled and verified from established, public databases/entities and scientific literature to expedite Class VI permitting in Sweetwater County within the Greater Green River Basin of southern Wyoming. The preliminary suite of compiled data from 14,000 wells includes 8,000 wells with logs and 7,250 wells with formation tops, ~70 wells with core data (e.g., X-Ray diffraction, petrographic, and petrophysical data), ~2,500 water analyses, ~740 seismic events data, and ~520 bottom-hole temperature measurements. Future work on—and stemming from—this project will include new core analyses, calculation and interpolation of subsurface temperature gradients, mechanical earth models, geochemical simulations, storage capacity estimation, stratigraphic column generation and correlation, and construction of subsurface maps. Finally, this work will help to inspire and facilitate subsurface data compilation and curation beyond Sweetwater County, Wyoming.

42 ENGINEERING↗

Physics-informed graph neural networks for predicting cetane number with systematic data quality analysis

Designing alternative fuels for advanced compression ignition engines necessitates a predictive model for cetane number (CN). In this study, the physics-informed graph neural networks are introduced for a reliable CN prediction by considering molecular features pertinent to the physical properties of molecules that affect CN. The reliability of measured data is another key factor to consider for improving the predictive model. Various experimental instruments for measuring CN exist, including standard and non-standard methods. In this regard, a systematic data quality analysis was carried out for the total 630 CNs collected from literature and new measurements in this study using Advanced Fuel Ignition Delay Analyzer (AFIDA). The results from this data curation process were reflected in the model by imposing lower sample weights on the data coming from less reliable measurement techniques. This approach effectively maximized the prediction accuracy while incorporating data from all available sources. Using the sample weights decreased the mean absolute error (MAE) up to 0.8 CN units. The accuracy was also improved by introducing the CN-related physical properties (the number of hydrogen bond donors and acceptors); the test set MAE is 5.74 and 7.01 for the model with and without such properties, respectively. Investigating molecular structural effects on CN was also carried out to gain chemical insights into factors used to design new fuel candidates. The dimensionality reduction analysis of feature vectors showed a clear clustering in terms of functional groups and CN and the structural effect derived from the model was consistent with the physicochemical insights. Finally, this physics-informed model and data curation would be helpful for accurate CN prediction and inform rational fuel design.

97 MATHEMATICS AND COMPUTING↗

Improving the Quality of Geothermal Data Through Data Standards and Pipelines Within the Geothermal Data Repository: Preprint

For machine learning outputs to be applicable to real world problems, high quality data are needed to ensure high quality results. With the more recent emphasis on machine learning in geothermal, there is an increasing need for greater focus on the quality of the data available for use in these projects. For example, Geothermal Operational Optimization Using Machine Learning (GOOML) utilized large quantities of geothermal power plant operational data to inform power plant operational configurations to maximize power generation. High quality datasets result from dependable sensors or devices collecting data, high frequency of measurements, sufficient data points, adequate metadata, reliable storage of data, and sufficient data curation. Another component that contributes to high quality data is reusability, which can be enhanced through data standardization. Data Standardization creates consistency in formatting and contents of like datasets, lessening preprocessing requirements and ensuring adequate information provided by a given dataset. The Geothermal Data Repository (GDR) aims to help improve data quality through automated data standardization for high-value datasets through the implementation of data pipelines alongside reliable and accessible long-term storage for datasets. As such, the GDR has decided to shift away from recommending the use of Excel-based content models and towards the implementation of automated data pipelines. This takes the burden of data standardization off the user and project team and will increase the availability of standardized geothermal data available through the GDR. A set of recommendations, or a data standard for each data type will exist with each data pipeline in order to advise data collection for maximum usability for future research. This paper serves to describe the GDR's proposed transition towards data standardization through automated data pipelines, to discuss the need for and value of such a shift, and to call for suggestions from the community regarding the most useful data standards and pipelines.

data↗