Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Trust Not Verify? The Critical Need for Data Curation Standards in Materials Informatics

The importance of data curation has been recognized in multiple areas of research; however, the discussion of this important issue is only beginning to emerge in materials science. In this Perspective, we highlight the benefits of using the standardized data curation protocols in materials science and discuss current gaps in accurate and reproducible data reporting using case studies drawn from high-impact materials science papers and well-known databases such as the Crystallography Open Database (COD) and the Cambridge Structural Database (CSD). We argue that both experimental and computational materials scientists need to embrace a culture of rigorous data curation as part of modern research data management. We propose a sample data curation pipeline for materials chemistry and illustrate its use by creating two new materials chemistry databases. Here, we hope that this perspective will serve to catalyze further discussion and promote the continuous development of rigorous data curation practices within the materials science research community. We posit that adherence to best practices of data curation will promote and enhance the reliability, reproducibility, and integrity of materials research and enable the development of reliable AI and machine learning models that critically depend on the use of quality data.

Chemical structure↗

DOE COVID-19 Data Curation Effort: Overview of Initial Data Collection Coverage (March - June 2020)

During the COVID-19 pandemic of 2020, major case reporting outlets quickly coalesced around two or three primary vendors. Johns Hopkins University and The New York Times were among the more prominent, and all were of great value to the nation, particularly during the uncertain early stages of the pandemic. They primarily focused on three major attributes: number of new cases, deaths, and recovery, but only at the state level. Recognizing that many states were reporting very detailed data sets (e.g., hospital beds) at a count level or finer, the ORNL Pandemic Modeling team embarked on a major data curation effort from March to June 2020 for the purpose of capturing this wealth of detailed data. The challenge of curating this data was daunting. The number of attributes reported by the states grew on almost on a weekly basis. States were routinely shifting their web tool strategies away from easily parsable HTML-based formatting to new Tableau and ArcGIS content. This growth in the sheer number of attributes combined with the unpredictable shifts in data format meant an aggressive and agile combination of automated scripting and manual scraping was required to capture new daily streams. To keep up, the team had to scale up staff and widen its approach for capture and storage. The DOE COVID-19 data collection effort resulted in over 11 million data points being collected, covering over 13,000 unique geographies and over 2,000 unique attributes that spanned predominantly from early March through the end of June 2020.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Data Curation for Machine Learning Applied to Geothermal Power Plant Operational Data for GOOML: Geothermal Operational Optimization with Machine Learning: Preprint

Geothermal Operational Optimization with Machine Learning (GOOML) is a transferable and extensible component-based geothermal asset modeling framework that considers complex steamfield relationships and identifies optimization prospects using a data-driven approach to physics-guided, data-centric machine learning. This framework has been used to develop digital twins that provide steamfield operators with operational environments to analyze and understand historical and forecasted power production, explore new steamfield configuration possibilities, and seek optimal asset management in real world applications. To create, test, and apply the GOOML framework, diverse time-series datasets spanning multiple years were sourced from various geothermal power plant components within several complex real-world geothermal operations. These operations are based in the United States and New Zealand and include a variety of technologies, end-uses and configurations, collectively covering nearly all relevant operating conditions for modern geothermal fields. Datasets were acquired from multiple sources to ensure that machine learning experiments generalized properly to various operating conditions. It was found that the data varied in quality, format, and completeness. To ensure consistency between the various datasets, a standardized data curation process was developed to reliably streamline data preparation. This paper will discuss best practices as learned from the GOOML data curation process which takes the following steps: 1) acquisition of large quantities of data from power plant operators, 2) digestion of data to gain an initial understanding of what is included, 3) data transformation, which includes converting the data into a standardized machine-readable format so that they can be visualized, quality checked, and cleaned, 4) quality assurance and quality control, involving identification of significant data gaps and apparent anomalies through mapping of data features to real world componentry via the GOOML historical model, followed by discussion with modelers and power plant operators to identify additional data needs and to resolve issues, 5) use in machine learning algorithms, and 6) repetition of steps one through five until all data needs are met and data are deemed suitable for producing trustworthy modeling results which may be disseminated, ideally along with the curated dataset. This iterative process is focused on improving the quality of the data rather than tuning machine learning model parameters and supports a shift towards data-centric AI as a means to improving real-world applicability of geothermal machine learning projects.

access↗

Use of Semantic Technology to Create Curated Data Albums

One of the continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available online. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the data sets they need can obtain the specific files using these systems. However, in cases where researchers are interested in studying an event of research interest, they must manually assemble a variety of relevant data sets by searching the different distributed data systems. Consequently, there is a need to design and build specialized search and discover tools in Earth science that can filter through large volumes of distributed online data and information and only aggregate the relevant resources needed to support climatology and case studies. This paper presents a specialized search and discovery tool that automatically creates curated Data Albums. The tool was designed to enable key elements of the search process such as dynamic interaction and sense-making. The tool supports dynamic interaction via different modes of interactivity and visual presentation of information. The compilation of information and data into a Data Album is analogous to a shoebox within the sense-making framework. This tool automates most of the tedious information/data gathering tasks for researchers. Data curation by the tool is achieved via an ontology-based, relevancy ranking algorithm that filters out nonrelevant information and data. The curation enables better search results as compared to the simple keyword searches provided by existing data systems in Earth science.

Ramachandran, Rahul↗

Use of Semantic Technology to Create Curated Data Albums

One of the continuing challenges in any Earth science investigation is the discovery and access of useful science content from the increasingly large volumes of Earth science data and related information available online. Current Earth science data systems are designed with the assumption that researchers access data primarily by instrument or geophysical parameter. Those who know exactly the data sets they need can obtain the specific files using these systems. However, in cases where researchers are interested in studying an event of research interest, they must manually assemble a variety of relevant data sets by searching the different distributed data systems. Consequently, there is a need to design and build specialized search and discovery tools in Earth science that can filter through large volumes of distributed online data and information and only aggregate the relevant resources needed to support climatology and case studies. This paper presents a specialized search and discovery tool that automatically creates curated Data Albums. The tool was designed to enable key elements of the search process such as dynamic interaction and sense-making. The tool supports dynamic interaction via different modes of interactivity and visual presentation of information. The compilation of information and data into a Data Album is analogous to a shoebox within the sense-making framework. This tool automates most of the tedious information/data gathering tasks for researchers. Data curation by the tool is achieved via an ontology-based, relevancy ranking algorithm that filters out non-relevant information and data. The curation enables better search results as compared to the simple keyword searches provided by existing data systems in Earth science.

Ramachandran, Rahul↗

VA EDH Data Curation Documentation FY22-Q2, Rev. 2

The health and well-being of the Nation’s men and women who have served in uniform is the highest priority for the U.S. Department of Veterans Affairs (VA). VA is committed to providing timely access to high-quality, recovery-oriented, evidence-based mental health care that anticipates and responds to Veterans’ needs and supports the reintegration of returning Service members into their communities. VA is working to eliminate suicide among all Veterans by developing and implementing innovative suicide prevention approaches and resources. Health outcomes, such as suicide are typically modeled as a function of genetics and environment, where environment refers to factors beyond medical, e.g., air quality, access to transportation and food, homelessness status, etc. Mental health outcomes for each individual are considered to be associated with multiple stressors that fall under a variety of categories socioeconomic, economic, physical environment. Understanding the relationships between these stressors, covariates and health outcomes, requires curated, standardized data that can be input into the VA’s Recovery Engagement and Coordination for Health - Veterans Enhanced Treatment (REACH VET) or other health outcomes model. Environmental Determinants of Health (EDH) as defined by the World Health Organization (WHO) is clean air, stable climate, adequate water, sanitation and hygiene, safe use of chemicals, protection from radiation, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved nature are all prerequisites for good health.

60 APPLIED LIFE SCIENCES↗

VA EDH Data Curation Documentation (FY22-Q3, Rev.2)

The health and well-being of the Nation’s men and women who have served in uniform is the highest priority for the U.S. Department of Veterans Affairs (VA). VA is committed to providing timely access to high-quality, recovery-oriented, evidence-based mental health care that anticipates and responds to Veterans’ needs and supports the reintegration of returning Service members into their communities. VA is working to eliminate suicide among all Veterans by developing and implementing innovative suicide prevention approaches and resources. Health outcomes, such as suicide are typically modeled as a function of genetics and environment, where environment refers to factors beyond medical, e.g., air quality, access to transportation and food, homelessness status, etc. Mental health outcomes for each individual are considered to be associated with multiple stressors that fall under a variety of categories socioeconomic, economic, physical environment. Understanding the relationships between these stressors, covariates and health outcomes, requires curated, standardized data that can be input into the VA’s Recovery Engagement and Coordination for Health - Veterans Enhanced Treatment (REACH VET) or other health outcomes model. Environmental Determinants of Health (EDH) as defined by the World Health Organization (WHO) is clean air, stable climate, adequate water, sanitation and hygiene, safe use of chemicals, protection from radiation, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved nature are all prerequisites for good health.

59 BASIC BIOLOGICAL SCIENCES↗

VA EDH Data Curation Documentation FY22-Q4

The health and well-being of the Nation’s men and women who have served in uniform is the highest priority for the U.S. Department of Veterans Affairs (VA). VA is committed to providing timely access to high-quality, recovery-oriented, evidence-based mental health care that anticipates and responds to Veterans’ needs and supports the reintegration of returning Service members into their communities. VA is working to eliminate suicide among all Veterans by developing and implementing innovative suicide prevention approaches and resources. Health outcomes, such as suicide are typically modeled as a function of genetics and environment, where environment refers to factors beyond medical, e.g., air quality, access to transportation and food, homelessness status, etc. Mental health outcomes for each individual are considered to be associated with multiple stressors that fall under a variety of categories socioeconomic, economic, physical environment. Understanding the relationships between these stressors, covariates, and health outcomes, requires curated, standardized data that can be input into the VA’s Recovery Engagement and Coordination for Health - Veterans Enhanced Treatment (REACH VET) or other health outcomes model. Environmental Determinants of Health (EDH) as defined by the World Health Organization (WHO) is clean air, stable climate, adequate water, sanitation and hygiene, safe use of chemicals, protection from radiation, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved nature are all prerequisites for good health.

99 GENERAL AND MISCELLANEOUS↗

VA EDH Data Curation Documentation FY23-Q2

The health and well-being of the Nation’s men and women who have served in uniform is the highest priority for the U.S. Department of Veterans Affairs (VA). VA is committed to providing timely access to high-quality, recovery-oriented, evidence-based mental health care that anticipates and responds to Veterans’ needs and supports the reintegration of returning Service members into their communities. Since its creation, VA has been working to eliminate suicide among all veterans by developing and implementing innovative suicide prevention approaches and resources. Health outcomes, such as suicide are typically modeled as a function of genetics and environment, where environment refers to factors beyond medical, e.g., air quality, access to transportation and food, homelessness status, etc. Mental health outcomes for each individual are considered to be associated with multiple stressors that fall under a variety of categories including socioeconomic, economic, physical environment. Understanding the relationships between these stressors, covariates and health outcomes requires curated, standardized data that can be input into the VA’s Recovery Engagement and Coordination for Health, Veterans Enhanced Treatment (REACH VET) or other health outcomes model. Environmental Determinants of Health (EDH) as defined by the World Health Organization (WHO) is clean air, stable climate, adequate water, sanitation and hygiene, safe use of chemicals, protection from radiation, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved nature are all prerequisites for good health.

99 GENERAL AND MISCELLANEOUS↗

VA EDH Data Curation Documentation FY23-Q3

The health and well-being of the Nation’s men and women who have served in uniform is the highest priority for the U.S. Department of Veterans Affairs (VA). VA is committed to providing timely access to high-quality, recovery-oriented, evidence-based mental health care that anticipates and responds to Veterans’ needs and supports the reintegration of returning Service members into their communities. Since its creation, VA has been working to eliminate suicide among all veterans by developing and implementing innovative suicide prevention approaches and resources. Health outcomes, such as suicide are typically modeled as a function of genetics and environment, where environment refers to factors beyond medical, e.g., air quality, access to transportation and food, homelessness status, etc. Mental health outcomes for each individual are considered to be associated with multiple stressors that fall under a variety of categories including socioeconomic, economic, physical environment. Understanding the relationships between these stressors, covariates and health outcomes requires curated, standardized data that can be input into the VA’s Recovery Engagement and Coordination for Health, Veterans Enhanced Treatment (REACH VET) or other health outcomes model. Environmental Determinants of Health (EDH) as defined by the World Health Organization (WHO) refers to clean air, stable climate, adequate water, sanitation and hygiene, safe use of chemicals, protection from radiation, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved nature are all prerequisites for good health.

97 MATHEMATICS AND COMPUTING↗

VA EDH Data Curation Documentation FY23-Q4

The health and well-being of the Nation’s men and women who have served in uniform is the highest priority for the U.S. Department of Veterans Affairs (VA). VA is committed to providing timely access to high-quality, recovery-oriented, evidence-based mental health care that anticipates and responds to Veterans’ needs and supports the reintegration of returning service members into their communities. Since its creation, the VA has been working to eliminate suicide among all veterans by developing and implementing innovative suicide prevention approaches and resources. Health outcomes, such as suicide, are typically modeled as a function of genetics and environment, where environment refers to factors beyond medical, e.g., air quality, access to transportation and food, homelessness status, etc. Mental health outcomes for each individual are considered to be associated with multiple stressors that fall under a variety of categories, including socioeconomic, economic, and physical environments. Understanding the relationships between these stressors, covariates, and health outcomes requires curated, standardized data that can be input into the VA’s Recovery Engagement and Coordination for Health, Veterans Enhanced Treatment (REACH VET) or other health outcomes model. Environmental Determinants of Health (EDH), as defined by the World Health Organization (WHO), refer to clean air, stable climate, adequate water, sanitation and hygiene, safe use of chemicals, protection from radiation, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved nature, which are all prerequisites for good health.

60 APPLIED LIFE SCIENCES↗

Improving the Acquisition and Management of Sample Curation Data

This paper discusses the current sample documentation processes used during and after a mission, examines the challenges and special considerations needed for designing effective sample curation data systems, and looks at the results of a simulated sample result mission and the lessons learned from this simulation. In addition, it introduces a new data architecture for an integrated sample Curation data system being implemented at the NASA Astromaterials Acquisition and Curation department and discusses how it improves on existing data management systems.

Todd, Nancy S.↗

The disCO2ver Platform: Curating Data and Tools for Geologic Carbon Sequestration and Deep Subsurface Research Systems

The U.S. DOE National Energy Technology Laboratory has invested 12+ years of development into the data repository and digital laboratory, the Energy Data eXchange (EDX, edx.netl.doe.gov). Supporting a variety of research areas across the DOE Office of Fossil Energy and Carbon Management, the platform has successfully curated and preserved thousands of data products from DOE research. The Carbon Storage Program has successfully supported data curation, upload, and publishing of data products on EDX for many years, demonstrating a success story of how resources like EDX can effectively help with long term preservation and publishing of DOE data products. EDX continues to shift towards cloud-supported infrastructure, taking a hybrid approach combining on-premises compute and storage integrated with cloud-hosted services. The integration of cloud compute and hybrid architecture enables the development of EDX-hosted platforms that tailor the data and tools hosted on them to a specific community, enables implementation of machine learning tools for data discovery and filtering, and enables the hosting of virtual (online user interface) tools. Geologic carbon sequestration (GCS) research continues to scale up in response to the current administration goals to reduce greenhouse gas emissions and transition the energy economy. Over the last year, EDX’s disCO2ver platform has been developed in response to the need for access to data products and tools to support the scaling up of GCS research. disCO2ver provides access to data resources and tools, produced by DOE and outside authoritative external resources. The platform also provides a user-access control component for the virtualization and cloud hosting of tools. Tools that need to be virtualized, to eliminate the need for users to download the tool and use local compute resources, is essential to supporting big-data analysis and machine learning that is becoming common place in carbon storage modeling, risk analysis, and data publishing practices. This talk will review the EDX’s disCO2ver platform and the current work ongoing to curate data and tools to support GCS and deep subsurface systems research.

Morkner, Paige↗

Virtual Collections: An Earth Science Data Curation Service

The role of Earth science data centers has traditionally been to maintain central archives that serve openly available Earth observation data. However, in order to ensure data are as useful as possible to a diverse user community, Earth science data centers must move beyond simply serving as an archive to offering innovative data services to user communities. A virtual collection, the end product of a curation activity that searches, selects, and synthesizes diffuse data and information resources around a specific topic or event, is a data curation service that improves the discoverability, accessibility, and usability of Earth science data and also supports the needs of unanticipated users. Virtual collections minimize the amount of the time and effort needed to begin research by maximizing certainty of reward and by providing a trustworthy source of data for unanticipated users. This presentation will define a virtual collection in the context of an Earth science data center and will highlight a virtual collection case study created at the Global Hydrology Resource Center data center.

Data↗

Collaborative Data Curation to Support the Multi-Mission Algorithm and Analysis Platform (MAAP)

Upcoming space-borne missions will offer unprecedented data about Earth but will also feature exponentially high data volumes. These high data volumes will change the way the scientific community works with data and will also create a unique need for improved data sharing and collaboration. NASA and ESA are working together to address these issues by collaboratively developing the Multi-Mission Algorithm and Analysis Platform (MAAP) to improve the understanding of global aboveground terrestrial carbon dynamics. The MAAP will support ESA’s BIOMASS mission, NASA’s GEDI mission and NASA/ISRO’s NISAR mission. The MAAP will be developed in two phases: a pilot phase and a full production phase. The pilot phase will demonstrate collaboration and basic capabilities. The pilot phase will focus on biomass relevant airborne and field campaign data. Two NASA teams are supporting the development of the MAAP. The MAAP engineering team is responsible for the development, maintenance and operations of the MAAP system while the MAAP data team ensures the ongoing quality of the data, metadata and other information provided in the MAAP. The MAAP data team also supports the ingest and archive of identified data to the MAAP platform. This poster describes the use case development process for the pilot MAAP and the data curated in support of those use cases. Additionally, this presentation will outline the pilot MAAP data ingest process and metadata curation effort along with efforts to ensure interoperability between ESA and NASA data and metadata.

Bugbee, Kaylin↗

EXFOR-NSR PDF database: a system for nuclear knowledge preservation and data curation

Current needs of nuclear science and technology include complete, well-documented, and easily verifiable nuclear data. The complete data records require supporting nuclear bibliography, presently stored in dedicated libraries, in addition, to actual data. Additionally, experimental nuclear reaction data (EXFOR) and Nuclear Science References (NSR) databases contain compilations based on primary (journals) and secondary (conference proceedings, theses, preprints, etc.) publications, and data received from authors via private communications. The secondary library materials and private communications often represent a bottleneck for nuclear data verification, compilation, evaluation, and dissemination activities. To address this issue, bibliographic materials were scanned into PDF (Portable Document Format) files and uploaded in a relational database. The traditional scope of nuclear databases that includes meta-data and numbers derived from data in specialized formats was broadened to accommodate the large volumes of original nuclear data publications. The complete PDF publication files were stored in a relational database as Binary Large OBjects (BLOB). This unique collection of nuclear data compilations and supporting publications generate many opportunities for machine learning applications. The Web interfaces for authorized and public access to the EXFOR-NSR nuclear publications database were implemented at the U.S. National Nuclear Data Center, https://www.nndc.bnl.gov/ and IAEA Nuclear Data Section, https://www-nds.iaea.org/ . The current system is complementary to major nuclear libraries and narrowly focused on nuclear data compilation and evaluation procedures. The contents of the PDF database, details of implementation, and Web interface are described. New capabilities for data curation, knowledge preservation, worldwide dissemination, and natural language processing (NLP) applications are given.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗