Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Improving the Discoverability and Availability of Sample Data and Imagery in NASA's Astromaterials Curation Digital Repository Using a New Common Architecture for Sample Databases

The Astromaterials Acquisition and Curation Office at NASA's Johnson Space Center (JSC) is the designated facility for curating all of NASA's extraterrestrial samples. The suite of collections includes the lunar samples from the Apollo missions, cosmic dust particles falling into the Earth's atmosphere, meteorites collected in Antarctica, comet and interstellar dust particles from the Stardust mission, asteroid particles from the Japanese Hayabusa mission, and solar wind atoms collected during the Genesis mission. To support planetary science research on these samples, NASA's Astromaterials Curation Office hosts the Astromaterials Curation Digital Repository, which provides descriptions of the missions and collections, and critical information about each individual sample. Our office is implementing several informatics initiatives with the goal of better serving the planetary research community. One of these initiatives aims to increase the availability and discoverability of sample data and images through the use of a newly designed common architecture for Astromaterials Curation databases.

Todd, N. S.↗

Science Validation for Dark Energy Research with Optical Imaging Surveys

The Universe has been expanding at an accelerating rate over the past several billion years, as though an unknown form of dark energy permeates all of space. The ultimate scientific goal of the proposed research is to distinguish between different physical mechanisms that could account for this observed accelerated expansion, for example, the zero-point energy of the vacuum, a dynamical form of energy that varies in time and/or space, or a modification to our theory of gravity. Wide- field optical imaging surveys of the night sky can test these competing models by measuring both the cosmic expansion history and the growth of large-scale structure. To this end, the Dark Energy Survey (DES) has cataloged several hundred million galaxies and thousands of supernovae. The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST), will enlarge the census to billions of galaxies and hundreds of thousands of supernovae. A critical question for these dark energy experiments is whether systematic uncertainties can continue to be controlled at a level to keep pace with the statistical precision offered by such enormous datasets. The immediate research objectives of this project were (1) to prepare and validate input datasets that are the foundation of cosmological analyses with DES, and (2) to prepare for value-added characterization of Rubin Observatory commissioning data to inform early operations and accelerate the realization of dark energy science from LSST data products. For DES, we assembled and curated cosmology-ready data releases that include value-added components such as enhanced photometric and astrometric calibrations, alternative source extraction algorithms, maps of the survey coverage and survey conditions, object classifications, object quality selections, galaxy shapes, and photometric redshifts. We used the galaxy clustering technique to validate the photometric redshift distributions of various galaxy samples to be used as lenses in combined studies of galaxy clustering and weak gravitational lensing. The galaxy clustering redshift analysis was enhanced by use of a larger sample of reference galaxies from the eBOSS spectroscopic survey that extends to higher redshifts. For LSST, we prepared for science validation studies of commissioning data aimed at dark energy science capability that extend beyond the normative system-level tests to be done by the Rubin Observatory Construction Project. We identified a set of proposed survey strategies and candidate target fields that could be observed during the commissioning period to enhance science validation activities related to studies of dark energy.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

PDBx/mmCIF Ecosystem: Foundational Semantic Tools for Structural Biology

PDBx/mmCIF, Protein Data Bank Exchange (PDBx) macromolecular Crystallographic Information Framework (mmCIF), has become the data standard for structural biology. With its early roots in the domain of small-molecule crystallography, PDBx/mmCIF provides an extensible data representation that is used for deposition, archiving, remediation, and public dissemination of experimentally determined three-dimensional (3D) structures of biological macromolecules by the Worldwide Protein Data Bank (wwPDB, wwpdb.org). Extensions of PDBx/mmCIF are similarly used for computed structure models by ModelArchive (modelarchive.org), integrative/hybrid structures by PDB-Dev (pdb-dev.wwpdb.org), small angle scattering data by Small Angle Scattering Biological Data Bank SASBDB (sasbdb.org), and for models computed generated with the AlphaFold 2.0 deep learning software suite (alphafold.ebi.ac.uk). Community-driven development of PDBx/mmCIF spans three decades, involving contributions from researchers, software and methods developers in structural sciences, data repository providers, scientific publishers, and professional societies. Having a semantically rich and extensible data framework for representing a wide range of structural biology experimental and computational results, combined with expertly curated 3D biostructure data sets in public repositories, accelerates the pace of scientific discovery. Herein, we describe the architecture of the PDBx/mmCIF data standard, tools used to maintain representations of the data standard, governance, and processes by which data content standards are extended, plus community tools/software libraries available for processing and checking the integrity of PDBx/mmCIF data. Use cases exemplify how the members of the Worldwide Protein Data Bank have used PDBx/mmCIF as the foundation for its pipeline for delivering Findable, Accessible, Interoperable, and Reusable (FAIR) data to many millions of users worldwide.

59 BASIC BIOLOGICAL SCIENCES↗

CoRE MOF DB: A curated experimental metal-organic framework database with machine-learned properties for integrated material-process screening

Here, we present an updated version of the Computation-Ready, Experimental (CoRE) Metal-Organic Framework (MOF) database, which includes a curated set of computation-ready MOF crystal structures designed for high-throughput computational materials discovery. Data collection and curation procedures were improved from the previous version to enable more frequent updates in the future. Machine-learning-predicted properties, such as stability metrics and heat capacities, are included in the dataset to streamline screening activities. An updated version of MOFid was developed to provide detailed information on metal nodes, organic linkers, and topologies of an MOF structure. DDEC6 partial atomic charges of MOFs were assigned based on a machine-learning model. Gibbs ensemble Monte Carlo simulations were used to classify the hydrophobicity of MOFs. The finalized dataset was subsequently used to perform integrated material-process screening for various carbon-capture conditions using high-fidelity temperature-swing adsorption (TSA) simulations. Our workflow identified multiple MOF candidates that are predicted to outperform CALF-20 for these applications.

CoRE MOF database↗

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES↗

Investigating the Future of Scientific Data Search [Slides]

Searching for usable, actionable, data in a trustworthy manner is a challenge across scientific communities. Artificial Intelligence (AI) and Machine Learning (ML) techniques may be leveraged to increase the utility of scientific data by: Demystify unstructured data to aid curation & sharing Surfacing hard to find datasets. User Experience (UX) Research can help uncover scientists needs & challenges finding data and using AI/ML enabled tools.

97 MATHEMATICS AND COMPUTING↗

Computational tools and data integration to accelerate vaccine development: challenges, opportunities, and future directions

The development of effective vaccines is crucial for combating current and emerging pathogens. Despite significant advances in the field of vaccine development there remain numerous challenges including the lack of standardized data reporting and curation practices, making it difficult to determine correlates of protection from experimental and clinical studies. Significant gaps in data and knowledge integration can hinder vaccine development which relies on a comprehensive understanding of the interplay between pathogens and the host immune system. In this review, we explore the current landscape of vaccine development, highlighting the computational challenges, limitations, and opportunities associated with integrating diverse data types for leveraging artificial intelligence (AI) and machine learning (ML) techniques in vaccine design. We discuss the role of natural language processing, semantic integration, and causal inference in extracting valuable insights from published literature and unstructured data sources, as well as the computational modeling of immune responses. Furthermore, we highlight specific challenges associated with uncertainty quantification in vaccine development and emphasize the importance of establishing standardized data formats and ontologies to facilitate the integration and analysis of heterogeneous data. Through data harmonization and integration, the development of safe and effective vaccines can be accelerated to improve public health outcomes. Looking to the future, we highlight the need for collaborative efforts among researchers, data scientists, and public health experts to realize the full potential of AI-assisted vaccine design and streamline the vaccine development process.

60 APPLIED LIFE SCIENCES↗

Apollo Lunar Sample Photographs: Digitizing the Moon Rock Collection

The Acquisition and Curation Office at JSC has undertaken a 4-year data restoration project effort for the lunar science community funded by the LASER program (Lunar Advanced Science and Exploration Research) to digitize photographs of the Apollo lunar rock samples and create high resolution digital images. These sample photographs are not easily accessible outside of JSC, and currently exist only on degradable film in the Curation Data Storage Facility

Lofgren, Gary E.↗

Automating data analysis for hydrogen/deuterium exchange mass spectrometry using data-independent acquisition methodology

We present a hydrogen/deuterium exchange workflow coupled to tandem mass spectrometry (HX-MS 2 ) that supports the acquisition of peptide fragment ions alongside their peptide precursors. The approach enables true auto-curation of HX data by mining a rich set of deuterated fragments, generated by collisional-induced dissociation (CID), to simultaneously confirm the peptide ID and authenticate MS 1 -based deuteration calculations. The high redundancy provided by the fragments supports a confidence assessment of deuterium calculations using a combinatorial strategy. The approach requires data-independent acquisition (DIA) methods that are available on most MS platforms, making the switch to HX-MS 2 straightforward. Importantly, we find that HX-DIA enables a proteomics-grade approach and wide-spread applications. Considerable time is saved through auto-curation and complex samples can now be characterized and at higher throughput. We illustrate these advantages in a drug binding analysis of the ultra-large protein kinase DNA-PKcs, isolated directly from mammalian cells.

59 BASIC BIOLOGICAL SCIENCES↗

Curating NASA's Past, Present, and Future Astromaterial Sample Collections

The Astromaterials Acquisition and Curation Office at NASA Johnson Space Center (hereafter JSC curation) is responsible for curating all of NASA's extraterrestrial samples. JSC presently curates 9 different astromaterials collections in seven different clean-room suites: (1) Apollo Samples (ISO (International Standards Organization) class 6 + 7); (2) Antarctic Meteorites (ISO 6 + 7); (3) Cosmic Dust Particles (ISO 5); (4) Microparticle Impact Collection (ISO 7; formerly called Space-Exposed Hardware); (5) Genesis Solar Wind Atoms (ISO 4); (6) Stardust Comet Particles (ISO 5); (7) Stardust Interstellar Particles (ISO 5); (8) Hayabusa Asteroid Particles (ISO 5); (9) OSIRIS-REx Spacecraft Coupons and Witness Plates (ISO 7). Additional cleanrooms are currently being planned to house samples from two new collections, Hayabusa 2 (2021) and OSIRIS-REx (2023). In addition to the labs that house the samples, we maintain a wide variety of infra-structure facilities required to support the clean rooms: HEPA-filtered air-handling systems, ultrapure dry gaseous nitrogen systems, an ultrapure water system, and cleaning facilities to provide clean tools and equipment for the labs. We also have sample preparation facilities for making thin sections, microtome sections, and even focused ion-beam sections. We routinely monitor the cleanliness of our clean rooms and infrastructure systems, including measurements of inorganic or organic contamination, weekly airborne particle counts, compositional and isotopic monitoring of liquid N2 deliveries, and daily UPW system monitoring. In addition to the physical maintenance of the samples, we track within our databases the current and ever changing characteristics (weight, location, etc.) of more than 250,000 individually numbered samples across our various collections, as well as more than 100,000 images, and countless "analog" records that record the sample processing records of each individual sample. JSC Curation is co-located with JSC's Astromaterials Research Office, which houses a world-class suite of analytical instrumentation and scientists. We leverage these labs and personnel to better curate the samples. Part of the cu-ration process is planning for the future, and we refer to these planning efforts as "advanced curation". Advanced Curation is tasked with developing procedures, technology, and data sets necessary for curating new types of collections as envi-sioned by NASA exploration goals. We are (and have been) planning for future cu-ration, including cold curation, extended curation of ices and volatiles, curation of samples with special chemical considerations such as perchlorate-rich samples, and curation of organically- and biologically-sensitive samples.

Zeigler, R. A.↗

The Astromaterials X-Ray Computed Tomography Laboratory at Johnson Space Center

The Astromaterials Acquisition and Curation Office at NASA's Johnson Space Center (hereafter JSC curation) is the past, present, and future home of all of NASA's astromaterials sample collections. JSC curation currently houses all or part of nine different sample collections: (1) Apollo samples (1969), (2) Lunar samples (1972), (3) Antarctic meteorites (1976), (4) Cosmic Dust particles (1981), (5) Microparticle Impact Collection (1985), (6) Genesis solar wind atoms (2004); (7) Stardust comet Wild-2 particles (2006), (8) Stardust interstellar particles (2006), and (9) Hayabusa asteroid Itokawa particles (2010). Each sample collection is housed in a dedicated clean room, or suite of clean rooms, that is tailored to the requirements of that sample collection. Our primary goals are to maintain the long-term integrity of the samples and ensure that the samples are distributed for scientific study in a fair, timely, and responsible manner, thus maximizing the return on each sample. Part of the curation process is planning for the future, and we also perform fundamental research in advanced curation initiatives. Advanced Curation is tasked with developing procedures, technology, and data sets necessary for curating new types of sample collections, or getting new results from existing sample collections [2]. We are (and have been) planning for future curation, including cold curation, extended curation of ices and volatiles, curation of samples with special chemical considerations such as perchlorate-rich samples, and curation of organically- and biologically-sensitive samples. As part of these advanced curation efforts we are augmenting our analytical facilities as well. A micro X-Ray computed tomography (micro-XCT) laboratory dedicated to the study of astromaterials will be coming online this spring within the JSC Curation office, and we plan to add additional facilities that will enable nondestructive (or minimally-destructive) analyses of astromaterials in the near future (micro-XRF, confocal imaging Raman Spectroscopy). These facilities will be available to: (1) develop sample handling and storage techniques for future sample return missions; (2) be utilized by PET for future sample return missions; (3) be used for retroactive PET (Positron Emission Tomography)-style analyses of our existing collections; and (4) for periodic assessments of the existing sample collections. Here we describe the new micro-XCT system, as well as some of the ongoing or anticipated applications of the instrument.

Zeigler, R. A.↗

Priority Science Targets for Future Sample Return Missions within the Solar System Out to the Year 2050

The Astromaterials Acquisition and Curation Office (henceforth referred to herein as NASA Curation Office) at NASA Johnson Space Center (JSC) is responsible for curating all of NASA's extraterrestrial samples. JSC presently curates 9 different astromaterials collections: (1) Apollo samples, (2) LUNA samples, (3) Antarctic meteorites, (4) Cosmic dust particles, (5) Microparticle Impact Collection [formerly called Space Exposed Hardware], (6) Genesis solar wind, (7) Star-dust comet Wild-2 particles, (8) Stardust interstellar particles, and (9) Hayabusa asteroid Itokawa particles. In addition, the next missions bringing carbonaceous asteroid samples to JSC are Hayabusa 2/ asteroid Ryugu and OSIRIS-Rex/ asteroid Bennu, in 2021 and 2023, respectively. The Hayabusa 2 samples are provided as part of an international agreement with JAXA. The NASA Curation Office plans for the requirements of future collections in an "Advanced Curation" program. Advanced Curation is tasked with developing procedures, technology, and data sets necessary for curating new types of collections as envisioned by NASA exploration goals. Here we review the science value and sample curation needs of some potential targets for sample return missions over the next 35 years.

McCubbin, F. M.↗

The Astromaterials X-Ray Computed Tomography Laboratory at Johnson Space Center

The Astromaterials Acquisition and Curation Office at NASA's Johnson Space Center (hereafter JSC curation) is the past, present, and future home of all of NASA's astromaterials sample collections. JSC curation currently houses all or part of nine different sample collections: (1) Apollo samples (1969), (2) Lunar samples (1972), (3) Antarctic meteorites (1976), (4) Cosmic Dust particles (1981), (5) Microparticle Impact Collection (1985), (6) Genesis solar wind atoms (2004); (7) Stardust comet Wild-2 particles (2006), (8) Stardust interstellar particles (2006), and (9) Hayabusa asteroid Itokawa particles (2010). Each sample collection is housed in a dedicated clean room, or suite of clean rooms, that is tailored to the requirements of that sample collection. Our primary goals are to maintain the long-term integrity of the samples and ensure that the samples are distributed for scientific study in a fair, timely, and responsible manner, thus maximizing the return on each sample. Part of the curation process is planning for the future, and we also perform fundamental research in advanced curation initiatives. Advanced Curation is tasked with developing procedures, technology, and data sets necessary for curating new types of sample collections, or getting new results from existing sample collections [2]. We are (and have been) planning for future curation, including cold curation, extended curation of ices and volatiles, curation of samples with special chemical considerations such as perchlorate-rich samples, and curation of organically- and biologically-sensitive samples. As part of these advanced curation efforts we are augmenting our analytical facilities as well. A micro X-Ray computed tomography (micro-XCT) laboratory dedicated to the study of astromaterials will be coming online this spring within the JSC Curation office, and we plan to add additional facilities that will enable nondestructive (or minimally-destructive) analyses of astromaterials in the near future (micro-XRF, confocal imaging Raman Spectroscopy). These facilities will be available to: (1) develop sample handling and storage techniques for future sample return missions; (2) be utilized by PET for future sample return missions; (3) be used for retroactive PET (Positron Emission Tomography)-style analyses of our existing collections; and (4) for periodic assessments of the existing sample collections. Here we describe the new micro-XCT system, as well as some of the ongoing or anticipated applications of the instrument.

Zeigler, R. A.↗

Priority Science Targets for Future Sample Return Missions Within the Inner Solar System Out to the Year 2061

The Astromaterials Acquisition and Curation Office at NASA Johnson Space Center (JSC) is re-sponsible for curating all of NASA's extraterrestrial samples. The NASA Curation Office plans for the requirements of future collections in an ''Advanced Curation'' program. Advanced Curation is tasked with developing procedures, technology, and data sets necessary for curating new types of collections as envisioned by NASA exploration goals. Here we review the science value of some potential targets for sample return missions from the inner solar system over the next 43 years.

McCubbin, F. M.↗

The Astromaterials X-Ray Computed Tomography Laboratory at Johnson Space Center

The Astromaterials Acquisition and Cura-tion Office at NASA's Johnson Space Center (hereafter JSC curation) is the past, present, and future home of all of NASA's astromaterials sample collections. JSC curation currently houses all or part of nine different sample collections. Our primary goals are to maintain the long-term integrity of the samples and ensure that the samples are distributed for scientific study in a fair, timely, and responsible manner, thus maximizing the return on each sample. Part of the curation process is planning for the future, thus we also perform funda-mental research in advanced curation initiatives. Ad-vanced Curation is tasked with developing procedures, technology, and data sets necessary for curating new types of sample collections, or getting new results from existing sample collections [1]. As part of these ad-vanced curation efforts we are augmenting our analyti-cal facilities.

Zeigler, R. A.↗

GeneLab for High Schools – Bioinformatic Training For Students And Educators

Modern biological sciences are increasingly based on high-throughput molecular techniques, including genomics, transcriptomics, and proteomics. NASA’s GeneLab program has collected extensive data from ‘omics’ studies, curated them into an accessible platform and provided data analysis/visualization tools to facilitate the generation of new hypotheses and research directions. GeneLab for High Schools (GL4HS), launched in 2017, has endeavored to utilize this database and provide tools for students to understand and analyze omics datasets whilst also learning about spaceflight research. The GL4HS program ran in person at Ames from 2017-2019 and has run virtually since 2020. Each year fifteen high school students are trained to analyze and interpret GeneLab transcriptomic data. Additionally, in the last several years we have expanded our “teacher training program” to include 10 teachers total in an effort to enable this program to be utilized in classrooms across the USA. Teachers also join the NASA GeneLab Education Working Group (EWG) enabling support as they implement custom GL4HS modules into their classrooms. The GL4HS program consists of three main components – (1) core learning modules, (2) networking and teamwork, and (3) an independent learning project. Students are also taught critical networking and science communication skills facilitating their ability to ‘sell their science’ in innovative and creative ways. This program has enabled students to learn about biology in space and to have a glimpse into the world of research for the first time. Many of the students in this program shared that the course was transformative to their perception about biological sciences and how it linked to other areas of STEM. The ultimate and long-term goal of GL4HS is to expand the program to multiple locations thereby facilitating the reach of NASA Space Biology beyond NASA-centric regions.

GeneLab↗

Multi-Species Complex and Standard Metabolomic Samples with Verified Truth Annotations Dataset

This dataset contains 4523251 (~6.35 GB) metabolite-spectra matches following identification with CoreMS. Data were manually curated as true positives, true negatives, or unknowns. Calculations for spectral similarity scores were carried out with two methods for a total of ~12.7 GB (6.35 * 2) of data. They are all .tsv files, though can easily be changed to .txt. The file types are: * human cerebrospinal fluid (CSF), human blood plasma human urine: already published here https://www.nature.com/articles/s41597-021-00894-y, • purchased FAMES standards • fungi species (A. niger, A. nidulans, T. reesei) • soil crust

59 BASIC BIOLOGICAL SCIENCES↗