Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Open-Source Science-led Development of the Atmosphere Observing System (AOS) Mission Science Data System (SDS)

The Earth System Observatory (ESO) Atmosphere Observing System (AOS) mission will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The AOS Science Data System (SDS) will be a system of systems developed within the Cloud to manage the research and operational processing of AOS mission orbital and suborbital sensors and curate these data for reprocessing (e.g., in near real-time or by collection) and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage. Further, AOS SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The AOS mission follows NASA’s lead in making a commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the AOS SDS system components will be developed with open-source concepts including components of SDS itself as well as AOS mission algorithms. Further, the AOS SDS assumes the role to lead and facilitate OSS activities for the AOS mission. This presentation describes the framework of the AOS SDS and its integral part in facilitating OSS within the AOS mission.

David Giles↗

Open-Source Science-led Development of the AOS Mission Science Data System (SDS)

The Earth System Observatory (ESO) Atmosphere Observing System (AOS) mission will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The AOS Science Data System (SDS) will be a system of systems developed within the Cloud to manage the research and operational processing of AOS mission orbital and suborbital sensors and curate these data for reprocessing (e.g., in near real-time or by collection) and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage. Further, AOS SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The AOS mission follows NASA’s lead in making a commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the AOS SDS system components will be developed with open-source concepts including components of SDS itself as well as AOS mission algorithms. Further, the AOS SDS assumes the role to lead and facilitate OSS activities for the AOS mission. This presentation describes the framework of the AOS SDS and its integral part in facilitating OSS within the AOS mission.

David M. Giles↗

Beyond Fair: Engagement, Data Usability, and Open Community Productivity through the NASA Open Science Data Repository

The FAIR principle (findable, accessible, interoperable, and reusable) governs the storage and sharing of NASA space biology and health data[1]. These guiding principles maximize reuse of data and the reproducibility of scientific findings. The NASA Open Science Data Repository (OSDR; an expansion of NASA GeneLab) was built on the FAIR principles and houses over 500 studies and close to 1000 datasets from decades of space life sciences experiments. OSDR embodies the FAIR principles through data governance that includes mediated, embargoed, and fully open access data. The FAIR data governance principles were recently proposed to be expanded to encompass a FAIREST framework for assessing research data repositories (FAIR + Engagement, Social connections, and Trust)[2]. FAIREST emphasizes the importance of data repositories engaging with the scientific community and gaining the trust of researchers regarding data quality. Trust also refers to the TRUST principles developed for assessment of digital repositories: Transparency, Responsibility, User Focus, Sustainability, Technology[3]. We present the “Open Science for Life in Space” Analysis Working Groups (AWGs) as evidence regarding the power of engagement, social connections, and trust which has enhanced OSDR’s capabilities and productivity. AWG members engage in two main activities. One, members provide feedback on OSDR scientific standards for data ingestion, curation, and reuse (study, subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability). Two, AWG members collaborate to mine-reuse OSDR data to conduct scientific analysis. With nearly 800 active members, the AWGs have resulted in 32 publications re-using OSDR data and contributed many papers in two major special issues in Cell (2020) and Nature (2024). AWGs also serve as networking groups, facilitate social connections between researchers at all levels of experience, and also have a social online ‘Forum’ used to keep members informed on projects and opportunities. This community-centric, productive, and trustworthy data culture has resulted in a broader effect with international space agencies, academics, and the commercial space sector wanting to submit their data to OSDR. Ten studies of Inspiration 4 data were recently publicly released by OSDR, as were some JAXA human data. Coming up soon in OSDR are data submissions from the European Space Agency, Virgin Galactic PIs, and SpaceX Polaris Dawn. A major benefit of OSDR is the array of standardized and uniformly formatted data (which was developed through AWG member consensus), from which visualization tools, analysis tools, and machine learning models can be built or trained. This talk will cover the Multi-Study Visualization Tool, the Environmental Data Application, RadLab, and a UCSF-NSF funded knowledge graph biomedical health discovery tool ‘SPOKE’ currently being integrated with OSDR. OSDR also provides training programs in bioinformatics and machine learning to improve the scientific community’s awareness of data availability and to boost their ability to perform data analysis. The increasing engagement of the scientific community and the public with technologies powered by artificial intelligence (AI) heightens the need for data analysis to be transparent. The AI for Life in Space initiative leverages the data products provided in OSDR to train AI models, with an emphasis on explainable and trustworthy AI, which would not be possible without FAIR data and metadata. Overall, here we will demonstrate the importance for NASA life sciences data repositories to adhere to the FAIREST framework, by providing examples and success stories from different aspects of OSDR.

data↗

Updates to the Alliance of Genome Resources central infrastructure

The Alliance of Genome Resources (Alliance) is an extensible coalition of knowledgebases focused on the genetics and genomics of intensively studied model organisms. The Alliance is organized as individual knowledge centers with strong connections to their research communities and a centralized software infrastructure, discussed here. Model organisms currently represented in the Alliance are budding yeast, Caenorhabditis elegans, Drosophila, zebrafish, frog, laboratory mouse, laboratory rat, and the Gene Ontology Consortium. The project is in a rapid development phase to harmonize knowledge, store it, analyze it, and present it to the community through a web portal, direct downloads, and application programming interfaces (APIs). Here, we focus on developments over the last 2 years. Specifically, we added and enhanced tools for browsing the genome (JBrowse), downloading sequences, mining complex data (AllianceMine), visualizing pathways, full-text searching of the literature (Textpresso), and sequence similarity searching (SequenceServer). We enhanced existing interactive data tables and added an interactive table of paralogs to complement our representation of orthology. To support individual model organism communities, we implemented species-specific “landing pages” and will add disease-specific portals soon; in addition, we support a common community forum implemented in Discourse software. We describe our progress toward a central persistent database to support curation, the data modeling that underpins harmonization, and progress toward a state-of-the-art literature curation system with integrated artificial intelligence and machine learning (AI/ML).

59 BASIC BIOLOGICAL SCIENCES↗

Organic Contamination Baseline Study: In NASA JSC Astromaterials Curation Laboratories. Summary Report

In preparation for OSIRIS-REx and other future sample return missions concerned with analyzing organics, we conducted an Organic Contamination Baseline Study for JSC Curation Labsoratories in FY12. For FY12 testing, organic baseline study focused only on molecular organic contamination in JSC curation gloveboxes: presumably future collections (i.e. Lunar, Mars, asteroid missions) would use isolation containment systems over only cleanrooms for primary sample storage. This decision was made due to limit historical data on curation gloveboxes, limited IR&D funds and Genesis routinely monitors organics in their ISO class 4 cleanrooms.

Calaway, Michael J.↗

Open-Source Science-Driven Development of the Science Data System (SDS) for Earth System Observatory (ESO) Atmospheric Missions

The NASA Earth System Observatory (ESO) atmospheric missions will provide space-based and suborbital observations of collocated cloud, dynamic, precipitation and aerosol processing leading to improved weather, air quality, and climate predictions. The Science Data System (SDS) will deploy the adaptive processing system (APS) developed within the Cloud to manage the research and operational processing of ESO atmospheric mission orbital and suborbital sensors and curate these data for near real-time and collection reprocessing and transfer them to a NASA Distributed Active Archive Center (DAAC) for long-term storage and distribution. Further, the SDS will follow guidelines provided by NASA Earth Science Data Systems (ESDS) program including standard conventions for data file formats, naming, and metadata to improve data interoperability, interpretability, usability, discovery, provenance, and spatiotemporal representativeness. The SDS follows NASA’s commitment to Open-Source Science (OSS) including the sharing of data, software, and knowledge in an open and timely manner. Each of the SDS system components will be developed with open-source concepts including components of APS itself as well as ESO atmospheric mission algorithms. This presentation describes the framework of the SDS and its integral part in facilitating OSS within the ESO atmospheric missions.

David M. Giles↗

A Comprehensive Northern Hemisphere Particle Microphysics Data Set From the Precipitation Imaging Package

Microphysical observations of precipitating particles are critical data sources for numerical weather prediction models and remote sensing retrieval algorithms. However, obtaining coherent data sets of particle microphysics is challenging as they are often unindexed, distributed across disparate institutions, and have not undergone a uniform quality control process. This work introduces a unified, comprehensive Northern Hemisphere particle microphysical data set from the National Aeronautics and Space Administration precipitation imaging package (PIP), accessible in a standardized data format and stored in a centralized, public repository. Data is collected from 10 measurement sites spanning 34° latitude (37°N–71°N) over 10 years (2014–2023), which comprise a set of 1,070,000 precipitating minutes. The provided data set includes measurements of a suite of microphysical attributes for both rain and snow, including distributions of particle size, vertical velocity, and effective density, along with higher-order products including an approximation of volume-weighted equivalent particle densities, liquid equivalent snowfall, and rainfall rate estimates. The data underwent a rigorous standardization and quality assurance process to filter out erroneous observations to produce a self-describing, scalable, and achievable data set. Case study analyses demonstrate the capabilities of the data set in identifying physical processes like precipitation phase-changes at high temporal resolution. Bulk precipitation characteristics from a multi-site intercomparison also highlight distinct microphysical properties unique to each location. This curated PIP data set is a robust database of high-quality particle microphysical observations for constraining future precipitation retrieval algorithms, and offers new insights toward better understanding regional and seasonal differences in bulk precipitation characteristics.

54 ENVIRONMENTAL SCIENCES↗

Lowering the barrier to access information-rich transient kinetic data for machine learning methods

Transient kinetic data contain a wealth of information about intrinsic features of a catalyst as well as the reaction mechanism. Currently, high volume transient data is underutilized, and data science methods could both increase the value of information that can be extracted from this data, integrate experimental with theoretical data sources, and accelerate the pace of catalyst technology advancement. Transient kinetic characterizations with simple probe molecules exhibiting reversible adsorption, irreversible adsorption and bulk-surface diffusion are presented as training components for similar experiments with more complex surface reactions. In conclusion, by increasing the availability and accessibility of transient kinetic data through details of its structure and acquisition, we aim to decrease the barrier for data scientists to apply machine learning methods to this valuable data source.

Catalysis↗

Laying the Foundations for FAIR-er Science: ISA and the LSDA Data Submission Process in NASA's Evolving Data Management Environment

The Life Sciences Data Archive (LSDA) archives data resulting from research on the effects of spaceflight on humans and the development of countermeasures to mitigate spaceflight hazards. Archivists work with researchers to ensure that unique and high value data products and their metadata are preserved and managed to support current and future research. Currently, LSDA is updating its procedures and data submission requirements in response to the evolving data preservation environment at NASA. LSDA is implementing best practices for research data management through the establishment of clear data submission guidelines, integration of the FAIR (Findability, Accessibility, Interoperability, Reusability) principles, and use of the ISA (Investigation, Study, Assay) research metadata framework for data discoverability and transparency into the data management processes. These changes directly impact LSDA’s requirements for research data submissions. The newly revised Research Data Submission Agreement (RDSA), formerly the Data Submission Agreement (DSA), introduces ISA-compatible metadata collection standards to LSDA’s process. Adherence to LSDA’s data submission guidelines enhances the FAIR-ness of the repository’s collections for future users. This presentation will discuss (1) how submission of research data and associated metadata are impacted by current data management policies, (2) benefits of the adoption of FAIR principles and the ISA metadata framework for retrospective studies utilizing existing LSDA datasets and historic data collections, and (3) the support LSDA will provide to researchers during this transition.

LSDA↗

Using National-Scale Data to Inform Coal Power Plant Redevelopment and Coal Community Revitalization Planning

In the United States, individual coal plants have high variation in their construction and operation. Coal plants commonly have multiple units, with different commissioning ages, operational position, and maintenance conditions – even multiple fuels. These units may be in changing states of utilization, with variation in productivity depending on economic contexts. Coal plants themselves may be temporarily offline, mothballed, retired, or decommissioned due to owner/operator portfolios or market conditions, and the current posture of the plant may be difficult to ascertain without local news sources (Tarekegne et al. 2021). Anticipated dates of plant closures are prospectively represented to regulators and to the Energy Information Administration (EIA), but may also experience uncertainty due to the engineering, environmental, economic and social complexity of closing a very large power plant. As noted in this report, closure of the first unit to the last within a single plant may span more than 20 years. Yet tracking and predicting the position of our nation’s coal fleets on a national scale is deeply important. In developing a case for federal datasets and an approach to building and curating these data, this report identified that data must have high fidelity at the plant-level to be meaningful in aggregate. Plant-scale detail is also a match for its application: national priorities in federal investment range from economic revitalization for coal plant communities to maintaining reliable electric grid operations. This document is organized into a statistical review of coal plant retirements, past and prospective, in the United States; trends in the relationship between closures and existing federal investment programs for community revitalization; and potential methods for accelerating site redevelopment through advanced geospatial analysis based on federal data.

01 COAL, LIGNITE, AND PEAT↗

Deep transfer learning for star cluster classification: I. application to the PHANGS– HST survey

ABSTRACT We present the results of a proof-of-concept experiment that demonstrates that deep learning can successfully be used for production-scale classification of compact star clusters detected in Hubble Space Telescope(HST) ultraviolet-optical imaging of nearby spiral galaxies ($D\lesssim 20\, \textrm{Mpc}$) in the Physics at High Angular Resolution in Nearby GalaxieS (PHANGS)–HST survey. Given the relatively small nature of existing, human-labelled star cluster samples, we transfer the knowledge of state-of-the-art neural network models for real-object recognition to classify star clusters candidates into four morphological classes. We perform a series of experiments to determine the dependence of classification performance on neural network architecture (ResNet18 and VGG19-BN), training data sets curated by either a single expert or three astronomers, and the size of the images used for training. We find that the overall classification accuracies are not significantly affected by these choices. The networks are used to classify star cluster candidates in the PHANGS–HST galaxy NGC 1559, which was not included in the training samples. The resulting prediction accuracies are 70 per cent, 40 per cent, 40–50 per cent, and 50–70 per cent for class 1, 2, 3 star clusters, and class 4 non-clusters, respectively. This performance is competitive with consistency achieved in previously published human and automated quantitative classification of star cluster candidate samples (70–80 per cent, 40–50 per cent, 40–50 per cent, and 60–70 per cent). The methods introduced herein lay the foundations to automate classification for star clusters at scale, and exhibit the need to prepare a standardized data set of human-labelled star cluster classifications, agreed upon by a full range of experts in the field, to further improve the performance of the networks introduced in this study.

Wei, Wei↗

Geolab in NASA's First Generation Pressurized Excursion Module: Operational Concepts

We are building a prototype laboratory for preliminary examination of geological samples to be integrated into a first generation Habitat Demonstration Unit-1/Pressurized Excursion Module (HDU1-PEM) in 2010. The laboratory GeoLab will be equipped with a glovebox for handling samples, and a suite of instruments for collecting preliminary data to help characterize those samples. The GeoLab and the HDU1-PEM will be tested for the first time as part of the 2010 Desert Research and Technology Studies (DRATS), NASAs annual field exercise designed to test analog mission technologies. The HDU1-PEM and GeoLab will participate in joint operations in northern Arizona with two Lunar Electric Rovers (LER) and the DRATS science team. Historically, science participation in DRATS exercises has supported the technology demonstrations with geological traverse activities that are consistent with preliminary concepts for lunar surface science Extravehicular Activities (EVAs). Next years HDU1-PEM demonstration is a starting point to guide the development of requirements for the Lunar Surface Systems Program and test initial operational concepts for an early lunar excursion habitat that would follow geological traverses along with the LER. For the GeoLab, these objectives are specifically applied to enable future geological surface science activities. The goal of our GeoLab is to enhance geological science returns with the infrastructure that supports preliminary examination, early analytical characterization of key samples, insight into special considerations for curation, and data for prioritization of lunar samples for return to Earth.

Evans, C. A.↗

Machine learning framework for predicting uranium enrichments from M400 CZT gamma spectra

A machine learning framework was developed for predicting uranium enrichments from M400 CZT gamma spectra. This framework leverages the availability of a large amount of measured M400 gamma spectra and uses a recently updated version of Gamma Detector Response and Analysis Software (GADRAS) for gamma spectrum analysis and generation. It also leverages the existing machine learning modules in Python for gamma spectrum data processing, curation, model training, benchmarking, and optimization of the deep machine learning models. The framework is used to develop a deep learning model to analyze gamma spectra from a set of U 3 O 8 samples with enrichments ranging from 0.31 to 93.17% and UF 6 cylinders with enrichments ranging from 0.2 to 4.95%, and the model performance is tested using a set of measured spectra and the respective declared enrichment values. Results show that the model can correctly classify 99.35% of the U 3 O 8 sample enrichments, and can predict the samples’ enrichments within an average absolute error of 0.099% (in percentage points of enrichment). For the UF 6 cylinders, the average absolute error was approximately 0.03%, with an accuracy of 98% in classifying discrete enrichment values of UF 6 samples. Finally, the results also show that the model has performed significantly better in terms of predicting enrichments in UF 6 cylinders based on measured gamma spectra than the GEM code, with a standard deviation (of the relative errors) of 2.23% (compared with the 11.51% value for the GEM code) based on results from a set of test data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Frictionless knowledge injection for few-shot learning

Cutting-edge machine learning methods often require large volumes of curated training data, precluding their use in national security problems with rare events in massive datasets. We present a method for incorporating abstract knowledge into models tailored for sparse data. A subject matter expert defines salient concepts using data examples, which are encoded in the model’s embedding space. Models are then trained to respect these concepts. This method enables knowledge injection, yielding effective models with limited labeled data and the ability to assess model sensitivity for subject matter expertise across the nonproliferation mission space, as demonstrated with Raman spectra analysis.

Stomps, Jordan [ORNL] (ORCID:0000000178114479)↗

A data-centric weak supervised learning for highway traffic incident detection

Using the data from loop detector sensors for near-real-time detection of traffic incidents on highways is crucial to averting major traffic congestion. While recent supervised machine learning methods offer solutions to incident detection by leveraging human-labeled incident data, the false alarm rate is often too high to be used in practice. Specifically, the inconsistency in the human labeling of the incidents significantly affects the performance of supervised learning models. To that end, we focus on a data-centric approach to improve the accuracy and reduce the false alarm rate of traffic incident detection on highways. We develop a weak supervised learning workflow to generate high-quality training labels for the incident data without the ground truth labels, and we use those generated labels in the supervised learning setup for final detection. This approach comprises three stages. First, we introduce a data preprocessing and curation pipeline that processes traffic sensor data to generate high-quality training data through leveraging labeling functions, which can be domain knowledge-related or simple heuristic rules. Second, we evaluate the training data generated by weak supervision using three supervised learning models-random forest, k-nearest neighbors, and a support vector machine ensemble-and long short-term memory classifiers. The results show that the accuracy of all of the models improves significantly after using the training data generated by weak supervision. Third, we develop an online real-time incident detection approach that leverages the model ensemble and the uncertainty quantification while detecting incidents. Finally, we show that our proposed weak supervised learning workflow achieves a high incident detection rate (0.90) and low false alarm rate (0.08).

97 MATHEMATICS AND COMPUTING↗

Mapping use cases and dataset needs for benchmarking buildings data

A perennial challenge in buildings research is the lack of high-quality datasets that can be relied upon for a wide array of tasks, including model calibration and improving energy efficiency and load flexibility. Instrumenting a building for data collection is resource intensive, so it is important to be methodical in the approach and ensure that resulting data are flexible and useful for a broad range of analyses. This study aims to fill the gaps in characterizing potential use cases for buildings datasets and mapping them to dataset needs using a well-defined data infrastructure. Here, we have developed a systematic mapping strategy between buildings dataset needs and use cases to help streamline the processes of efficiently targeting datasets, designing building sensing systems, and determining buildings research use cases. We selected 14 prospective use cases and 11 refined buildings data categories for developing the preliminary dataset-needs-to-use-cases mapping matrix (‘DN-UC mapping matrix’) with generic ‘Tags’—a detailed sub-level of data categories extracted by justifying the needs of an aspect of the datasets to use cases. We present two example applications of the developed mapping matrix to demonstrate use of the mapping matrix and its effectiveness.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Spectroscopic features of dissolved iodine in pristine and gamma-irradiated nitric acid solutions

While iodine speciation is important for a wide range of nuclear activities, understanding the mechanisms of the transformations of iodine between chemical forms and the sensitivity of these transitions to solution conditions and exposure to radiation remains an active area of research. This work curates spectroscopic data from several experimental techniques and establishes their sensitivity and limitations in detecting changes in iodine speciation in both neutral and acidic regimes. The techniques include Raman spectroscopy, Fourier Transform Infrared (FTIR) spectroscopy, 127 I NMR spectroscopy, and ultraviolet-visible (UV-Vis) spectroscopy. Analysis of these data indicates that these commonly accessible spectroscopies often have dynamic ranges of measurable concentrations that do not always overlap between all techniques. The experimental techniques are disparately sensitive to iodide (I - ), molecular iodine (I 2 ), iodate (IO 3 - ), and periodate (IO 4 - ) species. Raman, FTIR, and NMR spectra were subsequentially analyzed using two-dimensional correlation analyses to generate high-resolution autocorrelation spectra. Here, the use of these spectroscopies is then extended to tracking acidification-induced and gamma irradiation-induced transformations of dissolved sodium iodate in deionized water and concentrated nitric acid. Both dissolution into nitric acid and irradiation with a gamma source are demonstrated to perturb the iodine speciation promoting their assembly into molecular iodine (I 2 ) and/or triiodide (I 3 - ). While I 2 and I 3 - species are undetectable with FTIR spectroscopy and 127 I NMR spectroscopy, the species can be detected with UV-Vis spectroscopy, and in some instances, I 3 - can be detected with Raman spectroscopy in the low wavenumber region. Ultimately, the results of this work provide a path to designing optimal combinations of techniques to detect forms of iodine across a wide range of concentrations and conditions.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Cluster-Graph Fingerprinting: A Framework for Quantitative Analysis of Machine-Learned Interatomic Model Training and Simulation Data

Machine-learned interatomic models represent a significant advancement in simulation methods, extending the predictive ability of first-principles methods to previously inaccessible length and time scales. However, the data-driven nature of these models can lead to difficult-to-detect errors that can compromise prediction accuracy. To address this challenge, we introduce a novel fingerprinting approach based on the Chebyshev Interaction Model for Efficient Simulation (ChIMES) ML-IAM graph-based descriptor. Our strategy enables efficient and statistically rigorous analysis of system configurations used in ML-IAM training and those generated by their application, e.g., in molecular dynamics simulations. We demonstrate that these fingerprints can effectively assess novelty of a configuration relative to an existing data set and determine dissimilarity among individual configurations, which are two key tasks in workflows for active learning-based ML-IAM training, data set curation, and on-the-fly uncertainty quantification.

36 MATERIALS SCIENCE↗