Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metadata extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Biogeochemistry of Pond B (Savannah River Site, South Carolina, USA): Sediment Core, Total extraction data, Pond B Savannah River Site July 2019. Subsurface Biogeochemistry of Actinides SFA

Pond B at Savannah River Site (SRS, South Carolina) is a monomictic reservoir that received SRS R reactor cooling water from 1961–1964. Previous studies conducted between the 1980s–1990s on the water column and sediments of Pond B measured trace amounts of Pu (33 MBq 238Pu and 430 MBq 239,240Pu), 241Am, and 137Cs. Since then, the pond has been relatively isolated and the radionuclide concentrations have not been monitored over time. Herein, about 30 years after the last publication on Pond B, we are re-evaluating the geochemistry and radionuclide distribution within Pond B at four locations along a horizontal transect from the inlet to outlet.This study investigated the distribution of anthropogenic radionuclides Pu-239 and Cs-137 along with total organic carbon, iron, and trace element in contaminated sediments of Pond B at the Savannah River Site (SRS). Pond B received reactor cooling water from 1961 to 1964, and trace amounts of Pu-239 and Cs-137 during operations. Our study collected sediment cores to determine concentrations of Pu-239, Cs-137, and major and minor elements in solid phase, pore water and an electrochemical method was used on wet cores to determine dissolved elemental concentrations.

54 ENVIRONMENTAL SCIENCES↗

Data from a four-day long microcosm experiment addressing the destabilization of artificial mineral-associated organic matter by model root exudates embedded in a soil matrix from the Rocky Mountain Biological Laboratory (Gothic, CO, USA), 2019

This dataset provides data collected during a four-day long laboratory soil microcosm experiment testing the efficacy of root exudate-driven mineral-associated organic matter destabilization. This dataset contains four data files in comma-separate values (*.csv). The files provide the metadata and the experimental results on microbial respiration, MAOM-derived respiration, and sequential mineral-extractions. This data was used to produce the figures in Bölscher et al., 2026. The results of the experiment can be found in the open access article Bölscher et al., 2026 (https://doi.org/10.1016/j.soilbio.2026.110276). Abstract: Mineral-associated organic matter (MAOM) is often considered stable, but root exudates can destabilize MAOM via various pathways. Theory and model system studies suggest that direct MAOM destabilization by strong ligands, like oxalic acid, or reducing agents, like catechol, is more effective than indirect, microbial-mediated MAOM destabilization, stimulated by less reactive compounds like glucose. Here, we demonstrate that the presence of a soil matrix alters the efficacy of exudate-driven MAOM destabilization pathways. Glucose and catechol destabilized significantly greater amounts of MAOM from ferrihydrite and aluminum hydroxide (Al (OH)3) embedded in a soil matrix than oxalic acid. Our findings indicate that indirect, microbial-mediated MAOM destabilization may play a larger role than direct MAOM destabilization in soil environments.

Destabilization↗

Quarterly Soil Core and Root Analyses from the Missouri Ozarks AmeriFlux (MOFLUX) Site, Ashland, Missouri, 2017-2023

This dataset contains quarterly soil core measurements from the Missouri Ozarks AmeriFlux (MOFLUX) site located at the University of Missouri’s Thomas H. Baskett Wildlife Research and Education Area near Ashland, Missouri. These data will be used to parameterize an ensemble of MOFLUX-optimized soil carbon-nitrogen models, used to simulate carbon (C) and nitrogen (N) cycling responses to future hydroclimatic scenarios and the trajectory of soil C stocks with concomitant forest decline. Beginning in 2017, eight soil cores were collected approximately quarterly near plot 1 of the southeast transect, near the automated soil respiration flux chambers, from 0–15 cm depth. Data are currently available through 2023 (2017-06-14 to 2023-11-13); additional observations will be appended to this dataset as they become available. Cores were analyzed for gravimetric moisture content, pH, total carbon and nitrogen, texture, microbial biomass carbon and nitrogen, and extractable dissolved organic carbon and nitrogen. This dataset contains one data file in comma separate (*.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (*.csv) format and a user guide in PDF (*.pdf) format.

54 ENVIRONMENTAL SCIENCES↗

Application of a Dataset-Publication Knowledge Graph for Improving Earth Science Data Search

Finding a dataset at a NASA data center that is the best fit for the researcher’s application presents a challenge, not only for a novice user but for an experienced one, due to the data complexity and a multitude of choices of the existing data. Users often search for the data based on the application they are interested in, their research domain, phenomena, research topic, etc. As existing dataset metadata may not cover these search terms, the user may not obtain the most relevant results for their purpose. This problem was addressed by leveraging the content of the titles and abstracts of the research papers that utilize NASA datasets. For this, features from the paper titles and abstracts were extracted, and then a knowledge graph (KG) was used to link these features to the datasets used in that paper. The search for the datasets was tested by querying this knowledge graph through various terms extracted from Earth Science ontologies such as Semantic Web for Earth and Environment Technology (SWEET), and it was shown that this KG search outperforms the existing search that exclusively queries the dataset metadata.

Kristina Stoyanova↗

APPL Hyperspectral_Imaging_Dataset_for_Heritability_Analysis_in_Populus_trichocarpa

This dataset contains hyperspectral imaging data collected at the Advanced Plant Phenotyping Laboratory (APPL) at Oak Ridge National Laboratory. Natural variants of Populus trichocarpa were imaged using a high-throughput hyperspectral phenotyping pipeline to quantify spectral reflectance traits for downstream quantitative genetics analyses. The dataset includes hyperspectral image files and derived reflectance data products suitable for extracting spectral features across the measured wavelength range (e.g., VNIR and/or SWIR, depending on instrument configuration), along with associated sample metadata (e.g., genotype identifiers, experimental design factors, and imaging run identifiers). These data were generated to support analyses of broad-sense heritability of hyperspectral traits and their relationships with biochemical phenotypes (including lignin traits from Py-MBMS).

APPL↗

The Heliophysics Data Environment, Virtual Observatories, NSSDC, and SPASE

Heliophysics (the study of the Sun and its effects on the Solar System, especially the Earth) has an interesting data environment in that the data are often to be found in relatively small data sets widely scattered in archives around the world. Within the last decade there have been more concentrated efforts to organize the data access methods and create a Heliophysics Data and Model Consortium (HDMC). To provide data search and access capability a number of Virtual Observatories (VO's) have been established both via funding from the U.S. National Aeronautics and Space Administration (NASA) and through other funding agencies in the U.S. and worldwide. At least 15 systems can be labeled as Heliophysics Virtual Observatories, 9 of them funded by NASA. Other parts of this data environment include Resident Archives, and the final, or "deep" archive at the National Space Science Data Center (NSSDC). The problem is that different data search and access approaches are used by all of these elements of the HDMC and a search for data relevant to a particular research question can involve consulting with multiple VO's - needing to learn a different approach for finding and acquiring data for each. The Space Physics Archive Search and Extract (SPASE) project is intended to provide a common data model for Heliophysics data and therefore a common set of metadata for searches of the VO's and other data environment elements. The SPASE Data Model has been developed through the common efforts of the HDMC representatives over a number of years. We currently have released Version 2.1. of the Data Model. The advantages and disadvantages of the Data Model will be discussed along with the plans for the future. Recent changes requested by new members of the SPASE community indicate some of the directions for further development.

Thieman, James↗

Metadata Entry Optimization for NASA's Biological Institutional Scientific Collection (NBISC)

The NASA Biological Institutional Sample Collection (NBISC) at NASA’s Ames Research Center is a critical resource housing non-human samples collected from spaceflight missions and ground analog studies, primarily consisting of specimens from rats, mice, and select microbes. The primary objective of NBISC is to systematically receive, document, preserve, and facilitate access to these samples for the global scientific community. NBISC promotes international collaboration and maximizes the return on investment for precious tissues from spaceflight and analog experiments. Researchers can request physical samples through an online request form and subsequent written proposal review process. This study addresses two core research objectives: streamlining the NBISC sample lifecycle processes and strategizing for managing an influx of 50,000 tissue samples from a series of cosmic radiation analog experiments carried out at the NASA Space Radiation Laboratory (NSRL) by Drs. Eleanor Chang (Lawrence Berkeley Laboratory) and Polly Blakely (SRI). The Chang/Blakely studies investigated Harderian gland (HG) tumorigenesis in mice exposed to low dose and LET radiation comprising 8 different exposure protocols in over 4000 mice. NBISC sample metadata is stored in a Laboratory Information Management System (SLIMS). To streamline sample data entry, we customize python scripts using information extracted from the individual experimental protocols. The scripts automate entry into multiple SLIMS data fields including protocol name, unique sample barcode, tissue and sub-tissue information, freezer location, sample preservation method, etc. The semi-automated procedure significantly decreases the time spent on data entry by several orders of magnitude. Automation and data organization are essential, as they free up time for curation and promotion of the collection which, in turn, increase the accessibility of samples to the broader research community. NBISC benefits from streamlined data ingestion, and the methodologies developed here are applicable to other projects which use SLIMS including the NASA Biospecimen Sharing Program and GeneLab. As of Fall 2023, plans include transferring sample data from SLIMS to public facing repositories (OSDR and NLSP), expanding the reach of the Chang/Blakely sample collection. The Human Research Program Space Radiation Element plans to transfer non-human tissues from many more investigations to NBISC in the coming year.

Sample Repository↗

Metadata Entry Optimization For NASA's Biological Institutional Scientific Collection (NBISC)

The NASA Biological Institutional Sample Collection (NBISC) at NASA’s Ames Research Center is a critical resource housing non-human samples collected from spaceflight missions and ground analog studies, primarily consisting of specimens from rats, mice, and select microbes. The primary objective of NBISC is to systematically receive, document, preserve, and facilitate access to these samples for the global scientific community. NBISC promotes international collaboration and maximizes the return on investment for precious tissues from spaceflight and analog experiments. Researchers can request physical samples through an online request form and subsequent written proposal review process. This study addresses two core research objectives: streamlining the NBISC sample lifecycle processes and strategizing for managing an influx of 50,000 tissue samples from a series of cosmic radiation analog experiments carried out at the NASA Space Radiation Laboratory (NSRL) by Drs. Eleanor Chang (Lawrence Berkeley Laboratory) and Polly Blakely (SRI). The Chang/Blakely studies investigated Harderian gland (HG) tumorigenesis in mice exposed to low dose and LET radiation comprising 8 different exposure protocols in over 4000 mice. NBISC sample metadata is stored in a Laboratory Information Management System (SLIMS). To streamline sample data entry, we customize python scripts using information extracted from the individual experimental protocols. The scripts automate entry into multiple SLIMS data fields including protocol name, unique sample barcode, tissue and sub-tissue information, freezer location, sample preservation method, etc. The semi-automated procedure significantly decreases the time spent on data entry by several orders of magnitude. Automation and data organization are essential, as they free up time for curation and promotion of the collection which, in turn, increase the accessibility of samples to the broader research community. NBISC benefits from streamlined data ingestion, and the methodologies developed here are applicable to other projects which use SLIMS including the NASA Biospecimen Sharing Program and GeneLab. As of Fall 2023, plans include transferring sample data from SLIMS to public facing repositories (OSDR and NLSP), expanding the reach of the Chang/Blakely sample collection. The Human Research Program Space Radiation Element plans to transfer non-human tissues from many more investigations to NBISC in the coming year.

Biospecimen↗

CHESS 2025: Crown polygons and extracted reflectance for field sampling sites

This dataset contains (1) crown polygons for each tree, meadow, and shrub site sampled in the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign (in geojson format, .geojson) and (2) extracted reflectance, uncertainty, and shade estimates for each crown polygon from the 2018 National Ecological Observatory Network (NEON) and 2025 CHESS campaigns. (in CSV format, .csv). Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). Crown polygons were manually delineated for each site in the 2025 campaign using a combination of field-collected GPS data (doi:10.15485/3022418), RGB (red, green, blue) and false color reflectance mosaics (doi:10.15485/3013535), and LiDAR-derived (Light Detection and Ranging) canopy height (CHM) and digital surface (DSM) models (DOI and citation to be added upon publication). Where there was misalignment between the spectrometer- and LiDAR-derived data products, polygons prioritized alignment with the spectrometer-derived data products. Polygons were delineated conservatively to only select pixels representative of vegetation samples collected in the field. Crown polygons for 2018 are published at (doi:10.15485/1618130) and were developed using the same protocol. For each polygon, all pixels from all flightlines were extracted where the pixel centroid was contained within the polygon. For each pixel, we extracted the surface reflectance, uncertainty, and shade estimates. Details on the extracted datasets are available at (doi:10.15485/3013527, doi:10.15485/3013535). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

ImageLabler: Labeling and Managing Image Data for Machine Learning in the Earth Sciences

While machine learning techniques for image classification have been around for a long time, storing and managing the vast number of images required as training data is still a problem for scientists. This is especially true for the field of Earth science, where only recently have experts begun using machine learning techniques for image-based phenomena classification. Image Labeler, a fast and scalable cloud-based tagging platform for Earth science images, seeks to improve upon existing methods of managing images and associated metadata, such as maintaining categorized folders of images on a local machine, a process that can be cumbersome and difficult to scale. The platform facilitates rapid development of image-based Earth science phenomena training datasets by allowing scientists to upload their existing imagery as well as extract new samples from open satellite imagery services made available through NASA’s Global Imagery Browse Service (GIBS). Image Labeler also supports GeoTIFF data, with capabilities such as displaying GeoTIFFs on an interactive map, drawing shapefiles over them, and tagging them with additional metadata. This allows scientists to perform spatiotemporal subsetting with geographic information and develop training data more quickly. Built using modern web technologies, Image Labeler includes additional capabilities such as team collaboration for large-scale image tagging projects. Users can download their data in a machine-learning-ready format, allowing scientists to spend time on experimentation rather than on the collection of training data. In this presentation, we demonstrate how Image Labeler seeks to become a one-stop image data management solution for machine learning applications in Earth science.

Ashish Acharya↗

HydroDCM: Hydrological Domain-Conditioned Modulation for Cross-Reservoir Inflow Prediction

Deep learning models have shown promise in reservoir inflow prediction, yet their performance often deteriorates when applied to different reservoirs due to distributional differences, referred to as the domain shift problem. Domain generalization (DG) solutions aim to address this issue by extracting domain-invariant representations that mitigate errors in unseen domains. However, in hydrological settings, each reservoir exhibits unique inflow patterns, while some metadata beyond observations like spatial information exerts indirect but significant influence. This mismatch limits the applicability of conventional DG techniques to many-domain hydrological systems. To overcome these challenges, we propose HydroDCM, a scalable DG framework for cross-reservoir inflow forecasting. Spatial metadata of reservoirs is used to construct pseudo-domain labels that guide adversarial learning of invariant temporal features. During inference, HydroDCM adapts these features through light-weight conditioning layers informed by the target reservoir’s metadata, reconciling DG’s invariance with location-specific adaptation. Experiment results on 30 real-world reservoirs in the Upper Colorado River Basin demonstrate that our method substantially outperforms state-of-the-art DG baselines under many-domain conditions and remains computationally efficient.

Hu, Pengfei [ORNL] (ORCID:0009000367130950)↗

Soil physical and chemical measurements for topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The package is part of the DOE Watershed Function Science Focus Area (SFA) project and includes soil physical and chemical measurements from topsoils collected at the East River, Colorado, in conjunction with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey conducted in June 2018. The soil measurements include soil bulk density, soil volumetric water content, soil microbial biomass C (Carbon), N (Nitrogen) and C:N (C to N ratio), soil DNA yield, soil total extractable organic C, soil total extractable N, soil extractable nitrate, soil extractable ammonium, soil dissolved inorganic N, soil dissolved organic N, soil pH, soil TOC400 (total organic carbon at 400°C), soil ROC (residual oxidizable carbon), soil TIC (total inorganic carbon), soil TOC (total organic carbon), soil TC (total carbon), soil N, soil OM (organic matter) loss on ignition. Additional associated site metadata can be found in the ESS-DIVE package 10.15485/1618130. The dataset includes (1) 2018_NEON_soil_physical_chemical_measurements.csv: soil physical and chemical measurements indexed by soil sample IGSNs; (2) samples.csv: sample metadata file used to register International Generic Sample Numbers (IGSNs); (3) flmd.csv: file level metadata file; and (4) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS (Catchment Hydrology and ↗

Availability of Previously Unprocessed ALSEP Raw Instrument Data, Derivative Data, and Metadata Products

In year 2010, 440 original data archival tapes for the Apollo Lunar Science Experiment Package (ALSEP) experiments were found at the Washington National Records Center. These tapes hold raw instrument data received from the Moon for all the ALSEP instruments for the period of April through June 1975. We have recently completed extraction of binary files from these tapes, and we have delivered them to the NASA Space Science Data Cordinated Archive (NSSDCA). We are currently processing the raw data into higher order data products in file formats more readily usable by contemporary researchers. These data products will fill a number of gaps in the current ALSEP data collection at NSSDCA. In addition, we have estabilished a digital, searcheable archive of ALSEP document and metadata as part of the web portal of the Lunar and Planetary Institute. It currently holds approx. 700 documents totaling approx. 40,000 pages

ALSEP↗

Scripts and data associated with a manuscript linking soil and sediment elemental composition with dissolved organic matter chemistry across CONUS

This data package provides scripts and geochemical data for a manuscript titled “Linkages between mineral element composition of soils and sediments with hyporheic zone dissolved organic matter chemistry across the contiguous United States” (preprint: doi: 10.22541/essoar.169447343.31694990/v1). This data is associated with the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS, https://whondrs.pnnl.gov) and is an extension of the Summer 2019 Sampling campaign which crowdsourced samples from rivers and sediment across the continental United States. Data from this study can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719. The main objective of this manuscript was to couple sediment water extractable dissolved organic matter chemistry, defined by ultra-high resolution mass spectrometry, with localized sediment elemental composition and watershed scale soil elemental characteristics. This data package contains one main folder with four subfolders. The main data folder contains (1) readme; (2) data dictionary (dd); (3) file-level metadata (flmd); (4) an R markdown to reproduce manuscript figures and analyses; (5) a pdf of instructions to reproduce NGS interpolations with ArcGIS software; and (6) a python script to reproduce NGS extrapolations with python. The four subfolders contain files required to reproduce NGS extrapolations include (1) ‘CONUS_boundaries’ containing boundary layers (.shp) for the Continental United States; (2) ‘ngs_project’ containing files (.shp) with point level NGS soil elemental data (Grossman et al., 2004); (3) ‘raster_outputs’ containing the interpolated raster output files for various soil elements; and (4) ‘NGS_Chemistry_Final’ contain final extracted soil elemental data.

54 ENVIRONMENTAL SCIENCES↗

EXCHANGE Campaign Degradation (ECD): Understanding Decomposition Dynamics Across Mid-Atlantic and Great Lakes Coastal Ecosystems

The EXploration of Coastal Hydrobiogeochemistry Across a Network of Gradients and Experiments (EXCHANGE) Degradation Experiment (EXCHANGE-D) is an in situ experiment designed to assess organic matter decomposition rates across coastal terrestrial-aquatic interfaces (TAIs), from coastal uplands through transition zones to wetlands. Through a network of partner scientists and coastal sites, we are testing how environmental gradients shape decomposition and carbon dynamics across terrestrial-aquatic interfaces. Using standardized tea bag substrates deployed across a network of diverse coastal sites, we compare decomposition rates at different fresh- and salt-water TAIs to develop transferable knowledge that improves the representation of organic matter degradation in coastal ecosystem models. For more information, please see https://compass.pnnl.gov/FME/EXCHANGE. This is Version 1 of the data package, which includes: ecd_README.pdf flmd.csv dd.csv ecd_soil_weom_L2.csv ecd_soil_ph_conductivity_L2.csv ecd_soil_gwc_L2.csv ecd_soil_teabag_degradation_L2.csv ecd_readme.pdf

coastal soils↗

Requirements for Cataloging Hanford Geophysical Datasets

Environmental management activities at the Hanford Site produce extensive data about site conditions, contaminants, cleanup, and more. Managing and archiving that data requires a high degree of collaboration among site contractors and a high level of awareness by project managers and staff. Part of that effort is developing a Hanford Environmental Information and Data Index (HEIDI) to organize the data and maximize its value by making it findable and available for reuse. The objective is to catalog the disparate data sets collected to address the evolving needs of planning, executing, and documenting cleanup over several decades up to the present day, including links to active data sources when available. A properly implemented data catalog makes finding environmental datasets related to an area or theme a routine, reliable process, without requiring the searcher to have special knowledge that a data set exists and where it may be stored. In this project, a working group, including the U.S. Department of Energy, the Hanford Site contractors, and Pacific Northwest National Laboratory staff, identified needs and requirements for handling complex site data. Geophysical data was chosen as a test case because it can be large and complex and often involves multiple processing steps to extract the information incorporated into deliverables. The ability to document those steps was one of the requirements identified for the catalog. In addition to developing requirements, other activities included selecting a metadata schema and initial testing with the objective of determining whether the workflow and capabilities of selected data catalog software platforms were sufficient to implement and impose the identified requirements. This initial testing involved running the default catalog instance using the software platform of interest and altering the configuration to achieve each requirement, if possible. Where configuration alone was insufficient, the possibility of modifying the software by changing the code was examined, but not implemented. A follow-on task is planned to reprogram the code as necessary to implement requirements in a prototype catalog.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Capturing, Analyzing, Maintaining, and Disseminating Shape Memory Material Data Between Information Management Systems

With an increased demand on reducing the time, cost, and effort to develop new materials, Integrated Computational Materials Engineering (ICME) has received widespread attention in various engineering disciplines as a catalyst for significantly reducing experimental testing during the material design process. An ICME approach to design can enable ‘fit-for-purpose’ materials to be realized in engineering applications by incorporating well-understood process-property-performance relationships between the various length and time scales in a material’s structure, enabling material optimization. However, such an approach requires validated multiscale models at the various length scales for a material, which in turn requires a large amount of data, a robust means of storing the data, and the ability to link data to developed material models. The NASA Vision 2040 [1] has identified nine key elements to enabling ICME approaches in system level design, with one being “Data, Information, and Visualization”, thus outlining the importance of a robust information management system for ICME. As the relationship between microstructure, properties, and material performance become better understood and incorporated into multiscale models that can be leveraged in application design, the emergence of new materials with application-driven properties can be realized. One such new material class that has seen growing attention are shape memory materials (SMM), in which a material can transition between a deformed and undeformed state via a reversible phase transformation when subject to a thermal, mechanical, or magnetic load [2]. SMMs have been used widely in aerospace and biomedical industries, including applications such as actuators, low-shock mechanisms, medical staples, braces, and stents [3, 4]. These materials exhibit unique behavior due to their ability to transition between phases, and thus the mechanisms that enable this transition must be captured in a data information management system and incorporated into SMM material models. At NASA Glenn Research Center, the Shape Memory Materials Database (SMMD) Tool has been developed to capture the necessary information that governs SMM material behavior and provide users the ability to select and visualize various SMMs for a specific application [5]. The database contains point-wise data for published SMM materials, along with the pedigree metadata for traceability necessary for a robust information management system. The database is also capable of storing in-house test data performed at NASA GRC by interacting with the developed Shape Memory Alloy (SMA) Analytics tool to extract the necessary point-wise values and populate the database. Although the SMMD Tool offers its users a single, authoritative source for SMM material data that is critical for model development and material design, the full material pedigree of the in-house test data for SMMs is not currently captured and is out of the scope for the SMMD tool. In this work, the schema for capturing SMM test data within the larger NASA GRC ICME Schema [6, 7, 8, 9] will be developed and implemented for thermomechanical tests conducted at NASA GRC. The developed schema will not only store the relevant data needed for the SMMD tool, but also the material pedigree (i.e., production of the bulk material, bulk material analysis, sample cut-out diagrams, sample fabrication procedure, etc.), test pedigree (i.e., test equipment used, measurement systems used, raw test data), and analysis pedigree (i.e., how the data in the SMMD tool is calculated). Furthermore, a Python-based framework will be developed to seamlessly interact between the SMA Analytics and SMMD tools, which will write the full dataset and associated metadata to the GRC Information Management System before passing the required point-wise data to the SMMD tool. Data informatics is a key element of the NASA Vision 2040, which requires not only that data is stored and maintained throughout the material lifecycle, but that the data is also accessible and reusable such that material development efforts can be minimized. Therefore, for an ICME design approach to be realized, a centralized information management system that drives the ICME process must be able to communicate with other databases. The work that will be presented in this presentation will therefore not only demonstrate the ability of NASA GRC’s information management system to capture SMM data, but also its ability to interact with pre-existing tools specialized for such materials.

Data management↗

Challenges and Vision for Standardization of Biopolymer Data Sets for Machine Learning

Machine learning (ML) is transforming materials research, yet potential for biopolymer discovery remains constrained by fragmented data and nonstandardized reporting. Biopolymers differ significantly from synthetic polymers, requiring specialized approaches to represent their biosynthetic origins, hierarchical structures, and application-specific metrics. In this Perspective, we identify three core challenges limiting biopolymer representation: information encoding, data quality, and data sharing. We describe the most pressing issues and propose commensurate approaches to address each key challenge. Recommendations include the design and adoption of biopolymer-specific fingerprinting and representation frameworks, development of hybrid human-large language model (LLM) data extraction strategies, and expanding Findable, Accessible, Interoperable, Reusable (FAIR)-compliant repositories. We propose a robust foundation to define interoperable, high-quality data sets that capture the full context of biopolymer materials. Standardized metadata, shared ontologies, and community-driven infrastructure would enable scalable, reproducible workflows and accelerate the ML-driven development of biopolymers.

36 MATERIALS SCIENCE↗