Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Machine Learning Prediction of the Experimental Transition Temperature of Fe(II) Spin-Crossover Complexes

Spin-crossover (SCO) complexes are materials that exhibit changes in the spin state in response to external stimuli, with potential applications in molecular electronics. It is challenging to know a priori how to design ligands to achieve the delicate balance of entropic and enthalpic contributions needed to tailor a transition temperature close to room temperature. Here, we leverage the SCO complexes from the previously curated SCO-95 data set [Vennelakanti et al. J. Chem. Phys. 159, 024120 (2023)] to train three machine learning (ML) models for transition temperature (T 1/2 ) prediction using graph-based revised autocorrelations as features. We perform feature selection using random forest-ranked recursive feature addition (RF-RFA) to identify the features essential to model transferability. Of the ML models considered, the full feature set RF and recursive feature addition RF models perform best, achieving moderate correlation to experimental T 1/2 values. We then compare ML T 1/2 predictions to those from three previously identified best-performing density functional approximations (DFAs) which accurately predict SCO behavior across SCO-95, finding that the ML models predict T 1/2 more accurately than the best-performing DFAs. In addition, we study ML model predictions for a set of 18 SCO complexes for which only estimated T 1/2 values are available. Upon excluding outliers from this set, the RF-RFA RF model shows a strong correlation to estimated T 1/2 values with a Pearson’s r of 0.82. In contrast, DFA-predicted T 1/2 values have large errors and show no correlation to estimated T 1/2 values over the same set of complexes. Overall, our study demonstrates slightly superior performance of ML models in comparison with some of the best-performing DFAs, and we expect ML models to improve further as larger data sets of SCO complexes are curated and become available for model training.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Changes in Four Decades of Near‐CONUS Tropical Cyclones in an Ensemble of 12 km Thermodynamic Global Warming Simulations

We evaluate tropical cyclones (TCs) in a set of thermodynamic global warming (TGW) simulations over the continental United States (CONUS). A 12 km simulation forced by ERA5 provides a 40‐year historical (1980–2019) control. Four complimentary future scenarios are generated using thermodynamic deltas applied to lateral boundary, interior, and surface forcing. We curate a data set of 4,498 6‐hourly TC snapshots in the control and find a corresponding “twin” in each counterfactual, permitting a paired comparison. Warming results in an increase in mean dynamical TC intensity and moisture‐related quantities, with the latter being more pronounced. TC inner cores contract slightly but outer storm size remains unchanged. The frequency with which TCs become more intense is only moderately consistent, with snapshots having increased hazards ranging from 50% to 80% depending on warming level. The fractions of TCs undergoing rapid intensification and weakening both increase across all warming simulations, suggesting elevated short‐term intensity variability.

54 ENVIRONMENTAL SCIENCES↗

Edge AI-Enhanced Traffic Monitoring and Anomaly Detection Using Multimodal Large Language Models

This paper addresses the challenge of traffic monitoring and incident detection in remote areas, utilizing multimodal large language models (LLMs) deployed on edge AI devices. The key novelty of the LLM is to convert real-time video streams into descriptive texts, enabling low-bandwidth transmissions and reliable detection of anomalies and incidents in environments of intermittent connectivity. The model is developed based on fine-tuning open-source LLMs and extending it with multi-modal capabilities to analyze video frames. Our work also involves deploying this model on edge devices such as Nvidia IGX Orin and is planned to be tested in realistic environments in future work. The methodology includes data set curation, iterative model fine-tuning and compression, and hardware-based optimization. This approach aims to enhance traffic safety and response speed in remote areas, marking a significant advancement in the application of AI for traffic monitoring and safety management.

Peruski, Ryan [University of Tennessee, Knoxville ↗

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram↗

Seascape Interface Control Document (V.1)

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source codes, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Seascape Interface Control Document (V. 2)

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source codes, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Seascape Interface Control Document

This paper serves as the Interface Control Document (ICD) for the Seascape automated test harness developed at Sandia National Laboratories. The primary purposes of the Seascape system are: (1) provide a place for accruing large, curated, labeled data sets useful for developing and evaluating detection and classification algorithms (including, but not limited to, supervised machine learning applications) (2) provide an automated structure for specifying, running and generating reports on algorithm performance. Seascape uses GitLab, Nexus, Solr, and Banana, open source software, together with code written in the Python language, to automatically provision and configure computational nodes, queue up jobs to accomplish algorithms test runs against the stored data sets, gather the results and generate reports which are then stored in the Nexus artifact server.

97 MATHEMATICS AND COMPUTING↗

Workshop Summary: Bridging the Gap Between Atmospheric Science and Grid Integration

The need for dedicated, accurate, expertly curated weather data is increasingly important as the share of variable renewable energy increases on the power system. Projections for futures with very high (50+% annual energy) shares of variable generation require ongoing assessment of data requirements from industry stakeholders in their power system operation and planning contexts. In March 2024, NREL organized a workshop entitled "Bridging the Gap Between Atmospheric Science and Grid Integration Workshop", which brought atmospheric scientists and power system experts together to refine the requirements of atmospheric datasets for grid integration, and to describe a holistic approach to creating new and regularly updated national scale wind datasets for power system planning and operations. The results of this workshop are being used to inform the near-term development and a longer-term strategy for DOE to produce relevant wind resource datasets and inform wider use of wind/solar/load data sets in power system planning. This presentation provides an overview of a preworkshop survey, an assessment of current state of the art of national-scale datasets for wind resource assessment and grid integration, insights on appropriate uses of the WTK-LED, power system perspectives on data needs, as well as recommended next steps as discussed in the workshop and how these steps support longer-term strategies.

17 WIND ENERGY↗

Decadal Seasonal Shifts of Precipitation and Temperature in TRMM and AIRS Data

We present results from an analysis of seasonal phase shifts in the global precipitation and surface temperatures. We use data from the TRMM (Tropical Rainfall Measuring Mission) Multi-satellite Precipitation Algorithm (TMPA), and the Atmospheric Infrared Sounder (AIRS) on Aqua satellite, all hosted at NASA Goddard Earth Science Data and Information Services Center (GES DISC). We explore the information content and data usability by first aggregating daily grids from the entire records of both missions to pentad (5-day) series which are then processed using Singular Value Decomposition approach. A strength of this approach is the normalized principal components that can then be easily converted from real to complex time series. Thus, we can separate the most informative, the seasonal, components and analyze unambiguously for potential seasonal phase drifts. TMPA and AIRS records represent correspondingly 20 and 15 years of data, which allows us to run simple “phase learning†from the first 5 years of records and use it as reference. The most recent 5 years are then phase-compared with the reference. We demonstrate that the seasonal phase of global precipitation and surface temperatures has been stable in the past two decades. However, a small global trend of delayed precipitation, and earlier arrival of surface temperatures seasons, are detectable at 95% confidence level. Larger phase shifts are detectable at regional level, in regions recognizable from the Eigen vectors to having strong seasonal patterns. For instance, in Central North America, including the North American Monsoon region, confident phase shifts of 1-2 days per decade are detected at 95% confidence level. While seemingly symbolic, these shifts are indicative of larger changes in the Earth Climate System. We thus also demonstrate a potential usability scenario of Earth Science Data Records curated at the NASA GES DISC in partnership with Earth Science Missions.

surface temperatures↗

Empowering Open Science with the Science Discovery Engine

: This presentation will describe the work to date in building the SDE as well as what the team has learned about the SMD ecosystem, information curation, and data governance. A short demonstration of the SDE will be presented, and an overview of near- and long-term goals for future development will be shared. Community feedback will be welcomed about the interface, content, and other features to help inform actions to maximize the SDE’s performance and usability. Whether users aim to discover Earth-like atmospheres on planets outside of our solar system or better understand the impacts of solar energy on our own planet, the Science Discovery Engine provides a means for scientists and all curious individuals to find content to further their understanding of science across all time and space scales.

Emily Foshee↗

AssistTaxi: A Comprehensive Dataset for Taxiway Analysis and Autonomous Operations

The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems. This poster presents AssistTaxi, which is a comprehensive novel dataset which is a collection of images for runway and taxiway analysis. The dataset comprises of more than 300,000 frames of diverse and carefully collected data, gathered from Melbourne (MLB) and Grant-Valkaria (X59) general aviation airports. The importance of AssistTaxi lies in its potential to advance autonomous operations, enabling researchers and developers to train and evaluate algorithms for efficient and safe taxiing. Researchers can utilize AssistTaxi to benchmark their algorithms, assess performance, and explore novel approaches for runway and taxiway analysis. Additionally, the dataset serves as a valuable resource for validating and enhancing existing algorithms as well as facilitating innovation in autonomous operations for aviation. We also propose an initial approach to label the dataset using a contour based detection and line extraction technique.

Data Collection↗

Cislunar Trajectory Design and Maneuver Autonomy for NASA's Moon to Mars Architecture

NASA’s Moon to Mars architecture is an ambitious roadmap of manned cislunar and deep space exploration. The extensive amount of orbital assets required will place a significant burden on ground-based resources, such as communication networks and operations facilities. Spacecraft autonomy is essential for maintaining a vast number of complex missions beyond Earth orbit. To achieve full autonomy, spacecraft must be able to employ methods of robust maneuver design without an explicit dependence on commands sent from the ground. This level of autonomy is needed not only for stationkeeping, but also for outbound transfers. To address the need of spacecraft maneuver design autonomy, this work investigates the use of neural networks (NNs) in a supervised learning environment. A supervised learning approach for NNs allows for a curated training data set, consisting exclusively of perturbations applied to a desired mission concept of operations (ConOps). The proposed approach allows humans on the ground to design a specific mission ConOps before flight, then employ NNs to fly the mission robustly and autonomously. This investigation numerically tests maneuver autonomy in four highly sensitive regions of flight: orbit raising, translunar injection burns, powered lunar flybys, and invariant manifold insertion burns. These straining cases are contextualized by testing them in a demonstration mission, targeting an Earth-Moon L3 orbit. The study first establishes feasibility by automating impulsive burn maneuvers. However, some guidance algorithms will need more intensive commands, such as inertial pointing and angular rates. To validate this method, NN maneuver autonomy is applied to a finite burn model of the demonstration mission. The use of sequential, mission specific maneuvers provide an appropriate testbed to demonstrate the robustness of a NN trained on feasible perturbed states. Moreover, these scenarios provide preliminary proof-of-concept for fully autonomous missions that execute maneuvers without dependence upon explicit command uplinks. As a result, the technological advancement proposed in this work may significantly ease the strain on ground-based mission operations. This would enable complex and autonomous mission execution in cislunar and deep space regimes, filling a technology gap required to support future manned missions.

NASA↗

Cognitive Performance in ISS Astronauts on 6-Month Low Earth Orbit Missions

Introduction: Current and future astronauts will endure prolonged exposure to spaceflight hazards and environmental stressors that could compromise cognitive functioning, yet cognitive performance in current missions to the International Space Station remains critically under-characterized. We systematically assessed cognitive performance across 10 cognitive domains in astronauts on 6-month missions to the ISS. Methods: Twenty-five professional astronauts were administered the Cognition Battery as part of National Aeronautics and Space Administration (NASA) Human Research Program Standard Measures Cross-Cutting Project. Cognitive performance data were collected at five mission phases: pre-flight, early flight, late flight, early post-flight, and late post-flight. We calculated speed and accuracy scores, corrected for practice effects, and derived z-scores to represent deviations in cognitive performance across mission phases from the sample’s mean baseline (i.e., pre-flight) performance. Linear mixed models with random subject intercepts and pairwise comparisons examined the relationships between mission phase and cognitive performance. Results: Cognitive performance was generally stable over time with some differences observed across mission phases for specific subtests. There was slowed performance observed in early flight on tasks of processing speed, visual working memory, and sustained attention. We observed a decrease in risk-taking propensity during late flight and post-flight mission phases. Beyond examining group differences, we inspected scores that represented a significant shift from the sample’s mean baseline score, revealing that 11.8% of all flight and post-flight scores were at or below 1.5 standard deviations below the sample’s baseline mean. Finally, exploratory analyses yielded no clear pattern of associations between cognitive performance and either sleep or ratings of alertness. Conclusions: There was no evidence for a systematic decline in cognitive performance for astronauts on a 6-month missions to the ISS. Some differences were observed for specific subtests at specific mission phases, suggesting that processing speed, visual working memory, sustained attention, and risk-taking propensity may be the cognitive domains most susceptible to change in Low Earth Orbit for high performing, professional astronauts. We provide descriptive statistics of pre-flight cognitive performance from 25 astronauts, the largest published preliminary normative database of its kind to date, to help identify significant performance decrements in future samples.

Data curation↗

NASA GeneLab: Open Science for Life in Space

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 350 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab Sequencing Lab. The GLDS contains rich metadata about each experiment and has integrated radiation dosimetry data from experiments flown on the Space Shuttle, International Space Station, and Free Flying spacecrafts. With the increasing amount and complexity of omics data being generated, GeneLab utilizes community-defined, common models for metadata and terminology so that omics data and results are discoverable and reliably reproducible. GeneLab uses the ISA-Tab specification and semantic model for organizing and representing omics metadata. In addition to metadata standards, data files must be open-source file or common exchange formats to ensure accessibility and usability by all users. To ease data ingestion and transfer, the web-based submission tool allows PIs a user-friendly user interface to curate, organize, and publish their space relevant omics data. In the more recent years, data curation and submission portal has incorporated the FAIR principles making data findable, accessible, interoperable, and reusable. To increase reusability of data, GeneLab has implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 200 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. To train the next generation of scientists, NASA offers training programs such as GeneLab 4 High School (GL4HS) and GeneLab 4 Universities. NLM Curation at a Scale Workshop 2022 | NASA GeneLab (GL4U) to teach students bioinformatics and computational biology methods to analyze omics data. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

GeneLab↗

The Colorado East River Community Observatory Data Collection

Abstract The U.S. Department of Energy's (DOE) Colorado East River Community Observatory (ER) in the Upper Colorado River Basin was established in 2015 as a representative mountainous, snow‐dominated watershed to study hydrobiogeochemical responses to hydrological perturbations in headwater systems. The ER is characterized by steep elevation, geologic, hydrologic and vegetation gradients along floodplain, montane, subalpine, and alpine life zones, which makes it an ideal location for researchers to understand how different mountain subsystems contribute to overall watershed behaviour. The ER has both long‐term and spatially‐extensive observations and experimental campaigns carried out by the Watershed Function Scientific Focus Area (SFA), led by Lawrence Berkeley National Laboratory, and researchers from over 30 organizations who conduct cross‐disciplinary process‐based investigations and modelling of watershed behaviour. The heterogeneous data generated at the ER include hydrological, genomic, biogeochemical, climate, vegetation, geological, and remote sensing data, which combined with model inputs and outputs comprise a collection of datasets and value‐added products within a mountainous watershed that span multiple spatiotemporal scales, compartments, and life zones. Within 5 years of collection, these datasets have revealed insights into numerous aspects of watershed function such as factors influencing snow accumulation and melt timing, water balance partitioning, and impacts of floodplain biogeochemistry and hillslope ecohydrology on riverine geochemical exports. Data generated by the SFA are managed and curated through its Data Management Framework. The SFA has an open data policy, and over 70 ER datasets are publicly available through relevant data repositories. A public interactive map of data collection sites run by the SFA is available to inform the broader community about SFA field activities. Here, we describe the ER and the SFA measurement network, present the public data collection generated by the SFA and partner institutions, and highlight the value of collecting multidisciplinary multiscale measurements in representative catchment observatories.

54 ENVIRONMENTAL SCIENCES↗

A Proposed Geospatial Data Preservation Strategy for DOE's Office of Legacy Management

Identifying and planning preservation and curation activities associated with geospatial data will improve the ability of the U.S. Department of Energy Office of Legacy Management (LM) to support their core mission of protecting human health and the environment. This report documents the development LM's strategy for preserving and curating geospatial data within the context of LM's data-lifecycle-management framework. The strategy consists of preservation and curation elements, specific activities, and key enabling factors that ensure LM's geospatial data is maintained. Preservation elements enable the effective preservation of LM's geospatial data and recognizes that strategies need to be flexible to adapt to ongoing changes in scale, technology, and standards. Key enabling factors are intended to highlight critical data management responsibilities that must be addressed by LM to meet its preservation and curation objectives. A summary of best practices for geospatial data preservation is provided as part of the strategy.

97 MATHEMATICS AND COMPUTING↗