Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

SREDD: Framework to facilitate MLopS related to discovery of commercial satellite data

SREDD or Super Resolution Event Detection Dashboard is a comprehensive data discovery platform designed to facilitate the identification of multiple events. It is equipped to handle data of varying spatial resolutions and from various vendors, acquired by NASA’s Commercial Smallsat Data Acquisition (CSDA) Program, making it a convenient centralized hub for searching event-related information.

Sujit Roy

Causal discovery from data assisted by large language models

Knowledge-driven discovery of novel materials necessitates the development of causal models for property emergence. While in the classical physical paradigm, the causal relationships are deduced based on physical principles or via experiment, the rapid accumulation of observational data necessitates learning causal relationships between dissimilar aspects of material structure and functionalities based on observations. For this, it is essential to integrate experimental data with prior domain knowledge. Here, we demonstrate this approach by combining high-resolution scanning transmission electron microscopy data with insights derived from large language models (LLMs). By applying ChatGPT to domain-specific literature, such as arXiv papers on ferroelectrics, and combining the obtained information with data-driven causal discovery, we construct adjacency matrices for directed acyclic graphs that map the causal relationships between structural, chemical, and polarization degrees of freedom in Sm-doped BiFeO 3 . This approach enables us to hypothesize how synthesis conditions influence material properties and guides experimental validation. Furthermore, the ultimate objective of this work is to develop a unified framework that integrates LLM-driven literature analysis with data-driven discovery, facilitating the precise engineering of ferroelectric materials by establishing clear connections between synthesis conditions and their resulting material properties.

Causal inference

Bringing Research to New Heights: How CASEI Integrates Data Curation, Discovery, and Education in Earth and Atmospheric Science

A challenging aspect of any project is finding all the relevant data and information needed to address the research objective. Searching for data and its contextual metadata can become overwhelming for both undergraduate and graduate students, potentially hindering their work and affecting the scientific discoveries that could be made in the long run. To ease this, the NASA Airborne Data Management Group (ADMG), part of the Interagency Implementation and Advanced Concepts Team (IMPACT), has developed the new Catalog of Archived Suborbital Earth science Investigations (CASEI). CASEI includes a web portal that users, be they professionals or students, can use to search, browse, discover, and locate relevant observations associated with NASA’s airborne and field campaigns. Users are able to query data in a variety of ways (via keywords, locations, timeframe, etc) from one online portal, minimizing the amount of time needed to search. CASEI also allows access to key contextual metadata and data from a wide array of Earth and Atmospheric Science topics such as aerosols and boundary layer processes, as well as ice and glacial properties or processes. Users are able to access the data via DOI links to data set landing pages. This presentation will demonstrate how CASEI can be used for classwork and student research. Teachers can provide CASEI to their students as a tool for their studies, or use it to find data themselves while constructing their curriculums. Additionally, users can leverage CASEI to learn about NASA’s Earth and Atmospheric Science research efforts and to find data relevant for assignments or other research projects. The metadata in CASEI has been carefully curated, and highlights important information about the campaigns and their data. Students can explore and learn about the scientific objectives of the campaigns, as well as descriptions of the campaign’s best research days. Having access to contextual metadata in an easy to understand way can help plant the seeds of new ideas in students at any point in their academic journey. From class projects to theses/dissertations and other research, CASEI is a valuable emerging tool for data discovery, giving access to all users and guiding researchers to NASA’s unique airborne data to answer the burning Earth Science questions of our time.

education

Data Tips: Learn How to Discover, Access and Analyze NASA GES DISC Data

At the NASA Goddard Earth Sciences Data and Information Services (GES DISC), we strive to simplify data discovery and data access to our wide range of global climate data, concentrated primarily in the areas of atmospheric composition, atmospheric dynamics, global precipitation, solar irradiance, and several modeling data sets related to land surface hydrology. To help meet user needs, we will demonstrate how you can use the GES DISC knowledge-base resources (HowTo's) and we also encourage community contributions.

data tips

Data Democratization: Challenges and Opportunities

Democratizing Earth data is one of the challenges many organizations around the world face in order to maximize the use of their Earth data for research, applications, education, and societal benefits. For example, at the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC), over 1600 global and regional datasets in several NASA Earth science focus areas, including atmospheric composition, water and energy cycles, and climate variability, are archived and distributed to the public. Giovanni, the Geospatial Interactive Online Visualization and Analysis Infrastructure, was developed by GES DISC to facilitate data access and exploration, especially for novice users of Earth science. With Giovanni, users can analyze and visualize over 2000 Earth science variables (e.g., precipitation, aerosol, surface wind) without downloading data, software, the expert understanding of data formats and structures, and coding skills, lowering the barrier to data analysis/comparison by preprocessing and accessing to the data. Results of data analysis and visualization can be accessed in several popular formats (e.g., NetCDF, CSV). As a result of Giovanni's efforts, more than 3000 referral papers have been published in various fields. In spite of this, Giovanni is still difficult to use for some users. For instance, if one searches for "precipitation," it will return over 150 related variables. The question is, which one to use? Furthermore, variables from different data providers (e.g., satellites and models) are named differently with different units, further confusing users, especially those outside the communities. Data democratization is complex and multifaceted. Challenges include service and data discovery, user experiences, visualization, data quality, trustworthiness, and more. In this presentation, we will examine Giovanni as an example of challenges and opportunities in developing data democratization services.

data democratization

An Update on the CDDIS

The Crustal Dynamics Data Inforn1ation System (CoorS) supports data archiving and distribution activities for the space geodesy and geodynamics community. The main objectives of the system are to store space geodesy and geodynamics related data products in a central data bank, to maintain infom1ation about the archival of these data, and to disseminate these data and information in a timely mam1er to a global scientific research community. The archive consists of GNSS, laser ranging, VLBI, and OORIS data sets and products derived from these data. The coors is one of NASA's Earth Observing System Oata and Infom1ation System (EOSorS) distributed data centers; EOSOIS data centers serve a diverse user community and are tasked to provide facilities to search and access science data and products. The coors data system and its archive have become increasingly important to many national and international science communities, in pal1icular several of the operational services within the International Association of Geodesy (lAG) and its project the Global Geodetic Observing System (GGOS), including the International OORIS Service (IDS), the International GNSS Service (IGS), the International Laser Ranging Service (ILRS), the International VLBI Service for Geodesy and Astrometry (IVS), and the International Earth Rotation Service (IERS). The coors has recently expanded its archive to supp011 the IGS Multi-GNSS Experiment (MGEX). The archive now contains daily and hourly 3D-second and subhourly I-second data from an additional 35+ stations in RINEX V3 fOm1at. The coors will soon install an Ntrip broadcast relay to support the activities of the IGS Real-Time Pilot Project (RTPP) and the future Real-Time IGS Service. The coors has also developed a new web-based application to aid users in data discovery, both within the current community and beyond. To enable this data discovery application, the CDDIS is currently implementing modifications to the metadata extracted from incoming data and product files pushed to its archive. This poster will include background information about the system and its user communities, archive contents and updates, enhancements for data discovery, new system architecture, and future plans.

Noll, Carey

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777

From Science to e-Science to Semantic e-Science: A Heliosphysics Case Study

The past few years have witnessed unparalleled efforts to make scientific data web accessible. The Semantic Web has proven invaluable in this effort; however, much of the literature is devoted to system design, ontology creation, and trials and tribulations of current technologies. In order to fully develop the nascent field of Semantic e-Science we must also evaluate systems in real-world settings. We describe a case study within the field of Heliophysics and provide a comparison of the evolutionary stages of data discovery, from manual to semantically enable. We describe the socio-technical implications of moving toward automated and intelligent data discovery. In doing so, we highlight how this process enhances what is currently being done manually in various scientific disciplines. Our case study illustrates that Semantic e-Science is more than just semantic search. The integration of search with web services, relational databases, and other cyberinfrastructure is a central tenet of our case study and one that we believe has applicability as a generalized research area within Semantic e-Science. This case study illustrates a specific example of the benefits, and limitations, of semantically replicating data discovery. We show examples of significant reductions in time and effort enable by Semantic e-Science; yet, we argue that a "complete" solution requires integrating semantic search with other research areas such as data provenance and web services.

Narock, Thomas

Proto-Examples of Data Access and Visualization Components of a Potential Cloud-Based GEOSS-AI System

Once a research or application problem has been identified, one logical next step is to search for available relevant data products. Thus, an early component of a potential GEOSS-AI system, in the continuum between observations and end point research, applications, and decision making, would be one that enables transparent data discovery and access by users. Such a component might be effected via the systems data agents. Presumably, some kind of data cataloging has already been implemented, e.g., in the GEOSS Common Infrastructure (GCI). Both the agents and cataloging could also leverage existing resources external to the system. The system would have some means to accept and integrate user-contributed agents. The need or desirability for some data format internal to the system should be evaluated. Another early component would be one that facilitates browsing visualization of the data, as well as some basic analyses.Three ongoing projects at the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) provide possible proto-examples of potential data access and visualization components of a cloud-based GEOSS-AI system. 1. Reorganizing data archived as time-step arrays to point-time series (data rods), as well as leveraging the NASA Simple Subset Wizard (SSW), to significantly increase the number of data products available, at multiple NASA data centers, for production as on-the-fly (virtual) data rods. SSWs data discovery is based on OpenSearch. Both pre-generated and virtual data rods are accessible via Web services. 2. Developing Web Feature Services to publish the metadata, and expose the locations, of pre-generated and virtual data rods in the GEOSS Portal and enable direct access of the data via Web services. SSW is also leveraged to increase the availability of both NASA and non-NASA data.3.Federating NASA Giovanni (Geospatial Interactive Online Visualization and Analysis Interface), for multi-sensor data exploration, that would allow each cooperating data center, currently the NASA Distributed Active Archive Centers (DAACs), to configure its own Giovanni deployment, while also allowing all the deployments to incorporate each others data. A federated Giovanni comprises Giovanni Virtual Machines, which can be run on local servers or in the cloud.

access

Increasing Discovery and Usability of Earth Science Satellite Data with My NASA Data

For 20 years, the My NASA Data project at NASA Langley Research Center has developed innovative approaches to increase the use of NASA’s satellite data by learners. My NASA Data offers a variety of authentic Earth Science datasets and a data visualization tool, eliminating the need for educators and/or learners to obtain specialized knowledge of GIS data formats and software to access and use authentic Earth Science data. While there is no shortage of available data, as federal government agencies such as NASA house petabytes of freely accessible Earth Science datasets, much of the data are only available for download and visualization in specialized formats and software, limiting their accessibility to educators and learners, especially those in primary and secondary school. Using the Google Earth Engine platform, the My NASA Data team has recently reinvented their data visualization tool, called the Earth System Data Explorer (ESDE). The ESDE gives users the capability to explore over 60 Earth Science satellite datasets in a multitude of formats such as maps, graphs, and data table Its new and improved user interface design was developed based on the preferences of educators, whom the My NASA Data project has over 20 years’ experience working with. Earth Science and GIS Subject Matter Experts (SMEs) structured the data in a professional and scientific manner. During Fiscal Year 2023, the My NASA Data website received over 1 million digital engagements, with over one-third being visitors to the data visualization tool. These metrics highlight the interest in a visualization tool that is simple and free to use with reliable and trusted datasets. The ESDE empowers users to readily relate and analyze NASA Earth Science data within their area of interest. The team used a user-centered design (UCD) framework to receive and incorporate feedback into the application’s design. Core requested features include the ability to create time series graphs, comparative analysis of maps, and download the data as CSV file. Responses indicate that advances in data visualization tools such as the ESDE make authentic Earth Science data more accessible. This presentation will cover how the My NASA Data project develops tools to enhance data discovery and accessibility, as well as how SME and user suggestions are incorporated.

Desiray Wilson

Easing the Discovery of NASA and International Near-Real-Time Data Using the Global Change Master Directory

The Global Change Master Directory (GCMD) provides an extensive directory of descriptive and spatial information about data sets and data-related services, which are relevant to Earth science research. The directory's data discovery components include controlled keywords, free-text searches, and map/date searches. The GCMD portal for NASA's Land Atmosphere Near-real-time Capability for EOS (LANCE) data products leverages these discovery features by providing users a direct route to NASA's Near-Real-Time (NRT) collections. This portal offers direct access to collection entries by instrument name, informing users of the availability of data. After a relevant collection entry is found through the GCMD's search components, the "Get Data" URL within the entry directs the user to the desired data. http://gcmd.nasa.gov/r/p/gcmd_lance_nrt.

Olsen, Lola

A Relevancy Algorithm for Curating Earth Science Data Around Phenomenon

Earth science data are being collected for various science needs and applications, processed using different algorithms at multiple resolutions and coverages, and then archived at different archiving centers for distribution and stewardship causing difficulty in data discovery. Curation, which typically occurs in museums, art galleries, and libraries, is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest. Curating data sets around topics or areas of interest addresses some of the data discovery needs in the field of Earth science, especially for unanticipated users of data. This paper describes a methodology to automate search and selection of data around specific phenomena. Different components of the methodology including the assumptions, the process, and the relevancy ranking algorithm are described. The paper makes two unique contributions to improving data search and discovery capabilities. First, the paper describes a novel methodology developed for automatically curating data around a topic using Earthscience metadata records. Second, the methodology has been implemented as a standalone web service that is utilized to augment search and usability of data in a variety of tools.

earth science phenomena

Faults Discovery By Using Mined Data

Fault discovery in the complex systems consist of model based reasoning, fault tree analysis, rule based inference methods, and other approaches. Model based reasoning builds models for the systems either by mathematic formulations or by experiment model. Fault Tree Analysis shows the possible causes of a system malfunction by enumerating the suspect components and their respective failure modes that may have induced the problem. The rule based inference build the model based on the expert knowledge. Those models and methods have one thing in common; they have presumed some prior-conditions. Complex systems often use fault trees to analyze the faults. Fault diagnosis, when error occurs, is performed by engineers and analysts performing extensive examination of all data gathered during the mission. International Space Station (ISS) control center operates on the data feedback from the system and decisions are made based on threshold values by using fault trees. Since those decision-making tasks are safety critical and must be done promptly, the engineers who manually analyze the data are facing time challenge. To automate this process, this paper present an approach that uses decision trees to discover fault from data in real-time and capture the contents of fault trees as the initial state of the trees.

Lee, Charles

Enhancing Discoverability and Management of Atmospheric Data at Scale: Solutions from the ARM Data Center

The Atmospheric Radiation Measurement (ARM) is a multi-laboratory and multi-institutional U.S. Department of Energy (DOE) Office of Science National User Facility. The ARM Data Center (ADC), located at Oak Ridge National Laboratory, collects, archives, and shares vast atmospheric data crucial for climate research. The ADC manages over 7 PB of data from 460 instruments worldwide, processing it into more than 11,000 diverse data products using the Network Common Data Form (NetCDF) for machine-independent accessibility. The primary challenge addressed in this paper is the efficient management and distribution of vast and diverse datasets essential for the climate research community, enhancing accessibility through advanced tools like Data Discovery. The ADC has developed advanced infrastructure and software architecture to handle the continuous influx of heterogeneous data to enhance data discoverability, resulting in increased scientific collaboration. In 2023, users from over 34 countries downloaded and utilized ARM data, resulting in 1,455 publications. The ADC’s efforts have significantly improved the discoverability and usability of atmospheric data, fostering extensive scientific research and collaboration. This paper details the solutions implemented by the ADC team for efficient data discovery and distribution, and it demonstrates ARM’s capability of staging processed data for scientific analysis.

Shah, Chirag [ORNL] (ORCID:0000000203145737)

From Data to Discovery: AI's Transformative Role in Thin Film Research

The advancement of thin film technologies is pivotal for progress in numerous fields, including energy, electronics, and quantum computing. However, the traditional trial-and-error approach to materials discovery is inherently slow and inefficient. This presentation will showcase how artificial intelligence (AI) is transforming thin film research by enabling a data-driven paradigm shift. We will highlight our past successes in applying AI to understand radiation damage in thin film oxides, demonstrating how graph analytics can unravel complex material behavior. Additionally, we will provide insights into our current work at the National Renewable Energy Laboratory, where we are leading the charge in autonomous materials science. Backed by a $14M investment in our characterization facility, we are developing AI-guided workflows that seamlessly integrate experimentation and AI-guided decision-making. By harnessing the power of AI, we aim to accelerate the discovery and design of high-performance thin films, propelling innovation across a multitude of industries.

36 MATERIALS SCIENCE

Hybrid Data‐Driven Discovery of High‐Performance Silver Selenide‐Based Thermoelectric Composites

Optimizing material compositions often enhances thermoelectric performances. However, the large selection of possible base elements and dopants results in a vast composition design space that is too large to systematically search using solely domain knowledge. To address this challenge, a hybrid data-driven strategy that integrates Bayesian optimization (BO) and Gaussian process regression (GPR) is proposed to optimize the composition of five elements (Ag, Se, S, Cu, and Te) in AgSe-based thermoelectric materials. Data is collected from the literature to provide prior knowledge for the initial GPR model, which is updated by actively collected experimental data during the iteration between BO and experiments. Within seven iterations, the optimized AgSe-based materials prepared using a simple high-throughput ink mixing and blade coating method deliver a high power factor of 2100 µW m −1 K −2 , which is a 75% improvement from the baseline composite (nominal composition of Ag 2 Se 1 ). In conclusion, the success of this study provides opportunities to generalize the demonstrated active machine learning technique to accelerate the development and optimization of a wide range of material systems with reduced experimental trials.

36 MATERIALS SCIENCE

Data-Driven Discovery and Experimental Validation of Solvent Polarity Effects on Conjugated Polymer Solution-to-Film Assembly Pathways

Understanding how solvent properties influence the solution-to-film assembly of conjugated polymers remains a critical challenge due to the complex and intertwined nature of polymer–solvent interactions. In this study, we integrate a data-driven framework with experimental validation to identify key parameters influencing the assembly and performance of poly[2,5-(2-octyldodecyl)-3,6-diketopyrrolopyrrole-alt-5,5-(2,5-di(thien-2-yl)thieno[3,2-b]thiophene)] (DPP-DTT) in organic field-effect transistors (OFETs). A machine learning (ML) approach identified the normalized Reichardt polarity parameter (E T N ) as a significant descriptor correlated with DPP-DTT hole mobility (μ). Systematic DPP-DTT devices fabricated using solvents across a wide E T N range revealed that higher E T N solvents yield enhanced μ. To elucidate the structural origins of high μ, we conducted comprehensive analyses using UV–vis–NIR spectroscopy and grazing incidence wide angle X-ray scattering (GIWAXS) measurements. The results revealed that films processed from high E T N solvents exhibit reduced paracrystallinity. By analyzing the solution-state behavior using optical microscopy and solution WAXS, we revealed polymer solubility differences in the various solvents and associated distinct polymer assembly pathways, elucidating why the high E T N solvent produces long-range ordered films. Notably, the high E T N solvent shows a pronounced preference for liquid-crystal (LC)-mediated assembly, providing a mechanistic explanation for the enhanced structural order. Therefore, these results demonstrate that solvent polarity, as evaluated by E T N , serves as an important parameter that plays a significant role in the DPP-DTT assembly pathway and resultant solid-state morphology. This work provides a strategy for integrating data science with experiments to identify critical parameters associated with complex polymer systems and helps guide rational process design for high-performance organic electronics.

36 MATERIALS SCIENCE