Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Data-driven discovery of dynamics from time-resolved coherent scattering

Coherent X-ray scattering (CXS) techniques are capable of interrogating dynamics of nano- to mesoscale materials systems at time scales spanning several orders of magnitude. However, obtaining accurate theoretical descriptions of complex dynamics is often limited by one or more factors—the ability to visualize dynamics in real space, computational cost of high-fidelity simulations, and effectiveness of approximate or phenomenological models. In this work, we develop a data-driven framework to uncover mechanistic models of dynamics directly from time-resolved CXS measurements without solving the phase reconstruction problem for the entire time series of diffraction patterns. Our approach uses neural differential equations to parameterize unknown real-space dynamics and implements a computational scattering forward model to relate real-space predictions to reciprocal-space observations. This method is shown to recover the dynamics of several computational model systems under various simulated conditions of measurement resolution and noise. Moreover, the trained model enables estimation of long-term dynamics well beyond the maximum observation time, which can be used to inform and refine experimental parameters in practice. Finally, we demonstrate an experimental proof-of-concept by applying our framework to recover the probe trajectory from a ptychographic scan. Our proposed framework bridges the wide existing gap between approximate models and complex data.

36 MATERIALS SCIENCE

Data for Discovery, Characterization, and Application of Chromosomal Integration Sites for Stable Heterologous Gene Expression in Rhodotorula toruloides

Rhodotorula toruloides is a non-model, oleaginous yeast uniquely suited to produce acetyl-CoA-derived chemicals. However, the lack of well-characterized genomic integration sites has impeded the metabolic engineering of this organism. Here we report a set of computationally predicted and experimentally validated chromosomal integration sites in R. toruloides . We first implemented an in silico platform by integrating essential gene information and transcriptomic data to identify candidate sites that meet stringent criteria. We then conducted a full experimental characterization of these sites, assessing integration efficiency, gene expression levels, impact on cell growth, and long-term expression stability. Among the identified sites, 12 exhibited integration efficiencies of 50% or higher, making them sufficient for most metabolic engineering applications. Using selected high-efficiency sites, we achieved simultaneous double and triple integrations and efficiently integrated long functional pathways (up to 14.7 kb). Additionally, we developed a new inducible marker recycling system that allows multiple rounds of integration at our characterized sites. We validated this system by performing five sequential rounds of GFP integration and three sequential rounds of MaFAR integration for fatty alcohol production, demonstrating, for the first time, precise gene copy number tuning in R. toruloides . These characterized integration sites should significantly advance metabolic engineering efforts and future genetic tool development in R. toruloides .

Conversion

A high-throughput experimentation platform for data-driven discovery in electrochemistry

Automating electrochemical analyses combined with artificial intelligence is poised to accelerate discoveries in renewable energy sciences and technologies. This study presents an automated high-throughput electrochemical characterization (AHTech) platform as a cost-effective and versatile tool for rapidly assessing liquid analytes. The Python-controlled platform combines a liquid handling robot, potentiostat, and customizable microelectrode bundles for diverse, reproducible electrochemical measurements in microtiter plates, minimizing chemical consumption and manual effort. To showcase the capability of AHTech, we screened a library of 180 small molecules as electrolyte additives for aqueous zinc metal batteries, generating data for training machine learning models to predict Coulombic efficiencies. Key molecular features governing additive performance were elucidated using Shapley Additive exPlanations and Spearman’s correlation, pinpointing high-performance candidates like cis-4-hydroxy-d-proline, which achieved an average Coulombic efficiency of 99.52% over 200 cycles. The workflow established herein is highly adaptable, offering a powerful framework for accelerating the exploration and optimization of extensive chemical spaces across diverse energy storage and conversion fields.

Lin, Dian-Zhao [Johns Hopkins University, Baltimor

Data-Driven Discovery of Bimetallic Nanoparticles Catalysts for the Hydrogenolysis of Polyethylene

Supported platinum nanoparticles are known to convert polyolefins to high-quality liquid hydrocarbons with hydrogen under relatively mild conditions. However, no systematic study has been undertaken using bimetallic catalysts for polyethylene upcycling. Specifically, a total of 98 monometallic and bimetallic combinations (Ag, Cr, Co, Cu, Fe, Ga, In, Mn, Ni, Pd, Pt, Rh, Ru, Zr) on alumina were synthesized utilizing surface organometallic chemistry (SOMC) technique via robotic platform. These were investigated at a small scale (10 mg of catalyst and 50 mg of polyethylene) for their activity for the hydrogenolysis of polyethylene in a high-throughput batch reactor. Combinations of Ni and Co were selected as candidates with high activity toward conversion into paraffin oils. Reaction conditions were optimized with Ni/Co/Al 2 O 3 catalyst at a larger scale (300 mg catalyst and 3 g polyethylene) to obtain a high yield (93.1%) of paraffin wax with desired properties (M n = 380 Da) and low polydispersity (Đ = 1.2). Ni/Co/Al 2 O 3 was compared against Co/Ni/Al 2 O 3 to understand the role of the deposition sequence. When Co is deposited before Ni, a layer of cobalt aluminate is formed upon reduction, stabilizing the deposition of 5 nm metallic Ni particles. When nickel is deposited before Co, particles are larger (average >20 nm) and more oxidized (Ni δ+ in NiAl 2 O 4 ), decreasing the availability of the catalytically active metallic Ni. In conclusion, the difference in electronic environments was also described by DFT calculations, which revealed that smaller 3D clusters of Ni are preferred on CoAl2O4 over the 3D clusters on NiAl 2 O 4 and that these smaller clusters are more reducible, as confirmed experimentally.

Polymer

Adaptable Standards for Discovery, Access, and Usability of Oak Ridge National Laboratory’s Data Portals and Catalogs

Oak Ridge National Laboratory (ORNL) is leveraging its established capabilities and subject matter expertise in data curation, governance, management, national security, and risk assessment and mitigation to support the US Department of Energy (DOE) Grid Modernization Initiative. Using standards modeled by the National Institute of Standards and Technology (NIST), the Data Curation Network (DCN), the Oak Ridge Leadership Computing Facility (OLCF), and other leading organizations in the fields of energy research, high-performance computing, and national and homeland security, ORNL seeks to provide a federated approach to research data discovery, use, and interoperability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

The STARDUST Discovery Mission: Data from the Encounter with Comet Wild 2 and the Expected Sample Return

On January 2,2004, the STARDUST spacecraft made the closest ever flyby (236 km) of the nucleus of a comet - Comet Wild 2. During the fly by the spacecraft collected samples of dust from the coma of the comet. These samples will be returned to Earth on January 15,2006. After a brief preliminary examination to establish the nature of the returned samples, they will be made available to the general scientific community for study. In addition to its aerogel dust collector, the STARDUST spacecraft was also equipped with instruments that made in situ measurements of the comet during the flyby. These included several dust impact monitors, a mass spectrometer, and a camera. The spacecraft's communication system was also used to place dynamical constraints on the mass of the nucleus and the number of impacts the spacecraft had with large particles. The data taken by these instruments indicate that the spacecraft successfully captured coma samples. These instruments, particularly the camera, also demonstrated that Wild 2 is unlike any other object in the Solar System previously visited by a spacecraft. During my talk I will discuss the scientific goals of the STARDUST mission and provide an overview of its design and flight to date. I will then end with a description of the exciting data returned by the spacecraft during the recent encounter with Wild 2 and discuss what these data tell us about the nature of comets. It will probably come as no surprise that the encounter data raise as many (or more) new questions as they answer old ones.

Sandford, Scott A.

Exposing the Strategies that Can Reduce the Obstacles: Improving the Science User Experience

It is now well established that pursuing generic solutions to what seem are common problems in Earth science data access and use can often lead to disappointing results for both system developers and the intended users. This presentation focuses on real-world experience of managing a large and complex data system, NASAs Earth Science Data and Information Science System (EOSDIS), whose mission is to serve both broad user communities and those in smaller niche applications of Earth science data and services. In the talk, we focus on our experiences with known data user obstacles characterizing EOSDIS approaches, including various technological techniques, for engaging and bolstering, where possible, user experiences with EOSDIS. For improving how existing and prospective users discover and access NASA data from EOSDIS we introduce our cross-archive tool: Earthdata Search. This new search and order tool further empowers users to quickly access data sets using clever and intuitive features. The Worldview data visualization tool is also discussed highlighting how many users are now performing extensive data exploration without necessarily downloading data. Also, we explore our EOSDIS data discovery and access webinars, data recipes and short tutorials, targeted technical and data publications, user profiles and social media as additional tools and methods used for improving our outreach and communications to a diverse user community. These efforts have paid substantial dividends for our user communities by allowing us to target discipline specific community needs. The desired take-away from this presentation will be an improved understanding of how EOSDIS has approached, and in several instances achieved, removing or lowering the barriers to data access and use. As we look ahead to more complex Earth science missions, EOSDIS will continue to focus on our user communities, both broad and specialized, so that our overall data system can continue to serve the needs of science and applications users.

communications

NASA EOSDIS: Enabling Science by Improving User Knowledge

Lessons learned and impacts of applying these newer methods are explained and include several examples from our current efforts such as the interactive, on-line webinars focusing on data discovery and access including tool usage, informal and informative data chats with data experts across our EOSDIS community, data user profile interviews with scientists actively using EOSDIS data in their research, and improved conference and meeting interactions via EOSDIS data interactively used during hyper-wall talks and Worldview application. The suite of internet-based, interactive capabilities and technologies has allowed our project to expand our user community by making the data and applications from numerous Earth science missions more engaging, approachable and meaningful.

EOSDIS

ARM Metadata Entry and Data Upload Manual

The ARM Metadata Entry and Data Upload Tool, Online Metadata Editor (OME), makes it easy to describe ARM, ASR, and externally funded data products in a standardized way and enables these metadata records and uploaded data to be searchable in the ARM Data Discovery tool. The metadata records provide context for the data and facilitates the discovery and (re)use of the data.

54 ENVIRONMENTAL SCIENCES

The Future of NASA Earth Science in the Commercial Cloud: Challenges and Opportunities

NASA produces a large volume and variety of data products that are used every day to support research, decision making, and education. The widespread use of NASA’s Earth Science data is enabled by NASA’s Earth Science Data System (ESDS) program, which oversees the archiving and distribution of these data and invests in the development of new data systems and tools. However, NASA’s current approach to Earth Science data distribution — based on distributed institutional archives with individual on-premises high-performance computing capabilities — faces some significant challenges, including massive increases in data volume from upcoming missions, a greater need for transdisciplinary science that synthesizes many different kinds of observations, and a push to make science more open, inclusive, and accessible. To address these challenges, NASA is aggressively migrating its Earth Science data and related tools and services into the commercial cloud. Migration of data into the commercial cloud can significantly improve NASA’s existing data system capabilities by (1) providing more flexible options for storage and compute (including rapid, as-needed access to state-of-the-art capabilities); (2) by centralizing and standardizing data access, which gives all of NASA’s institutional data centers access to all of each other’s datasets; and (3) by facilitating “analysis-in-place”, whereby users can bring their own computational workflows and tools to the data rather than having to maintain their own copies of NASA datasets. However, migration to the commercial cloud also poses some significant challenges, including (1) managing costs under a “pay-as-you-go” model; (2) incompatibility with existing tools and data formats with object-based storage and network access; (3) vendor lock-in; (4) challenges with data access for workflows that mix on-premise and cloud computing; and (5) standardization for highly diverse data as is present in NASA’s data archive. I conclude with two examples of recent NASA activities showcasing capabilities enabled by the commercial cloud: An interactive analysis and development platform for analyzing airborne imaging spectroscopy data, and a new collection of tools and services for data discovery, analysis, publication, and data-driven storytelling (Visualization, Exploration, and Data Analysis, VEDA).

Alexey N Shiklomanov

Locating Biodiversity Data Through The Global Change Master Directory

The Global Change Master Directory (GCMD) presently holds descriptions for almost 7000 data sets held worldwide. The directory's primary purpose is for data discovery. The information provided through the GCMD's Directory Interchange Format (DIF) is the set of information that a researcher would need to determine if a particular data set could be of value. By offering data set descriptions worldwide in many scientific disciplines - including meteorology, oceanography, ecology, geology, hydrology, geophysics, remote sensing, paleoclimate, solar-terrestrial physics, and human dimensions of climate change - the GCMD simplifies the discovery of data sources. Direct linkages to many of the data sets are also provided. In addition, several data set registration tools are offered for populating the directory. To search the directory, one may choose the Guided Search or Free-Text Search. Two experimental interfaces were also made available with the latest software release - one based on a keyword search and another based on a graphical interface. The graphical interface was designed in collaboration with the Human Computer Interaction Laboratory at the University of Maryland. The latest version of the software, Version 6, was released in April, 1998. It features the implementation of a scheme to handle hierarchical data set collections (parent-child relationships); a hierarchical geospatial location search scheme; a Java-based geographic map for conducting geospatial searches; a Related-URL field for project-related data set collections, metadata extensions (such as more detailed inventory information), etc.; a new implementation of the Isite software; a new dataset language field; hyperlinked email addresses, and more. The key to the continued evolution of the GCMD is in the flexibility of the GCMD database, allowing modifications and additions to made relatively easily to maintain currency, thus providing the ability to capitalize on current technology while importing all existing records. Changes are discussed and approved through an online "interoperability" forum. The next major release of the GCMD is scheduled for early 1999 and will include the incorporation of a new matrix-based interface, a rapid valids-based query system; improvement in the operations facility - important for future distributed options; new streamlined code for greater performance and maintainability; improvements in the handling of seven current fields proposed through the interoperability forum (at no expense to the data providers); and the release of DOCmorph, a more robust version of DIFmorph to translate many 'standards' multi-directionally. Issues and actions will also be addressed.

Olsen, Lola M.

Use of Schema on Read in Earth Science Data Archives

Traditionally, NASA Earth Science data archives have file-based storage using proprietary data file formats, such as HDF and HDF-EOS, which are optimized to support fast and efficient storage of spaceborne and model data as they are generated. The use of file-based storage essentially imposes an indexing strategy based on data dimensions. In most cases, NASA Earth Science data uses time as the primary index, leading to poor performance in accessing data in spatial dimensions. For example, producing a time series for a single spatial grid cell involves accessing a large number of data files. With exponential growth in data volume due to the ever-increasing spatial and temporal resolution of the data, using file-based archives poses significant performance and cost barriers to data discovery and access. Storing and disseminating data in proprietary data formats imposes an additional access barrier for users outside the mainstream research community. At the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), we have evaluated applying the schema-on-read principle to data access and distribution. We used Apache Parquet to store geospatial data, and have exposed data through Amazon Web Services (AWS) Athena, AWS Simple Storage Service (S3), and Apache Spark. Using the schema-on-read approach allows customization of indexing spatially or temporally to suit the data access pattern. The storage of data in open formats such as Apache Parquet has widespread support in popular programming languages. A wide range of solutions for handling big data lowers the access barrier for all users. This presentation will discuss formats used for data storage, frameworks with This presentation will discuss formats used for data storage, frameworks with support for schema-on-read used for data access, and common use cases covering data usage patterns seen in a geospatial data archive.

cloud applications

An Integrated Data Analytics Platform

An Integrated Science Data Analytics Platform is an environment that enables the confluence of resources for scientific investigation. It harmonizes data, tools and computational resources which subsequently enable the research community to focus on the investigation rather than spending time on security, data preparation, management, etc. OceanWorks is a NASA technology integration project to establish a cloud-based Integrated Ocean Science Data Analytics Platform at NASA’s Physical Oceanography Distributed Active Archive Center (PO.DAAC) for big ocean science. It focuses on advancement and maturity by bringing together several NASA open-source, big data projects for parallel analytics, anomaly detection, in-situ to satellite data matchup, quality-screened data subsetting, search relevancy, and data discovery. Our communities are relying on data distributed through data centers such as the PO.DAAC, COAPS, NCAR, and many others to conduct their research. In typical investigations, scientists would engage in: search for data, evaluate the relevance of that data, download it, and then apply algorithms to identify trends. Such workflow cannot scale if the research involves a massive amount of data or multi-variate measurements. NASA’s Surface Water and Ocean Topography (SWOT) mission is expected to produce massive amount of observational data during its 3-year nominal mission. Collections like SWOT challenges all existing Earth Science data archival, distribution and analysis paradigms. In this paper, we will discuss how OceanWorks enhances the analysis of physical ocean data where the computation is done on an elastic cloud platform next to the archive to deliver fast, web-accessible services for working with oceanographic measurements.

Yang, Chaowei

Bridging the Gap: Enhancing Prominence and Provenance of NASA Datasets in Research Publications

Attribution of datasets that were used to generate research results described in peer-reviewed publications to the original source of these datasets (which are often archived at NASA Earth Science data centers) has been very challenging. Even though the data citation standard of citing datasets as research artifacts and citing them with Digital Object Identifiers (DOIs) was introduced over a decade ago, most authors do not properly reference the data used in their studies and merely mention them in the text. The lack of proper citations of datasets makes the peer-reviewed publication less transparent, imperils reproducibility, and impedes open science. We offer an open-source publication management methodology and a tool that can help to enhance usage-based data discovery, prominence, and provenance of the data; reproducibility of the research results; and potentially increase the return on investment on NASA-funded research.

open-source