Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

GES DISC Datalist Improves Earth Science Data Discoverability

At American Geophysical Union(AGU) 2016 Fall Meeting, Goddard Earth Sciences Data Information Services Center (GES DISC) unveiled a novel way to access data: Datalist. Currently, datalist is a collection of predefined data variables from one or more archived datasets, curated by our subject matter expert (SME). Our science support team has curated a predefined Hurricane Datalist and received very positive feedback from the user community. Datalist uses the same architecture our new website uses and have the same look and feel as other datasets on our web site. and also provides a one-stop shopping for data, metadata, citation, documentation, visualization and other available services. Since the last AGU Meeting, we have further developed a few new datalists corresponding to the Big Earth Data Initiative (BEDI) Societal Benefit Areas and A-Train data. We now have four datalists: Hurricane, Wind Energy, Greenhouse Gas and A-Train. We have also started working with our User Working Group members to create their favorite datalists and working with other DAAC to explore the possibility to include their products in our datalists that may also lead to a future of potential federated (cross-DAAC) datalists. Since our datalist prototype effort was a success, we are planning to make datalist operational. It's extremely important to have a common metadata model to support datalist, this will also be the foundation of federated datalist. We mapped our datalist metadata model to the unpublished UMM(Universal Metadata Model)-Var (Variable) (June version) and found that the UMM-var together with UMM-C (Collection) and possible UMM-S (Service) will meet our basic requirements. For example: Dataset shortname, and version are already specified in UMM-C, variable name, long name, units, dimensions are all specified in UMM-Var. UMM-Var also facilitates Science Keywords to allow tagging at variable level and Characteristics for optional variable characteristics. Measurements is useful for grouping of the variables and Set is promising to define datalist. And finally, the UMM-Service model to specify the available services for the variable will be very beneficial. In summary, UMM-Var, UMM-C and UMM-S are the basis of federated datalist and the development and deployment of datalist will contribute to the evolution of the UMM.

datalist↗

Examples of Mission-driven Data Science from Jefferson Lab and ACES

This presentation details mission-driven data science initiatives at Jefferson Lab and the Joint Institute for Advanced Computing on Environmental Studies (ACES). JLab, a U.S. Department of Energy Office of Science national laboratory, operates the Continuous Electron Beam Accelerator Facility (CEBAF), and is the lead institute for the new High Performance Data Facility (HPDF) Hub. The Joint Institute for ACES brings together interdisciplinary teams in health informatics, climate modeling, computer science, and physics to address environmental challenges, including flood modeling. The Hampton Roads region, particularly Norfolk and Virginia Beach, faces increasing flood risks, motivating the need for rapid, reliable, and risk-aware decision support. ACES’s flooding work has a focus on uncertainty quantification (UQ) and machine learning (ML) for coastal flood management. The work is motivated by the increasing vulnerability of communities such as Norfolk and Virginia Beach, Virginia, to frequent coastal flooding events, and the need for rapid, reliable decision support. The research develops computationally efficient ML surrogate models to forecast water levels and flooding risk. A central theme is the quantification and calibration of predictive uncertainty, especially for out-of-distribution (OOD) scenarios, using techniques such as Monte Carlo Dropout, Deep Ensembles, Gaussian Processes, and Deep Quantile Regression (DQR). The study demonstrates that distance-aware UQ is critical for reliable scientific AI, particularly in high-dimensional, safety-critical, and real-time applications.

McSpadden, Diana [Thomas Jefferson National Accele↗

NASA Open Science Data Repository: Maximizing Spaceflight Bioscience Data

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for data re-analysis and re-use via Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). To address the challenges posed by gaining new knowledge from a vast and diverse amount of biological, health and environmental data in space, the NASA Open Science Data Repository (OSDR - osdr.nasa.gov/bio) plays a crucial role in curating and openly publishing biological data from space-related experiments. Its design incorporates successes and lessons from NASA GeneLab, encompassing not only high-throughput sequencing data but also physiological, phenotypic, and telemetry data. The OSDR makes space biological data FAIR (findable, accessible, interoperable, reusable), and facilitates effective data ingestion, dissemination, and Open Science collaborations. The OSDR also has the capability to integrate human astronaut data with state-of-the-art security and accessibility procedures. We will discuss here several strategies that NASA’s Biological and Physical Science Division have put in place to maximize the return on investment for spaceflight bioscience data.

space biology↗

New Earth Science Data and Access Methods

NASA's Earth Science Enterprise, working with its domestic and international partners, provides scientific data and analysis to improve life here on Earth. NASA provides science data products that cover a wide range of physical, geophysical, biochemical and other parameters, as well as services for interdisciplinary Earth science studies. Management and distribution of these products is administered through the Earth Observing System Data and Information System (EOSDIS) Distributed Active Archive Centers (DAACs), which all hold data within a different Earth science discipline. This paper will highlight selected EOS datasets and will focus on how these observations contribute to the improvement of essential services such as weather forecasting, climate prediction, air quality, and agricultural efficiency. Emphasis will be placed on new data products derived from instruments on board Terra, Aqua and ICESat as well as new regional data products and field campaigns. A variety of data tools and services are available to the user community. This paper will introduce primary and specialized DAAC-specific methods for finding, ordering and using these data products. Special sections will focus on orienting users unfamiliar with DAAC resources, HDF-EOS formatted data and the use of desktop research and application tools.

Moses, John F.↗

Data Science Meets Physical Organic Chemistry

At the heart of synthetic chemistry is the holy grail of predictable catalyst design. In particular, researchers involved in reaction development in asymmetric catalysis have pursued a variety of strategies toward this goal. This is driven by both the pragmatic need to achieve high selectivities and the inability to readily identify why a certain catalyst is effective for a given reaction. While empiricism and intuition have dominated the field of asymmetric catalysis since its inception, enantioselectivity offers a mechanistically rich platform to interrogate catalyst-structure response patterns that explain the performance of a particular catalyst or substrate. In the early stages of an asymmetric reaction development campaign, the overarching mechanism of the reaction, catalyst speciation, the turnover limiting step, and many other details are unknown or posited based on related reactions. Considering the unclear details leading to a successful reaction, initial enantioselectivity data are often used to intuitively guide the ultimate direction of optimization. However, if the conditions of the Curtin-Hammett principle are satisfied, then measured enantioselectivity can be directly connected to the ensemble of diastereomeric transition states (TSs) that lead to the enantiomeric products, and the associated free energy difference between competing TSs (ΔΔ G ‡ = - RT ln[( S )/( R )], where ( S ) and ( R ) represent the concentrations of the enantiomeric products). We, and others, speculated that this important piece of information can be leveraged to guide reaction optimization in a quantitative way. Although traditional linear free energy relationships (LFERs), such as Hammett plots, have been used to illuminate important mechanistic features, we sought to develop data science derived tools to expand the power of LFERs in order to describe complex reactions frequently encountered in modern asymmetric catalysis. Specifically, we investigated whether enantioselectivity data from a reaction can be quantitatively connected to the attributes of reaction components, such as catalyst and substrate structural features, to harness data for asymmetric catalyst design. In this context, we developed a workflow to relate computationally derived features of reaction components to enantioselectivity using data science tools. The mathematical representation of molecules can incorporate many aspects of a transformation, such as molecular features from substrate, product, catalyst, and proposed transition states. Statistical models relating these features to reaction outputs can be used for various tasks, such as performance prediction of untested molecules. Perhaps most importantly, statistical models can guide the generation of mechanistic hypotheses that are embedded within complex patterns of reaction responses. Overall, merging traditional physical organic experiments with statistical modeling techniques creates a feedback loop that enables both evaluation of multiple mechanistic hypotheses and future catalyst design. In this Account, we highlight the evolution and application of this approach in the context of a collaborative program based on chiral phosphoric acid catalysts (CPAs) in asymmetric catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Leveraging Machine Learning and Geo-Tagged Citizen Science Data to Disentangle the Factors of Avian Mortality Events at the Species Level

Abrupt environmental changes can affect the population structures of living species and cause habitat loss and fragmentations in the ecosystem. During August–October 2020, remarkably high mortality events of avian species were reported across the western and central United States, likely resulting from winter storms and wildfires. However, the differences of mortality events among various species responding to the abrupt environmental changes remain poorly understood. In this study, we focused on three species, Wilson’s Warbler, Barn Owl, and Common Murre, with the highest mortality events that had been recorded by citizen scientists. We leveraged the citizen science data and multiple remotely sensed earth observations and employed the ensemble random forest models to disentangle the species responses to winter storm and wildfire. We found that the mortality events of Wilson’s Warbler were primarily impacted by early winter storms, with more deaths identified in areas with a higher average daily snow cover. The Barn Owl’s mortalities were more identified in places with severe wildfire-induced air pollution. Both winter storms and wildfire had relatively mild effects on the mortality of Common Murre, which might be more related to anomalously warm water. Our findings highlight the species-specific responses to environmental changes, which can provide significant insights into the resilience of ecosystems to environmental change and avian conservations. Additionally, the study emphasized the efficiency and effectiveness of monitoring large-scale abrupt environmental changes and conservation using remotely sensed and citizen science data.

47 OTHER INSTRUMENTATION↗

Life Sciences Data Archives (LSDA) in the Post-Shuttle Era

Now, more than ever before, NASA is realizing the value and importance of their intellectual assets. Principles of knowledge management-the systematic use and reuse of information, experience, and expertise to achieve a specific goal-are being applied throughout the agency. LSDA is also applying these solutions, which rely on a combination of content and collaboration technologies, to enable research teams to create, capture, share, and harness knowledge to do the things they do well, even better. In the early days of spaceflight, space life sciences data were collected and stored in numerous databases, formats, media-types and geographical locations. These data were largely unknown/unavailable to the research community. The Biomedical Informatics and Health Care Systems Branch of the Space Life Sciences Directorate at JSC and the Data Archive Project at ARC, with funding from the Human Research Program through the Exploration Medical Capability Element, are fulfilling these requirements through the systematic population of the Life Sciences Data Archive. This project constitutes a formal system for the acquisition, archival and distribution of data for HRP-related experiments and investigations. The general goal of the archive is to acquire, preserve, and distribute these data and be responsive to inquiries for the science communities. Information about experiments and data, as well as non-attributable human data and data from other species' are available on our public Web site http://lsda.jsc.nasa.gov. The Web site also includes a repository for biospecimens, and a utilization process. NASA has undertaken an initiative to develop a Shuttle Data Archive repository. The Shuttle program is nearing its end in 2010 and it is critical that the medical and research data related to the Shuttle program be captured, retained, and usable for research, lessons learned, and future mission planning. Communities of practice are groups of people who share a concern or a passion for something they do, and learn how to do it better as they interact regularly. LSDA works with the HRP community of practice to ensure that we are preserving the relevant research and data they need in the LSDA repository. An evidence-based approach to risk management is required in space life sciences. Evidence changes over time. LSDA has a pilot project with Collexis, a new type of Web-based search engine. Collexis differentiates itself from full-text search engines by making use of thesauri for information retrieval. The high-quality search is based on semantics that have been defined in a life sciences ontology. Additionally, Collexis' matching technology is unique, allowing discovery of partially matching dicuments. Users do not have to construct a complicated (Boolean) search query, but can simply enter a free text search without the risk of getting "no results". Collexis may address these issues by virtue of its retrieval and discovery capabilities across multiple repositories.

Fitts, Mary A.↗

Life Sciences Data Archive (LSDA) in the Post-Shuttle Era

Now, more than ever before, NASA is realizing the value and importance of their intellectual assets. Principles of knowledge management, the systematic use and reuse of information/experience/expertise to achieve a specific goal, are being applied throughout the agency. LSDA is also applying these solutions, which rely on a combination of content and collaboration technologies, to enable research teams to create, capture, share, and harness knowledge to do the things they do well, even better. In the early days of spaceflight, space life sciences data were been collected and stored in numerous databases, formats, media-types and geographical locations. These data were largely unknown/unavailable to the research community. The Biomedical Informatics and Health Care Systems Branch of the Space Life Sciences Directorate at JSC and the Data Archive Project at ARC, with funding from the Human Research Program through the Exploration Medical Capability Element, are fulfilling these requirements through the systematic population of the Life Sciences Data Archive. This project constitutes a formal system for the acquisition, archival and distribution of data for HRP-related experiments and investigations. The general goal of the archive is to acquire, preserve, and distribute these data and be responsive to inquiries from the science communities.

Fitts, Mary A.↗

Goddard Earth Science Data and Information Center (GES DISC)

The GES DIS is one of 12 NASA Earth science data centers. The GES DISC vision is to enable researchers and educators maximize knowledge of the Earth by engaging in understanding their goals, and by leading the advancement of remote sensing information services in response to satisfying their goals. This presentation will describe the GES DISC approach, successes, challenges, and best practices.

data management↗

National Space Science Data Center and World Data Center A for Rockets and Satellites - Ionospheric data holdings and services

The activities and services of the National Space Science data Center (NSSDC) and the World Data Center A for Rockets and Satellites (WDC-A-R and S) are described with special emphasis on ionospheric physics. The present catalog/archive system is explained and future developments are indicated. In addition to the basic data acquisition, archiving, and dissemination functions, ongoing activities include the Central Online Data Directory (CODD), the Coordinated Data Analysis Workshopps (CDAW), the Space Physics Analysis Network (SPAN), advanced data management systems (CD/DIS, NCDS, PLDS), and publication of the NSSDC News, the SPACEWARN Bulletin, and several NSSD reports.

Bilitza, D.↗

Hacking Limnology Workshop and DSOS22: Creating a Community of Practice for the Nexus of Data Science, Open Science, and the Aquatic Sciences

The 2nd Aquatic Ecosystem Modeling-Junior (AEMON-J) Hacking Limnology Workshop and 3rd Virtual Summit: Incorporating Data Science and Open Science in the Aquatic Sciences (DSOS) took place on 25–29 July 2022. These virtual events were developed to bring together researchers from diverse backgrounds to share developments in data-intensive research in the aquatic sciences and train participants in cutting-edge data analysis methods related to remote sensing, data pipelines, and modeling of aquatic ecosystems.

54 ENVIRONMENTAL SCIENCES↗

Quantum Computing for Biomedical Computational and Data Sciences: A Joint DOE-NIH Roundtable

The overlap of quantum computing and biomedical research, while less explored, presents significant near-term opportunities. The Department of Energy (DOE) and the National Institutes of Health (NIH) are interested in exploiting the DOE community’s capabilities and expertise in quantum computing to potentially advance biomedical research, targeting fundamental studies of biological and molecular structures, understanding of human health as well as mental and physical disorders and diseases, and deriving insights from clinical data. NIH’s approach to quantum computing is guided by its Strategic Plan for Data Science, emphasizing the importance of findable, accessible, interoperable, and reusable (FAIR) data assets, security and privacy of data, and efficient computing and storage. DOE’s Office of Science (SC), and more specifically the Advanced Scientific Computing Research (ASCR) program, supports quantum information science (QIS) research, contributing to a unique portfolio of quantum computing and communications expertise. This roundtable was assembled to consider the opportunities and challenges in the near-, medium-, and long-term at the intersection of quantum computing, data science, and biomedical research and how these could be addressed through inter-agency collaboration and multi-disciplinary partnerships.

59 BASIC BIOLOGICAL SCIENCES↗

Discovery of complex oxides via automated experiments and data science

Significance Automation is accelerating the discovery of useful materials, yet testing even a small fraction of the billions of possible materials for a desired property is beyond the reach of workflows involving resource-intensive property measurements. Due to relationships among composition, structure, and properties, identifying a complex material with one interesting property makes it the proverbial needle in a haystack that merits testing for additional properties. We accelerate materials synthesis and optical characterization by employing physics-aware data science to identify materials for further investigation. With this approach, one does not need high-throughput methods for measuring every material property of interest since a single ultra-high–throughput workflow can guide material selection for other properties, which is a new paradigm for accelerated materials discovery.

36 MATERIALS SCIENCE↗

NASA'S Earth Science Data Stewardship Activities

NASA has been collecting Earth observation data for over 50 years using instruments on board satellites, aircraft and ground-based systems. With the inception of the Earth Observing System (EOS) Program in 1990, NASA established the Earth Science Data and Information System (ESDIS) Project and initiated development of the Earth Observing System Data and Information System (EOSDIS). A set of Distributed Active Archive Centers (DAACs) was established at locations based on science discipline expertise. Today, EOSDIS consists of 12 DAACs and 12 Science Investigator-led Processing Systems (SIPS), processing data from the EOS missions, as well as the Suomi National Polar Orbiting Partnership mission, and other satellite and airborne missions. The DAACs archive and distribute the vast majority of data from NASA’s Earth science missions, with data holdings exceeding 12 petabytes The data held by EOSDIS are available to all users consistent with NASA’s free and open data policy, which has been in effect since 1990. The EOSDIS archives consist of raw instrument data counts (level 0 data), as well as higher level standard products (e.g., geophysical parameters, products mapped to standard spatio-temporal grids, results of Earth system models using multi-instrument observations, and long time series of Earth System Data Records resulting from multiple satellite observations of a given type of phenomenon). EOSDIS data stewardship responsibilities include ensuring that the data and information content are reliable, of high quality, easily accessible, and usable for as long as they are considered to be of value.

metadata↗

The crush of new earth science data knocking at our door

The reasons for collecting massive amounts of earth science data in the Earth Observing System (EOS) Project are discussed. A processing hierarchy for handling the data is described, and the prospects for adequate throughput and storage for operational data analysis in the EOS era are addressed. Needs for successful exploratory data analysis are examined, and the policy issues implicated by the large stream of EOS data are considered.

Kahn, Ralph↗

Earth Sciences Data and Information System (ESDIS) program planning and evaluation methodology development

An Earth Sciences Data and Information System (ESDIS) Project Management Plan (PMP) is prepared. An ESDIS Project Systems Engineering Management Plan (SEMP) consistent with the developed PMP is also prepared. ESDIS and related EOS program requirements developments, management and analysis processes are evaluated. Opportunities to improve the effectiveness of these processes and program/project responsiveness to requirements are identified. Overall ESDIS cost estimation processes are evaluated, and recommendations to improve cost estimating and modeling techniques are developed. ESDIS schedules and scheduling tools are evaluated. Risk assessment, risk mitigation strategies and approaches, and use of risk information in management decision-making are addressed.

Dickinson, William B.↗

The NPOESS Preparatory Project Science Data Segment (SDS) Data Depository and Distribution Element (SD3E) System Architecture

The National Polar-orbiting Operational Environmental Satellite System (NPOESS) Preparatory Project (NPP) Science Data Segment (SDS) will make daily data requests for approximately six terabytes of NPP science products for each of its six environmental assessment elements from the operational data providers. As a result, issues associated with duplicate data requests, data transfers of large volumes of diverse products, and data transfer failures raised concerns with respect to the network traffic and bandwidth consumption. The NPP SDS Data Depository and Distribution Element (SD3E) was developed to provide a mechanism for efficient data exchange, alleviate duplicate network traffic, and reduce operational costs.

Ho, Evelyn L.↗