Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data sciences”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Enabling Analytics in the Cloud for Earth Science Data

The purpose of this workshop was to hold interactive discussions where providers, users, and other stakeholders could explore the convergence of three main elements in the rapidly developing world of technology: Big Data, Cloud Computing, and Analytics, [for earth science data].

Analytics↗

A Robust, Low-Cost Virtual Archive for Science Data

Despite their expense tape silos are still often the only affordable option for petabytescale science data archives, particularly when other factors such as data reliability, floor space, power and cooling load are accounted for. However, the complexity, management software, hardware reliability and access latency of tape silos make online data storage ever more attractive. Drastic reductions in low-cost mass-market PC disk drivers help to make this more affordable (approx. 1$/GB), but are challenging to scale to the petabyte range and of questionable reliability for archival use, On the other hand, if much of the science archive could be "virtualized", i.e., produced on demand when requested by users, we would need store only a fraction of the data online, perhaps bringing an online-only system into in affordable range. Radiance data from the satellite-borne Moderate Resolution Imaging Spectroradiometer (MODIS) instrument provides a good opportunity for such a virtual archive: the raw data amount to 140 GB/day, but these are small relative to the 550 GB/day making up the radiance products. These data are routinely processed as inputs for geophysical parameter products and then archived on tape at the Goddard Earth Sciences Distributed Active Archive (GES DAAC) for distributing to users. Virtualizing them would be an immediate and signifcant reduction in the amount of data being stored in the tape archives and provide more customizable products. A prototype of such a virtual archive is being developed to prove the concept and develop ways of incorporating the robustness that a science data archive requires.

Lynnes, Christopher↗

Data Science and the Knowledge Discovery Adventure

This talk will cover the important steps involved in the data science and knowledge discovery process: • Initial fact gathering (interview domain experts, review reports, articles, state-of-the-art) • Identify the problem (prediction, classification, statistical analysis, etc.) • Survey supporting data sources • Understand the data (numerical, categorical, text, sampling rate, data quality issues, etc.) • Selecting relevant features and sources • Acquire the data (set up agreements with the data stewards, APIs to download, etc.) • Merge data sources (temporal, spatial, common key, other ontologies...) • Feature Engineering (non linear domain knowledge or physics-based relationships) • Build data processing pipeline (may need to tap into data stream, develop parallel processing algorithm, federated learning etc.) • Build model and test (tune hyper-parameters, cross validation.) • Analyze/Validate results (do the results make sense. Does it answer the original question). • Deploy/Publish (Monitor and assess benefits)

Data science↗

Smarter Earth Science Data System

The explosive growth in Earth observational data in the recent decade demands a better method of interoperability across heterogeneous systems. The Earth science data system community has mastered the art in storing large volume of observational data, but it is still unclear how this traditional method scale over time as we are entering the age of Big Data. Indexed search solutions such as Apache Solr (Smiley and Pugh, 2011) provides fast, scalable search via keyword or phases without any reasoning or inference. The modern search solutions such as Googles Knowledge Graph (Singhal, 2012) and Microsoft Bing, all utilize semantic reasoning to improve its accuracy in searches. The Earth science user community is demanding for an intelligent solution to help them finding the right data for their researches. The Ontological System for Context Artifacts and Resources (OSCAR) (Huang et al., 2012), was created in response to the DARPA Adaptive Vehicle Make (AVM) programs need for an intelligent context models management system to empower its terrain simulation subsystem. The core component of OSCAR is the Environmental Context Ontology (ECO) is built using the Semantic Web for Earth and Environmental Terminology (SWEET) (Raskin and Pan, 2005). This paper presents the current data archival methodology within a NASA Earth science data centers and discuss using semantic web to improve the way we capture and serve data to our users.

data center↗

NASA Earth Sciences Data Support System and Services for the Northern Eurasia Earth Science Partnership Initiative

The presentation describes the recently awarded ACCESS project to provide data management of NASA remote sensing data for the Northern Eurasia Earth Science Partnership Initiative (NEESPI). The project targets integration of remote sensing data from MODIS, and other NASA instruments on board US-satellites (with potential expansion to data from non-US satellites), customized data products from climatology data sets (e.g., ISCCP, ISLSCP) and model data (e.g., NCEP/NCAR) into a single, well-architected data management system. It will utilize two existing components developed by the Goddard Earth Sciences Data & Information Services Center (GES DISC) at the NASA Goddard Space Flight Center: (1) online archiving and distribution system, that allows collection, processing and ingest of data from various sources into the online archive, and (2) user-friendly intelligent web-based online visualization and analysis system, also known as Giovanni. The former includes various kinds of data preparation for seamless interoperability between measurements by different instruments. The latter provides convenient access to various geophysical parameters measured in the Northern Eurasia region without any need to learn complicated remote sensing data formats, or retrieve and process large volumes of NASA data. Initial implementation of this data management system will concentrate on atmospheric data and surface data aggregated to coarse resolution to support collaborative environment and climate change studies and modeling, while at later stages, data from NASA and non-NASA satellites at higher resolution will be integrated into the system.

Leptoukh, Gregory↗

Earth Science Data Archive and Access at the NASA/Goddard Space Flight Center Distributed Active Archive Center (DAAC)

The Goddard Distributed Active Archive Center (DAAC), as an integral part of the Earth Observing System Data and Information System (EOSDIS), is the official source of data for several important earth remote sensing missions. These include the Sea-viewing Wide-Field-of-view Sensor (SeaWiFS) launched in August 1997, the Tropical Rainfall Measuring Mission (TRMM) launched in November 1997, and the Moderate Resolution Imaging Spectroradiometer (MODIS) scheduled for launch in mid 1999 as part of the EOS AM-1 instrumentation package. The data generated from these missions supports a host of users in the hydrological, land biosphere and oceanographic research and applications communities. The volume and nature of the data present unique challenges to an Earth science data archive and distribution system such as the DAAC. The DAAC system receives, archives and distributes a large number of standard data products on a daily basis, including data files that have been reprocessed with updated calibration data or improved analytical algorithms. A World Wide Web interface is provided allowing interactive data selection and automatic data subscriptions as distribution options. The DAAC also creates customized and value-added data products, which allow additional user flexibility and reduced data volume. Another significant part of our overall mission is to provide ancillary data support services and archive support for worldwide field campaigns designed to validate the results from the various satellite-derived measurements. In addition to direct data services, accompanying documentation, WWW links to related resources, support for EOSDIS data formats, and informed response to inquiries are routinely provided to users. The current GDAAC WWW search and order system is being restructured to provide users with a simplified, hierarchical access to data. Data Browsers have been developed for several data sets to aid users in ordering data. These Browsers allow users to specify spatial, temporal, and other parameter criteria in searching for and previewing data.

Leptoukh, Gregory↗

Science Data Report for the Optical Properties Monitor (OPM) Experiment

This science data report describes the Optical Properties Monitor (OPM) experiment and the data gathered during its 9-mo exposure on the Mir space station. Three independent optical instruments made up OPM: an integrating sphere spectral reflectometer, vacuum ultraviolet spectrometer, and a total integrated scatter instrument. Selected materials were exposed to the low-Earth orbit, and their performance monitored in situ by the OPM instruments. Coinvestigators from four NASA Centers, five International Space Station contractors, one university, two Department of Defense organizations, and the Russian space company, Energia, contributed samples to this experiment. These materials included a number of thermal control coatings, optical materials, polymeric films, nanocomposites, and other state-of-the-art materials. Degradation of some materials, including aluminum conversion coatings and Beta cloth, was greater than expected. The OPM experiment was launched aboard the Space Shuttle on mission STS-81 in January 1997 and transferred to the Mir space station. An extravehicular activity (EVA) was performed in April 1997 to attach the OPM experiment to the outside of the Mir/Shuttle Docking Module for space environment exposure. OPM was retrieved during an EVA in January 1998 and was returned to Earth on board the Space Shuttle on mission STS-89.

Wilkes, D. R.↗

Interplanetary space science data base and access/display tool on the NSSDC heliospheric CD-ROM

The National Space Science Data Center (NSSDC) has accumulated a rich archive of heliospheric, magnetospheric, and ionospheric data, as well as data from most other NASA-involved science disciplines. To facilitate access to and use of these data, NSSDC has begun to put selected data onto CD-ROM's. This paper describes one such CD-ROM, and the access and display software developed at NSSDC to support its use. The data on the CD-ROM consist primarily of hourly solar wind magnetic field and plasma data from many near-Earth spacecraft (OMNI) and deep space spacecraft (Voyagers, Pioneers, Helios, Pioneer Venus Orbiter). In addition, 5-minute resolution IMP-8 and ISEE-3 magnetic field and plasma data are also included. Data are stored in both ASCII and CDF formats.

Papitashvili, N. E.↗

2020 ETI Annual Summer School: Data Science and Engineering

The Consortium for Enabling Technologies & Innovation (ETI) was established in 2019 to address emerging technologies within the context of nuclear nonproliferation. ETI creates a research and education environment to support cross-cutting technologies across three core disciplines: 1) computer and engineering science research specifically in a form of machine learning and high performance computing (HPC), 2) advanced manufacturing, and 3) nuclear detection technologies. For outreach and development, ETI hosted the first of three summer schools from August 24-28, 2020 with the theme of “Data Science and Engineering”. The school was hosted in an on-line format and had over 200 participants. The recorded content is available on-line as a resource for students. The summer school had four modules: 1) Fundamentals of data Applications, 2) Computational Machine Learning, 3) Bayesian Modeling and Inference, and 4) Data Science for Safeguards. Modules contained both lectures as well as student exercises. Poll Everywhere was utilized in some modules as an on-line method to engage large groups of students. Upcoming ETI Summer Schools include Novel Instrumentation in 2021 and Advanced Manufacturing in 2022.

Biegalski, Steven R.↗

GES DISC Datalist Improves Earth Science Data Discoverability

At American Geophysical Union(AGU) 2016 Fall Meeting, Goddard Earth Sciences Data Information Services Center (GES DISC) unveiled a novel way to access data: Datalist. Currently, datalist is a collection of predefined data variables from one or more archived datasets, curated by our subject matter expert (SME). Our science support team has curated a predefined Hurricane Datalist and received very positive feedback from the user community. Datalist uses the same architecture our new website uses and have the same look and feel as other datasets on our web site. and also provides a one-stop shopping for data, metadata, citation, documentation, visualization and other available services. Since the last AGU Meeting, we have further developed a few new datalists corresponding to the Big Earth Data Initiative (BEDI) Societal Benefit Areas and A-Train data. We now have four datalists: Hurricane, Wind Energy, Greenhouse Gas and A-Train. We have also started working with our User Working Group members to create their favorite datalists and working with other DAAC to explore the possibility to include their products in our datalists that may also lead to a future of potential federated (cross-DAAC) datalists. Since our datalist prototype effort was a success, we are planning to make datalist operational. It's extremely important to have a common metadata model to support datalist, this will also be the foundation of federated datalist. We mapped our datalist metadata model to the unpublished UMM(Universal Metadata Model)-Var (Variable) (June version) and found that the UMM-var together with UMM-C (Collection) and possible UMM-S (Service) will meet our basic requirements. For example: Dataset shortname, and version are already specified in UMM-C, variable name, long name, units, dimensions are all specified in UMM-Var. UMM-Var also facilitates Science Keywords to allow tagging at variable level and Characteristics for optional variable characteristics. Measurements is useful for grouping of the variables and Set is promising to define datalist. And finally, the UMM-Service model to specify the available services for the variable will be very beneficial. In summary, UMM-Var, UMM-C and UMM-S are the basis of federated datalist and the development and deployment of datalist will contribute to the evolution of the UMM.

datalist↗

Examples of Mission-driven Data Science from Jefferson Lab and ACES

This presentation details mission-driven data science initiatives at Jefferson Lab and the Joint Institute for Advanced Computing on Environmental Studies (ACES). JLab, a U.S. Department of Energy Office of Science national laboratory, operates the Continuous Electron Beam Accelerator Facility (CEBAF), and is the lead institute for the new High Performance Data Facility (HPDF) Hub. The Joint Institute for ACES brings together interdisciplinary teams in health informatics, climate modeling, computer science, and physics to address environmental challenges, including flood modeling. The Hampton Roads region, particularly Norfolk and Virginia Beach, faces increasing flood risks, motivating the need for rapid, reliable, and risk-aware decision support. ACES’s flooding work has a focus on uncertainty quantification (UQ) and machine learning (ML) for coastal flood management. The work is motivated by the increasing vulnerability of communities such as Norfolk and Virginia Beach, Virginia, to frequent coastal flooding events, and the need for rapid, reliable decision support. The research develops computationally efficient ML surrogate models to forecast water levels and flooding risk. A central theme is the quantification and calibration of predictive uncertainty, especially for out-of-distribution (OOD) scenarios, using techniques such as Monte Carlo Dropout, Deep Ensembles, Gaussian Processes, and Deep Quantile Regression (DQR). The study demonstrates that distance-aware UQ is critical for reliable scientific AI, particularly in high-dimensional, safety-critical, and real-time applications.

McSpadden, Diana [Thomas Jefferson National Accele↗

NASA Open Science Data Repository: Maximizing Spaceflight Bioscience Data

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for data re-analysis and re-use via Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). To address the challenges posed by gaining new knowledge from a vast and diverse amount of biological, health and environmental data in space, the NASA Open Science Data Repository (OSDR - osdr.nasa.gov/bio) plays a crucial role in curating and openly publishing biological data from space-related experiments. Its design incorporates successes and lessons from NASA GeneLab, encompassing not only high-throughput sequencing data but also physiological, phenotypic, and telemetry data. The OSDR makes space biological data FAIR (findable, accessible, interoperable, reusable), and facilitates effective data ingestion, dissemination, and Open Science collaborations. The OSDR also has the capability to integrate human astronaut data with state-of-the-art security and accessibility procedures. We will discuss here several strategies that NASA’s Biological and Physical Science Division have put in place to maximize the return on investment for spaceflight bioscience data.

space biology↗

New Earth Science Data and Access Methods

NASA's Earth Science Enterprise, working with its domestic and international partners, provides scientific data and analysis to improve life here on Earth. NASA provides science data products that cover a wide range of physical, geophysical, biochemical and other parameters, as well as services for interdisciplinary Earth science studies. Management and distribution of these products is administered through the Earth Observing System Data and Information System (EOSDIS) Distributed Active Archive Centers (DAACs), which all hold data within a different Earth science discipline. This paper will highlight selected EOS datasets and will focus on how these observations contribute to the improvement of essential services such as weather forecasting, climate prediction, air quality, and agricultural efficiency. Emphasis will be placed on new data products derived from instruments on board Terra, Aqua and ICESat as well as new regional data products and field campaigns. A variety of data tools and services are available to the user community. This paper will introduce primary and specialized DAAC-specific methods for finding, ordering and using these data products. Special sections will focus on orienting users unfamiliar with DAAC resources, HDF-EOS formatted data and the use of desktop research and application tools.

Moses, John F.↗

Data Science Meets Physical Organic Chemistry

At the heart of synthetic chemistry is the holy grail of predictable catalyst design. In particular, researchers involved in reaction development in asymmetric catalysis have pursued a variety of strategies toward this goal. This is driven by both the pragmatic need to achieve high selectivities and the inability to readily identify why a certain catalyst is effective for a given reaction. While empiricism and intuition have dominated the field of asymmetric catalysis since its inception, enantioselectivity offers a mechanistically rich platform to interrogate catalyst-structure response patterns that explain the performance of a particular catalyst or substrate. In the early stages of an asymmetric reaction development campaign, the overarching mechanism of the reaction, catalyst speciation, the turnover limiting step, and many other details are unknown or posited based on related reactions. Considering the unclear details leading to a successful reaction, initial enantioselectivity data are often used to intuitively guide the ultimate direction of optimization. However, if the conditions of the Curtin-Hammett principle are satisfied, then measured enantioselectivity can be directly connected to the ensemble of diastereomeric transition states (TSs) that lead to the enantiomeric products, and the associated free energy difference between competing TSs (ΔΔ G ‡ = - RT ln[( S )/( R )], where ( S ) and ( R ) represent the concentrations of the enantiomeric products). We, and others, speculated that this important piece of information can be leveraged to guide reaction optimization in a quantitative way. Although traditional linear free energy relationships (LFERs), such as Hammett plots, have been used to illuminate important mechanistic features, we sought to develop data science derived tools to expand the power of LFERs in order to describe complex reactions frequently encountered in modern asymmetric catalysis. Specifically, we investigated whether enantioselectivity data from a reaction can be quantitatively connected to the attributes of reaction components, such as catalyst and substrate structural features, to harness data for asymmetric catalyst design. In this context, we developed a workflow to relate computationally derived features of reaction components to enantioselectivity using data science tools. The mathematical representation of molecules can incorporate many aspects of a transformation, such as molecular features from substrate, product, catalyst, and proposed transition states. Statistical models relating these features to reaction outputs can be used for various tasks, such as performance prediction of untested molecules. Perhaps most importantly, statistical models can guide the generation of mechanistic hypotheses that are embedded within complex patterns of reaction responses. Overall, merging traditional physical organic experiments with statistical modeling techniques creates a feedback loop that enables both evaluation of multiple mechanistic hypotheses and future catalyst design. In this Account, we highlight the evolution and application of this approach in the context of a collaborative program based on chiral phosphoric acid catalysts (CPAs) in asymmetric catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Leveraging Machine Learning and Geo-Tagged Citizen Science Data to Disentangle the Factors of Avian Mortality Events at the Species Level

Abrupt environmental changes can affect the population structures of living species and cause habitat loss and fragmentations in the ecosystem. During August–October 2020, remarkably high mortality events of avian species were reported across the western and central United States, likely resulting from winter storms and wildfires. However, the differences of mortality events among various species responding to the abrupt environmental changes remain poorly understood. In this study, we focused on three species, Wilson’s Warbler, Barn Owl, and Common Murre, with the highest mortality events that had been recorded by citizen scientists. We leveraged the citizen science data and multiple remotely sensed earth observations and employed the ensemble random forest models to disentangle the species responses to winter storm and wildfire. We found that the mortality events of Wilson’s Warbler were primarily impacted by early winter storms, with more deaths identified in areas with a higher average daily snow cover. The Barn Owl’s mortalities were more identified in places with severe wildfire-induced air pollution. Both winter storms and wildfire had relatively mild effects on the mortality of Common Murre, which might be more related to anomalously warm water. Our findings highlight the species-specific responses to environmental changes, which can provide significant insights into the resilience of ecosystems to environmental change and avian conservations. Additionally, the study emphasized the efficiency and effectiveness of monitoring large-scale abrupt environmental changes and conservation using remotely sensed and citizen science data.

47 OTHER INSTRUMENTATION↗

Life Sciences Data Archives (LSDA) in the Post-Shuttle Era

Now, more than ever before, NASA is realizing the value and importance of their intellectual assets. Principles of knowledge management-the systematic use and reuse of information, experience, and expertise to achieve a specific goal-are being applied throughout the agency. LSDA is also applying these solutions, which rely on a combination of content and collaboration technologies, to enable research teams to create, capture, share, and harness knowledge to do the things they do well, even better. In the early days of spaceflight, space life sciences data were collected and stored in numerous databases, formats, media-types and geographical locations. These data were largely unknown/unavailable to the research community. The Biomedical Informatics and Health Care Systems Branch of the Space Life Sciences Directorate at JSC and the Data Archive Project at ARC, with funding from the Human Research Program through the Exploration Medical Capability Element, are fulfilling these requirements through the systematic population of the Life Sciences Data Archive. This project constitutes a formal system for the acquisition, archival and distribution of data for HRP-related experiments and investigations. The general goal of the archive is to acquire, preserve, and distribute these data and be responsive to inquiries for the science communities. Information about experiments and data, as well as non-attributable human data and data from other species' are available on our public Web site http://lsda.jsc.nasa.gov. The Web site also includes a repository for biospecimens, and a utilization process. NASA has undertaken an initiative to develop a Shuttle Data Archive repository. The Shuttle program is nearing its end in 2010 and it is critical that the medical and research data related to the Shuttle program be captured, retained, and usable for research, lessons learned, and future mission planning. Communities of practice are groups of people who share a concern or a passion for something they do, and learn how to do it better as they interact regularly. LSDA works with the HRP community of practice to ensure that we are preserving the relevant research and data they need in the LSDA repository. An evidence-based approach to risk management is required in space life sciences. Evidence changes over time. LSDA has a pilot project with Collexis, a new type of Web-based search engine. Collexis differentiates itself from full-text search engines by making use of thesauri for information retrieval. The high-quality search is based on semantics that have been defined in a life sciences ontology. Additionally, Collexis' matching technology is unique, allowing discovery of partially matching dicuments. Users do not have to construct a complicated (Boolean) search query, but can simply enter a free text search without the risk of getting "no results". Collexis may address these issues by virtue of its retrieval and discovery capabilities across multiple repositories.

Fitts, Mary A.↗

Life Sciences Data Archive (LSDA) in the Post-Shuttle Era

Now, more than ever before, NASA is realizing the value and importance of their intellectual assets. Principles of knowledge management, the systematic use and reuse of information/experience/expertise to achieve a specific goal, are being applied throughout the agency. LSDA is also applying these solutions, which rely on a combination of content and collaboration technologies, to enable research teams to create, capture, share, and harness knowledge to do the things they do well, even better. In the early days of spaceflight, space life sciences data were been collected and stored in numerous databases, formats, media-types and geographical locations. These data were largely unknown/unavailable to the research community. The Biomedical Informatics and Health Care Systems Branch of the Space Life Sciences Directorate at JSC and the Data Archive Project at ARC, with funding from the Human Research Program through the Exploration Medical Capability Element, are fulfilling these requirements through the systematic population of the Life Sciences Data Archive. This project constitutes a formal system for the acquisition, archival and distribution of data for HRP-related experiments and investigations. The general goal of the archive is to acquire, preserve, and distribute these data and be responsive to inquiries from the science communities.

Fitts, Mary A.↗

Goddard Earth Science Data and Information Center (GES DISC)

The GES DIS is one of 12 NASA Earth science data centers. The GES DISC vision is to enable researchers and educators maximize knowledge of the Earth by engaging in understanding their goals, and by leading the advancement of remote sensing information services in response to satisfying their goals. This presentation will describe the GES DISC approach, successes, challenges, and best practices.

data management↗