Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Information Theoretic Approaches to Rapid Discovery of Relationships in Large Climate Data Sets

Mutual information as the asymptotic Bayesian measure of independence is an excellent starting point for investigating the existence of possible relationships among climate-relevant variables in large data sets, As mutual information is a nonlinear function of of its arguments, it is not beholden to the assumption of a linear relationship between the variables in question and can reveal features missed in linear correlation analyses. However, as mutual information is symmetric in its arguments, it only has the ability to reveal the probability that two variables are related. it provides no information as to how they are related; specifically, causal interactions or a relation based on a common cause cannot be detected. For this reason we also investigate the utility of a related quantity called the transfer entropy. The transfer entropy can be written as a difference between mutual informations and has the capability to reveal whether and how the variables are causally related. The application of these information theoretic measures is rested on some familiar examples using data from the International Satellite Cloud Climatology Project (ISCCP) to identify relation between global cloud cover and other variables, including equatorial pacific sea surface temperature (SST), over seasonal and El Nino Southern Oscillation (ENSO) cycles.

Knuth, Kevin H.↗

Development of a Knowledge Graph for Dataset Discovery and Identification at a NASA Data Center

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) archives and distributes hundreds of Earth Science data collections to the public. These collections are used in research, resulting in the publication of thousands of scientific papers each year. As new users come to GES DISC for data, it is important for them to understand how prior research used the data. To help researchers, a knowledge graph (KG) was designed and implemented to connect publication citations with dataset metadata. The relationships created in the graph have the potential to allow the Web applications that utilize this information to directly connect the publication to the GES DISC datasets and services. These relationships are demonstrated using a web application prototype. In addition, the graph can also make connections between publications, datasets, and measurements based on the mentions of datasets and their attributes in the publications. To demonstrate this capability, a web application was created that takes the excerpt from the publication and returns a most likely dataset and measurement pairing, ranking the results based on how often these datasets and measurements were used in prior publications.

Nathaniel Crosby↗

G2PDeep-v2: A Web-Based Deep-Learning Framework for Phenotype Prediction and Biomarker Discovery for All Organisms Using Multi-Omics Data

Multi-omics data offers rich insights into complex traits across organisms, yet integrating and analyzing these datasets for phenotype prediction and marker discovery remains challenging. Researchers need accessible tools that combine deep learning, hyperparameter optimization, visualization, and downstream analysis in a unified web platform. To address this, we developed G2PDeep-v2, a web-based platform powered by deep learning for phenotype prediction and marker discovery from multi-omics data across a wide range of organisms, including humans and plants. The server provides multiple services for researchers to create deep-learning models through an interactive interface and train these models using an automated hyperparameter tuning algorithm on high-performance computing resources. Users can visualize the results of phenotype and markers predictions and perform Gene Set Enrichment Analysis for the significant markers to provide insights into the molecular mechanisms underlying complex diseases, conditions and other biological phenotypes being studied.

59 BASIC BIOLOGICAL SCIENCES↗

Smart Handoffs: Preserving User Context Between Tools and Services Related to NASA's EOSDIS Data Archive

NASA's Earth Observing System Data and Information System (EOSDIS) is tasked with archiving and distributing Earth Observation data across a range of disciplines, including atmospheric science, oceanography, land processes, natural hazards, solar radiance and even socioeconomic aspects relating to the environment. Given the breadth of disciplines and depth of data that EOSDIS provides, the efficient and intuitive discovery and usage of data by a scientist is of paramount importance. An effective data gathering workflow may involve switching from general use discovery tools to a more bespoke services designed specifically for the scientist's discipline. Providing concrete interoperability between such tools could vastly improve the efficiency of a scientist's workflow.

Analytics↗

QuakeSim 2.0

QuakeSim 2.0 improves understanding of earthquake processes by providing modeling tools and integrating model applications and various heterogeneous data sources within a Web services environment. QuakeSim is a multisource, synergistic, data-intensive environment for modeling the behavior of earthquake faults individually, and as part of complex interacting systems. Remotely sensed geodetic data products may be explored, compared with faults and landscape features, mined by pattern analysis applications, and integrated with models and pattern analysis applications in a rich Web-based and visualization environment. Integration of heterogeneous data products with pattern informatics tools enables efficient development of models. Federated database components and visualization tools allow rapid exploration of large datasets, while pattern informatics enables identification of subtle, but important, features in large data sets. QuakeSim is valuable for earthquake investigations and modeling in its current state, and also serves as a prototype and nucleus for broader systems under development. The framework provides access to physics-based simulation tools that model the earthquake cycle and related crustal deformation. Spaceborne GPS and Inter ferometric Synthetic Aperture (InSAR) data provide information on near-term crustal deformation, while paleoseismic geologic data provide longerterm information on earthquake fault processes. These data sources are integrated into QuakeSim's QuakeTables database system, and are accessible by users or various model applications. UAVSAR repeat pass interferometry data products are added to the QuakeTables database, and are available through a browseable map interface or Representational State Transfer (REST) interfaces. Model applications can retrieve data from Quake Tables, or from third-party GPS velocity data services; alternatively, users can manually input parameters into the models. Pattern analysis of GPS and seismicity data has proved useful for mid-term forecasting of earthquakes, and for detecting subtle changes in crustal deformation. The GPS time series analysis has also proved useful as a data-quality tool, enabling the discovery of station anomalies and data processing and distribution errors. Improved visualization tools enable more efficient data exploration and understanding. Tools provide flexibility to science users for exploring data in new ways through download links, but also facilitate standard, intuitive, and routine uses for science users and end users such as emergency responders.

Donnellan, Andrea↗

All-sky Neutrino Point-source Search with IceCube Combined Track and Cascade Data

Despite extensive efforts, discovery of high-energy astrophysical neutrino sources remains elusive. We present an event-level simultaneous maximum likelihood analysis of tracks and cascades using IceCube data collected from 2008 April 6 to 2022 May 23 to search the whole sky for neutrino sources, and using a source catalog, for coincidence of neutrino emission with gamma-ray emission. This is the first time a simultaneous fit of different detection channels is used to conduct a time-integrated all-sky scan with IceCube. Combining all-sky tracks, with superior pointing power and sensitivity in the northern sky, with all-sky cascades, with good energy resolution and sensitivity in the southern sky, we have developed the most sensitive point-source search to date by IceCube that targets the entire sky. The most significant point in the northern sky aligns with NGC 1068, a Seyfert II galaxy, which, from the catalog search, shows a 3.5σ excess over background after accounting for trials. The most significant point in the southern sky does not align with any source in the catalog and is not significant after accounting for trials. A search for the single most significant Gaussian flare at the locations of NGC 1068, PKS 1424+240, and the southern highest-significance point shows results consistent with expectations for steady emission. Notably, this is the first time that a flare shorter than four years has been excluded as being responsible for NGC 1068’s emergence as a neutrino source. Our results show that combining tracks and cascades when conducting neutrino source searches improves sensitivity and can lead to new discoveries.

Abbasi, R. [Loyola University, Chicago, IL (United↗

Opening Historical Airborne Data to Present Day Researchers

For more than 50 years, NASA has flown airborne sensors to carry out research, validate satellite sensors, and test new instrument capabilities. Data collected prior to 2000 are typically analog and difficult to locate and use. The Airborne Data Management Group (ADMG) facilitates rescue of these valuable data to ensure easier discovery, access, and use. But opening historical data comes at a cost of both time and money. Careful decisions are required in assessing the return on investment. - Is there interest in the science community? - Are there government data requirements? - What is the temporal / spatial value of the data? - Can data be transformed to a digital format? - What is cost of transformation? - What time period is needed for rescue? Converting the data to today’s digital storage standards increases value and provides data access. The addition of metadata makes the data easier to search for.

Deborah Smith↗

Privacy Preservation from High-Performance Computing to Autonomous Science [Industrial and Governmental Activities]

High-Performance Computing (HPC) and Leadership-Class Supercomputing are driving forces behind scientific advancements, enabling researchers to tackle complex challenges in physics, chemistry, biology, and engineering. These systems power vast simulations and data analyses, fueling discoveries in fields ranging from materials science to climate modeling. However, their use often involves processing sensitive data—such as proprietary industry simulations, biomedical records, and national security computations—posing significant privacy concerns. In conclusion, this issue is amplified in collaborative environments like Department of Energy (DOE) user facilities, where HPC resources are shared across institutions to foster innovation.

Kotevska, Olivera [Oak Ridge National Laboratory (↗

Discovery of Nine Gamma-Ray Pulsars in Fermi-Lat Data Using a New Blind Search Method

We report the discovery of nine previously unknown gamma-ray pulsars in a blind search of data from the Fermi Large Area Telescope (LAT). The pulsars were found with a novel hierarchical search method originally developed for detecting continuous gravitational waves from rapidly rotating neutron stars. Designed to find isolated pulsars spinning at up to kHz frequencies, the new method is computationally efficient, and incorporates several advances, including a metric-based gridding of the search parameter space (frequency, frequency derivative and sky location) and the use of photon probability weights. The nine pulsars have spin frequencies between 3 and 12 Hz, and characteristic ages ranging from 17 kyr to 3 Myr. Two of them, PSRs Jl803-2149 and J2111+4606, are young and energetic Galactic-plane pulsars (spin-down power above 6 x 10(exp 35) ergs per second and ages below 100 kyr). The seven remaining pulsars, PSRs J0106+4855, J010622+3749, Jl620-4927, Jl746-3239, J2028+3332,J2030+4415, J2139+4716, are older and less energetic; two of them are located at higher Galactic latitudes (|b| greater than 10 degrees). PSR J0106+4855 has the largest characteristic age (3 Myr) and the smallest surface magnetic field (2x 10(exp 11)G) of all LAT blind-search pulsars. PSR J2139+4716 has the lowest spin-down power (3 x l0(exp 33) erg per second) among all non-recycled gamma-ray pulsars ever found. Despite extensive multi-frequency observations, only PSR J0106+4855 has detectable pulsations in the radio band. The other eight pulsars belong to the increasing population of radio-quiet gamma-ray pulsars.

Celik-Tinmaz, Ozlem↗

CSW Best Practices

During the development of the CMR (Common Metadata Repository) (CMR) for the Earth Observing System Data and Information System (EOSDIS), CSW (Catalog Service for the Web) a number of best practices came to light. Given that the ESIP (Earth Science Information Partners) Discovery Cluster is committed to interoperability and standards in earth data discovery this seemed like a convenient moment to provide Best Practices to the organization in the same way we did for OpenSearch for this widely-used standard.

CMR↗

The NASA Open Science Data Repository: Biomedical Fair Data, Analysis Tools, User Communities, Publications, and Discoveries for Deep Space Missions

Increased biomedical risks and challenges associated with deep space missions require new knowledge discovery, new health countermeasures, and development of novel ecosystems, life support, crop production, and biomedical support capabilities. To meet NASA’s Moon to Mars strategic program goals for Human and Biological Sciences, findable, accessible, interoperable, reusable (FAIR), and maximally open-access data is going to be required to enable humanity to thrive in deep space. Indeed, this cornerstone perspective on FAIR and maximally open access data was also recommended in the recent 2023-2032 Decadal Survey from the National Academies of Sciences, Engineering, and Medicine. The NASA Open Science Data Repository (OSDR) is a maximally open access and FAIR database, and meets various scientific, technical, and operational spaceflight needs. It offers public users and submitters the ability to upload, download, search, share, analyze, and visualize data across ‘omics, physiological, phenotypic, behavioral, bioimaging, video, and environmental monitoring telemetry datasets. OSDR includes NASA GeneLab, NASA Ames Life Sciences Data Archive, and the NASA Biological Institutional Scientific Collection. OSDR has >455 studies with datasets from model organisms and non-NASA human astronauts. There are ~12 datasets from the Inspiration 4 (I4) mission, spanning metagenomics, comprehensive metabolic panels, clonal hematopoiesis, spatial transcriptomics, proteomics, and cytokine panels. In the interest of data privacy, two I4 datasets have raw FASTQ and FASTA files relating to the epitranscriptome, and a new request feature is live in OSDR (with a backend review process established) which was developed based on industry norms. OSDR also recently began a collaboration with the European Space Agency (ESA) to scientifically curate and make available >200 terabytes of human and model organism space-relevant data. The OSDR submission portal is designed to ingest and curate ~25 ‘omics assay data types, and ~50 physiological-phenotypic-imaging assay data types, spanning ultrasonography, micro-computed tomography, histology, morphometric photography, rebound tonometry, gait analysis, optical coherence tomography, novel object recognition, flow cytometry, and immunohistochemistry. A suite of analysis tools are available for OSDR users including: 1) an Environmental Data Application to compare radiation, CO2, relative humidity, temperature, and other telemetry across missions and subjects, 2) the RadLab database, a collaboration between NASA, ESA, the German and Italian Space Agencies, and the Bulgarian Academy of Sciences, which compiles radiation measurements relevant to human spaceflight and provides tools for accessing and manipulating the data, and 3) a Multi-study visualization tool which enables users to look across and combine GeneLab’s omics datasets across different experiments and missions. There are ~600 volunteer OSDR Analysis Working Group (AWG) members who: 1) provide feedback on scientific standards for reuse (subject and assay metadata; processing pipelines; dataset formats and uniformed structures for machine-readability), and 2) collaborate to mine-reuse OSDR data conducting scientific analysis. OSDR has enabled 60 publications as of September 2023, many directly from AWG collaborations most notably the Cell Press package in 2020. Lastly, there are at least 15 articles which mine OSDR data part of a package of ~50 articles across Nature Portfolio with research stemming from I4, the Japan Aerospace Exploration Agency, NASA Space Biology, and the NASA Human Research Program.

space biology↗

Pre-training Vision Models for the Classification of Alerts from Wide-field Time-domain Surveys

Modern wide-field time-domain surveys facilitate the study of transient, variable and moving phenomena by conducting image differencing and relaying alerts to their communities. Machine learning tools have been used on data from these surveys and their precursors for more than a decade, and convolutional neural networks (CNNs), which make predictions directly from input images, saw particularly broad adoption through the 2010s. Since then, continually rapid advances in computer vision have transformed the standard practices around using such models. It is now commonplace to use standardized architectures pre-trained on large corpora of everyday images (e.g., ImageNet). In contrast, time-domain astronomy studies still typically design custom CNN architectures and train them from scratch. Here, we explore the effects of adopting various pre-training regimens and standardized model architectures on the performance of alert classification. We find that the resulting models match or outperform a custom, specialized CNN like what is typically used for filtering alerts. Moreover, our results show that pre-training on galaxy images from Galaxy Zoo tends to yield better performance than pre-training on ImageNet or training from scratch. We observe that the design of standardized architectures are much better optimized than the custom CNN baseline, requiring significantly less time and memory for inference despite having more trainable parameters. On the eve of the Legacy Survey of Space and Time and other image-differencing surveys, these findings advocate for a paradigm shift in the creation of vision models for alerts, demonstrating that greater performance and efficiency, in time and in data, can be achieved by adopting the latest practices from the computer vision field.

79 ASTRONOMY AND ASTROPHYSICS↗

A Pride of Satellites in the Constellation Leo? Discovery of the Leo VI Milky Way Satellite Ultra-faint Dwarf Galaxy with DELVE Early Data Release 3

Abstract We report the discovery and spectroscopic confirmation of an ultra-faint Milky Way satellite in the constellation of Leo. This system was discovered as a spatial overdensity of resolved stars observed with Dark Energy Camera (DECam) data from an early version of the third data release of the DECam Local Volume Exploration (or DELVE) survey. The low luminosity ( M V = − 3.5 6 − 0.37 + 0.47 ; L V = 230 0 − 700 + 1200 L ⊙ ), large size ( R 1 / 2 = 9 0 − 30 + 30 pc), and large heliocentric distance ( D = 11 1 − 6 + 9 kpc) are all consistent with the population of ultra-faint dwarf galaxies (UFDs). Using Keck/DEIMOS observations of the system, we were able to spectroscopically confirm nine member stars, while measuring a tentative mass-to-light ratio of 70 0 − 500 + 1400 M ⊙ / L ⊙ and a nonzero metallicity dispersion of σ [ Fe / H ] = 0.1 9 − 0.11 + 0.14 , further confirming Leo VI’s identity as a UFD. While the system has a highly elliptical shape, ϵ = 0.5 4 − 0.29 + 0.19 , we do not find any conclusive evidence that it is tidally disrupting. Moreover, despite the apparent on-sky proximity of Leo VI to members of the proposed Crater-Leo infall group, its smaller heliocentric distance and inconsistent position in energy–angular momentum space make it unlikely that Leo VI is part of the proposed infall group.

79 ASTRONOMY AND ASTROPHYSICS↗

International Symposium on Remote Sensing of Environment, 15th, University of Michigan, Ann Arbor, MI, May 11-15, 1981, Proceedings. Volumes 1, 2 & 3

Developments related to advanced sensors and sensor systems are being examined, taking into account advanced aerospace remote sensing systems for global resource applications, spaceborne radar observation of the earth surface, a concept for an advanced earth resources satellite system, technologies for the multispectral mapping of earth resources, and the use of Landsat images and morphologic analogs in space exploration. Other topics discussed are related to modeling for terrain analysis, digital processing and analysis of remotely sensed data, microwave remote sensing, new discoveries from planetary remote sensing, and data base utilization. Advances in the area of luminescence are also considered along with future plans and prospects concerning the remote sensing of the earth from space.

Source record↗

Managing Sustainable Data Infrastructures: The Gestalt of EOSDIS

EOSDIS epitomizes a System of Systems, whose many varied and distributed parts are integrated into a single, highly functional organized science data system. A distributed architecture was adopted to ensure discipline-specific support for the science data, while also leveraging standards and establishing policies and tools to enable interdisciplinary research, and analysis across multiple scientific instruments. The EOSDIS is composed of system elements such as geographically distributed archive centers used to manage the stewardship of data. The infrastructure consists of underlying capabilities connections that enable the primary system elements to function together. For example, one key infrastructure component is the common metadata repository, which enables discovery of all data within the EOSDIS system. EOSDIS employs processes and standards to ensure partners can work together effectively, and provide coherent services to users.

remote sensing↗

Harvesting NASA's Common Metadata Repository (CMR)

As part of NASA's Earth Observing System Data and Information System (EOSDIS), the Common Metadata Repository (CMR) stores metadata for over 30,000 datasets from both NASA and international providers along with over 300M granules. This metadata enables sub-second discovery and facilitates data access. While the CMR offers a robust temporal, spatial and keyword search functionality to the general public and international community, it is sometimes more desirable for international partners to harvest the CMR metadata and merge the CMR metadata into a partner's existing metadata repository. This poster will focus on best practices to follow when harvesting CMR metadata to ensure that any changes made to the CMR can also be updated in a partner's own repository. Additionally, since each partner has distinct metadata formats they are able to consume, the best practices will also include guidance on retrieving the metadata in the desired metadata format using CMR's Unified Metadata Model translation software.

Earth Resources↗