Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Time-Distance Helioseismology Data-Analysis Pipeline for Helioseismic and Magnetic Imager Onboard Solar Dynamics Observatory (SDO-HMI) and Its Initial Results

The Helioseismic and Magnetic Imager onboard the Solar Dynamics Observatory (SDO/HMI) provides continuous full-disk observations of solar oscillations. We develop a data-analysis pipeline based on the time-distance helioseismology method to measure acoustic travel times using HMI Doppler-shift observations, and infer solar interior properties by inverting these measurements. The pipeline is used for routine production of near-real-time full-disk maps of subsurface wave-speed perturbations and horizontal flow velocities for depths ranging from 0 to 20 Mm, every eight hours. In addition, Carrington synoptic maps for the subsurface properties are made from these full-disk maps. The pipeline can also be used for selected target areas and time periods. We explain details of the pipeline organization and procedures, including processing of the HMI Doppler observations, measurements of the travel times, inversions, and constructions of the full-disk and synoptic maps. Some initial results from the pipeline, including full-disk flow maps, sunspot subsurface flow fields, and the interior rotation and meridional flow speeds, are presented.

Sun: helioseismology

IN13B-1660: Analytics and Visualization Pipelines for Big Data on the NASA Earth Exchange (NEX) and OpenNEX

We are developing capabilities for an integrated petabyte-scale Earth science collaborative analysis and visualization environment. The ultimate goal is to deploy this environment within the NASA Earth Exchange (NEX) and OpenNEX in order to enhance existing science data production pipelines in both high-performance computing (HPC) and cloud environments. Bridging of HPC and cloud is a fairly new concept under active research and this system significantly enhances the ability of the scientific community to accelerate analysis and visualization of Earth science data from NASA missions, model outputs and other sources. We have developed a web-based system that seamlessly interfaces with both high-performance computing (HPC) and cloud environments, providing tools that enable science teams to develop and deploy large-scale analysis, visualization and QA pipelines of both the production process and the data products, and enable sharing results with the community. Our project is developed in several stages each addressing separate challenge - workflow integration, parallel execution in either cloud or HPC environments and big-data analytics or visualization. This work benefits a number of existing and upcoming projects supported by NEX, such as the Web Enabled Landsat Data (WELD), where we are developing a new QA pipeline for the 25PB system.

visualization

Technology Cost and Schedule Estimation (TCASE) Final Report

During the 2014-2015 project year, the focus of the TCASE project has shifted from collection of historical data from many sources to securing a data pipeline between TCASE and NASA's widely used TechPort system. TCASE v1.0 implements a data import solution that was achievable within the project scope, while still providing the basis for a long-term ability to keep TCASE in sync with TechPort. Conclusion: TCASE data quantity is adequate and the established data pipeline will enable future growth. Data quality is now highly dependent the quality of data in TechPort. Recommendation: Technology development organizations within NASA should continue to work closely with project/program data tracking and archiving efforts (e.g. TechPort) to ensure that the right data is being captured at the appropriate quality level. TCASE would greatly benefit, for example, if project cost/budget information was included in TechPort in the future.

Wallace, Jon

The Kepler Science Data Processing Pipeline Source Code Road Map

We give an overview of the operational concepts and architecture of the Kepler Science Processing Pipeline. Designed, developed, operated, and maintained by the Kepler Science Operations Center (SOC) at NASA Ames Research Center, the Science Processing Pipeline is a central element of the Kepler Ground Data System. The SOC consists of an office at Ames Research Center, software development and operations departments, and a data center which hosts the computers required to perform data analysis. The SOC's charter is to analyze stellar photometric data from the Kepler spacecraft and report results to the Kepler Science Office for further analysis. We describe how this is accomplished via the Kepler Science Processing Pipeline, including, the software algorithms. We present the high-performance, parallel computing software modules of the pipeline that perform transit photometry, pixel-level calibration, systematic error correction, attitude determination, stellar target management, and instrument characterization.

Kepler pipeline software

Surface Biology & Geology Pathfinder Data Analysis Pipeline

NASA's future global orbital mission, currently in development as the Surface Biology and Geology (SBG) Designated Observable study, will acquire relatively high resolution solar-reflected spectroscopy and thermal infrared observations. Innovative processes must be utilized for handling the high volume of data anticipated to be collected, which is anticipated to exceed 100 terabytes/day, greater than NASA's total extant airborne hyperspectral data collection. Collecting, processing/re-processing, disseminating, and exploiting this volume of data presents new challenges. To begin addressing them, NASA is drawing upon the expertise developed from its astrophysics programs to address Earth science and applications. Specifically, NASA is adapting the science processing operations technology developed for the Kepler and TESS planet-hunting missions for imaging spectroscopy data processing. This technology development has been the foundation for the remarkable scientific successes of Kepler and TESS. The Kepler/TESS data processing technology provides a scalable architecture for robust, repeatable, and replicable science and application products while enabling the Earth science community to develop, test, and implement new algorithms. Our effort to leverage this existing capability has begun by ingesting data and applying workflows from the EO-1/Hyperion 17-year mission archive that provides globally sampled visible through shortwave infrared spectra that are representative of SBG data types and volumes. This pathfinding data processing system will help define the solutions to processing SBG data volumes and will enable the scientific community to interact with the data and processing pipeline to create new science products.

Jenkins, Jon

Selection of a Pair of Experiments to Optimally Reduce Uncertainty in Targeted Nuclear Data

We propose a novel process to select a pair of differential and integral experiments that best reduce uncertainties in targeted 239 ⁢Pu nuclear data while compressing the current nuclear data pipeline from 20 to 3 years. 239⁢ Pu nuclear data are poorly understood for neutrons in the intermediate energy range due to sparsity and uncertainty in historical experiments. New experiments targeting this range will enable better understanding of these nuclear data, but choosing the ideal experiments to conduct is challenging. Beginning with a prior distribution represented by samples of nuclear data generated from theory, generalized least squares adjustments are made to incorporate data from historical experiments. To quantify potential uncertainty reduction obtainable from a pair of candidate experiments, we compute the D-optimality criterion of the posterior covariance of intermediate energy range nuclear data compared to the equivalent covariance after additional adjustment to the pair of candidate experiments. Repeating the process for each of many candidate pairs facilitates the final selection. Results support 63⁢ Cu total cross section measurements for differential experiments and alumina and alumina/graphite configurations for integral experiments. This analysis enables choosing differential and integral experiments to be executed concurrently while shortening decision times relative to the current nuclear data pipeline.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Establishing Data Analysis Pipeline for Bulk ATAC-Seq Datasets

We developed an analysis pipeline for transposase-accessible chromatin sequencing (ATAC-Seq) data derived from bulk samples, which brings together publicly available R packages in addition to command-line tools designed for analysis of bulk ATAC-Seq data and can be run on any computer running a Linux-like operating system such as Ubuntu or Apple OSX.

97 MATHEMATICS AND COMPUTING

A Proposal to Investigate Outstanding Problems in Astronomy

During the period leading up to the spectacular launch of the Space Shuttle Columbia (STS-109) on 1 March 2002 6:22 am EST, the team worked hard on a myriad of tasks to be ready for launch. Our launch support included preparations and rehearsals for the support during the mission, preparation for the SMOV and ERO program, and work to have the science team's data pipeline (APSIS) and data archive (SDA) ready by launch. A core of the team that was at the GSFC during the EVA that installed ACS monitored the turn-on and aliveness tests of ACS. One hour after installation of ACS in the HST George Hartig was showing those of us at Goddard the telemetry which demonstrated that the HRC and WFC CCDs were cooling to their preset temperatures. The TECs had survived launch! After launch, the team had several immediate and demanding tasks. We had to process the ERO observations through our pipeline and understand the limitations of the ground based-based calibrations, and simultaneously prepare the EROs for public release. The ERO images and the SMOV calibrations demonstrated that ACS met or exceeded its specifications for image quality and sensitivity. It is the most sensitive instrument that Hubble has had. The ERO images themselves made the front page of all of the major newspapers in the US. During the months after launch we have worked on the SMOV observations, and are analyzing the data from our science program.

Ford, Holland

Ground System for Solar Dynamics Observatory (SDO) Mission

NASA s Goddard Space Flight Center (GSFC) has recently completed its Critical Design Review (CDR) of a new dual Ka and S-band ground system for the Solar Dynamics Observatory (SDO) Mission. SDO, the flagship mission under the new Living with a Star Program Office, is one of GSFC s most recent large-scale in-house missions. The observatory is scheduled for launch in August 2008 from the Kennedy Space Center aboard an Atlas-5 expendable launch vehicle. Unique to this mission is an extremely challenging science data capture requirement. The mission is required to capture 99.99% of available science over 95% of all observation opportunities. Due to the continuous, high volume (150 Mbps) science data rate, no on-board storage of science data will be implemented on this mission. With the observatory placed in a geo-synchronous orbit at 36,000 kilometers within view of dedicated ground stations, the ground system will in effect implement a "real-time" science data pipeline with appropriate data accounting, data storage, data distribution, data recovery, and automated system failure detection and correction to keep the science data flowing continuously to three separate Science Operations Centers (SOCs). Data storage rates of approx. 45 Tera-bytes per month are expected. The Mission Operations Center (MOC) will be based at GSFC and is designed to be highly automated. Three SOCs will share in the observatory operations, each operating their own instrument. Remote operations of a multi-antenna ground station in White Sands, New Mexico from the MOC is part of the design baseline.

Tann, Hun K.

Data Accountability and Uncertainty Analysis for the Mars Science Laboratory

This paper presents machine learning-based approaches to automate and optimize the detection of volume loss for the downlink process of telemetry data from the Mars Curiosity Rover. The Curiosity observes volume loss and data corruption, requiring re-transmits from the rover and Ground Data System Analysts (GDSA) to monitor the data flow. To resolve this issue, we created a data pipeline to accumulate data from various data sources in the downlink process and detect where the data is missed. In this paper, we benchmarked different methodologies based on the accuracy and excitability of them to identify whether a downlink data that is received to the ground system is complete or incomplete. Our results show that machine learning methods can improve the performance of the GDSA by 55% while the user can diagnose why data is missed and provide an explanation for the data accountability problem.

Chowdhury, Ameera

Generalizing a Data Analysis Pipeline in the Cloud to Handle Diverse Use Cases in NASA's EOSDIS

NASA's Earth Observing System Data and Information System (EOSDIS) is tasked with archiving and distributing Earth Observation data across a range of disciplines, including atmospheric science, oceanography, land processes, natural hazards, solar radiance and even socioeconomic aspects relating to the environment. Driven by rapidly rising data volumes, EOSDIS is migrating to a cloud computing based archive over the next few years. Although this simplifies data management somewhat, the main aim is to provide the data in an environment where end users can bring their analysis to the data rather than attempting to download and manage ever-increasing volumes. To that end, a cloud-based analysis platform is being constructed to enable data transformations, analyses and visualization without egressing the data from the cloud. In this endeavor, we expect a wide variety of users, algorithms and use cases. Consequently, the architecture of this cloud analytics platform is expressly designed to be based on open services, thus fostering an ecosystem that enables the efficient combination of common components with data-specific or analysis-specific components. Reviewed and approved by Andrew Mitchell, ESDIS project manager.

Cloud computing

Data readiness pipeline patterns for scientific AI at scale: Insights from climate, fusion, life sciences, and materials

This article examines how data readiness for AI principles apply to large scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, life sciences, and materials—to identify common preprocessing patterns and domain‐specific constraints. We introduce a two‐dimensional readiness model that combines canonical preprocessing patterns with a five‐level operational readiness scale, both tailored to high‐performance computing (HPC) environments. This construct helps outline key challenges in transforming large‐scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross‐domain support for scalable and reproducible AI for science. Finally, we evaluate this maturity matrix in the context of case studies including ClimaX (climate), AFLOW (materials), OpenFold (proteomics), and DIII‐D fusion disruption‐prediction workflows, from which we distill lessons learned and provide recommendations to guide practitioners in developing robust AI‐readiness pipelines. Finally, we discuss remaining cross‐cutting challenges that persist across scientific domains.

97 MATHEMATICS AND COMPUTING

TESS Science Processing Operations Center Pipeline and Data Products

TESS launched 18 April 2018 to conduct a two-year, near all-sky survey for at least 50 small, nearby exoplanets for which masses can be ascertained and whose atmospheres can be characterized by ground- and space-based follow-on observations. TESS just completed its survey of the southern hemisphere, identifying >600 candidate exoplanets and unveiling a plethora of exciting non-exoplanet astrophysics results, such as asteroseismology, asteroids, and supernova. The TESS Science Processing Operations Center (SPOC) processes the data downlinked every two weeks to generate a range of data products hosted at the Mikulski Archive for Space Telescopes (MAST). For each sector (~1 month) of observations, the SPOC calibrates the image data for both 30-min Full Frame Images (FFIs) and up to 20,000 pre-selected 2-min target star postage stamps. Data products for the 2-min targets include simple aperture photometry and systematic error-corrected flux time series. The SPOC also conducts searches for transiting exoplanets in the 2-min data for each sector and generates Data Validation time series and associated reports for each transit-like feature identified in the search. Multi-sector searches for exoplanets are conducted periodically to discover longer period planets, including those in the James Webb Continuous Viewing Zone (CVZ), which are observed for up to one year. Data products also include co-trending basis vectors (CBVs) and calibration files, such as the Pixel Response Functions across the field of view of each of TESS's four cameras. To maximize the usability, the TESS science data products are modeled after those for Kepler, including Target Pixel Files and Light Curve files.In this talk, I describe the SPOC pipeline and the chief differences between the TESS and the Kepler pipelines, and the major updates to the SPOC pipeline (4.0) available now to the community at MAST. I also discuss the documentation available to the community to help them in properly interpreting and analyzing the TESS data products.The TESS Mission is funded by NASA's Science Mission Directorate as an Astrophysics Explorer Mission.

Jenkins, Jon M.

New Techniques for High-Contrast Imaging with ADI: The ACORNS-ADI SEEDS Data Reduction Pipeline

We describe Algorithms for Calibration, Optimized Registration, and Nulling the Star in Angular Differential Imaging (ACORNS-ADI), a new, parallelized software package to reduce high-contrast imaging data, and its application to data from the Strategic Exploration of Exoplanets and Disks (SEEDS) survey. We implement seyeral new algorithms, includbg a method to centroid saturated images, a trimmed mean for combining an image sequence that reduces noise by up to approx 20%, and a robust and computationally fast method to compute the sensitivitv of a high-contrast obsen-ation everywhere on the field-of-view without introducing artificial sources. We also include a description of image processing steps to remove electronic artifacts specific to Hawaii2-RG detectors like the one used for SEEDS, and a detailed analysis of the Locally Optimized Combination of Images (LOCI) algorithm commonly used to reduce high-contrast imaging data. ACORNS-ADI is efficient and open-source, and includes several optional features which may improve performance on data from other instruments. ACORNS-ADI is freely available for download at www.github.com/t-brandt/acorns_-adi under a BSD license

Brandt, Timothy D.

Laboratory Testing and Performance Verification of the CHARIS Integral Field Spectrograph

The Coronagraphic High Angular Resolution Imaging Spectrograph (CHARIS) is an integral field spectrograph (IFS) that has been built for the Subaru telescope. CHARIS has two imaging modes; the high-resolution mode is R82, R69, and R82 in J, H, and K bands respectively while the low-resolution discovery mode uses a second low-resolution prism with R19 spanning 1.15-2.37 microns (J+H+K bands). The discovery mode is meant to augment the low inner working angle of the Subaru Coronagraphic Extreme Adaptive Optics (SCExAO) adaptive optics system, which feeds CHARIS a coronagraphic image. The goal is to detect and characterize brown dwarfs and hot Jovian planets down to contrasts five orders of magnitude dimmer than their parent star at an inner working angle as low as 80 milliarcseconds. CHARIS constrains spectral crosstalk through several key aspects of the optical design. Additionally, the repeatability of alignment of certain optical components is critical to the calibrations required for the data pipeline. Specifically the relative alignment of the lens let array, prism, and detector must be highly stable and repeatable between imaging modes. We report on the measured repeatability and stability of these mechanisms, measurements of spectral crosstalk in the instrument, and the propagation of these errors through the data pipeline. Another key design feature of CHARIS is the prism, which pairs Barium Fluoride with Ohara L-BBH2 high index glass. The dispersion of the prism is significantly more uniform than other glass choices, and the CHARIS prisms represent the first NIR astronomical instrument that uses L-BBH2as the high index material. This material choice was key to the utility of the discovery mode, so significant efforts were put into cryogenic characterization of the material. The final performance of the prism assemblies in their operating environment is described in detail. The spectrograph is going through final alignment, cryogenic cycling, and is being delivered to the Subaru telescope in April 2016. This paper is a report on the laboratory performance of the spectrograph, and its current status in the commissioning process so that observers will better understand the instrument capabilities. We will also discuss the lessons learned during the testing process and their impact on future high-contrast imaging spectrographs for wavefront control.

Coronagraphic High Angular Resolution Imaging Spec

TESS Science Processing Operations Center Pipeline and Data Products

TESS (Transiting Exoplanet Survey Satellite) launched on 18-4-2018 to conduct a two-year, near all-sky survey for at least 50 nearby exoplanets for which masses can be obtained. TESS just completed surveying the southern hemisphere, identifying hundreds of candidate exoplanet systems and unveiling a plethora of exciting non-exoplanet astrophysics results, such as asteroseismology, asteroids, and supernova. The TESS Science Processing Operations Center (SPOC) at NASA Ames Research Center processes the image data downlinked from TESS every two weeks to generate a variety of data products hosted at the Mikulski Archive for Space Telescopes (MAST). For each approximately 1-month sector, the SPOC calibrates the image data for both 30-minute Full Frame Images (FFIs) and up to 20,000 pre-selected 2-minute target star postage stamps. Simple aperture photometry and systematic error-corrected flux-time series are generated for the 2-minute data. The data products also include co-trending basis vectors (CBVs) and calibration files, such as the Pixel Response Functions (PRF). The archival files are modeled after Kepler's for ease of use, and include Target Pixel Files (TPFs) containing original and calibrated 2-minute image data, Light Curve files (LCs) containing the photometric time series for each 2-minute target, as well as the Data Validation products. New products derived from the FFIs include light curves for the 2-minute targets and CBVs. The TESS Mission is funded by NASA's Science Mission Directorate as an Astrophysics Explorer Mission.

Science Pipeline