Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Automated Downlink Pipeline for Scientific Data Using TReK

ISS users generate scientific data files on orbit that require console operators to retrieve and deliver them for analysis. Automating the downlink process using TReK CFDP and DTN provides ground flight controllers and PD teams increased efficiency, reducing workload and resulting in cost savings without reduced services provided.

TReK↗

Kepler Data Release 3 Notes

This describes the collection of data and the processing done on it so when researchers around the world get the Kepler data sets (which are a set of pixels from the telescope of a particular target (star, galaxy or whatever) over a 3 month period) they can adjust their algorithms fro things that were done (like subtracting all of one particular wavelength for example). This is used to calibrate their own algorithms so that they know what it is they are starting with. It is posted so that whoever is accessing the publicly available data (not all of it is made public) can understand it .. (most of the Kepler data is under restriction for 1 - 4 years and is not available, but the handbook is for everyone (public and restricted) The Data Analysis Working Group have released long and short cadence materials, including FFls and Dropped Targets for the Public. The Kepler Science Office considers Data Release 3 to provide "browse quality" data. These notes have been prepared to give Kepler users of the Multimission Archive at STScl (MAST) a summary of how the data were collected and prepared, and how well the data processing pipeline is functioning on flight data. They will be updated for each release of data to the public archive and placed on MAST along with other Kepler documentation, at http:// archive.stsci.edu/kepler/documents.html .Data release 3 is meant to give users the opportunity to examine the data for possibly interesting science and to involve the users in improving the pipeline for future data releases. To perform the latter service, users are encouraged to notice and document artifacts, either in the raw or processed data, and report them to the Science Office.

Cleve, Jeffrey E.↗

Kepler Data Release 4 Notes

The Data Analysis Working Group have released long and short cadence materials, including FFIs and Dropped Targets for the Public. The Kepler Science Office considers Data Release 4 to provide "browse quality" data. These notes have been prepared to give Kepler users of the Multimission Archive at STScl (MAST) a summary of how the data were collected and prepared, and how well the data processing pipeline is functioning on flight data. They will be updated for each release of data to the public archive and placed on MAST along with other Kepler documentation, at http://archive.stsci.edu/kepler/documents.html. Data release 3 is meant to give users the opportunity to examine the data for possibly interesting science and to involve the users in improving the pipeline for future data releases. To perform the latter service, users are encouraged to notice and document artifacts, either in the raw or processed data, and report them to the Science Office.

Van Cleve, Jeffrey↗

Evolving Multi-hazard Machine Learning Modeling for Advanced Risk-Informed Infrastructure Resilience Assessment

The socioeconomic impacts of pipeline incidents have escalated over the past three decades, revealing the limitation of traditional risk modeling methods when applied to extensive pipeline networks. This research aims to develop machine learning (ML) models that effectively identify, rank, and predict the diverse hazards and socioeconomic consequences associated with pipeline incidents. Utilizing historical data on pipeline incidents alongside weather and oceanographic data from the 1980s onward, the Houston metropolitan area serves as a testbed for the proposed methodologies. The research segments the combined datasets into three consecutive periods, demonstrating the efficacy of the updated model in predicting future events, particularly concerning precipitation rate data. Despite the challenges posed by a relatively limited dataset, local-level ML modeling offers valuable insights into the spatial and temporal dynamics of multiple hazards that contribute to pipeline incidents. These findings hold significant implications for future research, particularly in understanding and mitigating risks in various locations across the Gulf Coast and other coastal regions.

42 ENGINEERING↗

Improving Sim-to-Real Transfer in Vision-Based Robot Navigation Via Instance-Level GAN-Based Data Augmentation

Achieving robust vision-based robotic tasks requires large amounts of data, which are often difficult to obtain in real-world scenarios. Simulators and synthetic data offer a cost-effective alternative, but the visual gap between simulation and reality hinders the performance of models when deployed in real-world environments. In this paper, we present a data augmentation pipeline that integrates a foundation model (Segment Anything Model) with an unsupervised image-to-image translation model (CycleGAN) for instance-level domain transfer from simulation to reality. This pipeline enables the generation of realistic labeled data from synthetic images for training supervised machine learning models in vision-based navigation tasks. We evaluate our approach on real-world data for ego-vehicle pose estimation, a critical autonomous navigation task involving the prediction of cross-track position and heading angle relative to road center line markings. The results of our tests show that our GAN-based data augmentation pipeline significantly outperforms models trained solely on simulation data or on data processed with standard image augmentation methods for sim-to-real transfer, enhancing model robustness and generalizability in real-world scenarios. Our method provides a scalable and flexible data augmentation tool for leveraging large synthetic datasets to enhance vision-based robotic navigation tasks.

artificial intelligence↗

RectifHyd Version 2.0: Historical and counterfactual-climatological hydropower monthly generation totals for CONUS plants, 1980 – 2019

This dataset will contain the following files: - RectifHydV2.zip – the RectifHydV2 dataset, split into two files: o RectifHydV2_Actual_MWh.csv: Estimated actual monthly net generation from 590 Hydropower Plants (>10MW nameplate) in CONUS, 1980 – 2019 o RectifHydV2_Counterfactual_MWh.csv: Counterfactual (climate-only) monthly net generation from 590 Hydropower Plants (>10MW nameplate) in CONUS, 1980 – 2019 - RectifHydV2_code.zip: Full data processing pipeline, coded using R {targets} framework. This is a snapshot release (v1.0) of the code repository stored at https://code.ornl.gov/turnersw/rectifhydv2 - RectifHydV2_inputs.zip: Complete set of input data used to create RectifHydV2, organized for direct entry into “/data/” directory of the RectifHydV2 reproducible data pipeline - RectifHydV2_misc.zip: - RectifHydV2_dailyRelease.csv

13 HYDRO ENERGY↗

Kepler Planet Detection Metrics: Pixel-Level Transit Injection Tests of Pipeline Detection Efficiency for Data Release 25

This document describes the results of the fourth pixel-level transit injection experiment, which was designed to measure the detection efficiency of both the Kepler pipeline (Jenkins 2002, 2010; Jenkins et al. 2017) and the Robovetter (Coughlin 2017). Previous transit injection experiments are described in Christiansen et al. (2013, 2015a,b, 2016).In order to calculate planet occurrence rates using a given Kepler planet catalogue, produced with a given version of the Kepler pipeline, we need to know the detection efficiency of that pipeline. This can be empirically determined by injecting a suite of simulated transit signals into the Kepler data, processing the data through the pipeline, and examining the distribution of successfully recovered transits. This document describes the results for the pixel-level transit injection experiment performed to accompany the final Q1-Q17 Data Release 25 (DR25) catalogue (Thompson et al. 2017)of the Kepler Objects of Interest. The catalogue was generated using the SOC pipeline version 9.3 and the DR25 Robovetter acting on the uniformly processed Q1-Q17 DR25 light curves (Thompson et al. 2016a) and assuming the Q1-Q17 DR25 Kepler stellar properties (Mathur et al. 2017).

Pixel-Level Transit Injection↗

The Zwicky Transient Facility: Data Processing, Products, and Archive

The Zwicky Transient Facility (ZTF) is a new robotic time-domain survey currently in progress using the Palomar 48-inch Schmidt Telescope. ZTF uses a 47 square degree field with a 600 megapixel camera to scan the entire northern visible sky at rates of ∼3760 square degrees/hour to median depths of g ~ 20.8 and r ~ 20.6 mag (AB, 5σ in 30 sec). We describe the Science Data System that is housed at IPAC, Caltech. This comprises the data-processing pipelines, alert production system, data archive, and user interfaces for accessing and analyzing the products. The real-time pipeline employs a novel image-differencing algorithm, optimized for the detection of point-source transient events. These events are vetted for reliability using a machine-learned classifier and combined with contextual information to generate data-rich alert packets. The packets become available for distribution typically within 13 minutes (95th percentile) of observation. Detected events are also linked to generate candidate moving-object tracks using a novel algorithm. Objects that move fast enough to streak in the individual exposures are also extracted and vetted. We present some preliminary results of the calibration performance delivered by the real-time pipeline. The reconstructed astrometric accuracy per science image with respect to Gaia DR1 is typically 45 to 85 milliarcsec. This is the RMS per-axis on the sky for sources extracted with photometric S/N ≥10 and hence corresponds to the typical astrometric uncertainty down to this limit. The derived photometric precision (repeatability) at bright unsaturated fluxes varies between 8 and 25 millimag. The high end of these ranges corresponds to an airmass approaching ∼2—the limit of the public survey. Photometric calibration accuracy with respect to Pan-STARRS1 is generally better than 2%. The products support a broad range of scientific applications: fast and young supernovae; rare flux transients; variable stars; eclipsing binaries; variability from active galactic nuclei; counterparts to gravitational wave sources; a more complete census of Type Ia supernovae; and solar-system objects.

Frank J. Masci↗

Demystifying Kepler Data: A Primer for Systematic Artifact Mitigation

The Kepler spacecraft has collected data of high photometric precision and cadence almost continuously since operations began on 2009 May 2. Primarily designed to detect planetary transits and asteroseismological signals from solar-like stars, Kepler has provided high quality data for many areas of investigation. Unconditioned simple aperture time-series photometry are however affected by systematic structure. Examples of these systematics are differential velocity aberration, thermal gradients across the spacecraft, and pointing variations. While exhibiting some impact on Kepler's primary science, these systematics can critically handicap potentially ground-breaking scientific gains in other astrophysical areas, especially over long timescales greater than 10 days. As the data archive grows to provide light curves for 10(exp 5) stars of many years in length, Kepler will only fulfill its broad potential for stellar astrophysics if these systematics are understood and mitigated. Post-launch developments in the Kepler archive, data reduction pipeline and open source data analysis software have occurred to remove or reduce systematic artifacts. This paper provides a conceptual primer for users of the Kepler data archive to understand and recognize systematic artifacts within light curves and some methods for their removal. Specific examples of artifact mitigation are provided using data available within the archive. Through the methods defined here, the Kepler community will find a road map to maximizing the quality and employment of the Kepler legacy archive.

Kinemuchi, K.↗

Automated Detection and Analysis of Resident Space Objects with the 1.3-Meter Eugene Stansbery-Meter Class Autonomous Telescope

Optical telescopes dedicated to the detection of orbital debris employ large-area detectors that generate a large number of images each night. Such surveys require automated data analysis pipelines that process the images and detect moving objects. We present an overview of the data analysis pipeline employed by the 1.3-meter Eugene Stansbery-Meter Class Autonomous Telescope (ES-MCAT) on Ascension Island, operated by NASA’s Orbital Debris Program Office. The pipeline enfolds the astrometric and photometric calibration of the images, star-trail removal, object detection, correlation over multiple sequential image frames, and orbital parameter estimation. The performance of the pipeline was investigated by means of Monte-Carlo simulations in which simulated object tracks were inserted into ES-MCAT images and then processed by the pipeline. This technique allows one to confidently estimate the completeness for the detection of resident space objects as a function of apparent magnitude and angular velocity. This paper discusses these techniques and provides examples using actual data.

Paul Hickson↗

Automated Detection and Analysis of Resident Space Objects with the 1.3-Meter Eugene Stansbery-Meter Class Autonomous Telescope

Optical telescopes dedicated to the detection of orbital debris employ large-area detectors that generate a large number of images each night. Such surveys require automated data analysis pipelines that process the images and detect moving objects. We present an overview of the data analysis pipeline employed by the 1.3-meter Eugene Stansbery-Meter Class Autonomous Telescope (ES-MCAT) on Ascension Island, operated by NASA’s Orbital Debris Program Office. The pipeline enfolds the astrometric and photometric calibration of the images, star-trail removal, object detection, correlation over multiple sequential image frames, and orbital parameter estimation. The performance of the pipeline was investigated by means of Monte-Carlo simulations in which simulated object tracks were inserted into ES-MCAT images and then processed by the pipeline. This technique allows one to confidently estimate the completeness for the detection of resident space objects as a function of apparent magnitude and angular velocity. This paper discusses these techniques and provides examples using actual data.

Paul Hickson↗

Scalable Hybrid Learning Techniques for Scientific Data Compression

Data compression is becoming critical for storing scientific data because many scientific applications need to store large amounts of data and post process this data for scientific discovery. Unlike image and video compression algorithms that limit errors to primary data (PD), scientists require compression techniques that accurately preserve derived quantities of interest (QoIs). Here, this article presents a physics-informed compression technique implemented as an end-to-end, scalable, GPU-based pipeline for data compression that addresses this requirement. Our hybrid compression technique combines machine learning techniques and standard compression methods. Specifically, we combine an autoencoder, an error-bounded lossy compressor to provide guarantees on raw data error, and a constraint satisfaction post-processing step to preserve the QoIs within a minimal error (generally less than floating point error). The effectiveness of the data compression pipeline is demonstrated by compressing nuclear fusion simulation data generated by a large-scale fusion code, XGC, which produces hundreds of terabytes of data in a single day. Our approach works within the ADIOS framework and results in compression by a factor of more than 150 while requiring only a few percent of the computational resources necessary for generating the data, making the overall approach highly effective for practical scenarios.

ITER↗

Kepler Data Validation II–Transit Model Fitting and Multiple-Planet Search

This paper discusses the transit model-fitting and multiple-planet search algorithms and performance of the Kepler Science Data Processing Pipeline, developed by the Kepler Science Operations Center (SOC). Threshold crossing events (TCEs), which are transit candidate events, are generated by the Transiting Planet Search (TPS) component of the pipeline and subsequently processed in the data validation (DV) component. The transit model is used in DV to fit TCEs to characterize planetary candidates and to derive parameters that are used in various diagnostic tests to classify them. After the signature associated with the TCE is removed from the light curve of the target star, the residual light curve goes through TPS again to search for additional TCEs. The iterative process of transit model fitting and multiple-planet search continues until no TCE is generated from the residual light curve or an upper limit is reached. The transit model-fitting and multiple-planet search performance of the final release (9.3, 2016January) of the pipeline is demonstrated with the results of the processing of four years (17 quarters) of flight data from the primary Kepler Mission. The transit model-fitting results are accessible from the NASA Exoplanet Archive. The final version of the SOC codebase is available through GitHub.

Threshold crossing events (TCEs↗

The DEEP2 Galaxy Redshift Survey: Design, Observations, Data Reduction, and Redshifts

We describe the design and data analysis of the DEEP2 Galaxy Redshift Survey, the densest and largest high-precision redshift survey of galaxies at z approx. 1 completed to date. The survey was designed to conduct a comprehensive census of massive galaxies, their properties, environments, and large-scale structure down to absolute magnitude MB = −20 at z approx. 1 via approx.90 nights of observation on the Keck telescope. The survey covers an area of 2.8 Sq. deg divided into four separate fields observed to a limiting apparent magnitude of R(sub AB) = 24.1. Objects with z approx. < 0.7 are readily identifiable using BRI photometry and rejected in three of the four DEEP2 fields, allowing galaxies with z > 0.7 to be targeted approx. 2.5 times more efficiently than in a purely magnitude-limited sample. Approximately 60% of eligible targets are chosen for spectroscopy, yielding nearly 53,000 spectra and more than 38,000 reliable redshift measurements. Most of the targets that fail to yield secure redshifts are blue objects that lie beyond z approx. 1.45, where the [O ii] 3727 Ang. doublet lies in the infrared. The DEIMOS 1200 line mm(exp −1) grating used for the survey delivers high spectral resolution (R approx. 6000), accurate and secure redshifts, and unique internal kinematic information. Extensive ancillary data are available in the DEEP2 fields, particularly in the Extended Groth Strip, which has evolved into one of the richest multiwavelength regions on the sky. This paper is intended as a handbook for users of the DEEP2 Data Release 4, which includes all DEEP2 spectra and redshifts, as well as for the DEEP2 DEIMOS data reduction pipelines. Extensive details are provided on object selection, mask design, biases in target selection and redshift measurements, the spec2d two-dimensional data-reduction pipeline, the spec1d automated redshift pipeline, and the zspec visual redshift verification process, along with examples of instrumental signatures or other artifacts that in some cases remain after data reduction. Redshift errors and catastrophic failure rates are assessed through more than 2000 objects with duplicate observations. Sky subtraction is essentially photon-limited even under bright OH sky lines; we describe the strategies that permitted this, based on high image stability, accurate wavelength solutions, and powerful B-spline modeling methods. We also investigate the impact of targets that appear to be single objects in ground-based targeting imaging but prove to be composite in Hubble Space Telescope data; they constitute several percent of targets at z approx. 1, approaching approx. 5%-10% at z > 1.5. Summary data are given that demonstrate the superiority of DEEP2 over other deep high-precision redshift surveys at z approx. 1 in terms of redshift accuracy, sample number density, and amount of spectral information. We also provide an overview of the scientific highlights of the DEEP2 survey thus far.

Galaxy↗

Modeling performance of data collection systems for high-energy physics

Exponential increases in scientific experimental data are outpacing silicon technology progress, necessitating heterogeneous computing systems—particularly those utilizing machine learning (ML)—to meet future scientific computing demands. The growing importance and complexity of heterogeneous computing systems require systematic modeling to understand and predict the effective roles for ML. We present a model that addresses this need by framing the key aspects of data collection pipelines and constraints and combining them with the important vectors of technology that shape alternatives, computing metrics that allow complex alternatives to be compared. For instance, a data collection pipeline may be characterized by parameters such as sensor sampling rates and the overall relevancy of retrieved samples. Alternatives to this pipeline are enabled by development vectors including ML, parallelization, advancing CMOS, and neuromorphic computing. By calculating metrics for each alternative such as overall F1 score, power, hardware cost, and energy expended per relevant sample, our model allows alternative data collection systems to be rigorously compared. We apply this model to the Compact Muon Solenoid experiment and its planned high luminosity-large hadron collider upgrade, evaluating novel technologies for the data acquisition system (DAQ), including ML-based filtering and parallelized software. The results demonstrate that improvements to early DAQ stages significantly reduce resources required later, with a power reduction of 60% and increased relevant data retrieval per unit power (from 0.065 to 0.31 samples/kJ). However, we predict that further advances will be required in order to meet overall power and cost constraints for the DAQ.

Olin-Ammentorp, Wilkie (ORCID:0000000224729862)↗

Analytical methods for online data quality assessment

This chapter provides a comprehensive overview of the main steps for algorithmic sensor signal quality assessment, which can enhance the decision-making process for water resource recovery facility (WRRF) operation and optimization. It introduces the concept of redundancy as the basis for data quality assessment. It also explains the typical data processing pipeline, which consists of preliminary analysis, data pre-processing, and specific algorithmic approaches. Each of these processes is presented and discussed in three separate sections. Importantly, this chapter introduces the main approaches for data quality assessment, provides guidelines for selecting the most suitable one and the key performance indicators to evaluate them and explains how to collect metadata through such an algorithmic approach.

Aguado, Daniel↗

HydroForecast Long-term: Improving hydropower’s resilience to climate change through accurate climate-scale

With hydrologic patterns and water availability across the globe shifting due to climate change, advancements in hydrologic prediction systems can help significantly reduce the uncertainties that utilities and water supply entities have in their decision making. Understanding and estimating hydrology at the climate scale is critical for managing water resources under changing climate scenarios. This project focuses on integrating state-of-the-art neural network modeling with downscaled climate projections to deliver the reliable water supply projections decades into the future to meet an urgent need from hydropower operators and water utilities. In this Phase 1 DOE SBIR proposal, we developed and validated a theory-guided neural network model, HydroForecast Long-term, for climate-scale hydrology and implemented the model within existing HydroForecast infrastructure. HydroForecast Long-term combines the most accurate streamflow modeling system with a flexible and scalable data architecture to generate water supply projections out to the year 2100. This report illustrates that we have achieved our four objectives: 1) create a prototype of HydroForecast Long-term, building the neural network prediction model, 2) build an automated data input pipeline that processes large amounts of data from the latest global temperature and precipitation climate models; 3) benchmark the accuracy of the hydrologic model over the recent two decades over a large set of diverse basins, and 4) create a set of output visuals and summary metrics informed by customer feedback that connect the data to critical decision points. This work empowers water users to make data-informed decisions supporting a resilient, renewable-powered grid and water system. The results advance the Department of Energy’s mission by addressing critical gaps in water supply planning under climate change.

13 HYDRO ENERGY↗

The Kepler End-to-End Model: Creating High-Fidelity Simulations to Test Kepler Ground Processing

The Kepler mission is designed to detect the transit of Earth-like planets around Sun-like stars by observing 100,000 stellar targets. Developing and testing the Kepler ground-segment processing system, in particular the data analysis pipeline, requires high-fidelity simulated data. This simulated data is provided by the Kepler End-to-End Model (ETEM). ETEM simulates the astrophysics of planetary transits and other phenomena, properties of the Kepler spacecraft and the format of the downlinked data. Major challenges addressed by ETEM include the rapid production of large amounts of simulated data, extensibility and maintainability.

Bryson, Stephen T.↗