Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data pipeline”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Nuclear data sensitivity and uncertainty study of copper-reflected integral experiments [Slides]

This presentation touches on reducing uncertainties in intermediate-energy actinide nuclear data and this continues to be a high priority for many applications. The goal of PARADIGM (PARallel Approach of Differential and InteGral Measurements): accelerate efforts to reduce biases and uncertainties in nuclear data through improvements to the nuclear data pipeline. This presentation includes integral experiments and a summarization of existing copper nuclear data.

Cu63↗

Leveraging AI and Spatial Data to Unlock Pipeline Integrity Insights: NETL’s Advanced Infrastructure Integrity Model (AIIM)

Maintaining the integrity of natural gas infrastructure plays a critical role in ensuring energy security. Robust, data-driven foundational AI models for pipeline integrity can help address risk management and mitigation issues. Trusted foundational models can help with industry adoption and accelerate innovation by enhancing integrity predictions, reduce costs, and informing infrastructure build-out. The AIIM dashboard was released in 2022 and utilizes multi-ML models for ensemble-type insights. It was expanded to include analytics on reported incidents. It was developed as an ESRI Dashboard to support data visualization & interrogation and contains pipeline data and model results.

Advanced Infrastructure Integrity Model (AIIM)↗

FloodPlanet: High-Resolution Commercial Imagery for Training and Validation of Deep Learning-Based Models of Inundation Extent

Flooding events are becoming increasingly frequent worldwide and are known to cause extensive damage. Public optical and radar satellite imagery can be used to detect large areas of inundation in rural areas, however, long revisit times and coarse spatial resolution limit applications for short-lived events and urban areas. Commercial constellations such as those operated by Planet offer increased spatial and temporal resolution and can supplement mapping efforts to provide more information to disaster response, relief, and mitigation efforts. Deep learning requires high quality labeled data for training across coincident sensors. The FloodPlanet dataset presented here contains labeled surface water for 18 events across the world based on Planetscope imagery with coincident Harmonized Landsat Sentinel-2 ( HLS) or Sentinel-1 and builds upon the previously existing Sen1Floods11, xBD, and NASA Sentinel-1 datasets. Sen1Floods11 includes 4,831 512x512 pixel overlapping tiles of coincident Sentinel-1 and Sentinel-2 data observing 11 flood events across the world from 2017-2019. The dataset contains a combination of automated and hand-labeled surface water for use in training and validation of inundation modeling efforts. The xBD dataset identifies flood-damaged buildings and indicates the scale of damage to each (none, minor, moderate, and major) from four flood events which occurred in the United States, India, Nepal, and Bangladesh from the same time period. The NASA dataset contains hand-labeled water bodies observed in Sentinel-1 imagery during five flood events within the 2017-2019 period. The effort presented here utilizes observations from these previously investigated flood events to generate labels of surface water at the 3-5m spatial resolution provided by Planetscope and facilitate the comparison between public and commercial data. A data pipeline was built which uses clustering algorithms to pick the most suitable overlapping chips between the public data and PlanetScope data for manual labeling. Labels were created manually using NASA’s ImageLabeler tool and include areas of high- and low-confidence water. The high confidence designation is reserved for areas of open, unobstructed water while low confidence is used for areas of suspected water beneath vegetation, clouds, or cloud shadows. Expected to be released in late 2022, the FloodPlanet dataset will include tiled imagery with a unique ID for each 1024x1024 pixel tile, 7 bands of HLS data, and high- and low-confidence flood labels in both shapefile and tiff formats. The authors will follow Spatial Temporal Access Catalog (STAC) guidelines to release FloodPlanet on the Radiant Earth ML hub, which hosts public datasets for machine learning.

Alexander Melancon↗

Kepler Mission's Focal Plane Characterization Models Implementation

The Kepler Mission photometer is an unusually complex array of CCDs. A large number of time-varying instrumental and systemic effects must be modeled and removed from the Kepler pixel data to produce light curves of sufficiently high quality for the mission to be successful in its planet-finding objective. After the launch of the spacecraft, many of these effects are difficult to remeasure frequently, and various interpolations over a small number of sample measurements must be used to determine the correct value of a given effect at different points in time. A library of software modules, called Focal Plane Characterization (FC) Models, is the element of the Kepler Science Data Pipeline (hereafter "pipeline") that handles this. FC, or products generated by FC, are used by nearly every element of the SOC processing chain. FC includes Java components: database persistence classes, operations classes, model classes, and data importers; and MATLAB code: model classes, interpolation methods, and wrapper functions. These classes, their interactions, and the database tables they represent, are discussed. This paper describes how these data and the FC software work together to provide the pipeline with the correct values to remove non-photometric effects caused by the photometer and its electronics from the Kepler light curves. The interpolation mathematics is reviewed, as well as the special case of the sky-to-pixel,pixel-to-sky coordinate transformation code, which incorporates a compound model that is unique in the SOC software.

mission↗

The Zwicky Transient Facility: System Overview, Performance, and First Results

The Zwicky Transient Facility (ZTF) is a new optical time-domain survey that uses the Palomar 48 inch Schmidt telescope. A custom-built wide-field camera provides a 47 deg ^(2) field of view and 8 s readout time, yielding more than an order of magnitude improvement in survey speed relative to its predecessor survey, the Palomar Transient Factory. We describe the design and implementation of the camera and observing system. The ZTF data system at the Infrared Processing and Analysis Center provides near-real-time reduction to identify moving and varying objects. We outline the analysis pipelines, data products, and associated archive. Finally, we present on-sky performance analysis and first scientific results from commissioning and the early survey. ZTF’s public alert stream will serve as a useful precursor for that of the Large Synoptic Survey Telescope.

Eric C. Bellm↗

Ramdb: The NASA Raman Spectral Database (version 1.00).

Given that, in most instances, minimal sample preparation is required and due to its contactless instrument design, Raman spectroscopy is one of the most versatile vibrational spectroscopic techniques for the chemical analysis of environmental and biological specimens. The diversity of applications of Raman spectroscopy ranges anywhere from art [1] to planetary science missions [2]. The advancement in the use of Raman spectroscopy in Solar System missions, notably in post-mission sample return analysis, requires a spectral library holding the broad range of specimens that could be found in Solar System sources. For this purpose, we have initiated the development of a Raman spectral database (Ramdb) at NASA Ames Research Center. Currently, the database includes experimental and theoretical Raman spectra of PAHs [3, 4], as well as laboratory Raman spectra of amino acids, carbon allotropes, minerals, and analogs relevance to Earth Sciences [5], Exobiology [6], Planetary [7], and Astrochemistry [8] to name just a few examples. Ramdb can be found on the web at www.astrochemistry.org/ramdb, where raw and processed Raman spectra can be downloaded in CSV format. The laboratory Raman spectra are measured using a laser Raman spectrometer (JASCO NRS-5500-532QRI). The Raman instrument is equipped with three excitation lasers, with wavelengths of 405, 532, and 785 nm. A clean silicon substrate is used as the internal standard for wavenumber calibration. Powdered samples were prepared (microscopic >10 um, grounded microscopic < 10 um) on glass slides. Some raw data exhibited a background signal arising as a combination of laser-induced fluorescence from the sample. To correct this background, we developed a Python pipeline that uses open-source Python libraries. Ramdb provides both raw and processed (using Python pipeline) data, which includes tabulated Raman shift transitions and other measurement details. The theoretical Raman band positions of PAHs (pyrene monomers and tetramer clusters) were computed using density functional theory (DFT) with the help of the Gaussian 16 suite of programs [9]. In the near future, Ramdb will serve as a repository of Raman spectral data from Laboratory Astrophysics and Planetary Science experiments involving the irradiation of organic compounds under simulated space and planetary conditions. In addition, online and offline tools will be developed for utilising the database for comparison to the user’s sample.

N Punnakayathil↗

Navigating Exascale Operational Data Analytics: From Inundation to Insight

In this paper, we address the challenges in achieving sustainable data-driven efficiency by providing a detailed exploration of the end-to-end operational data analytics (ODA) framework that evolved through two generations of supercomputer systems at the Oak Ridge Leadership Computing Facility (OLCF). This framework addresses large data streams ingested from heavily instrumented HPC environment that accumulates multi-terabytes per day. We outline the multifaceted data life cycle across HPC procurement, operations, and research & development, identifying key obstacles and design decisions that shape effective strategies in building and supporting data pipelines end-to-end. By sharing key insights and lessons learned from our experience, we offer recommendations for the HPC community on enabling sustainable operational data analytics and beyond. Our contributions aim to bridge the gap between potential and real benefits of operational data, guiding future efforts towards integrated and sustainable operational intelligence in high-performance computing environments.

Shin, Woong↗

Avoiding and tolerating latency in large-scale next-generation shared-memory multiprocessors

A scalable solution to the memory-latency problem is necessary to prevent the large latencies of synchronization and memory operations inherent in large-scale shared-memory multiprocessors from reducing high performance. We distinguish latency avoidance and latency tolerance. Latency is avoided when data is brought to nearby locales for future reference. Latency is tolerated when references are overlapped with other computation. Latency-avoiding locales include: processor registers, data caches used temporally, and nearby memory modules. Tolerating communication latency requires parallelism, allowing the overlap of communication and computation. Latency-tolerating techniques include: vector pipelining, data caches used spatially, prefetching in various forms, and multithreading in various forms. Relaxing the consistency model permits increased use of avoidance and tolerance techniques. Each model is a mapping from the program text to sets of partial orders on program operations; it is a convention about which temporal precedences among program operations are necessary. Information about temporal locality and parallelism constrains the use of avoidance and tolerance techniques. Suitable architectural primitives and compiler technology are required to exploit the increased freedom to reorder and overlap operations in relaxed models.

Probst, David K.↗

Real-Time Reed-Solomon Decoder

RS decoder uses dedicated hardware and data pipelining for high-speed operation. Parallel processing techniques provide equivalent of over one billion operations per second at one step in decoding. Decoder finds commercial application in data encoding/decoding, telemetry, and radio communications.

Lahmeyer, C. R.↗

Primary Mission Threshold Crossing Events in the TESS SPOC Transit Search

We present an overview of the single- and multiple-sector results of the Science Processing Operations Center (SPOC) transit search in the primary Transiting Exoplanet Survey Satellite (TESS) mission. TESS was designed to survey bright stars in the greater Solar neighborhood in search of transiting exoplanets. Data were acquired at a 2-minute cadence for 16,000-20,000 pre-selected target stars in each 28-day observation sector and processed in the SPOC pipeline at NASA Ames Research Center. The photometry pipeline produced a systematic error corrected light curve for each target star. Light curves were searched for transiting planet signatures by sector for all target stars, and separately for target stars observed in multiple sectors. Potential transit signals for which the transiting planet detection threshold was exceeded and a series of transit consistency tests were passed are referred to as Threshold Crossing Events (TCEs). We highlight the full TCE population and the population of SPOC TCEs that were later identified as TESS Objects of Interest (TOIs). Characteristics of the TCE populations implied by limb-darkened transiting planet model fits are also presented. SPOC pipeline data products are delivered to the Mikulski Archive for Space Telescopes (MAST)(http://archive.stsci.edu/missions-and-data/tess) for access by the community. Funding for the TESS Mission has been provided by the NASA Science Mission Directorate.

TESS↗

Online Detection of Power Grid Anomalies via Federated Learning

Data from sensors is critical for advanced applica- tions that support efficient, reliable, and resilient electric grid operations. Historically, data from phasor measurement units (PMU) has been utilized to develop a wide variety of wide area control and protection applications suitable for power grid control centers. However, until now, most of these could not be deployed for automated operations due to a set of data corruption challenges and uncertainty in the incoming data pipeline. In this paper, we address the problem of detecting different variety of anomalies that are evident in different high- speed power grid measurements. The paper discusses a workflow for handling problems with data acquisition and highlights some of the key findings suitable for anomaly detection in a centralized and distributed environment. The effectiveness of the proposed method was demonstrated with results utilizing realistic PMU datasets

Shinkle, Matthew W.↗

Automated pipeline processing X-ray diffraction data from dynamic compression experiments on the Extreme Conditions Beamline of PETRA III

Presented and discussed here is the implementation of a software solution that provides prompt X-ray diffraction data analysis during fast dynamic compression experiments conducted within the dynamic diamond anvil cell technique. It includes efficient data collection, streaming of data and metadata to a high-performance cluster (HPC), fast azimuthal data integration on the cluster, and tools for controlling the data processing steps and visualizing the data using the DIOPTAS software package. This data processing pipeline is invaluable for a great number of studies. The potential of the pipeline is illustrated with two examples of data collected on ammonia–water mixtures and multiphase mineral assemblies under high pressure. The pipeline is designed to be generic in nature and could be readily adapted to provide rapid feedback for many other X-ray diffraction techniques, e.g. large-volume press studies, in situ stress/strain studies, phase transformation studies, chemical reactions studied with high-resolution diffraction etc.

97 MATHEMATICS AND COMPUTING↗

A statistical-based scheduling algorithm in automated data path synthesis

In this paper, we propose a new heuristic scheduling algorithm based on the statistical analysis of the cumulative frequency distribution of operations among control steps. It has a tendency of escaping from local minima and therefore reaching a globally optimal solution. The presented algorithm considers the real world constraints such as chained operations, multicycle operations, and pipelined data paths. The result of the experiment shows that it gives optimal solutions, even though it is greedy in nature.

Jeon, Byung Wook↗

Processors, Pipelines, and Protocols for Advanced Modeling Networks

Predictive capabilities arise from our understanding of natural processes and our ability to construct models that accurately reproduce these processes. Although our modeling state-of-the-art is primarily limited by existing computational capabilities, other technical areas will soon present obstacles to the development and deployment of future predictive capabilities. Advancement of our modeling capabilities will require not only faster processors, but new processing algorithms, high-speed data pipelines, and a common software engineering framework that allows networking of diverse models that represent the many components of Earth's climate and weather system. Development and integration of these new capabilities will pose serious challenges to the Information Systems (IS) technology community. Designers of future IS infrastructures must deal with issues that include performance, reliability, interoperability, portability of data and software, and ultimately, the full integration of various ES model systems into a unified ES modeling network.

Coughlan, Joseph↗

The DECADE cosmic shear project III: validation of analysis pipeline using spatially inhomogeneous data

We present the pipeline for the cosmic shear analysis of the Dark Energy Camera All Data Everywhere (DECADE) weak lensing dataset: a catalog consisting of 107 million galaxies observed by the Dark Energy Camera (DECam) in the northern Galactic cap. The catalog derives from a large number of disparate observing programs and is therefore more inhomogeneous across the sky compared to existing lensing surveys. First, we use simulated data-vectors to show the sensitivity of our constraints to different analysis choices in our inference pipeline, including sensitivity to residual systematics. Next we use simulations to validate our covariance modeling for inhomogeneous datasets. Finally, we show that our choices in the end-to-end cosmic shear pipeline are robust against inhomogeneities in the survey, by extracting relative shifts in the cosmology constraints across different subsets of the footprint/catalog and showing they are all consistent within 1σ to 2σ. This is done for forty-six subsets of the data and is carried out in a fully consistent manner: for each subset of the data, we re-derive the photometric redshift estimates, shear calibrations, survey transfer functions, the data vector, measurement covariance, and finally, the cosmological constraints. Our results show that existing analysis methods for weak lensing cosmology can be fairly resilient towards inhomogeneous datasets. This also motivates exploring a wider range of image data for pursuing such cosmological constraints.

79 ASTRONOMY AND ASTROPHYSICS↗

Using Quality Attributes to Bridge Systems Engineering Gaps : A Juno Ground Data Systems Case Study

The Juno Mission to Jupiter is the second mission selected by the NASA New Frontiers Program. Juno launched August 2011 and will reach Jupiter July 2016. Juno's payload system is composed of nine instruments plus a gravity science experiment. One of the primary functions of the Juno Ground Data System (GDS) is the assembly and distribution of the CFDP (CCSDS File Delivery Protocol) product telemetry, also referred to as raw science data, for eight out of the nine instruments. The GDS accomplishes this with the Instrument Data Pipeline (IDP). During payload integration, the first attempt to exercise the IDP in a flight like manner revealed that although the functional requirements were well understood, the system was unable to meet latency requirements with the as-is heritage design. A systems engineering gap emerged between Juno instrument data delivery requirements and the assumptions behind the heritage flight-ground interactions. This paper describes the use of quality attributes to measure and overcome this gap by introducing a new systems engineering activity, and a new monitoring service architecture that successfully delivered the performance metrics needed to validate Juno IDP.

Ground Data Systems (GDS)↗

Embedded FPGA developments in 130 nm and 28 nm CMOS for machine learning in particle detector readout

Embedded field programmable gate array (eFPGA) technology allows the implementation of reconfigurable logic within the design of an application-specific integrated circuit (ASIC). This approach offers the low power and efficiency of an ASIC along with the ease of FPGA configuration, particularly beneficial for the use case of machine learning in the data pipeline of next-generation collider experiments. An open-source framework called "FABulous" was used to design eFPGAs using 130 nm and 28 nm CMOS technology nodes, which were subsequently fabricated and verified through testing. The capability of an eFPGA to act as a front-end readout chip was assessed using simulation of high energy particles passing through a silicon pixel sensor. A machine learning-based classifier, designed for reduction of sensor data at the source, was synthesized and configured onto the eFPGA. A successful proof-of-concept was demonstrated through reproduction of the expected algorithm result on the eFPGA with perfect accuracy. Finally, further development of the eFPGA technology and its application to collider detector readout is discussed.

47 OTHER INSTRUMENTATION↗

Importance of Higher Fidelity Model Geometries during Optimization of Critical Experiments

PARADIGM, PARallel Approach of Differential and InteGral Measurements, is a cross-collaborative effort at Los Alamos National Laboratory between nuclear data theorists, differential and integral experimenters, as well as machine learning statisticians to tackle uncertainties in the intermediate region of 239 Pu. In essence, the idea behind PARADIGM is to remove the linear conceptualization of the nuclear data pipeline, shown in Figure 1, and replace it with a far more parallelized approach. The novel approach leverages machine learning to guide which differential measurements and integral experiments will result in the largest decrease in uncertain ties for a nuclide reaction pair in a given energy range. The concept builds off earlier work, EUCLID, which focused on the fast region of 239 Pu. The practical benefit of having evaluation, differential measurement, and integral experiment personnel in collaboration with machine learning is to represent the entire nuclear data in one snapshot. This enable large reduction in the time to deliver improved nuclear data, which using the PARADIGM approach could be done in 3 years. A general outline of PARADIGM and specific topics are available in other papers. The discussion here will pertain directly to the integral experiment design. More specifically, the process of taking a rough design and transforming it into a finalized neutronic model will be discussed.

97 MATHEMATICS AND COMPUTING↗