Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data reduction pipelines”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Infrared Spectroscopy of Star Formation in Galactic and Extragalactic Regions

In this program we proposed to perform a series of spectroscopic studies, including data analysis and modeling, of star formation regions using an ensemble of archival space-based data from the Infrared Space Observatory's Long Wavelength Spectrometer and Short Wavelength Spectrometer, and to take advantage of other spectroscopic databases including the first results from SIRTF. Our emphasis has been on star formation in external, bright IR galaxies, but other areas of research have included young, low or high mass pre-main sequence stars in star formation regions, and the galactic center. The OH lines in the far infrared were proposed as one key focus of this inquiry, because the Principal Investigator (H. Smith) had a full set of OH IR lines from IS0 observations. It was planned that during the proposed 2-1/2 year timeframe of the proposal other data (including perhaps from SIRTF) would become available, and we intended to be responsive to these and other such spectroscopic data sets. The program has the following goals: 1) Refine the data analysis of IS0 observations to obtain deeper and better SNR results on selected sources. The IS0 data itself underwent pipeline 10 reductions in early 2001, and the more 'hands-on data reduction packages' have been released. The IS0 Fabry-Perot database is particularly sensitive to noise and can have slight calibration errors, and improvements are anticipated. We plan to build on these deep analysis tools and contribute to their development. Model the atomic and molecular line shapes, in particular the OH lines, using revised montecarlo techniques developed by the Submillimeter Wave Astronomy Satellite (SWAS) team at the Center for Astrophysics. 2) 3) Use newly acquired space-based SIRTF or SOFIA spectroscopic data as they become available, and contribute to these observing programs as appropriate. 4) Attend scientific meetings and workshops. 5) E&PO activities, especially as related to infrared astrophysics and/or spectroscopy.

Smith, Howard A.

A long-term space astrophysics research program: An x-ray perspective of the components and structure of galaxies

X-ray studies of galaxies by the Smithsonian Astrophysical Observatory (SAO) and MIT are described. Activities at SAO include ROSAT PSPC x-ray data reduction and analysis pipeline; x-ray sources in nearby Sc galaxies; optical, x-ray, and radio study of ongoing galactic merger; a radio, far infrared, optical, and x-ray study of the Sc galaxy NGC247; and a multiparametric analysis of the Einstein sample of early-type galaxies. Activities at MIT included continued analysis of observations with ROSAT and ASCA, and continued development of new approaches to spectral analysis with ASCA and AXAF. Also, a new method for characterizing structure in galactic clusters was developed and applied to ROSAT images of a large sample of clusters. An appendix contains preprints generated by the research.

Fabbiano, G.

Wide-field Hard X-ray Survey with ART-XC

The Astronomical Roentgen Telescope X-ray Concentrator (ART-XC) instrument onboard the Spectrum Röntgen Gamma (SRG) mission began the 4-year all-sky hard X-ray survey since December 2019 and has already completed four scans of the sky. The observations of the ecliptic pole regions will reach exceptional depth thanks to the survey design of overlapping exposure in these regions. We will discuss the progress of the ART-XC survey in the North Ecliptic Pole (NEP) region so far, including the status of the data reduction and analysis pipeline as well as the survey results.

Chien-Ting Chen

The Caltech-NRAO Stripe 82 Survey (CNSS) Paper. I. The Pilot Radio Transient Survey in 50 Deg.(exp. 2)

We have commenced a multiyear program, the Caltech-NRAO Stripe 82 Survey (CNSS), to search for radio transients with the Jansky VLA in the Sloan Digital Sky Survey Stripe 82 region. The CNSS will deliver five epochs over the entire approx. 270 deg.(exp. 2) of Stripe 82, an eventual deep combined map with an rms noise of approx. 40 proper motion epoch y and catalogs at a frequency of 3 GHz, and having a spatial resolution of 3 inches. This first paper presents the results from an initial pilot survey of a 50 deg.(exp. 2) region of Stripe 82, involving four epochs spanning logarithmic timescales between 1 week and 1.5 yr, with the combined map having a median rms noise of 35 proper motion epoch y. This pilot survey enabled the development of the hardware and software for rapid data processing, as well as transient detection and follow-up, necessary for the full 270 deg.(exp. 2) survey. Data editing, calibration, imaging, source extraction, cataloging, and transient identification were completed in a semi-automated fashion within 6 hr of completion of each epoch of observations, using dedicated computational hardware at the NRAO in Socorro and custom-developed data reduction and transient detection pipelines. Classification of variable and transient sources relied heavily on the wealth of multiwavelength legacy survey data in the Stripe 82 region, supplemented by repeated mapping of the region by the Palomar Transient Factory. A total of 3.9(+0.5%/-0.9%) of the few thousand detected point sources werefound to vary by greater than 30%, consistent with similar studies at 1.4 and 5 GHz. Multiwavelength photometric data and light curves suggest that the variability is mostly due to shock-induced flaring in the jets of active galactic nuclei (AGNs). Although this was only a pilot survey, we detected two bona fide transients, associated with an RS CVn binary and a dKe star. Comparison with existing legacy survey data (FIRST, VLA-Stripe 82) revealed additional highly variable and transient sources on timescales between 5 and 20 yr, largely associated with renewed AGN activity. The rates of such AGNs possibly imply episodes of enhanced accretion and jet activity occurring once every approx. 40,000 yr in these galaxies. We compile the revised radio transient rates and make recommendations for future transient surveys and joint radio-optical experiments.

galaxies: active – radio continuum: galaxies –

The HEASARC Swift Gamma-Ray Burst Archive: The Pipeline and the Catalog

Since its launch in late 2004, the Swift satellite triggered or observed an average of one gamma-ray burst (GRB) every 3 days, for a total of 771 GRBs by 2012 January. Here, we report the development of a pipeline that semi automatically performs the data-reduction and data-analysis processes for the three instruments on board Swift (BAT, XRT, UVOT). The pipeline is written in Perl, and it uses only HEAsoft tools and can be used to perform the analysis of a majority of the point-like objects (e.g., GRBs, active galactic nuclei, pulsars) observed by Swift. We run the pipeline on the GRBs, and we present a database containing the screened data, the output products, and the results of our ongoing analysis. Furthermore, we created a catalog summarizing some GRB information, collected either by running the pipeline or from the literature. The Perl script, the database, and the catalog are available for downloading and querying at the HEASARC Web site.

BURST ARCHIVE

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)

The Zwicky Transient Facility: System Overview, Performance, and First Results

The Zwicky Transient Facility (ZTF) is a new optical time-domain survey that uses the Palomar 48 inch Schmidt telescope. A custom-built wide-field camera provides a 47 deg ^(2) field of view and 8 s readout time, yielding more than an order of magnitude improvement in survey speed relative to its predecessor survey, the Palomar Transient Factory. We describe the design and implementation of the camera and observing system. The ZTF data system at the Infrared Processing and Analysis Center provides near-real-time reduction to identify moving and varying objects. We outline the analysis pipelines, data products, and associated archive. Finally, we present on-sky performance analysis and first scientific results from commissioning and the early survey. ZTF’s public alert stream will serve as a useful precursor for that of the Large Synoptic Survey Telescope.

Eric C. Bellm

PSTN-019: The LSST Science Pipelines Software: Optical Survey Pipeline Reduction and Analysis Environment

The NSF-DOE Vera C. Rubin Observatory is executing the Legacy Survey of Space and Time (LSST) as its prime mission, producing a series of data releases over the ten-year survey. The LSST Science Pipelines Software will be used to create these data releases and to perform the nightly prompt processing and alert production. This paper provides an overview of the LSST Science Pipelines Software, describing the components and their integration into pipelines that generate science-ready data products.

79 ASTRONOMY AND ASTROPHYSICS

pySimpleMask

SF-26-118 pySimpleMask is a tool for creating masks and Q-partition maps for X-ray scattering patterns, supporting SAXS, WAXS, and XPCS data reduction. It ships both a desktop GUI and a headless Python API that can drive the full pipeline from scripts.

Chu, Miaoqi [Argonne National Laboratory (ANL), Ar

Pixel-Level Calibration in the Kepler Science Operations Center Pipeline

We present an overview of the pixel-level calibration of flight data from the Kepler Mission performed within the Kepler Science Operations Center Science Processing Pipeline. This article describes the calibration (CAL) module, which operates on original spacecraft data to remove instrument effects and other artifacts that pollute the data. Traditional CCD data reduction is performed (removal of instrument/detector effects such as bias and dark current), in addition to pixel-level calibration (correcting for cosmic rays and variations in pixel sensitivity), Kepler-specific corrections (removing smear signals which result from the lack of a shutter on the photometer and correcting for distortions induced by the readout electronics), and additional operations that are needed due to the complexity and large volume of flight data. CAL operates on long (~30 min) and short (~1 min) sampled data, as well as full-frame images, and produces calibrated pixel flux time series, uncertainties, and other metrics that are used in subsequent Pipeline modules. The raw and calibrated data are also archived in the Multi-mission Archive at Space Telescope at the Space Telescope Science Institute for use by the astronomical community.

Calibration

Validation and Calibration of Energy Models with Real Vehicle Data from Chassis Dynamometer Experiments

Accurate estimation of vehicle fuel consumption typically requires detailed modeling of complex internal powertrain dynamics, often resulting in computationally intensive simulations. However, many transportation applications-such as traffic flow modeling, optimization, and control-require simplified models that are fast, interpretable, and easy to implement, while still maintaining fidelity to physical energy behavior. This work builds upon a recently developed model reduction pipeline that derives physics-like energy models from high-fidelity Autonomie vehicle simulations. These reduced models preserve essential vehicle dynamics, enabling realistic fuel consumption estimation with minimal computational overhead. While the reduced models have demonstrated strong agreement with their Autonomie counterparts, previous validation efforts have been confined to simulation environments. This study extends the validation by comparing the reduced energy model's outputs against real-world vehicle data. Focusing on the MidSUV category, we tune the baseline Autonomie model to closely replicate the characteristics of a Toyota RAV4. We then assess the accuracy of the resulting reduced model in estimating fuel consumption under actual drive conditions. Our findings suggest that, when the reference Autonomie model is properly calibrated, the simplified model produced by the reduction pipeline can provide reliable, semi-principled fuel rate estimates suitable for large-scale transportation applications.

42 ENGINEERING

Data handling for the geometric correction of large images

Several geometric distortions are present in remotely sensed images depending on the type of sensors and the object being observed. It is often desirable to compensate for these distortions and store the images in reference to a standard coordinate system. Digital techniques for correction are versatile and introduce a minimum of radiometric errors. The main problems to be considered in this area are the determination of the corrective transformation, resampling, and the management of the large quantities of data. It is shown that, by a judicious rearrangement of the input data, considerable reductions in the required memory capacity can be achieved. The rearrangement can be accomplished in several stages. The method presented here is amenable to pipeline implementation for processing a continuous stream of images.

Ramapriyan, H. K.

A 32-bit Ultrafast Parallel Correlator using Resonant Tunneling Devices

An ultrafast 32-bit pipeline correlator has been implemented using resonant tunneling diodes (RTD) and hetero-junction bipolar transistors (HBT). The negative differential resistance (NDR) characteristics of RTD's is the basis of logic gates with the self-latching property that eliminates pipeline area and delay overheads which limit throughput in conventional technologies. The circuit topology also allows threshold logic functions such as minority/majority to be implemented in a compact manner resulting in reduction of the overall complexity and delay of arbitrary logic circuits. The parallel correlator is an essential component in code division multi-access (CDMA) transceivers used for the continuous calculation of correlation between an incoming data stream and a PN sequence. Simulation results show that a nano-pipelined correlator can provide and effective throughput of one 32-bit correlation every 100 picoseconds, using minimal hardware, with a power dissipation of 1.5 watts. RTD plus HBT based logic gates have been fabricated and the RTD plus HBT based correlator is compared with state of the art complementary metal oxide semiconductor (CMOS) implementations.

Kulkarni, Shriram

Early-type galaxies: Automated reduction and analysis of ROSAT PSPC data

Preliminary results of early-type galaxies that will be part of a galaxy catalog to be derived from the complete Rosat data base are presented. The stored data were reduced and analyzed by an automatic pipeline. This pipeline is based on a command language scrip. The important features of the pipeline include new data time screening in order to maximize the signal to noise ratio of faint point-like sources, source detection via a wavelet algorithm, and the identification of sources with objects from existing catalogs. The pipeline outputs include reduced images, contour maps, surface brightness profiles, spectra, color and hardness ratios.

Mackie, G.

Central Stars of Planetary Nebulae in the SMC

In FUSE cycle 3's program C056 we studied four Central Stars of Planetary Nebulae (CSPN) in the Small Magellanic Could. All FUSE observations have been successfully completed and have been reduced and analyzed. The observation of one object (SMP SMC 5) appeared to be off-target and no useful stellar flux was gathered. For another observation (SMP SMC 1) the voltage problems resulted in the loss of data from one of the SiC detectors, but we were still able to analyze the remaining data. The analysis and the results are summarized below. The FUSE data were reduced using the latest available version of the FUSE calibration pipeline (CALFUSE v2.4). The flux of these SMC post-AGB objects is at the threshold of FUSE S sensitivity, and the targets required many orbit-long exposures, each of which typically had low (target) count-rates. The background subtraction required special care during the reduction, and was done in a similar manner to our FUSE cycle 2 BOO1 objects. The resulting calibrated data from the different channels were compared in the overlapping regions for consistency. The final combined, extracted spectra of each target was then modeled to determine the stellar and nebular parameters. The FUSE spectra, combined with archival HST spectra, have been analyzed using stellar atmospheres codes such as TLUSTY and CMFGEN to derive photospheric and wind parameters of the central stars, and with ISM models to determine the amount and temperature of the surrounding atomic and molecular hydrogen. We have combined these results with those of our cycle 4 (D034) program (CSPN of the LMC) in Herald & Bianchi 2004a (paper in preparation, will be submitted to ApJ in June 2004). Two of the three SMC objects analyzed were found to have significantly lower stellar temperatures than had been predicted using nebular photoionization models, indicating either a hotter ionizing companion or the existence of strong shocks in the nebular environment. The analysis also revealed that some objects are surrounded by significant quantities of hot (e.g., 1000- 2000 K) molecular H2, similar to what we found for some LMC and Galactic CSPN (Herald & Bianchi 2002, 2003a, b, 2004b, c, d).

Bianchi, Luciana

The European Southern Observatory-MIDAS table file system

The new and substantially upgraded version of the Table File System in MIDAS is presented as a scientific database system. MIDAS applications for performing database operations on tables are discussed, for instance, the exchange of the data to and from the TFS, the selection of objects, the uncertainty joins across tables, and the graphical representation of data. This upgraded version of the TFS is a full implementation of the binary table extension of the FITS format; in addition, it also supports arrays of strings. Different storage strategies for optimal access of very large data sets are implemented and are addressed in detail. As a simple relational database, the TFS may be used for the management of personal data files. This opens the way to intelligent pipeline processing of large amounts of data. One of the key features of the Table File System is to provide also an extensive set of tools for the analysis of the final results of a reduction process. Column operations using standard and special mathematical functions as well as statistical distributions can be carried out; commands for linear regression and model fitting using nonlinear least square methods and user-defined functions are available. Finally, statistical tests of hypothesis and multivariate methods can also operate on tables.

Peron, M.

Embedded FPGA developments in 130 nm and 28 nm CMOS for machine learning in particle detector readout

Embedded field programmable gate array (eFPGA) technology allows the implementation of reconfigurable logic within the design of an application-specific integrated circuit (ASIC). This approach offers the low power and efficiency of an ASIC along with the ease of FPGA configuration, particularly beneficial for the use case of machine learning in the data pipeline of next-generation collider experiments. An open-source framework called "FABulous" was used to design eFPGAs using 130 nm and 28 nm CMOS technology nodes, which were subsequently fabricated and verified through testing. The capability of an eFPGA to act as a front-end readout chip was assessed using simulation of high energy particles passing through a silicon pixel sensor. A machine learning-based classifier, designed for reduction of sensor data at the source, was synthesized and configured onto the eFPGA. A successful proof-of-concept was demonstrated through reproduction of the expected algorithm result on the eFPGA with perfect accuracy. Finally, further development of the eFPGA technology and its application to collider detector readout is discussed.

47 OTHER INSTRUMENTATION

LCLS Big Data Handling – How I Learned to Stop Worrying and Love the Data Deluge

Advanced data and computing systems are vital to Linac Coherent Light Source (LCLS) operations, data interpretation and overall scientific productivity. The transition to MHz-era operation marks a fundamental change in scale that requires new infrastructure and architectures to link LCLS to the required scale of computing needed for scientific interpretation. The LCLS-II Data System meets big data challenges by implementing configurable data reduction that can adapt to multiple science areas, real-time analysis frameworks to provide visualization and fast feedback, and the ability to transfer data to local and remote computational facilities for near real time analysis at the appropriate scale. Feature extracted information generated in the data analysis pipeline - at the edge, local compute, or remote High-Performance Computing (HPC) resources - can be used to steer experiments and inform user decisions during beam time. Artificial Intelligence and Machine Learning (AI/ML) techniques present new opportunities to rapidly analyse large datasets and direct experiments, but create new challenges in scaling, adaptability, complexity, and trustworthiness. We describe how the LCLS-II Data System architecture addresses its data-driven challenges in the areas of data acquisition, data processing, data management, and workflow orchestration to decrease the overall time-to-science and provide a vision for future developments.

artificial intelligence