Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Astronomy data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Inference of the optical depth to reionization τ from Planck CMB maps with convolutional neural networks

The optical depth to reionization, τ, is the least constrained parameter of the cosmological Λ cold dark matter (ΛCDM) model. To date, its most precise value is inferred from large-scale polarized cosmic microwave background (CMB) power spectra from the High Frequency Instrument (HFI) aboard the Planck satellite. These maps are known to contain significant contamination by residual non-Gaussian systematic effects, which are hard to model analytically. Therefore, robust constraints on τ are currently obtained through an empirical cross-spectrum likelihood built from simulations. In this paper, we present a likelihood-free inference of τ from polarized Planck HFI maps which, for the first time, is fully based on neural networks (NNs). NNs have the advantage of not requiring an analytical description of the data and can be trained on state-of-the-art simulations, combining the information from multiple channels. By using Gaussian sky simulations and Planck SRoll2 simulations, including CMB, noise, and residual instrumental systematic effects, we trained, tested, and validated NN models considering different setups. We inferred the value of τ directly from Stokes Q and U maps at ~4° pixel resolution, without computing angular power spectra. On Planck data, we obtained τ NN = 0.0579 ± 0.0082, which is compatible with current EE cross-spectrum results but with a ~30% larger uncertainty, which can be assigned to the inherent nonoptimality of our estimator and to the retraining procedure applied to avoid biases. While this paper does not improve on current cosmological constraints on τ, our analysis represents a first robust application of NN-based inference on real data, and highlights its potential as a promising tool for complementary analysis of near-future CMB experiments, also in view of the ongoing challenge to achieve the first detection of primordial gravitational waves.

79 ASTRONOMY AND ASTROPHYSICS↗

The eROSITA Final Equatorial-Depth Survey (eFEDS): Identification and characterization of the counterparts to point-like sources

In November 2019, eROSITA on board of the Spektrum-Roentgen-Gamma (SRG) observatory started to map the entire sky in X-rays. After the four-year survey program, it will reach a flux limit that is about 25 times deeper than ROSAT. During the SRG performance verification phase, eROSITA observed a contiguous 140 deg 2 area of the sky down to the final depth of the eROSITA all-sky survey (eROSITA Final Equatorial-Depth Survey; eFEDS), with the goal of obtaining a census of the X-ray emitting populations (stars, compact objects, galaxies, clusters of galaxies, and active galactic nuclei) that will be discovered over the entire sky. This paper presents the identification of the counterparts to the point sources detected in eFEDS in the main and hard samples and their multi-wavelength properties, including redshift. To identify the counterparts, we combined the results from two independent methods (NWAY and ASTROMATCH), trained on the multi-wavelength properties of a sample of 23k XMM-Newton sources detected in the DESI Legacy Imaging Survey DR8. Then spectroscopic redshifts and photometry from ancillary surveys were collated to compute photometric redshifts. Of the eFEDS sources, 24 774 of 27 369 have reliable counterparts (90.5%) in the main sample and 231 of 246 sources (93.9%) have counterparts in the hard sample, including 2514 (3) sources for which a second counterpart is equally likely. By means of reliable spectra, Gaia parallaxes, and/or multi-wavelength properties, we have classified the reliable counterparts in both samples into Galactic (2695) and extragalactic sources (22 079). For about 340 of the extragalactic sources, we cannot rule out the possibility that they are unresolved clusters or belong to clusters. Inspection of the distributions of the X-ray sources in various optical/IR colour-magnitude spaces reveal a rich variety of diverse classes of objects. The photometric redshifts are most reliable within the KiDS/VIKING area, where deep near-infrared data are also available. This paper accompanies the eROSITA early data release of all the observations performed during the performance and verification phase. Together with the catalogues of primary and secondary counterparts to the main and hard samples of the eFEDS survey, this paper releases their multi-wavelength properties and redshifts.

79 ASTRONOMY AND ASTROPHYSICS↗

Informed total-error-minimizing priors: Interpretable cosmological parameter constraints despite complex nuisance effects

While Bayesian inference techniques are standard in cosmological analyses, it is common to interpret resulting parameter constraints with a frequentist intuition. This intuition can fail, for example, when marginalizing high-dimensional parameter spaces onto subsets of parameters, because of what has come to be known as projection effects or prior volume effects. We present the method of informed total-error-minimizing (ITEM) priors to address this problem. An ITEM prior is a prior distribution on a set of nuisance parameters, such as those describing astrophysical or calibration systematics, intended to enforce the validity of a frequentist interpretation of the posterior constraints derived for a set of target parameters (e.g., cosmological parameters). Our method works as follows. For a set of plausible nuisance realizations, we generate target parameter posteriors using several different candidate priors for the nuisance parameters. We reject candidate priors that do not accomplish the minimum requirements of bias (of point estimates) and coverage (of confidence regions among a set of noisy realizations of the data) for the target parameters on one or more of the plausible nuisance realizations. Of the priors that survive this cut, we select the ITEM prior as the one that minimizes the total error of the marginalized posteriors of the target parameters. As a proof of concept, we applied our method to the density split statistics measured in Dark Energy Survey Year 1 data. We demonstrate that the ITEM priors substantially reduce prior volume effects that otherwise arise and that they allow for sharpened yet robust constraints on the parameters of interest.

79 ASTRONOMY AND ASTROPHYSICS↗

Fast inference of Boosted Decision Trees in FPGAs for particle physics

We describe the implementation of Boosted Decision Trees in the hls4ml library, which allows the translation of a trained model into FPGA firmware through an automated conversion process. Thanks to its fully on-chip implementation, hls4ml performs inference of Boosted Decision Tree models with extremely low latency. With a typical latency less than 100 ns, this solution is suitable for FPGA-based real-time processing, such as in the Level-1 Trigger system of a collider experiment. These developments open up prospects for physicists to deploy BDTs in FPGAs for identifying the origin of jets, better reconstructing the energies of muons, and enabling better selection of rare signal processes.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Optimizing a magnitude-limited spectroscopic training sample for photometric classification of supernovae

ABSTRACT In preparation for photometric classification of transients from the Legacy Survey of Space and Time (LSST) we run tests with different training data sets. Using estimates of the depth to which the 4-m Multi-Object Spectroscopic Telescope (4MOST) Time Domain Extragalactic Survey (TiDES) can classify transients, we simulate a magnitude-limited sample reaching rAB ≈ 22.5 mag. We run our simulations with the software snmachine, a photometric classification pipeline using machine learning. The machine-learning algorithms struggle to classify supernovae when the training sample is magnitude limited, in contrast to representative training samples. Classification performance noticeably improves when we combine the magnitude-limited training sample with a simulated realistic sample of faint high-redshift supernovae observed from larger spectroscopic facilities; the algorithms’ range of average area under receiver operator characteristic curve (AUC) scores over 10 runs increases from 0.547–0.628 to 0.946–0.969 and purity of the classified sample reaches 95 per cent in all runs for two of the four algorithms. By creating new, artificial light curves using the augmentation software avocado, we achieve a purity in our classified sample of 95 per cent in all 10 runs performed for all machine-learning algorithms considered. We also reach a highest average AUC score of 0.986 with the artificial neural network algorithm. Having ‘true’ faint supernovae to complement our magnitude-limited sample is a crucial requirement in optimization of a 4MOST spectroscopic sample. However, our results are a proof of concept that augmentation is also necessary to achieve the best classification results.

79 ASTRONOMY AND ASTROPHYSICS↗

The miniJPAS survey quasar selection – II. Machine learning classification with photometric measurements and uncertainties

Astrophysical surveys rely heavily on the classification of sources as stars, galaxies, or quasars from multiband photometry. Surveys in narrow-band filters allow for greater discriminatory power, but the variety of different types and redshifts of the objects present a challenge to standard template-based methods. In this work, which is part of a larger effort that aims at building a catalogue of quasars from the miniJPAS survey, we present a machine learning-based method that employs convolutional neural networks (CNNs) to classify point-like sources including the information in the measurement errors. We validate our methods using data from the miniJPAS survey, a proof-of-concept project of the Javalambre Physics of the Accelerating Universe Astrophysical Survey (J-PAS) collaboration covering ∼1 deg 2 of the northern sky using the 56 narrow-band filters of the J-PAS survey. Due to the scarcity of real data, we trained our algorithms using mocks that were purpose-built to reproduce the distributions of different types of objects that we expect to find in the miniJPAS survey, as well as the properties of the real observations in terms of signal and noise. We compare the performance of the CNNs with other well-established machine learning classification methods based on decision trees, finding that the CNNs improve the classification when the measurement errors are provided as inputs. The predicted distribution of objects in miniJPAS is consistent with the putative luminosity functions of stars, quasars, and unresolved galaxies. Our results are a proof of concept for the idea that the J-PAS survey will be able to detect unprecedented numbers of quasars with high confidence.

79 ASTRONOMY AND ASTROPHYSICS↗

A measurement of the distance to the Galactic centre using the kinematics of bar stars

The distance to the Galactic centre R 0 is a fundamental parameter for understanding the Milky Way, because all observations of our Galaxy are made from our heliocentric reference point. The uncertainty in R 0 limits our knowledge of many aspects of the Milky Way, including its total mass and the relative mass of its major components, and any orbital parameters of stars employed in chemo-dynamical analyses. While measurements of R 0 have been improving over a century, measurements in the past few years from a variety of methods still find a wide range of R 0 being somewhere within 8.0 to $8.5\, \mathrm{kpc}$. The most precise measurements to date have to assume that Sgr A* is at rest at the Galactic centre, which may not be the case. In this paper, we use maps of the kinematics of stars in the Galactic bar derived from APOGEE DR17 and Gaia EDR3 data augmented with spectrophotometric distances from the astroNN neural-network method. These maps clearly display the minimum in the rotational velocity v T and the quadrupolar signature in radial velocity v R expected for stars orbiting in a bar. From the minimum in v T , we measure $R_0 = 8.23\pm 0.12\, \mathrm{kpc}$. We validate our measurement using realistic N-body simulations of the Milky Way. We further measure the pattern speed of the bar to be $\Omega _\mathrm{bar} = 40.08\pm 1.78\, \mathrm{km\, s}^{-1}\,\mathrm{kpc}^{-1}$. Because the bar forms out of the disc, its centre is manifestly the barycentre of the bar+disc system and our measurement is therefore one of the most robust and accurate measurements of R 0 to date.

79 ASTRONOMY AND ASTROPHYSICS↗

Photometric calibration in u -band using blue halo stars

ABSTRACT We develop a method to calibrate u-band photometry based on the observed colour of blue Galactic halo stars. The Galactic halo stars belong to an old stellar population of the Milky Way and have relatively low metallicity. The ‘blue tip’ of the halo population – the main sequence turn-off (MSTO) stars – is known to have a relatively uniform intrinsic edge u-g colour with only slow spatial variation. In SDSS data, the observed variation is correlated with Galactic Latitude, which we attribute to contamination by higher metallicity disc stars and fit with an empirical curve. This curve can then be used to calibrate u-band imaging if g-band imaging of matching depth is available. Our approach can be applied to single-field observations at |b| > 30°, and removes the need for standard star observations or overlap with calibrated u-band imaging. We include in our method the calibration of g-band data with ATLAS-Refcat2. We test our approach on stars in KiDS DR 4, ATLAS DR 4, and DECam imaging from the NOIRLab Source Catalog (NSC DR2), and compare our calibration with SDSS. For this process, we use synthetic magnitudes to derive the colour equations between these data sets, in order to improve zero-point accuracy. We find an improvement for all data sets, reaching a zero-point precision of 0.016 mag for KiDS (compared to the original 0.033 mag), 0.020 mag for ATLAS (originally 0.027 mag), and 0.016 mag for DECam (originally 0.041 mag). Thus, this method alone reaches the goal of 0.02 mag photometric precision in u-band for the Rubin Observatory’s Legacy Survey of Space and Time (LSST).

79 ASTRONOMY AND ASTROPHYSICS↗

First detection of the BAO signal from early DESI data

We present the first detection of the baryon acoustic oscillations (BAOs) signal obtained using unblinded data collected during the initial 2 months of operations of the Stage-IV ground-based Dark Energy Spectroscopic Instrument (DESI). From a selected sample of 261 291 luminous red galaxies spanning the redshift interval 0.4 < z < 1.1 and covering 1651 square degrees with a 57.9 per cent completeness level, we report a ∼5σ level BAO detection and the measurement of the BAO location at a precision of 1.7 per cent. Using a bright galaxy sample of 109 523 galaxies in the redshift range 0.1 < z < 0.5, over 3677 square degrees with a 50.0 per cent completeness, we also detect the BAO feature at ∼3σ significance with a 2.6 per cent precision. These first BAO measurements represent an important milestone, acting as a quality control on the optimal performance of the complex robotically actuated, fibre-fed DESI spectrograph, as well as an early validation of the DESI spectroscopic pipeline and data management system. Based on these first promising results, we forecast that DESI is on target to achieve a high-significance BAO detection at sub-per cent precision with the completed 5-yr survey data, meeting the top-level science requirements on BAO measurements. This exquisite level of precision will set new standards in cosmology and confirm DESI as the most competitive BAO experiment for the remainder of this decade.

79 ASTRONOMY AND ASTROPHYSICS↗

Curifactory: A research experiment manager

Curifactory is a command line tool and framework for organizing Python experiment code, configuration parameters, and results. It is an opinionated and lightweight approach to workflow management infrastructure and is primarily intended to support researchers conducting experiments on one machine. This software was developed to support the reproducibility of results for several data science projects in the Nuclear Nonproliferation Division at Oak Ridge National Laboratory. Curifactory is intended to be a general framework and is not specific to machine learning or data science. It can aid in any field in which experiments are primarily computation-based studies and can be implemented in Python (e.g., high-energy physics, astronomy, computational chemistry). Here, the design emphasizes the automated caching of intermediate data analysis artifacts to speed up development involving computationally intensive tasks. It also allows for data provenance and experiment reproduction. Individual experiment runs are tracked through logs and their output reports, and entire copies of a run with all cached data and metadata can be exported for others to run using Curifactory on another machine. Curifactory experiments can either be integrated into a project from the beginning or can be written on top of an existing codebase without needing significant modification. A few important views of the Curifactory library can be seen in Figure 1.

97 MATHEMATICS AND COMPUTING↗

Validation of standardized data formats and tools for ground-level particle-based gamma-ray observatories

Context. Ground-based γ-ray astronomy is still a rather young field of research, with strong historical connections to particle physics. This is why most observations are conducted by experiments with proprietary data and analysis software, as is usual in the particle physics field. However, in recent years, this paradigm has been slowly shifting toward the development and use of open-source data formats and tools, driven by upcoming observatories such as the Cherenkov Telescope Array (CTA). In this context, a community-driven, shared data format (the gamma-astro-data-format, or GADF) and analysis tools such as Gammapy and ctools have been developed. So far, these efforts have been led by the Imaging Atmospheric Cherenkov Telescope community, leaving out other types of ground-based γ-ray instruments. Aims. We aim to show that the data from ground particle arrays, such as the High-Altitude Water Cherenkov (HAWC) observatory, are also compatible with the GADF and can thus be fully analyzed using the related tools, in this case, Gammapy. Methods. We reproduced several published HAWC results using Gammapy and data products compliant with GADF standard. We also illustrate the capabilities of the shared format and tools by producing a joint fit of the Crab spectrum including data from six different γ-ray experiments. Results. We find excellent agreement with the reference results, a powerful confirmation of both the published results and the tools involved. Conclusions. The data from particle detector arrays such as the HAWC observatory can be adapted to the GADF and thus analyzed with Gammapy. A common data format and shared analysis tools allow multi-instrument joint analysis and effective data sharing. To emphasize this, a sample of Crab nebula event lists is made public with this paper. Because of the complementary nature of pointing and wide-field instruments, this synergy will be distinctly beneficial for the joint scientific exploitation of future observatories such as the Southern Wide-field Gamma-ray Observatory and CTA.

79 ASTRONOMY AND ASTROPHYSICS↗

Optimal Frequency-domain Analysis for Spacecraft Time Series: Introducing the Missing-data Multitaper Power Spectrum Estimator

While the Lomb–Scargle periodogram is foundational to astronomy, it has a significant shortcoming: the variance in the estimated power spectrum does not decrease as more data are acquired. Statisticians have a 60 yr history of developing variance-suppressing power spectrum estimators, but most are not used in astronomy because they are formulated for time series with uniform observing cadence and without seasonal or daily gaps. Here we demonstrate how to apply the missing-data multitaper power spectrum estimator to spacecraft data with uniform time intervals between observations but missing data during thruster fires or momentum dumps. The F-test for harmonic components may be applied to multitaper power spectrum estimates to identify statistically significant oscillations that would not rise above a white noise–based false alarm probability. Multitapering improves the dynamic range of the power spectrum estimate and suppresses spectral window artifacts. We show that the multitaper–F-test combination applied to Kepler observations of KIC 6102338 detects differential rotation without requiring iterative sinusoid fitting and subtraction. Significant signals reside at harmonics of both fundamental rotation frequencies and suggest an antisolar rotation profile. Next we use the missing-data multitaper power spectrum estimator to identify the oscillation modes responsible for the complex "scallop-shell" shape of the K2 light curve of EPIC 203354381. We argue that multitaper power spectrum estimators should be used for all time series with regular observing cadence.

79 ASTRONOMY AND ASTROPHYSICS↗

Impact of point spread function higher moments error on weak gravitational lensing

ABSTRACT Weak gravitational lensing is one of the most powerful tools for cosmology, while subject to challenges in quantifying subtle systematic biases. The point spread function (PSF) can cause biases in weak lensing shear inference when the PSF model does not match the true PSF that is convolved with the galaxy light profile. Although the effect of PSF size and shape errors – i.e. errors in second moments – is well studied, weak lensing systematics associated with errors in higher moments of the PSF model require further investigation. The goal of our study is to estimate their potential impact for LSST weak lensing analysis. We go beyond second moments of the PSF by using image simulations to relate multiplicative bias in shear to errors in the higher moments of the PSF model. We find that the current level of errors in higher moments of the PSF model in data from the Hyper Suprime-Cam survey can induce a ∼0.05 per cent shear bias, making this effect unimportant for ongoing surveys but relevant at the precision of upcoming surveys such as LSST.

79 ASTRONOMY AND ASTROPHYSICS↗

A general framework for removing point-spread function additive systematics in cosmological weak lensing analysis

ABSTRACT Cosmological weak lensing measurements rely on a precise measurement of the shear two-point correlation function (2PCF) along with a deep understanding of systematics that affect it. In this work, we demonstrate a general framework for detecting and modelling the impact of PSF systematics on the cosmic shear 2PCF and mitigating its impact on cosmological analysis. Our framework can detect PSF leakage and modelling error from all spin-2 quantities contributed by the PSF second and higher moments, rather than just the second moments, using the cross-correlations between galaxy shapes and PSF moments. We interpret null tests using the HSC Year 3 (Y3) catalogs with this formalism and find that leakage from the spin-2 combination of PSF fourth moments is the leading contributor to additive shear systematics, with total contamination that is an order-of-magnitude higher than that contributed by PSF second moments alone. We conducted a mock cosmic shear analysis for HSC Y3 and find that, if uncorrected, PSF systematics can bias the cosmological parameters Ωm and S8 by ∼0.3σ. The traditional second moment-based model can only correct for a 0.1σ bias, leaving the contamination largely uncorrected. We conclude it is necessary to model both PSF second and fourth moment contaminations for HSC Y3 cosmic shear analysis. We also reanalyse the HSC Y1 cosmic shear analysis with our updated systematics model and identify a 0.07σ bias on Ωm when using the more restricted second moment model from the original analysis. We demonstrate how to self-consistently use the method in both real space and Fourier space, assess shear systematics in tomographic bins, and test for PSF model overfitting.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Dark Energy Studies with LSST using the Photon Simulator (Final Report)

This grant funded the development and dissemination of the Photon Simulator (PhoSim) for the purpose of studying dark energy at high precision with the upcoming Large Synoptic Survey Telescope (LSST) astronomical survey. The work was in collaboration with the Dark Energy Science Collaboration (DESC). Several detailed physics improvements were made in the optics, atmosphere and sensor, a number of validation studies were performed, and a significant number of usability features were implemented. Future work in DESC will use PhoSim as the image simulation tool for data challenges used by the analysis groups.

79 ASTRONOMY AND ASTROPHYSICS↗

Photometric redshift-aided classification using ensemble learning

We present SHEEP, a new machine learning approach to the classic problem of astronomical source classification, which combines the outputs from the XGBoost, LightGBM, and CatBoost learning algorithms to create stronger classifiers. A novel step in our pipeline is that prior to performing the classification, SHEEP first estimates photometric redshifts, which are then placed into the data set as an additional feature for classification model training; this results in significant improvements in the subsequent classification performance. SHEEP contains two distinct classification methodologies: (i) Multi-class and (ii) one versus all with correction by a meta-learner. We demonstrate the performance of SHEEP for the classification of stars, galaxies, and quasars using a data set composed of SDSS and WISE photometry of 3.5 million astronomical sources. The resulting F1 -scores are as follows: 0.992 for galaxies; 0.967 for quasars; and 0.985 for stars. In terms of the F1-scores for the three classes, SHEEP is found to outperform a recent RandomForest-based classification approach using an essentially identical data set. Our methodology also facilitates model and data set explainability via feature importances; it also allows the selection of sources whose uncertain classifications may make them interesting sources for follow-up observations.

79 ASTRONOMY AND ASTROPHYSICS↗

Discovering strongly lensed quasar candidates with catalogue-based methods from DESI Legacy Surveys

The Hubble tension, revealed by a ~5σ discrepancy between measurements of the Hubble-Lemaitre constant among observations of the early and local Universe, is one of the most significant problems in modern cosmology. In order to better understand the origin of this mismatch, independent techniques to measure H 0 , such as strong lensing time delays, are required. Notably, the sample size of such systems is key to minimising the statistical uncertainties and cosmic variance, which can be improved by exploring the datasets of large-scale sky surveys such as Dark Energy Spectroscopic Instrument (DESI). We identify possible strong lensing time-delay systems within DESI by selecting candidate multiply imaged lensed quasars from a catalogue of 24 440 816 candidate QSOs contained in the ninth data release of the DESI Legacy Imaging Surveys (DESI-LS). Using a friend-of-friends-like algorithm on spatial co-ordinates, our method generates an initial list of compact quasar groups. This list is subsequently filtered using a measure of the similarity of colours among a group’s members and the likelihood that they are quasars. A visual inspection finally selects candidate strong lensing systems based on the spatial configuration of the group members. We identified 620 new candidate multiply imaged lensed quasars (101 grade-A, 214 grade-B, 305 grade-C). This number excludes 53 known spectroscopically confirmed systems and existing candidate systems identified in other similar catalogues. When available, these new candidates will be further checked by combining the spectroscopic and photometric data from DESI.

79 ASTRONOMY AND ASTROPHYSICS↗

DeepSZ: identification of Sunyaev–Zel’dovich galaxy clusters using deep learning

Galaxy clusters identified via the Sunyaev–Zel’dovich (SZ) effect are a key ingredient in multiwavelength cluster cosmology. In this paper, we present and compare three methods of cluster identification: the standard matched filter (MF) method in SZ cluster finding, a convolutional neural networks (CNN), and a ‘combined’ identifier. We apply the methods to simulated millimeter maps for several observing frequencies for a survey similar to SPT-3G, the third-generation camera for the South Pole Telescope. The MF requires image pre-processing to remove point sources and a model for the noise, while the CNN requires very little pre-processing of images. Additionally, the CNN requires tuning of hyperparameters in the model and takes cut-out images of the sky as input, identifying the cut-out as cluster-containing or not. We compare differences in purity and completeness. The MF signal-to-noise ratio depends on both mass and redshift. Our CNN, trained for a given mass threshold, captures a different set of clusters than the MF, some with signal-to-noise-ratio below the MF detection threshold. However, the CNN tends to mis-classify cut-out whose clusters are located near the edge of the cut-out, which can be mitigated with staggered cut-out. We leverage the complementarity of the two methods, combining the scores from each method for identification. The purity and completeness are both 0.61 for MF, and 0.59 and 0.61 for CNN. The combined method yields 0.60 and 0.77, a significant increase for completeness with a modest decrease in purity. We advocate for combined methods that increase the confidence of many low signal-to-noise clusters.

79 ASTRONOMY AND ASTROPHYSICS↗