Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Astronomy data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Scaling pair count to next galaxy surveys

ABSTRACT Counting pairs of galaxies or stars according to their distance is at the core of real-space correlation analyses performed in astrophysics and cosmology. Upcoming galaxy surveys (LSST, Euclid) will measure properties of billions of galaxies challenging our ability to perform such counting in a minute-scale time relevant for the usage of simulations. The problem is only limited by efficient access to the data, hence belongs to the big data category. We use the popular Apache Spark framework to address it and design an efficient high-throughput algorithm to deal with hundreds of millions to billions of input data. To optimize it, we revisit the question of non-hierarchical sphere pixelization based on cube symmetries and develop a new one dubbed the ‘Similar Radius Sphere Pixelization’ (SARSPix) with very close to square pixels. It provides the most adapted indexing over the sphere for all distance-related computations. Using LSST-like fast simulations, we compute autocorrelation functions on tomographic bins containing between a hundred million to one billion data points. In each case, we achieve the construction of a standard pair-distance histogram in about 2 min, using a simple algorithm that is shown to scale, over a moderate number of nodes (16–64). This illustrates the potential of this new techniques in the field of astronomy where data access is becoming the main bottleneck. They can be easily adapted to other use-cases as nearest-neighbours search, catalogue cross-match or cluster finding. The software is publicly available from https://github.com/astrolabsoftware/SparkCorr.

79 ASTRONOMY AND ASTROPHYSICS↗

Optimizing the shape of photometric redshift distributions with clustering cross-correlations

We present an optimization method for the assignment of photometric galaxies to a chosen set of redshift bins. This is achieved by combining simulated annealing, an optimization algorithm inspired by solid-state physics, with an unsupervised machine learning method, a self-organizing map (SOM) of the observed colours of galaxies. Starting with a sample of galaxies that is divided into redshift bins based on a photometric redshift point estimate, the simulated annealing algorithm repeatedly reassigns SOM-selected subsamples of galaxies, which are close in colour, to alternative redshift bins. We optimize the clustering cross-correlation signal between photometric galaxies and a reference sample of galaxies with well-calibrated redshifts. Depending on the effect on the clustering signal, the reassignment is either accepted or rejected. By dynamically increasing the resolution of the SOM, the algorithm eventually converges to a solution that minimizes the number of mismatched galaxies in each tomographic redshift bin and thus improves the compactness of their corresponding redshift distribution. This method is demonstrated on the synthetic Legacy Survey of Space and Time cosmoDC2 catalogue. We find a significant decrease in the fraction of catastrophic outliers in the redshift distribution in all tomographic bins, most notably in the highest redshift bin with a decrease in the outlier fraction from 57 percent to 16 percent.

79 ASTRONOMY AND ASTROPHYSICS↗

Using action space clustering to constrain the recent accretion history of Milky Way-like galaxies

ABSTRACT In the currently favoured cosmological paradigm galaxies form hierarchically through the accretion of satellites. Since a satellite is less massive than the host, its stars occupy a smaller volume in action space. Actions are conserved when the potential of the host halo changes adiabatically, so stars from an accreted satellite would remain clustered in action space as the host evolves. In this paper, we identify recently disrupted accreted satellites in three Milky Way-like disc galaxies from the cosmological baryonic FIRE-2 simulations by tracking satellites through simulation snapshots. We try to recover these satellites by applying the cluster analysis algorithm Enlink to the orbital actions of accreted star particles in the z = 0 snapshot. Even with completely error-free mock data we find that only 35 per cent (14/39) satellites are well recovered while the rest (25/39) are poorly recovered (i.e. either contaminated or split up). Most (10/14 ∼70 per cent) of the well-recovered satellites have infall times <7.1 Gyr ago and total mass >4 × 108M⊙ (stellar mass more than 1.2 × 106 M⊙, although our upper mass limit is likely to be resolution dependent). Since cosmological simulations predict that stellar haloes include a population of in situ stars, we test our ability to recover satellites when the data include 10–50 per cent in situ contamination. We find that most previously well-recovered satellites stay well recovered even with 50 per cent contamination. With the wealth of 6D phase space data becoming available we expect that cluster analysis in action space will be useful in identifying the majority of recently accreted and moderately massive satellites in the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS↗

Map-based cosmology inference with lognormal cosmic shear maps

ABSTRACT Most cosmic shear analyses to date have relied on summary statistics (e.g. ξ+ and ξ−). These types of analyses are necessarily suboptimal, as the use of summary statistics is lossy. In this paper, we forward-model the convergence field of the Universe as a lognormal random field conditioned on the observed shear data. This new map-based inference framework enables us to recover the joint posterior of the cosmological parameters and the convergence field of the Universe. Our analysis properly accounts for the covariance in the mass maps across tomographic bins, which significantly improves the fidelity of the maps relative to single-bin reconstructions. We verify that applying our inference pipeline to Gaussian random fields recovers posteriors that are in excellent agreement with their analytical counterparts. At the resolution of our maps – and to the extent that the convergence field can be described by the lognormal model – our map posteriors allow us to reconstruct all summary statistics (including non-Gaussian statistics). We forecast that a map-based inference analysis of LSST-Y10 data can improve cosmological constraints in the σ8–Ωm plane by $\approx\!{30}{{\ \rm per\ cent}}$ relative to the currently standard cosmic shear analysis. This improvement happens almost entirely along the $S_8=\sigma _8\Omega _{\rm m}^{1/2}$ directions, meaning map-based inference fails to significantly improve constraints on S8.

79 ASTRONOMY AND ASTROPHYSICS↗

Characterizing the Sample Selection for Supernova Cosmology

Type Ia supernovae (SNe Ia) are used as distance indicators to infer the cosmological parameters that specify the expansion history of the universe. Parameter inference depends on the criteria by which the analysis SN sample is selected. Only for the simplest selection criteria and population models can the likelihood be calculated analytically, otherwise it needs to be determined numerically, a process that inherently has error. Numerical errors in the likelihood lead to errors in parameter inference. This article presents toy examples where the distance modulus is inferred given a set of SNe at a single redshift. Parameter estimators and their uncertainties are calculated using Monte Carlo techniques. The relationship between the number of Monte Carlo realizations and numerical errors is presented. The procedure can be applied to more realistic models and used to determine the computational and data management requirements of the transient analysis pipeline.

79 ASTRONOMY AND ASTROPHYSICS↗

Constraining low redshift [C II ] emission by cross-correlating FIRAS and BOSS data

ABSTRACT We perform a tomographic cross-correlation analysis of archival FIRAS data and the BOSS galaxy redshift survey to constrain the amplitude of [C II] 2P3/2 → 2P1/2 fine structure emission. Our analysis employs spherical harmonic tomography (SHT), which is based on the angular cross-power spectrum between FIRAS maps and BOSS galaxy over-densities at each pair of redshift bins, over a redshift range of 0.24 < z < 0.69. We develop the SHT approach for intensity mapping, where it has several advantages over existing power spectral estimators. Our analysis constrains the product of the [C II] bias and [C II] specific intensity, $b_{\rm [C \small{\rm II}]}I_{\rm [C \small{\rm II}]}$, to be <0.31 MJy/sr at z ≈ 0.35 and <0.28 MJy/sr at z ≈ 0.57 at $95{{\ \rm per\ cent}}$ confidence. These limits are consistent with most current models of the [C II] signal, as well as with higher-redshift [C II] cross-power spectrum measurements from the Planck satellite and BOSS quasars. We also show that our analysis, if applied to data from a more sensitive instrument such as the proposed PIXIE satellite, can detect pessimistic [C II] models at high significance.

79 ASTRONOMY AND ASTROPHYSICS↗

TDCOSMO. X. Automated modeling of nine strongly lensed quasars and comparison between lens-modeling software

When strong gravitational lenses are to be used as an astrophysical or cosmological probe, models of their mass distributions are often needed. We present a new, time-efficient automation code for the uniform modeling of strongly lensed quasars with GLEE, a lens-modeling software for multiband data. By using the observed positions of the lensed quasars and the spatially extended surface brightness distribution of the host galaxy of the lensed quasar, we obtain a model of the mass distribution of the lens galaxy. We applied this uniform modeling pipeline to a sample of nine strongly lensed quasars for which images were obtained with the Wide Field Camera 3 of the Hubble Space Telescope. The models show well-reconstructed light components and a good alignment between mass and light centroids in most cases. We find that the automated modeling code significantly reduces the input time during the modeling process for the user. The time for preparing the required input files is reduced by a factor of 3 from ~3 h to about one hour. The active input time during the modeling process for the user is reduced by a factor of 10 from ~ 10 h to about one hour per lens system. This automated uniform modeling pipeline can efficiently produce uniform models of extensive lens-system samples that can be used for further cosmological analysis. A blind test that compared our results with those of an independent automated modeling pipeline based on the modeling software Lenstronomy revealed important lessons. Quantities such as Einstein radius, astrometry, mass flattening, and position angle are generally robustly determined. Other quantities, such as the radial slope of the mass density profile and predicted time delays, depend crucially on the quality of the data and on the accuracy with which the point spread function is reconstructed. Better data and/or a more detailed analysis are necessary to elevate our automated models to cosmography grade. Nevertheless, our pipeline enables the quick selection of lenses for follow-up and further modeling, which significantly speeds up the construction of cosmography-grade models. This important step forward will help us to take advantage of the increase in the number of lenses that is expected in the coming decade, which is an increase of several orders of magnitude.

79 ASTRONOMY AND ASTROPHYSICS↗

A Bayesian approach to strong lens finding in the era of wide-area surveys

ABSTRACT The arrival of the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST), Euclid-Wide and Roman wide-area sensitive surveys will herald a new era in strong lens science in which the number of strong lenses known is expected to rise from $\mathcal {O}(10^3)$ to $\mathcal {O}(10^5)$. However, current lens-finding methods still require time-consuming follow-up visual inspection by strong lens experts to remove false positives which is only set to increase with these surveys. In this work, we demonstrate a range of methods to produce calibrated probabilities to help determine the veracity of any given lens candidate. To do this we use the classifications from citizen science and multiple neural networks for galaxies selected from the Hyper Suprime-Cam survey. Our methodology is not restricted to particular classifier types and could be applied to any strong lens classifier which produces quantitative scores. Using these calibrated probabilities, we generate an ensemble classifier, combining citizen science, and neural network lens finders. We find such an ensemble can provide improved classification over the individual classifiers. We find a false-positive rate of 10−3 can be achieved with a completeness of 46 per cent, compared to 34 per cent for the best individual classifier. Given the large number of galaxy–galaxy strong lenses anticipated in LSST, such improvement would still produce significant numbers of false positives, in which case using calibrated probabilities will be essential for population analysis of large populations of lenses and to help prioritize candidates for follow-up.

79 ASTRONOMY AND ASTROPHYSICS↗

The Planetary Ephemeris Program: Capability, Comparison, and Open Source Availability

We describe for the first time in scientific literature the Planetary Ephemeris Program (PEP), an open-source general-purpose astrometric data-analysis program. We discuss, in particular, the implementation of pulsar timing analysis, which was recently upgraded in PEP to handle more options. This implementation was done independently of other pulsar programs, with minor exceptions that we discuss. We illustrate the implementation of this capability by comparing the postfit residuals from the analyses of time-of-arrival observations by both PEP and Tempo2. The comparison shows substantial agreement: 22 ns rms differences for 1065 pulse time-of-arrival measurements for the millisecond pulsar in a binary system, PSR J1909-3744 (pulse period 2.947108 ms; full width half maximum of pulse 43 μs), for epochs in the interval from 2002 December to 2011 February.

42 ENGINEERING↗

Active anomaly detection for time-domain discoveries

Aims. We present the first evidence that adaptive learning techniques can boost the discovery of unusual objects within astronomical light curve data sets. Methods. Our method follows an active learning strategy where the learning algorithm chooses objects which can potentially improve the learner if additional information about them is provided. This new information is subsequently used to update the machine learning model, allowing its accuracy to evolve with each new information. For the case of anomaly detection, the algorithm aims to maximize the number of scientifically interesting anomalies presented to the expert by slightly modifying the weights of a traditional Isolation Forest (IF) at each iteration. In order to demonstrate the potential of such techniques, we apply the Active Anomaly Discovery (AAD) algorithm to 2 data sets: simulated light curves from the Photometric LSST Astronomical Time-Series Classification Challenge (PLAsTiCC) and real light curves from the Open Supernova Catalog. We compare the AAD results to those of a static IF. For both methods, we performed a detailed analysis for all objects with the ~2% highest anomaly scores. Results. We show that, in the real data scenario , AAD was able to identify ~80% more true anomalies than the IF. This result is the first evidence that AAD algorithms can play a central role in the search for new physics in the era of large scale sky surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

TDCOSMO - XVII. New time delays in 22 lensed quasars from optical monitoring with the ESO-VST 2.6m and MPG 2.2m telescopes

We present new time delays, the main ingredient of time delay cosmography, for 22 lensed quasars resulting from high-cadence r-band monitoring on the 2.6 m ESO VLT Survey Telescope and Max-Planck-Gesellschaft 2.2 m telescope. Each lensed quasar was typically monitored for one to four seasons, often shared between the two telescopes to mitigate the interruptions forced by the COVID-19 pandemic. The sample of targets consists of 19 quadruply and 3 doubly imaged quasars, which received a total of 1918 hours of on-sky time split into 21 581 wide-field frames, each 320 seconds long. In a given field, the 5-σ depth of the combined exposures typically reaches the 27th magnitude, while that of single visits is 24.5 mag – similar to the expected depth of the upcoming Vera-Rubin LSST. The fluxes of the different lensed images of the targets were reliably de-blended, providing not only light curves with photometric precision down to the photon noise limit, but also high-resolution models of the targets whose features and astrometry were systematically confirmed in Hubble Space Telescope imaging. This was made possible thanks to a new photometric pipeline, lightcurver, and the forward modelling method STARRED. Finally, the time delays between pairs of curves and their uncertainties were estimated, taking into account the degeneracy due to microlensing, and for the first time the full covariance matrices of the delay pairs are provided. Of note, this survey, with 13 square degrees, has applications beyond that of time delays, such as the study of the structure function of the multiple high-redshift quasars present in the footprint at a new high in terms of both depth and frequency. The reduced images will be available through the European Southern Observatory Science Portal.Key words: methods: data analysis / surveys / distance scale

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A machine learning approach to galaxy properties: joint redshift–stellar mass probability distributions with Random Forest

We demonstrate that highly accurate joint redshift–stellar mass probability distribution functions (PDFs) can be obtained using the Random Forest (RF) machine learning (ML) algorithm, even with few photometric bands available. As an example, we use the Dark Energy Survey (DES), combined with the COSMOS2015 catalogue for redshifts and stellar masses. We build two ML models: one containing deep photometry in the griz bands, and the second reflecting the photometric scatter present in the main DES survey, with carefully constructed representative training data in each case. We validate our joint PDFs for 10 699 test galaxies by utilizing the copula probability integral transform and the Kendall distribution function, and their univariate counterparts to validate the marginals. Benchmarked against a basic set-up of the template-fitting code bagpipes, our ML-based method outperforms template fitting on all of our predefined performance metrics. In addition to accuracy, the RF is extremely fast, able to compute joint PDFs for a million galaxies in just under 6 min with consumer computer hardware. Such speed enables PDFs to be derived in real time within analysis codes, solving potential storage issues. As part of this work we have developed galpro 1, a highly intuitive and efficient python package to rapidly generate multivariate PDFs on-the-fly. galpro is documented and available for researchers to use in their cosmology and galaxy evolution studies.

79 ASTRONOMY AND ASTROPHYSICS↗

A variational encoder–decoder approach to precise spectroscopic age estimation for large Galactic surveys

Constraints on the formation and evolution of the Milky Way Galaxy require multidimensional measurements of kinematics, abundances, and ages for a large population of stars. Ages for luminous giants, which can be seen to large distances, are an essential component of studies of the Milky Way, but they are traditionally very difficult to estimate precisely for a large data set and often require careful analysis on a star-by-star basis in asteroseismology. Because spectra are easier to obtain for large samples, being able to determine precise ages from spectra allows for large age samples to be constructed, but spectroscopic ages are often imprecise and contaminated by abundance correlations. Here we present an application of a variational encoder–decoder on cross-domain astronomical data to solve these issues. The model is trained on pairs of observations from APOGEE and Kepler of the same star in order to reduce the dimensionality of the APOGEE spectra in a latent space while removing abundance information. The low dimensional latent representation of these spectra can then be trained to predict age with just ∼1000 precise seismic ages. We demonstrate that this model produces more precise spectroscopic ages (∼ 22 per cent overall, ∼ 11 per cent for red-clump stars) than previous data-driven spectroscopic ages while being less contaminated by abundance information (in particular, our ages do not depend on [α/M]). We create a public age catalogue for the APOGEE DR17 data set and use it to map the age distribution and the age-[Fe/H]-[α/M] distribution across the radial range of the Galactic disc.

79 ASTRONOMY AND ASTROPHYSICS↗

The GAPS programme at TNG

Neptunes represent one of the main types of exoplanets and have chemical-physical characteristics halfway between rocky and gas giant planets. Therefore, their characterization is important for understanding and constraining both the formation mechanisms and the evolution patterns of planets. We investigate the exoplanet candidate TOI-1422 b, which was discovered by the TESS space telescope around the high proper-motion G2 V star TOI-1422 (V = 10.6 mag), 155 pc away, with the primary goal of confirming its planetary nature and characterising its properties. We monitored TOI-1422 with the HARPS-N spectrograph for 1.5 yr to precisely quantify its radial velocity (RV) variation. We analyse these RV measurements jointly with TESS photometry and check for blended companions through high-spatial resolution images using the AstraLux instrument. We estimate that the parent star has a radius of R$_\star$ = 1.019 -0.013 +0.014 R ⊙ , and a mass of M$_\star$ = 1.019 -0.013 +0.014 M ⊙ . Our analysis confirms the planetary nature of TOI-1422 b and also suggests the presence of a Neptune-mass planet on a more distant orbit, the candidate TOI-1422 c, which is not detected in TESS light curves. The inner planet, TOI-1422 b, orbits on a period of P b = 12.9972 ± 0.0006 days and has an equilibrium temperature of T eq,b = 867 ± 17 K. With a radius of R b = 3.96 -0.11 +0.13 R ⊕ , a mass of M b = 9.0 -2.0 +2.3 M ⊕ and, consequently, a density of ρ b = 0.795 -0.235 +0.290 g cm -3 , it can be considered a warm Neptune-sized planet. Compared to other exoplanets of a similar mass range, TOI-1422 b is among the most inflated, and we expect this planet to have an extensive gaseous envelope that surrounds a core with a mass fraction around 10% – 25% of the total mass of the planet. The outer non-transiting planet candidate, TOI-1422 c, has an orbital period of P c = 29.29 -0.20 +0.21 days, a minimum mass, M c sin i, of 11.1 -2.3 +2.6 M ⊕ , an equilibrium temperature of T eq,c = 661 ± 13 K and, therefore, if confirmed, could be considered as another warm Neptune.

79 ASTRONOMY AND ASTROPHYSICS↗

Photometric redshift estimation with convolutional neural networks and galaxy images: Case study of resolving biases in data-driven methods

Deep-learning models have been increasingly exploited in astrophysical studies, but these data-driven algorithms are prone to producing biased outputs that are detrimental for subsequent analyses. In this work, we investigate two main forms of biases: class-dependent residuals, and mode collapse. We do this in a case study, in which we estimate photometric redshift as a classification problem using convolutional neural networks (CNNs) trained with galaxy images and associated spectroscopic redshifts. We focus on point estimates and propose a set of consecutive steps for resolving the two biases based on CNN models, involving representation learning with multichannel outputs, balancing the training data, and leveraging soft labels. The residuals can be viewed as a function of spectroscopic redshift or photometric redshift, and the biases with respect to these two definitions are incompatible and should be treated individually. We suggest that a prerequisite for resolving biases in photometric space is resolving biases in spectroscopic space. Experiments show that our methods can better control biases than benchmark methods, and they are robust in various implementing and training conditions with high-quality data. Our methods hold promises for future cosmological surveys that require a good constraint of biases, and they may be applied to regression problems and other studies that make use of data-driven models. Nonetheless, the bias-variance tradeoff and the requirement of sufficient statistics suggest that we need better methods and optimized data usage strategies.

79 ASTRONOMY AND ASTROPHYSICS↗

VarIabiLity seLection of AstrophysIcal sources iN PTF (VILLAIN): I. Structure function fits to 71 million objects

Light-curve variability is well-suited to characterizing objects in surveys with high cadence and a long baseline. This is especially relevant in view of the large datasets to be produced by the Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST). We aim to determine variability parameters for objects in the Palomar Transient Factory (PTF) and explore differences between quasars (QSOs), stars, and galaxies. We relate variability and colour information in preparation for future surveys. We fit joint likelihoods to structure functions (SFs) of 71 million PTF light curves with a Markov chain Monte Carlo method. For each object, we assume a power-law SF and extract two parameters: the amplitude on timescales of one year, A, and a power-law index, γ. With these parameters and colours in the optical (Pan-STARRS1) and mid-infrared (WISE), we identify regions of parameter space dominated by different types of spectroscopically confirmed objects from SDSS. Candidate QSOs, stars, and galaxies are selected to show their parameter distributions. QSOs show high-amplitude variations in the R band, and the highest γ values. Galaxies have a broader range of amplitudes and their variability shows relatively little dependency on timescale. With variability and colours, we achieve a photometric selection purity of 99.3% for QSOs. Even though hard cuts in monochromatic variability alone are not as effective as seven-band magnitude cuts, variability is useful in characterizing object subclasses. Through variability, we also find QSOs that were erroneously classified as stars in the SDSS. We discuss perspectives and computational solutions in view of the upcoming LSST.

79 ASTRONOMY AND ASTROPHYSICS↗

Stellar loci IV. red giant stars

In the fourth paper of this series, we present the metallicity-dependent Sloan Digital Sky Survey (SDSS) stellar color loci of red giant stars, using a spectroscopic sample of red giants in the SDSS Stripe 82 region. The stars span a range of 0.55 – 1.2 mag in color g – i, –0.3 – –2.5 in metallicity [Fe/H], and have values of surface gravity log g smaller than 3.5 dex. As in the case of main-sequence (MS) stars, the intrinsic widths of loci of red giants are also found to be quite narrow, a few mmag at maximum. There are however systematic differences between the metallicity-dependent stellar loci of red giants and MS stars. The colors of red giants are less sensitive to metallicity than those of MS stars. With good photometry, photometric metallicities of red giants can be reliably determined by fitting the u – g, g – r, r – i, and i – z colors simultaneously to an accuracy of 0.2 – 0.25 dex, comparable to the precision achievable with low-resolution spectroscopy for a signal-to-noise ratio of 10. By comparing fitting results to the stellar loci of red giants and MS stars, we propose a new technique to discriminate between red giants and MS stars based on the SDSS photometry. The technique achieves completeness of ~70 per cent and efficiency of ~80 per cent in selecting metal-poor red giant stars of [Fe/H] ≤ –1.2. It thus provides an important tool to probe the structure and assemblage history of the Galactic halo using red giant stars.

79 ASTRONOMY AND ASTROPHYSICS↗