Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data analysis methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Deep learning methods for obtaining photometric redshift estimations from images

ABSTRACT Knowing the redshift of galaxies is one of the first requirements of many cosmological experiments, and as it is impossible to perform spectroscopy for every galaxy being observed, photometric redshift (photo-z) estimations are still of particular interest. Here, we investigate different deep learning methods for obtaining photo-z estimates directly from images, comparing these with ‘traditional’ machine learning algorithms which make use of magnitudes retrieved through photometry. As well as testing a convolutional neural network (CNN) and inception-module CNN, we introduce a novel mixed-input model that allows for both images and magnitude data to be used in the same model as a way of further improving the estimated redshifts. We also perform benchmarking as a way of demonstrating the performance and scalability of the different algorithms. The data used in the study comes entirely from the Sloan Digital Sky Survey (SDSS) from which 1 million galaxies were used, each having 5-filtre (ugriz) images with complete photometry and a spectroscopic redshift which was taken as the ground truth. The mixed-input inception CNN achieved a mean squared error (MSE) =0.009, which was a significant improvement ($30{{\ \rm per\ cent}}$) over the traditional random forest (RF), and the model performed even better at lower redshifts achieving a MSE = 0.0007 (a $50{{\ \rm per\ cent}}$ improvement over the RF) in the range of z < 0.3. This method could be hugely beneficial to upcoming surveys, such as Euclid and the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST), which will require vast numbers of photo-z estimates produced as quickly and accurately as possible.

79 ASTRONOMY AND ASTROPHYSICS↗

Empirically Driven multiwavelength K-c orrections at low redshift

K -corrections – a necessary ingredient for converting between flux in observed bands to flux in rest-frame bands – are critical for comparing galaxies at differing redshifts. These corrections often rely on fits to empirical or theoretical spectral energy distribution (SED) templates of galaxies. However, templates can only produce reliable K-corrections in regimes where SED models are robust. For instance, the templates utilized in some popular software packages are not well-constrained in some bands (e.g. WISE W4 in KCORRECT ), which results in ill-behaved K -corrections. We address this shortcoming by developing an empirically driven approach to K -corrections that limits the dependence on SED templates. We perform a polynomial fit for the K -correction as a function of a galaxy’s rest-frame colour determined in a pair of well-constrained bands (e.g. 0 (g − r)) and redshift, exploiting the fact that galaxy SEDs can be approximated as a one-parameter family at low redshift. For bands well-constrained by SED templates, our empirically driven K -corrections yield results comparable to the SED fitting methods used by KCORRECT and the GSWLC-M2 catalogue (the updated medium-deep GALEX–SDSS–WISE Legacy Catalogue). However, our method dramatically outperforms Kcorrect derived K -corrections for WISE W4 . Our method is also robust to incorrect template assumptions outside of the optical bands and enforces that the K -correction must be zero at z = 0. Our K -corrected photometry and code are publicly available.

79 ASTRONOMY AND ASTROPHYSICS↗

What drives the variance of galaxy spectra?

We present a study aimed at understanding the physical phenomena underlying the formation and evolution of galaxies following a data-driven analysis of spectroscopic data based on the variance in a carefully selected sample. We apply principal component analysis (PCA) independently to three subsets of continuum-subtracted optical spectra, segregated into their nebular emission activity as quiescent, star-forming, and active galactic nuclei (AGNs). We emphasize that the variance of the input data in this work only relates to the absorption lines in the photospheres of the stellar populations. The sample is taken from the Sloan Digital Sky Survey (SDSS) in the stellar velocity dispersion range 100–150 km s −1 , to minimize the ‘blurring’ effect of the stellar motion. We restrict the analysis to the first three principal components (PCs) and find that PCA segregates the three types with the highest variance mapping SSP-equivalent age, along with an inextricable degeneracy with metallicity, even when all three PCs are included. Spectral fitting shows that stellar age dominates PC1, whereas PC2 and PC3 have a mixed dependence of age and metallicity. The trends support – independently of any model fitting – the hypothesis of an evolutionary sequence from star formation to AGN to quiescence. As a further test of the consistency of the analysis, we apply the same methodology in different spectral windows, finding similar trends, but the variance is maximal in the blue wavelength range, roughly around the 4000 Å break.

79 ASTRONOMY AND ASTROPHYSICS↗

Data-driven selection and spectral classification of white dwarf stars

The next generation of spectroscopic surveys is expected to provide spectra for hundreds of thousands of white dwarf (WD) candidates in the upcoming years. Currently, spectroscopic classification of white dwarfs is mostly done by visual inspection, requiring substantial amounts of expert attention. We propose a data-driven pipeline for fast, automatic selection, and spectroscopic classification of WD candidates, trained using spectroscopically confirmed objects with available Gaia astrometry, photometry, and Sloan Digital Sky Survey (SDSS) spectra with signal-to-noise ratios ≥9. The pipeline selects WD candidates with improved accuracy and completeness over existing algorithms, classifies their primary spectroscopic type with ≥ 90 % accuracy, and spectroscopically detects main sequence companions with similar performance. We apply our pipeline to the Gaia Data Release 3 cross-matched with the SDSS Data Release 17 (DR17), identifying 424 096 high-confidence WD candidates and providing the first catalogue of automated and quantifiable classification for 36 523 WD spectra. Both the catalogue and pipeline are made available online. Such a tool will prove particularly useful for the undergoing SDSS-V survey, allowing for rapid classification of thousands of spectra at every data release.

79 ASTRONOMY AND ASTROPHYSICS↗

Robust clustering of the local Milky Way stellar kinematic substructures with Gaia eDR3

Understanding local stellar kinematic substructures in the solar neighbourhood helps build a complete picture of the formation of the Milky Way, as well as an empirical phase space distribution of dark matter that would inform detection experiments. We apply the clustering algorithm HDBSCAN on the Gaia early third data release to identify a list of stable clusters in velocity space and action-angle space by taking into account the measurement uncertainties and studying the stability of the clustering results. We find 1405 (497) stars in 23 (6) robust clusters in velocity space (action-angle space) that are consistently not associated with noise. We discuss the kinematic properties of these structures and study whether many of the small clusters belong to a similar larger cluster based on their chemical abundances. They are attributed to the known structures: the Gaia Sausage-Enceladus, the Helmi Stream, and globular cluster NGC 3201 are found in both spaces, while NGC 104 and the thick disc (Sequoia) are identified in velocity space (action-angle space). Although we do not identify any new structures, we find that the HDBSCAN member selection of already known structures is unstable to input kinematics of the stars when resampled within their uncertainties. We therefore present the stable subset of local kinematic structures, which are consistently identified by the clustering algorithm, and emphasize the need to take into account error propagation during both the manual and automated identification of stellar structures, both for existing ones as well as future discoveries.

79 ASTRONOMY AND ASTROPHYSICS↗

Measurements of the z > 5 Lyman-α forest flux autocorrelation functions from the extended XQR-30 data set

We present the first observational measurements of the Lyman-α (Ly α) forest flux autocorrelation functions in ten redshift bins from 5.1 ≤ z ≤ 6.0. We use a sample of 35 quasar sightlines at z > 5.7 from the extended XQR-30 data set; these data have signal-to-noise ratios of >20 per spectral pixel. We carefully account for systematic errors in continuum reconstruction, instrumentation, and contamination by damped Ly α systems. With these measurements, we introduce software tools to generate autocorrelation function measurements from any simulation. Our measurements of the smallest bin of the autocorrelation function increase with redshift when normalizing by the mean flux, $\langle{F}\rangle$. This increase may come from decreasing $\langle{F}\rangle$ or increasing mean free path of hydrogen-ionizing photons, λmfp. Recent work has shown that the autocorrelation function from simulations at z > 5 is sensitive to λmfp, a quantity that contains vital information on the ending of reionization. For an initial comparison, we show our autocorrelation measurements with simulation models for recently measured λmfp values and find good agreements. Further work in modelling and understanding the covariance matrices of the data is necessary to get robust measurements of λmfp from this data.

79 ASTRONOMY AND ASTROPHYSICS↗

Detection of the large-scale tidal field with galaxy multiplet alignment in the DESI Y1 spectroscopic survey

We explore correlations between the orientations of small galaxy groups, or ‘multiplets’, and the large-scale gravitational tidal field. Using data from the Dark Energy Spectroscopic Instrument (DESI) Y1 survey, we detect the intrinsic alignment (IA) of multiplets to the galaxy-traced matter field out to separations of $100\,h^{-1}$ Mpc. Unlike traditional IA measurements of individual galaxies, this estimator is not limited by imaging of galaxy shapes and allows for direct IA detection beyond redshift $z=1$. Multiplet alignment is a form of higher order clustering, for which the scale-dependence traces the underlying tidal field and amplitude is a result of small-scale ($\lt 1h^{-1}$ Mpc) dynamics. Within samples of bright galaxies, luminous red galaxies (LRG) and emission-line galaxies, we find similar scale-dependence regardless of intrinsic luminosity or colour. This is promising for measuring tidal alignment in galaxy samples that typically display no IA. DESI’s LRG mock galaxy catalogues created from the A BACUS S UMMIT N -body simulations produce a similar alignment signal, though with a 33 per cent lower amplitude at all scales. An analytic model using a non-linear power spectrum (NLA) only matches the signal down to 20 $h^{-1}$ Mpc. Our detection demonstrates that galaxy clustering in the non-linear regime of structure formation preserves an interpretable memory of the large-scale tidal field. Multiplet alignment complements traditional two-point measurements by retaining directional information imprinted by tidal forces, and contains additional line-of-sight information compared to weak lensing. This is a more effective estimator than the alignment of individual galaxies in dense, blue, or faint galaxy samples.

79 ASTRONOMY AND ASTROPHYSICS↗

Redshift-dependent RSD bias from intrinsic alignment with DESI Year 1 spectra

ABSTRACT We estimate the redshift-dependent, anisotropic clustering signal in the Dark Energy Spectroscopic Instrument (DESI) Year 1 Survey created by tidal alignments of Luminous Red Galaxies (LRGs) and a selection-induced galaxy orientation bias. To this end, we measured the correlation between LRG shapes and the tidal field with DESI’s Year 1 redshifts, as traced by LRGs and Emission-Line Galaxies. We also estimate the galaxy orientation bias of LRGs caused by DESI’s aperture-based selection, and find it to increase by a factor of seven between redshifts 0.4−1.1 due to redder, fainter galaxies falling closer to DESI’s imaging selection cuts. These effects combine to dampen measurements of the quadrupole of the correlation function (ξ2) caused by structure growth on scales of 10–80 h−1 Mpc by about 0.15 per cent for low redshifts (0.4 < z < 0.6) and 0.8 per cent for high (0.8 < z < 1.1), a significant fraction of DESI’s error budget. We provide estimates of the ξ2 signal created by intrinsic alignments that can be used to correct this effect, which is necessary to meet DESI’s forecasted precision on measuring the growth rate of structure. While imaging quality varies across DESI’s footprint, we find no significant difference in this effect between imaging regions in the Legacy Imaging Survey.

79 ASTRONOMY AND ASTROPHYSICS↗

Tracing back the birth environments of Type Ia supernova progenitor stars: a pilot study based on 44 early-type host galaxies

The environmental dependence of Type Ia supernova (SN Ia) luminosities is well established, and efforts are being made to find its origin. Previous studies typically use the currently observed status of the host galaxy. However, given the delay time between the birth of the progenitor star and the SN Ia explosion, the currently observed status may differ from the birth environment of the SN Ia progenitor star. In this paper, employing the chemical evolution and accurately determined stellar population properties of 44 early-type host galaxies, we, for the first time, estimate the SN Ia progenitor star birth environment, specifically [Fe/H] Birth and [α/Fe] Birth . We show that [α/Fe] Birth has a $30.4^{\text{+10.6}}_{-10.1}{{\ \rm per\ cent}}$ wider range than the currently observed [α/Fe] Current , while the range of [Fe/H] Birth is not statistically different ($17.9^{\text{+26.0}}_{-27.1}{{\ \rm per\ cent}}$) to that of [Fe/H] Current . The birth and current environments of [Fe/H] and [α/Fe] are sampled from different populations (p-values of the Kolmogorov–Smirnov test <0.01). We find that light-curve fit parameters are insensitive to [Fe/H] Birth (<0.9σ for the non-zero slope), while a linear trend is observed with Hubble residuals (HRs) at the 2.4σ significance level. With [α/Fe] Birth , no linear trends (<1.1σ) are observed. Interestingly, we find that [α/Fe] Birth clearly splits the SN Ia sample into two groups: SN Ia exploded in [α/Fe] Birth -rich or [α/Fe ]Birth -poor environments. SNe Ia exploded in different [α/Fe] Birth groups have different weighted-means of light-curve shape parameters: 0.81 ± 0.33 (2.5σ). They are thought to be drawn from different populations (p-value = 0.01). Regarding SN Ia colour and HRs, there is no difference (<1.0σ) in the weighted-means and distribution (p-value > 0.27) of each [α/Fe] Birth group.

79 ASTRONOMY AND ASTROPHYSICS↗

Galaxy-multiplet clustering from DESI DR2

We present an efficient estimator for higher-order galaxy clustering using small groups of nearby galaxies, or multiplets. Using the Luminous Red Galaxy (LRG) sample from the Dark Energy Spectroscopic Instrument (DESI) Data Release 2, we identify galaxy multiplets as discrete objects and measure their cross-correlations with the general galaxy field. Our results show that the multiplets exhibit stronger clustering bias as they trace more massive dark matter halos than individual galaxies. When comparing the observed clustering statistics with the mock catalogs generated from the N-body simulation AbacusSummit, we find that the mocks underpredict multiplet clustering despite reproducing the galaxy two-point auto-correlation reasonably well. This discrepancy indicates that the standard Halo Occupation Distribution (HOD) model is insufficient to describe the properties of galaxy multiplets, revealing the greater constraining power of this higher-order statistic on galaxy-halo connection and the possibility that multiplets are specific to additional assembly bias. We demonstrate that incorporating secondary biases into the HOD model improves agreement with the observed multiplet statistics, specifically by allowing galaxies to preferentially occupy halos in denser environments. Our results highlight the potential of utilizing multiplet clustering, beyond traditional two-point correlation measurements, to break degeneracies in models describing the galaxy-dark matter connection.

cosmology↗

On the morphology of the gamma-ray galactic centre excess

ABSTRACT The characteristics of the galactic centre excess (GCE) emission observed in gamma-ray energies – especially the morphology of the GCE – remain a hotly debated subject. The manner in which the dominant diffuse gamma-ray background is modelled has been claimed to have a determining effect on the preferred morphology. In this work, we compare two distinct approaches to the galactic diffuse gamma-ray emission background: the first approach models this emission through templates calculated from a sequence of well-defined astrophysical assumptions, while the second approach divides surrogates for the background gamma-ray emission into cylindrical galactocentric rings with free independent normalizations. At the latitudes that we focus on, we find that the former approach works better, and that the overall best fit is obtained for an astrophysically motivated fit when the GCE follows the morphology expected of dark matter annihilation. Quantitatively, the improvement compared with the best ring-based fits is roughly 6500 in the χ2 and roughly 4000 in the log of the Bayesian evidence.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Using cloud radar to investigate the effect of rainfall on migratory insect flight

The fate of migrating insects that encounter rainfall in flight is a critical consideration when modelling insect movement, but few field observations of this common phenomenon have ever been collected due to the logistical challenges of witnessing these encounters. Operational cloud radars have been deployed around the world by meteorological agencies to study precipitation physics, and as a byproduct, provide a rich database of insect observations that is freely available to researchers. Although considered unwanted ‘clutter’ by the meteorologists who collect the data, the analysis method presented here enables ecologists to delineate co-occurring signals from insects and raindrops. We present a method that uses image processing techniques on cloud radar velocity spectra to examine the fate of migrating insects when they encounter precipitation. By analysing velocity spectra, we can distinguish flying insects from falling rain and compare the relative density of insects in flight before, during and after the rainfall. We demonstrate the method on a case of insect migration in Oklahoma, USA. Using this method, we show the first reconstructed images of migrating insect layers in flight during rainfall. Our analysis shows that mild to moderate rainfall diminishes the number of insects aloft but does not cause full termination of migratory flight, as has previously been suggested. We hope this technique will spur further investigations of how changing weather conditions impact insect migration, and enable some of the first of such studies in regions of the world that are underrepresented in the literature.

54 ENVIRONMENTAL SCIENCES↗

Quantitative Scanning Transmission Electron Microscopy for Materials Science: Imaging, Diffraction, Spectroscopy, and Tomography

Scanning transmission electron microscopy (STEM) is one of the most powerful characterization tools in materials science research. Due to instrumentation developments such as highly coherent electron sources, aberration correctors, and direct electron detectors, STEM experiments can examine the structure and properties of materials at length scales of functional devices and materials down to single atoms. STEM encompasses a wide array of flexible operating modes, including imaging, diffraction, spectroscopy, and 3D tomography experiments. This review outlines many common STEM experimental methods with a focus on quantitative data analysis and simulation methods, especially those enabled by open source software. The hope is to introduce both classic and new experimental methods to materials scientists and summarize recent progress in STEM characterization. The review also discusses the strengths and weaknesses of the various STEM methodologies and briefly considers promising future directions for quantitative STEM research.

36 MATERIALS SCIENCE↗

Field Validation of MVA Technology for Offshore CCS: Novel Ultra-High-Resolution 3D Marine Seismic Technology (P-Cable) (Final Report)

The objectives of the proposed study were to deploy and validate a specific monitoring technology, high-resolution 3D marine seismic (HR3D), appropriate for large-demonstration and commercial-scale offshore CCS sites. The project accomplished successful acquisition two HR3D seismic surveys. The first HR3D dataset was over the offshore injection site of the Tomakomai, Japan integrated pilot CCS project, which at the time of survey acquisition was actively injecting CO 2 . The first survey also represented a successful international collaboration between the DOE NETL program and Japan’s national CCS program and was the first successful acquisition and use of HR3D over an active CO 2 injection site (Meckel, Feng et al. 2019). The Tomakomai HR3D survey successfully tested a novel 4-streamer HR3D system array in which, for the first time, no cross-cable (aka “P-Cable”) was utilized and only four GeoEel streamers were used instead of the standard 12-streamer configuration. Consequently, this was not, strictly speaking, a deployment of the “P-Cable” system of (Planke and Berndt 2004) but rather a modified version, thereof, and it is the first known demonstration of the modified system configuration. One very positive outcome from the Japanese collaboration earlier in the project was the ability to learn from the Japanese how they used tail buoys with GPS to determine the position of the seismic source and receivers in time and space. Based on that experience, GCCC designed and built six GPS receivers that could be used to position the streamer receivers and the seismic source via tail buoys. A fundamental advance that was made on the original design, was the ability to directly power the tail buoy GPS units and transfer data through the streamers (i.e., vs. the batteries used at Tomakomai). The bulkiness of the GPS batteries caused drag and episodic surging of the buoys, which affected data quality by lifting up the tail end of the streamers so the receivers were not at the same depth. The units were tested onshore for accuracy and functionality, and the design was subsequently and successfully tested in marine acquisition mode during the SLP survey acquisition. The marine acquisition test and survey satisfied Subtasks 2.2.2, Novel Positioning Technology Selection and Subtask 2.2.3, Novel Positioning Technology Deployment. Results of the novel positioning technology selection (Subtask 2.2.2) were considered successful and will be incorporated in future HR3D seismic acquisition projects to reduce costs, improve deployment safety at sea, and integrate both seismic and data recording via a single data transfer through the streamers to the recording system. The project also established a permitting process through NETL NEPA compliance, which included an Environmental Assessment in a marine setting and is required for conducting these types of surveys using Federal funding. The permitting process charted a “boilerplate,” which can allow future surveys related to other funded projects to move forward more expeditiously. Future improvements that could be considered are more robust seals on the GPS module and stronger materials (especially joints) on tail buoy fabrication. These would increase fixed costs, but would be advisable and probably more economic long-term if multiple HR3D surveys are planned. Project Accomplishments include: • Pre-survey Sensitivity Study • Marine geochemistry methods and data analysis • Successful HR3D seismic dataset acquired @ Tomakomai active CO 2 injection marine site • Developed advanced seismic processing techniques • No NRMS anomalies detected in overburden; Demonstration of containment • Repeatability study • Second survey collected @ San Luis Pass, TX • 4D application using positioning techniques developed in the project for monitoring were successful

3D seismic GPS positioning↗

Developing ML/AI Methods for High-Throughput Characterization of Multiple-Sensor Streams of Tokamak Dynamics for High-Speed Control (Final Report)

This project evaluated and developed new mathematical and algorithmic techniques capable of handling (in real-time) the growing amounts of data generated by modern fusion research. While existing numerical linear algebra (NLA) methods provide the backbone to classical data analysis and algorithms, these methods fundamentally do not port to distributed architectures nor do they allow low-latency data reduction for control. Motivated by the needs for modern fusion reactors, this project explored and implemented new numerical methods to characterize plasma dynamics, respond in real-time to discharge evolution, and to process massive-scale data accurately and rapidly more fully. This project links expertise in multiple-sensor diagnostics of tokamak plasma dynamics from Columbia University’s Plasma Physics Laboratory with expertise in massive-scale data reduction and extreme data control algorithms at Columbia University’s Data Science Institute. This interdisciplinary project (i) applied machine learning methods, (ii) implemented a properly-trained neural-network for very fast processing of high-speed plasma videography, and (ii) developed the applied mathematical methods, based on randomized-NLA (rNLA) routines, for data analysis, reduction, and real-time control. The Columbia University High Beta Tokamak-Extended Pulse (HBT-EP) facility provided data to test new algorithms and partnership with Columbia University's Data Sciences Institute evaluated the broader use of new algorithms for many challenging control applications.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

From Images to Dark Matter: End-to-end Inference of Substructure from Hundreds of Strong Gravitational Lenses

Abstract Constraining the distribution of small-scale structure in our universe allows us to probe alternatives to the cold dark matter paradigm. Strong gravitational lensing offers a unique window into small dark matter halos (<10 10 M ⊙ ) because these halos impart a gravitational lensing signal even if they do not host luminous galaxies. We create large data sets of strong lensing images with realistic low-mass halos, Hubble Space Telescope (HST) observational effects, and galaxy light from HST’s COSMOS field. Using a simulation-based inference pipeline, we train a neural posterior estimator of the subhalo mass function (SHMF) and place constraints on populations of lenses generated using a separate set of galaxy sources. We find that by combining our network with a hierarchical inference framework, we can both reliably infer the SHMF across a variety of configurations and scale efficiently to populations with hundreds of lenses. By conducting precise inference on large and complex simulated data sets, our method lays a foundation for extracting dark matter constraints from the next generation of wide-field optical imaging surveys.

79 ASTRONOMY AND ASTROPHYSICS↗