Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “maximum likelihood”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Analysis of overlapping count data

Counts of a specific characteristic were obtained within regions defined on an object that was manufactured in a proprietary setting. The count regions were altered during production and resulted in misaligned or overlapping count data. A closed-formula maximum likelihood estimator (MLE) of the new region means is derived using all of the available count data and an independent Poisson model. The MLE is shown to be preferable to estimators constructed using generalized linear models for the overlapping data setting. This closed-form estimator extends to over-dispersed overlapping count data as the quasi-MLE and also performs well with correlated overlapping count data. Standard errors for the estimator are approximated and are validated with a simulation study. Additionally, the methods are extended to overlapping multinomial data. Illustrative examples of the methods are provided throughout the paper and are reproducible with the supplemental R code. Additionally, proofs of the paper’s results are also included in the supplemental material.

97 MATHEMATICS AND COMPUTING↗

Tensor decompositions for count data that leverage stochastic and deterministic optimization

There is growing interest to extend low-rank matrix decompositions to multi-way arrays, or tensors. One fundamental low-rank tensor decomposition is the canonical polyadic decomposition (CPD). The challenge of fitting a low-rank, nonnegative CPD model to Poisson-distributed count data is of particular interest. Several popular algorithms use local search methods to approximate the maximum likelihood estimator (MLE) of the Poisson CPD model. Here, this work presents two new algorithms that extend state-of-the-art local methods for Poisson CPD. Hybrid GCP-CPAPR combines Generalized Canonical Decomposition (GCP) with stochastic optimization and CP Alternating Poisson Regression (CPAPR), a deterministic algorithm, to increase the probability of converging to the MLE over either method used alone. Restarted CPAPR with SVDrop uses a heuristic based on the singular values of the CPD model unfoldings to identify convergence toward optimizers that are not the MLE and restarts within the feasible domain of the optimization problem, thus reducing overall computational cost when using a multi-start strategy. We provide empirical evidence that indicates our approaches outperform existing methods with respect to converging to the Poisson CPD MLE.

CPAPR↗

Characterization of a novel HIV-1 circulating recombinant form, CRF91_cpx, comprising CRF02_AG, G, J, and U, mostly among men who have sex with men

Prospective molecular studies of HIV-1 pol region (2253–5250 in HXB2 genome) sequences from sequenced samples of 269 HIV-1-infected patients in Cyprus (2017–2021) revealed a transmission cluster of 14 unknown HIV-1 recombinants that were not classified as previously established CRFs. The earliest recombinant was collected in September 2017, and the transmission cluster continued to grow until November 2020. Near full-length HIV-1 genome sequences of the 11 of the 14 recombinants were successfully obtained (790–8795 in HXB2 genome) and aligned against a reference dataset of HIV-1 subtypes and CRFs. We employed MEGAX for maximum-likelihood tree construction (GTR model, 1000 bootstrap replicates), Cluster-Picker for phylogenetic clustering analysis (genetic distance ≤0.045, bootstrap support value ≥70%), and REGA-3.0 for subtype determination. Bootscan and similarity plot analyses (sliding window of 400 nucleotides overlapped by 40 nucleotides) were conducted using SimPlot-v3.5.1, and subregion confirmatory neighbour-joining tree analyses were conducted using MEGAX (Kimura two-parameter model, 1000 bootstrap replicates, ≥70% bootstrap-support value). Exclusive clustering of the HIV-1 recombinants revealed their uniqueness. The recombination analyses illustrated the same unique mosaic pattern with six putative intersubtype recombination breakpoints, seven fragments of subtypes CRF02_AG, G, J and an unclassified fragment. We conclusively characterized the mosaic structure of the novel HIV-1 CRF, named CRF91_cpx, by the Los Alamos HIV Sequence Database. Additionally, we identified a URF of CRF91_cpx with two additional recombination sites, generated by a recombination event between subtype B and CRF91_cpx. Since the identification of CRF91_cpx, two additional patient samples have been entered into the CRF91_cpx transmission cluster, demonstrating active growth.

60 APPLIED LIFE SCIENCES↗

Impurity gas detection for SNF canisters using probabilistic deep learning and acoustic sensing *

Abstract Monitoring impurity gases in spent nuclear fuel (SNF) canisters is a novel structural health monitoring approach for SNF in dry storage. The SNF canisters are sealed containers that do not facilitate visual access to the inside. Acoustic sensing can be deployed by taking advantage of the pathways unobstructed by internal hardware. Although the ultrasonic time-of-flight measurement can provide valuable information, it is limited in its ability to discern the concentration of only one impurity gas. As such, deep learning algorithms, particularly convolutional neural networks (CNNs), offer a promising solution. In this study, CNN-based probabilistic deep learning models were implemented to detect and quantify multiple impurity gases in helium. An experimental platform was established to simulate canister conditions, and ultrasonic test data were collected. The presence of argon and air in helium at concentrations ranging from 0% to 1.2% at increments of 0.05% was considered. The multi-layer perceptron, decision tree, and logistic regression classifiers achieved high accuracies when distinguishing pure helium from helium with impurities. CNN with dropout layers and CNN using maximum likelihood estimation showed a similar performance, indicating their ability to capture uncertainties. The ensemble CNN model exhibited improved predictions and the ability to balance individual gas concentration by integrating 1D- and 2D-CNN models. These findings contribute probabilistic deep learning solutions for impurity gas detection and analysis within SNF canisters, thus ensuring safe storage and management of SNFs.

47 OTHER INSTRUMENTATION↗

The Atacama Cosmology Telescope: arcminute-resolution maps of 18 000 square degrees of the microwave sky from ACT 2008–2018 data combined with Planck

This paper presents a maximum-likelihood algorithm for combining sky maps with disparate sky coverage, angular resolution and spatially varying anisotropic noise into a single map of the sky. We use this to merge hundreds of individual maps covering the 2008–2018 ACT observing seasons, resulting in by far the deepest ACT maps released so far. We also combine the maps with the full Planck maps, resulting in maps that have the best features of both Planck and ACT: Planck’s nearly white noise on intermediate and large angular scales and ACT’s high-resolution and sensitivity on small angular scales. The maps cover over 18 000 square degrees, nearly half the full sky, at 100, 150 and 220 GHz. Furthermore, they reveal 4 000 optically-confirmed clusters through the Sunyaev Zel’dovich effect (SZ) and 18 500 point source candidates at > 5σ, the largest single collection of SZ clusters and millimeter wave sources to date. The multi-frequency maps provide millimeter images of nearby galaxies and individual Milky Way nebulae, and even clear detections of several nearby stars. Other anticipated uses of these maps include, for example, thermal SZ and kinematic SZ cluster stacking, CMB cluster lensing and galactic dust science. The method itself has negligible bias. However, due to the preliminary nature of some of the component data sets, we caution that these maps should not be used for precision cosmological analysis. The maps are part of ACT DR5, and will be made available on LAMBDA no later than three months after the journal publication of this article, along with an interactive sky atlas.

79 ASTRONOMY AND ASTROPHYSICS↗

A possible mass distribution of primordial black holes implied by LIGO-Virgo

The LIGO-Virgo Collaboration has so far detected around 90 black holes, some of which have masses larger than what were expected from the collapse of stars. The mass distribution of LIGO-Virgo black holes appears to have a peak at ~ 30 M ⊙ and two tails on the ends. By assuming that they all have a primordial origin, we analyze the GWTC-1 (O1&O2) and GWTC-2 (O3a) datasets by performing maximum likelihood estimation on a broken power law mass function f ( m ), with the result f ∝ m 1.2 for m < 35 M ⊙ and f ∝ m -4 for m > 35 M ⊙ . This appears to behave better than the popular log-normal mass function. Surprisingly, such a simple and unique distribution can be realized in our previously proposed mechanism of PBH formation, where the black holes are formed by vacuum bubbles that nucleate during inflation via quantum tunneling. Moreover, this mass distribution can also provide an explanation to supermassive black holes formed at high redshifts.

Astronomy & Astrophysics↗

Effects of overlapping sources on cosmic shear estimation: Statistical sensitivity and pixel-noise bias

The next generation of dark-energy imaging surveys — so called “Stage-IV” surveys, such as that of the Rubin Observatory Legacy Survey of Space and Time (LSST) — will cross a threshold in the number density of detected sources on the sky that requires qualitatively different image analysis and measurement techniques compared to the current generation of Stage-III surveys. In Stage-IV surveys, a significant amount of the cosmologically useful information is due to sources whose images overlap with those of other sources on the sky. Here, we focus on the weak gravitational lensing probe, for which we expect the largest impact since the cosmic shear signal is primarily encoded in the estimated shapes of observed galaxies and thus directly impacted by overlaps. We introduce a framework based on the Fisher formalism to analyze the effect of the overlapping sources (“blending”) on the estimation of cosmic shear. This method gives concrete predictions for the minimum loss of information due to noise and blending for any choice of “deblending” scheme and shape-measurement algorithm. Our studies account for undetected sources but do not address their full effects and biases they may introduce. We use simulated images and predict this impact of blending for three surveys: the Dark Energy Survey (DES), the Hyper-Suprime Cam Subaru Strategic Program (HSC-SSP), and the Rubin LSST. Our methodology successfully estimates the statistical sensitivity to weak lensing for DES and HSC early results. For LSST, we present the expected loss in statistical sensitivity for the ten-year survey due to blending. We find that for approximately 62% of galaxies that are likely to be detected in full-depth LSST images, at least 1% of the flux in their pixels is from overlapping sources. We also find that the statistical correlations between measures of overlapping galaxies and, to a much lesser extent (0.2%) the higher shot noise level due to their presence, decrease the effective number density of galaxies, N eff , by ~ 18%. We calculate an upper limit on N eff of 39.4 galaxies per arcmin 2 in r band. We study the impact of stars on as a function of stellar density and illustrate the diminishing returns of extending the survey into lower Galactic latitudes. We extend the simulation-based Fisher formalism to predict the expected increase in pixel-noise bias due to blending for maximum-likelihood (ML) shape estimators. We find that noise bias depends sensitively on the particular shape estimator and measure of ensemble-average shape that is used, and properties of the galaxy that include redshift-dependent quantities such as size and luminosity. The source code for these studies is available online.[The documented software developed for the catalog-level studies are available in the open-source LSST DESC github repository https://github.com/LSSTDESC/WeakLensingDeblending. The software for analyzing one or two galaxies with user-defined parameters is in the open-source github repository https://github.com/ismael-mendoza/ShapeMeasurementFisherFormalism.]

79 ASTRONOMY AND ASTROPHYSICS↗

DESI DR1 Lyα 1D power spectrum: the optimal estimator measurement

The one-dimensional power spectrum P 1D of Lyα forest offers rich insights into cosmological and astrophysical parameters, including constraints on the sum of neutrino masses, warm dark matter models, and the thermal state of the intergalactic medium. We present the measurement of P 1D using the optimal quadratic maximum likelihood estimator applied to over 300,000 Lyα quasars from Data Release 1 (DR1) of the Dark Energy Spectroscopic Instrument (DESI) survey. This sample represents the largest to date for P 1D measurements and is larger than the Extended Baryon Oscillation Spectroscopic Survey (eBOSS) by a factor of 1.7. We conduct a meticulous investigation of instrumental and analysis systematics and quantify their impact on P 1D . This includes the development of a cross-exposure estimator that eliminates the need to model the pipeline noise and has strong potential for future P 1D measurements. We also present new insights into metal contamination through the 1D correlation function. Using a fitting function we measure the evolution of the Lyα forest bias with high precision: b F (z) = (-0.218 ± 0.002) × ((1 + z)/4) 2.96±0.06 . In a companion validation paper, we substantially extend our previous suite of CCD image simulations to quantify the pipeline's exquisite performance accurately. In another companion paper, we present DR1 P 1D measurements using the Fast Fourier Transform (FFT) approach to power spectrum estimation. These two measurements produce a forest bias parameter that differs by 2.2 sigma. However, our model is simplistic, so this disagreement will be investigated in future work.

Lyman alpha forest↗

DESI DR1 Ly α 1D power spectrum: the Fast Fourier Transform estimator measurement

Here, we present the one-dimensional Lyman-α forest power spectrum measurement derived from the data release 1 (DR1) of the Dark Energy Spectroscopic Instrument (DESI). The measurement of the Lyman-α forest power spectrum along the line of sight from high-redshift quasar spectra provides information on the shape of the linear matter power spectrum, neutrino masses, and the properties of dark matter. In this work, we use a Fast Fourier Transform (FFT)-based estimator, which is validated on synthetic data in a companion paper. Compared to the FFT measurement performed on the DESI early data release, we improve the noise characterization with a cross-exposure estimator and test the robustness of our measurement using various data splits. We also refine the estimation of the uncertainties and now present an estimator for the covariance matrix of the measurement. Furthermore, we compare our results to previous high-resolution and eBOSS measurements. In another companion paper, we present the same DR1 measurement using the Quadratic Maximum Likelihood Estimator (QMLE). These two measurements are consistent with each other and constitute the most precise one-dimensional power spectrum measurement to date, while being in good agreement with results from the DESI early data release.

Lyman alpha forest↗

Avoiding biases in binned fits

Abstract Binned maximum likelihood fits are an attractive option when analysing large datasets, but require care when computing likelihoods of continuous PDFs in bins.For many years the widely used statistical modelling package evaluated probabilities at the bin centre, leading to significant biases for strongly curved probability density functions.We demonstrate the biases with real-world examples, and introduce a PDF class to that removes these biases.The physics and computation performance of this new class are discussed.

Instruments & Instrumentation↗

Monte Carlo method for constructing confidence intervals with unconstrained and constrained nuisance parameters in the NOvA experiment

Measuring observables to constrain models using maximum-likelihood estimation is fundamental to many physics experiments. Wilks' theorem provides a simple way to construct confidence intervals on model parameters, but it only applies under certain conditions. These conditions, such as nested hypotheses and unbounded parameters, are often violated in neutrino oscillation measurements and other experimental scenarios. Monte Carlo methods can address these issues, albeit at increased computational cost. In the presence of nuisance parameters, however, the best way to implement a Monte Carlo method is ambiguous. Furthermore, this paper documents the method selected by the NOvA experiment, the profile construction. It presents the toy studies that informed the choice of method, details of its implementation, and tests performed to validate it. It also includes some practical considerations which may be of use to others choosing to use the profile construction.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Predicting the single-site and multi-site event discrimination power of dual-phase time projection chambers

Dual-phase xenon time projection chambers (TPCs) are widely used in searches for rare dark matter and neutrino interactions, in part because of their excellent position reconstruction capability in 3D. Despite their millimeter-scale resolution along the charge drift axis, xenon TPCs face challenges in resolving single-site (SS) and multi-site (MS) interactions in the transverse plane. In this paper, we build a generic TPC model with an idealized signal readout, and use Fisher Information (FI) to predict its theoretical capability of differentiating SS and MS events using the electroluminescence signal. We also demonstrate via simulation that, when only statistical photon noise is present, the theoretical limits can be approached with conventional reconstruction algorithms like maximum likelihood estimation, and with a convolutional neural network classifier. The implications of this study on future TPC experiments will be discussed.

Physics↗

Fast decay of classification error in variational quantum circuits

Variational quantum circuits (VQCs) have shown great potential in near-term applications. However, the discriminative power of a VQC, in connection to its circuit architecture and depth, is not understood. To unleash the genuine discriminative power of a VQC, we propose a VQC system with the optimal classical post-processing—maximum-likelihood estimation on measuring all VQC output qubits. Via extensive numerical simulations, we find that the error of VQC quantum data classification typically decays exponentially with the circuit depth, when the VQC architecture is extensive—the number of gates does not shrink with the circuit depth. This fast error suppression ends at the saturation towards the ultimate Helstrom limit of quantum state discrimination. On the other hand, non-extensive VQCs such as quantum convolutional neural networks are sub-optimal and fail to achieve the Helstrom limit, demonstrating a trade-off between ansatz complexity and classification performance in general. To achieve the best performance for a given VQC, the optimal classical post-processing is crucial even for a binary classification problem. To simplify VQCs for near-term implementations, we find that utilizing the symmetry of the input properly can improve the performance, while oversimplification can lead to degradation.

Zhang, Bingzhi↗

Hypothesis tests on Rayleigh wave radiation pattern shapes: a theoretical assessment of idealized source screening

SUMMARY Shallow seismic sources excite Rayleigh wave ground motion with azimuthally dependent radiation patterns. We place binary hypothesis tests on theoretical models of such radiation patterns to screen cylindrically symmetric sources (like explosions) from non-symmetric sources (like non-vertical dip-slip or non-VDS faults). These models for data include sources with several unknown parameters, contaminated by Gaussian noise and embedded in a layered half-space. The generalized maximum likelihood ratio tests that we derive from these data models produce screening statistics and decision rules that depend on measured, noisy ground motion at discrete sensor locations. We explicitly quantify how the screening power of these statistics increase with the size of any dip-slip and strike-slip components of the source, relative to noise (faulting signal strength) and how they vary with network geometry. As applications of our theory, we apply these tests to (1) find optimal sensor locations that maximize the probability of screening non-circular radiation patterns and (2) invert for the largest non-VDS faulting signal that could be mistakenly attributed to an explosion with damage, at a particular attribution probability. Finally, we quantify how certain errors that are sourced by opening cracks increase screening rate errors. While such theoretical solutions are ideal and require future validation, they remain important in underground explosion monitoring scenarios because they provide fundamental physical limits on the discrimination power of tests that screen explosive from non-VDS faulting sources.

58 GEOSCIENCES↗

Graph-learning approach to combine multiresolution seismic velocity models

SUMMARY The resolution of velocity models obtained by tomography varies due to multiple factors and variables, such as the inversion approach, ray coverage, data quality, etc. Combining velocity models with different resolutions can enable more accurate ground motion simulations. Toward this goal, we present a novel methodology to fuse multiresolution seismic velocity maps with probabilistic graphical models (PGMs). The PGMs provide segmentation results, corresponding to various velocity intervals, in seismic velocity models with different resolutions. Further, by considering physical information (such as ray path density), we introduce physics-informed probabilistic graphical models (PIPGMs). These models provide data-driven relations between subdomains with low (LR) and high (HR) resolutions. Transferring (segmented) distribution information from the HR regions enhances the details in the LR regions by solving a maximum likelihood problem with prior knowledge from HR models. When updating areas bordering HR and LR regions, a patch-scanning policy is adopted to consider local patterns and avoid sharp boundaries. To evaluate the efficacy of the proposed PGM fusion method, we tested the fusion approach on both a synthetic checkerboard model and a fault zone structure imaged from the 2019 Ridgecrest, CA, earthquake sequence. The Ridgecrest fault zone image consists of a shallow (top 1 km) high-resolution shear-wave velocity model obtained from ambient noise tomography, which is embedded into the coarser Statewide California Earthquake Center Community Velocity Model version S4.26-M01. The model efficacy is underscored by the deviation between observed and calculated traveltimes along the boundaries between HR and LR regions, 38 per cent less than obtained by conventional Gaussian interpolation. The proposed PGM fusion method can merge any gridded multiresolution velocity model, a valuable tool for computational seismology and ground motion estimation.

Geochemistry & Geophysics↗

Removing imaging systematics from galaxy clustering measurements with Obiwan: application to the SDSS-IV extended Baryon Oscillation Spectroscopic Survey emission-line galaxy sample

This article presents the application of a new tool, Obiwan, which uses image simulations to determine the selection function of a galaxy redshift survey and calculate three-dimensional (3D) clustering statistics. Obiwan relies on a forward model of the process by which images of the night sky are transformed into a 3D large-scale structure catalogue, and offers several advantages over more traditional map-based techniques – such as operating on individual exposures and adopting a maximum likelihood approach. The photometric pipeline automatically detects and models galaxies and then generates a catalogue of such galaxies with detailed information for each one of them, including their location, redshift, and so on. Systematic biases in the imaging data are therefore imparted into the catalogues and must be accounted for in any scientific analysis of their information content. Obiwan simulates this process for samples selected from the Legacy Surveys imaging data. This imaging data will be used to select target samples for the next-generation Dark Energy Spectroscopic Instrument (DESI) experiment. Here, we apply Obiwan to a portion of the SDSS-IV extended Baryon Oscillation Spectroscopic Survey emission-line galaxies (ELGs). Systematic biases in the data are clearly identified and removed. We compare the 3D clustering results to those obtained by the map-based approach applied to the complete eBOSS Data Release 16 (DR16) sample. We find the results are consistent, thereby validating the eBOSS DR16 ELG catalogues, which is used to obtain cosmological results.

79 ASTRONOMY AND ASTROPHYSICS↗

Joint inference of multiplicative and additive systematics in galaxy density fluctuations and clustering measurements

Galaxy clustering measurements are a key probe of the matter density field in the Universe. With the era of precision cosmology upon us, surveys rely on precise measurements of the clustering signal for meaningful cosmological analysis. However, the presence of systematic contaminants can bias the observed galaxy number density, and thereby bias the galaxy two-point statistics. As the statistical uncertainties get smaller, correcting for these systematic contaminants becomes increasingly important for unbiased cosmological analysis. We present and validate a new method for understanding and mitigating both additive and multiplicative systematics in galaxy clustering measurements (two-point function) by joint inference of contaminants in the galaxy overdensity field (one-point function) using a maximum-likelihood estimator (MLE). We test this methodology with Kilo-Degree Survey-like mock galaxy catalogues and synthetic systematic template maps. We estimate the cosmological impact of such mitigation by quantifying uncertainties and possible biases in the inferred relationship between the observed and the true galaxy clustering signal. Our method robustly corrects the clustering signal to the sub-percent level and reduces numerous additive and multiplicative systematics from 1.5σ to less than 0.1σ for the scenarios we tested. In addition, we provide an empirical approach to identifying the functional form (additive, multiplicative, or other) by which specific systematics contaminate the galaxy number density. Even though this approach is tested and geared towards systematics contaminating the galaxy number density, the methods can be extended to systematics mitigation for other two-point correlation measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

GalaxyFlow: upsampling hydrodynamical simulations for realistic mock stellar catalogues

ABSTRACT Cosmological N-body simulations of galaxies operate at the level of ‘star particles’ with a mass resolution on the scale of thousands of solar masses. Turning these simulations into stellar mock catalogues requires ‘upsampling’ the star particles into individual stars following the same phase-space density. In this paper, we introduce two new upsampling methods. First, we describe GalaxyFlow, a sophisticated upsampling method that utilizes normalizing flows to both estimate the stellar phase-space density and sample from it. Secondly, we improve on existing upsamplers based on adaptive kernel density estimation (KDE), using maximum likelihood estimation to fine-tune the bandwidth for such algorithms in a way that improves both the density estimation accuracy and upsampling results. We demonstrate our upsampling techniques on a neighbourhood of the Solar location in two simulated galaxies: Auriga 6 and h277. Both yield smooth stellar distributions that closely resemble the stellar densities seen in the Gaia DR3 catalogue. Furthermore, we introduce a novel multimodel classifier test to compare the accuracy of different upsampling methods quantitatively. This test confirms that GalaxyFlow more accurately estimates the density of the underlying star particles than methods based on KDE, at the cost of being more computationally intensive.

Lim, Sung Hak (ORCID:0000000330981092)↗