Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gaussian process fitting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

How “hot” are hotspots: Statistically localizing the high-activity areas on soil and rhizosphere images

The topic of microbial hotspots in soil requires not only visualizing their spatial distribution and biochemical analyses, but also statistical approaches to identify these hotspots and separate them from the surrounding activities (background). We hypothesized that each hotspot type (e.g. enzyme activities in the rhizosphere, root exudation, localization of herbicide accumulation) is a result of local process driven by biotic and/or abiotic factors, and the process rates in the hotspots are much faster than those in the soil background. We further hypothesized that the background and hotspot activities in soil belong to different statistical distributions. Consequently, hotspot determination should be based on statistical separation of activities significantly higher than the background. We analyzed for the statistical distributions of grey values on three groups of published images: 1) 14 C images of carbon input by roots into the rhizosphere, 2) 14 C glyphosate accumulation in the plant, and 3) zymogram of leucine aminopeptidase activity in rooted soil. The two Gaussian distributions were fit (the first representing the background, the second the hotspots) to the distribution of grey values in the images, the parameters (means and standard deviations, SD) of the fitted distributions were calculated, and the background was removed. Thus, we identified hotspots as areas outside of the Mean+2SD image intensity (corresponding to the upper ~ 2.5% of activity, being over 97.5% of background values) and finally, visualized images of solely hotspot locations. Finally, these results were compared with previously used decisions on hotspot intensity thresholding (i.e. Top-25% and 17 standard thresholding approaches in ImageJ) and discussed the advantages of the Mean+2SD as well as Mean+3SD approaches. These advantages include: i) simple unification of the thresholding approach for several imaging methods with various principles of activity distribution, ii) identification of hotspots with various activity levels, iii) analysis of “time-specific” hotspots in temporal sequences of images. Compared with 17 standard thresholding methods, we concluded that objectively elucidating and separating the hotspots should be based on statistical distribution analysis, e.g. using the Mean+2SD or Mean+3SD approaches. Furthermore, this simple Mean+2SD approach delivered suitable results for three groups of images and so, helps to understand the processes responsible for the highest activities and elucidate hotspots.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling the East-West Asymmetry of Energetic Particle Fluence in Large Solar Energetic Particle Events Using the iPATH Model

It has been noted that in large solar energetic particle (SEP) events, the peak intensities show an East-West asymmetry with respect to the source flare locations. Using the 2D improved Particle Acceleration and Transport in the Heliosphere (iPATH) model, we investigate the origin of this longitudinal trend. We consider multiple cases with different solar wind speeds and eruption speeds of the coronal mass ejections (CMEs) and fit the longitudinal distributions of time-averaged fluence by symmetric/asymmetric Gaussian functions with three time intervals of 8, 24 and 48 hr after the flare onset time respectively. The simulation results are compared with a statistical study of three-spacecraft events. We suggest that the East-West asymmetry of SEP fluence and peak intensity can be primarily caused the combined effect of an extended shock acceleration process and the evolution of magnetic field connection to the shock front. Our simulations show that the solar wind speed and the CME speed are important factors determining the East-West fluence asymmetry.

Zheyi Ding↗

A Machine Learning based Approach of Estimating Equivalent Circuit Model Parameters at Different SoCs of Li-ion Batteries from Voltage Relaxation

Abstract: In this study, an approach of estimating the equivalent circuit model (ECM) parameters for Li-ion batteries (LIBs) is proposed based on the voltage value at different intervals while relaxing the LIB after discharge. The typical approach for estimating ECM parameters of a LIB is to conduct electrochemical impedance spectroscopy (EIS) measurements at different frequencies and fit them to a predefined circuit model, which requires additional measuring arrangements and specialized devices. The proposed methodology utilizes four different voltages at 0s, 60s, 360s, and 1800s alongside the specific state of charge (SoC) value for a specific constant discharge current value of ~1C until the relaxation stage to train and evaluate three regression-based machine learning models— Support Vector Regression (SVR), Extreme Gradient Boosting (XGBoost), and Gaussian Process Regression (GPR)—for estimating the ECM parameters of the selected model. Bayesian optimization is employed for hyperparameter tuning to achieve optimal performance for all the regressor models, among which, the GPR provided the best performance with the root-mean-squared error (RMSE) of less than 4x10-4 on average for the resistive components and less than 0.27 for capacitive components with excellent R2 scores. The simplicity of the approach enables it to eliminate the need for sophisticated measuring equipment and computation power.

Sagar, Md. Samiul [The University of Alabama (UA)]↗

Physics-Based Machine Learning Methods for U-235 Forensics Signatures

Signatures of low-intensity U-235 sources have been recently studied by utilizing a variety of machine learning (ML) classifiers using features derived from gamma spectral measurements collectedunder structured campaigns. Several ML classifiers, such as ensemble of tress and classification trees, revealed misleadingly-optimistic training error due to over-fitting, and furthermore,their performance is not directly relatable to the physical properties due to their data-driven, opaque designs. We present a regression-based ML method that first estimates the inverse distanceto the source and then utilizes a threshold to infer its presence, by representing the background as a source located at an infinite distance. For the inverse distance estimation, we study the ensembleof trees and Gaussian process regression methods, and a hyper parameter auto-tuning and selection method that employs five regression estimators. These methods avoid the over-fittingobserved in several ML classifiers, while providing the classification error nearly comparable to them based on independent test data. Their error is directly related to estimates of the inversephysical distance to source, and the precision of error determines the seperability property that determines the false alarm and missed detection rates. The property of monotonic decrease of thesource strength with increasing detector distance combined with Poisson distribution of measurements is utilized to analytically validate these methods by deriving the generalization equations ofunderlying regression methods.

Rao, Nageswara↗

The Asymmetric Inner Disk of the Herbig Ae Star HD 163296 in the Eyes of VLTI/MATISSE: Evidence for a Vortex?

Context.A complex environment exists in the inner few astronomical units of planet-forming disks. High-angular-resolution observa-tions play a key role in our understanding of the disk structure and the dynamical processes at work.Aims.In this study we aim to characterize the mid-infrared brightness distribution of the inner disk of the young intermediate-massstar HD 163296 from early VLTI/MATISSE observations taken in theL- andN-bands. We put special emphasis on the detection ofpotential disk asymmetries.Methods.We use simple geometric models to fit the interferometric visibilities and closure phases. Our models include a smoothedring, a flat disk with an inner cavity, and a 2D Gaussian. The models can account for disk inclination and for azimuthal asymmetriesas well. We also perform numerical hydrodynamical simulations of the inner edge of the disk.Results.Our modeling reveals a significant brightness asymmetry in theL-band disk emission. The brightness maximum of the asym-metry is located at the NW part of the disk image, nearly at the position angle of the semimajor axis. The surface brightness ratio inthe azimuthal variation is3.5±0.2. Comparing our result on the location of the asymmetry with other interferometric measurements,we confirm that the morphology of ther<0.3au disk region is time-variable. We propose that this asymmetric structure, located in ornear the inner rim of the dusty disk, orbits the star. To find the physical origin of the asymmetry, we tested a hypothesis where a vortexis created by Rossby wave instability, and we find that a unique large-scale vortex may be compatible with our data. The half-lightradius of theL-band-emitting region is0.33±0.01au, the inclination is52◦+5◦−7◦, and the position angle is143◦±3◦. Our models predictthat a non-negligible fraction of theL-band disk emission originates inside the dust sublimation radius forμm-sized grains. Refractorygrains or large (&10μm-sized) grains could be the origin of this emission.N-band observations may also support a lack of smallsilicate grains in the innermost disk (r.0.6au), in agreement with our findings fromL-band data.

J Varga↗

A machine learning pipeline for membrane segmentation of cryo-electron tomograms

We describe how to use several machine learning techniques organized in a learning pipeline to segment and identify cell membrane structures from cryo electron tomograms. These tomograms are difficult to analyze with traditional segmentation tools. The learning pipeline in our approach starts from supervised learning via a special convolutional neural network trained with simulated data. It continues with semi-supervised reinforcement learning and/or a region merging technique that tries to piece together disconnected components belonging to the same membrane structure. A parametric or non-parametric fitting procedure is then used to enhance the segmentation results and quantify uncertainties in the fitting. Domain knowledge is used in generating the training data for the neural network and in guiding the fitting procedure through the use of appropriately chosen priors and constraints. We demonstrate that the approach proposed here works well for extracting membrane surfaces in two real tomogram datasets.

97 MATHEMATICS AND COMPUTING↗

Multiple emission components in the Cygnus cocoon detected from Fermi -LAT observations

Star-forming regions may play an important role in the life cycle of Galactic cosmic rays (CRs), notably as home to specific acceleration mechanisms and transport conditions. Gamma-ray observations of Cygnus X have revealed the presence of an excess of hard-spectrum gamma-ray emission, possibly related to a cocoon of freshly accelerated particles. We seek an improved description of the gamma-ray emission from the cocoon using ~13 yr of observations with the Fermi-Large Area Telescope (LAT) and use it to further constrain the processes and objects responsible for the young CR population. We developed an emission model for a large region of interest, including a description of interstellar emission from the background population of CRs and recent models for other gamma-ray sources in the field. Thus, we performed an improved spectro-morphological characterisation of the residual emission including the cocoon. The best-fit model for the cocoon includes two main emission components: an extended component FCES G78.74+1.56, described by a 2D Gaussian of extension r 68 = 4.4° ± 0.1° -0.1° +0.1° and a smooth broken power law spectrum with spectral indices 1.67 ± 0.05 -0.01 +0.02 and 2.12 ± 0.02 -0.01 +0.00 below and above 3.0 ± 0.6 -0.2 +0.0 GeV, respectively; and a central component FCES G80.00+0.50, traced by the distribution of ionised gas within the borders of the photo-dissociation regions and with a power law spectrum of index 2.19 ± 0.03 -0.01 +0.00 that is significantly different from the spectrum of FCES G78.74+1.56. An additional extended emission component FCES G78.83+3.57, located on the edge of the central cavities in Cygnus X and with a spectrum compatible with that of FCES G80.00+0.50, is likely related to the cocoon. For the two brightest components FCES G80.00+0.50 and FCES G78.74+1.56, spectra and radial-azimuthal profiles of the emission can be accounted for in a diffusion-loss framework involving one single population of non-thermal particles with a flat injection spectrum. Particles span the full extent of FCES G78.74+1.56 as a result of diffusion from a central source, and give rise to source FCES G80.00+0.50 by interacting with ionised gas in the innermost region. For this simple diffusion-loss model, viable setups can be very different in terms of energetics, transport conditions, and timescales involved, and both hadronic and leptonic scenarios are possible. The solutions range from long-lasting particle acceleration, possibly in prominent star clusters such as Cyg OB2 and NGC 6910, to a more recent and short-lived release of particles within the last 10–100 kyr, likely from a supernova remnant. The observables extracted from our analysis can be used to perform detailed comparisons with advanced models of particle acceleration and transport in star-forming regions.

79 ASTRONOMY AND ASTROPHYSICS↗

Spatiotemporal Downscaling Model for Solar Irradiance Forecast Using Nearest-Neighbor Random Forest and Gaussian Process

Accurate solar photovoltaic (PV) capacity estimation requires high-resolution, site-specific solar irradiance data to account for localized variability. However, global datasets, such as the National Solar Radiation Database (NSRDB), provide regional averages that fail to capture the fine-scale fluctuations critical for large-scale grid integration. This limitation is particularly relevant in the context of increasing distributed energy resources (DERs) penetration, such as rooftop PV. Additionally, it is critical to the implementation of the U.S. Federal Energy Regulatory Commission (FERC) Order 2222, which facilitates DER participation in U.S. bulk power markets. To address this challenge, this study evaluates Nearest-Neighbor Random Forest (NNRF) and Nearest-Neighbor Gaussian Process (NNGP) models for spatiotemporal downscaling of global solar irradiance data. By leveraging historical irradiance and meteorological data, these models incorporate spatial, temporal, and feature-based correlations to enhance local irradiance predictions. The NNRF model, a machine-learning approach, prioritizes computational efficiency and predictive accuracy, while the NNGP model offers a level of interpretability and prediction uncertainty by numerically quantifying correlations and dependencies in the data. Model validation was conducted using day-ahead predictions. The results showed that the average Goodness of Fit (GoF) of the NNRF model of 90.61% across all eight sites outperformed the GoF of the NNGP of 85.88%. Additionally, the computational speed of NNRF was 2.5 times faster than the NNGP. Finally, the NNGP displayed polynomial scaling while the NNRF scaled linearly with increasing number of nearest neighbors. Additional validation of the model on five sites in Puerto Rico further confirmed the superiority of the NNRF model over the NNGP model. These findings highlight the robustness and computational efficiency of NNRF for large-scale solar irradiance downscaling, making it a strong candidate for improving PV capacity estimation and real-time electricity market integration for DERs.

Asiedu, Shadrack (ORCID:0009000646004826)↗

Emulating the Lyman-Alpha forest 1D power spectrum from cosmological simulations: new models and constraints from the eBOSS measurement

We present the Lyssa suite of high-resolution cosmological simulations of the Lyman-α forest designed for cosmological analyses. These 18 simulations have been run using the Nyx code with 40963 hydrodynamical cells in a 120 Mpc (∼ 81 Mpc/h) comoving box and individually provide sub-percent level convergence of the Lyman-α forest 1d flux power spectrum. We build a Gaussian process emulator for the Lyssa simulations in the lym1d likelihood framework to interpolate the power spectrum at arbitrary parameter values. We validate this emulator based on leave-one-out tests and based on the parameter constraints for simulations outside of the training set. We also perform comparisons with a previous emulator, showing a percent level accuracy and a good recovery of the expected cosmological parameters. Using this emulator we derive constraints on the linear matter power spectrum amplitude and slope parameters A Lyα and n Lyα . While the best-fit Planck ΛCDM model has A Lyα = 8.79 and n Lyα = -2.363, from DR14 eBOSS data we find that A Lyα < 7.6 (95% CI) and n Lyα = -2.369 ± 0.008. The low value of A Lyα , in tension with Planck, is driven by the correlation of this parameter with the mean transmission of the Lyman-α forest. This tension disappears when imposing a well-motivated external prior on this mean transmission, in which case we find A Lyα = 9.8 ± 1.1 in accordance with Planck.

Walther, Michael↗

Laser transit anemometer software development program

Algorithms were developed for the extraction of two components of mean velocity, standard deviation, and the associated correlation coefficient from laser transit anemometry (LTA) data ensembles. The solution method is based on an assumed two-dimensional Gaussian probability density function (PDF) model of the flow field under investigation. The procedure consists of transforming the data ensembles from the data acquisition domain (consisting of time and angle information) to the velocity space domain (consisting of velocity component information). The mean velocity results are obtained from the data ensemble centroid. Through a least squares fitting of the transformed data to an ellipse representing the intersection of a plane with the PDF, the standard deviations and correlation coefficient are obtained. A data set simulation method is presented to test the data reduction process. Results of using the simulation system with a limited test matrix of input values is also given.

Abbiss, John B.↗

Planck 2018 results

We analyse the Planck full-mission cosmic microwave background (CMB) temperature and E -mode polarization maps to obtain constraints on primordial non-Gaussianity (NG). We compare estimates obtained from separable template-fitting, binned, and optimal modal bispectrum estimators, finding consistent values for the local, equilateral, and orthogonal bispectrum amplitudes. Our combined temperature and polarization analysis produces the following final results: f NL local = -0.9 ± 5.1; f NL equil = -26 ± 47; and f NL ortho = -38 ± 24 (68% CL, statistical). These results include low-multipole (4 ≤ ℓ < 40) polarization data that are not included in our previous analysis. The results also pass an extensive battery of tests (with additional tests regarding foreground residuals compared to 2015), and they are stable with respect to our 2015 measurements (with small fluctuations, at the level of a fraction of a standard deviation, which is consistent with changes in data processing). Polarization-only bispectra display a significant improvement in robustness; they can now be used independently to set primordial NG constraints with a sensitivity comparable to WMAP temperature-based results and they give excellent agreement. In addition to the analysis of the standard local, equilateral, and orthogonal bispectrum shapes, we consider a large number of additional cases, such as scale-dependent feature and resonance bispectra, isocurvature primordial NG, and parity-breaking models, where we also place tight constraints but do not detect any signal. The non-primordial lensing bispectrum is, however, detected with an improved significance compared to 2015, excluding the null hypothesis at 3.5 σ . Beyond estimates of individual shape amplitudes, we also present model-independent reconstructions and analyses of the Planck CMB bispectrum. Our final constraint on the local primordial trispectrum shape is g NL local = (-5.8 ± 6.5) × 10 4 (68% CL, statistical), while constraints for other trispectrum shapes are also determined. Exploiting the tight limits on various bispectrum and trispectrum shapes, we constrain the parameter space of different early-Universe scenarios that generate primordial NG, including general single-field models of inflation, multi-field models (e.g. curvaton models), models of inflation with axion fields producing parity-violation bispectra in the tensor sector, and inflationary models involving vector-like fields with directionally-dependent bispectra. Our results provide a high-precision test for structure-formation scenarios, showing complete agreement with the basic picture of the ΛCDM cosmology regarding the statistics of the initial conditions, with cosmic structures arising from adiabatic, passive, Gaussian, and primordial seed perturbations.

79 ASTRONOMY AND ASTROPHYSICS↗

Near-Infrared Spectroscopy can Predict Anatomical Abundance in Corn Stover

Feedstock heterogeneity is a key challenge impacting the deconstruction and conversion of herbaceous lignocellulosic biomass to biobased fuels, chemicals, and materials. Upstream processing to homogenize biomass feedstock streams into their anatomical components via air classification allows for a more tailored approach to subsequent mechanical and chemical processing. Here, we show that differing corn stover anatomical tissues respond differently to pretreatment and enzymatic hydrolysis and therefore, a one-size-fits-all approach to chemical processing biomass is inappropriate. To inform on-line downstream processing, a robust and high-throughput analytical technique is needed to quantitatively characterize the separated biomass. Predictive correlation of near-infrared spectra to biomass chemical composition is such a technique. Here, we demonstrate the capability of models developed using an “off-the-shelf,” industrially relevant spectrometer with limited spectral range to make strong predictions of both cell wall chemical composition and the relative abundance of anatomical components of the corn stover, the latter for the first time ever. Gaussian process regression (GPR) yields stronger correlations (average R 2 v = 88% for chemical composition and 95% for anatomical relative abundance) than the more commonly used partial least squares (PLS) regression (average R 2 v = 84% for chemical composition and 92% for anatomical relative abundance). In nearly all cases, both GPR and PLS outperform models generated using neural networks. These results highlight the potential for coupling NIRS with predictive models based on GPR due to the potential to yield more robust correlations.

09 BIOMASS FUELS↗

Distributed memory, GPU accelerated Fock construction for hybrid, Gaussian basis density functional theory

With the growing reliance of modern supercomputers on accelerator-based architecture such a graphics processing units (GPUs), the development and optimization of electronic structure methods to exploit these massively parallel resources has become a recent priority. While significant strides have been made in the development GPU accelerated, distributed memory algorithms for many modern electronic structure methods, the primary focus of GPU development for Gaussian basis atomic orbital methods has been for shared memory systems with only a handful of examples pursing massive parallelism. Here in this work, we present a set of distributed memory algorithms for the evaluation of the Coulomb and exact exchange matrices for hybrid Kohn–Sham DFT with Gaussian basis sets via direct density-fitted (DF-J-Engine) and seminumerical (sn-K) methods, respectively. The absolute performance and strong scalability of the developed methods are demonstrated on systems ranging from a few hundred to over one thousand atoms using up to 128 NVIDIA A100 GPUs on the Perlmutter supercomputer.

97 MATHEMATICS AND COMPUTING↗

The Langdon effect in laser plasmas: Absorption and conduction

A plasma heated by inverse bremsstrahlung absorption of laser light develops a non-Maxwellian electron distribution function, called the Langdon effect [A. B. Langdon, Phys. Rev. Lett. 44, 575 (1980)]. These non-Maxwellian distributions are sufficiently long-lived to impact the absorption processes itself as well as the transport of heat by electrons. The theory of the Langdon effect in a homogeneous plasma is reviewed to clarify some aspects of Langdon's derivation as well as to confirm that the widely used super-Gaussian approximation works fairly well to describe the shape of the distribution function and reduction of the absorption rate. The Langdon effect on thermal conduction in an inhomogeneous plasma is developed by considering perturbations in a homogeneous absorbing plasma, which develops a heat flux due to both temperature and density gradients. A practical theory of the heat flux is developed by fitting the results of Vlasov–Fokker–Planck simulations, which avoids several approximations that compromised the usefulness of past theoretical predictions, most critically, the effect of electron–electron collisions on the fluxes. The present fits parameterize the coefficients of the temperature gradient (thermal conductivity) and the density gradient for a plasma of any ionization state and for any laser intensity where the theory of the Langdon effect remains locally valid. It is expected that this generalized theory of heat flow in an absorbing plasma will improve the predictive capability of radiation-hydrodynamics simulations of laser-produced plasmas, especially those formed in inertial confinement fusion experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Application of machine learning and artificial intelligence to extend EFIT equilibrium reconstruction

Recent progress in the application of machine learning (ML)/artificial intelligence (AI) algorithms to improve the Equilibrium Fitting (EFIT) code equilibrium reconstruction for fusion data analysis applications is presented. A device-independent portable core equilibrium solver capable of computing or reconstructing equilibrium for different tokamaks has been created to facilitate adaptation of ML/AI algorithms. A large EFIT database comprising of DIII-D magnetic, motional Stark effect, and kinetic reconstruction data has been generated for developments of EFIT model-order-reduction (MOR) surrogate models to reconstruct approximate equilibrium solutions. Furthermore, a neural-network MOR surrogate model has been successfully trained and tested using the magnetically reconstructed datasets with encouraging results. Other progress includes developments of a Gaussian process Bayesian framework that can adapt its many hyperparameters to improve processing of experimental input data and a 3D perturbed equilibrium database from toroidal full magnetohydrodynamic linear response modeling using the Magnetohydrodynamic Resistive Spectrum - Feedback (MARS-F) code for developments of 3D-MOR surrogate models.

Gaussian process↗

Score-Based Physics-Informed Neural Networks for High-Dimensional Fokker–Planck Equations

The Fokker-Planck (FP) equation is a foundational partial differential equation (PDE) in stochastic processes involving Brownian motions. However, the curse of dimensionality (CoD) poses a formidable challenge when dealing with high-dimensional FP equations. Although Monte Carlo simulation and (vanilla) Physics-Informed Neural Networks (PINNs) have shown the potential to tackle CoD, both methods exhibit significant numerical errors in high dimensions when dealing with the probability density function (PDF) associated with Brownian motion. The point-wise PDF values tend to decrease exponentially as dimensionality increases, surpassing the precision of numerical simulations and resulting in substantial errors. In addition, due to its massive sampling, Monte Carlo fails to offer fast sampling. Modeling the logarithm likelihood (LL) via vanilla PINNs transforms the FP equation into a notoriously difficult Hamilton-Jacobi-Bellman (HJB) equation, which is impractical for PINN learning, whose error grows rapidly with dimension. To this end, we propose a novel approach utilizing a score-based solver to fit the score function in stochastic differential equations (SDEs). The score function, defined as the gradient of the LL, plays a fundamental role in inferring LL and PDF and enables fast SDE sampling, offering an effective means to overcome the CoD. Three fitting methods, Score Matching (SM), Sliced Score Matching (SSM), and Score-PINN, are introduced, each contributing unique advantages in computational complexity, accuracy, and generality. The proposed score-based SDE solver operates in two stages: first, employing score matching or Score-PINN to acquire the score function; and second, solving the LL via an ordinary differential equation (ODE) using the obtained score function. Comparative evaluations across these methods showcase varying trade-offs. The proposed methodology is evaluated across diverse SDEs, including anisotropic Ornstein-Uhlenbeck processes, geometric Brownian motion, and Brownian motion with varying eigenspace. We also test various distributions, including Gaussian, Log-normal, Laplace, and Cauchy distributions. The numerical results demonstrate the score-based SDE solver’s stability, speed, and performance across different experimental settings, solidifying its potential as a solution to CoD for high-dimensional FP equations.

97 MATHEMATICS AND COMPUTING↗

Learning Model Structural Uncertainty with Gaussian Processes

The advent of commercially available quantum computers has marked the beginning of quantum computing as a reality. Both quantum gate and annealing computers have been released by major computer hardware companies. In this work, the D-Wave 2XTM quantum annealing computer housed at the NASA Advanced Systems computational facility is investigated to accelerate Machine Learning (ML) for image registration. NASA collects large amounts of images over the globe remotely using space-based monitoring. Images of a fixed areas of the land surface are taken over time. Due to the orbit of the sensors, the viewing angles deviate slightly, and it is necessary to align or register the images precisely to create image time series over the land surface. Unaligned images can lead to substantial analysis errors. These time-series are then used in modeling Earth Systems models such as hydrological, weather, and carbon monitoring models. In this work, we consider the Moderate Resolution Image Spectrometer (MODIS) data collected by the NASA's terra satellite. Artificial Neural Networks (ANNs) is a natural fit for ML modelling of images. Several successes have been reported using machine learning related to image processing. We investigate the use of ML to register MODIS images. ANNs are investigated in combination with a Restricted Boltzmann Machines (RBM) as an auto-encoder. We will present results showing the accuracy and efficiency of this approach.The D-Wave 2XTM quantum annealer samples the ground-state wave-function of a spin-Ising systems with quadratic interactions between qubits and a Chimera connectivity. The system sits in a ~15 mK thermal bath. One can think of the system as being placed in the ground state initially and subject to thermal excitations governed by Boltzmann statistics. If this is assumed true, one can use the statistics from the D-Wave 2XTM to train RBMs. Generating statistics for training Boltzmann machines is an NP-hard problem and constitutes the largest compute cost. We investigate the use of the D-Wave 2XTM to accelerate the training of the RBMs in our ANNs and report on the results.

Kouatchou, Jules↗

Multitaper Spectral Analysis and Wavelet Denoising Applied to Helioseismic Data

Estimates of solar normal mode frequencies from helioseismic observations can be improved by using Multitaper Spectral Analysis (MTSA) to estimate spectra from the time series, then using wavelet denoising of the log spectra. MTSA leads to a power spectrum estimate with reduced variance and better leakage properties than the conventional periodogram. Under the assumption of stationarity and mild regularity conditions, the log multitaper spectrum has a statistical distribution that is approximately Gaussian, so wavelet denoising is asymptotically an optimal method to reduce the noise in the estimated spectra. We find that a single m-upsilon spectrum benefits greatly from MTSA followed by wavelet denoising, and that wavelet denoising by itself can be used to improve m-averaged spectra. We compare estimates using two different 5-taper estimates (Stepian and sine tapers) and the periodogram estimate, for GONG time series at selected angular degrees l. We compare those three spectra with and without wavelet-denoising, both visually, and in terms of the mode parameters estimated from the pre-processed spectra using the GONG peak-fitting algorithm. The two multitaper estimates give equivalent results. The number of modes fitted well by the GONG algorithm is 20% to 60% larger (depending on l and the temporal frequency) when applied to the multitaper estimates than when applied to the periodogram. The estimated mode parameters (frequency, amplitude and width) are comparable for the three power spectrum estimates, except for modes with very small mode widths (a few frequency bins), where the multitaper spectra broadened the modest compared with the periodogram. We tested the influence of the number of tapers used and found that narrow modes at low n values are broadened to the extent that they can no longer be fit if the number of tapers is too large. For helioseismic time series of this length and temporal resolution, the optimal number of tapers is less than 10.

Komm, R. W.↗