Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Gaussian process fitting”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Deep Learning of Dark Energy Spectroscopic Instrument Mock Spectra to Find Damped Lyα Systems

We have updated and applied a convolutional neural network (CNN) machine-learning model to discover and characterize damped Ly α systems (DLAs) based on Dark Energy Spectroscopic Instrument (DESI) mock spectra. We have optimized the training process and constructed a CNN model that yields a DLA classification accuracy above 99% for spectra that have signal-to-noise ratios (S/N) above 5 per pixel. The classification accuracy is the rate of correct classifications. This accuracy remains above 97% for lower S/N ≈1 spectra. This CNN model provides estimations for redshift and H i column density with standard deviations of 0.002 and 0.17 dex for spectra with S/N above 3 pixel -1 . Also, this DLA finder is able to identify overlapping DLAs and sub-DLAs. Further, the impact of different DLA catalogs on the measurement of baryon acoustic oscillations (BAO) is investigated. The cosmological fitting parameter result for BAO has less than 0.61% difference compared to analysis of the mock results with perfect knowledge of DLAs. This difference is lower than the statistical error for the first year estimated from the mock spectra: above 1.7%. We also compared the performances of the CNN and Gaussian Process (GP) models. Our improved CNN model has moderately 14% higher purity and 7% higher completeness than an older version of the GP code, for S/N > 3. Both codes provide good DLA redshift estimates, but the GP produces a better column density estimate by 24% less standard deviation. A credible DLA catalog for the DESI main survey can be provided by combining these two algorithms.

79 ASTRONOMY AND ASTROPHYSICS↗

Reducing Ground-based Astrometric Errors with Gaia and Gaussian Processes

Stochastic field distortions caused by atmospheric turbulence are a fundamental limitation to the astrometric accuracy of ground-based imaging. This distortion field is measurable at the locations of stars with accurate positions provided by the Gaia DR2 catalog; we develop the use of Gaussian process regression (GPR) to interpolate the distortion field to arbitrary locations in each exposure. We introduce an extension to standard GPR techniques that exploits the knowledge that the 2D distortion field is curl-free. Applied to several hundred 90 s exposures from the Dark Energy Survey as a test bed, we find that the GPR correction reduces the variance of the turbulent astrometric distortions ≈12× , on average, with better performance in denser regions of the Gaia catalog. The rms per-coordinate distortion in the riz bands is typically ≈7 mas before any correction and ≈2 mas after application of the GPR model. The GPR astrometric corrections are validated by the observation that their use reduces, from 10 to 5 mas rms, the residuals to an orbit fit to riz-band observations over 5 yr of the r = 18.5 trans-Neptunian object Eris. We also propose a GPR method, not yet implemented, for simultaneously estimating the turbulence fields and the 5D stellar solutions in a stack of overlapping exposures, which should yield further turbulence reductions in future deep surveys.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Model selection and signal extraction using Gaussian Process regression

We present a novel computational approach for extracting localized signals from smooth background distributions. We focus on datasets that can be naturally presented as binned integer counts, demonstrating our procedure on the CERN open dataset with the Higgs boson signature, from the ATLAS collaboration at the Large Hadron Collider. Our approach is based on Gaussian Process (GP) regression — a powerful and flexible machine learning technique which has allowed us to model the background without specifying its functional form explicitly and separately measure the background and signal contributions in a robust and reproducible manner. Unlike functional fits, our GP-regression-based approach does not need to be constantly updated as more data becomes available. We discuss how to select the GP kernel type, considering trade-offs between kernel complexity and its ability to capture the features of the background distribution. We show that our GP framework can be used to detect the Higgs boson resonance in the data with more statistical significance than a polynomial fit specifically tailored to the dataset. Finally, we use Markov Chain Monte Carlo (MCMC) sampling to confirm the statistical significance of the extracted Higgs signature.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Stochastic Approach to Reconstruct Gamma-Ray-burst Light Curves

Gamma-ray bursts (GRBs), as they are observed at high redshift ( z = 9.4), are vital to cosmological studies and investigating Population III stars. To tackle these studies, we need correlations among relevant GRB variables with the requirement of small uncertainties on their variables. Thus, we must have good coverage of GRB light curves (LCs). However, gaps in the LC hinder the precise determination of GRB properties and are often unavoidable. Therefore, extensive categorization of GRB LCs remains a hurdle. We address LC gaps using a stochastic reconstruction, wherein we fit two preexisting models (the Willingale model; W07; and a broken power law; BPL) to the observed LC, then use the distribution of flux residuals from the original data to generate data to fill in the temporal gaps. We also demonstrate a model-independent LC reconstruction via Gaussian processes. At 10% noise, the uncertainty of the end time of the plateau, its correspondent flux, and the temporal decay index after the plateau decreases by 33.3%, 35.03%, and 43.32% on average for the W07, and by 33.3%, 30.78%, 43.9% for the BPL, respectively. The uncertainty of the slope of the plateau decreases by 14.76% in the BPL. After using the Gaussian process technique, we see similar trends of a decrease in uncertainty for all model parameters for both the W07 and BPL models. These improvements are essential for the application of GRBs as standard candles in cosmology, for the investigation of theoretical models, and for inferring the redshift of GRBs with future machine-learning analyses.

79 ASTRONOMY AND ASTROPHYSICS↗

O’Hare Airport Short-Term Ground Transportation Modal Demand Forecast Using Gaussian Processes

Here, the principal objective of this study is to analyze the spatial and temporal variation of ground transportation airport demand and provide demand forecast to inform planning capability and explore alternatives for investments to accommodate airport growth. Because of its good adaptability and strong generalization ability for dealing with high-dimensional input, small-sample, and nonlinear spatial data, Gaussian process (GP) regression is used to provide forecast estimates using data from transportation network company (TNC) trips and urban rail passengers at Chicago's O'Hare International Airport. TNC airport trips differ significantly, with three times more distance, more than twice the travel time, and half of the share requests compared with nonairport trips. This highlights the need for separate demand models. Hourly analysis of the rail service indicates that this is likely heavily used by airport workers, whereas TNC services focus on travelers because of variations in the peak demand hours. Heteroscedastic GP regression is implemented because of differences in trip variance between night and day hours. Estimates are given for weekdays and weekend trips, and the 95% confidence intervals are calculated. The introduction of flight schedule information into the models shows marginal improvements in their performance. However, fitting a GP regression becomes computationally expensive with increased sample size and the introduction of spatial components. Transportation planners and policymakers can use the results and methods implemented in this study to optimize transportation assets and provide long-range simulations of the current and future conditions in the area.

42 ENGINEERING↗

TEMImageNet training library and AtomSegNet deep-learning models for high-precision atom segmentation, localization, denoising, and deblurring of atomic-resolution images

Abstract Atom segmentation and localization, noise reduction and deblurring of atomic-resolution scanning transmission electron microscopy (STEM) images with high precision and robustness is a challenging task. Although several conventional algorithms, such has thresholding, edge detection and clustering, can achieve reasonable performance in some predefined sceneries, they tend to fail when interferences from the background are strong and unpredictable. Particularly, for atomic-resolution STEM images, so far there is no well-established algorithm that is robust enough to segment or detect all atomic columns when there is large thickness variation in a recorded image. Herein, we report the development of a training library and a deep learning method that can perform robust and precise atom segmentation, localization, denoising, and super-resolution processing of experimental images. Despite using simulated images as training datasets, the deep-learning model can self-adapt to experimental STEM images and shows outstanding performance in atom detection and localization in challenging contrast conditions and the precision consistently outperforms the state-of-the-art two-dimensional Gaussian fit method. Taking a step further, we have deployed our deep-learning models to a desktop app with a graphical user interface and the app is free and open-source. We have also built a TEM ImageNet project website for easy browsing and downloading of the training data.

25 ENERGY STORAGE↗

A general Bayesian algorithm for the autonomous alignment of beamlines

Autonomous methods to align beamlines can decrease the amount of time spent on diagnostics, and also uncover better global optima leading to better beam quality. The alignment of these beamlines is a high-dimensional expensive-to-sample optimization problem involving the simultaneous treatment of many optical elements with correlated and nonlinear dynamics. Bayesian optimization is a strategy of efficient global optimization that has proved successful in similar regimes in a wide variety of beamline alignment applications, though it has typically been implemented for particular beamlines and optimization tasks. In this paper, we present a basic formulation of Bayesian inference and Gaussian process models as they relate to multi-objective Bayesian optimization, as well as the practical challenges presented by beamline alignment. We show that the same general implementation of Bayesian optimization with special consideration for beamline alignment can quickly learn the dynamics of particular beamlines in an online fashion through hyperparameter fitting with no prior information. We present the implementation of a concise software framework for beamline alignment and test it on four different optimization problems for experiments on X-ray beamlines at the National Synchrotron Light Source II and the Advanced Light Source, and an electron beam at the Accelerator Test Facility, along with benchmarking on a simulated digital twin. We discuss new applications of the framework, and the potential for a unified approach to beamline alignment at synchrotron facilities.

47 OTHER INSTRUMENTATION↗

Dynamical Mass Estimates of the β Pictoris Planetary System through Gaussian Process Stellar Activity Modeling

Nearly 15 yr of radial velocity (RV) monitoring and direct imaging enabled the detection of two giant planets orbiting the young, nearby star β Pictoris. The δ Scuti pulsations of the star, which overwhelm planetary signals, need to be carefully suppressed. In this work, we independently revisit the analysis of the RV data following a different approach than available in the literature to model the activity of the star. We show that a Gaussian process (GP) with a stochastically driven damped harmonic oscillator kernel can model the δ Scuti pulsations. It provides similar results to parametric models but with a simpler framework, using only three hyperparameters. It also enables us to model poorly sampled RV data that were excluded from previous analyses, hence extending the RV baseline by nearly five years. Altogether, the orbit and mass of both planets can be constrained from RV only, which was not possible with the parametric modeling. To characterize the system more accurately, we also perform a joint fit of all available relative astrometry and RV data. Our orbital solutions for β Pic b favor a low eccentricity of 0.029$_{−0.024}^{+0.061}$ and a relatively short period of 21.1$_{−0.8}^{+2.0}$ yr. The orbit of β Pic c is eccentric with 0.206$_{−0.063}^{+0.074}$ with a period of 3.36 ± 0.03 yr. We find model-independent masses of 11.7 ± 1.4 and 8.5 ± 0.5 M Jup for β Pic b and c, respectively, assuming coplanarity. The mass of β Pic b is consistent with the hottest start evolutionary models, at an age of 25 ± 3 Myr. A direct detection of β Pic c would provide a second calibration measurement in a coeval system.

79 ASTRONOMY AND ASTROPHYSICS↗

Lithium-Ion Battery Diagnostics Using Electrochemical Impedance via Machine-Learning

Diagnosing battery states such as health, state-of-charge, or temperature is crucial for ensuring the safety and reliability of electrochemical energy storage systems. While some states, such as temperature, may be measured using cheap sensors, accurate diagnosis of battery health metrics usually requires time-consuming performance measurements, making them infeasible for use in real-world operation. These health metrics can be measured during lab-testing and then estimated on-line using predictive life models or via state observer algorithms such as Kalman filters, but these predictive methods should be supplemented by actual measurement of battery health whenever possible to ensure reliability. Rapid measurement of battery health may be done by various types of fast diagnostic techniques such as electrochemical impedance spectroscopy (EIS), which can be performed in only a few minutes and require only a fraction of the energy and power needed for a full charge and discharge measurement. But there is a substantial challenge for estimating battery health using EIS data, as EIS is sensitive to cell temperature, state-of-charge, current, and resting time in addition to health. Thus, utilizing EIS data to predict battery capacity requires correcting for all these additional variables, a task that is extremely difficult to handle analytically. This talk utilizes machine-learning methods to estimate the effectiveness of battery capacity prediction from EIS data, leveraging a data set of hundreds of EIS measurements recorded at varying temperature and state-of-charge throughout a 500-day aging study of 32 commercial, large-format NMC-Graphite lithium-ion batteries. Using EIS as input to machine-learning models is complicated by the nonlinear response of impedance to battery health, temperature, and state-of-charge, as well as the collinearity between the impedance response at neighboring frequencies, which can easily lead to overfit models. To train robust models, features from EIS data need to be extracted from the data or some subset of critical frequencies selected. Many approaches for extracting and selecting features from EIS data from electrochemical analysis and machine-learning fields were identified for analysis: using the entire raw spectra; selection of one, two, or many frequencies from the entire spectra; selecting interesting points from the EIS measurement using domain knowledge; fitting EIS with an equivalent-circuit model; calculating statistics on the raw impedance values; and reducing the dimensionality of the data using unsupervised linear (principal component analysis) and non-linear (uniform manifold approximation and projection) methods. These approaches were rigorously compared using a machine-learning pipeline approach, training linear, Gaussian process, and random forest regression models and quantifying performance using cross-validation as well as a held-out test set. An artificial neural network model trained on the raw spectra was also tested. Promising pipelines were fine-tuned via Bayesian hyperparameter optimization using cross-validation loss and training with class-specific weights to counter data set imbalance. The most reliable method for utilizing impedance in this work was the selection of two optimal frequencies through an exhaustive search, resulting in about 2% mean absolute error on test data for both Gaussian process and random forest model architectures. Interrogation of a variety of models reveals critical frequencies of 100 Hz and 103 Hz for this data set, though the optimal set of frequencies is not necessarily intuitive, i.e., the best performing models are not simply those that use impedance at frequencies that have the highest correlation to the relative discharge capacity. The best performing model is an ensemble model, which is able to predict battery capacity with 1.9% mean absolute error for unseen cells using impedance recorded at a variety of temperatures and states-of-charge.

battery↗

Apodization Specific Fitting for Improved Resolution, Charge Measurement, and Data Analysis Speed in Charge Detection Mass Spectrometry

Short-time Fourier transforms with short segment lengths are typically used to analyze single ion charge detection mass spectrometry (CDMS) data either to overcome effects of frequency shifts that may occur during the trapping period or to more precisely determine the time at which an ion changes mass or charge, or enters an unstable orbit. The short segment lengths can lead to scalloping loss unless a large number of zero-fills are used, making computational time a significant factor in real-time analysis of data. Apodization specific fitting leads to a 9-fold reduction in computation time compared to zero-filling to a similar extent of accuracy. This makes possible real-time data analysis using a standard desktop computer. Rectangular apodization leads to higher resolution than the more commonly used Gaussian or Hann apodization and makes it possible to separate ions with similar frequencies, a significant advantage for experiments in which the masses of many individual ions are measured simultaneously. Equally important is a >20% increase in S/N obtained with rectangular apodization compared to Gaussian or Hann, which directly translates to a corresponding improvement in accuracy of both charge measurements and ion energy measurements that rely on the amplitudes of the fundamental and harmonic frequencies. Finally, combined with computing the fast Fourier transform in a lower-level language, this fitting procedure eliminates computational barriers and should enable real-time processing of CDMS data on a laptop computer.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Detection Limits of Low-mass, Long-period Exoplanets Using Gaussian Processes Applied to HARPS-N Solar Radial Velocities

Radial velocity (RV) searches for Earth-mass exoplanets in the habitable zone around Sun-like stars are limited by the effects of stellar variability on the host star. In particular, suppression of convective blueshift and brightness inhomogeneities due to photospheric faculae/plage and starspots are the dominant contribution to the variability of such stellar RVs. Gaussian process (GP) regression is a powerful tool for statistically modeling these quasi-periodic variations. We investigate the limits of this technique using 800 days of RVs from the solar telescope on the High Accuracy Radial velocity Planet Searcher for the Northern hemisphere (HARPS-N) spectrograph. These data provide a well-sampled time series of stellar RV variations. Into this data set, we inject Keplerian signals with periods between 100 and 500 days and amplitudes between 0.6 and 2.4 m s{sup −1}. We use GP regression to fit the resulting RVs and determine the statistical significance of recovered periods and amplitudes. We then generate synthetic RVs with the same covariance properties as the solar data to determine a lower bound on the observational baseline necessary to detect low-mass planets in Venus-like orbits around a Sun-like star. Our simulations show that discovering planets with a larger mass (∼0.5 m s{sup −1}) using current-generation spectrographs and GP regression will require more than 12 yr of densely sampled RV observations. Furthermore, even with a perfect model of stellar variability, discovering a true exo-Venus (∼0.1 m s{sup −1}) with current instruments would take over 15 yr. Therefore, next-generation spectrographs and better models of stellar variability are required for detection of such planets.

47 OTHER INSTRUMENTATION↗

How “hot” are hotspots: Statistically localizing the high-activity areas on soil and rhizosphere images

The topic of microbial hotspots in soil requires not only visualizing their spatial distribution and biochemical analyses, but also statistical approaches to identify these hotspots and separate them from the surrounding activities (background). We hypothesized that each hotspot type (e.g. enzyme activities in the rhizosphere, root exudation, localization of herbicide accumulation) is a result of local process driven by biotic and/or abiotic factors, and the process rates in the hotspots are much faster than those in the soil background. We further hypothesized that the background and hotspot activities in soil belong to different statistical distributions. Consequently, hotspot determination should be based on statistical separation of activities significantly higher than the background. We analyzed for the statistical distributions of grey values on three groups of published images: 1) 14 C images of carbon input by roots into the rhizosphere, 2) 14 C glyphosate accumulation in the plant, and 3) zymogram of leucine aminopeptidase activity in rooted soil. The two Gaussian distributions were fit (the first representing the background, the second the hotspots) to the distribution of grey values in the images, the parameters (means and standard deviations, SD) of the fitted distributions were calculated, and the background was removed. Thus, we identified hotspots as areas outside of the Mean+2SD image intensity (corresponding to the upper ~ 2.5% of activity, being over 97.5% of background values) and finally, visualized images of solely hotspot locations. Finally, these results were compared with previously used decisions on hotspot intensity thresholding (i.e. Top-25% and 17 standard thresholding approaches in ImageJ) and discussed the advantages of the Mean+2SD as well as Mean+3SD approaches. These advantages include: i) simple unification of the thresholding approach for several imaging methods with various principles of activity distribution, ii) identification of hotspots with various activity levels, iii) analysis of “time-specific” hotspots in temporal sequences of images. Compared with 17 standard thresholding methods, we concluded that objectively elucidating and separating the hotspots should be based on statistical distribution analysis, e.g. using the Mean+2SD or Mean+3SD approaches. Furthermore, this simple Mean+2SD approach delivered suitable results for three groups of images and so, helps to understand the processes responsible for the highest activities and elucidate hotspots.

59 BASIC BIOLOGICAL SCIENCES↗

A Machine Learning based Approach of Estimating Equivalent Circuit Model Parameters at Different SoCs of Li-ion Batteries from Voltage Relaxation

Abstract: In this study, an approach of estimating the equivalent circuit model (ECM) parameters for Li-ion batteries (LIBs) is proposed based on the voltage value at different intervals while relaxing the LIB after discharge. The typical approach for estimating ECM parameters of a LIB is to conduct electrochemical impedance spectroscopy (EIS) measurements at different frequencies and fit them to a predefined circuit model, which requires additional measuring arrangements and specialized devices. The proposed methodology utilizes four different voltages at 0s, 60s, 360s, and 1800s alongside the specific state of charge (SoC) value for a specific constant discharge current value of ~1C until the relaxation stage to train and evaluate three regression-based machine learning models— Support Vector Regression (SVR), Extreme Gradient Boosting (XGBoost), and Gaussian Process Regression (GPR)—for estimating the ECM parameters of the selected model. Bayesian optimization is employed for hyperparameter tuning to achieve optimal performance for all the regressor models, among which, the GPR provided the best performance with the root-mean-squared error (RMSE) of less than 4x10-4 on average for the resistive components and less than 0.27 for capacitive components with excellent R2 scores. The simplicity of the approach enables it to eliminate the need for sophisticated measuring equipment and computation power.

Sagar, Md. Samiul [The University of Alabama (UA)]↗

Physics-Based Machine Learning Methods for U-235 Forensics Signatures

Signatures of low-intensity U-235 sources have been recently studied by utilizing a variety of machine learning (ML) classifiers using features derived from gamma spectral measurements collectedunder structured campaigns. Several ML classifiers, such as ensemble of tress and classification trees, revealed misleadingly-optimistic training error due to over-fitting, and furthermore,their performance is not directly relatable to the physical properties due to their data-driven, opaque designs. We present a regression-based ML method that first estimates the inverse distanceto the source and then utilizes a threshold to infer its presence, by representing the background as a source located at an infinite distance. For the inverse distance estimation, we study the ensembleof trees and Gaussian process regression methods, and a hyper parameter auto-tuning and selection method that employs five regression estimators. These methods avoid the over-fittingobserved in several ML classifiers, while providing the classification error nearly comparable to them based on independent test data. Their error is directly related to estimates of the inversephysical distance to source, and the precision of error determines the seperability property that determines the false alarm and missed detection rates. The property of monotonic decrease of thesource strength with increasing detector distance combined with Poisson distribution of measurements is utilized to analytically validate these methods by deriving the generalization equations ofunderlying regression methods.

Rao, Nageswara↗

A machine learning pipeline for membrane segmentation of cryo-electron tomograms

We describe how to use several machine learning techniques organized in a learning pipeline to segment and identify cell membrane structures from cryo electron tomograms. These tomograms are difficult to analyze with traditional segmentation tools. The learning pipeline in our approach starts from supervised learning via a special convolutional neural network trained with simulated data. It continues with semi-supervised reinforcement learning and/or a region merging technique that tries to piece together disconnected components belonging to the same membrane structure. A parametric or non-parametric fitting procedure is then used to enhance the segmentation results and quantify uncertainties in the fitting. Domain knowledge is used in generating the training data for the neural network and in guiding the fitting procedure through the use of appropriately chosen priors and constraints. We demonstrate that the approach proposed here works well for extracting membrane surfaces in two real tomogram datasets.

97 MATHEMATICS AND COMPUTING↗

Multiple emission components in the Cygnus cocoon detected from Fermi -LAT observations

Star-forming regions may play an important role in the life cycle of Galactic cosmic rays (CRs), notably as home to specific acceleration mechanisms and transport conditions. Gamma-ray observations of Cygnus X have revealed the presence of an excess of hard-spectrum gamma-ray emission, possibly related to a cocoon of freshly accelerated particles. We seek an improved description of the gamma-ray emission from the cocoon using ~13 yr of observations with the Fermi-Large Area Telescope (LAT) and use it to further constrain the processes and objects responsible for the young CR population. We developed an emission model for a large region of interest, including a description of interstellar emission from the background population of CRs and recent models for other gamma-ray sources in the field. Thus, we performed an improved spectro-morphological characterisation of the residual emission including the cocoon. The best-fit model for the cocoon includes two main emission components: an extended component FCES G78.74+1.56, described by a 2D Gaussian of extension r 68 = 4.4° ± 0.1° -0.1° +0.1° and a smooth broken power law spectrum with spectral indices 1.67 ± 0.05 -0.01 +0.02 and 2.12 ± 0.02 -0.01 +0.00 below and above 3.0 ± 0.6 -0.2 +0.0 GeV, respectively; and a central component FCES G80.00+0.50, traced by the distribution of ionised gas within the borders of the photo-dissociation regions and with a power law spectrum of index 2.19 ± 0.03 -0.01 +0.00 that is significantly different from the spectrum of FCES G78.74+1.56. An additional extended emission component FCES G78.83+3.57, located on the edge of the central cavities in Cygnus X and with a spectrum compatible with that of FCES G80.00+0.50, is likely related to the cocoon. For the two brightest components FCES G80.00+0.50 and FCES G78.74+1.56, spectra and radial-azimuthal profiles of the emission can be accounted for in a diffusion-loss framework involving one single population of non-thermal particles with a flat injection spectrum. Particles span the full extent of FCES G78.74+1.56 as a result of diffusion from a central source, and give rise to source FCES G80.00+0.50 by interacting with ionised gas in the innermost region. For this simple diffusion-loss model, viable setups can be very different in terms of energetics, transport conditions, and timescales involved, and both hadronic and leptonic scenarios are possible. The solutions range from long-lasting particle acceleration, possibly in prominent star clusters such as Cyg OB2 and NGC 6910, to a more recent and short-lived release of particles within the last 10–100 kyr, likely from a supernova remnant. The observables extracted from our analysis can be used to perform detailed comparisons with advanced models of particle acceleration and transport in star-forming regions.

79 ASTRONOMY AND ASTROPHYSICS↗

Spatiotemporal Downscaling Model for Solar Irradiance Forecast Using Nearest-Neighbor Random Forest and Gaussian Process

Accurate solar photovoltaic (PV) capacity estimation requires high-resolution, site-specific solar irradiance data to account for localized variability. However, global datasets, such as the National Solar Radiation Database (NSRDB), provide regional averages that fail to capture the fine-scale fluctuations critical for large-scale grid integration. This limitation is particularly relevant in the context of increasing distributed energy resources (DERs) penetration, such as rooftop PV. Additionally, it is critical to the implementation of the U.S. Federal Energy Regulatory Commission (FERC) Order 2222, which facilitates DER participation in U.S. bulk power markets. To address this challenge, this study evaluates Nearest-Neighbor Random Forest (NNRF) and Nearest-Neighbor Gaussian Process (NNGP) models for spatiotemporal downscaling of global solar irradiance data. By leveraging historical irradiance and meteorological data, these models incorporate spatial, temporal, and feature-based correlations to enhance local irradiance predictions. The NNRF model, a machine-learning approach, prioritizes computational efficiency and predictive accuracy, while the NNGP model offers a level of interpretability and prediction uncertainty by numerically quantifying correlations and dependencies in the data. Model validation was conducted using day-ahead predictions. The results showed that the average Goodness of Fit (GoF) of the NNRF model of 90.61% across all eight sites outperformed the GoF of the NNGP of 85.88%. Additionally, the computational speed of NNRF was 2.5 times faster than the NNGP. Finally, the NNGP displayed polynomial scaling while the NNRF scaled linearly with increasing number of nearest neighbors. Additional validation of the model on five sites in Puerto Rico further confirmed the superiority of the NNRF model over the NNGP model. These findings highlight the robustness and computational efficiency of NNRF for large-scale solar irradiance downscaling, making it a strong candidate for improving PV capacity estimation and real-time electricity market integration for DERs.

Asiedu, Shadrack (ORCID:0009000646004826)↗

Emulating the Lyman-Alpha forest 1D power spectrum from cosmological simulations: new models and constraints from the eBOSS measurement

We present the Lyssa suite of high-resolution cosmological simulations of the Lyman-α forest designed for cosmological analyses. These 18 simulations have been run using the Nyx code with 40963 hydrodynamical cells in a 120 Mpc (∼ 81 Mpc/h) comoving box and individually provide sub-percent level convergence of the Lyman-α forest 1d flux power spectrum. We build a Gaussian process emulator for the Lyssa simulations in the lym1d likelihood framework to interpolate the power spectrum at arbitrary parameter values. We validate this emulator based on leave-one-out tests and based on the parameter constraints for simulations outside of the training set. We also perform comparisons with a previous emulator, showing a percent level accuracy and a good recovery of the expected cosmological parameters. Using this emulator we derive constraints on the linear matter power spectrum amplitude and slope parameters A Lyα and n Lyα . While the best-fit Planck ΛCDM model has A Lyα = 8.79 and n Lyα = -2.363, from DR14 eBOSS data we find that A Lyα < 7.6 (95% CI) and n Lyα = -2.369 ± 0.008. The low value of A Lyα , in tension with Planck, is driven by the correlation of this parameter with the mean transmission of the Lyman-α forest. This tension disappears when imposing a well-motivated external prior on this mean transmission, in which case we find A Lyα = 9.8 ± 1.1 in accordance with Planck.

Walther, Michael↗