Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Adaptive sampling for accelerating neutron diffraction-based strain mapping *

Abstract Neutron diffraction is a useful technique for mapping residual strains in dense metal objects. The technique works by placing an object in the path of a neutron beam, measuring the diffracted signals and inferring the local lattice strain values from the measurement. In order to map the strains across the entire object, the object is stepped one position at a time in the path of the neutron beam, typically in raster order, and at each position a strain value is estimated. Typical dwell times at neutron diffraction instruments result in an overall measurement that can take several hours to map an object that is several tens of centimeters in each dimension at a resolution of a few millimeters, during which the end users do not have an estimate of the global strain features and are at risk of incomplete information in case of instruments outages. In this paper, we propose an object adaptive sampling strategy to measure the significant points first. We start with a small initial uniform set of measurement points across the object to be mapped, compute the strain in those positions and use a machine learning technique to predict the next position to measure in the object. Specifically, we use a Bayesian optimization based on a Gaussian process regression method to infer the underlying strain field from a sparse set of measurements and predict the next most informative positions to measure based on estimates of the mean and variance in the strain fields estimated from the previously measured points. We demonstrate our real-time measure-infer-predict workflow on additively manufactured steel parts—demonstrating that we can get an accurate strain estimate even with 30%–40% of the typical number of measurements—leading the path to faster strain mapping with useful real-time feedback. We emphasize that the proposed method is general and can be used for fast mapping of other material properties such as phase fractions from time-consuming point-wise neutron measurements.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Estimation of Surface Air Temperature Over Central and Eastern Eurasia from MODIS Land Surface Temperature

Surface air temperature (T(sub a)) is a critical variable in the energy and water cycle of the Earth.atmosphere system and is a key input element for hydrology and land surface models. This is a preliminary study to evaluate estimation of T(sub a) from satellite remotely sensed land surface temperature (T(sub s)) by using MODIS-Terra data over two Eurasia regions: northern China and fUSSR. High correlations are observed in both regions between station-measured T(sub a) and MODIS T(sub s). The relationships between the maximum T(sub a) and daytime T(sub s) depend significantly on land cover types, but the minimum T(sub a) and nighttime T(sub s) have little dependence on the land cover types. The largest difference between maximum T(sub a) and daytime T(sub s) appears over the barren and sparsely vegetated area during the summer time. Using a linear regression method, the daily maximum T(sub a) were estimated from 1 km resolution MODIS T(sub s) under clear-sky conditions with coefficients calculated based on land cover types, while the minimum T(sub a) were estimated without considering land cover types. The uncertainty, mean absolute error (MAE), of the estimated maximum T(sub a) varies from 2.4 C over closed shrublands to 3.2 C over grasslands, and the MAE of the estimated minimum Ta is about 3.0 C.

Shen, Suhung↗

Development of the Ames Global Hyperspectral Synthetic Dataset

This study develops the surface BRDF (bidirectional reflectance distribution function) product of the Ames Global Hyperspectral Synthetic Dataset (AGHSD), based on the corresponding MODIS products, to support the NASA Surface Biology and Geology mission development. A main challenge in deriving a hyperspectral dataset from the multi-band satellite products is how to identify a succinct yet robust algorithm that allow us to infer BRDF at unobserved wavelengths based on the few observed bands. Using the theories of radiative transfer in vegetation canopies, we arrive at a simple equation that accurately approximates hyperspectral surface BRDF as the weighted sum of components from the soil and the vegetation. Each of the components is modeled by the product of the spectrally-dependent optical properties of a surface element (the spectra of the soil surface reflectance, the leaf single albedo, or the canopy scattering coefficient) and a spectrally-independent bidirectional scattering function. The optical properties of the soil and the vegetation can be obtained from existing spectral libraries or model simulations. The bidirectional scattering functions are represented by the Ross-Thick-Li-Sparse BRDF model, where the linear coefficients are estimated with regression analysis from the multi-band MODIS data. We validate the algorithm with simulations by Monte Carlo Ray Tracing model experiments, and the results are highly consistent with the theoretic derivation. We apply the algorithm to generate the AGHSD BRDF product at 1km and 8-day resolutions for the year of 2019. The results are biogeochemically and physically coherent and consistent, and thus serve the goal to support the science and application development of the SBG community.

Hyperspectral↗

Bayesian Gaussian process inference for neutron spin echo measurement

Neutron spin echo (NSE) spectroscopy provides unique access to microscopic dynamics, but its application is often constrained by low neutron flux, long acquisition times, and significant noise. Here, we present a Bayesian inference approach based on Gaussian process regression (GPR) to reconstruct high-quality spin echo signals from sparse and noisy data by exploiting correlations in reciprocal space. Benchmarks on synthetic datasets and validation with experimental NSE measurements of dendrimers show that GPR suppresses noise, interpolates missing intensity values, and accommodates irregular observations. The method improves accuracy, shortens acquisition times, and enables high-throughput and real-time studies. Beyond NSE, the framework is broadly applicable to other low signal-to-noise ratio scattering techniques, thereby extending the scope of neutron spectroscopy.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

Identifying Entangled Physics Relationships Through Sparse Matrix Decomposition to Inform Plasma Fusion Design

We report a sustainable burn platform through inertial confinement fusion (ICF) has been an ongoing challenge for over 50 years. Mitigating engineering limitations and improving the current design involves an understanding of the complex coupling of physical processes. While sophisticated simulation codes are used to model ICF implosions, these tools contain necessary numerical approximation but miss physical processes that limit predictive capability. Identification of relationships between controllable design inputs to ICF experiments and measurable outcomes (e.g., neutron yield, neutron velocity, areal density) from performed experiments can help guide the future design of experiments and development of simulation codes, to potentially improve the accuracy of the computational models used to simulate ICF experiments. We use sparse matrix decomposition methods to identify clusters of a few related design variables. Sparse principal component analysis (SPCA) identifies groupings that are related to the physical origin of the variables (laser, hohlraum, and capsule). A variable importance analysis finds that in addition to variables highly correlated with neutron yield, such as picket power and laser energy, variables that represent a dramatic change of the ICF design, such as number of pulse steps, are also very important. The obtained sparse components are then used to train a random forest (RF) regression surrogate for predicting total yield. The RF performance on the training and testing data compares with the performance of the RF trained using all the design variables considered. This work is intended to inform design changes in future ICF experiments by augmenting the expert intuition and simulation results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Interhemispheric comparison of atmospheric circulation features as evaluated from NIMBUS satellite data

Findings are presented for IRIS data from NIMBUS 3 in mapping the global ozone distribution. The seasonal and regional variations of ozone, especially in the Southern Hemisphere, reveal features that were not evident from the sparse ground-based ozone observation network in this hemisphere. A regression analysis was undertaken for temperature and height fields on radiance data. Spectrum analyses of upper wind data from the North American section and Australia were completed.

Reiter, E. R.↗

A convergence metric for counting statistics in time-resolved small angle neutron scattering

Here, this work introduces a model-independent, dimensionless metric for predicting optimal measurement duration in time-resolved small-angle neutron scattering using early-time data. Built on a Gaussian process regression framework, the method reconstructs scattering profiles with quantified uncertainty, even from sparse or noisy measurements. Demonstrated on the EQ-SANS instrument at the Spallation Neutron Source, the approach generalizes to general SANS instruments with a two-dimensional detector. A key result is the discovery of a dimensionless convergence metric revealing a universal power-law scaling in profile evolution across soft matter systems. When time is normalized by a system-specific characteristic time t*, the variation in inferred profiles collapses onto a single curve with an exponent between −2 and −1. This trend emerges within the first ten time steps, enabling early prediction of measurement sufficiency. The method supports real-time experimental optimization and is especially valuable for maximizing efficiency in low-flux environments such as compact accelerator-based neutron sources.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

Sparse-Data Deep Learning Strategies for Radiographic Non-Destructive Testing

Radiography is an imaging technique used in a variety of applications, such as medical diagnosis, airport security, and nondestructive testing. We present a deep learning system for extracting information from radiographic images. We perform various prediction tasks using our system, including material classification and regression on the dimensions of a given object that is being radiographed. Our system is designed to address the sparse-data issue for radiographic nondestructive testing applications. It uses a radiographic simulation tool for synthetic data augmentation, and it uses transfer learning with a pre-trained convolutional neural network model. Using this system, our preliminary results indicate that the object geometry regression task saw an improvement of 70% in the R-squared value when using a multi-regime model. In addition, we increase the performance of the object material classification tasks by utilizing data from different imaging systems. In particular, using neutron imaging improved the material classification accuracy by 20% when compared to x-ray imaging.

convolutional neural networks↗

Gaussian Process Regression under Computational and Epistemic Misspecification

Gaussian process regression is a classical kernel method for function estimation and data interpolation. In large data applications, computational costs can be reduced using low-rank or sparse approximations of the kernel. This paper investigates the effect of such kernel approximations on the interpolation error. We introduce a unified framework to analyze Gaussian process regression under important classes of computational misspecification: Karhunen-Loève expansions that result in low-rank kernel approximations, multiscale wavelet expansions that induce sparsity in the covariance matrix, and finite element representations that induce sparsity in the precision matrix. Furthermore, our theory also accounts for epistemic misspecification in the choice of kernel parameters.

Gaussian process regression↗

Prediction and compression of lattice QCD data using machine learning algorithms on quantum annealer

We present regression and compression algorithms for lattice QCD data utilizing the efficient binary optimization ability of quantum annealers. In the regression algorithm, we encode the correlation between the input and output variables into a sparse coding machine learning algorithm. The trained correlation pattern is used to predict lattice QCD observables of unseen lattice configurations from other observables measured on the lattice. In the compression algorithm, we define a mapping from lattice QCD data of floating-point numbers to the binary coefficients that closely reconstruct the input data from a set of basis vectors. Since the reconstruction is not exact, the mapping defines a lossy compression, but, a reasonably small number of binary coefficients are able to reconstruct the input vector of lattice QCD data with the reconstruction error much smaller than the statistical fluctuation. In both applications, we use D-Wave quantum annealers to solve the NP-hard binary optimization problems of the machine learning algorithms.

79 ASTRONOMY AND ASTROPHYSICS↗

LAI inversion from optical reflectance using a neural network trained with a multiple scattering model

The inversion of the leaf area index (LAI) canopy parameter from optical spectral reflectance measurements is obtained using a backpropagation artificial neural network trained using input-output pairs generated by a multiple scattering reflectance model. The problem of LAI estimation over sparse canopies (LAI < 1.0) with varying soil reflectance backgrounds is particularly difficult. Standard multiple regression methods applied to canopies within a single homogeneous soil type yield good results but perform unacceptably when applied across soil boundaries, resulting in absolute percentage errors of >1000 percent for low LAI. Minimization methods applied to merit functions constructed from differences between measured reflectances and predicted reflectances using multiple-scattering models are unacceptably sensitive to a good initial guess for the desired parameter. In contrast, the neural network reported generally yields absolute percentage errors of <30 percent when weighting coefficients trained on one soil type were applied to predicted canopy reflectance at a different soil background.

Smith, James A.↗

Conditioned Simulation of Ground-Motion Time Series at Uninstrumented Sites Using Gaussian Process Regression

Ground-motion time series are essential input data in seismic analysis and performance assessment of the built environment. Because instruments to record free-field ground motions are generally sparse, methods are needed to estimate motions at locations with no available ground-motion recording instrumentation. In this study, given a set of observed motions, ground-motion time series at target sites are constructed using a Gaussian process regression (GPR) approach, which treats the real and imaginary parts of the Fourier spectrum as random Gaussian variables. Model training, verification, and applicability studies are carried out using the physics-based simulated ground motions of the 1906 Mw 7.9 San Francisco earthquake and Mw 7.0 Hayward fault scenario earthquake in northern California. Additionally, the method’s performance is further evaluated using the 2019 Mw 7.1 Ridgecrest earthquake ground motions recorded by the Community Seismic Network stations located in southern California. These evaluations indicate that the trained GPR model is able to adequately estimate the ground-motion time series for frequency ranges that are pertinent for most earthquake engineering applications. The trained GPR model exhibits proper performance in predicting the long-period content of the ground motions as well as directivity pulses.

58 GEOSCIENCES↗

Corrosion of rebar in concrete. Part III: Artificial Neural Network analysis of chloride threshold data

Active corrosion of carbon steel in reinforced concrete occurs when the chloride ion concentration exceeds the chloride threshold (CT). CT is affected by many physical and chemical parameters including primary variables (pH, corrosion potential, breakdown potential, and temperature) and secondary variables (cemenet composition, concreete porosity, and water/cement ratio). A Kohonen-self organized map (KSOM) and regression artificial neural network (ANN) coupled methodology was developed to find the missing values of independent variables in the sparse database and for quantitatively evaluating the effects of these variables on CT values that are expressed in %TotalCl/cem (or %FreeCl/cem), and [Cl - ]/[OH - ].

36 MATERIALS SCIENCE↗

Dynamic Ensemble Prediction of Cognitive Performance in Space

Astronauts are exposed to a unique set of stressors in spaceflight. Microgravity, isolation, confinement, and environmental and operational hazards: all of these can impact sleep, vigilant attention, and alertness, which are critical to mission success. In this paper, we seek to understand the most important predictors of alertness over the course of a space mission, using self-reported, cognitive, and environmental data collected from 24 astronauts on 6-month missions to the International Space Station (ISS). Alertness was repeatedly and objectively assessed on the ISS with a brief 3-minute Psychomotor Vigilance Test (PVT) that is highly sensitive to sleep deprivation. To relate PVT performance to time-varying and sparsely-measured environmental, operational, and psychological covariates, we propose a n ensemble prediction model comprising of linear mixed effects regression, random forest, and functional concurrent regression models. An extensive cross-validation procedure reveals that this ensemble outperforms any one of its components alone. We also discover that a participant’s past performance, reported fatigue and stress, and temperature and radiation exposure were among the most important variables associated with alertness. This method is broadly applicable to environmental studies where the main goal is accurate, individualized prediction involving a mixture of person-level traits and irregularly measured time series.

Danni Tu↗

A Bayesian framework for adsorption energy prediction on bimetallic alloy catalysts

Abstract For high-throughput screening of materials for heterogeneous catalysis, scaling relations provides an efficient scheme to estimate the chemisorption energies of hydrogenated species. However, conditioning on a single descriptor ignores the model uncertainty and leads to suboptimal prediction of the chemisorption energy. In this article, we extend the single descriptor linear scaling relation to a multi-descriptor linear regression models to leverage the correlation between adsorption energy of any two pair of adsorbates. With a large dataset, we use Bayesian Information Criteria (BIC) as the model evidence to select the best linear regression model. Furthermore, Gaussian Process Regression (GPR) based on the meaningful convolution of physical properties of the metal-adsorbate complex can be used to predict the baseline residual of the selected model. This integrated Bayesian model selection and Gaussian process regression, dubbed as residual learning, can achieve performance comparable to standard DFT error (0.1 eV) for most adsorbate system. For sparse and small datasets, we propose an ad hoc Bayesian Model Averaging (BMA) approach to make a robust prediction. With this Bayesian framework, we significantly reduce the model uncertainty and improve the prediction accuracy. The possibilities of the framework for high-throughput catalytic materials exploration in a realistic setting is illustrated using large and small sets of both dense and sparse simulated dataset generated from a public database of bimetallic alloys available in Catalysis-Hub.org.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Study on the Potential Applications of Satellite Data in Air Quality Monitoring and Forecasting

In this study we explore the potential applications of MODIS (Moderate Resolution Imaging Spectroradiometer) -like satellite sensors in air quality research for some Asian regions. The MODIS aerosol optical thickness (AOT), NCEP global reanalysis meteorological data, and daily surface PM(sub 10) concentrations over China and Thailand from 2001 to 2009 were analyzed using simple and multiple regression models. The AOT-PM(sub 10) correlation demonstrates substantial seasonal and regional difference, likely reflecting variations in aerosol composition and atmospheric conditions, Meteorological factors, particularly relative humidity, were found to influence the AOT-PM(sub 10) relationship. Their inclusion in regression models leads to more accurate assessment of PM(sub 10) from space borne observations. We further introduced a simple method for employing the satellite data to empirically forecast surface particulate pollution, In general, AOT from the previous day (day 0) is used as a predicator variable, along with the forecasted meteorology for the following day (day 1), to predict the PM(sub 10) level for day 1. The contribution of regional transport is represented by backward trajectories combined with AOT. This method was evaluated through PM(sub 10) hindcasts for 2008-2009, using ohservations from 2005 to 2007 as a training data set to obtain model coefficients. For five big Chinese cities, over 50% of the hindcasts have percentage error less than or equal to 30%. Similar performance was achieved for cities in northern Thailand. The MODIS AOT data are responsible for at least part of the demonstrated forecasting skill. This method can be easily adapted for other regions, but is probably most useful for those having sparse ground monitoring networks or no access to sophisticated deterministic models. We also highlight several existing issues, including some inherent to a regression-based approach as exemplified by a case study for Beijing, Further studies will be necessa1Y before satellite data can see more extensive applications in the operational air quality monitoring and forecasting.

Li, Can↗

Generative Ensemble Regression: Learning Particle Dynamics from Observations of Ensembles with Physics-informed Deep Generative Models

Here, we propose a new method for inferring the governing stochastic ordinary differential equations (SODEs) by observing particle ensembles at discrete and sparse time instants, i.e., multiple “snapshots.” Particle coordinates at a single time instant, possibly noisy or truncated, are recorded in each snapshot but are unpaired across the snapshots. By training a physics-informed generative model that generates “fake” sample paths, we aim to fit the observed particle ensemble distributions with a curve in the probability measure space, which is induced from the inferred particle dynamics. We employ different metrics to quantify the differences between distributions, e.g., the sliced Wasserstein distances and the adversarial losses in generative adversarial networks. We refer to this method as generative “ensemble-regression” (GER), in analogy to the classic “point-regression,” where we infer the dynamics by performing regression in the Euclidean space. We illustrate the GER by learning the drift and diffusion terms of particle ensembles governed by SODEs with Brownian motions and Lévy processes up to 100 dimensions. We also discuss how to treat cases with noisy or truncated observations. Apart from systems consisting of independent particles, we also tackle nonlocal interacting particle systems with unknown interaction potential parameters by constructing a physics-informed loss function. Finally, we investigate scenarios of paired observations and discuss how to reduce the dimensionality in such cases by proving a convergence theorem that provides theoretical support.

97 MATHEMATICS AND COMPUTING↗

Sparse Cholesky factorization for solving nonlinear PDEs via Gaussian processes

In recent years, there has been widespread adoption of machine learning-based approaches to automate the solving of partial differential equations (PDEs). Among these approaches, Gaussian processes (GPs) and kernel methods have garnered considerable interest due to their flexibility, robust theoretical guarantees, and close ties to traditional methods. They can transform the solving of general nonlinear PDEs into solving quadratic optimization problems with nonlinear, PDE-induced constraints. However, the complexity bottleneck lies in computing with dense kernel matrices obtained from pointwise evaluations of the covariance kernel, and its partial derivatives, a result of the PDE constraint and for which fast algorithms are scarce. The primary goal of this paper is to provide a near-linear complexity algorithm for working with such kernel matrices. We present a sparse Cholesky factorization algorithm for these matrices based on the near-sparsity of the Cholesky factor under a novel ordering of pointwise and derivative measurements. The near-sparsity is rigorously justified by directly connecting the factor to GP regression and exponential decay of basis functions in numerical homogenization. We then employ the Vecchia approximation of GPs, which is optimal in the Kullback-Leibler divergence, to compute the approximate factor. This enables us to compute ϵ-approximate inverse Cholesky factors of the kernel matrices with complexity O(N log d (N/ϵ)) in space and O(N log 2d (N/ϵ)) in time. We integrate sparse Cholesky factorizations into optimization algorithms to obtain fast solvers of the nonlinear PDE. We numerically illustrate our algorithm’s near-linear space/time complexity for a broad class of nonlinear PDEs such as the nonlinear elliptic, Burgers, and Monge-Ampère equations. In summary, we provide a fast, scalable, and accurate method for solving general PDEs with GPs and kernel methods.

97 MATHEMATICS AND COMPUTING↗