Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multivariate regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Quantification of Rare Earth Elements in the Parts Per Million Range: A Novel Approach in the Application of Laser-Induced Breakdown Spectroscopy

This work extends a previous percentage level concentration study of the optical emission spectra for six rare earth elements, europium (Eu), gadolinium (Gd), lanthanum (La), praseodymium (Pr), neodymium (Nd), and samarium (Sm), along with the transition metal, yttrium (Y) using laser-induced breakdown spectroscopy (LIBS). The concentration of these six rare earth elements and yttrium has been attempted for the first time systematically down to parts per million (ppm) concentration levels ranging from 30 to 300 ppm. In this study, the authors have developed multivariate models for each element capable of predicting concentration with acceptable to excellent levels of accuracy. Additionally, partial least squares regression coefficients were used to identify key spectral features able to be used in this lower concentration regime. This study has demonstrated that it is conceivable to quantify the six rare earth elements along with yttrium at low concentrations in the parts per million levels.

59 BASIC BIOLOGICAL SCIENCES↗

Leveraging visible and near-infrared spectroelectrochemistry to calibrate a robust model for Vanadium(IV/V) in varying nitric acid and temperature levels

Spectroelectrochemistry and optimal design of experiments can be used to rapidly build accurate models for species quantification and enable a greater level of process awareness. Optical spectroscopy can provide vital elemental and molecular information, but several hurdles must be overcome before it can become a widely adopted analytical method for remote analysis in the nuclear field. Analytes with varying oxidation state, acid concentration, and fluctuating temperature must be efficiently accounted for to minimize time and resources in restrictive hot cell environments. The classic one-factor-at-a-time approach is not suitable for frequent calibration/maintenance operations in this setting. Therefore, a novel alternative was developed to characterize a system containing vanadium(IV/V) (0.01–0.1 M), nitric acid (0.1–4 M), and varying temperatures (20–45 °C). Here, spectroelectrochemistry methods were used to acquire a sample set selected by optimal design of experiments. This new approach allows for the accurate analysis of vanadium and HNO 3 concentration by leveraging UV–Vis–NIR absorption spectroscopy with robust and accurate chemometric models. The top model's root mean squared error of prediction percent values were 3.47%, 4.06%, 3.40%, and 10.9% for V(IV), V(V), HNO 3 , and temperature, respectively. These models, efficiently developed using the designed approach, exhibited strong predictive accuracy for vanadium and acid with varying oxidation states and temperature using only spectrophotometry, which advances current technology for real-world hot cell applications. Additionally, Nernstian analysis of the V(IV/V) standard potential was performed using traditional absorbance methods and multivariate curve resolution (MCR). The successful tests demonstrated that MCR Nernst tests may be valuable in highly convoluted spectral systems to better understand the redox processes' behavior.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Latent Stochastic Differential Equations for Modeling Quasar Variability and Inferring Black Hole Properties

Quasars are bright and unobscured active galactic nuclei (AGN) thought to be powered by the accretion of matter around supermassive black holes at the centers of galaxies. The temporal variability of a quasar’s brightness contains valuable information about its physical properties. The UV/optical variability is thought to be a stochastic process, often represented as a damped random walk described by a stochastic differential equation (SDE). Upcoming wide-field telescopes such as the Rubin Observatory Legacy Survey of Space and Time (LSST) are expected to observe tens of millions of AGN in multiple filters over a ten year period, so there is a need for efficient and automated modeling techniques that can handle the large volume of data. Latent SDEs are machine learning models well suited for modeling quasar variability, as they can explicitly capture the underlying stochastic dynamics. In this work, we adapt latent SDEs to jointly reconstruct multivariate quasar light curves and infer their physical properties such as the black hole mass, inclination angle, and temperature slope. Our model is trained on realistic simulations of LSST ten year quasar light curves, and we demonstrate its ability to reconstruct quasar light curves even in the presence of long seasonal gaps and irregular sampling across different bands, outperforming a multioutput Gaussian process regression baseline. Our method has the potential to provide a deeper understanding of the physical properties of quasars and is applicable to a wide range of other multivariate time series with missing data and irregular sampling.

79 ASTRONOMY AND ASTROPHYSICS↗

From multivariate to functional data analysis: Fundamentals, recent developments, and emerging areas

Functional data analysis (FDA), which is a branch of statistics on modeling infinite dimensional random vectors resided in functional spaces, has become a major research area for Journal of Multivariate Analysis. We review some fundamental concepts of FDA, their origins and connections from multivariate analysis, and some of its recent developments, including multi-level functional data analysis, high-dimensional functional regression, and dependent functional data analysis. Here, we also discuss the impact of these new methodology developments on genetics, plant science, wearable device data analysis, image data analysis, and business analytics. Two real data examples are provided to motivate our discussions.

97 MATHEMATICS AND COMPUTING↗

Identification and correction of temporal and spatial distortions in scanning transmission electron microscopy

Scanning transmission electron microscopy (STEM) has become the technique of choice for quantitative characterization of atomic structure of materials, where the minute displacements of atomic columns from high-symmetry positions can be used to map strain, polarization, octahedra tilts, and other physical and chemical order parameter fields. The latter can be used as inputs into mesoscopic and atomistic models, providing insight into the correlative relationships and generative physics of materials on the atomic level. However, these quantitative applications of STEM necessitate understanding the microscope induced image distortions and developing the pathways to compensate them both as part of a rapid calibration procedure for in situ imaging, and the post-experimental data analysis stage. Here, we explore the spatiotemporal structure of the microscopic distortions in STEM using multivariate analysis of the atomic trajectories in the image stacks. Based on the behavior of principal component analysis (PCA), we develop the Gaussian process (GP)-based regression method for quantification of the distortion function. The limitations of such an approach and possible strategies for implementation as a part of in-line data acquisition in STEM are discussed. Here, the analysis workflow is summarized in a Jupyter notebook that can be used to retrace the analysis and analyze the reader's data.

36 MATERIALS SCIENCE↗

Comparing Calibration Algorithms for the Rapid Characterization of Pretreated Corn Stover Using Near-Infrared Spectroscopy

Rapid characterization of biomass composition is a key enabling technology for biorefineries—the ability to measure the chemical composition of biomass materials entering the biorefinery as well as the composition of key process intermediate streams would allow real-time process control and the development of robust models to predict process performance. The utility of near-infrared (NIR) spectroscopy for rapid characterization requires multivariate algorithms for building calibration models. The most prevalent algorithm used for building calibration models using NIR spectra is the linear modeling algorithm Partial Least Squares Regression (PLS). Nonlinear regression algorithms (which are typically more computationally intensive than linear modeling approaches) have gained popularity in recent years due to their ability to solve a wide variety of classification and regression problems and the dramatic increase in available computational resources. In this work, we demonstrate that a calibration model can predict the composition of corn stover process intermediate samples pretreated with three different treatments—hot water (HW), dilute acid (DA), and deacetylation followed by dilute acid (DDA). We quantitatively compare three different algorithms for building prediction models based on near-infrared spectroscopy—partial least squares (PLS), support vector machines (SVM), and random forests (RF). We demonstrate the utility of improving model performance by accounting for instrument performance variability using repeated measurements of standard materials (e.g., the “repeatability file” strategy) and investigate its performance with nonlinear regression techniques, and we discuss methods for quantifying the uncertainties of specific predictions among the three methods.

09 BIOMASS FUELS↗

Hierarchical Modeling to Enhance Spectrophotometry Measurements—Overcoming Dynamic Range Limitations for Remote Monitoring of Neptunium

A robust hierarchical model has been demonstrated for monitoring a wide range of neptunium concentrations (0.75–890 mM) and varying temperatures (10–80 °C) using chemometrics and feature selection. The visible–near infrared electronic absorption spectrum (400–1700 nm) of monocharged neptunyl dioxocation (Np(V) = NpO2+) includes many bands, which have molar absorption coefficients that differ by nearly 2 orders of magnitude. The shape, position, and intensity of these bands differ with chemical interactions and changing temperature. These challenges make traditional quantification by univariate methods unfeasible. Measuring Np(V) concentration over several orders of magnitude would typically necessitate cells with varying path length, optical switches, and/or multiple spectrophotometers. Alternatively, the differences in the molar extinction coefficients for multiple absorption bands can be used to quantify Np(V) concentration over 3 orders of magnitude with a single optical path length (1 mm) and a hierarchical multivariate model. In this work, principal component analysis was used to distinguish the concentration regime of the sample, directing it to the relevant partial least squares regression submodels. Each submodel was optimized with unique feature selection filters that were selected by a genetic algorithm to enhance predictions. Through this approach, the percent root mean square error of prediction values were ≤1.05% for Np(V) concentrations and ≤4% for temperatures. This approach may be applied to other nuclear fuel cycle and environmental applications requiring real-time spectroscopic measurements over a wide range of conditions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Combining Emergent Constraints for Climate Sensitivity

A method is proposed for combining information from several emergent constraints into a probabilistic estimate for a climate sensitivity proxy Y such as equilibrium climate sensitivity (ECS). The method is based on fitting a multivariate Gaussian PDF for Y and the emergent constraints using an ensemble of global climate models (GCMs); it can be viewed as a form of multiple linear regression of Y on the constraints. The method accounts for uncertainties in sampling this multidimensional PDF with a small number of models, for observational uncertainties in the constraints, and for overconfidence about the correlation of the constraints with the climate sensitivity. Its general form (Method C) accounts for correlations between the constraints. Method C becomes less robust when some constraints are too strongly related to each other; this can be mitigated using regularization approaches such as ridge regression. An illuminating special case, Method U, neglects any correlations between constraints except through their mutual relationship to the climate proxy; it is more robust to small GCM sample size and is appealingly interpretable. These methods are applied to ECS and the climate feedback parameter using a previously published set of 11 possible emergent constraints derived from climate models in the Coupled Model Intercomparison Project (CMIP). The ±2σ posterior range of ECS for Method C with no overconfidence adjustment is 4.3 ± 0.7 K. For Method U with a large overconfidence adjustment, it is 4.0 ± 1.3 K. This study adds confidence to past findings that most constraints predict higher climate sensitivity than the CMIP mean.

54 ENVIRONMENTAL SCIENCES↗

Online evolutionary neural architecture search for multivariate non-stationary time series forecasting

Time series forecasting (TSF) is one of the most important tasks in data science. TSF models are usually pre-trained with historical data and then applied on future unseen datapoints. However, real-world time series data is usually non-stationary and models trained offline usually face problems from data drift. Models trained and designed in an offline fashion can not quickly adapt to changes quickly or be deployed in real-time. To address these issues, this work presents the Online NeuroEvolution-based Neural Architecture Search (ONE-NAS) algorithm, which is a novel neural architecture search method capable of automatically designing and dynamically training recurrent neural networks (RNNs) for online forecasting tasks. Without any pre-training, ONE-NAS utilizes populations of RNNs that are continuously updated with new network structures and weights in response to new multivariate input data. ONE-NAS is tested on real-world, large-scale multivariate wind turbine data as well as the univariate Dow Jones Industrial Average (DJIA) dataset. These results demonstrate that ONE-NAS outperforms traditional statistical time series forecasting methods, including online linear regression, fixed long short-term memory (LSTM) and gated recurrent unit (GRU) models trained online, as well as state-of-the-art, online ARIMA strategies. Additionally, results show that utilizing multiple populations of RNNs which are periodically repopulated provide significant performance improvements, allowing this online neural network architecture design and training to be successful.

97 MATHEMATICS AND COMPUTING↗

Effect of vaccination on the case fatality rate for COVID-19 infections 2020–2021: multivariate modelling of data from the US Department of Veterans Affairs

Objectives: To evaluate the benefits of vaccination on the case fatality rate (CFR) for COVID-19 infections. Design, setting and participants: The US Department of Veterans Affairs has 130 medical centres. We created multivariate models from these data—339 772 patients with COVID-19—as of 30 September 2021. Outcome measures: The primary outcome for all models was death within 60 days of the diagnosis. Logistic regression was used to derive adjusted ORs for vaccination and infection with Delta versus earlier variants. Models were adjusted for confounding factors, including demographics, comorbidity indices and novel parameters representing prior diagnoses, vital signs/baseline laboratory tests and outpatient treatments. Patients with a Delta infection were divided into eight cohorts based on the time from vaccination to diagnosis. A common model was used to estimate the odds of death associated with vaccination for each cohort relative to that of unvaccinated patients. Results: 9.1% of subjects were vaccinated. 21.5% had the Delta variant. 18 120 patients (5.33%) died within 60 days of their diagnoses. The adjusted OR for a Delta infection was 1.87±0.05, which corresponds to a relative risk (RR) of 1.78. The overall adjusted OR for prior vaccination was 0.280±0.011 corresponding to an RR of 0.291. Raw CFR rose steadily after 10–14 weeks. The OR for vaccination remained stable for 10–34 weeks. Conclusions: Our CFR model controls for the severity of confounding factors and priority of vaccination, rather than solely using the presence of comorbidities. Our results confirm that Delta was more lethal than earlier variants and that vaccination is an effective means of preventing death. After adjusting for major selection biases, we found no evidence that the benefits of vaccination on CFR declined over 34 weeks. We suggest that this model can be used to evaluate vaccines designed for emerging variants.

59 BASIC BIOLOGICAL SCIENCES↗

Joint Modeling of Quasar Variability and Accretion Disk Reprocessing Using Latent Stochastic Differential Equations

Quasars are bright active galactic nuclei powered by the accretion of matter around supermassive black holes at the center of galaxies. Their stochastic brightness variability depends on the physical properties of the accretion disk and black hole. The upcoming Rubin Observatory Legacy Survey of Space and Time (LSST) is expected to observe tens of millions of quasars, so there is a need for efficient techniques like machine learning that can handle the large volume of data. Quasar variability is believed to be driven by an X-ray corona, which is reprocessed by the accretion disk and emitted as UV/optical variability. We are the first to introduce an auto-differentiable simulation of the accretion disk and reprocessing. We use the simulation as a direct component of our neural network to jointly model the driving variability and reprocessing, trained with supervised learning on simulated LSST-like 10 yr quasar light curves. We encode the light curves using a transformer encoder, and the driving variability is reconstructed using latent stochastic differential equations, a physically motivated generative deep learning method that can model continuous-time stochastic dynamics. By embedding the physical processes of the driving signal and reprocessing into our network, we achieve a model that is more robust and interpretable. We demonstrate that our model outperforms a Gaussian process regression baseline and can infer accretion disk parameters and time delays between wave bands, even for out-of-distribution driving signals. Our approach provides a powerful framework that can be adapted to solve other inverse problems in multivariate time series.

Fagin, Joshua [City Univ. of New York (CUNY), NY (↗

Insights into Tetravalent Np Speciation in HNO 3 through Spectroelectrochemistry and Multivariate Analysis

In situ optical spectroscopy, spectropotentiometry, and multivariate analysis were applied to the Np(IV) nitrate system to better understand speciation and quantify HNO 3 concentration. Thin-layer spectropotentiometry, or spectroelectrochemistry, was leveraged to isolate and stabilize Np(IV) without compromising the solution conditions and generate representative Vis-NIR absorption spectra from 0.5 to 10 M HNO 3 and benchmark the corresponding Np(IV) molar absorptivity coefficients. Spectra were described with principal component analysis (PCA) to identify the purest Np(IV) absorbance spectra among other oxidation states [e.g., Np(V/VI)] at each acid concentration and then to identify the primary sources of variance within each Np(IV) spectrum with respect to Np(IV) nitrate complexes. Then, partial least-squares regression (PLSR) and support vector regression (SVR) models were built to predict HNO 3 concentration from the Np(IV) spectral data. The nonlinear SVR model outperformed the linear PLSR model for the HNO 3 concentration predictions. Finally, the inclusion of spectra collected in edge and center point HNO 3 concentrations in the calibration set was determined to be crucial for producing models with strong predictive capabilities. The multivariate approach used in this study makes it possible to quantify HNO 3 concentration solely based on Np(IV) absorption spectra, which is essential to quantifying processing streams in various online monitoring applications.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Quantifying neptunium oxidation states in nitric acid through spectroelectrochemistry and chemometrics

Controlled-potential in situ thin-layer spectropotentiometry was leveraged to generate visible/near-infrared (VIS/NIR) absorption spectral data sets for the development of chemometric models to quantify Np(III/IV/V/VI) oxidation states in HNO 3 . This technology would be valuable in laboratory studies and when monitoring process solutions to guide feed adjustments for radiochemical separations—the performance of which depends on oxidation state. This approach successfully isolated and stabilized Np species in pure (~99%) oxidation states without compromising solution optical properties. Multivariate curve resolution–alternating least squares models were evaluated to resolve spectral and component concentrations from a scan that sequentially produced Np(VI), Np(V), Np(IV), and Np(III) spectra with mixtures of two valences at a time. Although it provided a useful approximation, the method was not able to quantitively resolve each component likely because of rotational ambiguity. Additionally, partial least squares regression models were built from artificial and electrochemically generated VIS/NIR spectral training sets to study the effect of interionic interactions on spectral characteristics. Models built with true Bi-chemical mixtures of coexisting Np oxidation states and spectra generated from additive combinations of pure end points had similar prediction performance. This methodology can be used to directly quantify Np concentration and the ratio of Np oxidation states and other actinides in remote settings such as hot cells.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

Mapping wall-to-wall fractional cover of Arctic tundra plant functional types in Alaska using 20-m spatial resolution satellite imagery and harmonized plot observations

Estimates of fractional cover (fCover) across given land surfaces are used to assess, and often model, vegetation composition and diversity, which are crucial for understanding the health and functioning of terrestrial ecosystems. Remote sensing provides a useful means for scaling local, plot-measured fCover estimates to regional scales. Leveraging a recently synthesized and harmonized plot database, this study generated wall-to-wall maps of fCover for six Alaskan-Arctic plant functional types (PFT), including non-vascular plants, forbs, graminoids, and deciduous and evergreen shrubs, using 20-m satellite data (Sentinel-1, Sentinel-2, ArcticDEM) using a machine learning regression approach, specifically the random forest (RF) algorithm, which is well-suited for handling nonlinear relationships and high-dimensional satellite datasets. This study additionally addressed the spatio-temporal inconsistencies e.g., sampling scale, plot size, and collection year in plot measured fCover by adopting a multivariate outlier detection approach—Cook’s distance—to identify high-quality plots for model training and validation. Our approach achieves high accuracy (R 2 = 0.59–0.93, root mean squared errors = 0.02–0.10 for all PFTs) between plot-observed and satellite-derived fCover when using high-quality plot samples. The mapped fCover characterizes the spatial patterns of different PFTs across the tundra biome at a 20-m resolution, providing key information needed for improved representation of Arctic tundra vegetation in terrestrial biosphere models to better understand climate-vegetation feedback across the Arctic tundra.

Arctic tundra↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Supplemental Data for the Manuscript: Quantification of Manganese for ChemCam Mars and Laboratory Spectra Using a Multivariate Model

This dataset includes all of the data needed to validate and/or reproduce the manganese calibration model described in the manuscript. The reference database contains metadata for the new Mn-bearing standards, minerals, and mixtures that are > 2.9 wt.% MnO. In addition, the files include the MnO composition data for all standards used, non-normalized spectral data, mean peak area spectrum, results of outlier determination, RMSECV data, regression vectors, and Test Set predictions.

58 GEOSCIENCES↗

Regularized Differentiation for Bioburden Density Estimation in Planetary Protection

In this paper, we propose and investigate the performance of two novel shrinkage estimators for bioburden density estimation in planetary protection. The estimators are based on the regularized differentiation of a cumulative count of colony forming units collected throughout the data collecting session or the life cycle of the entire mission. The regularized differentiation recasts the problem of bioburden density estimation as a linear least squares problem. The least squares problem is then solved through regularization techniques, such as truncated singular value decomposition and penalized least squares. The regularization is necessary to avoid noise amplification during the differentiation of noisy data. The two regularization estimators are compared with four other commonly used estimators to simultaneously evaluate the means of multivariable independent Poisson distributions: the maximum likelihood, noninformative Bayes estimator with Jeffreys prior, Empirical Bayes using conjugate gamma-Poisson model with gamma parameters selected by method of moments, and the Clevenson-Zidek estimator. It is shown through computer-simulated data that the regularized differentiation based on ridge regression has the smallest mean-squared error among all estimators. The analysis of shrinkage mechanism implemented by regularized differentiation is performed, and it is shown that the regularized differentiation amounts to performing a weighted averaging of all the samples. The weights are determined by the regularization parameter automatically selected by the L-curve technique. Since the method of least squares makes no distributional assumptions about the data, it presents an attractive technique for bioburden density estimation when there are concerns about the misspecification of the distributional model. The paper concludes with the analysis of the bioburden data collected during InSight mission and directions for future work.

97 - MATHEMATICS AND COMPUTING↗