Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multivariate regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Supplemental Data for the Manuscript: Quantification of Manganese for ChemCam Mars and Laboratory Spectra Using a Multivariate Model

This dataset includes all of the data needed to validate and/or reproduce the manganese calibration model described in the manuscript. The reference database contains metadata for the new Mn-bearing standards, minerals, and mixtures that are > 2.9 wt.% MnO. In addition, the files include the MnO composition data for all standards used, non-normalized spectral data, mean peak area spectrum, results of outlier determination, RMSECV data, regression vectors, and Test Set predictions.

58 GEOSCIENCES↗

Regularized Differentiation for Bioburden Density Estimation in Planetary Protection

In this paper, we propose and investigate the performance of two novel shrinkage estimators for bioburden density estimation in planetary protection. The estimators are based on the regularized differentiation of a cumulative count of colony forming units collected throughout the data collecting session or the life cycle of the entire mission. The regularized differentiation recasts the problem of bioburden density estimation as a linear least squares problem. The least squares problem is then solved through regularization techniques, such as truncated singular value decomposition and penalized least squares. The regularization is necessary to avoid noise amplification during the differentiation of noisy data. The two regularization estimators are compared with four other commonly used estimators to simultaneously evaluate the means of multivariable independent Poisson distributions: the maximum likelihood, noninformative Bayes estimator with Jeffreys prior, Empirical Bayes using conjugate gamma-Poisson model with gamma parameters selected by method of moments, and the Clevenson-Zidek estimator. It is shown through computer-simulated data that the regularized differentiation based on ridge regression has the smallest mean-squared error among all estimators. The analysis of shrinkage mechanism implemented by regularized differentiation is performed, and it is shown that the regularized differentiation amounts to performing a weighted averaging of all the samples. The weights are determined by the regularization parameter automatically selected by the L-curve technique. Since the method of least squares makes no distributional assumptions about the data, it presents an attractive technique for bioburden density estimation when there are concerns about the misspecification of the distributional model. The paper concludes with the analysis of the bioburden data collected during InSight mission and directions for future work.

97 - MATHEMATICS AND COMPUTING↗

Removing the Effects of Tropical Dynamics from North Pacific Climate Variability

Teleconnections from the Tropics energize variations of the North Pacific climate, but detailed diagnosis of this relationship has proven difficult. Simple univariate methods, such as regression on El Niño-Southern Oscillation (ENSO) indices, may be inadequate since the key dynamical processes involved -- including ENSO diversity in the Tropics, re-emergence of mixed layer thermal anomalies, and oceanic Rossby wave propagation in the North Pacific -- have a variety of overlapping spatial and temporal scales. Here we use a multivariate Linear Inverse Model to quantify tropical and extra-tropical multi-scale dynamical contributions to North Pacific variability, in both observations and CMIP6 models. In observations, we find that the Tropics are responsible for almost half of the seasonal variance, and almost three quarters of the decadal variance, along the North American coast and within the subtropical front region northwest of Hawaii. SST anomalies that are generated by local dynamics within the Northeast Pacific have much shorter time scales, consistent with transient weather forcing by Aleutian low anomalies. Variability within the Kuroshio-Oyashio Extension (KOE) region is considerably less impacted by the Tropics, on all time scales. Consequently, without tropical forcing the dominant pattern of North Pacific variability would be a KOE pattern, rather than the Pacific Decadal Oscillation (PDO). In contrast to observations, most CMIP6 historical simulations produce North Pacific variability that maximizes in the KOE region, with amplitude significantly higher than observed. Correspondingly, the simulated North Pacific in all CMIP6 models is shown to be relatively insensitive to the Tropics, with a dominant spatial pattern generally resembling the KOE pattern, not the PDO.

54 ENVIRONMENTAL SCIENCES↗

Modeling and managing risk early in software development

In order to improve the quality of the software development process, we need to be able to build empirical multivariate models based on data collectable early in the software process. These models need to be both useful for prediction and easy to interpret, so that remedial actions may be taken in order to control and optimize the development process. We present an automated modeling technique which can be used as an alternative to regression techniques. We show how it can be used to facilitate the identification and aid the interpretation of the significant trends which characterize 'high risk' components in several Ada systems. Finally, we evaluate the effectiveness of our technique based on a comparison with logistic regression based models.

Briand, Lionel C.↗

Active Learning-driven Quantitative Synthesis-Structure-Property Relations for Improving Performance and Revealing Active Sites of Nitrogen-Doped Carbon for the Hydrogen Evolution Reaction

While quantitative structure-properties relations (QSPRs) have been developed successfully in multiple fields, catalyst synthesis affects structure and in turn performance, making simple QSPRs inadequate. Furthermore, catalysts often have multiple active sites preventing one from obtaining insights into structure-property relations. Here, we develop a data-driven quantitative synthesis-structure-property relation (QS2PRs) methodology to elucidate correlations between catalyst synthesis conditions, structural properties as well as observed performance and to provide fundamental insights into active sites and a systematic way to optimize practical catalysts. Here, we demonstrate the approach to the synthesis of nitrogen-doped catalysts (NDC) made via pyrolysis for the performance of the electrochemical hydrogen evolution reaction (HER), quantified by the onset potential and the current density. We determine crystallinity, nitrogen species type and fraction, surface area, and pore structure of the NDC’s using XRD, XPS, and BET characterization. We demonstrated that an active learning-based optimization combined with various elementary machine learning tools (regression, principal component analysis, partial least squares) can efficiently identify optimum pyrolysis conditions to tune structural characteristics and performance with concomitant savings in materials and experimental time. Unlike previous reports on the importance of pyridinic or graphitic nitrogen, we discover that the electrochemical performance is not driven by a single catalyst property; rather, it arises from a multivariate influence of nitrogen dopants, pore structure and disorder in the NDC materials. Identification of active sites can help mechanistic understanding and further catalyst improvement.

42 ENGINEERING↗

The LANDSAT-1 multispectral scanner as a tool in the classification of inland lakes

Relationships between LANDSAT-1 multispectral scanner (MSS) data and the trophic status of a group of lakes in the north-northeastern part of the United States were studied by predicting the magnitudes of two trophic state indicators, estimating lake position on a multivariate trophic scale, and automatically classifying lakes according to their trophic state. Initially, the principal component ordination was employed with 100 lakes. MSS data for some 20 lakes was then extracted from computer-compatible tapes (CCT) using a binary marking technique. The output was in the form of descriptive statistics and photographic concatenations. Color ratios were incorporated into regression models for the prediction of Secchi disc transparency, chlorophyll a, and lake position on the tropic scale. Results indicate that the LANDSAT-1 system, although handicapped by low spectral and spatial resolutions as well as excessive cloud cover, can be used as a supplemental data source in lake survey programs.

Boland, D. H. P.↗

Monitoring the Caustic Dissolution of Aluminum Alloy in a Radiochemical Hot Cell Using Raman Spectroscopy

Chemical processing of highly radioactive materials commonly takes place in heavily shielded hot cells. The remote, real-time monitoring of chemical processing streams via optical spectroscopic techniques in hot cells may be particularly useful. Here in this paper, we describe the implementation of Raman spectroscopy and chemometric analysis to monitor the dissolution of aluminum-clad targets containing irradiated aluminum–neptunium oxide cermet pellets in caustic solutions in a hot cell environment. Partial least squares regression analysis was used to generate calibration models to quantify the concentration of dissolved aluminum, nitrate, and hydroxide in solutions within the radiochemical hot cell. This work explored a systematic approach to optimize a matrix of calibration standards using a D-optimal experimental design. The Design of Experiments-based regression model, in comparison to more traditional analytical approaches, was found to be the more practical method for building calibration models, with fewer samples, to obtain informative analytical data from Raman spectra.

36 MATERIALS SCIENCE↗

Understanding the relation between wind- and pressure-driven sea level variability

Sea surface adjustment to combined wind and pressure forcing is examined using numerical solutions to the shallow water equations. The experiments use coastal geometry and bottom topography representative of the North Atlantic and are forced by realistic barometric pressure and wind stress fields. The repsonse to pressure is essentially static or close to the inverted barometer solution at periods longer than a few days and dominates the sea level variability, with wind-driven sea level signals being relatively small. With regard to the dynamic signals, wind-driven fluctuations dominate at long periods, as expected from quasi-geostrophic theory. Pressure becomes more important than wind stress as a source of dynamic signals only at periods shorter than approximately three days. Wind- and pressure-driven sea level fluctuations are anticorrelated over most regions. Hence, regressions of sea level on barometric pressure yield coefficients generally smaller than expected for the inverted barometer response known to be the case in the model. In the regions of significant wind-pressure correlation effects, to infer the correct pressure reponse using statistical methods, input fields must include winds as well as pressure. Because of the nonlocal character of the wind response, multivariate statistical models with local wind driving as input are not very successful. Inclusion of nonlocal wind variability over extensive regions is necessary to extract the correct pressure response. Implications of these results to the interpretation of sea level observations are discussed.

Ponte, Rui M.↗

Linear Least Squares for Correlated Data

Throughout the literature authors have consistently discussed the suspicion that regression results were less than satisfactory when the independent variables were correlated. Camm, Gulledge, and Womer, and Womer and Marcotte provide excellent applied examples of these concerns. Many authors have obtained partial solutions for this problem as discussed by Womer and Marcotte and Wonnacott and Wonnacott, which result in generalized least squares algorithms to solve restrictive cases. This paper presents a simple but relatively general multivariate method for obtaining linear least squares coefficients which are free of the statistical distortion created by correlated independent variables.

Dean, Edwin B.↗

ClimSim: A large multi-scale dataset for hybrid physics-ML climate emulation

Modern climate projections lack adequate spatial and temporal resolution due to computational constraints. A consequence is inaccurate and imprecise predictions of critical processes such as storms. Hybrid methods that combine physics with machine learning (ML) have introduced a new generation of higher fidelity climate simulators that can sidestep Moore’s Law by outsourcing compute-hungry, short, high-resolution simulations to ML emulators. However, this hybrid ML-physics simulation approach requires domain-specific treatment and has been inaccessible to MLexperts because of lack of training data and relevant, easy-to-use workflows. Wepresent ClimSim, the largest-ever dataset designed for hybrid ML-physics research. It comprises multi-scale climate simulations, developed by a consortium of climate scientists and ML researchers. It consists of 5.7 billion pairs of multivariate input and output vectors that isolate the influence of locally-nested, high-resolution, high-fidelity physics on a host climate simulator’s macro-scale physical state. The dataset is global in coverage, spans multiple years at high sampling frequency, and is designed such that resulting emulators are compatible with downstream coupling into operational climate simulators. We implement a range of deterministic and stochastic regression baselines to highlight the ML challenges and their scoring. The data (https://huggingface.co/datasets/LEAP/ClimSim_high-res2) and code(https://leap-stc.github.io/ClimSim)arereleasedopenlytosupport the development of hybrid ML-physics and high-fidelity climate simulations for the benefit of science and society.

artificial intelligence, machine learning↗

Rapid monitoring of fermentations: a feasibility study on biological 2,3-butanediol production

2,3-butanediol (2,3-BDO) is an economically important platform chemical that can be produced by the fermentation of sugars using an engineered strain of Zymomonas mobilis . These fermentations require continuous monitoring and modification of fermentation conditions to maximize 2,3-BDO yields and minimize the production of the undesired coproducts glycerol and acetoin. Because of the time required for sampling and off-line chromatographic measurement of fermentation samples, the ability of fermentation scientists to modify fermentation conditions in a timely manner is limited. The goal of this study was to test if near-infrared spectroscopy (NIRS) along with multivariate statistics could reduce the time needed for this analysis and enable real-time monitoring and control of the fermentation. In this work we developed partial least squares (PLS) calibration models to predict the concentrations of glucose, xylose, 2,3-BDO, acetoin, and glycerol in fermentations via NIRS using two different spectrometers and two different spectroscopy modalities. We first evaluated the feasibility of rapid NIRS monitoring through experiments where we measured the signals from each analyte of interest and built NIRS-based PLS models using spectra from synthetic samples containing uncorrelated concentrations of these analytes. All analytes showed unique spectral signatures, and this initial modeling showed that all analytes could be detected simultaneously. We then began work with samples from laboratory fermentation experiments and tested the feasibility of regression model development across two spectral collection modalities (at-line and on-line) and two instruments: a laboratory-grade instrument and a low-cost instrument with a more limited spectral range. All modalities showed promise in the ability to monitor Z. mobilis fermentations of glucose and xylose to 2,3-BDO. The low-cost instrument displayed a lower signal-to-noise ratio than the laboratory-grade instrument, which led to comparatively lower performance overall, but still provided sufficient accuracy to monitor fermentation trends. While the ease of use of on-line monitoring systems was favored as compared to at-line systems due to the lack of sampling required and potential for automated process control, we observed some decrease in performance due to the additional complexity of the sample matrix. We have demonstrated that NIRS combined with multivariate analysis can be used for at-line and on-line monitoring of the concentrations of glucose, xylose, 2,3-BDO, acetoin, and glycerol during Z. mobilis fermentations. The decrease in signal-to-noise ratio when using a low-cost spectrometer led to greater prediction error than the laboratory-grade spectrometer for at-line monitoring. The on-line monitoring modality showed great promise for real time process control via NIRS.

09 BIOMASS FUELS↗

Using foreground/background analysis to determine leaf and canopy chemistry

Spectral Mixture Analysis (SMA) has become a well established procedure for analyzing imaging spectrometry data, however, the technique is relatively insensitive to minor sources of spectral variation (e.g., discriminating stressed from unstressed vegetation and variations in canopy chemistry). Other statistical approaches have been tried e.g., stepwise multiple linear regression analysis to predict canopy chemistry. Grossman et al. reported that SMLR is sensitive to measurement error and that the prediction of minor chemical components are not independent of patterns observed in more dominant spectral components like water. Further, they observed that the relationships were strongly dependent on the mode of expressing reflectance (R, -log R) and whether chemistry was expressed on a weight (g/g) or are basis (g/sq m). Thus, alternative multivariate techniques need to be examined. Smith et al. reported a revised SMA that they termed Foreground/Background Analysis (FBA) that permits directing the analysis along any axis of variance by identifying vectors through the n-dimensional spectral volume orthonormal to each other. Here, we report an application of the FBA technique for the detection of canopy chemistry using a modified form of the analysis.

Pinzon, J. E.↗

Machine-Learning Assisted Identification of Battery Life Models

Predictive battery life models are commonly utilized to extrapolate degradation trends observed during accelerated aging tests for simulation of degradation in real-world applications. Thus, fitting accelerated aging data as accurately as possible and with low uncertainty is crucial for making believable projections of battery lifetime, but it is challenging to identify algebraic expressions that accurately fit multivariate degradation trends. A review of models published in literature reveal some common expressions for fitting calendar aging data, which is only dependent on temperature and state-of-charge, but no consistency across many models for fitting cycle aging data, indicating the need for a statistically rigorous data driven approach for developing empirical models. This talk will describe a machine-learning assisted method for identification of predictive battery life models utilizing bilevel optimization and symbolic regression. Bilevel optimization with cross-validation is used to statistically determine cell- and stress-dependent model parameters, while symbolic regression identifies both linear and multiplicative candidate expressions to predict stress-dependent degradation rates by selecting low-order subsets of features from a generated feature library. Because model expressions are identified empirically, it is crucial to ensure resulting models behave according to physical expectations, so the stability of models for interpolation or extrapolation is interrogated qualitatively through simulation and quantitatively through cross-validation and uncertainty quantification via bootstrap resampling. This model identification approach substantially improves upon models identified purely using expert judgement in terms of both accuracy and uncertainty. Model simulation and validation is then conducted by deriving a state-equation form of the predictive model, enabling simulation of battery aging under dynamic stresses. This enables validation of the predictive battery model on lab-based tests with varying conditions or on drive-cycle or application-cycle testing protocols. Parameter uncertainty can be carried forward into model simulation, giving lifetime estimates and confidence windows for cell- or system-level lifetime. The financial impact of battery model uncertainty can be estimated by incorporating uncertainty into a technoeconomic model.

battery↗

A Southern Photometric Quasar Catalog from the Dark Energy Survey Data Release 2

We present a catalog of 1.4 million photometrically selected quasar candidates in the southern hemisphere over the ~5000 deg 2 Dark Energy Survey (DES) wide survey area. We combine optical photometry from the DES second data release (DR2) with available near-infrared (NIR) and the all-sky unWISE mid-infrared photometry in the selection. We build models of quasars, galaxies, and stars with multivariate skew-t distributions in the multidimensional space of relative fluxes as functions of redshift (or color for stars) and magnitude. Our selection algorithm assigns probabilities for quasars, galaxies, and stars and simultaneously calculates photometric redshifts (photo-z) for quasar and galaxy candidates. Benchmarking on spectroscopically confirmed objects, we successfully classify (with photometry) 94.7% of quasars, 99.3% of galaxies, and 96.3% of stars when all IR bands (NIR YJHK and WISE W1W2) are available. The classification and photo-z regression success rates decrease when fewer bands are available. Our quasar (galaxy) photo-z quality, defined as the fraction of objects with the difference between the photo-z z p and the spectroscopic redshift z s , |Δz| ≡ |z s - z p |/(1 + z s ) ≤ 0.1, is 92.2% (98.1%) when all IR bands are available, decreasing to 72.2% (90.0%) using optical DES data only. Our photometric quasar catalog achieves an estimated completeness of 89% and purity of 79% at r < 21.5 (0.68 million quasar candidates), with reduced completeness and purity at 21.5 < r ≲ 24. Among the 1.4 million quasar candidates, 87,857 have existing spectra, and 84,978 (96.7%) of them are spectroscopically confirmed quasars. Finally, we provide quasar, galaxy, and star probabilities for all (0.69 billion) photometric sources in the DES DR2 coadded photometric catalog.

79 ASTRONOMY AND ASTROPHYSICS↗

Areal Distribution of the Oxygen-Isotope Ratio in Greenland

Mean values of the oxygen-isotope ratio relative to standard mean ocean water reported for 46 sites on the Greenland ice sheet are compiled together with data on mean annual surface temperature, latitude, 6180 elevation, and mean annual shortest distance to the open ocean denoted by the 10% sea-ice concentration boundary. Stepwise regression analyses, with 6180 as the dependent variable, define two robust models. In the forward mode at the 99.9% confidence level, only temperature enters the model. In the backward mode at the 95% confidence level, only temperature, latitude, and distance to the open ocean remain in the model. Inversions of the models on the basis of 160 gridpoint locations 100 km apart in the area delimited by the surface equilibrium line produce four contoured distributions of 6"0. Two distributions are based on the bivariate model and two on the multivariate model. The second distribution for each model is obtained substituting mean annual surface-temperature values obtained from the Nimbus-7 Temperature Humidity Infrared Radiometer (THIR) database. All four distributions are considered valid, and differences between them are evaluated using contoured anomaly maps. It is suggested that the inversion of the multivariate model using THIR data provides the more reliable pattern for studies of atmospheric advection or for the derivation of ice-flow adjustments for 6180 series obtained from deep-core or ablation-zone sites.

Zwally, H. Jay↗

Development of a Multivariable Parametric Cost Analysis for Space-Based Telescopes

Over the past 400 years, the telescope has proven to be a valuable tool in helping humankind understand the Universe around us. The images and data produced by telescopes have revolutionized planetary, solar, stellar, and galactic astronomy and have inspired a wide range of people, from the child who dreams about the images seen on NASA websites to the most highly trained scientist. Like all scientific endeavors, astronomical research must operate within the constraints imposed by budget limitations. Hence the importance of understanding cost: to find the balance between the dreams of scientists and the restrictions of the available budget. By logically analyzing the data we have collected for over thirty different telescopes from more than 200 different sources, statistical methods, such as plotting regressions and residuals, can be used to determine what drives the cost of telescopes to build and use a cost model for space-based telescopes. Previous cost models have focused their attention on ground-based telescopes due to limited data for space telescopes and the larger number and longer history of ground-based astronomy. Due to the increased availability of cost data from recent space-telescope construction, we have been able to produce and begin testing a comprehensive cost model for space telescopes, with guidance from the cost models for ground-based telescopes. By separating the variables that effect cost such as diameter, mass, wavelength, density, data rate, and number of instruments, we advance the goal to better understand the cost drivers of space telescopes.. The use of sophisticated mathematical techniques to improve the accuracy of cost models has the potential to help society make informed decisions about proposed scientific projects. An improved knowledge of cost will allow scientists to get the maximum value returned for the money given and create a harmony between the visions of scientists and the reality of a budget.

Dollinger, Courtnay↗

Application of Partial Least Squares Approaches to Pyroprocessing ER Data

Multivariate approaches show promise for application to process monitoring for safeguards of pyroprocessing. Past MPACT work explored the application of Principal Component Analysis (PCA) to detect off-normal conditions in pyroprocessing electrorefiner (ER) data from in the Hot Fuel Examination Facility (HFEF) at Idaho National Laboratory (INL) known as the Scalable Pyrochemical Recycling testbed (SPyRe) ER. PCA, however, does not consider the output variables. In FY24, multivariate analysis was extended from PCA to Partial Least Squares (PLS) analysis. PLS maximizes the variance between both the input signals and output variables. In the case of this work, PLS was applied in two different manners: Predictive PLS and Discriminant PLS. Predictive PLS maximizes the covariance between the process variables of the ER and the measured U concentration from in-situ voltammetry. Discriminant PLS maximizes the covariance between the process variables and a set of training process “states” such as known off-normal conditions. By projecting into the latent variable space in PLS, the process variables can be regressed onto the outputs and predictions can be made for new data sets. In this work, by applying predictive PLS, a penalized non-linear PLS approach was able to make predictions of concentration based on test and training data and detect when operations were off-normal. However, the predictive PLS does not classify the signals to which off-normal operations are attributable. Discriminant PLS can be used to classify off-normal operations but is inadequate to properly classify specific off-normal classes like power supply faults when the Discriminant PLS model is only specifically trained to detect that off-normal class. When all faults are trained against the observation data, all three operational classes are accurately classified and distinguished. Thus, future application of latent variable techniques should not select any given method, but should use a mixture of PCA, Predictive PLS, and Discriminant PLS.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗