Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “robust regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Taming nuclear mass models with Gaussian processes

We propose a new set of nuclear mass predictions based on multiple theoretical mass models. By employing Gaussian process regression with the Matérn kernel, we achieved root-mean-square (rms) deviations below 100 keV for the training dataset. The best-performing mass models achieved rms deviations below 150 keV for the new precise mass data from AME2020, whereas the ensemble average showed robust performance across the nuclear chart. Our approach uniquely combines: (1) systematic refinement of eight mass models through their residuals, (2) physics-informed features, including magic numbers, nucleon parity numbers, neutron excess, and nuclear collectivity, and (3) theory-to-theory validation demonstrating robust extrapolation capability. We find that the Matérn kernel provides superior uncertainty quantification compared to the RBF kernel, with a length-scale analysis revealing enhanced inter-nuclei correlations. We provide complete mass predictions for all unknown nuclides in AME2020, offering valuable constraints for nuclear structure studies and astrophysical modeling when used with proper uncertainty propagation.

Gaussian processes↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗

Predicting the evolution of biomass bulk density through feedstock preprocessing: Discrete element modeling, regression analysis, and pilot-scale validation

Bulk density is an important material property of biomass feedstocks, influencing handling, storage, transport costs, and conversion efficiency. In this study, predictive regression models for loose and tapped bulk densities of Alamo and Cave-in-Rock switchgrass are developed using a comprehensive dataset generated via calibrated bonded-sphere discrete element method (DEM) simulations. Here, a key contribution of this study is the use of a DEM-based approach, which correlates density with moisture content and particle size distribution parameters and enables analysis across a continuous particle size range, overcoming limitations of purely experimental data. For comparison, regression models are also developed using only experimental data from pilot-scale runs at the Biomass Feedstock National User Facility at Idaho National Laboratory. Validation against pilot-scale data showed reasonable prediction accuracy for both model types, particularly for smaller particle sizes (post-secondary grinding). While the experimental model showed slightly better performance matching the validation data in some cases, the DEM-based model benefits from a much larger dataset, reduced predictor multicollinearity, and continuous parameter coverage, highlighting the utility of validated simulation models for developing robust predictive tools for biomass preprocessing applications.

09 - BIOMASS FUELS↗

A nonparametric software reliability growth model

Miller and Sofer have presented a nonparametric method for estimating the failure rate of a software program. The method is based on the complete monotonicity property of the failure rate function, and uses a regression approach to obtain estimates of the current software failure rate. This completely monotone software model is extended. It is shown how it can also provide long-range predictions of future reliability growth. Preliminary testing indicates that the method is competitive with parametric approaches, while being more robust.

Miller, Douglas R.↗

NASA Experimental Program to Stimulate Competitive Research: South Carolina

The use of an appropriate relationship model is critical for reliable prediction of future urban growth. Identification of proper variables and mathematic functions and determination of the weights or coefficients are the key tasks for building such a model. Although the conventional logistic regression model is appropriate for handing land use problems, it appears insufficient to address the issue of interdependency of the predictor variables. This study used an alternative approach to simulation and modeling urban growth using artificial neural networks. It developed an operational neural network model trained using a robust backpropagation method. The model was applied in the Myrtle Beach region of South Carolina, and tested with both global datasets and areal datasets to examine the strength of both regional models and areal models. The results indicate that the neural network model not only has many theoretic advantages over other conventional mathematic models in representing the complex urban systems, but also is practically superior to the logistic model in its capability to predict urban growth with better - accuracy and less variation. The neural network model is particularly effective in terms of successfully identifying urban patterns in the rural areas where the logistic model often falls short. It was also found from the area-based tests that there are significant intra-regional differentiations in urban growth with different rules and rates. This suggests that the global modeling approach, or one model for the entire region, may not be adequate for simulation of a urban growth at the regional scale. Future research should develop methods for identification and subdivision of these areas and use a set of area-based models to address the issues of multi-centered, intra- regionally differentiated urban growth.

Sutton, Michael A.↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Prediction of electric and magnetic fields from spectral data using machine learning algorithms for Doppler-free saturation spectroscopy diagnostics

The prediction of electric and magnetic field amplitudes from atomic spectral data is critical for plasma control in fusion devices such as tokamaks. Conventional approaches that rely on physics-based models are computationally expensive and unsuitable for real-time applications. In this work, we develop and benchmark three machine learning algorithms—simulation-based inference (SBI), fully connected neural networks (FCNN), and histogram-based gradient boosting regression (GBR-Hist)—to infer field intensities directly from Doppler-free saturation spectroscopy (DFSS) spectra. Synthetic datasets of spectra were generated using the EZSSS code and evaluated both with and without added Poisson noise to mimic experimental conditions. We find that SBI achieves the highest accuracy and robustness, FCNN provides a strong balance of accuracy and computational efficiency for real-time applications, and GBR-Hist offers the fastest inference but is more sensitive to noise. Furthermore, these results demonstrate the potential of machine learning to accelerate DFSS analysis and enhance its utility for plasma diagnostics and control.

Doppler-free saturation spectroscopy↗

Observations of Coastal Wind Momentum Flux: Dependence on Fetch and Waves with Comparisons to COARE

Here, using observations from the Martha’s Vineyard Coastal Observatory, this paper investigates how momentum flux in the marine atmospheric surface layer over the coastal ocean varies with fetch, wave age, and wave slope and assesses the performance of the COARE 3.5 bulk parameterization. Long-fetch (at least 300 km) and short-fetch (3–6 km away from land) conditions have very similar momentum flux, with the latter being just 15% higher. The COARE 3.5 wind speed–dependent formulation closely matches the observations. The sea state dependence of wave age and wave slope is analyzed by considering both peak frequency and mean frequency in a wave spectrum. The slope of regression lines between normalized roughness and wave age is sensitive to wind speed ranges and the scatter of momentum flux, potentially explaining why earlier studies did not find a universal formula to characterize the wave dependence. Although the observed momentum flux exhibits an obvious dependence on wave age, a robust quadratic fit between momentum flux and the 10-m neutral wind speed exists only for young waves. For the short-fetch conditions, the momentum flux does not increase as wave slope increases. This may be because waves generated by local wind are still weak and swell that propagates from other areas dominates the wave spectra. In other words, the wave slope computed by considering the spectra does not properly reflect the local wind–wave interaction.

Air-sea interaction↗

Performance of a dynamic single bubbler in single and two-phase immiscible liquids

Ensuring nonproliferation and safeguards of special nuclear materials (SNM) is a critical aspect of advancing the nuclear fuel cycle. Traditional bubbler systems used to estimate liquid levels and densities in nuclear recycling processes have limitations, particularly in harsh environments where dip-tube corrosion and buildup necessitate frequent maintenance and recalibration. This study explores the Dynamic Single Bubbler (DSB) method, which utilizes a single dip-tube attached to a linear actuator to estimate liquid properties dynamically. This approach is extended to estimate liquid-liquid interfaces in immiscible liquids and employs a linear regression method to reduce uncertainties and improve accuracy. The DSB method achieved density estimate uncertainties of less than 0.5% and surface level estimate uncertainties typically under 0.5%, across various fluids including water, acetone, methanol, mineral oil, glycerol, and aqueous salt solutions. Results indicate that the DSB method provides accurate and robust estimates of liquid density and surface levels with minimal maintenance and without the need for calibration. Additionally, the method's applicability to immiscible liquids and various dip-tube geometries was demonstrated, showing promise for widespread use in nuclear and other industrial applications.

98 - NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL↗

Estimation of multidimensional precipitation parameters by areal estimates of oceanic rainfall

The parameters of the multidimensional precipitation model proposed by Waymire et al. (1984) are estimated using the areal-averaged radar measurements of precipitation of the Global Atlantic Tropical Experiment (GATE) data set. The procedure followed was the fitting of the first- and second-order moments at different aggregation scales by nonlinear regression techniques. The numerical estimates of the parameters using different subsets of GATE information were reasonably stable, i.e., they were not affected by changes of the area-averaging size, temporal length of the records, and percentage of areal coverage of rainfall. This suggests that the estimation procedure is relatively robust and suitable to estimate the parameters of the multidimensional model in areas of sparse density of rain gages. The use of the space-time spectrum of rainfall to help in the determination of sampling errors due to intermittent visits of future space-borne low-altitude sensors of precipitation is also discussed.

Valdes, J. B.↗

Fundamental Insights into Cathode Stability: Linking Compositional Tuning and Local Coordination in Complex Metal Oxides under Aqueous Transformations

Compositional tuning of complex metal oxides in Li-ion battery materials influences their performance as well as their end-of-life behavior, in particular, the tendency to release toxic metal cations in aqueous solution. We modeled ternary variants of a parent LiCoO 2 delafossite structure by varying the metal identity and relative amounts. This yielded ten model formulations of Li(A 4/6 B 1/6 C 1/6 )O 2 , where the material is enriched with the A metal and doped with B and C, with Ni, Mn, Co, Fe, Al, V, and Ti as constituent metals. To assess their stability in aqueous conditions, metal release energetics were calculated using a combination of Density Functional Theory calculations and thermodynamics. Metal release in ternary oxides is dictated by subtle variations in the coordination environment of the leaving group. To identify governing chemical features across diverse compositions with varying local coordination environments, we leverage random forest regression and descriptor importance analysis. A key result is that metal–oxygen orbital hybridization, quantified using a projected density-of-states-derived descriptor, H d/p , provides a physically grounded measure of interaction strength that governs metal release energetics. This refined perspective goes beyond conventional oxidation state considerations and offers more robust insights for materials science. Finally, we model defect surface-bound O 2 dimer formation as a proxy for reactive oxygen species (ROS) generation. The results show that Ni-rich compositions more readily stabilize spin-polarized O 2 dimers, corroborating experimental reports of an increased ROS-driven biological response. In conclusion, our results establish a compositional and electronic basis for metal release and surface oxygen reactivity that form a rationale for complex metal oxide design principles.

36 MATERIALS SCIENCE↗

Advancing Glaciological Applications of Remote Sensing with EO-1: (1) Mapping Snow Grain Size and Albedo on the Greenland Ice Sheet Using an Imaging Spectrometer, and (2) ALI Evaluation for Subtle Surface Topographic Mapping via Shape-from Shading

The Hyperion sensor, onboard NASA's Earth Observing-1 (EO-1) satellite,is an imaging spectroradiometer with 220 spectral bands over the spectral range from 0.4 - 2.5 microns. Over the course of summer 2001, the instrument acquired numerous images over the Greenland ice sheet. Our main motivation is to develop an accurate and robust approach for measuring the broadband albedo of snow from satellites. Satellite-derived estimates of broadband have typically been plagued with three problems: errors resulting from inaccurate atmospheric correction, particularly in the visible wavelengths from the conversion of reflectance to albedo (accounting for snow BRDE); and errors resulting from regression-based approaches used to convert narrowband albedo to broadband albedo. A typerspectral method has been developed that substantially reduces these three main sources of error and produces highly accurate estimates of snow albedo. This technique uses hyperspectral data from 0.98 - 1.06 microns, spanning a spectral absorption feature centered at 1.03 microns. A key aspect of this work is that this spectral range is within an atmospheric transmission window and reflectances are largely unaffected by atmospheric aerosols, water vapor, or ozone. In this investigation, we make broadband albedo measurements at four sites on the Greenland ice sheet: Summit, a high altitude station in central Greenland; the ETH/CU camp, a camp on the equilibrium line in western Greenland; Crawford Point, a site located between Summit and the ETH/CU camp; and Tunu, a site located in northeastern Greenland at 2000 m. altitude. Each of these sites has an automated weather station (AWS) that continually measures broadband albedo thereby providing validation data.

Source record↗

Predictive Model for Starlink Maritime Performance Using Multi-Horizon RandomForest

Low Earth orbit (LEO) satellite systems have become a crucial enabler of broadband access for maritime industries, where traditional networks are unavailable. However, the high mobility of LEO constellations and constantly changing weather conditions result in unpredictable link fluctuations, limiting the ability of maritime platforms to plan bandwidth usage proactively. To the best of our knowledge, no prior work has developed a short-term predictive model for maritime LEO connectivity using real experimental field measurements. This paper proposes a data-driven forecasting model that predicts future downlink throughput using multi-horizon RandomForest regression. The model is trained using real experimental coastal measurement data incorporating recent throughput history, network-layer indicators, and environmental variables. The proposed approach reduces mean absolute error by approximately 31% compared to a persistence baseline for 15-minute horizons. It maintains a measurable improvement at 30 minutes, despite increased stochasticity. These findings confirm that proactive bandwidth awareness is feasible on maritime platforms and can effectively support operational decisions such as adaptive streaming, routing, and resource scheduling. The performance gap between forecasting horizons also highlights the need for expanded offshore datasets to improve prediction robustness under harsher maritime environments.

97 MATHEMATICS AND COMPUTING↗

Anomaly detection in collider physics via factorized observables

To maximize the discovery potential of high-energy colliders, experimental searches should be sensitive to unforeseen new physics scenarios. This goal has motivated the use of machine learning for unsupervised anomaly detection. In this paper, we introduce a new anomaly detection strategy called : factorized observables for regressing conditional expectations. Our approach is based on the inductive bias of factorization, which is the idea that the physics governing different energy scales can be treated as approximately independent. Assuming factorization holds separately for signal and background processes, the appearance of nontrivial correlations between low- and high-energy observables is a robust indicator of new physics. Under the most restrictive form of factorization, a machine-learned model trained to identify such correlations will in fact converge to the optimal new physics classifier. We test on a benchmark anomaly detection task for the Large Hadron Collider involving collimated sprays of particles called jets. By teasing out correlations between the kinematics and substructure of jets, our method can reliably extract percent-level signal fractions. This strategy for uncovering new physics adds to the growing toolbox of anomaly detection methods for collider physics with a complementary set of assumptions. Published by the American Physical Society 2024

Astronomy & Astrophysics↗

Monitoring Sulfuric Acid and Temperature Using Raman Spectroscopy and Multivariate Chemometrics

Multivariate regression models were optimized for the quantification of sulfuric acid (H 2 SO 4 ) [0–8 M] and temperature (20 °C–80 °C) in the presence of ammonium sulfate ((NH 4 ) 2 SO 4 [0–0.6 M]) using Raman spectroscopy. Optical vibrational spectroscopy is a useful nondestructive technique for the in situ analysis of complex chemical systems notoriously difficult to monitor in situ and in real-time. Multivariate analysis, a chemometrics method, can be paired with these nondestructive optical methods for determining analyte concentration and speciation in complex solutions, such as dissociated species in polyprotic acids, e.g., H 2 SO 4 . The effect of temperature is often overlooked although it can have a major influence on speciation and the corresponding Raman spectra. Here, in this study, partial least squares regression models were optimized for the quantification of H 2 SO 4 and its two deprotonated forms as a function of temperature. Measuring bisulfate as a function of temperature is particularly challenging owing to changes in the second dissociation constant. A designed training set effectively minimized the sample set size and trained a robust predictive model with percent root mean square error of <3% for H 2 SO 4 . The practical strategy employed here was demonstrated to be effective for building chemometric models that directly account for dynamic temperatures with static samples and is shown to be amenable to flow cell analysis applications with a simple calibration transfer for process monitoring applications.

D-optimal design↗

Performance Metrics for the Assessment of Satellite Data Products: An Ocean Color Case Study

Performance assessment of ocean color satellite data has generally relied on statistical metrics chosen for their common usage and the rationale for selecting certain metrics is infrequently explained. Commonly reported statistics based on mean squared errors, such as the coefficient of determination (r2), root mean square error, and regression slopes, are most appropriate for Gaussian distributions without outliers and, therefore, are often not ideal for ocean color algorithm performance assessment, which is often limited by sample availability. In contrast, metrics based on simple deviations, such as bias and mean absolute error, as well as pair-wise comparisons, often provide more robust and straightforward quantities for evaluating ocean color algorithms with non-Gaussian distributions and outliers. This study uses a SeaWiFS chlorophyll-a validation data set to demonstrate a framework for satellite data product assessment and recommends a multimetric and user-dependent approach that can be applied within science, modeling, and resource management communities.

remote sensing↗

Case Study: Analysis of Autonomous Center line Tracking Neural Networks

Deep neural networks have gained widespread usage in a number of applications. However, limitations such as lack of explainability and robustness inhibit building trust in their behavior, which is crucial in safety critical applications such as autonomous driving. Therefore, techniques which aid in understanding and providing guarantees for neural network behavior are the need of the hour. In this paper, we present a case study applying a recently proposed technique, Prophecy, to analyze the behavior of a neural network model, provided by our industry partner and used for autonomous guiding of airplanes on taxi runways. This regression model takes as input an image of the runway and produces two outputs, cross-track error and heading error, which represent the position of the plane relative to the center line. We use the Prophecy tool to extract neuron activation patterns for the correctness and safety properties of the model. We show the use of these patterns to identify features of the input that explain correct and incorrect behavior. We also use the patterns to provide guarantees of consistent behavior. We explore a novel idea of using sequences of images (instead of single images) to obtain good explanations and identify regions of consistent behavior.

Deep Neural Networks↗

Power generation forecasting for solar plants based on Dynamic Bayesian networks by fusing multi-source information

A Dynamic Bayesian network (DBN) model for solar power generation forecasting in solar plants is proposed in this paper. The key idea is to fuse sensor data, operational indicators, meteorological data, lagged output power information, and model errors for more accurate short-term (e.g., hours) and mid-term (e.g., days to weeks) power generation forecasting. The proposed DBN augments automated data-driven structure learning with expert knowledge encoding using continuous and categorical data given constraints to represent causal relationships within a solar inverter system. Additionally, an error compensation mechanism is proposed to capture temporal fluctuation. The effectiveness of the DBN on solar power generation forecasting was evaluated by rolling window analysis with one-year testing data collected from a local solar plant. The proposed DBN is compared with four state-of-art methods including support-vector regression (SVR), k-nearest neighbors (kNN), artificial neural network (ANN), and long short-term memory (LSTM) models. The result show that the proposed DBN achieves better accuracy in general, and it is not as data-hungry as some neural network-based models. The proposed DBN is also shown to have robust and consistent forecasting power with different forecasting horizons. The accuracy is 92% - 95% from one hour to one week ahead forecasting.

14 SOLAR ENERGY↗