Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “robust regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Elliptically-Contoured Tensor-variate Distributions with Application to Image Learning

Statistical analysis of tensor-valued data has largely used the tensor-variate normal (TVN) distribution that may be inadequate for data arising from distributions with heavier or lighter tails. We study a general family of elliptically contoured (EC) TV distributions and derive its characterizations, moments, marginal, and conditional distributions. We describe procedures for maximum likelihood estimation from data that are (1) uncorrelated draws from an EC distribution, (2) from a scale mixture of the TVN distribution, and (3) from an underlying but unknown EC distribution, for which we extend Tyler’s robust estimator. A detailed simulation study highlights the benefits of choosing an EC distribution over the TVN for heavier-tailed data. We develop TV classification rules using discriminant analysis and EC errors and show that they better predict cats and dogs from images in the Animal Faces-HQ dataset than the TVN-based rules. A novel tensor-on-tensor regression and TV analysis of variance (TANOVA) framework under EC errors is also demonstrated to better characterize gender, age, and ethnic origin than the usual TVN-based TANOVA in the celebrated labeled faces of the wild dataset.

97 MATHEMATICS AND COMPUTING↗

Enhancing fire emissions inventories for acute health effects studies: integrating high spatial and temporal resolution data

Daily fire progression information is crucial for public health studies that examine the relationship between population-level smoke exposures and subsequent health events. Issues with remote sensing used in fire emissions inventories (FEI) lead to the possibility of missed exposures that impact the results of acute health effects studies. This paper provides a method for improving an FEI dataset with readily available information to create a more robust dataset with daily fire progression. High temporal and spatial resolution burned area information from two FEI products are combined into a single dataset, and a linear regression model fills gaps in daily fire progression. The combined dataset provides up to 71% more PM 2.5 emissions, 69% more burned area, and 367% more fire days per year than using a single source of burned area information. The FEI combination method results in improved FEI information with no gaps in daily fire emissions estimates. The combined dataset provides a functional improvement to FEI data that can be achieved with currently available data.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Leveraging Artificial Intelligence to Predict Novel Eutectic Alloys

The goal of this project was to train an artificial neural network (ANN) to predict the fractional composition and melting point of eutectic alloys using fundamental atomic properties as inputs. The fundamental properties considered include atomic number, atomic weight, atomic radius, valence electron concentration, electronegativity, and electron affinity. The project involved several phases, starting with data preparation, where phase diagram data was harvested from the ASM International database. Approximately 1300 binary eutectics were collected and cleaned to ensure relevance and accuracy. A regression model was selected for training, utilizing a rectified linear unit as the activation function. Various model configurations were evaluated for predictive accuracy, with validation techniques employed to ensure robustness. The model demonstrated predictive capabilities above random guessing and was able to achieve up to 11% accuracy under certain conditions. An ablative test identified atomic radius and valence electron concentration as critical inputs for model performance. Incorporating the melting point of atomic constituents improved accuracy significantly, although ultimately the model’s predictive capability still fell short of the 80% target. This report details the methodology, results, and implications of the research, contributing to the understanding of employing artificial intelligence to predict the phase transition behavior of eutectic alloys.

36 MATERIALS SCIENCE↗

Machine learning without a processor: Emergent learning in a nonlinear analog network

Standard deep learning algorithms require differentiating large nonlinear networks, a process that is slow and power-hungry. Electronic contrastive local learning networks (CLLNs) offer potentially fast, efficient, and fault-tolerant hardware for analog machine learning, but existing implementations are linear, severely limiting their capabilities. These systems differ significantly from artificial neural networks as well as the brain, so the feasibility and utility of incorporating nonlinear elements have not been explored. Here, we introduce a nonlinear CLLN—an analog electronic network made of self-adjusting nonlinear resistive elements based on transistors. We demonstrate that the system learns tasks unachievable in linear systems, including XOR (exclusive or) and nonlinear regression, without a computer. We find our decentralized system reduces modes of training error in order (mean, slope, curvature), similar to spectral bias in artificial neural networks. The circuitry is robust to damage, retrainable in seconds, and performs learned tasks in microseconds while dissipating only picojoules of energy across each transistor. This suggests enormous potential for fast, low-power computing in edge systems like sensors, robotic controllers, and medical devices, as well as manufacturability at scale for performing and studying emergent learning.

Science & Technology - Other Topics↗

When does global attention help: a unified empirical study on atomistic graph learning

Graph neural networks (GNNs) are widely used as surrogates for costly experiments and first-principles simulations to study the behavior of compounds at atomistic scale, and their architectural complexity is constantly increasing to enable the modeling of complex physics. While most recent GNNs combine more traditional message passing neural networks (MPNNs) layers to model short-range interactions with more advanced graph transformers (GTs) with global attention mechanisms to model long-range interactions, it is still unclear when global attention mechanisms provide real benefits over well-tuned MPNN layers due to inconsistent implementations, features, or hyperparameter tuning. We introduce the first unified, reproducible benchmarking framework–built on HydraGNN–that enables seamless switching among four controlled model classes: MPNN, MPNN with chemistry/topology encoders, GPS-style hybrids of MPNN with global attention, and fully fused localglobal models with encoders. Using seven diverse open-source datasets for benchmarking across regression and classification tasks, we systematically isolate the contributions of message passing, global attention, and encoder-based feature augmentation. Our study shows that encoder-augmented MPNNs form a robust baseline, while fused localglobal models yield the clearest benefits for properties governed by long-range interaction effects. We further quantify the accuracycompute trade-offs of attention, reporting its overhead in memory. Together, these results establish the first controlled evaluation of global attention in atomistic graph learning and provide a reproducible testbed for future model development.

Equivariant graph neural networks↗

Reduced-Order Modeling of Multigroup Neutron Cross Sections for High-Temperature Gas-cooled Reactors

Abstract – Deterministic neutronics calculations rely on multigroup neutron cross section libraries, which consist of databases of tabulated values, used to calculate the neutron cross sections through multivariate linear interpolation. However, interpolation of the multidimensional cross section data becomes memory inefficient and time consuming as the number of tabulations increases, significantly slowing down the neutronics calculation, especially in the case of microscopic cross section libraries where every isotope (on the order of hundreds) has its own set of specific reactions and cross sections. In order to address this challenge, this work constructs efficient and robust reduced-order models (ROMs) of the multi-group cross sections to support the Griffin simulation of high-temperature gas-cooled reactors (HTGRs). The first part of the study investigates the linearity of the multigroup cross section data across isotopes, reaction types, and energy groups on pre-generated datasets for the purpose of dimensionality reduction. Secondly, a down-selection of ROM techniques is presented on representative classical machine learning (ML) techniques, including variants of linear regression, kernel-based methods, tree-based algorithms, and artificial neural networks. The selection criteria jointly consider the memory efficiency, predictive accuracy, prediction speed, and scalability in comparison to the multidimensional interpolation. Among all the ML techniques, deep neural networks (DNNs) have proven to be the best selection with sufficient accuracy, high robustness, good memory efficiency, great scalability, and superior flexibility. DNNs have been trained for all isotopes in this work and systematic Griffin testing is ongoing to ensure the feasibility of this ROM technique for predicting cross section and reducing memory requirements without a significant sacrifice in computational performance.

42 - ENGINEERING↗

Advanced Cross Section Library Generation using Reduced Order Models

Deterministic neutronics calculations rely on multigroup neutron cross section libraries, which consist of databases of tabulated values, used to calculate the neutron cross sections through multivariate linear interpolation. However, interpolation of the multidimensional cross section data becomes memory inefficient and time consuming as the number of tabulations increases, significantly slowing down the neutronics calculation, especially in the case of microscopic cross section libraries where every isotope (on the order of hundreds) has its own set of specific reactions and cross sections. In order to address this challenge, this work constructs efficient and robust reduced-order models (ROMs) of the multi-group cross sections to support the Griffin simulation of high-temperature gas-cooled reactors (HTGRs). The first part of the study investigates the linearity of the multigroup cross section data across isotopes, reaction types, and energy groups on pre-generated datasets for the purpose of dimensionality reduction. Secondly, a down-selection of ROM techniques is presented on representative classical machine learning (ML) techniques, including variants of linear regression, kernel-based methods, tree-based algorithms, and artificial neural networks. The selection criteria jointly consider the memory efficiency, predictive accuracy, prediction speed, and scalability in comparison to the multidimensional interpolation. Among all the ML techniques, deep neural networks (DNNs) have proven to be the best selection with sufficient accuracy, high robustness, good memory efficiency, great scalability, and superior flexibility. DNNs have been trained for all isotopes in this work and systematic Griffin testing is ongoing to ensure the feasibility of this ROM technique for predicting cross section and reducing memory requirements without a significant sacrifice in computational performance.

42 - ENGINEERING↗

Reduce-Order Modeling of Multigroup Neutron Cross Sections for High-Temperature Gas-cooled Reactors

Deterministic neutronics calculations rely on multigroup neutron cross section libraries, which usually consists of a database of tabulated values, used to calculate the cross sections through multivariate linear interpolation. However, interpolation of the multidimensional cross section data becomes memory inefficient and time consuming as the number of tabulations increases, significantly slowing down the neutronics calculation, especially in the case of micro cross section libraries where every isotope (on the order of hundreds) has its own set of specific reactions and cross sections. To address this challenge, this work constructs efficient and robust reduced-order models (ROMs) of the multi-group cross sections to support the Griffin simulation of high-temperature gas-cooled reactors (HTGRs). The first part of the study investigates the linearity of the multi-group cross section data across isotopes, reaction types and energy groups on pre-generated datasets for the purpose of dimensionality reduction. Secondly, a down-selection of ROM techniques is presented on representative classical machine learning (ML) techniques, including variants of linear regression, kernel-based methods, tree-based algorithms, and artificial neural networks. The selection criteria jointly consider the memory efficiency, predictive accuracy, prediction speed, and scalability in comparison to the multidimensional interpolation. Among all the ML techniques, deep neural networks (DNNs) have proven to be the best selection with sufficient accuracy, high robustness, good memory efficiency, great scalability, and superior flexibility. DNNs for have been trained for all isotopes in this work and systematic Griffin testing is ongoing at this moment to ensure the feasibility of this ROM technique for cross section predictions.

42 - ENGINEERING↗

Using an Isotope Enabled Mass Balance to Evaluate Existing Land Surface Models

Abstract Land surface models (LSMs) play a crucial role in elucidating water and carbon cycles by simulating processes such as plant transpiration and evaporation from bare soil, yet calibration often relies on comparing LSM outputs of landscape total evapotranspiration ( ET ) and discharge with measured bulk fluxes. Discrepancies in partitioning into component fluxes predicted by various LSMs have been noted, prompting the need for improved evaluation methods. Stable water isotopes serve as effective tracers of component hydrologic fluxes, but data and model integration challenges have hindered their widespread application. Leveraging National Ecological Observation Network measurements of water isotope ratios at 16 US sites over 3 years combined with LSM‐modeled fluxes, we employed an isotope‐enabled mass balance framework to simulate ET isotope values ( δET ) within three operational LSMs (Mosaic, Noah, and VIC) to evaluate their partitioning. Models simulating δET values consistent with observations were deemed more reflective of water cycling in these ecosystems. Mosaic exhibited the best overall performance (Kling‐Gupta Efficiency of 0.28). For both Mosaic and Noah there were robust correlations between bare soil evaporation fraction and error (negative) as well as transpiration fraction and error (positive). We found the point at which errors are smallest ( x ‐intercept of the multi‐site regression) is at a higher transpiration fraction than is currently specified in the models. Which means that transpiration fraction is underestimated on average. Stable isotope tracers offer an additional tool for model evaluation and identifying areas for improvement, potentially enhancing LSM simulations and our understanding of land‐surface hydrologic processes.

58 GEOSCIENCES↗

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING↗

Taming nuclear mass models with Gaussian processes

We propose a new set of nuclear mass predictions based on multiple theoretical mass models. By employing Gaussian process regression with the Matérn kernel, we achieved root-mean-square (rms) deviations below 100 keV for the training dataset. The best-performing mass models achieved rms deviations below 150 keV for the new precise mass data from AME2020, whereas the ensemble average showed robust performance across the nuclear chart. Our approach uniquely combines: (1) systematic refinement of eight mass models through their residuals, (2) physics-informed features, including magic numbers, nucleon parity numbers, neutron excess, and nuclear collectivity, and (3) theory-to-theory validation demonstrating robust extrapolation capability. We find that the Matérn kernel provides superior uncertainty quantification compared to the RBF kernel, with a length-scale analysis revealing enhanced inter-nuclei correlations. We provide complete mass predictions for all unknown nuclides in AME2020, offering valuable constraints for nuclear structure studies and astrophysical modeling when used with proper uncertainty propagation.

Gaussian processes↗

Efficient Decision Trees for Tensor Regressions

Here, we proposed the tensor-input tree (TT) method for scalar-on-tensor and tensor-on-tensor regression problems. We first address scalar-on-tensor problem by proposing scalar-output regression tree models whose input variables are tensors (i.e., multi-way arrays). We devised and implemented fast randomized and deterministic algorithms for efficient fitting of scalar-on-tensor trees, making TT competitive against tensor-input GP models (Yu, Li, and Liu; Sun et al.). Based on scalar-on-tensor tree models, we extend our method to tensor-on-tensor problems using additive tree ensemble approaches. Theoretical justification and extensive experiments, including testing robustness to entrywise input tensor noise, are provided on real and synthetic datasets to illustrate the performance of TT. Our implementation is provided at https://github.com/hrluo/TensorDecisionTreeRegressor. Supplementary materials for this article are available online.

Decision tree regressions↗

Predicting the evolution of biomass bulk density through feedstock preprocessing: Discrete element modeling, regression analysis, and pilot-scale validation

Bulk density is an important material property of biomass feedstocks, influencing handling, storage, transport costs, and conversion efficiency. In this study, predictive regression models for loose and tapped bulk densities of Alamo and Cave-in-Rock switchgrass are developed using a comprehensive dataset generated via calibrated bonded-sphere discrete element method (DEM) simulations. Here, a key contribution of this study is the use of a DEM-based approach, which correlates density with moisture content and particle size distribution parameters and enables analysis across a continuous particle size range, overcoming limitations of purely experimental data. For comparison, regression models are also developed using only experimental data from pilot-scale runs at the Biomass Feedstock National User Facility at Idaho National Laboratory. Validation against pilot-scale data showed reasonable prediction accuracy for both model types, particularly for smaller particle sizes (post-secondary grinding). While the experimental model showed slightly better performance matching the validation data in some cases, the DEM-based model benefits from a much larger dataset, reduced predictor multicollinearity, and continuous parameter coverage, highlighting the utility of validated simulation models for developing robust predictive tools for biomass preprocessing applications.

09 - BIOMASS FUELS↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

Prediction of electric and magnetic fields from spectral data using machine learning algorithms for Doppler-free saturation spectroscopy diagnostics

The prediction of electric and magnetic field amplitudes from atomic spectral data is critical for plasma control in fusion devices such as tokamaks. Conventional approaches that rely on physics-based models are computationally expensive and unsuitable for real-time applications. In this work, we develop and benchmark three machine learning algorithms—simulation-based inference (SBI), fully connected neural networks (FCNN), and histogram-based gradient boosting regression (GBR-Hist)—to infer field intensities directly from Doppler-free saturation spectroscopy (DFSS) spectra. Synthetic datasets of spectra were generated using the EZSSS code and evaluated both with and without added Poisson noise to mimic experimental conditions. We find that SBI achieves the highest accuracy and robustness, FCNN provides a strong balance of accuracy and computational efficiency for real-time applications, and GBR-Hist offers the fastest inference but is more sensitive to noise. Furthermore, these results demonstrate the potential of machine learning to accelerate DFSS analysis and enhance its utility for plasma diagnostics and control.

Doppler-free saturation spectroscopy↗

Observations of Coastal Wind Momentum Flux: Dependence on Fetch and Waves with Comparisons to COARE

Here, using observations from the Martha’s Vineyard Coastal Observatory, this paper investigates how momentum flux in the marine atmospheric surface layer over the coastal ocean varies with fetch, wave age, and wave slope and assesses the performance of the COARE 3.5 bulk parameterization. Long-fetch (at least 300 km) and short-fetch (3–6 km away from land) conditions have very similar momentum flux, with the latter being just 15% higher. The COARE 3.5 wind speed–dependent formulation closely matches the observations. The sea state dependence of wave age and wave slope is analyzed by considering both peak frequency and mean frequency in a wave spectrum. The slope of regression lines between normalized roughness and wave age is sensitive to wind speed ranges and the scatter of momentum flux, potentially explaining why earlier studies did not find a universal formula to characterize the wave dependence. Although the observed momentum flux exhibits an obvious dependence on wave age, a robust quadratic fit between momentum flux and the 10-m neutral wind speed exists only for young waves. For the short-fetch conditions, the momentum flux does not increase as wave slope increases. This may be because waves generated by local wind are still weak and swell that propagates from other areas dominates the wave spectra. In other words, the wave slope computed by considering the spectra does not properly reflect the local wind–wave interaction.

Air-sea interaction↗

Performance of a dynamic single bubbler in single and two-phase immiscible liquids

Ensuring nonproliferation and safeguards of special nuclear materials (SNM) is a critical aspect of advancing the nuclear fuel cycle. Traditional bubbler systems used to estimate liquid levels and densities in nuclear recycling processes have limitations, particularly in harsh environments where dip-tube corrosion and buildup necessitate frequent maintenance and recalibration. This study explores the Dynamic Single Bubbler (DSB) method, which utilizes a single dip-tube attached to a linear actuator to estimate liquid properties dynamically. This approach is extended to estimate liquid-liquid interfaces in immiscible liquids and employs a linear regression method to reduce uncertainties and improve accuracy. The DSB method achieved density estimate uncertainties of less than 0.5% and surface level estimate uncertainties typically under 0.5%, across various fluids including water, acetone, methanol, mineral oil, glycerol, and aqueous salt solutions. Results indicate that the DSB method provides accurate and robust estimates of liquid density and surface levels with minimal maintenance and without the need for calibration. Additionally, the method's applicability to immiscible liquids and various dip-tube geometries was demonstrated, showing promise for widespread use in nuclear and other industrial applications.

98 - NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL↗

Fundamental Insights into Cathode Stability: Linking Compositional Tuning and Local Coordination in Complex Metal Oxides under Aqueous Transformations

Compositional tuning of complex metal oxides in Li-ion battery materials influences their performance as well as their end-of-life behavior, in particular, the tendency to release toxic metal cations in aqueous solution. We modeled ternary variants of a parent LiCoO 2 delafossite structure by varying the metal identity and relative amounts. This yielded ten model formulations of Li(A 4/6 B 1/6 C 1/6 )O 2 , where the material is enriched with the A metal and doped with B and C, with Ni, Mn, Co, Fe, Al, V, and Ti as constituent metals. To assess their stability in aqueous conditions, metal release energetics were calculated using a combination of Density Functional Theory calculations and thermodynamics. Metal release in ternary oxides is dictated by subtle variations in the coordination environment of the leaving group. To identify governing chemical features across diverse compositions with varying local coordination environments, we leverage random forest regression and descriptor importance analysis. A key result is that metal–oxygen orbital hybridization, quantified using a projected density-of-states-derived descriptor, H d/p , provides a physically grounded measure of interaction strength that governs metal release energetics. This refined perspective goes beyond conventional oxidation state considerations and offers more robust insights for materials science. Finally, we model defect surface-bound O 2 dimer formation as a proxy for reactive oxygen species (ROS) generation. The results show that Ni-rich compositions more readily stabilize spin-polarized O 2 dimers, corroborating experimental reports of an increased ROS-driven biological response. In conclusion, our results establish a compositional and electronic basis for metal release and surface oxygen reactivity that form a rationale for complex metal oxide design principles.

36 MATERIALS SCIENCE↗