Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “robust regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Predicting Damages to Remainder Parcels in Right-of-Way Acquisitions for Expanding Transportation Infrastructure: Using a Truncated Finite-Mixture Model

Right-of-way acquisition is a critical component of transportation infrastructure development. Transportation infrastructure projects cannot proceed without proper right-of-way acquisition or may face significant delays. State Departments of Transportation frequently acquire parcels of land for roadway expansion projects. A majority of these acquisitions can be partial takings, referring to a portion of a parcel that is acquired. The remainder of the property usually suffers economic changes due to the partial acquisition, which can be calculated as damage percentages. The damage percentage represents the extent to which the remaining land or property value has been diminished due to the acquisition. It reflects the remaining property value percentage that may have been lost or compromised due to the acquisition. Here, this study aims to provide a robust model to estimate damage percentages to the remainder parcels that may help state Departments of Transportation appraisers make early predictions about the damages in cases involving partial takings. The research uses 509 appraisal reports from the Tennessee Department of Transportation to identify the key parcel attributes that influence the percentage of damages. Three regression models are developed: a linear regression model, a finite-mixture model (FMM), and a truncated FMM with two latent classes. The modeling results show that the truncated FMM with two classes outperforms the other models. To validate the models, actual sales data is collected and analyzed for 59 properties, and the results suggest that the model predictions are fairly accurate. A predictive tool is developed based on the models to help appraisers anticipate right-of-way damages under different scenarios and can provide early predictions about the damages.

42 ENGINEERING↗

Steady (robust) conditionally effective estimation of parameters

The concept of conditionally-effective estimation which provides optimum estimates for a given criterion in cases of given limitations is examined. The concept has the best accuracy for given limitation on the suitability of the algorithm concerning the deviation of the law governing error distribution from the proposed law. It is concluded that there must be a greater difference between individual algorithms in terms of difficulty, and that for the linear regression problem algorithms based on excluding lost points should be studied.

Gurin, L. S.↗

Mapping Rare Earths and Toxics in E-Waste via Hyperspectral Imaging and Machine Learning

Electronic waste (e-waste) presents a mounting challenge to environmental sustainability due to its complex composition, which includes high-value rare earth elements, hazardous organic compounds, and non-recyclable plastics. Accurate and scalable material classification is essential for enabling efficient resource recovery and safe recycling practices. This study introduces a confidence-aware classification pipeline that combines mid-infrared hyperspectral imaging (HSI), spectral angle mapping (SAM), and iterative machine learning to perform pixel-level material identification across e-waste devices. A curated spectral library encompassing artificial materials (e.g., plastic iron oxide, galvanized metals), minerals (e.g., allanite, hematite), and organic compounds (e.g., benzanthracene, toluene) was used to generate pseudo-labels, each assigned a confidence score based on SAM-derived spectral similarity. High-confidence samples from seven consumer electronics—digital cameras, keyboards, laptop fans, modems, motherboards, TV remotes, and speakers—were iteratively expanded and classified using models such as Support Vector Machine (SVM), Random Forest, Gradient Boosting Classifier, Partial Least Squares Discriminant Analysis (PLSDA) and Logistic Regression. The best-performing classifiers achieved macro F1 scores approaching 1.0. Results revealed widespread plastic content (dominated by plastic iron oxide), the presence of rare earth-bearing minerals like cerium-containing allanite, and pervasive detection of hazardous organics such as benzanthracene. Principal Component Analysis (PCA) visualizations and confusion matrices confirmed high separability and robust classification performance. This methodology enables precise, non-destructive, and scalable classification of heterogeneous e-waste streams. It supports automated, hazard-aware sorting in recycling workflows, facilitating selective recovery of critical materials and compliance with circular economy goals. The confidence-aware framework provides a foundation for real-time deployment in industrial settings, offering significant implications for smart e-recycling infrastructure and policy-driven material stewardship.

Circular economy↗

Can Selforganizing Maps Accurately Predict Photometric Redshifts?

We present an unsupervised machine-learning approach that can be employed for estimating photometric redshifts. The proposed method is based on a vector quantization called the self-organizing-map (SOM) approach. A variety of photometrically derived input values were utilized from the Sloan Digital Sky Survey's main galaxy sample, luminous red galaxy, and quasar samples, along with the PHAT0 data set from the Photo-z Accuracy Testing project. Regression results obtained with this new approach were evaluated in terms of root-mean-square error (RMSE) to estimate the accuracy of the photometric redshift estimates. The results demonstrate competitive RMSE and outlier percentages when compared with several other popular approaches, such as artificial neural networks and Gaussian process regression. SOM RMSE results (using delta(z) = z(sub phot) - z(sub spec)) are 0.023 for the main galaxy sample, 0.027 for the luminous red galaxy sample, 0.418 for quasars, and 0.022 for PHAT0 synthetic data. The results demonstrate that there are nonunique solutions for estimating SOM RMSEs. Further research is needed in order to find more robust estimation techniques using SOMs, but the results herein are a positive indication of their capabilities when compared with other well-known methods

Way, Michael J.↗

Bayesian Model Selection for Reducing Bloat and Overfitting in Genetic Programming for Symbolic Regression

When performing symbolic regression using genetic programming, overfitting and bloat can negatively impact generalizability and interpretability of the resulting equations as well as increase computation times. A Bayesian fitness metric is introduced and its impact on bloat and overfitting during population evolution is studied and compared to common alternatives in the literature. The proposed approach was found to be more robust to noise and data sparsity in numerical experiments, guiding evolution to a level of complexity appropriate to the dataset. Further evolution of the population resulted not in overfitting or bloat, but rather in slight simplifications in model form. The ability to identify an equation of complexity appropriate to the scale of noise in the training data was also demonstrated. In general, the Bayesian model selection algorithm was shown to be an effective means of regularization which resulted in less bloat and overfitting when any amount of noise was present in the training data.

G F Bomarito↗

The CAI Database: 26 Al– 26 Mg Isotope Systematics

We present a publicly available calcium–aluminum-rich inclusion (CAI) database that focuses on the initial 26 Al/ 27 Al 0 ratio in CAIs, designed in a way that researchers in cosmochemistry and astrophysics may find useful. To date, the database contains 497 CAIs from 75 peer-reviewed papers. The CAIs are from all chondrite groups and cover different CAI types, textures, and sizes. The database includes the paper; the host meteorite; the CAI name and type; the 26 Al/ 27 Al 0 , δ 26 Mg$^*_0$, and δ 25 Mg values and their uncertainties; the number of regression points; the maximum 27 Al/ 24 Mg; the mean-squared weighted deviation; the CAI size; and CAI descriptions. We grouped the CAIs in different ways to discuss 26 Al/ 27 Al 0 ratio distributions with implications for the CAI formation timeline. Overall, we agree with previous authors that CAIs have a bimodal 26 Al distribution: CAIs with robust isochrons (n = 151) have a median 26 Al/ 27 Al 0 = 4.8 × 10 −5 (with a 1σ standard error of 0.1), while those with isotopic anomalies (n = 87) have a median 26 Al/ 27 Al 0 = 0.3 × 10 −5 (with a 1σ standard error of 0.2). However, the large standard deviation of both groups (1.3 and 2.3, respectively) indicates that the 26 Al/ 27 Al 0 values scatter significantly within each population. CAI types and groups can have distinct 26 Al/ 27 Al 0 and δ 26 Mg$^*_0$, but the unmelted inclusions (n = 33) have the highest median 26 Al/ 27 Al 0 = 5.1 × 10 −5 and a low median δ 26 Mg$^*_0$ = −0.05‰. We find slightly different 26 Al/ 27 Al 0 distributions between CAI chondrite types, but no differences between petrographic types or sizes. These observations can help us to understand CAI formation in the context of astrophysical models.

Astronomy and AstroPhysics↗

STag. II. Classification of Serendipitous Supernovae Observed by Galaxy Redshift Surveys

With the number of supernovae observed expected to drastically increase thanks to large-scale surveys like the Dark Energy Spectroscopic Instrument (DESI), it is necessary that the tools we use to classify these objects keep up with this increase. We previously created Supernova Tagging and Classification (STag) to address this problem by employing machine learning techniques alongside logistic regression in order to assign “tags” to spectra based on spectral features. STag II is a continuation of this work, which now makes use of model supernova spectra combined with real DESI spectra in order to train STag to better deal with realistic data. Furthermore, we also make use of the rlap score as a trustworthiness cut, making for a more robust and accurate supernova classifier than before.

Astrostatistics techniques↗

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

Remote Americium Detection Using an Optical Sensor: A D-Optimal Strategy for Efficient PLS-Based Modeling

A fiber-optic visible–near-infrared absorption spectroscopy system in a glove box was demonstrated for remote quantification of Am(III) (0–500 µM) and HNO 3 (0.1–9 M) using partial least squares regression (PLSR) models. The sensor platform, featuring a simple plug-and-play spectrophotometer, can enable noninvasive, real-time monitoring of actinide process solutions. To establish a flexible PLSR model calibration strategy, a D-optimal design developed using Nd(III) in previous studies was successfully extended to an actinide system with Am(III) to effectively minimize sample set size while maintaining robust prediction performance. The results suggest strong spectral similarities between Nd(III) and Am(III) and validate Nd(III) as an effective optical surrogate for trivalent actinide species. This work also supports the generalizability of a D-optimal training set selection approach for two-factor systems. The PLS1 models for Am(III) and HNO 3 outperformed a PLS2 model and maintained reasonable performance in the presence of interfering U(VI). The resulting sensor system and multivariate approach provides a flexible and scalable solution for process monitoring, control, and safety in diverse nuclear applications.

actinide↗

Gaussian processes for inferring parton distributions

The extraction of parton distribution functions (PDFs) from experimental or lattice QCD data is an ill-posed inverse problem, where regularization strongly impacts both systematic uncertainties and the reliability of the results. We study a framework based on Gaussian Process Regression (GPR) to reconstruct PDFs from lattice QCD matrix elements. Within a Bayesian framework, Gaussian processes serve as flexible priors that encode uncertainties, correlations, and constraints without imposing rigid functional forms. We investigate a wide range of kernel choices, mean functions, and hyperparameter treatments. We quantify information gained from the data using the Kullback-Leibler divergence. Synthetic data tests demonstrate the consistency and robustness of the method. Our study establishes GPR as a systematic and non-parametric approach to PDF reconstruction, offering controlled uncertainty estimates and reduced model bias in lattice QCD analyses.

hadronic spectroscopy↗

Pu(IV) quantification via visible–near-infrared absorption spectroscopy: tackling interferences using D-optimal design and partial least squares

Here, this study presents a novel analytical approach for quantifying Pu(IV) in glove box environments using fiber-optic-based visible–near-infrared absorption spectroscopy in combination with partial least squares regression (PLSR) and design of experiments. The method addresses significant challenges posed by overlapping spectral features arising from Nd(III), which is a common fission product impurity, and the speciation variability of Pu(IV) nitrato complexes in HNO 3 concentrations ranging from 2.5 to 11 M. A curated training set consisting of data from 20 samples was developed via D-optimal design to enable robust PLSR model calibration for Pu(IV) using the near-infrared band near 1050 nm. The training set was acquired from samples in cuvettes with a 1-cm path length and was used to build the PLSR model. The robustness of the model was validated with data collected using a dip probe with a 1-cm path length and varying Pu(IV) concentrations. The strong performance of the model indicates good model transfer from cuvette to dip probe and highlights the potential for in situ measurements and online monitoring of reactions in a crystallization reactor vessel. The results demonstrate that this combined spectroscopic and chemometric approach can accurately and simultaneously quantify Pu(IV) and HNO 3 , thereby offering a promising tool for real-time monitoring in process environments.

Actinide↗

Sparse Cholesky factorization for solving nonlinear PDEs via Gaussian processes

In recent years, there has been widespread adoption of machine learning-based approaches to automate the solving of partial differential equations (PDEs). Among these approaches, Gaussian processes (GPs) and kernel methods have garnered considerable interest due to their flexibility, robust theoretical guarantees, and close ties to traditional methods. They can transform the solving of general nonlinear PDEs into solving quadratic optimization problems with nonlinear, PDE-induced constraints. However, the complexity bottleneck lies in computing with dense kernel matrices obtained from pointwise evaluations of the covariance kernel, and its partial derivatives, a result of the PDE constraint and for which fast algorithms are scarce. The primary goal of this paper is to provide a near-linear complexity algorithm for working with such kernel matrices. We present a sparse Cholesky factorization algorithm for these matrices based on the near-sparsity of the Cholesky factor under a novel ordering of pointwise and derivative measurements. The near-sparsity is rigorously justified by directly connecting the factor to GP regression and exponential decay of basis functions in numerical homogenization. We then employ the Vecchia approximation of GPs, which is optimal in the Kullback-Leibler divergence, to compute the approximate factor. This enables us to compute ϵ-approximate inverse Cholesky factors of the kernel matrices with complexity O(N log d (N/ϵ)) in space and O(N log 2d (N/ϵ)) in time. We integrate sparse Cholesky factorizations into optimization algorithms to obtain fast solvers of the nonlinear PDE. We numerically illustrate our algorithm’s near-linear space/time complexity for a broad class of nonlinear PDEs such as the nonlinear elliptic, Burgers, and Monge-Ampère equations. In summary, we provide a fast, scalable, and accurate method for solving general PDEs with GPs and kernel methods.

97 MATHEMATICS AND COMPUTING↗

Online LIBS–ML Framework for Dynamic Characterization of Heterogeneous Waste-Derived Gasification Feedstocks

LIBS−ML framework for real time feedstock characterization during continuous conveyor transport Heterogeneous waste derived feedstocks (e.g., waste coal, biomass and blends) introduce rapid variability in heating value and ash chemistry that affect gasifier operation, yet conventional laboratory characterization techniques are too slow to support proactive control. To address this gap, this study reports on an online, in situ, dynamic characterization framework that couple’s laser-induced breakdown spectroscopy (LIBS) with leakage safe machine learning (ML) regression to deliver real time, decision quality predictions of gasifier relevant properties. A controlled sample matrix spanning two different waste coals, two different biomasses, and engineered blends under two particle size conditions were constructed and benchmarked using standardized laboratory analyses for proximate/ultimate properties and ash composition. LIBS spectra were acquired dynamically as material flowed on a conveyor belt, using high energy 1064 nm laser ablation and shot averaging to improve repeatability and precision. Supervised regression models (multi layer perceptron (MLP) /artificial neural network (ANN), random forest (RF), and support vector regression (SVR)) and an optimized weighted ensemble were trained on emission line feature sets using nested cross validation with Bayesian hyperparameter tuning and validated against an independent hold out set. The proposed LIBS−ML workflow achieves near laboratory predictive fidelity across parametric targets (including higher heating value (HHV), ash content, fixed carbon, sulfur, major ash forming oxides, and initial deformation temperature (IDT)), with the weighted ensemble providing a robust default predictor under dynamic measurement conditions. These results demonstrate a practical pathway for real time feedstock characterization that can enable feedforward adjustments and more resilient gasifier operation for variable quality waste derived fuels.

Biomass↗

Predicting initial trans-membrane pressure across cycles in the ultrafiltration process using random forest

With growing freshwater scarcity, direct potable reuse (DPR) systems that reclaim wastewater for drinking are becoming increasingly important for sustainable water supply. Reliable operation requires minimizing downtime in ultrafiltration (UF) units, where membrane fouling leads to elevated trans-membrane pressure (TMP). This study develops data-driven regression models based on random forest (RF) and autoregressive (AR) approaches to forecast the initial TMP at the start of each UF filtration cycle in a pilot-scale DPR system. The RF model consistently outperforms baseline methods, including historical mean, last observation carried forward, and AR models, across multiple forecast horizons, achieving the lowest root mean square error. To evaluate how different classes of process variables contribute to TMP dynamics over time, we examine the feature importance of independent input variables across multiple forecast horizons. This analysis provides insight into the temporal relevance of operational and sensor-derived features, guiding control and monitoring strategies. Additionally, the impact of hyperparameter tuning on TMP prediction performance is assessed for both direct and recursive RF modelling approaches. The proposed RF framework establishes a robust foundation for predictive monitoring and real-time optimization of UF operations, supporting sustainable and reliable water reuse.

direct potable reuse↗

Data-Efficient Methods for Determining Flory–Huggins χ Parameters in Multicomponent Polymer Formulations

Polymer formulations are essential in diverse applications including personal care products, coatings, paints, adhesives, and plastic materials. Designing these formulations requires navigating large, complex design spaces, where phase and self-assembly behavior critically impact performance. The Flory–Huggins χ parameter, which quantifies segmental miscibility, is widely used to parametrize the excess free energy of mixing in formulation models. In this work, we introduce two data-efficient, top-down methods for estimating χ parameters using the Random Phase Approximation (RPA): (i) Boundary Nonlinear Regression (Boundary-NLR), which fits theoretical spinodal boundaries to experimental phase boundaries, and (ii) Surrogate Model Inverse Parameter Estimation (SMIPE), which uses a Gaussian Process Classifier to fit sparse phase maps via a surrogate model. Both methods allow rapid parametrization of polymer field-theoretic models without the need for additional experiments. We evaluate these approaches on data sets involving polymer–solvent–nonsolvent ternary mixtures and block copolymer–solvent systems, demonstrating their robustness to experimental noise and their relevance for real-world formulation design.

copolymers↗

The empirical accuracy of uncertain inference models

Uncertainty is a pervasive feature of the domains in which expert systems are designed to function. Research design to test uncertain inference methods for accuracy and robustness, in accordance with standard engineering practice is reviewed. Several studies were conducted to assess how well various methods perform on problems constructed so that correct answers are known, and to find out what underlying features of a problem cause strong or weak performance. For each method studied, situations were identified in which performance deteriorates dramatically. Over a broad range of problems, some well known methods do only about as well as a simple linear regression model, and often much worse than a simple independence probability model. The results indicate that some commercially available expert system shells should be used with caution, because the uncertain inference models that they implement can yield rather inaccurate results.

Vaughan, David S.↗

Gila Water Resources III - Modeling the Impacts of Post-fire Restoration Methods on Vegetation Recovery in the Gila National Forest

In recent years, wildfires in New Mexico’s Gila National Forest have become increasingly common and more severe. Wildfires can have powerful impacts on hydrology and soil stability, including erosion, flooding, and debris-flows that threaten lives and infrastructure downstream. Vegetation restoration treatments like seeding and mulching can mitigate these effects and facilitate ecosystem recovery. Understanding the effectiveness of various restoration methods is vital to planning a cost-effective and successful post-fire recovery strategy. The immediate response to a fire on US Forest Service land is coordinated by a Burned Area Emergency Response (BAER) team, a group responsible for mitigating immediate post-fire risks to human life, property, and critical natural and cultural resources. This study created a proof-of-concept methodology for a decision-support tool designed to help BAER teams identify the restoration treatments most likely to succeed in a given burned area. Leveraging random forest regression, Google Earth Engine, and Landsat 7 and 8 Earth observations, this study modeled vegetation recovery after the 2013 Silver Fire for seeded areas, seeded/mulched areas, and untreated areas. Treatment type and initial burn severity were the largest drivers of vegetation recovery across the landscape. Seeded/mulched areas showed higher recovery levels than untreated areas three months post-fire, but by four years post-fire, treated and untreated areas displayed similar recovery levels. To produce a robust predictive tool for the Gila National Forest, the model should be trained on many more fires and incorporate post-fire weather conditions into the process. Such a model will help partners ensure efficient resource use and plan effective post-fire restoration strategies.

DEVELOP Project Summary↗

Predicting Initial Trans-Membrane Pressure for Optimized Operations in UF Unit Using Random Forest

With the growing scarcity of freshwater, innovative process design mechanisms like Ultra-filtration(UF) units are increasingly gaining attention among water treatment utilities to address the rising demand. Ensuring reliable water production necessitates efficient resource utilization, minimizing downtime in UF systems. Recent advancements in machine learning (ML) have enabled the development of accurate data-driven models for Model Predictive Control (MPC), often requiring minimal prior knowledge of underlying physical processes. In this study, we present predictive regression models based on Random Forest (RF) and Auto-Regressive (AR) approaches to forecast the initial Trans-Membrane Pressure (TMP) for each filtration cycle in data generated by Direct Potable Reuse (DPR) systems. The proposed RF-based model demonstrates superior performance compared to baseline methods, including historical mean, Last Observation Carried Forward (LOCF), and naïve AR models, across various forecasting horizons in terms of root mean square (RMSE) metric. Accurate prediction of initial TMP is critical for optimizing CCRO operations, as it enables the development of robust modelling frameworks that enhance process efficiency and reliability. The demonstrated efficacy of the RF-based approach highlights its potential as a tool for real-time decision-making in water treatment systems, paving the way for advanced process optimization and sustainable water resource management.

Mukherjee, Subrata [ORNL] (ORCID:0000000309930338)↗