Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Metrics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

PeakQC: A Software Tool for Omics-Agnostic Automated Quality Control of Mass Spectrometry Data

Mass spectrometry is broadly employed to study complex molecular mechanisms in various biological and environmental fields, enabling 'omics' research such as proteomics, metabolomics, and lipidomics. As study cohorts grow larger and more complex with dozens to hundreds of samples, the need for robust quality control (QC) measures through automated software tools becomes paramount to ensure the integrity, high quality, and validity of scientific conclusions from downstream analyses and minimize the waste of resources. Since existing QC tools are mostly dedicated to proteomics, automated solutions supporting metabolomics are needed. To address this need, we developed the software PeakQC, a tool for automated QC of MS data that is independent of omics molecular types (i.e., omics-agnostic). It allows automated extraction and inspection of peak metrics of precursor ions (e.g., errors in mass, retention time, arrival time) and supports various instrumentations and acquisition types, from infusion experiments or using liquid chromatography and/or ion mobility spectrometry front-end separations and with/without fragmentation spectra from data-dependent or independent acquisition analyses. Diagnostic plots for fragmentation spectra are also generated. Here, in this paper, we describe and illustrate PeakQC’s functionalities using different representative data sets, demonstrating its utility as a valuable tool for enhancing the quality and reliability of omics mass spectrometry analyses.

47 OTHER INSTRUMENTATION↗

Rapid quantitative analysis of trace elements in plutonium alloys using a handheld laser-induced breakdown spectroscopy (LIBS) device coupled with chemometrics and machine learning

Here, we present the first reported quantification of trace elements in plutonium via a portable laser-induced breakdown spectroscopy (LIBS) device and demonstrate the use of chemometric analysis to enhance the handheld device's sensitivity and precision. Quantification of trace elements such as iron and nickel in plutonium metal via LIBS is a challenging problem due to the complex nature of the plutonium optical emission spectra. While rapid analysis of plutonium alloys has been demonstrated using portable LIBS devices, such as the SciAps Z300, their detection limits for trace elements are severely constrained by their achievable pulse power and length, light collection optics, and detectors. In this paper, analytical methods are evaluated as a means to circumvent the detection constraints. Three chemometric methods often used in analytical spectroscopy are evaluated; principal component regression, partial least-squares regression, and artificial neural networks. These models are evaluated based on goodness-of-fit metrics, root mean-squared error, and their achievable limits of detection (LoDs). Partial least squares proved superior for determining content of iron and nickel in plutonium metal, yielding LoDs of 15 and 20 ppm, respectively. These results of identifying the undesirable trace elements in plutonium components are critical for applications such as fabricating radioisotope thermoelectric generators or nuclear fuel.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A comparison of deep-learning-based inpainting techniques for experimental X-ray scattering

The implementation is proposed of image inpainting techniques for the reconstruction of gaps in experimental X-ray scattering data. The proposed methods use deep learning neural network architectures, such as convolutional autoencoders, tunable U-Nets, partial convolution neural networks and mixed-scale dense networks, to reconstruct the missing information in experimental scattering images. In particular, the recovered pixel intensities are evaluated against their corresponding ground-truth values using the mean absolute error and the correlation coefficient metrics. The results demonstrate that the proposed methods achieve better performance than traditional inpainting algorithms such as biharmonic functions. Overall, tunable U-Net and mixed-scale dense network architectures achieved the best reconstruction performance among all the tested algorithms, with correlation coefficient scores greater than 0.9980.

97 MATHEMATICS AND COMPUTING↗

Effective Missing Value Imputation Methods for Building Monitoring Data

To understand behaviors of natural and man-made events, such as energy consumption of buildings, which accounts for 40% of energy uses in the US, we deploy automated monitoring devices to record periodic observations. However, such experimental and observation data often contains problems and irregularities that have to be cleaned up before analyses. Due to various conditions affecting sensor operations, the communication channels, recording steps, or the recording media, the recorded data might have missing values, errors, or anomalous values. An effective way to clean up these problems is to replace these missing values, errors and anomalous values with expected values, a process generally known as imputation. In this work, we survey commonly used missing value imputation techniques and compare their performance on a set of building monitoring data. To compare the different types of sensor measurements with widely varying characteristics, we use normalized root mean squared error (NRMSE) as the key metric for the effectiveness of the imputation methods. We additionally consider periodicity and run time when considering comparing methods. Through extensive testing, we find that for small gap sizes, up to 8 consecutive missing values, linear interpolation performs the best; for larger gaps stretching up to 48 consecutive missing values, K-nearest neighbors provides the most accurate imputations; for even larger gaps, more computational intensive methods, such as matrix factorization, achieve the smallest NRMSE. Additionally, we observe that these computationally intensive algorithms not only provide accurate imputations for large gaps, but are also more robust across all types of sensors.

Cho, B↗

Visual Analytics of Performance of Quantum Computing Systems and Circuit Optimization

Driven by potential exponential speedups in business, security, and scientific scenarios, interest in quantum computing is surging. This interest feeds the development of quantum computing hardware, but several challenges arise in optimizing application performance for hardware metrics (e.g., qubit coherence and gate fidelity). In this work, we describe a visual analytics approach for analyzing the performance properties of quantum devices and quantum circuit optimization. Our approach allows users to explore spatial and temporal patterns in quantum device performance data and it computes similarities and variances in key performance metrics. Detailed analysis of the error properties characterizing individual qubits is also supported. We also describe a method for visualizing the optimization of quantum circuits. The resulting visualization tool allows researchers to design more efficient quantum algorithms and applications by increasing the interpretability of quantum computations.

Chae, Junghoon↗

Designing a parallel Feel-the-Way clustering algorithm on HPC systems

This paper introduces a new parallel clustering algorithm, named Feel-the-Way clustering algorithm, that provides better or equivalent convergence rate than the traditional clustering methods by optimizing the synchronization and communication costs. Our algorithm design centers on how to optimize three factors simultaneously: reduced synchronizations, improved convergence rate, and retained same or comparable optimization cost. To compare the optimization cost, we use the Sum of Square Error (SSE) cost as the metric, which is the sum of the square distance between each data point and its assigned clusters. Compared with the traditional MPI k-means algorithm, the new Feel-the-Way algorithm requires less communications among participating processes. As for the convergence rate, the new algorithm requires fewer number of iterations to converge. As for the optimization cost, it obtains the SSE costs that are close to the k-means algorithm. In the paper, we first design the full-step Feel-the-Way k-means clustering algorithm that can significantly reduce the number of iterations that are required by the original k-means clustering method. Next, we improve the performance of the full-step algorithm by adopting an optimized sampling-based approach, named reassignment-history-aware sampling. Our experimental results show that the optimized sampling-based Feel-the-Way method is significantly faster than the widely used k-means clustering method, and can provide comparable optimization costs. More extensive experiments with several synthetic datasets and real-world datasets (e.g., MNIST, CIFAR-10, ENRON, and PLACES-2) show that the new parallel algorithm can outperform the open source MPI k-means library by up to 110% on a high-performance computing system using 4,096 CPU cores. In addition, the new algorithm can take up to 51% fewer iterations to converge than the k-means clustering algorithm.

97 MATHEMATICS AND COMPUTING↗

Validation of the High-Resolution Salish Sea Tidal Hydrodynamic Model

In this study, a tidal hydrodynamic model was developed and validated to simulate tidal currents in Puget Sound, Washington, to support tidal energy resource characterization using the unstructured-grid, Finite Volume Community Ocean Model (FVCOM). The Salish Sea tidal hydrodynamic model was driven by tides along two open boundaries at the entrance of the Strait of Juan de Fuca and north end of Georgia Strait, and river flows from 19 major rivers in the Salish Sea. To simulate the tidal current in Puget Sound, a high-resolution model grid is required to accurately represent the complex coastlines and bathymetry. The spatial resolution of the model grid varies from ~10 m near river boundaries and ~30 m in small tidal channels and estuaries to near 1000 m inside Georgia Strait and at the open boundaries. Model validation was carried out by comparing simulated and observed water levels at 12 tidal stations and currents at 135 Acoustic Doppler Current Profiler stations in the model domain. A set of model performance metrics, including root mean square error, scatter index, bias, and linear correlation coefficient, were used to quantify the model skills in simulating the tidal hydrodynamics in Puget Sound. Error statistics showed an overall good agreement between simulated and observed tidal elevations and currents, which demonstrated that the Puget Sound tidal model can be used to accurately characterize the tidal stream energy resource in Puget Sound.

16 TIDAL AND WAVE POWER↗

Enter Gaussian Mixture Modeling Extensions for Improved False Discovery Rate Estimation in GC-MS Metabolomics

Identifying small molecules (e.g., metabolites) is key towards driving scientific advancement in metabolomics, and gas chromatography–mass spectrometry (GC-MS) is an analytic method that may be applied to facilitate this process. The typical GC-MS identification workflow involves quantifying the similarity of an observed sample spectrum and other features (e.g. retention index) to that of several references, noting the compound of the best-matching reference spectrum as the identified metabolite. While a deluge of similarity metrics exists, none characterize the error rate of generated identifications, thereby presenting an unknown risk of false identification or discovery. To quantify this unknown risk, we propose a model-based framework for estimating the false discovery rate (FDR) among a set of identifications. Extending the traditional mixture modeling framework, our method incorporates both similarity score and experimental information in estimating the FDR. We apply these models to identification lists derived from across 548 samples of varying complexity and sample type (e.g., fungal species, standard mixtures, etc.), comparing their performance to that of the traditional Gaussian mixture model (GMM). Through simulation, we additionally assess the impact of reference library size on the accuracy of FDR estimates. In comparing the best performing model extensions to the GMM, our results indicate relative decreases in median absolute estimation error (MAE) ranging from 12% to 70%, based on comparisons of the median MAEs across all hit-lists. Results indicate that these relative performance improvements generally hold despite library size, however FDR estimation error typically worsens as the set of reference compounds diminishes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Semi-supervised Hybrid Machine Learning Framework for the Qualification of Resistance Spot Welds

• Industries requiring high structural integrity, including automotive, aerospace, and construction, place considerable significance on weld quality classification. • The inspection normally involves human expertise through predefined quality metrics that are subjective, error-prone, and time-intensive • The challenge to classification model development is the scarcity of labeled data and imbalanced distributions in the data that are labeled. • This work develops a new hybrid methodology that achieves clustering using KMeans++ together with supervised classification to overcome these challenges. • The ensemble-based classifiers were identified as optimal, with accuracy enhancements of up to 8% using the pseudo-labeled dataset. • The work provides practical insight into feature engineering and machine learning integration in industrial quality assurance applications.

Rogers, Jeremy K. [Savannah River National Laborat↗

Machine Learning Analysis of Temperature-Strain Relationships for Structural Health Monitoring of Pipes: Self-powered wireless sensor system for health monitoring of liquid-sodium cooled fast reactors

This report presents machine learning (ML) analysis of temperature-strain relationships for structural health monitoring of nuclear reactor stainless steel (SS) pipes with the strain gauge sensor directly printed on the pipe with a 3D conformal aerosol jet printer. We investigate correlations for two sensor pairs installed on the same SS304 pipe: commercial K-type thermocouple with a printed gold strain gauge (TC3-SG3), and commercial K-type thermocouple with commercial Kyowa strain gauge (TC0-SG0). The temperature ranges for the sensor pairs TC0-SG0 and TC3-SG3 are 20.00°C to 266.37°C and 39.95°C to 219.28°C respectively. ML algorithms in this study include Linear Regression (baseline method), Ridge Regression, Lasso Regression, and Gradient Boosting. Performance evaluation metrics include Root Mean Square Error (RMSE), Mean Square Error (MSE), Mean Absolute Error (MAE), R 2 Score, and Explained Variance. Using advanced feature engineering techniques, we extracted 27 temperature-based features and 30 strategic inclusion features. The best performance was obtained with the Gradient Boosting method, which achieves prediction accuracy of R 2 = 0.9999 and RMSE = 7.69 μStrain for TC0-SG0, and R 2 = 0.9998 and RMSE = 18.03 μStrain for TC3-SG3. While the temperature-strain correlations are weaker for the gauge directly printed on the pipe than for the commercial strain gauge, deployment-ready performance exceeding industry standards is achieved for both sensor pairs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A Study of the Effects of Atmospheric Phenomena on Mars Science Laboratory Entry Performance

At Earth during entry the shuttle has experienced what has come to be known as potholes in the sky or regions of the atmosphere where the density changes suddenly. Because of the small data set of atmospheric information where the Mars Science Laboratory (MSL) parachute deploys, the purpose of this study is to examine the effect similar atmospheric pothole characteristics, should they exist at Mars, would have on MSL entry performance. The study considers the sensitivity of entry design metrics, including altitude and range error at parachute deploy and propellant use, to pothole like density and wind phenomena.

Cianciolo, Alicia D.↗

Study of the Effect of Temporal Sampling Frequency on DSCOVR Observations Using the GEOS-5 Nature Run Results: Cloud Coverage - Part II

This is the second part of a study on how temporal sampling frequency affects satellite retrievals in support of the Deep Space Climate Observatory (DSCOVR) mission. Continuing from Part 1, which looked at Earth's radiation budget, this paper presents the effect of sampling frequency on DSCOVR-derived cloud fraction. The output from NASA's Goddard Earth Observing System version 5 (GEOS-5) Nature Run is used as the "truth". The effect of temporal resolution on potential DSCOVR observations is assessed by subsampling the full Nature Run data. A set of metrics, including uncertainty and absolute error in the subsampled time series, correlation between the original and the subsamples, and Fourier analysis have been used for this study. Results show that, for a given sampling frequency, the uncertainties in the annual mean cloud fraction of the sunlit half of the Earth are larger over land than over ocean. Analysis of correlation coefficients between the subsamples and the original time series demonstrates that even though sampling at certain longer time intervals may not increase the uncertainty in the mean, the subsampled time series is further and further away from the "truth" as the sampling interval becomes larger and larger. Fourier analysis shows that the simulated DSCOVR cloud fraction has underlying periodical features at certain time intervals, such as 8, 12, and 24 h. If the data is subsampled at these frequencies, the uncertainties in the mean cloud fraction are higher. These results provide helpful insights for the DSCOVR temporal sampling strategy.

GEOS-5↗

Landslide Likelihood Prediction using Machine Learning Algorithms

The supply of electricity via power plants is criticalto the operation of many critical infrastructure systems in mod-ern society. Natural hazards can disrupt the power supply, causepower outages that can halt economic growth, and impede emer-gency response until power is restored. The proposed work aimsto predict the landslides likelihood in these critical infrastructurelocations in the Northeastern USA using integrated databases ofexplanatory variables and machine learning algorithms. First,data related to landslides are obtained and merged, includingtopographic, soil moisture, and precipitation-related data. Fiveregression algorithms, namely: Random Forest, Extreme Gradi-ent Boosting (XGBoost), K-Nearest Neighbor regression (KNN),Linear Support Vector Regressor (SVR), and Linear regression,are utilized to predict the landslide probability and evaluatedon the dataset. The accuracy of the models is assessed by usingstatistical metrics such as mean absolute error (MAE), meansquared error (MSE), and root mean squared error (RMSE).The study results show that Random Forest outperformed othermodels with the mutual information feature selection method.It achieved an MSE of 0.0011 with mutual information-basedfeature selection and an MSE of 0.00157 without feature selection.KNN regressor outperformed the other models with an MSEof 0.00139 with correlation-based information selection. Theproposed landslide identification model with Random Forestalgorithm shows outstanding robustness and great potential intackling the landslide likelihood prediction by employing MLalgorithms.

Vasundhara Acharya↗

Data-Driven Computation of Probabilistic Marching Cubes for Efficient Visualization of Level-Set Uncertainty

Uncertainty visualization is an important emerging research area. Being able to visualize data uncertainty can help scientists improve trust in analysis and decision-making. However, visualizing uncertainty can add computational overhead, which can hinder the efficiency of analysis. In this paper, we propose novel data-driven techniques to reduce the computational requirements of the probabilistic marching cubes (PMC) algorithm. PMC is an uncertainty visualization technique that studies how uncertainty in data affects level-set positions. However, the algorithm relies on expensive Monte Carlo (MC) sampling for the multivariate Gaussian uncertainty model because no closed-form solution exists for the integration of multivariate Gaussian. In this work, we propose the eigenvalue decomposition and adaptive probability model techniques that reduce the amount of MC sampling in the original PMC algorithm and hence speed up the computations. Our proposed methods produce results that show negligible differences compared with the original PMC algorithm demonstrated through metrics, including root mean squared error, maximum error, and difference images. We demonstrate the performance and accuracy evaluations of our data-driven methods through experiments on synthetic and real datasets.

Athawale, Tushar↗

Estimation of coarse dead wood stocks in intact and degraded forests in the Brazilian Amazon using airborne lidar

Coarse dead wood is an important component of forest carbon stocks, but it is rarely measured in Amazon forests and is typically excluded from regional forest carbon budgets. Our study is based on line intercept sampling for fallen coarse dead wood conducted along 103 transects with a total length of 48 km matched with forest inventory plots where standing coarse dead wood was measured in the footprints of larger areas of airborne lidar acquisitions. We developed models to relate lidar metrics and Landsat time series variables to coarse dead wood stocks for intact, logged, burned, or logged and burned forests. Canopy characteristics such as gap area produced significant individual relations for logged forests. For total fallen plus standing coarse dead wood (hereafter defined as total coarse dead wood), the relative root mean square error for models with only lidar metrics ranged from 33 % in logged forest to up to 36 % in burned forests. The addition of historical information improved model performance slightly for intact forests (31 % against 35 % relative root mean square error), not justifying the use of a number of disturbance events from historical satellite images (Landsat) with airborne lidar data. Lidar-derived estimates of total coarse dead wood compared favorably with independent ground-based sampling for areas up to several hundred hectares. The relations found between total coarse dead wood and variables quantifying forest structure derived from airborne lidar highlight the opportunity to quantify this important but rarely measured component of forest carbon over large areas in tropical forests.

Brazilian Amazon↗

Data efficiency and extrapolation trends in neural network interatomic potentials

Abstract Recently, key architectural advances have been proposed for neural network interatomic potentials (NNIPs), such as incorporating message-passing networks, equivariance, or many-body expansion terms. Although modern NNIP models exhibit small differences in test accuracy, this metric is still considered the main target when developing new NNIP architectures. In this work, we show how architectural and optimization choices influence the generalization of NNIPs, revealing trends in molecular dynamics (MD) stability, data efficiency, and loss landscapes. Using the 3BPA dataset, we uncover trends in NNIP errors and robustness to noise, showing these metrics are insufficient to predict MD stability in the high-accuracy regime. With a large-scale study on NequIP, MACE, and their optimizers, we show that our metric of loss entropy predicts out-of-distribution error and data efficiency despite being computed only on the training set. This work provides a deep learning justification for probing extrapolation and can inform the development of next-generation NNIPs.

36 MATERIALS SCIENCE↗

Advanced Communications Technology Satellite (ACTS) Fade Compensation Protocol Impact on Very Small-Aperture Terminal Bit Error Rate Performance

The Advanced Communications Technology Satellite (ACTS) communications system operates at Ka band. ACTS uses an adaptive rain fade compensation protocol to reduce the impact of signal attenuation resulting from propagation effects. The purpose of this paper is to present the results of an analysis characterizing the improvement in VSAT performance provided by this protocol. The metric for performance is VSAT bit error rate (BER) availability. The acceptable availability defined by communication system design specifications is 99.5% for a BER of 5E-7 or better. VSAT BER availabilities with and without rain fade compensation are presented. A comparison shows the improvement in BER availability realized with rain fade compensation. Results are presented for an eight-month period and for 24 months spread over a three-year period. The two time periods represent two different configurations of the fade compensation protocol. Index Terms-Adaptive coding, attenuation, propagation, rain, satellite communication, satellites.

Cox, Christina B.↗

Predicting Elastic Constants of Refractory Complex Concentrated Alloys Using Machine Learning Approach

Refractory complex concentrated alloys (RCCAs) have drawn increasing attention recently owing to their balanced mechanical properties, including excellent creep resistance, ductility, and oxidation resistance. The mechanical and thermal properties of RCCAs are directly linked with the elastic constants. However, it is time consuming and expensive to obtain the elastic constants of RCCAs with conventional trial-and-error experiments. The elastic constants of RCCAs are predicted using a combination of density functional theory simulation data and machine learning (ML) algorithms in this study. The elastic constants of several RCCAs are predicted using the random forest regressor, gradient boosting regressor (GBR), and XGBoost regression models. Based on performance metrics R-squared, mean average error and root mean square error, the GBR model was found to be most promising in predicting the elastic constant of RCCAs among the three ML models. Additionally, GBR model accuracy was verified using the other four RHEAs dataset which was never seen by the GBR model, and reasonable agreements between ML prediction and available results were found. The present findings show that the GBR model can be used to predict the elastic constant of new RHEAs more accurately without performing any expensive computational and experimental work.

36 MATERIALS SCIENCE↗