Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Flood Susceptibility Mapping Using Machine Learning and Geospatial-Sentinel-1 SAR Integration for Enhanced Early Warning Systems

This study presents a comprehensive framework for flood susceptibility mapping by integrating geospatial factors with both statistical and machine learning models. Thirteen Flood-related factors, including DEM, slope, TWI, NDVI, etc., are extracted as features of models, and historical flood data derived from Sentinel-1 SAR from 2018 to 2023 are used as the target variables of the models. These datasets are analyzed using a frequency-based statistical model and three machine learning models, including Random Forest, XGBoost, and CNN, to generate flood susceptibility maps. The performance of each model is evaluated through AUC; and SHAP scores are separately generated for Machine learning (ML) models to explain each feature contribution in the ML model. The generated susceptibility maps are validated by high-flood-risk locations monitored by flood sensors, BLE inundation models, and flood-prone areas suggested by the Local Community Task Force. The results indicate that the XGBoost model outperforms all other models, with an AUC of 0.92 and demonstrates the highest alignment with recommended high-flood-risk locations, while the frequency-based statistical model showed the weakest performance with an AUC of 0.65. SHAP value graphs highlight the elevation, slope, and TWI as the most influential features across all models. The susceptibility maps generated by the machine learning model show strong agreement with the BLE map and high-flood-risk areas identified by the local Community Task Force.

Google Engine

Long-Term Statistical Process Monitoring of an Ultrafiltration Water Treatment Process

As water treatment technology has improved, the amount of available process data has substantially increased, making real-time, data-driven fault detection a reality. One shortcoming of the fault detection literature is that methods are usually evaluated by comparing their performance on hand-picked, short-term case studies, which yields no insight into long-term performance. In this work, we first evaluate multiple statistical and machine learning approaches for detrending process data. Then, we evaluate the performance of a PCA-based fault detection approach, applied to the detrended data, to monitor influent water quality, filtrate quality, and membrane fouling of an ultrafiltration membrane system for indirect potable reuse. Based on two short case studies, the adaptive lasso detrending method is selected, and the performance of the multivariate approach is evaluated over more than a year. The method is tested for different sets of three critical tuning parameters, and we find that for long-term, autonomous monitoring to be successful, these parameters should be carefully evaluated. However, in comparison with industry standards of simpler, univariate monitoring or daily pressure decay tests, multivariate monitoring produces substantial benefits in long-term testing.

ammonia

Introduction to Analysis Methods for Big Earth Data

Big Earth Data are too big to be tractable to simple data inspection. Thus, they typically require models to make sense of all the data. Useful models for Big Earth Data may be physical, statistical, or machine learning based. While physical models are ideal for understanding the data, they are not always feasible, particularly when our ability to observe at finer scales exceeds our ability to incorporate the physics. Statistical models are more generalized, but computationally intensive for many Earth Observation datasets. Machine Learning models generally scale well but are sometimes limited in the physical understanding they can offer. Hybrid models combine attributes—and advantages—of two or more of these types.

Christopher Lynnes

Low-Cost Sensor Performance Intercomparison, Correction Factor Development, and 2+ Years of Ambient PM2.5 Monitoring in Accra, Ghana

Particulate matter air pollution is a leading cause of global mortality, particularly in Asia and Africa. Addressing the high and wide-ranging air pollution levels requires ambient monitoring, but many low- and middle-income countries (LMICs) remain scarcely monitored. To address these data gaps, recent studies have utilized low-cost sensors. These sensors have varied performance, and little literature exists about sensor intercomparison in Africa. By colocating 2 QuantAQ Modulair-PM, 2 PurpleAir PA-II SD, and 16 Clarity Node-S Generation II monitors with a reference-grade Teledyne monitor in Accra, Ghana, we present the first intercomparisons of different brands of low-cost sensors in Africa, demonstrating that each type of low-cost sensor PM2.5 is strongly correlated with reference PM2.5, but biased high for ambient mixture of sources found in Accra. When compared to a reference monitor, the QuantAQ Modulair-PM has the lowest mean absolute error at 3.04 μg/m3, followed by PurpleAir PA-II (4.54 μg/m3) and Clarity Node-S (13.68 μg/m3). We also compare the usage of 4 statistical or machine learning models (Multiple Linear Regression, Random Forest, Gaussian Mixture Regression, and XGBoost) to correct low-cost sensors data, and find that XGBoost performs the best in testing (R2: 0.97, 0.94, 0.96; mean absolute error: 0.56, 0.80, and 0.68 μg/m3 for PurpleAir PA-II, Clarity Node-S, and Modulair-PM, respectively), but tree-based models do not perform well when correcting data outside the range of the colocation training. Therefore, we used Gaussian Mixture Regression to correct data from the network of 17 Clarity Node-S monitors deployed around Accra, Ghana, from 2018 to 2021. We find that the network daily average PM2.5 concentration in Accra is 23.4 μg/m3, which is 1.6 times the World Health Organization Daily PM2.5 guideline of 15 μg/m3. While this level is lower than those seen in some larger African cities (such as Kinshasa, Democratic Republic of the Congo), mitigation strategies should be developed soon to prevent further impairment to air quality as Accra, and Ghana as a whole, rapidly grow.

Humidity

A cross-dimensional analysis of data-driven short-term load forecasting methods with large-scale smart meter data

Electricity load forecasting is essential to utility operation and power grid stability. A wide spectrum of data-driven methods, ranging from linear regression models to more recent deep learning models have been adopted to forecast electric load over the years. However, there still lacks a holistic evaluation of the applicability of conventional statistical and machine learning based algorithms with respect to different temporal and spatial scopes, computational requirements, and sensitivity of model-tuning. Enabled by a large-scale electricity load profile dataset of over 40,000 residential customers in a utility region, we conducted a cross-dimensional analysis of data-driven load forecasting methods. Three regression-based and seven deep learning algorithms with different model configurations were evaluated in terms of their overall and peak load prediction accuracy, and training burdens, across spatial aggregation levels ranging from the transformer, feeder, substation, to neighborhood. We found, first, the load forecasting accuracy is constrained by a predictability boundary, influenced by the forecasting horizon and spatial aggregation level. Specifically, RandomForest, XGBoost, TFT, TSMixer, and TiDE models achieved less than 10 % prediction error for up to 96-h ahead forecasting for district, substation, and feeder levels, while other models struggle at long-horizon predictions; Second, for winter and summer peak load dates, most models were able to predict the peak demand timing within ± 1 h, but the prediction percentage error varied by models, with TFT and TiDE models being the top performers; Third, models with similar prediction accuracy can differ in training burden by an order of magnitude. Therefore, choosing model configurations that balance prediction performance and computational resource is an important practical consideration for large-scale deployment of the machine learning based load forecasting. The outcome of this study can guide researchers and practitioners to choose the proper load forecasting algorithms based on their problem scope, required accuracy, and available resources. The predictability boundary can serve as a benchmark for electricity load forecasting problems with new algorithms and datasets.

Li, Han

sOPTICS: a modified density-based algorithm for identifying galaxy groups/clusters and brightest cluster galaxies

A direct approach to studying the galaxy–halo connection is to analyse groups and clusters of galaxies that trace the underlying dark matter haloes, emphasizing the importance of identifying galaxy clusters and their associated brightest cluster galaxies (BCGs). In this work, we test and propose a robust density-based clustering algorithm that outperforms the traditional Friends-of-Friends (FoF) algorithm in the currently available galaxy group/cluster catalogues. Our new approach is a modified version of the Ordering Points To Identify the Clustering Structure (OPTICS) algorithm, which accounts for line-of-sight positional uncertainties due to redshift space distortions by incorporating a scaling factor, and is thereby referred to as sOPTICS. When tested on both a galaxy group catalogue based on semi-analytic galaxy formation simulations and observational data, our algorithm demonstrated robustness to outliers and relative insensitivity to hyperparameter choices. In total, we compared the results of eight clustering algorithms. The proposed density-based clustering method, sOPTICS, outperforms FoF in accurately identifying giant galaxy clusters and their associated BCGs in various environments with higher purity and recovery rate, also successfully recovering 115 BCGs out of 118 reliable BCGs from a large galaxy sample. Furthermore, when applied to an independent observational catalogue without extensive re-tuning, sOPTICS maintains high recovery efficiency, confirming its flexibility and effectiveness for large-scale astronomical surveys.

79 ASTRONOMY AND ASTROPHYSICS

SigTime: Learning and Visually Explaining Time Series Signatures

Understanding and distinguishing temporal patterns in time series data is essential for scientific discovery and decision-making. For example, in biomedical research, uncovering meaningful patterns in physiological signals can improve diagnosis, risk assessment, and patient outcomes. However, existing methods for time series pattern discovery face major challenges, including high computational complexity, limited interpretability, and difficulty in capturing meaningful temporal structures. Here, to address these gaps, we introduce a novel learning framework that jointly trains two Transformer models using complementary time series representations: shapelet-based representations to capture localized temporal structures and traditional feature engineering to encode statistical properties. The learned shapelets serve as interpretable signatures that differentiate time series across classification labels. Additionally, we develop a visual analytics system—SigTime—with coordinated views to facilitate exploration of time series signatures from multiple perspectives, aiding in useful insights generation. We quantitatively evaluate our learning framework on eight publicly available datasets and one proprietary clinical dataset. Additionally, we demonstrate the effectiveness of our system through two usage scenarios along with the domain experts: one involving public ECG data and the other focused on preterm labor analysis.

97 MATHEMATICS AND COMPUTING

A high-throughput approach for statistical process optimization in Laser Powder Bed Fusion

Process variability is inherent in metal additive manufacturing (AM). However, it is often overlooked in process optimization frameworks, constraining the understanding of process uncertainties and their influence on parameter selection. To address this, we present an integrated framework that combines high-throughput single-track experiments, GAN-based melt pool geometry extraction, robust statistical and machine learning modeling, and uncertainty-quantified process mapping. Process variability is characterized through single-track melt pool behaviors, and its influence on defect formation is systematically quantified to enable statistically guided process parameter optimization. This approach is demonstrated on Laser Powder Bed Fusion (L-PBF) of stainless steel 316L, effectively capturing the interplay between process parameters, melt pool variability, and defect probability. By integrating uncertainty quantification into process optimization, this study provides a structured methodology for addressing variability challenges in AM quality control, ultimately contributing to enhanced manufacturing reliability.

Laser Powder Bed Fusion

Applications of Principled Search Methods in Climate Influences and Mechanisms

Forest and grass fires cause economic losses in the billions of dollars in the U.S. alone. In addition, boreal forests constitute a large carbon store; it has been estimated that, were no burning to occur, an additional 7 gigatons of carbon would be sequestered in boreal soils each century. Effective wildfire suppression requires anticipation of locales and times for which wildfire is most probable, preferably with a two to four week forecast, so that limited resources can be efficiently deployed. The United States Forest Service (USFS), and other experts and agencies have developed several measures of fire risk combining physical principles and expert judgment, and have used them in automated procedures for forecasting fire risk. Forecasting accuracies for some fire risk indices in combination with climate and other variables have been estimated for specific locations, with the value of fire risk index variables assessed by their statistical significance in regressions. In other cases, the MAPSS forecasts [23, 241 for example, forecasting accuracy has been estimated only by simulated data. We describe alternative forecasting methods that predict fire probability by locale and time using statistical or machine learning procedures trained on historical data, and we give comparative assessments of their forecasting accuracy for one fire season year, April- October, 2003, for all U.S. Forest Service lands. Aside from providing an accuracy baseline for other forecasting methods, the results illustrate the interdependence between the statistical significance of prediction variables and the forecasting method used.

Glymour, Clark

Measuring the Weibull modulus of microscope slides

The objectives are that students will understand why a three-point bending test is used for ceramic specimens, learn how Weibull statistics are used to measure the strength of brittle materials, and appreciate the amount of variation in the strength of brittle materials with low Weibull modulus. They will understand how the modulus of rupture is used to represent the strength of specimens in a three-point bend test. In addition, students will learn that a logarithmic transformation can be used to convert an exponent into the slope of a straight line. The experimental procedures are explained.

Sorensen, Carl D.

High-Throughput Discovery Illuminates Design Principles and Limits for Long-Lived Charged Species in Organic Electrolytes

The chemical stability of charged molecules in all-organic redox flow batteries (RFBs) is required for the prolonged operation of these devices. Molecular engineering and electrolyte optimization are used to mitigate parasitic reactions and extend the lifetimes of the charge carriers. However, how much can structural variation extend the lifetime? To probe this query, we designed a high-throughput kinetic study of the radical cation of N-methylphenothiazinium, guided by statistical sampling and learning algorithms. Using Argonne’s autonomous discovery facility, we conducted over 6,000 kinetic experiments with robotic sample preparation, parallel kinetic measurements, and machine learning inputs, testing 188 solvent molecules selected from a space of over 540 candidates from 11 chemical classes. Algorithmic selections guided us to stable solvent candidates, which were further tested in high concentration with and without supporting electrolyte. Our findings reveal the inherent difficulty of exceeding the current state of the art through solvent variation. The desired stability is statistically rare and poorly predictable. Among the many tested, only three solvents significantly outperformed our baseline, acetonitrile─and none by more than a factor of 3─suggesting a general challenge in achieving the necessary techno-economic targets. Furthermore, we suggest that self-discharge through solvent homolysis is the cause of the observed limitations. Several structural motifs contribute to >1,000 h half-life stability including molecular simplicity, symmetry, oxidation complement, and strategic fluorination. Importantly, this workflow establishes effective assays for diagnosing and predicting oxidative stress for highly stable liquid electrolytes in all batteries.

Batteries

Low responsiveness of machine learning models to critical or deteriorating health conditions

Machine learning (ML) based mortality prediction models can be immensely useful in intensive care units. Such a model should generate warnings to alert physicians when a patient’s condition rapidly deteriorates, or their vitals are in highly abnormal ranges. Before clinical deployment, it is important to comprehensively assess a model’s ability to recognize critical patient conditions. We develop multiple medical ML testing approaches, including a gradient ascent method and neural activation map. We systematically assess these machine learning models’ ability to respond to serious medical conditions using additional test cases, some of which are time series. Guided by medical doctors, our evaluation involves multiple machine learning models, resampling techniques, and four datasets for two clinical prediction tasks. We identify serious deficiencies in the models’ responsiveness, with the models being unable to recognize severely impaired medical conditions or rapidly deteriorating health. For in-hospital mortality prediction, the models tested using our synthesized cases fail to recognize 66% of the injuries. In some instances, the models fail to generate adequate mortality risk scores for all test cases. Our study identifies similar kinds of deficiencies in the responsiveness of 5-year breast and lung cancer prediction models. Using generated test cases, we find that statistical machine-learning models trained solely from patient data are grossly insufficient and have many dangerous blind spots. Most of the ML models tested fail to respond adequately to critically ill patients. How to incorporate medical knowledge into clinical machine learning models is an important future research direction.

60 APPLIED LIFE SCIENCES

Changes in ribbon synapses and rough endoplasmic reticulum of rat utricular macular hair cells in weightlessness

This study combined ultrastructural and statistical methods to learn the effects of weightlessness on rat utricular maculae. A principle aim was to determine whether weightlessness chiefly affects ribbon synapses of type II cells, since the cells communicate predominantly with branches of primary vestibular afferent endings. Maculae were microdissected from flight and ground control rat inner ears collected on day 13 of a 14-day spaceflight (F13), landing day (R0) and day 14 postflight (R14) and were prepared for ultrastructural study. Ribbon synapses were counted in hair cells examined in a Zeiss 902 transmission electron microscope. Significance of synaptic mean differences was determined for all hair cells contained within 100 section series, and for a subset of complete hair cells, using SuperANOVA software. The synaptic mean for all type II hair cells of F13 flight rats increased by 100%, and that for complete cells by 200%. Type I cells were less affected, with synaptic mean differences statistically insignificant in complete cells. Synapse deletion began within 8 h upon return to Earth. Additionally, hair cell laminated rough endoplasmic reticulum of flight rats was reversibly disorganized on R0. Results support the thesis that synapses in type II hair cells are uniquely affected by altered gravity. Type II hair cells may be chiefly sensors of gravitational and type I cells of translational linear accelerations.

short duration

MARGInS: Model-Based Analysis of Realizable Goals in Systems

Under NASAs Constellation effort, the Exploration Technology Development Program funded research toward a system validation capability that applied machine learning and test-case generation techniques to the analysis of black-box system behavior. The behavior analysis capability scaled to spaces of hundreds of input parameters and tens of thousands of test cases. Aerospace systems at the vehicle level, especially those systems which contain some level of autonomy, are best described by hybrid and non-linear mathematics. Even simplified models of such systems need parameter dimensionalities in the hundreds or thousands of parameters in order to capture sufficient fidelity. The System Safety Assessments (such as those described in the SAE ARP 4761A Safety Assessment Process guidelines) for these systems are prone to errorinteractions between the vehicles subsystems are complex, and can display emergent behaviors. NASA captured this new analysis in the Model-based Analysis of Realizable Goals in Systems (MARGInS) tool and applied it to the Pad Abort 1 (PA-1) simulation as part of the independent validation and verification cycle before the PA-1 flight test in May of 2010. MARGInS evaluated the adherence of the high-fidelity simulation to its requirements, and deter- mined the margins to failure from the expected nominal input conditions. Following the PA-1 test, the capabilities within the MARGInS framework have been extended with sophisticated statistical and white-box test case generation techniques and applied to other NASA missions. The frame- work now includes a critical factors analysis that was applied to NASAs Orion simulation and design. NASAs Aeronautics Research Mission Directorate (ARMD) leveraged the existing MARGInS framework for work on aviation safety for civil transport vehicles and for research on autonomy issues. The NASA ARMD effort created a time series output prediction capability that has been used to characterize trajectories for a plane with an adaptive control system, and a safety boundary detection capability that has been applied to an air traffic control concept of operation for the Federal Aviation Administration. The statistical and machine- learning based techniques within MARGInS have been successfully combined with concolic execution to improve the coverage of a critical unit by driving system-level inputs. The use case driving the concolic execution and MARGInS integration was inspired by the Air France 447 disaster in which the loss of a critical functionality (the airspeed calculation from the pitot tubes) led to loss of the entire plane with the people aboard. To illustrate capabilities and limitations, we will highlight the analyses for the applications listed above. We will then discuss the future plans for MARGInS and its interfaces with other tools.

Validation

ASCoT 3: Nonlinear Principal Components Analysis and Uncertainty Quantification in Early Concept Spacecraft Flight Software Cost Estimation

For mission planners and evaluators alike, value in cost models comes from a mean or median prediction, an understanding of the uncertainty on that prediction, and an understanding of model performance. Here we apply advanced statistical and machine learning methods to spacecraft flight software cost, effort, and SLOC estimation, and present the results in the latest version of the Analogy Software Cost Tool (ASCoT). We present in- and out-of-sample performance metrics for our models, each of which incorporate some amount of epistemic uncertainty. ASCoT, hosted on the One NASA Cost Engineering (ONCE) database via the Online NASA Space Estimation Tool (ONSET), was first showcased in 2016 as a number of analogy-based models and methods (kNN and Clustering) to support early project formulation. This ASCoT update improves upon the previous analogic methods by incorporating uncertainty in the data transformations. In particular, we use a Nonlinear Principal Components Analysis (NLPCA) to deal with ordinal data.

Robotic Spacecraft

A kinetic-based regularization method for data science applications

We propose a physics-based regularization technique for function learning, inspired by statistical mechanics. By drawing an analogy between optimizing the parameters of an interpolator and minimizing the energy of a system, we introduce corrections that impose constraints on the lower-order moments of the data distribution. This minimizes the discrepancy between the discrete and continuum representations of the data, in turn allowing to access more favorable energy landscapes, thus improving the accuracy of the interpolator. Our approach improves performance in both interpolation and regression tasks, even in high-dimensional spaces. Unlike traditional methods, it does not require empirical parameter tuning, making it particularly effective for handling noisy data. We also show that thanks to its local nature, the method offers computational and memory efficiency advantages over Radial Basis Function interpolators, especially for large datasets.

97 MATHEMATICS AND COMPUTING