Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Emerging anomaly detection techniques for electronic health records: A survey

Background Anomaly detection in electronic health records (EHRs) is a cornerstone of biomedical informatics, with direct implications for patient safety, clinical decision-making, and the prevention of healthcare fraud. Once guided primarily by simple rule-based methods, the field has advanced rapidly, driven by increased computing power, richer and more detailed health data, and the rise of machine learning and deep learning techniques. The objective of this paper is to provide a comprehensive overview of modern approaches to detecting anomalies in EHRs, outlining their strengths, limitations, and relevance to key healthcare challenges. We review traditional statistical methods alongside newer ML- and DL-based strategies and hybrid models, with particular attention to how these techniques support transparency and build clinical trust. Methods This paper presents a thorough and critical survey through systematic review (PRISMA-based) of the latest anomaly detection strategies in time-sequence data domains within electronic health record systems. Results We explore a broad spectrum of methodologies, including statistical models, supervised and unsupervised learning approaches, hybrid frameworks, and state-of-the-art ML-based techniques that collectively advance the precision and scalability of detecting anomalies in complex clinical datasets. In addition to mapping current capabilities, we address the enduring challenges that hinder widespread implementation and provide a forward-looking perspective on the future of anomaly detection in the data-rich landscape of modern healthcare. Summary The advancement in AI-based approaches is reported along with the basic principles of the individual approaches and their applicability. The increased availability of high-quality data, advancements in DL approaches, and enhanced computation power are leading to more frequent adaptation of DL-based approaches. Emerging DL-based approaches that have been adapted in other domains or recently applied in the EHR domain are also discussed in detail. Although DL-based approaches can improve model predictions by incorporating comorbidities, their application is limited in low-frequency data domains (e.g., when the total available data remains in the single digits). Therefore, the user must carefully consider the application based on data availability.

Anomaly detection↗

Cross-Market Price Difference Forecast Using Deep Learning for Electricity Markets

Price forecasting is in the center of decision making in electricity markets. Many researches have been done in forecasting energy prices while little research has been reported on forecasting price difference between day-ahead and realtime markets due to its high volatility, which however plays a critical role in virtual trading. To this end, this paper takes the first attempt to employ novel deep learning architecture with Bidirectional Long-Short Term Memory (LSTM) units to forecast the price difference between day-ahead and real-time markets for the same node. The raw data is collected from PJM market, processed and fed into the proposed network. The Root Mean Squared Error (RMSE) and customized performance metric are used to evaluate the performance of the proposed method. Case studies show that it outperforms the traditional statistical models like ARIMA, and machine learning models like XGBoost and SVR methods in both RMSE and the capability of forecasting the sign of price difference. Additionally, to cross-market price difference forecast, the proposed approach has the potential to be applied to solve other forecasting problems such as price spread forecast in DA market for Financial Transmission Right (FTR) trading purpose.

DA/RT price difference↗

High bias machine learning for antineutrino-based safeguards for small reactors

The statistical methods used for antineutrino detection will need to be improved to effectively monitor the inventory of next-generation nuclear reactors. In this sensitivity study, we evaluate machine learning models compared to previously used statistical approaches to identify diversion scenarios in a simulated Advanced Fast Reactor (AFR)-100. A chi-square goodness-of-fit technique, which individually compares the simulated antineutrino yields to the expected antineutrino yield, resulted in precise but low diversion detection probability. Various support vector machine (SVM) models were applied with diverse training datasets to evaluate the robustness of the method towards unexpected or “unseen” diversion scenarios. Furthermore, our results indicate that while the SVM models significantly improved the detection probability of near-field antineutrino-based safeguards, up to a probability of ~0.04, for the simulated small reactor, the detection system still needs improvements to reach the 0.2 detection limit established by the International Atomic Energy Agency.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Using kernel-based statistical distance to study the dynamics of charged particle beams in particle-based simulation codes

Measures of discrepancy between probability distributions (statistical distance) are widely used in the fields of artificial intelligence and machine learning. We describe how certain measures of statistical distance can be implemented as numerical diagnostics for simulations involving charged-particle beams. Related measures of statistical dependence are also described. The resulting diagnostics provide sensitive measures of dynamical processes important for beams in nonlinear or high-intensity systems, which are otherwise difficult to characterize. Here, the focus is on kernel-based methods such as maximum mean discrepancy, which have a well-developed mathematical foundation and reasonable computational complexity. Several benchmark problems and examples involving intense beams are discussed. While the focus is on charged-particle beams, these methods may also be applied to other many-body systems such as plasmas or gravitational systems.

47 OTHER INSTRUMENTATION↗

Learning Universal Authorship Representations

This code contains all the utilities required to reproduce the results of our EMNLP 2021 paper "Learning Universal Authorship Representations". It contains the utilities required for download the datasets, training our model, and performing all evaluations necessary for reproducing the results in the paper. Here's the abstract of our work: Determining whether two documents were composed by the same author, also known as authorship verification, has traditionally been tackled using statistical methods. Recently, authorship representations learned using neural networks have been found to outperform alternatives, particularly in large-scale settings involving hundreds of thousands of authors. But do such representations learned in a particular domain transfer to other domains? Or are these representations inherently entangled with domain-specific features? To study these questions, we conduct the first large-scale study of cross-domain transfer for authorship verification considering zero-shot transfers involving three disparate domains: Amazon reviews, fanfiction short stories, and Reddit comments. We find that although a surprising degree of transfer is possible between certain do- mains, it is not so successful between others. We examine properties of these domains that influence generalization and propose simple but effective methods to improve transfer.

Rivera Soto, RafaelA↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Diagnostics, Prognostics, and Optimization for Lithium-Ion Battery Systems

Health management of lithium-ion battery systems presents a host of challenges due to their complex physics, large numbers of components, and a wide variety of degradation behaviors across different battery types. Dr. Paul Gasper will present on research from the Electrochemical Energy Storage Group on Lithium-ion battery diagnostics, prognostics, and optimization. Diagnostics research, including state-estimation via machine-learning from electrochemical impedance spectroscopy and DC pulses as well as continuous state-estimation via Kalman filters, will highlight the ongoing challenges for accurately measuring the state of batteries without performing time-consuming characterization tests. NLR's industry-recognized battery prognostics work, which predicts real-world battery degradation by identifying degradation rate models from accelerated aging data using statistical modeling and machine-learning, will be used to demonstrate the critical impact of battery controls, thermal management, and operating strategy on durability and lifetime. Finally, the use of prognostic models for financial or lifetime optimization will be discussed.

25 ENERGY STORAGE↗

Probabilistic learning and updating of a digital twin for composite material systems

This paper presents an approach for characterizing and estimating statistical dependence between a large number of observables in a composite material system. Conditional regression is carried out using the estimated joint density function, permitting a systematic exploration of interdependence between fine scale and coarse observables that can be used for both prognosis and design of complex material systems. An example demonstrates the integration of experimental data with a computational database. The statistical approach is based on the probabilistic learning on manifolds recently developed by the authors. This approach leverages intrinsic structure detected through diffusion on graphs with projected stochastic differential equations to generate samples constrained to that structure.

42 ENGINEERING↗

Loosely Conditioned Emulation of Global Climate Models With Generative Adversarial Networks

Climate models encapsulate our best understanding of the Earth system, allowing research to be conducted on its future under alternative assumptions of how human-driven climate forces are going to evolve. An important application of climate models is to provide metrics of mean and extreme climate changes, particularly under these alternative future scenarios, as these quantities drive the impacts of climate on society and natural systems. Because of the need to explore a wide range of alternative scenarios and other sources of uncertainties in a computationally efficient manner, climate models can only take us so far, as they require significant computational resources, especially when attempting to characterize extreme events, which are rare and thus demand long and numerous simulations in order to accurately represent their changing statistics. Here we use deep learning in a proof of concept that lays the foundation for emulating global climate model output for different scenarios. We train two "loosely conditioned" Generative Adversarial Networks (GANs) that emulate daily precipitation output from a fully coupled Earth system model: one GAN modeling Fall-Winter behavior and the other Spring-Summer. Our GANs are trained to produce spatiotemporal samples: 32 days of precipitation over a 64x128 regular grid discretizing the globe. We evaluate the generator with a set of related performance metrics based upon KL divergence, and find the generated samples to be nearly as well matched to the test data as the validation data is to test. We also find the generated samples to accurately estimate the mean number of dry days and mean longest dry spell in the 32 day samples. Our trained GANs can rapidly generate numerous realizations at a vastly reduced computational expense, compared to large ensembles of climate models, which greatly aids in estimating the statistics of extreme events.

climate emulation, extreme climate, impacts, machi↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative data sets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student t process regression. We apply Student t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via colouring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics and classic machine learning clustering algorithms.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predictive Battery Lifetime Modeling at NREL [Slides]

Battery lifetime models are used to extrapolate data from accelerated aging tests to simulate degradation in real-world applications such as electric vehicles and battery energy storage systems. Methods developed at NREL utilize both expert domain-knowledge and machine-learning to identify models, using statistical methods such as cross-validation and bootstrap resampling to interrogate model performance and quantify uncertainty. These models can be utilized in systems level simulations to predict battery performance or technoeconomic models to estimate the lifetime cost of battery systems.

25 ENERGY STORAGE↗

Accelerating phase-field-based microstructure evolution predictions via surrogate models trained by machine learning methods

Abstract The phase-field method is a powerful and versatile computational approach for modeling the evolution of microstructures and associated properties for a wide variety of physical, chemical, and biological systems. However, existing high-fidelity phase-field models are inherently computationally expensive, requiring high-performance computing resources and sophisticated numerical integration schemes to achieve a useful degree of accuracy. In this paper, we present a computationally inexpensive, accurate, data-driven surrogate model that directly learns the microstructural evolution of targeted systems by combining phase-field and history-dependent machine-learning techniques. We integrate a statistically representative, low-dimensional description of the microstructure, obtained directly from phase-field simulations, with either a time-series multivariate adaptive regression splines autoregressive algorithm or a long short-term memory neural network. The neural-network-trained surrogate model shows the best performance and accurately predicts the nonlinear microstructure evolution of a two-phase mixture during spinodal decomposition in seconds, without the need for “on-the-fly” solutions of the phase-field equations of motion. We also show that the predictions from our machine-learned surrogate model can be fed directly as an input into a classical high-fidelity phase-field model in order to accelerate the high-fidelity phase-field simulations by leaping in time. Such machine-learned phase-field framework opens a promising path forward to use accelerated phase-field simulations for discovering, understanding, and predicting processing–microstructure–performance relationships.

36 MATERIALS SCIENCE↗

ORION: Operational FoRecastIng Of INduced Seismicity

The Operational Forecasting of Induced Seismicity toolkit “ORION” (ORION) is an open-source, observation-based ensemble forecasting toolkit which is geared towards helping operators understand the seismic hazard (i.e., probabilistic assessment of the magnitude and frequency of induced seismic events) at a site. ORION analyzes how the seismic hazard evolves during injection and suggests possible mitigation strategies to employ if an earthquake that exceeds certain threshold is observed. Through its ensemble modeling approach, ORION leverages the benefits of statistical-, physics-, and machine learning-based forecasting methodologies, while reducing the impact of each model’s respective limitations. The ORION toolkit consists of an easy-to-use GUI interface that affords a user as much or as little interaction as desired. Advanced capabilities allow the user to upload local, high-precision earthquake catalogs, projected injection profiles and/or spatiotemporal estimates of pressure/stress, and to tune various model parameters. ORION will then provide a spatial and temporal ensemble forecast of seismicity defined as the probability of exceedance of a given earthquake magnitude over a forecast period. Additionally, ORION will provide probability distribution of the statistically derived maximum possible earthquake magnitude that may be expected. Finally, ORION will provide suggested operational management strategies (e.g. reduce injection volumes at specific wells) based on the level of hazard.

NRAP,NRAP Tools,ORION,Operational Forcasting of In↗

Evaluating proxies for the drivers of natural gas productivity using machine-learning models

We report the extensive development of unconventional reservoirs using horizontal drilling and multistage hydraulic fracturing has generated large volumes of reservoir characterization and production data. The analysis of this abundant data using statistical methods and advanced machine-learning (ML) techniques can provide data-driven insights into well performance. Most predictive modeling studies have focused on the impact that different well completion and stimulation strategies have on well production but have not fully exploited the available in situ rock property data to determine its role in reservoir productivity. We have used machine-learning techniques to rank rock mechanical properties, microseismic attributes, and stimulation parameters in the order of their significance for predicting natural gas production from an unconventional reservoir. The data for this study came from a hydraulically fractured well in the Marcellus Shale in Monongalia County, West Virginia. The data classes included measurements aggregated by well completion stage that included (1) gas production, (2) well-log-derived measurements including bulk density, elastic moduli, shear impedance, compressional impedance, brittleness, and gamma measurements, (3) microseismic attributes, (4) long-period long-duration (LPLD) event counts, (5) fracture counts, and (6) stimulation parameters that included the fluid injection volume and average pumping pressure. To identify observable proxies for the drivers of gas production, we evaluated five commonly used ML approaches including multivariate adaptive regression spline, Gaussian mixture model, random forest, gradient boosting, and neural network. We selected five variables including LPLD event count, seismogenic b-value, hydraulic diffusivity, cumulative moment, and fluid volume as the features most likely to impact gas productivity at the stage level in the study area. The data-driven selection of these parameters for their importance in determining gas production can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs. Plain language summary: We use machine-learning methods and data-driven selection of reservoir parameters to rank and better understand their importance in determining gas production, which can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs.

58 GEOSCIENCES↗

Spatiotemporal Learning in Power Modules: Wavelet-Enhanced Forecasting of Thermomechanical Degradation

Detecting internal defects in power electronics packages is critical for their performance and reliability, especially under extreme operating conditions, as these defects can lead to catastrophic failure if not properly addressed. Confocal scanning acoustic microscopy (C-SAM) plays a key role in the nondestructive evaluation of bond layer degradation within a power electronics package by detecting defects such as delamination, voids, and cracks. However, accurately quantifying and predicting these defects from C-SAM images remains a significant challenge due to the low noise-to-signal ratio, which typically arises from both imaging process and bond patterns itself. In this paper, we explore machine learning strategies for processing C-SAM images and providing predictive models of defect growth. We use C-SAM images of sintered copper and sintered silver samples, which are obtained under accelerated thermal experiments, as the representative dataset for our study. We investigate the effect of Fourier transforms and wavelet transforms on these datasets to remove high-frequency noise and address noise across multiple scales with histogram equalization to enhance the contrast and improve the visibility of defects. As a result, defect boundaries can be clearly distinguished, enabling more accurate tracking of their growth over time. We then employ different time-series forecasting algorithms on the denoised images to formulate an image-based lifetime prediction model. Statistical models and deep-learning techniques are trained on images obtained in the early stages of thermal shock, and defect growth in the later stages is predicted. Our work serves as a preliminary attempt to improve the accuracy of lifetime prediction models of power electronics packages, which is critical under extreme operating environments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Probabilistic-learning-based stochastic surrogate model from small incomplete datasets for nonlinear dynamical systems

We consider a high-dimensional nonlinear computational model of a dynamical system, parameterized by a vector-valued control parameter, in the presence of uncertainties represented by an uncontrolled parameter modeled by a vector-valued random variable, and possibly with stochastic excitation. The objective is to construct a statistical surrogate model where the input is any deterministic value of the control parameter, and the output is a vector-valued observation of the computational model, which is a random vector whose probability measure is updated using a target dataset. To construct this statistical surrogate model, the stochastic response of the computational model must be built, which is a vector-valued time-discretized stochastic process in high dimension, depending on the control parameter. It is assumed that the computational cost of a single evaluation of the deterministic model is high. For the probabilistic updating, we consider a subset of the components of the observation of the computational model, defined as the “identification observation” of the computational model, for which a small target dataset is available. Therefore, the target dataset is associated with partial observability, corresponding to an incomplete data case. Given a prior probability model of the random control and uncontrolled parameters, a training dataset is constructed, consisting of realizations of the random triplet composed of the stochastic response, the random identification observation, and the random control parameter. Since the computational cost of a single evaluation of the deterministic model is assumed to be large, the training dataset is also of small size. The main challenges in this problem are the high dimensionality, partial observability leading to incomplete data in the target dataset for the identification observation of the computational model (which is not sufficient to identify the computational stochastic responses), and the availability of a small training dataset. To address these challenges, we propose a methodology based on statistical methods for constructing necessary reduced representations, direct probabilistic learning under constraints using probabilistic learning on manifolds (PLoM) constrained by the target dataset, and the use of a weak formulation of the Fourier transform of probability measures. Statistical conditioning is also employed to explore the learned dataset. The constructed predictive statistical surrogate model can be implemented in the context of online computation. Here, we apply this approach to a problem of nonlinear stochastic dynamics in high dimensions within the framework of deformable solids mechanics.

Engineering↗

Personalized Tucker Decomposition: Modeling Commonality and Peculiarity on Tensor Data

In this paper, we propose a personalized Tucker decomposition (perTucker) to address the limitations of traditional tensor decomposition methods in capturing heterogeneity across different datasets. perTucker decomposes tensor data into shared global components and personalized local components. We introduce an order orthogonality assumption and develop a proximal gradient regularized block coordinate descent algorithm guaranteed to converge to a stationary point. The unique and common representations learned by perTucker reveal intrinsic statistical patterns in data and provide valuable information for a wide range of downstream analytics, including anomaly detection, source classification, and clustering. We demonstrate perTucker’s effectiveness through a simulation study and two case studies on solar flare detection and tonnage signal classification.

14 SOLAR ENERGY↗

A review of computing-based automated fault detection and diagnosis of heating, ventilation and air conditioning systems

We report faults in Heating, Ventilation, and Air Conditioning (HVAC) systems of buildings result in significant energy waste in building operation. With fast-growing sensing data availability and advancement in computing, computational modeling has demonstrated strong capability to detect and diagnose HVAC system faults, hence, ensuring efficient building operation. This paper comprehensively reviews the state-of-the-art computing-based fault detection and diagnosis (FDD) for HVAC systems. Overall, the reviewed computing-based FDD methods are classified as two major approaches: knowledge-based and data-driven approaches. We then identify multiple important topics, including data availability, training data size, data quality, approach generality, capability, interpretability, and required modeling efforts, along with corresponding metrics to summarize the most updated FDD development. Generally, the knowledge-based approaches are further divided as physics-based modeling, Diagnostic Bayesian Network, and performance indicator-based methods while data-driven approaches include supervised learning, unsupervised learning, and regression and statistics-based methods. State-of-the-art FDD development, remaining challenges, and future research directions are further discussed to push forward FDD in practice. Availability of fault data, capability of existing methods to deal with complex fault situations (such as simultaneous faults), modeling interpretability for data-driven methods, and required engineering efforts for physics-based methods are identified as remaining challenges in FDD development. Improving modeling fidelity and reducing modeling efforts are essential for applying physics-based methods in real buildings. Meanwhile, addressing fault data availability, increasing algorithm adaptability, and handling multiple faults are essential to further enhance the applicability of data-driven FDD approaches.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗