Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Statistical Learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Report on the AAPM grand challenge on deep generative modeling for learning medical image statistics

Abstract Background The findings of the 2023 AAPM Grand Challenge on Deep Generative Modeling for Learning Medical Image Statistics are reported in this Special Report. Purpose The goal of this challenge was to promote the development of deep generative models for medical imaging and to emphasize the need for their domain‐relevant assessments via the analysis of relevant image statistics. Methods As part of this Grand Challenge, a common training dataset and an evaluation procedure was developed for benchmarking deep generative models for medical image synthesis. To create the training dataset, an established 3D virtual breast phantom was adapted. The resulting dataset comprised about 108 000 images of size 512 512. For the evaluation of submissions to the Challenge, an ensemble of 10 000 DGM‐generated images from each submission was employed. The evaluation procedure consisted of two stages. In the first stage, a preliminary check for memorization and image quality (via the Fréchet Inception Distance [FID]) was performed. Submissions that passed the first stage were then evaluated for the reproducibility of image statistics corresponding to several feature families including texture, morphology, image moments, fractal statistics, and skeleton statistics. A summary measure in this feature space was employed to rank the submissions. Additional analyses of submissions was performed to assess DGM performance specific to individual feature families, the four classes in the training data, and also to identify various artifacts. Results Fifty‐eight submissions from 12 unique users were received for this Challenge. Out of these 12 submissions, 9 submissions passed the first stage of evaluation and were eligible for ranking. The top‐ranked submission employed a conditional latent diffusion model, whereas the joint runners‐up employed a generative adversarial network, followed by another network for image superresolution. In general, we observed that the overall ranking of the top 9 submissions according to our evaluation method (i) did not match the FID‐based ranking, and (ii) differed with respect to individual feature families. Another important finding from our additional analyses was that different DGMs demonstrated similar kinds of artifacts. Conclusions This Grand Challenge highlighted the need for domain‐specific evaluation to further DGM design as well as deployment. It also demonstrated that the specification of a DGM may differ depending on its intended use.

Radiology, Nuclear Medicine & Medical Imaging

ForceFinder

SAND2025-11750O ForceFinder extends the Structural Dynamics Python Libraries (SDynPy) with comprehensive tools for inverse source estimation (ISE) tasks via frequency response function (FRF) matrix inversion. The software is designed for transfer path analysis and multiple-input/multiple-output (MIMO) vibration control problems. It allows users to estimate sources through various algorithms, from the basic Moore-Penrose pseudo-inverse to statistical learning methods such as Tikhonov regularization via an L-curve and elastic net regularization via an information criterion. ForceFinder uses an object-oriented framework, where all components of the ISE problem—such as FRFs, responses, and transformations—are stored in a "SourcePathReceiver" object. This software can be applied to any noise and vibration problem. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Carter, Steven [Sandia National Lab. (SNL-CA), Liv

Investigating Aerosol and Meteorological Influences on Convective Clouds in Houston, Texas, during the TRACER/ESCAPE Field Campaigns

Aerosols serve as cloud condensation nuclei, shaping the microphysical properties of cloud droplets. Aerosol effects on convective clouds are complex and remain controversial. The debate centers around the process of aerosol-induced invigoration of deep convection, a phenomenon that could significantly affect convective cloud properties but lacks robust evidence due to methodological limitations in observational approaches and questions about the robustness of modeling studies. Resolving these discrepancies is crucial for understanding how aerosols affect the atmosphere. Here, this study examines the effects of meteorological and aerosol parameters in a weakly synoptic-driven convective environment, where the influence of aerosols may be more pronounced and observable. Daily atmospheric soundings and aerosol concentrations from several ground instruments collected during the summer of 2022 in Houston, Texas, as part of the Tracking Aerosol Convection interactions Experiment (TRACER) and Experiment of Sea Breeze Convection, Aerosols, Precipitation, and Environment (ESCAPE) field campaigns are analyzed. Statistical learning methods are applied to uncover the complex relationships between aerosols, meteorology, and convective cloud characteristics, such as cell area and echo-top height. The findings reveal that higher aerosol concentrations are associated with narrower convective cells, which we argue contradicts the idea of stronger convection with increased aerosol loading. However, once the data are clustered by the synoptic environment, the relationship between aerosol loading and convective cell area diminishes, indicating that the covariablity between synoptic-scale weather patterns, local thermodynamics, and aerosol loading makes it challenging to draw definitive conclusions about the specific impacts of aerosols on convective cloud properties.

54 ENVIRONMENTAL SCIENCES

Hybrid NN/SVM Computational System for Optimizing Designs

A computational method and system based on a hybrid of an artificial neural network (NN) and a support vector machine (SVM) (see figure) has been conceived as a means of maximizing or minimizing an objective function, optionally subject to one or more constraints. Such maximization or minimization could be performed, for example, to optimize solve a data-regression or data-classification problem or to optimize a design associated with a response function. A response function can be considered as a subset of a response surface, which is a surface in a vector space of design and performance parameters. A typical example of a design problem that the method and system can be used to solve is that of an airfoil, for which a response function could be the spatial distribution of pressure over the airfoil. In this example, the response surface would describe the pressure distribution as a function of the operating conditions and the geometric parameters of the airfoil. The use of NNs to analyze physical objects in order to optimize their responses under specified physical conditions is well known. NN analysis is suitable for multidimensional interpolation of data that lack structure and enables the representation and optimization of a succession of numerical solutions of increasing complexity or increasing fidelity to the real world. NN analysis is especially useful in helping to satisfy multiple design objectives. Feedforward NNs can be used to make estimates based on nonlinear mathematical models. One difficulty associated with use of a feedforward NN arises from the need for nonlinear optimization to determine connection weights among input, intermediate, and output variables. It can be very expensive to train an NN in cases in which it is necessary to model large amounts of information. Less widely known (in comparison with NNs) are support vector machines (SVMs), which were originally applied in statistical learning theory. In terms that are necessarily oversimplified to fit the scope of this article, an SVM can be characterized as an algorithm that (1) effects a nonlinear mapping of input vectors into a higher-dimensional feature space and (2) involves a dual formulation of governing equations and constraints. One advantageous feature of the SVM approach is that an objective function (which one seeks to minimize to obtain coefficients that define an SVM mathematical model) is convex, so that unlike in the cases of many NN models, any local minimum of an SVM model is also a global minimum.

Rai, Man Mohan

Prognostics

Knowledge discovery, statistical learning, and more specifically an understanding of the system evolution in time when it undergoes undesirable fault conditions, are critical for an adequate implementation of successful prognostic systems. Prognosis may be understood as the generation of long-term predictions describing the evolution in time of a particular signal of interest or fault indicator, with the purpose of estimating the remaining useful life (RUL) of a failing component/subsystem. Predictions are made using a thorough understanding of the underlying processes and factor in the anticipated future usage.

Systems Health Management

K-Means Cluster Study for Radiofrequency Propagation Characterization

The objective of this study is to design a simple method for mining radio frequency (RF) propagation data. The study explored the characteristics of a large dataset of propagation experiments conducted over the span of years and using several ground stations around the world. Furthermore, this study developed simple predictive models that can be used for link characterization and overall propagation behavior description, without the need for physical measurements on-site. It is understood that such statistical learning has several drawbacks in terms of accuracy and precision. K-means clustering was used to characterize the data set in a way never explored before in an attempt to create useful tools that reduce cost, time and risk. K-means clustering was used to characterize the data set. Cosine distance was used as a method to determine the optimal number for clustering each feature. Dependence and independence analysis was performed to explore intra and inter-sensitivity between the presented features, with respect to each other and time. Several predicative models were generated and evaluated with respect to a test set to assess a measure of prediction accuracy and precision. A simple method for data analysis was developed and tested as the basis for further studies and future refinement to produce optimal performing models.

Cognitive

Reducing V&V Cost of Flight Critical Systems: Myth or Reality?

This paper presents an overview of NASA research program on the V&V of flight critical systems. Five years ago, NASA started an effort to reduce the cost and possibly increase the effectiveness of V&V for flight critical systems. It is the right time to take a look back and realize what progress has been made. This paper describes our overall approach and the tools introduced to address different phases of the software lifecycle. For example, we have improved testing by developing a statistical learning approach tor defining test cases. The tool automatically identifies possible unsafe conditions by analyzing outliers in output data; using an iterative learning process, it can then generate more test cases that represent potentially unsafe regions of operation. At the code level, we have developed and made available as open source a static analyzer for C and C++ programs called IKOS. We have shown that IKOS is very precise in the analysis of embedded C programs (very few false positives) and a bit less for regular C and C++ code. At the design level, in collaboration with our NRA partners, we have developed a suite of analysis tools for Simulink models. The analysis is done in a compositional framework for scalability.

Brat, Guillaume P.

Autonomous Contingency Management In Urban Air Mobility: The Communication Network Awareness Machine System

Next Generation Air Transportation System (NextGen) has begun the modernization of the nation’s air transportation system (NAS), with goals to improve system safety, increase operation efficiency and capacity, provide enhanced predictability, resilience and robustness [1]. The overall objective of the Air Traffic Management-eXploration (ATM-X) project is to facilitate the goals of NextGen by conducting research to enable the growing demand of new, mission variant, air vehicles with safe access to the NAS. The implementation and utilization of new and burgeoning technologies that are both flexible, scalable, and systematically user-focused are requisite for ATM-X to achieve its intention of NAS safe entry [2]. Researchers from NASA Langley’s Flight Deck Integration Team have developed a system architecture that would allow ATM-X to leverage the necessary capabilities of an Increasingly Autonomous System (IAS), machine-agent that will promote the safe access and operation of air vehicles within what has become the byproduct of NextGen modernization, a Net-Centric airspace architecture and an Urban Air Mobility (UAM) community. Conducting flight operations within this type of architecture constrains the human-agent’s natural ability by data management. When the massive volume of data, its types, and the acquisition speed at which the data is ingested is observed it becomes evident that the human-agent will be functioning at an operational disadvantage. Therefore, the development and integration of intelligent machine-agents into the flight deck are a necessary implementation to achieve ATM-X overall objective of safe access and operation in the NAS.

Urban Air Mobility

Communication Network Awareness Machine System Phase I Development: The Intelligent Party-Line Schema

As NextGen continues toward the full implementation of a Net-Centric Architecture (N-CA)it will inherently provide a continuous increase to the Three-Vs components (Volume, Velocity, and Variety) of big data . This will create an insurmountable environment for direct-action aviation personnel (DAAP)as the DAAP’s natural abilities to manage and process data into actionable information will be overmatched by the Three-Vs. Therefore, conducting operations within a N-CA requires that new tools and applications be researched and developed to aid the DAAP’s ability to understand and manage data, mitigate non-normals, create contingency plans and actions. This paper will describe a research area at NASA Langley Research Center known as the Intelligent Party-Line (IPL).

Intelligent Party-Line

Planning Bias: Planning as a Source of Sampling Bias

Many data-driven planning methods are trained on data generated by planners. It is well known that many statistical learning methods are sensitive to sampling bias, and yet there has been little or no attention to planning as a sampling method and its role in introducing sampling bias into planner-generated training data. Recently, it has been demonstrated that A**,* in the presence of problems with variable heuristic error, prefers some solutions over other equally cost-optimal solutions. But, as we discuss in this paper, mitigation may not be as simple as resolving arbitrary tie-breaking by sampling from ties uniformly at random. In this paper, we formalize an intuition of planning bias. We focus on problems which output a single solution. Diverse planning only complicates the problem by generalizing it to bias in the set of sets; we show how it is subject to bias in the single solution. We make some useful observations about deterministic algorithms in contrast to non-deterministic algorithms. We explain how information entropy may be a good way to measure planning bias, and discuss some issues in evaluating practical approaches to measurement. We address the intuition that uniform random tiebreaking should mitigate bias; and sketch a novel approach to constructing an appropriate random distribution for duplicate detection during forward search for unbiased A*. Finally, we suggest directions for future work.

Planning Scheduling Algorithms

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS

Introduction to Analysis Methods for Big Earth Data

Big Earth Data are too big to be tractable to simple data inspection and require models to make sense of all the data. Useful models for Big Earth Data may be physical, statistical, or machine learning based. In many cases, hybrid models combine attributes of two or more of these types.

parallel processing (computers)

Fuzzy self-learning control for magnetic servo system

It is known that an effective control system is the key condition for successful implementation of high-performance magnetic servo systems. Major issues to design such control systems are nonlinearity; unmodeled dynamics, such as secondary effects for copper resistance, stray fields, and saturation; and that disturbance rejection for the load effect reacts directly on the servo system without transmission elements. One typical approach to design control systems under these conditions is a special type of nonlinear feedback called gain scheduling. It accommodates linear regulators whose parameters are changed as a function of operating conditions in a preprogrammed way. In this paper, an on-line learning fuzzy control strategy is proposed. To inherit the wealth of linear control design, the relations between linear feedback and fuzzy logic controllers have been established. The exercise of engineering axioms of linear control design is thus transformed into tuning of appropriate fuzzy parameters. Furthermore, fuzzy logic control brings the domain of candidate control laws from linear into nonlinear, and brings new prospects into design of the local controllers. On the other hand, a self-learning scheme is utilized to automatically tune the fuzzy rule base. It is based on network learning infrastructure; statistical approximation to assign credit; animal learning method to update the reinforcement map with a fast learning rate; and temporal difference predictive scheme to optimize the control laws. Different from supervised and statistical unsupervised learning schemes, the proposed method learns on-line from past experience and information from the process and forms a rule base of an FLC system from randomly assigned initial control rules.

Tarn, J. H.

MARGInS Model-Based Analysis of Realizable Goals in Systems

The high complexity of modern aircraft and spacecraft requires elaborate Verification and Validation (V&V) approaches to make sure that such complex systems work properly and reliably. MARGInS is a framework for the analysis, understanding, and prediction of the behavior of a complex, hybrid system. MARGInS contains a set of machine learning and statistical algorithms for multivariate clustering, treatment learning, critical factor determination, time-series analysis, event prediction, and safety-boundary detection and characterization. The framework supports system testing and can be configured to find novel features in test suites, determine classes of behavior, propose new experiments that can efficiently explore and characterize the boundaries between classes of system behavior, and to create visualizations and reports.

He, Yuning