Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Functional data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

From multivariate to functional data analysis: Fundamentals, recent developments, and emerging areas

Functional data analysis (FDA), which is a branch of statistics on modeling infinite dimensional random vectors resided in functional spaces, has become a major research area for Journal of Multivariate Analysis. We review some fundamental concepts of FDA, their origins and connections from multivariate analysis, and some of its recent developments, including multi-level functional data analysis, high-dimensional functional regression, and dependent functional data analysis. Here, we also discuss the impact of these new methodology developments on genetics, plant science, wearable device data analysis, image data analysis, and business analytics. Two real data examples are provided to motivate our discussions.

97 MATHEMATICS AND COMPUTING↗

Covariate Dependent Sparse Functional Data Analysis

This study proposes a method to incorporate covariate information into sparse functional data analysis. The method aims at cases where each subject has a limited number of longitudinal measurements and is associated with static covariates. This research is motivated by several use cases in practice. One representative example is void swelling, a nuclear-specific material degradation mechanism. Void swelling is affected by many covariates, including alloy composition and irradiation type. How to accurately model the complicated joint effects of such covariates on the swelling process is the key to mitigating the effect of swelling and ensuring safe operation. Unlike most of the existing methods, the proposed method can handle high-dimensional covariates with the informative covariate identification procedure and sparse and irregularly spaced measurements, that is, does not require complete or dense observations. The main innovation of the proposed method is that we model the variation coming from covariates and the variation left conditioned on covariates, such that the functional principal component analysis and Gaussian process can be conducted in a unified manner. Further, we also propose a systematic approach to identify important covariates in the hypothesis testing context. The methodology is demonstrated on applications in nuclear engineering and healthcare and simulation studies.

42 ENGINEERING↗

Harnessing Uncertainty through Functional Data Analysis in Gas Breakthrough Data

Detecting subsurface explosions from radionuclide gas migration through rock fractures is an effective way to identify nuclear activity. Los Alamos National Laboratory (LANL) has developed simulation methods, based on data from the 1962 Hardhat underground nuclear test, to predict gas breakthrough times at the surface. However, these methods rely on an imperfect understanding of the relationship between rock damage and fracture permeability. Our clinic project studies methods for predicting breakthrough curves that characterize total mass produced as a function of time, as well as quantifying the uncertainty associated with these predictions. The model that is currently employed to relate damage to permeability uses an empirically motivated power-law expression, with a range of parameter values that are compatible with the experimental Hardhat data. We develop emulators, built from functional data analysis techniques and trained on simulation data, that rapidly predict the gas breakthrough curve given a damage field and given the parameter values of the power-law equation. Using Bayesian regression, we address the problem of uncertainty quantification in our emulators. Finally, in order to test the robustness of the model, we further validate it on a damage field representing different physical conditions.

58 GEOSCIENCES↗

Functional Data Analysis in Wearable Body Sensor Networks

Improving response time of indirect room-size calorimeters is still an outstanding problem in metabolic research. Accurate estimates of instantaneous rates of gaseous exchange require numerical differentiation of measured gaseousgas concentrations. We propose a new method to estimate the instantaneous gaseousgas exchange rates in indirect calorimetry. In contrast to the previously developed techniques, the method addresses the problem of differentiation of gaseous concentrations as an ill-posed problem. By applying the method of regularization, the problem of differentiation is converted into a well-posed problem resulting in smooth and consistent gaseous exchange rates. The validity of the method is tested on a large dataset of calorimeter experiments which included 313 human experiments along with 231 alcohol combustion experiments. It is demonstrated that the method is able to reliably differentiate between the “unphysiological” process of alcohol combustion and physiological variations produced by human metabolism. The method also allowed unraveling the previously unreported relative kinetics of O2 consumption and Respiratory Quotient (RQ) in humans. It was found that the kinetics of oxidative fuel selection lags behind the energy expenditure in humans exhibiting some sort of oxidative inertia. The time lag varies from 2-3 min up to 30 min, depending on particular individual. No such lag was found in alcohol combustion experiments. In addition to the relative kinetics of substrate oxidation, two statistical indexes reflecting variability of minute-by-minute RQ were estimated. The indexes were the RQ’s standard deviation and RQ’s first-order derivative. Both indexes showed statistically significant difference between human experiments and alcohol combustion experiments. We conclude that the proposed method can consistently extract physiologically-relevant information from noisy calorimetry data and the aforesaid information can provide additional insights into the mechanism of metabolic fuel selection in humans.

54 ENVIRONMENTAL SCIENCES↗

Functional Data Analysis in Wearable Body Sensor Networks

Improving response time of indirect room-size calorimeters is still an outstanding problem in metabolic research. Accurate estimates of instantaneous rates of gaseous exchange require numerical differentiation of measured gaseousgas concentrations. We propose a new method to estimate the instantaneous gaseousgas exchange rates in indirect calorimetry. In contrast to the previously developed techniques, the method addresses the problem of differentiation of gaseous concentrations as an ill-posed problem. By applying the method of regularization, the problem of differentiation is converted into a well-posed problem resulting in smooth and consistent gaseous exchange rates. The validity of the method is tested on a large dataset of calorimeter experiments which included 313 human experiments along with 231 alcohol combustion experiments. It is demonstrated that the method is able to reliably differentiate between the “unphysiological” process of alcohol combustion and physiological variations produced by human metabolism. The method also allowed unraveling the previously unreported relative kinetics of O2 consumption and Respiratory Quotient (RQ) in humans. It was found that the kinetics of oxidative fuel selection lags behind the energy expenditure in humans exhibiting some sort of oxidative inertia. The time lag varies from 2-3 min up to 30 min, depending on particular individual. No such lag was found in alcohol combustion experiments. In addition to the relative kinetics of substrate oxidation, two statistical indexes reflecting variability of minute-by-minute RQ were estimated. The indexes were the RQ’s standard deviation and RQ’s first-order derivative. Both indexes showed statistically significant difference between human experiments and alcohol combustion experiments. We conclude that the proposed method can consistently extract physiologically-relevant information from noisy calorimetry data and the aforesaid information can provide additional insights into the mechanism of metabolic fuel selection in humans.

60 - APPLIED LIFE SCIENCES↗

Functional Data Analysis for Extracting the Intrinsic Dimensionality of Spectra: Application to Chemical Homogeneity in the Open Cluster M67

High-resolution spectroscopic surveys of the Milky Way have entered the Big Data regime and have opened avenues for solving outstanding questions in Galactic archeology. However, exploiting their full potential is limited by complex systematics, whose characterization has not received much attention in modern spectroscopic analyses. In this work, we present a novel method to disentangle the component of spectral data space intrinsic to the stars from that due to systematics. Using functional principal component analysis on a sample of 18,933 giant spectra from APOGEE, we find that the intrinsic structure above the level of observational uncertainties requires ≈10 functional principal components (FPCs). Our FPCs can reduce the dimensionality of spectra, remove systematics, and impute masked wavelengths, thereby enabling accurate studies of stellar populations. To demonstrate the applicability of our FPCs, we use them to infer stellar parameters and abundances of 28 giants in the open cluster M67. We employ Sequential Neural Likelihood, a simulation-based Bayesian inference method that learns likelihood functions using neural density estimators, to incorporate non-Gaussian effects in spectral likelihoods. By hierarchically combining the inferred abundances, we limit the spread of the following elements in M67: Fe ≲ 0.02 dex; C ≲ 0.03 dex; O, Mg, Si, Ni ≲ 0.04 dex; Ca ≲ 0.05 dex; N, Al ≲ 0.07 dex (at 68% confidence). Our constraints suggest a lack of self-pollution by core-collapse supernovae in M67, which has promising implications for the future of chemical tagging to understand the star formation history and dynamical evolution of the Milky Way.

79 ASTRONOMY AND ASTROPHYSICS↗

Symmetry-mode analysis for local structure investigations using pair distribution function data

Symmetry-adapted distortion modes provide a natural way of describing distorted structures derived from higher-symmetry parent phases. Structural refinements using symmetry-mode amplitudes as fit variables have been used for at least ten years in Rietveld refinements of the average crystal structure from diffraction data; more recently, this approach has also been used for investigations of the local structure using real-space pair distribution function (PDF) data. Here, the value of performing symmetry-mode fits to PDF data is further demonstrated through the successful application of this method to two topical materials: TiSe2, where a subtle but long-range structural distortion driven by the formation of a charge-density wave is detected, and MnTe, where a large but highly localized structural distortion is characterized in terms of symmetry-lowering displacements of the Te atoms. Here, the analysis is performed using fully open-source code within the DiffPy framework via two packages developed for this work: isopydistort, which provides a scriptable interface to the ISODISTORT web application for group theoretical calculations, and isopytools, which converts the ISODISTORT output into a DiffPy-compatible format for subsequent fitting and analysis. These developments expand the potential impact of symmetry-adapted PDF analysis by enabling high-throughput analysis and removing the need for any commercial software.

36 MATERIALS SCIENCE↗

Elastic functional changepoint detection of climate impacts from localized sources

Detecting changepoints in functional data has become an important problem as interest in monitoring of climate phenomenon has increased, where the data is functional in nature. Here, the observed data often contains both amplitude (y-axis) and phase (x-axis) variability. If not accounted for properly, true changepoints may be undetected, and the estimated underlying mean change functions will be incorrect. In this article, an elastic functional changepoint method is developed which properly accounts for these types of variability. The method can detect amplitude and phase changepoints which current methods in the literature do not, as they focus solely on the amplitude changepoint. This method can easily be implemented using the functions directly or can be computed via functional principal component analysis to ease the computational burden. We apply the method and its nonelastic competitors to both simulated data and observed data to show its efficiency in handling data with phase variation with both amplitude and phase changepoints. We use the method to evaluate potential changes in stratospheric temperature due to the eruption of Mt. Pinatubo in the Philippines in June 1991. Using an epidemic changepoint model, we find evidence of a increase in stratospheric temperature during a period that contains the immediate aftermath of Mt. Pinatubo, with most detected changepoints occurring in the tropics as expected.

54 ENVIRONMENTAL SCIENCES↗

Analysis of Variance of Functional Data (F-ANOVA) [Slides]

Goals: What is Analysis of Variance (ANOVA); Extending Analysis of Variance to Function Data (F-ANOVA); The role of stochastic processes in F-ANOVA; One and Two Sample Problems for F-ANOVA; One-Way F-ANOVA.

97 MATHEMATICS AND COMPUTING↗

Inverse prediction of PuO2 processing conditions using Bayesian seemingly unrelated regression with functional data

Over the past decade, a variety of innovative methodologies have been developed to better characterize the relationships between processing conditions and the physical, morphological, and chemical features of special nuclear material (SNM). Different processing conditions generate SNM products with different features, which are known as “signatures” because they are indicative of the processing conditions used to produce the material. These signatures can potentially allow a forensic analyst to determine which processes were used to produce the SNM and make inferences about where the material originated. This article investigates a statistical technique for relating processing conditions to the morphological features of PuO 2 particles. We develop a Bayesian implementation of seemingly unrelated regression (SUR) to inverse-predict unknown PuO 2 processing conditions from known PuO 2 features. Model results from simulated data demonstrate the usefulness of the technique. Applied to empirical data from a bench-scale experiment specifically designed with inverse prediction in mind, our model successfully predicts nitric acid concentration, while results for Pu concentration and precipitation temperature were equivalent to a simple mean model. Our technique compliments other recent methodologies developed for forensic analysis of nuclear material and can be generalized across the field of chemometrics for application to other materials.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Predicting fatigue from heart rate signatures using functional logistic regression

Physical fatigue can have adverse effects on humans in extreme environments. Therefore, being able to predict fatigue using easy to measure metrics such as heart rate (HR) signatures has potential to have an impact in real-life scenarios. We apply a functional logistic regression model that uses HR signatures to predict physical fatigue, where physical fatigue is defined in a data-driven manner. Data were collected using commercially available wearable devices on 47 participants hiking the 20.7-mile Grand Canyon rim-to-rim trail in a single day. Fitted model provides good predictions and interpretable parameters for real-life application.

60 APPLIED LIFE SCIENCES↗

Autonomous phase mapping of gold nanoparticles synthesis with differentiable models of spectral shape

Autonomous experimentation–or self-driving labs–offers a systematic approach to accelerate materials discovery by integrating automated synthesis, characterization, and data-driven decision-making. We present a closed-loop workflow for the on-demand synthesis and structural characterization of colloidal gold nanoparticles, enabling direct mapping from composition to nanoscale structure. Our framework leverages differentiable models of spectral shape to address two central tasks in self-driving labs: (a) phase mapping, or identifying compositional regions with distinct structural behavior; and (b) material retrosynthesis, or optimizing compositions for target structure. Using functional data analysis, we develop a data-driven model with generative pre-training, active learning, and high-throughput experiments to predict spectral responses across composition space. We demonstrate the approach on seed-mediated growth of gold nanoparticles, showcasing its ability to extract design rules, reveal secondary interactions, and efficiently navigate morphology space. Gradient-based optimization of the models enables inverse design, making this a unified platform.

36 MATERIALS SCIENCE↗

Visualisation and outlier detection for probability density function ensembles

Abstract Exploratory data analysis (EDA) for functional data—data objects where observations are entire functions—is a difficult problem that has seen significant attention in recent literature. This surge in interest is motivated by the ubiquitous nature of functional data, which are prevalent in applications across fields such as meteorology, biology, medicine and engineering. Empirical probability density functions (PDFs) can be viewed as constrained functional data objects that must integrate to one and be nonnegative. They show up in contexts such as yearly income distributions, zooplankton size structure in oceanography and in connectivity patterns in the brain, among others. While PDF data are certainly common in modern research, little attention has been given to EDA specifically for PDFs. In this paper, we extend several methods for EDA on functional data for PDFs and compare them on simulated data that exhibit different types of variation, designed to mimic that seen in real‐world applications. We then use our new methods to perform EDA on the breakthrough curves observed in gas transport simulations for underground fracture networks.

97 MATHEMATICS AND COMPUTING↗