Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Nearest-Neighbor Machine Learning Feature Selection for Interpretation of Microbial Molecular Signatures from Isotope Ratio Mass Spectrometry Data

Mass spectrometry (MS) promises to be a powerful tool for potential biosignature detection during astrobiological missions on ocean worlds in our solar system. Accurate and generalizable machine learning methods could enhance science return on investment by predicting seawater chemistry and classifying isotopic biosignatures, either as a signature consistent with microbial life (biotic) or as a novelty (unclassified/unique). However, machine learning models are likely to be complex and involve interactions between MS features, making biosignatures difficult to interpret. Feature selection methods provide biological and chemical context that help interpret the mechanisms of machine learning models, but these methods also need the ability to detect complex interactions. Previously, we developed a machine learning feature selection algorithm called nearest-neighbor projected distance regression (NPDR) that has the ability to identify important model features that involve complex interactions and automatically reduce correlation and the dimensionality in a high-dimensional variable space. The standard distance metrics used in NPDR – Manhattan and Euclidean – assume the multivariate data are isotropic, which is often violated in real data due to differences in the covariance between variables. Thus, we extend NPDR to include a random forest distance, and other anisotropic distance metrics, for computing nearest neighbors. We also augment the isotope-ratio MS data with time-series features from the raw MS signal to improve biotic classification. We test NPDR on our novel experimental ocean world seawater analog MS data. We measure isotope fractionations of volatile CO 2 that could be measured in exospheres or plumes. Samples include baseline abiotic conditions using a range of possible seawater chemistry consistent with Europa and Enceladus, and biotic samples that include microbes in these seawaters. We use penalized NPDR with random forest proximity to identify interpretable microbial molecular signatures. We compare features with random forest importance, and we train a classifier that discriminates between biotic and abiotic samples with high accuracy. These ML-trained ocean-world analog MS data could be used to assist in identifying biosignatures during future missions.

geochemistry↗

Validation of MODIS FLH and In Situ Chlorophyll a from Tampa Bay, Florida (USA)

Satellite observation of phytoplankton concentration or chlorophyll-a (chla) is an important characteristic, critically integral to monitoring coastal water quality. However, the optical properties of estuarine and coastal waters are highly variable and complex and pose a great challenge for accurate analysis. Constituents such as suspended solids and dissolved organic matter and the overlapping and uncorrelated absorptions in the blue region of the spectrum renders the blue-green ratio algorithms for estimating chl-a inaccurate. Measurement of suninduced chlorophyll fluorescence, on the other hand, which utilizes the near infrared portion of the electromagnetic spectrum may, provide a better estimate of phytoplankton concentrations. While modelling and laboratory studies have illustrated both the utility and limitations of satellite algorithms based on the sun induced chlorophyll fluorescence signal, few have examined the empirical validity of these algorithms or compared their accuracy against bluegreen ratio algorithms . In an unprecedented analysis using a long term (2003-2011) in situ monitoring data set from Tampa Bay, Florida (USA), we assess the validity of the FLH product from the Moderate Resolution Imaging Spectrometer against a suite of water quality parameters taken in a variety of conditions throughout this large optically complex estuarine system. . Overall, the results show a 106% increase in the validity of chla concentration estimation using FLH over the standard chla estimate from the blue-green OC3M algorithm. Additionally, a systematic analysis of sampling sites throughout the bay is undertaken to understand how the FLH product responds to varying conditions in the estuary and correlations are conducted to see how the relationships between satellite FLH and in situ chlorophyll-a change with depth, distance from shore, from structures like bridges, and nutrient concentrations and turbidity. Such analysis illustrates that the correlations between FLH and in situ chla measurements increases with increasing distance between monitoring sites and structures like bridges and shore. Due probably to confounding factors, expected improvement in the FLH- chla relationship was not clearly noted when increasing depth and distance from shore alone (not including bridges). Correlations between turbidity and nutrient concentrations are discussed further and principle component analyses are employed to address the relationships between the multivariate data sets. A thorough understanding of how satellite FLH algorithms relate to in situ water quality parameters will enhance our understanding of how MODIS s global FLH algorithm can be used empirically to monitor coastal waters worldwide.

Fischer, Andrew↗

Intrinsic Kinetics of Polyethylene Terephthalate Pyrolysis via Micropyrolysis and Multivariate Chromatographic Analysis

This study provides an in-depth investigation of the primary decomposition of polyethylene terephthalate (PET) via pyrolysis, employing an experimental-analytic workflow that integrates design of experiments (DoE), micropyrolysis coupled with comprehensive two-dimensional gas chromatography (GC×GC), and multivariate data analysis to verify intrinsic kinetic conditions and elucidate evolving product distributions for mapping key reaction pathways. Peaks that could not be identified using commercial spectral libraries were assigned using Mass Frontier simulations, enabling the identification of divinyl terephthalate, ethyl vinyl terephthalate, and 2-(benzoyloxy)ethyl vinyl terephthalate. A polar×polar (non-orthogonal) column set tailored for the detection of carboxylic acids enhanced the quantification of benzoic acid, 4-vinylbenzoic acid, 4-ethylbenzoic acid, and methylbenzoic acid by up to 6-fold relative to an orthogonal column combination (non-polar×mid-polar). Moreover, pyrolysis variables were systematically evaluated using a Box- Behnken design (BBD), encompassing pyrolysis temperature (500−600 °C), sample weight (50−150 μg), and carrier gas flow rate (100−300 mL min −1 ). Among these, pyrolysis temperature was the only statistically significant factor influencing product yields, ranging from 58.78 to 84.26 wt %. In contrast, neither the sample weight nor the carrier gas flow rate had a significant effect on product yields within the evaluated experimental space. At 600 °C, the major pyrolysis products were benzoic acid (up to 20.20 ± 1.46 wt %) and CO 2 (up to 21.28 ± 1.46 wt %), which can be produced through decarboxylation reactions. These findings underscore the critical importance of selecting appropriate analytical columns for the accurate quantification of heteroatomcontaining products such as carboxylic acids, which may otherwise be underestimated or undetected due to their reactivity with the stationary phase of non-polar and mid-polar columns, as well as other GC components. They also highlight the importance of selecting pyrolysis conditions for investigating the primary decomposition of PET under an isothermal kinetically limited regime.

aromatic compounds↗

Node Distortion as a Tunable Mechanism for Negative Thermal Expansion in Metal–Organic Frameworks

Chemically functionalized series of metal–organic frameworks (MOFs), with subtle differences in local structure but divergent properties, provide a valuable opportunity to explore how local chemistry can be coupled to long-range structure and functionality. Using in situ synchrotron X-ray total scattering, with powder diffraction and pair distribution function (PDF) analysis, we investigate the temperature dependence of the local- and long-range structure of MOFs based on NU-1000, in which Zr 6 O 8 nodes are coordinated by different capping ligands (H 2 O/OH, Cl – ions, formate, acetylacetonate, and hexafluoroacetylacetonate). We show that the local distortion of the Zr 6 nodes depends on the lability of the ligand and contributes to a negative thermal expansion (NTE) of the extended framework. Using multivariate data analyses, involving non-negative matrix factorization (NMF), we demonstrate a new mechanism for NTE: progressive increase in the population of a smaller, distorted node state with increasing temperature leads to global contraction of the framework. The transformation between discrete node states is noncooperative and not ordered within the lattice, i.e., a solid solution of regular and distorted nodes. Density functional theory calculations show that removal of ligands from the node can lead to distortions consistent with the Zr···Zr distances observed in the experiment PDF data. Control of the node distortion imparted by the nonlinker ligand in turn controls the NTE behavior. Furthermore, these results reveal a mechanism to control the dynamic structure of MOFs based on local chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Inferring Plant Acclimation and Improving Model Generalizability With Differentiable Physics‐Informed Machine Learning of Photosynthesis

Net photosynthesis (A N ) is a key component of the global carbon cycle influencing climate feedback over decadal scales. Although plant acclimation to environmental changes can modify A N , traditional vegetation models in Earth system models (ESMs) often rely on plant functional type (PFT)-specific parameterizations or simplified acclimation assumptions limiting generalizability across time, space, and PFTs. In this study, we developed a differentiable photosynthesis model to learn the environmental dependencies of V c,max25 (maximum carboxylation rate at 25°C, representing photosynthetic capacity), as this genre of hybrid physics-informed machine learning can seamlessly train neural networks and process-based equations together. Compared to PFT-specific parameterization of V c,max25 , learning the environment dependencies of key photosynthetic parameters improved model spatiotemporal generalizability. Applying environmental acclimation to V c,max25 led to substantial variations in global mean A N indicating the need to address acclimation in ESMs. The model effectively captured multivariate observations (V c,max25 , A N , and stomatal conductance (g s )) simultaneously with multivariate constraints, improving generalization across space and PFTs. It also learned sensible acclimation relationships of V c,max25 to different environmental conditions. The model explained more than 54%, 57%, and 62% of the variance of A N , g s , and V c,max25 , respectively, presenting a first global-scale spatial test benchmark of A N and g s . These results highlight the potential for differentiable modeling to enhance process-based modules in ESMs and effectively leverage information from large, multivariate data sets.

54 ENVIRONMENTAL SCIENCES↗

Causal interaction in high frequency turbulence at the biosphere–atmosphere interface: Structural behavior

High-frequency (e.g., 10 Hz) eddy covariance measurements are typically used to estimate fluxes at the land–atmosphere interface at timescales of 15–60 min. These multivariate data contain information about the interdependency at high frequency between the interacting variables such as wind, humidity, temperature, and CO 2⁠ . We use data at 10 Hz from an eddy covariance instrument located at 25 m above agricultural land in the Midwestern US, which offers an opportunity to move beyond the traditional spectral analyses to explore causal dependency among variables. In this study, we quantify the structure of inter-dependencies of interacting variables at high frequency represented by a directed acyclic graph (DAG). We compare DAGs to investigate changes in structural differences in causal interactions. We then apply a distance-based classification and -means clustering approach to identify the evolution of the causal structure represented by a DAG. Our method selects an unbiased number of clusters of similar structures and characterizes the similarities and differences between them. We explore a range of dynamic behavior using data from a clear sky day and during a solar eclipse in 2017. Our results show well-defined clusters of similar causal dependencies as the system evolves. Furthermore, our approach provides a methodological framework to understand how causal dependence in turbulence manifests in high-frequency data when represented through a DAG.

54 ENVIRONMENTAL SCIENCES↗

Causal interaction in high frequency turbulence at the biosphere–atmosphere interface: Structure–function coupling

At the biosphere–atmosphere interface, nonlinear interdependencies among components of an ecohydrological complex system can be inferred using multivariate high frequency time series observations. Information flow among these interacting variables allows us to represent the causal dependencies in the form of a directed acyclic graph (DAG). Here, we use high frequency multivariate data at 10 Hz from an eddy covariance instrument located at 25 m above agricultural land in the Midwestern US to quantify the evolutionary dynamics of this complex system using a sequence of DAGs by examining the structural dependency of information flow and the associated functional response. We investigate whether functional differences correspond to structural differences or if there are no functional variations despite the structural differences. We base our analysis on the hypothesis that causal dependencies are instigated through information flow, and the resulting interactions sustain the dynamics and its functionality. To test our hypothesis, we build upon causal structure analysis in the companion paper to characterize the information flow in similarly clustered DAGs from 3-min non-overlapping contiguous windows in the observational data. We characterize functionality as the nature of interactions as discerned through redundant, unique, and synergistic components of information flow. Through this analysis, we find that in turbulence at the biosphere–atmosphere interface, the variables that control the dynamic character of the atmosphere as well as the thermodynamics are driven by non-local conditions, while the scalar transport associated with CO and H 2 O is mainly driven by short-term local conditions.

58 GEOSCIENCES↗

Sparsifying priors for Bayesian uncertainty quantification in model discovery

We propose a probabilistic model discovery method for identifying ordinary differential equations governing the dynamics of observed multivariate data. Our method is based on the sparse identification of nonlinear dynamics (SINDy) framework, where models are expressed as sparse linear combinations of pre-specified candidate functions. Promoting parsimony through sparsity leads to interpretable models that generalize to unknown data. Instead of targeting point estimates of the SINDy coefficients, we estimate these coefficients via sparse Bayesian inference. The resulting method, uncertainty quantification SINDy (UQ-SINDy), quantifies not only the uncertainty in the values of the SINDy coefficients due to observation errors and limited data, but also the probability of inclusion of each candidate function in the linear combination. UQ-SINDy promotes robustness against observation noise and limited data, interpretability (in terms of model selection and inclusion probabilities) and generalization capacity for out-of-sample forecast. Sparse inference for UQ-SINDy employs Markov chain Monte Carlo, and we explore two sparsifying priors: the spike and slab prior, and the regularized horseshoe prior. UQ-SINDy is shown to discover accurate models in the presence of noise and with orders-of-magnitude less data than current model discovery methods, thus providing a transformative method for real-world applications which have limited data.

97 MATHEMATICS AND COMPUTING↗

Machine-Learning Assisted Identification of Accurate Battery Lifetime Models with Uncertainty

Reduced-order battery lifetime models, which consist of algebraic expressions for various aging modes, are widely utilized for extrapolating degradation trends from accelerated aging tests to real-world aging scenarios. Identifying models with high accuracy and low uncertainty is crucial for ensuring that model extrapolations are believable, however, it is difficult to compose expressions that accurately predict multivariate data trends; a review of cycling degradation models from literature reveals a wide variety of functional relationships. Here, a machine-learning assisted model identification method is utilized to fit degradation in a stand-out LFP-Gr aging data set, with uncertainty quantified by bootstrap resampling. The model identified in this work results in approximately half the mean absolute error of a human expert model. Models are validated by converting to a state-equation form and comparing predictions against cells aging under varying loads. Parameter uncertainty is carried forward into an energy storage system simulation to estimate the impact of aging model uncertainty on system lifetime. The new model identification method used here reduces life-prediction uncertainty by more than a factor of three (86% ± 5% relative capacity at 10 years for human-expert model, 88.5% ± 1.5% for machine-learning assisted model), empowering more confident estimates of energy storage system lifetime.

25 ENERGY STORAGE↗

Editorial: Applications of spectroscopy and chemometrics in nuclear materials analysis

Optical analysis techniques, including spectroscopy and image analysis, have many advantages when applied to the study of nuclear materials. They require small sample sizes, can be performed remotely, and can be proceduralized through consistent practice. Most importantly, they provide a wealth of information by generating multivariate data. For example, ultraviolet–visible–near-infrared absorbance spectroscopy of actinides in aqueous and organic solutions is dependent on the oxidation state, anionic complexation, and temperature. These variables are important for solution-based separation processes, and sensitivity to these factors, combined with online monitoring, can drive the efficiency and control of these processes. The morphology and chemical composition of actinide particles can also provide a vital clue to the mechanisms by which the particles were formed, providing forensic information on the origins of the particles.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Cholesky-based experimental design for Gaussian process and kernel-based emulation and calibration.

Gaussian processes and other kernel-based methods are used extensively to construct approximations of multivariate data sets. The accuracy of these approximations is dependent on the data used. This paper presents a computationally efficient algorithm to greedily select training samples that minimize the weighted L p error of kernel-based approximations for a given number of data. The method successively generates nested samples, with the goal of minimizing the error in high probability regions of densities specified by users. The algorithm presented is extremely simple and can be implemented using existing pivoted Cholesky factorization methods. Training samples are generated in batches which allows training data to be evaluated (labeled) in parallel. For smooth kernels, the algorithm performs comparably with the greedy integrated variance design but has significantly lower complexity. Numerical experiments demonstrate the efficacy of the approach for bounded, unbounded, multi-modal and non-tensor product densities. We also show how to use the proposed algorithm to efficiently generate surrogates for inferring unknown model parameters from data using Bayesian inference.

97 MATHEMATICS AND COMPUTING↗

A method of determining spectral dye densities in color films

A mathematical analysis technique called characteristic vector analysis, reported by Simonds (1963), is used to determine spectral dye densities in multiemulsion film such as color or color-IR imagery. The technique involves examining a number of sets of multivariate data and determining linear transformations of these data to a smaller number of parameters which contain essentially all of the information contained in the original set of data. The steps involved in the actual procedure are outlined. It is shown that integral spectral density measurements of a large number of different color samples can be accurately reconstructed from the calculated spectral dye densities.

Friederichs, G. A.↗

Computer program documentation: ISOCLS iterative self-organizing clustering program, program C094

The author has identified the following significant results. This program implements an algorithm which, ideally, sorts a given set of multivariate data points into similar groups or clusters. The program is intended for use in the evaluation of multispectral scanner data; however, the algorithm could be used for other data types as well. The user may specify a set of initial estimated cluster means to begin the procedure, or he may begin with the assumption that all the data belongs to one cluster. The procedure is initiatized by assigning each data point to the nearest (in absolute distance) cluster mean. If no initial cluster means were input, all of the data is assigned to cluster 1. The means and standard deviations are calculated for each cluster.

Minter, R. T.↗

A fast routine for computing

A routine for calculating multidimensional histograms of multivariate data using a combination table look up and search procedure is described. The software was originally developed to computer four-dimensional histograms from LANDSAT multispectral imagery, but the concept can be used on other types of data and the program can be modified for the desired type of output information.

Jayroe, R. R., Jr.↗

Computer program documentation for the patch subsampling processor

The programs presented are intended to provide a way to extract a sample from a full-frame scene and summarize it in a useful way. The sample in each case was chosen to fill a 512-by-512 pixel (sample-by-line) image since this is the largest image that can be displayed on the Integrated Multivariant Data Analysis and Classification System. This sample size provides one megabyte of data for manipulation and storage and contains about 3% of the full-frame data. A patch image processor computes means for 256 32-by-32 pixel squares which constitute the 512-by-512 pixel image. Thus, 256 measurements are available for 8 vegetation indexes over a 100-mile square.

Nieves, M. J.↗

Configuration space representation in parallel coordinates

By means of a system of parallel coordinates, a nonprojective mapping from R exp N to R squared is obtained for any positive integer N. In this way multivariate data and relations can be represented in the Euclidean plane (embedded in the projective plane). Basically, R squared with Cartesian coordinates is augmented by N parallel axes, one for each variable. The N joint variables of a robotic device can be represented graphically by using parallel coordinates. It is pointed out that some properties of the relation are better perceived visually from the parallel coordinate representation, and that new algorithms and data structures can be obtained from this representation. The main features of parallel coordinates are described, and an example is presented of their use for configuration space representation of a mechanical arm (where Cartesian coordinates cannot be used).

Fiorini, Paolo↗

User's manual for the Gaussian windows program

'Gaussian Windows' is a method for exploring a set of multivariate data, in order to estimate the shape of the underlying density function. The method can be used to find and describe structural features in the data. The method is described in two earlier papers. I assume that the reader has access to both of these papers, so I will not repeat material from them. The program described herein is written in BASIC and it runs on an IBM PC or PS/2 with the DOS 3.3 operating system. Although the program is slow and has limited memory space, it is adequate for experimenting with the method. Since it is written in BASIC, it is relatively easy to modify. The program and some related files are available on a 3-inch diskette. A listing of the program is also available. This user's manual explains the use of the program. First, it gives a brief tutorial, illustrating some of the program's features with a set of artificial data. Then, it describes the results displayed after the program does a Gaussian window, and it explains each of the items on the various menus.

Jaeckel, Louis A.↗

The Grand Tour via Geodesic Interpolation of 2-frames

Grand tours are a class of methods for visualizing multivariate data, or any finite set of points in n-space. The idea is to create an animation of data projections by moving a 2-dimensional projection plane through n-space. The path of planes used in the animation is chosen so that it becomes dense, that is, it comes arbitrarily close to any plane. One of the original inspirations for the grand tour was the experience of trying to comprehend an abstract sculpture in a museum. One tends to walk around the sculpture, viewing it from many different angles. A useful class of grand tours is based on the idea of continuously interpolating an infinite sequence of randomly chosen planes. Visiting randomly (more precisely: uniformly) distributed planes guarantees denseness of the interpolating path. In computer implementations, 2-dimensional orthogonal projections are specified by two 1-dimensional projections which map to the horizontal and vertical screen dimensions, respectively. Hence, a grand tour is specified by a path of pairs of orthonormal projection vectors. This paper describes an interpolation scheme for smoothly connecting two pairs of orthonormal vectors, and thus for constructing interpolating grand tours. The scheme is optimal in the sense that connecting paths are geodesics in a natural Riemannian geometry.

Asimov, Daniel↗