Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Data and code from: Multivariate bayesian regression model for predicting disposed ash composition at U.S. coal fired power stations

This dataset contains the code and data files needed for implementation of a Multivariate Bayesian Regression model, described in Jin et al. (2025), for the historical prediction of the chemical composition of disposed coal ash at U.S. coal fired power plants as a function of annualized coal purchase data. The integrated coal supply data file (CoalSupplyDataset.csv) represents a compilation of monthly fuel purchase records for the period 1973-2022 at major U.S. power stations. These records were obtained from the U.S. Energy Information Administration. The CSV file also contains, for each coal purchase record, the coal region of the mine as defined by the U.S. Geological Survey. Data entry errors and data gaps in the EIA records were corrected as described in Jin et al. This CSV file represents the integrated coal supply data after corrections were made. The model structure and fitting parameters are encoded in pickle file format (Bayesian.pkl). The model was developed with the coal supply data and coal ash composition data, apportioned according to the Stratified Shuffle Split for training and testing subsets. The model was built using Python and the PyMC library. Reference Publication: Jin, Z.; Huang, J.; Hower, J.C.; Hsu-Kim, H.(2025). Predictive Assessment of the Chemical Composition of Coal Ash in Reserve at U.S. Disposal Sites. Environmental Science & Technology.

Coal ash composition↗

High-dimensional multivariate autoregressive model estimation of human electrophysiological data using fMRI priors

Multivariate autoregressive (MVAR) model estimation enables assessment of causal interactions in brain networks. However, accurately estimating MVAR models for high-dimensional electrophysiological recordings is challenging due to the extensive data requirements. Hence, the applicability of MVAR models for study of brain behavior over hundreds of recording sites has been very limited. Prior work has focused on different strategies for selecting a subset of important MVAR coefficients in the model to reduce the data requirements of conventional least-squares estimation algorithms. Here we propose incorporating prior information, such as resting state functional connectivity derived from functional magnetic resonance imaging, into MVAR model estimation using a weighted group least absolute shrinkage and selection operator (LASSO) regularization strategy. The proposed approach is shown to reduce data requirements by a factor of two relative to the recently proposed group LASSO method of Endemann et al (Neuroimage 254:119057, 2022) while resulting in models that are both more parsimonious and more accurate. The effectiveness of the method is demonstrated using simulation studies of physiologically realistic MVAR models derived from intracranial electroencephalography (iEEG) data. The robustness of the approach to deviations between the conditions under which the prior information and iEEG data is obtained is illustrated using models from data collected in different sleep stages. This approach allows accurate effective connectivity analyses over short time scales, facilitating investigations of causal interactions in the brain underlying perception and cognition during rapid transitions in behavioral state.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Anomaly Detection, Localization and Classification using Drifting Synchrophasor Data Streams

With ongoing automation and digitization of the electric power system, several Phasor Measurement Units(PMUs) have been deployed for monitoring and control. PMU data can have multiple anomalies, and many of the researchers in the past have concentrated on training machine/deep learning algorithms offline for anomaly detection over PMU data (i.e., not in real time). These machine/deep learning algorithms, when trained offline on a sample rather than a population of the dataset, fail to consider the dynamic behavior of the power grid in real-time, resulting in low accuracy. Considering the dynamic behavior of the power grid (e.g., change in load, generation, distributed energy resources (DERs) switching, network, controls), the definition of data anomalies varies in time and requires online training. A fundamental challenge is to enable online (i.e., real-time) training of machine/deep learning algorithms for anomaly detection over streaming PMU data. While machine/deep learning is often desirable to manage data streams, training a deep learning algorithm over streaming PMU data is nontrivial due to changes in data statistics caused by dynamic streaming data. This paper proposes PMUNET: a novel device-level deep learning-based data-driven approach for anomaly detection, localization, and classification over streaming PMU data, using online learning and multivariate data-drift detection algorithm .Two variants of PMUNET, Dynamic data Change Driven Learning (DCDL) and Continuity Driven Learning (CDL), are proposed and compared. DCDL aims to train the deep learning algorithm whenever the definition of anomaly changes due to the power grid dynamics. On the other hand, CDL continuously trains the deep learning algorithm over the PMU data-stream. The experimental results verify that DCDL outperforms CDL and other efficient anomaly detection methods over multiple events such as faults and load/ generator/capacitor/DERs variations/switching for IEEE 14 and 39 Bus test system as well as real PMU industrial data. The result verifies that DCDL variant of PMUNET improves over existing approach with a gain of 2% - 10% in terms of accuracy, false-positive rate, and false-negative rate.

adversarial deep learning↗

Large-Scale Inference of Multivariate Regression for Heavy-Tailed and Asymmetric Data

Large-scale multivariate regression is a fundamental statistical tool with a wide range of applications. Here, this study considers the problem of simultaneously testing a large number of general linear hypotheses, encompassing covariate-effect analysis, analysis of variance, and model comparisons. The challenge that accompanies a large number of tests is the ubiquitous presence of heavy-tailed and/or highly skewed measurement noise, which is the main reason for the failure of conventional least squares-based methods. For large-scale multivariate regression, we develop a set of robust inference methods to explore data features such as heavy tailedness and skewness, which are not visible to least squares methods. The new testing procedure is based on the data-adaptive Huber regression and a new covariance estimator of regression estimates. Under mild conditions, we show that our methods produce consistent estimates of the false discovery proportion. Extensive numerical experiments and an empirical study on quantitative linguistics demonstrate the advantage of the proposed method over many state-of-the-art methods when the data are generated from heavy-tailed and/or skewed distributions.

97 MATHEMATICS AND COMPUTING↗

Advances in Land Data Assimilation Systems

Assimilation of remotely sensed land surface observations into regional to global scale numerical models have the potential to significantly advance our ability, to assess, understand, and predict surface water, energy, and carbon cycles. This session seeks to assess the state-of-the-art in data assimilation methods for integrating land surface remote sensing and modeling, with a focus on practical applications and techniques. Assimilated land surface variables of interest include (but are not limited to, soil moisture, surface temperature, snowpack, streamflow, vegetation dynamics, and carbon storage. Contributions describing the development of practical land surface data assimilation methods, multivariate land surface data assimilation strategies, evaluation of the required accuracy and resolution of remote sensing observations, the effects of scale, process complexity, and uncertainty on data assimilation, and the optimal treatment of model and observation errors are encouraged.

Houser, Paul R.↗

A multivariate variational objective analysis-assimilation method. Part 2: Case study results with and without satellite data

The variational multivariate assimilation method described in a companion paper by Achtemeier and Ochs is applied to conventional and conventional plus satellite data. Ground-based and space-based meteorological data are weighted according to the respective measurement errors and blended into a data set that is a solution of numerical forms of the two nonlinear horizontal momentum equations, the hydrostatic equation, and an integrated continuity equation for a dry atmosphere. The analyses serve first, to evaluate the accuracy of the model, and second to contrast the analyses with and without satellite data. Evaluation criteria measure the extent to which: (1) the assimilated fields satisfy the dynamical constraints, (2) the assimilated fields depart from the observations, and (3) the assimilated fields are judged to be realistic through pattern analysis. The last criterion requires that the signs, magnitudes, and patterns of the hypersensitive vertical velocity and local tendencies of the horizontal velocity components be physically consistent with respect to the larger scale weather systems.

Achtemeier, Gary L.↗

TICC Clustering Library v.1.0

SAND2024-01234O TICC is a clustering algorithm that labels a sequence of data points according to numerical properties. This library is a Python implementation of the algorithm described in "Toeplitz Inverse Covariance-Based Clustering of Multivariate Time Series Data" (Hallac et al. 2017). It includes documentation, performance improvements, examples, and test coverage. This library allows users to automatically segment a series of multivariate data points according to their covariance—that is, the way the values at each data point are changing in relation to one another. This is useful for identifying periods in which a system is behaving. For example, if a sensor is measuring a car's velocity, steering wheel angle, braking and acceleration, TICC can determine when the car was stopped, beginning/exiting a turn, slowing or accelerating at an intersection, or driving on straight or curved roads. TICC can be applied to measure multiple quantities at known times. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525

Dalbey, Keith↗

STSR-INR: Spatiotemporal super-resolution for multivariate time-varying volumetric data via implicit neural representation

Implicit neural representation (INR) has surfaced as a promising direction for solving different scientific visualization tasks due to its continuous representation and flexible input and output settings. We present STSR-INR, an INR solution for generating simultaneous spatiotemporal super-resolution for multivariate time-varying volumetric data. Inheriting the benefits of the INR-based approach, STSR-INR supports unsupervised learning and permits data upscaling with arbitrary spatial and temporal scale factors. Unlike existing GAN- or INR-based super-resolution methods, STSR-INR focuses on tackling variables or ensembles and enabling joint training across datasets of various spatiotemporal resolutions. Here we achieve this capability via a variable embedding scheme that learns latent vectors for different variables. In conjunction with a modulated structure in the network design, we employ a variational auto-decoder to optimize the learnable latent vectors to enable latent-space interpolation. To combat the slow training of INR, we leverage a multi-head strategy to improve training and inference speed with significant speedup. We demonstrate the effectiveness of STSR-INR with multiple scalar field datasets and compare it with conventional tricubic+linear interpolation and state-of-the-art deep-learning-based solutions (STNet and CoordNet).

97 MATHEMATICS AND COMPUTING↗

Mixture Models for Dependent Observations

Parametric mixture models appropriate for data presented in homogeneous blocks of varying sizes from several unidentified source populations are considered. For most applications, the data elements within each block are dependent. Models are proposed for multivariate normal data incorporating two types of dependence, exchangeability of elements within blocks, and a Markov structure for blocks. The consequences of assuming exchangeability, when in fact the Markov structure holds, are explored. Computational problems for each model are considered, and results of a simple test of the exchangeability hypothesis for LANDSAT data are presented.

Peters, C.↗

Quantitative image processing in fluid mechanics

The current status of digital image processing in fluid flow research is reviewed. In particular, attention is given to a comprehensive approach to the extraction of quantitative data from multivariate databases and examples of recent developments. The discussion covers numerical simulations and experiments, data processing, generation and dissemination of knowledge, traditional image processing, hybrid processing, fluid flow vector field topology, and isosurface analysis using Marching Cubes.

Hesselink, Lambertus↗

Data Mining of Polymer Phase Transitions upon Temperature Changes by Small and Wide-Angle X-ray Scattering Combined with Raman Spectroscopy

The complex physical transformations of polymers upon external thermodynamic changes are related to the molecular length of the polymer and its associated multifaceted energetic balance. The understanding of subtle transitions or multistep phase transformation requires real-time phenomenological studies using a multi-technique approach that covers several length-scales and chemical states. A combination of X-ray scattering techniques with Raman spectroscopy and Differential Scanning Calorimetry was conducted to correlate the structural changes from the conformational chain to the polymer crystal and mesoscale organization. Current research applications and the experimental combination of Raman spectroscopy with simultaneous SAXS/WAXS measurements coupled to a DSC is discussed. In particular, we show that in order to obtain the maximum benefit from simultaneously obtained high-quality data sets from different techniques, one should look beyond traditional analysis techniques and instead apply multivariate analysis. Data mining strategies can be applied to develop methods to control polymer processing in an industrial context. Crystallization studies of a PVDF blend with a fluoroelastomer, known to feature complex phase transitions, were used to validate the combined approach and further analyzed by MVA.

36 MATERIALS SCIENCE↗

Calibration or inverse regression: Which is appropriate for crop surveys using LANDSAT data?

Calibration and inverse regression estimators of crop proportions are investigated where the auxiliary variable is obtained from binary classification of multivariate LANDSAT data. The appropriate model relating classifier proportions and ground observed proportions for a given crop type is the calibration model. Under this model the inverse regression estimator is superior to the calibration estimator in estimating the crop acreage or proportion for a region of interest.

Chhikara, R. S.↗

Deep anomaly detection for industrial systems: a case study

We explore the use of deep neural networks for anomaly detection of industrial systems where the data are multivariate time series measurements. We formulate the problem as a self-supervised learning where data under normal operation are used to train a deep neural network autoregressive model, i.e., use a window of time series data to predict future data values. The aim of such a model is to learn to represent the system dynamic behavior under normal conditions, while expect higher model vs. measurement discrepancies under faulty conditions. In real world applications, many control settings are discrete in nature. In this paper, vector embedding and joint losses are employed to deal with such situations. Both LSTM and CNN based deep neural network backbones are studied on the Secure Water Treatment (SWaT) testbed datasets. Also, Support Vector Data Description (SVDD) method is adapted to such anomaly detection settings with deep neural networks. Evaluation methods and results are discussed based on the SWaT dataset along with potential pitfalls.

anomaly detection, deep neural network↗