Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Editorial: Applications of spectroscopy and chemometrics in nuclear materials analysis

Optical analysis techniques, including spectroscopy and image analysis, have many advantages when applied to the study of nuclear materials. They require small sample sizes, can be performed remotely, and can be proceduralized through consistent practice. Most importantly, they provide a wealth of information by generating multivariate data. For example, ultraviolet–visible–near-infrared absorbance spectroscopy of actinides in aqueous and organic solutions is dependent on the oxidation state, anionic complexation, and temperature. These variables are important for solution-based separation processes, and sensitivity to these factors, combined with online monitoring, can drive the efficiency and control of these processes. The morphology and chemical composition of actinide particles can also provide a vital clue to the mechanisms by which the particles were formed, providing forensic information on the origins of the particles.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Investigation of the Performance and Explainability Tradeoffs for Machine-Learning Models for Predictive Maintenance of Circulating Water Systems in Nuclear Power Plants

Predictive maintenance (PdM) has shown great potential for achieving substantial cost savings and enhancing the economic competitiveness of nuclear power plants (NPPs) in today's energy market. Among the different modeling approaches that exist, machine learning (ML) tools in particular have a demonstrated ability to handle high dimensional and multivariate data and to extract hidden relationships within data in industrial environments. While ML methods show great potential, their lack of explainability---especially for black-box models---is a major hurdle to their adoption. Moreover, considering the supposed trade-off between explainability and performance challenges, careful consideration must be made as to which of these quality aspects takes precedence in light of multiple modeling options, resource availability, and domain characteristics. The present work evaluates the performance of six ML models, each with a different degree of explainability, in classifying the conditions of circulating water pumps (CWPs) by utilizing sensor data from nuclear power plants. To determine the drivers behind the trade-offs presented by this array of models, this work also tests different combinations of CWP units as the training and testing data, degrees of data imbalance, and objective functions for hyperparameter tuning. It was found that black-box models tend to afford superior performance in cases where there are far more instances of one type of labeled data than of any other type. It is recommended that a guided procedure be followed for designing and delivering an ML system that is sufficiently explainable to all involved stakeholders.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Effect of vaccination on the case fatality rate for COVID-19 infections 2020–2021: multivariate modelling of data from the US Department of Veterans Affairs

Objectives: To evaluate the benefits of vaccination on the case fatality rate (CFR) for COVID-19 infections. Design, setting and participants: The US Department of Veterans Affairs has 130 medical centres. We created multivariate models from these data—339 772 patients with COVID-19—as of 30 September 2021. Outcome measures: The primary outcome for all models was death within 60 days of the diagnosis. Logistic regression was used to derive adjusted ORs for vaccination and infection with Delta versus earlier variants. Models were adjusted for confounding factors, including demographics, comorbidity indices and novel parameters representing prior diagnoses, vital signs/baseline laboratory tests and outpatient treatments. Patients with a Delta infection were divided into eight cohorts based on the time from vaccination to diagnosis. A common model was used to estimate the odds of death associated with vaccination for each cohort relative to that of unvaccinated patients. Results: 9.1% of subjects were vaccinated. 21.5% had the Delta variant. 18 120 patients (5.33%) died within 60 days of their diagnoses. The adjusted OR for a Delta infection was 1.87±0.05, which corresponds to a relative risk (RR) of 1.78. The overall adjusted OR for prior vaccination was 0.280±0.011 corresponding to an RR of 0.291. Raw CFR rose steadily after 10–14 weeks. The OR for vaccination remained stable for 10–34 weeks. Conclusions: Our CFR model controls for the severity of confounding factors and priority of vaccination, rather than solely using the presence of comorbidities. Our results confirm that Delta was more lethal than earlier variants and that vaccination is an effective means of preventing death. After adjusting for major selection biases, we found no evidence that the benefits of vaccination on CFR declined over 34 weeks. We suggest that this model can be used to evaluate vaccines designed for emerging variants.

59 BASIC BIOLOGICAL SCIENCES↗

Large Deviations Anomaly Detection (LAD) for collection of multivariate time series data: Applications to COVID-19 data

Time series anomaly detection is frequently used to identify extreme behaviors within a single time series. Identifying extreme trends in relation to a collection of other time series, on the other hand, is frequently of significant interest, such as in public health policy, social justice, and pandemic propagation. Using concepts from large deviations theory , we propose an algorithm that can scale to large collections of time series data. This paper expands on the LAD algorithm presented in Guggilam et al. (2022). The proposed algorithm is an online anomaly detection method for identifying anomalies in a collection of multivariate time series that takes advantage of the algorithm’s ability to scale to high-dimensional data. We show how the proposed Large Deviations Anomaly Detection (LAD) algorithm can be used to identify regions with anomalous trends in COVID-19 cases, deaths, biweekly growth rates, vaccinations, and fatality rates. Several of the observed anomalous trends are associated with regions that have demonstrated poor response to the COVID pandemic.

97 MATHEMATICS AND COMPUTING↗

GMT: A deep learning approach to generalized multivariate translation for scientific data analysis and visualization

In scientific visualization, despite the significant advances of deep learning for data generation, researchers have not thoroughly investigated the issue of data translation. We present a new deep learning approach called generalized multivariate translation (GMT) for multivariate time-varying data analysis and visualization. Like V2V, GMT assumes a preprocessing step that selects suitable variables for translation. However, unlike V2V, which only handles one-to-one variable translation during training and inference, GMT enables one-to-many and many-to-many variable translation in the same framework. We leverage the recent StarGAN design from multi-domain image-to-image translation to achieve this generalization capability. We experiment with different loss functions and injection strategies to explore the best choices and leverage pre-training for performance improvement. We compare GMT with other state-of-the-art methods (i.e., Pix2Pix, V2V, StarGAN). Furthermore, the results demonstrate the overall advantage of GMT in translation quality and generalization ability.

97 MATHEMATICS AND COMPUTING↗

Data and code from: Multivariate bayesian regression model for predicting disposed ash composition at U.S. coal fired power stations

This dataset contains the code and data files needed for implementation of a Multivariate Bayesian Regression model, described in Jin et al. (2025), for the historical prediction of the chemical composition of disposed coal ash at U.S. coal fired power plants as a function of annualized coal purchase data. The integrated coal supply data file (CoalSupplyDataset.csv) represents a compilation of monthly fuel purchase records for the period 1973-2022 at major U.S. power stations. These records were obtained from the U.S. Energy Information Administration. The CSV file also contains, for each coal purchase record, the coal region of the mine as defined by the U.S. Geological Survey. Data entry errors and data gaps in the EIA records were corrected as described in Jin et al. This CSV file represents the integrated coal supply data after corrections were made. The model structure and fitting parameters are encoded in pickle file format (Bayesian.pkl). The model was developed with the coal supply data and coal ash composition data, apportioned according to the Stratified Shuffle Split for training and testing subsets. The model was built using Python and the PyMC library. Reference Publication: Jin, Z.; Huang, J.; Hower, J.C.; Hsu-Kim, H.(2025). Predictive Assessment of the Chemical Composition of Coal Ash in Reserve at U.S. Disposal Sites. Environmental Science & Technology.

Coal ash composition↗

High-dimensional multivariate autoregressive model estimation of human electrophysiological data using fMRI priors

Multivariate autoregressive (MVAR) model estimation enables assessment of causal interactions in brain networks. However, accurately estimating MVAR models for high-dimensional electrophysiological recordings is challenging due to the extensive data requirements. Hence, the applicability of MVAR models for study of brain behavior over hundreds of recording sites has been very limited. Prior work has focused on different strategies for selecting a subset of important MVAR coefficients in the model to reduce the data requirements of conventional least-squares estimation algorithms. Here we propose incorporating prior information, such as resting state functional connectivity derived from functional magnetic resonance imaging, into MVAR model estimation using a weighted group least absolute shrinkage and selection operator (LASSO) regularization strategy. The proposed approach is shown to reduce data requirements by a factor of two relative to the recently proposed group LASSO method of Endemann et al (Neuroimage 254:119057, 2022) while resulting in models that are both more parsimonious and more accurate. The effectiveness of the method is demonstrated using simulation studies of physiologically realistic MVAR models derived from intracranial electroencephalography (iEEG) data. The robustness of the approach to deviations between the conditions under which the prior information and iEEG data is obtained is illustrated using models from data collected in different sleep stages. This approach allows accurate effective connectivity analyses over short time scales, facilitating investigations of causal interactions in the brain underlying perception and cognition during rapid transitions in behavioral state.

62 RADIOLOGY AND NUCLEAR MEDICINE↗

Large-Scale Inference of Multivariate Regression for Heavy-Tailed and Asymmetric Data

Large-scale multivariate regression is a fundamental statistical tool with a wide range of applications. Here, this study considers the problem of simultaneously testing a large number of general linear hypotheses, encompassing covariate-effect analysis, analysis of variance, and model comparisons. The challenge that accompanies a large number of tests is the ubiquitous presence of heavy-tailed and/or highly skewed measurement noise, which is the main reason for the failure of conventional least squares-based methods. For large-scale multivariate regression, we develop a set of robust inference methods to explore data features such as heavy tailedness and skewness, which are not visible to least squares methods. The new testing procedure is based on the data-adaptive Huber regression and a new covariance estimator of regression estimates. Under mild conditions, we show that our methods produce consistent estimates of the false discovery proportion. Extensive numerical experiments and an empirical study on quantitative linguistics demonstrate the advantage of the proposed method over many state-of-the-art methods when the data are generated from heavy-tailed and/or skewed distributions.

97 MATHEMATICS AND COMPUTING↗

TICC Clustering Library v.1.0

SAND2024-01234O TICC is a clustering algorithm that labels a sequence of data points according to numerical properties. This library is a Python implementation of the algorithm described in "Toeplitz Inverse Covariance-Based Clustering of Multivariate Time Series Data" (Hallac et al. 2017). It includes documentation, performance improvements, examples, and test coverage. This library allows users to automatically segment a series of multivariate data points according to their covariance—that is, the way the values at each data point are changing in relation to one another. This is useful for identifying periods in which a system is behaving. For example, if a sensor is measuring a car's velocity, steering wheel angle, braking and acceleration, TICC can determine when the car was stopped, beginning/exiting a turn, slowing or accelerating at an intersection, or driving on straight or curved roads. TICC can be applied to measure multiple quantities at known times. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525

Dalbey, Keith↗

STSR-INR: Spatiotemporal super-resolution for multivariate time-varying volumetric data via implicit neural representation

Implicit neural representation (INR) has surfaced as a promising direction for solving different scientific visualization tasks due to its continuous representation and flexible input and output settings. We present STSR-INR, an INR solution for generating simultaneous spatiotemporal super-resolution for multivariate time-varying volumetric data. Inheriting the benefits of the INR-based approach, STSR-INR supports unsupervised learning and permits data upscaling with arbitrary spatial and temporal scale factors. Unlike existing GAN- or INR-based super-resolution methods, STSR-INR focuses on tackling variables or ensembles and enabling joint training across datasets of various spatiotemporal resolutions. Here we achieve this capability via a variable embedding scheme that learns latent vectors for different variables. In conjunction with a modulated structure in the network design, we employ a variational auto-decoder to optimize the learnable latent vectors to enable latent-space interpolation. To combat the slow training of INR, we leverage a multi-head strategy to improve training and inference speed with significant speedup. We demonstrate the effectiveness of STSR-INR with multiple scalar field datasets and compare it with conventional tricubic+linear interpolation and state-of-the-art deep-learning-based solutions (STNet and CoordNet).

97 MATHEMATICS AND COMPUTING↗

Advanced Interactive 3D Visualization Tool for Customizable Analyses of Tomography Datasets in Material Science

Current methods for visualizing and analyzing 3D tomography datasets in materials science often lack the interactivity and depth required for detailed structural insights. This limitation restricts a researchers' ability to accurately interpret complex data, which is critical for advancing material innovations and understanding structural properties. To address this issue, we have developed a novel, web-based interactive 3D visualization and analysis tool from the Trame framework that offers customizable features to enhance data interpretability. The tool allows users to adjust parameters such as visible range, slice planes, data rotation, and layering, providing a more detailed and dynamic view of complex structures. Its user-friendly web interface increases the accessibility and ease of use for both novice and experienced researchers, to visualize large volumetric datasets. The tool supports a diverse range of data formats, making it versatile for various research applications. Unique capabilities include real-time data manipulation, automated feature detection, context-sensitive feedback, and real-time volume calculations and distributions per sliced region or layer, alongside the ability to quickly generate high-quality screenshots and videos for presentations and reports. These advancements offer a comprehensive solution for enhanced 3D data exploration, significantly improving the analysis process and communication of results in materials science.

36 - MATERIALS SCIENCE↗

CHMMPY: A python package for constrained Hidden Markov Models

SAND2025-11909O chmmpy software analyzes multivariate timeseries data to detect patterns. It uses a Hidden Markov Model (HMM) and application-specific constraints that reflect known relationships among hidden states to accomplish this. The chmmpy software provides a generic framework for expressing application-specific constraints and supporting constrained HMM inference using optimization solvers. chmmpy is available on GitHub. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Hart, William↗

Multidimensional scaling informed by F -statistic: Visualizing grouped microbiome data with inference

Multidimensional scaling (MDS) is a widely used dimensionality reduction technique in microbial ecology data analysis that captures the multivariate structure of the data while preserving pairwise distances between samples. While improvements in MDS have enhanced the ability to reveal group-specific data patterns, these MDS-based methods require prior assumptions for inference, limiting their application in general microbiome analysis. Here, in this study, we introduce a new MDS-based ordination method, “F-informed MDS,” which configures the data distribution based on the F-statistic, the ratio of dispersion between groups sharing common and different characteristics. Using semisynthetic datasets, we demonstrate that the proposed method is robust to hyperparameter selection while maintaining statistical significance throughout the ordination process. Various quality metrics for evaluating dimensionality reduction confirm that F-informed MDS is comparable to state-of-the-art methods in preserving both local and global data structures. Its application to a diatom-associated bacterial community suggests the role of this new method in interpreting the community’s response to the host. Our approach offers a well-founded refinement of MDS that aligns with statistical test results, which can be beneficial for broader multidimensional data analyses in microbiology and ecology. This new visualization tool can be incorporated into standard microbiome data analyses.

Biological and medical sciences↗

Foundations of machine learning for low-temperature plasmas: methods and case studies

Abstract Machine learning (ML) and artificial intelligence have proven to be an invaluable tool in tackling a vast array of scientific, engineering, and societal problems. The main drivers behind the recent proliferation of ML in practically all aspects of science and technology can be attributed to: (a) improved data acquisition and inexpensive data storage; (b) exponential growth in computing power; and (c) availability of open-source software and resources that have made the use of state-of-the-art ML algorithms widely accessible. The impact of ML on the field of low-temperature plasmas (LTPs) could be particularly significant in the emerging applications that involve plasma treatment of complex interfaces in areas ranging from the manufacture of microelectronics and processing of quantum materials, to the LTP-driven electrification of the chemical industry, and to medicine and biotechnology. This is primarily due to the complex and poorly-understood nature of the plasma-surface interactions in these applications that pose unique challenges to the modeling, diagnostics, and predictive control of LTPs. As the use of ML is becoming more prevalent, it is increasingly paramount for the LTP community to be able to critically analyze and assess the concepts and techniques behind data-driven approaches. To this end, the goal of this paper is to provide a tutorial overview of some of the widely-used ML methods that can be useful, amongst others, for discovering and correlating patterns in the data that may be otherwise impractical to decipher by human intuition alone, for learning multivariable nonlinear data-driven prediction models that are capable of describing the complex behavior of plasma interacting with interfaces, and for guiding the design of experiments to explore the parameter space of plasma-assisted processes in a systematic and resource-efficient manner. We illustrate the utility of various supervised, unsupervised and active learning methods using LTP datasets consisting of commonly-available, information-rich measurements (e.g. optical emission spectra, current–voltage characteristics, scanning electron microscope images, infrared surface temperature measurements, Fourier transform infrared spectra). All the ML demonstrations presented in this paper are carried out using open-source software; the datasets and codes are made publicly available. The FAIR guiding principles for scientific data management and stewardship can accelerate the adoption and development of ML in the LTP community.

Physics↗

Long-term missing value imputation for time series data using deep neural networks

We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron, for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g., multiple months of missing daily observations) rather than on individual randomly missing observations. Our proposed gap filling algorithm uses an automated method for determining the optimal MLP model architecture, thus allowing for optimal prediction performance for the given time series. We tested our approach by filling gaps of various lengths (three months to three years) in three environmental datasets with different time series characteristics, namely daily groundwater levels, daily soil moisture, and hourly Net Ecosystem Exchange. We compared the accuracy of the gap-filled values obtained with our approach to the widely used R-based time series gap filling methods ImputeTS and mtsdi. The results indicate that using an MLP for filling a large gap leads to better results, especially when the data behave nonlinearly. Thus, our approach enables the use of datasets that have a large gap in one variable, which is common in many long-term environmental monitoring observations.

97 MATHEMATICS AND COMPUTING↗

Search for ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production in the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state using proton-proton collisions at $\sqrt{s}=13\,\text {Te}\hspace{-.08em}\text {V} $

A search for ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production in the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state is presented, where H is the standard model (SM) Higgs boson. The search uses an event sample of proton-proton collisions corresponding to an integrated luminosity of 133$\,\text {fb}^{-1}$ collected at a center-of-mass energy of 13$\,\text {Te}\hspace{-.08em}\text {V}$ with the CMS detector at the CERN LHC. The analysis introduces several novel techniques for deriving and validating a multi-dimensional background model based on control samples in data. A multiclass multivariate classifier customized for the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state is developed to derive the background model and extract the signal. The data are found to be consistent, within uncertainties, with the SM predictions. The observed (expected) upper limits at 95% confidence level are found to be 3.8 (3.8) and 5.0 (2.9) times the SM prediction for the ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production cross sections, respectively.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative data sets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student t process regression. We apply Student t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via colouring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics and classic machine learning clustering algorithms.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗