Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Parametric Testing of Launch Vehicle FDDR Models

For the safe operation of a complex system like a (manned) launch vehicle, real-time information about the state of the system and potential faults is extremely important. The on-board FDDR (Failure Detection, Diagnostics, and Response) system is a software system to detect and identify failures, provide real-time diagnostics, and to initiate fault recovery and mitigation. The ERIS (Evaluation of Rocket Integrated Subsystems) failure simulation is a unified Matlab/Simulink model of the Ares I Launch Vehicle with modular, hierarchical subsystems and components. With this model, the nominal flight performance characteristics can be studied. Additionally, failures can be injected to see their effects on vehicle state and on vehicle behavior. A comprehensive test and analysis of such a complicated model is virtually impossible. In this paper, we will describe, how parametric testing (PT) can be used to support testing and analysis of the ERIS failure simulation. PT uses a combination of Monte Carlo techniques with n-factor combinatorial exploration to generate a small, yet comprehensive set of parameters for the test runs. For the analysis of the high-dimensional simulation data, we are using multivariate clustering to automatically find structure in this high-dimensional data space. Our tools can generate detailed HTML reports that facilitate the analysis.

Schumann, Johann↗

Survey on stochastic distribution systems: A full probability density function control theory with potential applications

Complex systems seen either in general engineering practice or economics are subjected to ever increased uncertainties that are mostly represented as random variables or parameters, and the characteristics of random variables are represented by their probability density functions (PDFs). Controlling their PDFs means to shape their stochastic distributions and in general it would provide a full treatment for system analysis and operational control and optimization. This leads to the development of stochastic distribution control (SDC) systems theory in the past decades, where the original aim of the controller design is to realize a shape control of the distributions of certain random variables in their PDFs sense for some engineering processes. Indeed, once the PDFs of these random variables or parameters are used to describe their distribution characters, the control task is to obtain control signals so that the output PDFs of stochastic systems are made to follow their target PDFs. The subject of SDC was initially originated for non-Gaussian stochastic control systems design but has found a wide spectrum of applications in general systems in terms of data-driven modeling, analysis, signal processing (filtering), data mining via multivariable statistics, decision-making (optimization) for systems subjected to uncertainties and even in economics. In this context, SDC constitutes an effective primer tool for complex system analysis, control and operational optimizations. In this review paper, a detailed survey of the developments on the research of SDC systems will be made together with their wide spectrum applications and future perspectives.

42 ENGINEERING↗

CHMMPP: A c++ library for constrained Hidden Markov Models

SAND2024-13027O The CHMMPP: A c++ Library for Constrained Hidden Markov Models (HMM) software supports the analysis of multivariate time series data to detect patterns using HMM. Many applications involve the detection and characterization of hidden or latent states in a complex system using observable states and variables. This software supports inference of latent states integrating both an HMM and application-specific constraints that reflect known relationships in hidden states. The CHMMPP software supports application-specific and generic methods for constrained inference. This includes a framework for customized Viterbi methods, constrained inference of hidden states with A* and integer programming methods, and various constraint-informed methods for learning HMM model parameters. CHMMPP focuses on supporting generic methods that enable the agile expression of complex sets of constraints that naturally arise in many real-world applications.

Hart, William↗

Natural Resources Inventory and Land Evaluation in Switzerland

The author has identified the following significant results. A system was developed to operationally map and measure the areal extent of various land use categories for updating existing and producing new and actual thematic maps showing the latest state of rural and urban landscapes and its changes. The processing system includes: (1) preprocessing steps for radiometric and geometric corrections; (2) classification of the data by a multivariate procedure, using a stepwise linear discriminant analysis based on carefully selected training cells; and (3) output in form of color maps by printing black and white theme overlays of a selected scale with photomation system and its coloring and combination into a color composite.

Haefner, H.↗

Fitting Nonlinear Curves by use of Optimization Techniques

MULTIVAR is a FORTRAN 77 computer program that fits one of the members of a set of six multivariable mathematical models (five of which are nonlinear) to a multivariable set of data. The inputs to MULTIVAR include the data for the independent and dependent variables plus the user s choice of one of the models, one of the three optimization engines, and convergence criteria. By use of the chosen optimization engine, MULTIVAR finds values for the parameters of the chosen model so as to minimize the sum of squares of the residuals. One of the optimization engines implements a routine, developed in 1982, that utilizes the Broydon-Fletcher-Goldfarb-Shanno (BFGS) variable-metric method for unconstrained minimization in conjunction with a one-dimensional search technique that finds the minimum of an unconstrained function by polynomial interpolation and extrapolation without first finding bounds on the solution. The second optimization engine is a faster and more robust commercially available code, denoted Design Optimization Tool, that also uses the BFGS method. The third optimization engine is a robust and relatively fast routine that implements the Levenberg-Marquardt algorithm.

Hill, Scott A.↗

El Nino-Southern Oscillation Correlated Aerosol Angstrom Exponent Anomaly Over the Tropical Pacific Discovered in Satellite Measurements

El Nino.Southern Oscillation (ENSO) is the dominant mode of interannual variability in the tropical atmosphere. ENSO could potentially impact local and global aerosol properties through atmospheric circulation anomalies and teleconnections. By analyzing aerosol properties, including aerosol optical depth (AOD) and Angstrom exponent (AE; often used as a qualitative indicator of aerosol particle size) from the Moderate Resolution Imaging Spectrometer, the Multiangle Imaging Spectroradiometer and the Sea ]viewing Wide Field ]of ]view Sensor for the period 2000.2011, we find a strong correlation between the AE data and the multivariate ENSO index (MEI) over the tropical Pacific. Over the western tropical Pacific (WTP), AE increases during El Nino events and decreases during La Nina events, while the opposite is true over the eastern tropical Pacific (ETP). The difference between AE anomalies in the WTP and ETP has a higher correlation coefficient (>0.7) with the MEI than the individual time series and could be considered another type of ENSO index. As no significant ENSO correlation is found in AOD over the same region, the change in AE (and hence aerosol size) is likely to be associated with aerosol composition changes due to anomalous meteorological conditions induced by the ENSO. Several physical parameters or mechanisms that might be responsible for the correlation are discussed. Preliminary analysis indicates surface wind anomaly might be the major contributor, as it reduces sea ]salt production and aerosol transport during El Nino events. Precipitation and cloud fraction are also found to be correlated with tropical Pacific AE. Possible mechanisms, including wet removal and cloud shielding effects, are considered. Variations in relative humidity, tropospheric ozone concentration, and ocean color during El Nino have been ruled out. Further investigation is needed to fully understand this AE ]ENSO covariability and the underlying physical processes responsible for it.

Li, Jing↗

Practical Guide to Chemometric Analysis of Optical Spectroscopic Data

The methodology and mathematical treatment of several classic multivariate methods for the analysis of spectroscopic data is demonstrated in a straightforward way that can be used as a basis for teaching an undergraduate introductory course on chemometric analysis. The multivariate techniques of classical least squares (CLS), principal component regression (PCR), and partial least squares (PLS), as well as the univariate Beer’s law method have been described and compared, building students’ understanding by starting with the univariate method and progressing step by step into the multivariate methods. Equations for the production of regression vectors from training set spectral data is described and their use demonstrated for the prediction of constituent concentrations on a separate validation set of spectra. Extreme care is taken to ensure consistency in variable formatting of data matrices. This provides a key foundation to understanding how spectral data are manipulated using these different mathematical approaches for building quantitative regression models. Each method is applied to a real-world data set, and the results are discussed to show students the types of information that can be gleaned from each method. A training set comprised of 20 infrared absorbance spectra containing 3 constituents (benzene, polystyrene, and gasoline) of known composition are used to demonstrate the matrix operations for each regression method. A separate set of 12 real-world napalm samples (containing benzene, polystyrene and gasoline) are used as a validation set to demonstrate the ability to utilize the regression models on an unknown dataset. A toolbox (PNNL Chemometric Toolbox) written in MATLAB language is supplied in the Supplemental Information file and can be used as a companion for understanding the development and deployment of the chemometric algorithms described in this paper. The datasets of the infrared spectra are also supplied, allowing users to build and inspect the chemometric models on their own. Finally, the Toolbox includes scripts to assist users in loading their own datasets into MATLAB and performing CLS, PCR, and PLS on their data.

Upper-Division Undergraduate, Analytical Chemistry↗

Emissions mitigation technology for advanced water-lean solvent-based CO 2 capture processes

This technical final report submitted to DOE/NETL presents all the research activities performed during the entirety of DE-FE0031660 project-Emissions Mitigation Technology for Advanced Water-Lean Solvent-Based CO 2 Capture Processes which spans from October 2018 through March 2022. RTI International has been conducting studies from fundamental and operational aspects to reduce the overall amine emissions from the advanced Water-Lean Solvent (WLS) systems, specifically RTI’s Non-Aqueous Solvent (NAS). This technical final report will highlight the key findings from project which align closely to the project objectives which are: Identify the contribution of vapor loss, entrainment, and aerosols to the overall emissions of water-lean systems; Determine the significance of CO 2 capture system operating parameters to the amine emissions; Develop an emissions model based on critical operating parameters; Evaluate the effectiveness of emissions mitigation devices to reduce the amine emissions to <1 ppm under flue coal-fired flue gas; and, Determine the contribution of the ECTs to the overall CO 2 capture cost. The following are the key findings based on numerous tests using both lab-scale setups and parametric testing performed at RTI’s Bench-scale Gas Absorption System (BsGAS). During the BP1, the aerosol generation system and monitoring equipment were installed at BsGAS to produce and determine the aerosol characteristics during the NAS CO 2 capture process. The aerosol produced by this setup produced aerosols with the peak diameter and concentration of 50 micron and 1.2E10 7 cm -3 , respectively. These particle sizes and concentrations are matched to those observed in the actual coal-fired power plant flue gases and expected to be found at the absorber inlet of the CO 2 capture system. Over 1,300 hours of parametric testing have been conducted to evaluate the impact of the aerosols and operating conditions during the CO 2 capture with NAS on the overall amine emissions in the treated flue gas. At the worse condition tested, the presence of the aerosols in the flue gas could increase the overall emissions by 10X compared to the baseline emissions from NAS’s vapor pressure. CO 2 capture rate was found to be a main factor impacting the overall emissions as well as aerosol size and concentrations in the absorber off-gas. The higher CO 2 capture rate, the higher amine emissions in the treated gas. The temperature difference between the temperature bulge seen in the absorber and the water wash temperature also impacts the particle growth where the larger the temperature difference, the more amine emissions from aerosols in the treated gas. The majority of the aerosols did not grow substantially in the system, and the particle concentrations remained nearly constant between the absorber inlet and wash outlet. Only a small portion of the particles were found to grow significantly. The high efficiency demister with mesh size of 5-10 micron can be installed to remove a portion of the aerosols from the gas stream leaving the water wash. Overall, these results from parametric testing have established the emission baseline and validate our assumption on the need of emission control technologies (ECT) in order to minimize the emissions from the baseline NAS CO 2 capture process. Over 2,000 of BsGAS operating hours was used to investigate a handful of process improvements which led to a selection of the vital few changes that effectively control the amine emissions. These process improvements are lime-coated-filters for absorber gas inlet, advanced demister at the top of the absorber, a second water wash with amine recovery unit were designed, installed, and tested at BsGAS at the end of BP1. The result showed that the NAS CO 2 capture process with these additional emission control devices could lower the amine emission in the treated gas to about 1 ppm using a simulated coal-fire flue gas stream. The main contributor in lowering the amine emission came from the second water wash with amine recovery unit where the amine concentration in the scrubbing water was kept below 2 wt% through a continuous amine removal via an adsorbent bed, resulting in a low amine vapor pressure. The adsorbent bed was regenerated via a direct steam regeneration and the recover amine was returned to the absorber to minimize wastewater and makeup amine. A flue gas generation system was designed and installed during the first half of BP2 to support the emission testing using a real coal-derived flue gas. The system is capable of generating both coal- and natural gas- derived flue gases with the composition of the gaseous species highly resemble to that of the power plant flue gases. The particulates detected in the coal-derived flue gas showed the mean diameter of 1 micron. The CO 2 capture operating was then proceed using the real coal-derived flue gas where the amine emission was controlled to be about 0-3 ppm for the total run time of about 200 hours. Similar testing was conducted with natural gas-derived flue gas and the result showed a highly amine emission of 30 ppm under the total run time of 200 hours. The Principal Component Analysis (PCA) and the Partial Least Squares Projection to Latent Structures (PLS) techniques were applied to the parametric testing data to derive a multivariate statistical model. The model was validated and trained with half of the data collected, and the predictive ability of the model was evaluated using the remaining half of the data. The resulting empirical model was capable of predicting the overall emissions from the NAS process without the ECTs with ±15% accuracy (average absolute deviation, AAD) in BP1. As more emission data were obtained under the real coal-flue gas in the BP2, the model incorporated these new set of data to reflect the final process configuration, operating parameters, and amine emission. This results in the updated empirical model predicting the amine emission from the NAS CO 2 capture process with 84% goodness-of-fit (R 2 ), 85% predictability (Q 2 ), and 15% AAD. The study evaluates the use of RTI’s Non-Aqueous Solvent technology for 90% CO 2 capture from a net 650 MWe pulverized coal power plant, downstream of the flue-gas desulfurization unit. The captured CO 2 has a purity of > 95% CO 2 , and is dried, compressed to 15.3 MPa (2,215 psia), ready for sequestration. The analysis uses Case B12B from the DOE Baseline study on Bituminous Coal, Revision 4 where the Cansolv CO 2 capture plant is replaced by the RTI CO 2 Capture plant. The CO 2 capture plant has been sized to capture >90% CO 2 from flue gas derived from a net 650 MWe supercritical pulverized coal power plant. The CO 2 capture plant is equipped with emission control technologies that limits the amine emissions to < 1 ppm. Two different cases were evaluated for the technoeconomic study. The key difference between the two cases is the regenerator pressure. In Case 1, the regenerator operates at 0.195 MPa (28.3 psia), whereas in Case 2, the regenerator pressure is 0.44 MPa (64 psia) thus removing the need for the first stage of compression of the eight-stage compression train. Results from the TEA are compared against the DOE reference cases for SCPC plant with and without CO 2 Capture (Case B12A and Case B12B of the DOE Baseline study, respectively). Case 2 with CO 2 regeneration at higher pressure results in the lower cost of CO 2 capture. The total capital cost of the capture process has been estimated using 2018 dollars in Aspen Process Economic Analyzer and was estimated to be $579 MM. The capture plant operation leads to a total parasitic power loss rate of 96 MWe, resulting in a decrease in pulverized coal power plant efficiency of 7.8% points. The resulting cost of electric power increases from 64.4 mills/kWh, for no capture, to 97.5 mills/kWh, with 90% capture, an increase of 51% in the COE. The cost of capturing 90% CO 2 was estimated to be $38.2/tonne-CO 2 , and meets the DOE target of $40/t-CO 2 . Emission control technologies (ECT) investigated in this project includes a second water wash with use of activated carbon beds for removal of amine from the wash water prior to recirculation in the water wash. These ECT allow operation of the CO 2 capture plant with < 1 ppm amine emissions with the treated flue gas and contributes to $2.4/t-CO 2 captured. Amine emissions derived from thermal and oxidative degradations were investigated under this project along with the emissions derived from aerosols for the NAS system. The thermally degraded of the lean NAS showed less than 4% decreased of the original total amine content in the NAS at 150 °C while the result obtained at 120 °C showed no drop in total amine content, suggesting that thermal degradation of the NAS is minimal. These results also suggested that the thermally degraded species are not likely formed and contributed to the emissions due to the low regeneration temperature of the NAS at 90-105 °C. The oxidative degradation, on the other hand, could become problematic as some of these oxidative degraded species were observed during the NAS-5 testing at National Carbon Capture Center (NCCC) and SINTEF in our previous project. The rapid screening of selected inhibitors suggested that oxidative degradation of NAS can be suppressed using thiol containing compounds in amounts of at least 1 mol%. The detailed mechanistic degradation pathway was conceived for a specific amine used in NAS formulation during BP2. he reduction of the nitrosamines caused by the NO x present in the flue gas was also examined. The study suggested that the thermo-chemical treatment of the NAS solvent would be a more effective and economically viable compared to removing NO x at the DCC.

01 COAL, LIGNITE, AND PEAT↗

RadVolViz: An Information Display-Inspired Transfer Function Editor for Multivariate Volume Visualization

In volume visualization transfer functions are widely used for mapping voxel properties to color and opacity. Typically, volume density data are scalars which require simple 1D transfer functions to achieve this mapping. If the volume densities are vectors of three channels, one can straightforwardly map each channel to either red, green or blue, which requires a trivial extension of the 1D transfer function editor. Here, we devise a new method that applies to volume data with more than three channels. These types of data often arise in scientific scanning applications, where the data are separated into spectral bands or chemical elements. Our method expands on prior work in which a multivariate information display, RadViz, was fused with a radial color map, in order to visualize multi-band 2D images. In this work, we extend this joint interface to blended volume rendering. The information display allows users to recognize the presence and value distribution of the multivariate voxels and the joint volume rendering display visualizes their spatial distribution. We design a set of operators and lenses that allow users to interactively control the mapping of the multivariate voxels to opacity and color. This enables users to isolate or emphasize volumetric structures with desired multivariate properties. Furthermore, it turns out that our method also enables more insightful displays even for RGB data. We demonstrate our method with three datasets obtained from spectral electron microscopy, high energy X-ray scanning, and atmospheric science.

36 MATERIALS SCIENCE↗

Multivariate statistical analysis software technologies for astrophysical research involving large data bases

The existing and forthcoming data bases from NASA missions contain an abundance of information whose complexity cannot be efficiently tapped with simple statistical techniques. Powerful multivariate statistical methods already exist which can be used to harness much of the richness of these data. Automatic classification techniques have been developed to solve the problem of identifying known types of objects in multi parameter data sets, in addition to leading to the discovery of new physical phenomena and classes of objects. We propose an exploratory study and integration of promising techniques in the development of a general and modular classification/analysis system for very large data bases, which would enhance and optimize data management and the use of human research resources.

Djorgovski, Stanislav↗

Multivariate statistical analysis software technologies for astrophysical research involving large data bases

The existing and forthcoming data bases from NASA missions contain an abundance of information whose complexity cannot be efficiently tapped with simple statistical techniques. Powerful multivariate statistical methods already exist which can be used to harness much of the richness of these data. Automatic classification techniques have been developed to solve the problem of identifying known types of objects in multiparameter data sets, in addition to leading to the discovery of new physical phenomena and classes of objects. We propose an exploratory study and integration of promising techniques in the development of a general and modular classification/analysis system for very large data bases, which would enhance and optimize data management and the use of human research resource.

Djorgovski, George↗

Experiments with a three-dimensional statistical objective analysis scheme using FGGE data

A three-dimensional (3D), multivariate, statistical objective analysis scheme (referred to as optimum interpolation or OI) has been developed for use in numerical weather prediction studies with the FGGE data. Some novel aspects of the present scheme include: (1) a multivariate surface analysis over the oceans, which employs an Ekman balance instead of the usual geostrophic relationship, to model the pressure-wind error cross correlations, and (2) the capability to use an error correlation function which is geographically dependent. A series of 4-day data assimilation experiments are conducted to examine the importance of some of the key features of the OI in terms of their effects on forecast skill, as well as to compare the forecast skill using the OI with that utilizing a successive correction method (SCM) of analysis developed earlier. For the three cases examined, the forecast skill is found to be rather insensitive to varying the error correlation function geographically. However, significant differences are noted between forecasts from a two-dimensional (2D) version of the OI and those from the 3D OI, with the 3D OI forecasts exhibiting better forecast skill. The 3D OI forecasts are also more accurate than those from the SCM initial conditions. The 3D OI with the multivariate oceanic surface analysis was found to produce forecasts which were slightly more accurate, on the average, than a univariate version.

Baker, Wayman E.↗

Multivariate statistical analysis software technologies for astrophysical research involving large data bases

We developed a package to process and analyze the data from the digital version of the Second Palomar Sky Survey. This system, called SKICAT, incorporates the latest in machine learning and expert systems software technology, in order to classify the detected objects objectively and uniformly, and facilitate handling of the enormous data sets from digital sky surveys and other sources. The system provides a powerful, integrated environment for the manipulation and scientific investigation of catalogs from virtually any source. It serves three principal functions: image catalog construction, catalog management, and catalog analysis. Through use of the GID3* Decision Tree artificial induction software, SKICAT automates the process of classifying objects within CCD and digitized plate images. To exploit these catalogs, the system also provides tools to merge them into a large, complete database which may be easily queried and modified when new data or better methods of calibrating or classifying become available. The most innovative feature of SKICAT is the facility it provides to experiment with and apply the latest in machine learning technology to the tasks of catalog construction and analysis. SKICAT provides a unique environment for implementing these tools for any number of future scientific purposes. Initial scientific verification and performance tests have been made using galaxy counts and measurements of galaxy clustering from small subsets of the survey data, and a search for very high redshift quasars. All of the tests were successful, and produced new and interesting scientific results. Attachments to this report give detailed accounts of the technical aspects for multivariate statistical analysis of small and moderate-size data sets, called STATPROG. The package was tested extensively on a number of real scientific applications, and has produced real, published results.

Djorgovski, S. George↗

Nominal and adversarial synthetic PMU data for standard IEEE test systems

GridSTAGE (Spatio-Temporal Adversarial scenario GEneration) is a framework for the simulation of adversarial scenarios and the generation of multivariate spatio-temporal data in cyber-physical systems. GridSTAGE is developed based on Matlab and leverages Power System Toolbox (PST) where the evolution of the power network is governed by nonlinear differential equations. Using GridSTAGE, one can create several event scenarios that correspond to several operating states of the power network by enabling or disabling any of the following: faults, AGC control, PSS control, exciter control, load changes, generation changes, and different types of cyber-attacks. Standard IEEE bus system data is used to define the power system environment. GridSTAGE emulates the data from PMU and SCADA sensors. The rate of frequency and location of the sensors can be adjusted as well. Detailed instructions on generating data scenarios with different system topologies, attack characteristics, load characteristics, sensor configuration, control parameters are available in the Github repository - https://github.com/pnnl/GridSTAGE. There is no existing adversarial data-generation framework that can incorporate several attack characteristics and yield adversarial PMU data. The GridSTAGE framework currently supports simulation of False Data Injection attacks (such as a ramp, step, random, trapezoidal, multiplicative, replay, freezing) and Denial of Service attacks (such as time-delay, packet-loss) on PMU data. Furthermore, it supports generating spatio-temporal time-series data corresponding to several random load changes across the network or corresponding to several generation changes. A Koopman mode decomposition (KMD) based algorithm to detect and identify the false data attacks in real-time is proposed in https://ieeexplore.ieee.org/document/9303022. Machine learning-based predictive models are developed to capture the dynamics of the underlying power system with a high level of accuracy under various operating conditions for IEEE 68 bus system. The corresponding machine learning models are available at https://github.com/pnnl/grid_prediction.

99 GENERAL AND MISCELLANEOUS↗

Non-Gaussianity in the weak lensing correlation function likelihood – implications for cosmological parameter biases

ABSTRACT We study the significance of non-Gaussianity in the likelihood of weak lensing shear two-point correlation functions, detecting significantly non-zero skewness and kurtosis in 1D marginal distributions of shear two-point correlation functions in simulated weak lensing data. We examine the implications in the context of future surveys, in particular LSST, with derivations of how the non-Gaussianity scales with survey area. We show that there is no significant bias in 1D posteriors of Ωm and σ8 due to the non-Gaussian likelihood distributions of shear correlations functions using the mock data (100 deg2). We also present a systematic approach to constructing approximate multivariate likelihoods with 1D parametric functions by assuming independence or more flexible non-parametric multivariate methods after decorrelating the data points using principal component analysis (PCA). While the use of PCA does not modify the non-Gaussianity of the multivariate likelihood, we find empirically that the 1D marginal sampling distributions of the PCA components exhibit less skewness and kurtosis than the original shear correlation functions. Modelling the likelihood with marginal parametric functions based on the assumption of independence between PCA components thus gives a lower limit for the biases. We further demonstrate that the difference in cosmological parameter constraints between the multivariate Gaussian likelihood model and more complex non-Gaussian likelihood models would be even smaller for an LSST-like survey. In addition, the PCA approach automatically serves as a data compression method, enabling the retention of the majority of the cosmological information while reducing the dimensionality of the data vector by a factor of ∼5.

79 ASTRONOMY AND ASTROPHYSICS↗

GenAI-Based Digital Twins Aided Data Augmentation Increases Accuracy in Real-Time Cokurtosis-Based Anomaly Detection of Wearable Data

Early detection of potential infectious disease outbreaks is crucial for developing effective interventions. In this study, we introduce advanced anomaly detection methods tailored for health datasets collected from wearables, offering insights at both individual and population levels. Leveraging real-world physiological data from wearables, including heart rate and activity, we developed a framework for the early detection of infection in individuals. Despite the availability of data from recent pandemics, substantial gaps remain in data collection, hindering method development. To bridge this gap, we utilized Wasserstein Generative Adversarial Networks (WGANs) to generate realistic synthetic wearable data, augmenting our dataset for training. Subsequently, we use these augmented datasets to implement a cokurtosis-based technique for anomaly detection in multivariate time-series data. Our approach includes a comprehensive assessment of uncertainties in synthetic data compared to the actual data upon which it was modeled, as well as the uncertainty associated with fine-tuning anomaly detection thresholds in physiological measurements. Through our work, we present an enhanced method for early anomaly detection in multivariate datasets, with promising applications in healthcare and beyond. This framework could revolutionize early detection strategies and significantly impact public health response efforts in future pandemics.

Data-Driven Digital Twins↗

Multivariate Statistical Analysis Software Technologies for Astrophysical Research Involving Large Data Bases

We developed a package to process and analyze the data from the digital version of the Second Palomar Sky Survey. This system, called SKICAT, incorporates the latest in machine learning and expert systems software technology, in order to classify the detected objects objectively and uniformly, and facilitate handling of the enormous data sets from digital sky surveys and other sources. The system provides a powerful, integrated environment for the manipulation and scientific investigation of catalogs from virtually any source. It serves three principal functions: image catalog construction, catalog management, and catalog analysis. Through use of the GID3* Decision Tree artificial induction software, SKICAT automates the process of classifying objects within CCD and digitized plate images. To exploit these catalogs, the system also provides tools to merge them into a large, complex database which may be easily queried and modified when new data or better methods of calibrating or classifying become available. The most innovative feature of SKICAT is the facility it provides to experiment with and apply the latest in machine learning technology to the tasks of catalog construction and analysis. SKICAT provides a unique environment for implementing these tools for any number of future scientific purposes. Initial scientific verification and performance tests have been made using galaxy counts and measurements of galaxy clustering from small subsets of the survey data, and a search for very high redshift quasars. All of the tests were successful and produced new and interesting scientific results. Attachments to this report give detailed accounts of the technical aspects of the SKICAT system, and of some of the scientific results achieved to date. We also developed a user-friendly package for multivariate statistical analysis of small and moderate-size data sets, called STATPROG. The package was tested extensively on a number of real scientific applications and has produced real, published results.

Djorgovski, S. G.↗

Explainable AI for Multivariate Time Series Pattern Exploration: Latent Space Visual Analytics With Temporal Fusion Transformer and Variational Autoencoders in Power Grid Event Diagnosis

Detecting and analyzing complex patterns in multivariate time-series data is crucial for decision-making in urban and environmental system operations. However, challenges arise from the high dimensionality, intricate complexity, and interconnected nature of complex patterns, which hinder the understanding of their underlying physical processes. Existing AI methods often face limitations in interpretability, computational efficiency, and scalability, reducing their applicability in real-world scenarios. This paper proposes a novel visual analytics framework that integrates two generative AI models, Temporal Fusion Transformer (TFT) and Variational Autoencoders (VAEs), to reduce complex patterns into lower-dimensional latent spaces and visualize them in 2D using dimensionality reduction techniques such as PCA, t-SNE, and UMAP with DBSCAN. These visualizations, presented through coordinated and interactive views and tailored glyphs, enable intuitive exploration of complex multivariate temporal patterns, identifying patterns’ similarities and uncover their potential correlations for a better interpretability of the AI outputs. The framework is demonstrated through a case study on power grid signal data, where it identifies multi-label grid event signatures, including faults and anomalies with diverse root causes. Additionally, novel metrics and visualizations are introduced to validate the models and assess the performance, efficiency, and consistency of latent maps generated by VAE, which have been utilized in prior studies for latent space cartography and used as a benchmark in this study, and the emerging TFT architecture under various configurations. These analyses provide actionable insights for model parameter tuning and reliability improvements. Comparative results highlight that TFT achieves shorter run times and superior scalability to diverse time-series data shapes compared to VAE. This work advances fault diagnosis in multivariate time series, fostering explainable AI to support critical system operations.

Explainable AI↗