Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Logistic Risk Model for the Unique Effects of Inherent Aerobic Capacity on (+)G(sub z) Tolerance Before and After Simulated Weightlessness

Small sample size (n less than 1O) and inappropriate analysis of multivariate data have hindered previous attempts to describe which physiologic and demographic variables are most important in determining how long humans can tolerate acceleration. Data from previous centrifuge studies conducted at NASA/Ames Research Center, utilizing a 7-14 d bed rest protocol to simulate weightlessness, were included in the current investigation. After review, data on 25 women and 22 men were available for analysis. Study variables included gender, age, weight, height, percent body fat, resting heart rate, mean arterial pressure, Vo(sub 2)max and plasma volume. Since the dependent variable was time to greyout (failure), two contemporary biostatistical modeling procedures (proportional hazard and logistic discriminant function) were used to estimate risk, given a particular subject's profile. After adjusting for pro-bed-rest tolerance time, none of the profile variables remained in the risk equation for post-bed-rest tolerance greyout. However, prior to bed rest, risk of greyout could be predicted with 91% accuracy. All of the profile variables except weight, MAP, and those related to inherent aerobic capacity (Vo(sub 2)max, percent body fat, resting heart rate) entered the risk equation for pro-bed-rest greyout. A cross-validation using 24 new subjects indicated a very stable model for risk prediction, accurate within 5% of the original equation. The result for the inherent fitness variables is significant in that a consensus as to whether an increased aerobic capacity is beneficial or detrimental has not been satisfactorily established. We conclude that tolerance to +Gz acceleration before and after simulated weightlessness is independent of inherent aerobic fitness.

Ludwig, David A.↗

Low power, lightweight vapor sensing using arrays of conducting polymer composite chemically-sensitive resistors

Arrays of broadly responsive vapor detectors can be used to detect, identify, and quantify vapors and vapor mixtures. One implementation of this strategy involves the use of arrays of chemically-sensitive resistors made from conducting polymer composites. Sorption of an analyte into the polymer composite detector leads to swelling of the film material. The swelling is in turn transduced into a change in electrical resistance because the detector films consist of polymers filled with conducting particles such as carbon black. The differential sorption, and thus differential swelling, of an analyte into each polymer composite in the array produces a unique pattern for each different analyte of interest, Pattern recognition algorithms are then used to analyze the multivariate data arising from the responses of such a detector array. Chiral detector films can provide differential detection of the presence of certain chiral organic vapor analytes. Aspects of the spaceflight qualification and deployment of such a detector array, along with its performance for certain analytes of interest in manned life support applications, are reviewed and summarized in this article.

NASA Discipline Life Sciences Technologies↗

Interval Predictor Models for Robust System Identification

This paper proposes a framework for the identification and uncertainty quantification of plant models according to multivariable data. The only restriction imposed upon such models is for their outputs to depend continuously on their parameters. An Interval Predictor Model (IPM) prescribes the parameters of a computational model as a path-connected set thereby making each predicted output an interval-valued function of its inputs. The formulation proposed seeks the parameter set for which the predicted outputs tightly enclose the data. This set, which is modeled as a semi-algebraic set of low-degree polynomials, enables the characterization of possibly strong parameter dependencies commonly found in practice. This uncertainty characterization makes the resulting plant model amenable to robust control approaches using polynomial optimization. Furthermore, we use non-convex scenario theory to assess the reliability of the resulting IPM. This assessment yields a distribution-free upper bound on the probability that future data will fall outside the predicted intervals.

interval↗

Development of A Multidecadal Land Reanalysis Over High Mountain Asia

Anthropogenic and climatic changes affect the water and energy cycles in High Mountain Asia (HMA), home to over two billion people and the largest reservoirs of freshwater outside the polar zone. Despite their significant importance for water management, consistent and reliable estimates of water storage and fluxes over the region are lacking because of the high uncertainties associated with the estimates of atmospheric conditions and human management. Here, we relied on multivariate data assimilation (MVDA) to provide estimates of energy and water storage and fluxes that reflect the processes occurring in the region such as greening and irrigation-driven groundwater depletion. We developed and employed an ensemble precipitation estimate by blending different precipitation products thereby reducing the uncertainties and inconsistencies associated with precipitation in HMA. Then, we assimilated five variables that capture the changes in hydrology in response to climate change and anthropogenic activities. Overall, our results have shown that MVDA has allowed a better representation of the land surface processes including greening and irrigation-driven groundwater depletion in HMA.

Fadji Z. Maina↗

Visual HPC Workflows for the Analysis of System Dynamics Models

Visual analytics supported by high performance computing (HPC) accelerates and enhances the discovery, exploration, and analysis of causal patterns in complex system dynamics (SD) models. We present a suite of visualization-assisted ensemble-based techniques for hypothesis generation and testing, and for sensitivity analysis. By employing HPC to provide parallel, on-demand simulation of SD models, one can “steer” an ensemble of simulated scenarios in real time as one first formulates and then informally tests those hypotheses: this provides rapid feedback for analysts to refine their understanding of the causal relationships emergent from a model. Such understandings can be followed and augmented by rigorous application of statistical methods, namely global variance-based sensitivity analysis, Monte-Carlo filtering, adaptive regional sensitivity analysis, and self-organized maps: here timely computation relies on HPC, while effective presentation emphasizes high-dimensional multivariate data visualization. Immersive visualization in virtual 3D environments provides an excellent adjunct to the traditional 2D graphics typically used for SD models, as it generates an embodied understanding of model behavior and facilitates an active, collaborative critique of model structure and output. Finally, we summarize prospects for HPC-enabled visual analytics applied to SD modeling.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Investigation of the Performance and Explainability Tradeoffs for Machine-Learning Models for Predictive Maintenance of Circulating Water Systems in Nuclear Power Plants

Predictive maintenance (PdM) has shown great potential for achieving substantial cost savings and enhancing the economic competitiveness of nuclear power plants (NPPs) in today's energy market. Among the different modeling approaches that exist, machine learning (ML) tools in particular have a demonstrated ability to handle high dimensional and multivariate data and to extract hidden relationships within data in industrial environments. While ML methods show great potential, their lack of explainability---especially for black-box models---is a major hurdle to their adoption. Moreover, considering the supposed trade-off between explainability and performance challenges, careful consideration must be made as to which of these quality aspects takes precedence in light of multiple modeling options, resource availability, and domain characteristics. The present work evaluates the performance of six ML models, each with a different degree of explainability, in classifying the conditions of circulating water pumps (CWPs) by utilizing sensor data from nuclear power plants. To determine the drivers behind the trade-offs presented by this array of models, this work also tests different combinations of CWP units as the training and testing data, degrees of data imbalance, and objective functions for hyperparameter tuning. It was found that black-box models tend to afford superior performance in cases where there are far more instances of one type of labeled data than of any other type. It is recommended that a guided procedure be followed for designing and delivering an ML system that is sufficiently explainable to all involved stakeholders.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Identification of multivariable high performance turbofan engine dynamics from closed loop data

The multivariable instrumental variable/approximate maximum likelihood (IV/AML) method or recursive time-series analysis is used to identify the multivariable (four inputs-three outputs) dynamics of the Pratt and Whitney F100 engine. A detailed nonlinear engine simulation is used to determine linear engine model structures and parameters at an operating point using open loop data. Also, the IV/AML method is used in a direct identification mode to identify models from actual closed loop engine test data. Models identified from simulated and test data are compared to determine a final model structure and parameterization that can predict engine response for a wide class of inputs. The ability of the IV/AML algorithm to identify useful dynamic models from engine test data is assessed.

Merrill, W.↗

Identification of multivariable high performance turbofan engine dynamics from closed loop data

The multivariable instrumental variable/approximate maximum likelihood (IV/AML) method of recursive time-series analysis is used to identify the multivariable (four inputs-three outputs) dynamics of the Pratt and Whitney F100 engine. A detailed nonlinear engine simulation is used to determine linear engine model structures and parameters at an operating point using open loop data. Also, the IV/AML method is used in a direct identification made to identify models from actual closed loop engine test data. Models identified from simulated and test data are compared to determine a final model structure and parameterization that can predict engine response for a wide class of inputs. The ability of the IV/AML algorithm to identify useful dynamic models from engine test data is assessed. Previously announced in STAR as N82-20339

Merrill, W.↗

Linear Multivariable Regression Models for Prediction of Eddy Dissipation Rate from Available Meteorological Data

Linear multivariable regression models for predicting day and night Eddy Dissipation Rate (EDR) from available meteorological data sources are defined and validated. Model definition is based on a combination of 1997-2000 Dallas/Fort Worth (DFW) data sources, EDR from Aircraft Vortex Spacing System (AVOSS) deployment data, and regression variables primarily from corresponding Automated Surface Observation System (ASOS) data. Model validation is accomplished through EDR predictions on a similar combination of 1994-1995 Memphis (MEM) AVOSS and ASOS data. Model forms include an intercept plus a single term of fixed optimal power for each of these regression variables; 30-minute forward averaged mean and variance of near-surface wind speed and temperature, variance of wind direction, and a discrete cloud cover metric. Distinct day and night models, regressing on EDR and the natural log of EDR respectively, yield best performance and avoid model discontinuity over day/night data boundaries.

MCKissick, Burnell T.↗

A Network-Based Algorithm for Clustering Multivariate Repeated Measures Data

The National Aeronautics and Space Administration (NASA) Astronaut Corps is a unique occupational cohort for which vast amounts of measures data have been collected repeatedly in research or operational studies pre-, in-, and post-flight, as well as during multiple clinical care visits. In exploratory analyses aimed at generating hypotheses regarding physiological changes associated with spaceflight exposure, such as impaired vision, it is of interest to identify anomalies and trends across these expansive datasets. Multivariate clustering algorithms for repeated measures data may help parse the data to identify homogeneous groups of astronauts that have higher risks for a particular physiological change. However, available clustering methods may not be able to accommodate the complex data structures found in NASA data, since the methods often rely on strict model assumptions, require equally-spaced and balanced assessment times, cannot accommodate missing data or differing time scales across variables, and cannot process continuous and discrete data simultaneously. To fill this gap, we propose a network-based, multivariate clustering algorithm for repeated measures data that can be tailored to fit various research settings. Using simulated data, we demonstrate how our method can be used to identify patterns in complex data structures found in practice.

Koslovsky, Matthew↗

Effect of vaccination on the case fatality rate for COVID-19 infections 2020–2021: multivariate modelling of data from the US Department of Veterans Affairs

Objectives: To evaluate the benefits of vaccination on the case fatality rate (CFR) for COVID-19 infections. Design, setting and participants: The US Department of Veterans Affairs has 130 medical centres. We created multivariate models from these data—339 772 patients with COVID-19—as of 30 September 2021. Outcome measures: The primary outcome for all models was death within 60 days of the diagnosis. Logistic regression was used to derive adjusted ORs for vaccination and infection with Delta versus earlier variants. Models were adjusted for confounding factors, including demographics, comorbidity indices and novel parameters representing prior diagnoses, vital signs/baseline laboratory tests and outpatient treatments. Patients with a Delta infection were divided into eight cohorts based on the time from vaccination to diagnosis. A common model was used to estimate the odds of death associated with vaccination for each cohort relative to that of unvaccinated patients. Results: 9.1% of subjects were vaccinated. 21.5% had the Delta variant. 18 120 patients (5.33%) died within 60 days of their diagnoses. The adjusted OR for a Delta infection was 1.87±0.05, which corresponds to a relative risk (RR) of 1.78. The overall adjusted OR for prior vaccination was 0.280±0.011 corresponding to an RR of 0.291. Raw CFR rose steadily after 10–14 weeks. The OR for vaccination remained stable for 10–34 weeks. Conclusions: Our CFR model controls for the severity of confounding factors and priority of vaccination, rather than solely using the presence of comorbidities. Our results confirm that Delta was more lethal than earlier variants and that vaccination is an effective means of preventing death. After adjusting for major selection biases, we found no evidence that the benefits of vaccination on CFR declined over 34 weeks. We suggest that this model can be used to evaluate vaccines designed for emerging variants.

59 BASIC BIOLOGICAL SCIENCES↗

Large Deviations Anomaly Detection (LAD) for collection of multivariate time series data: Applications to COVID-19 data

Time series anomaly detection is frequently used to identify extreme behaviors within a single time series. Identifying extreme trends in relation to a collection of other time series, on the other hand, is frequently of significant interest, such as in public health policy, social justice, and pandemic propagation. Using concepts from large deviations theory , we propose an algorithm that can scale to large collections of time series data. This paper expands on the LAD algorithm presented in Guggilam et al. (2022). The proposed algorithm is an online anomaly detection method for identifying anomalies in a collection of multivariate time series that takes advantage of the algorithm’s ability to scale to high-dimensional data. We show how the proposed Large Deviations Anomaly Detection (LAD) algorithm can be used to identify regions with anomalous trends in COVID-19 cases, deaths, biweekly growth rates, vaccinations, and fatality rates. Several of the observed anomalous trends are associated with regions that have demonstrated poor response to the COVID pandemic.

97 MATHEMATICS AND COMPUTING↗

From multivariate to functional data analysis: Fundamentals, recent developments, and emerging areas

Functional data analysis (FDA), which is a branch of statistics on modeling infinite dimensional random vectors resided in functional spaces, has become a major research area for Journal of Multivariate Analysis. We review some fundamental concepts of FDA, their origins and connections from multivariate analysis, and some of its recent developments, including multi-level functional data analysis, high-dimensional functional regression, and dependent functional data analysis. Here, we also discuss the impact of these new methodology developments on genetics, plant science, wearable device data analysis, image data analysis, and business analytics. Two real data examples are provided to motivate our discussions.

97 MATHEMATICS AND COMPUTING↗

Wind Turbine Gearbox Failure Detection Through Cumulative Sum of Multivariate Time Series Data

The wind energy industry is continuously improving their operational and maintenance practice for reducing the levelized costs of energy. Anticipating failures in wind turbines enables early warnings and timely intervention, so that the costly corrective maintenance can be prevented to the largest extent possible. It also avoids production loss owing to prolonged unavailability. One critical element allowing early warning is the ability to accumulate small-magnitude symptoms resulting from the gradual degradation of wind turbine systems. Inspired by the cumulative sum control chart method, this study reports the development of a wind turbine failure detection method with such early warning capability. Specifically, the following key questions are addressed: what fault signals to accumulate, how long to accumulate, what offset to use, and how to set the alarm-triggering control limit. We apply the proposed approach to 2 years’ worth of Supervisory Control and Data Acquisition data recorded from five wind turbines. We focus our analysis on gearbox failure detection, in which the proposed approach demonstrates its ability to anticipate failure events with a good lead time.

17 WIND ENERGY↗

A distributed system for visualizing and analyzing multivariate and multidisciplinary data

The Linked Windows Interactive Data System (Link Winds) is being developed with NASA support. The objective of this proposal is to adapt and apply that system in a complex network environment containing elements to be found by scientists working multidisciplinary teams on very large scale and distributed data sets. The proposed three year program will develop specific visualization and analysis tools, to be exercised locally and remotely in the Link Winds environment, to demonstrate visual data analysis, interdisciplinary data analysis and cooperative and interactive televisualization and analysis of data by geographically separated science teams. These demonstrations will involve at least two science disciplines with the aim of producing publishable results.

Jacobson, Allan S.↗

A distributed system for visualizing and analyzing multivariate and multidisciplinary data

THe Linked Windows Interactive Data System (LinkWinds) is being developed with NASA support. The objective of this proposal is to adapt and apply that system in a complex network environment containing elements to be found by scientists working multidisciplinary teams on very large scale and distributed data sets. The proposed three year program will develop specific visualization and analysis tools, to be exercised locally and remotely in the LinkWinds environment, to demonstrate visual data analysis, interdisciplinary data analysis and cooperative and interactive televisualization and analysis of data by geographically separated science teams. These demonstrators will involve at least two science disciplines with the aim of producing publishable results.

Jacobson, Allan S.↗

GMT: A deep learning approach to generalized multivariate translation for scientific data analysis and visualization

In scientific visualization, despite the significant advances of deep learning for data generation, researchers have not thoroughly investigated the issue of data translation. We present a new deep learning approach called generalized multivariate translation (GMT) for multivariate time-varying data analysis and visualization. Like V2V, GMT assumes a preprocessing step that selects suitable variables for translation. However, unlike V2V, which only handles one-to-one variable translation during training and inference, GMT enables one-to-many and many-to-many variable translation in the same framework. We leverage the recent StarGAN design from multi-domain image-to-image translation to achieve this generalization capability. We experiment with different loss functions and injection strategies to explore the best choices and leverage pre-training for performance improvement. We compare GMT with other state-of-the-art methods (i.e., Pix2Pix, V2V, StarGAN). Furthermore, the results demonstrate the overall advantage of GMT in translation quality and generalization ability.

97 MATHEMATICS AND COMPUTING↗

Analysis/forecast experiments with a multivariate statistical analysis scheme using FGGE data

A three-dimensional, multivariate, statistical analysis method, optimal interpolation (OI) is described for modeling meteorological data from widely dispersed sites. The model was developed to analyze FGGE data at the NASA-Goddard Laboratory of Atmospherics. The model features a multivariate surface analysis over the oceans, including maintenance of the Ekman balance and a geographically dependent correlation function. Preliminary comparisons are made between the OI model and similar schemes employed at the European Center for Medium Range Weather Forecasts and the National Meteorological Center. The OI scheme is used to provide input to a GCM, and model error correlations are calculated for forecasts of 500 mb vertical water mixing ratios and the wind profiles. Comparisons are made between the predictions and measured data. The model is shown to be as accurate as a successive corrections model out to 4.5 days.

Baker, W. E.↗