Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Parametric Testing of Launch Vehicle FDDR Models

For the safe operation of a complex system like a (manned) launch vehicle, real-time information about the state of the system and potential faults is extremely important. The on-board FDDR (Failure Detection, Diagnostics, and Response) system is a software system to detect and identify failures, provide real-time diagnostics, and to initiate fault recovery and mitigation. The ERIS (Evaluation of Rocket Integrated Subsystems) failure simulation is a unified Matlab/Simulink model of the Ares I Launch Vehicle with modular, hierarchical subsystems and components. With this model, the nominal flight performance characteristics can be studied. Additionally, failures can be injected to see their effects on vehicle state and on vehicle behavior. A comprehensive test and analysis of such a complicated model is virtually impossible. In this paper, we will describe, how parametric testing (PT) can be used to support testing and analysis of the ERIS failure simulation. PT uses a combination of Monte Carlo techniques with n-factor combinatorial exploration to generate a small, yet comprehensive set of parameters for the test runs. For the analysis of the high-dimensional simulation data, we are using multivariate clustering to automatically find structure in this high-dimensional data space. Our tools can generate detailed HTML reports that facilitate the analysis.

Schumann, Johann↗

Using Gaussian windows to explore a multivariate data set

In an earlier paper, I recounted an exploratory analysis, using Gaussian windows, of a data set derived from the Infrared Astronomical Satellite. Here, my goals are to develop strategies for finding structural features in a data set in a many-dimensional space, and to find ways to describe the shape of such a data set. After a brief review of Gaussian windows, I describe the current implementation of the method. I give some ways of describing features that we might find in the data, such as clusters and saddle points, and also extended structures such as a 'bar', which is an essentially one-dimensional concentration of data points. I then define a distance function, which I use to determine which data points are 'associated' with a feature. Data points not associated with any feature are called 'outliers'. I then explore the data set, giving the strategies that I used and quantitative descriptions of the features that I found, including clusters, bars, and a saddle point. I tried to use strategies and procedures that could, in principle, be used in any number of dimensions.

Jaeckel, Louis A.↗

The State of the Art in Visualizing Dynamic Multivariate Networks

Abstract Most real‐world networks are both dynamic and multivariate in nature, meaning that the network is associated with various attributes and both the network structure and attributes evolve over time. Visualizing dynamic multivariate networks is of great significance to the visualization community because of their wide applications across multiple domains. However, it remains challenging because the techniques should focus on representing the network structure, attributes and their evolution concurrently. Many real‐world network analysis tasks require the concurrent usage of the three aspects of the dynamic multivariate networks. In this paper, we analyze current techniques and present a taxonomy to classify the existing visualization techniques based on three aspects: temporal encoding, topology encoding, and attribute encoding. Finally, we survey application areas and evaluation methods; and discuss challenges for future research.

Kale, Bharat↗

Modern CACSD using the Robust-Control Toolbox

The Robust-Control Toolbox is a collection of 40 M-files which extend the capability of PC/PRO-MATLAB to do modern multivariable robust control system design. Included are robust analysis tools like singular values and structured singular values, robust synthesis tools like continuous/discrete H(exp 2)/H infinity synthesis and Linear Quadratic Gaussian Loop Transfer Recovery methods and a variety of robust model reduction tools such as Hankel approximation, balanced truncation and balanced stochastic truncation, etc. The capabilities of the toolbox are described and illustated with examples to show how easily they can be used in practice. Examples include structured singular value analysis, H infinity loop-shaping and large space structure model reduction.

Chiang, Richard Y.↗

Using Clustering to Establish Climate Regimes from PCM Output

A multivariate statistical clustering technique--based on the k-means algorithm of Hartigan has been used to extract patterns of climatological significance from 200 years of general circulation model (GCM) output. Originally developed and implemented on a Beowulf-style parallel computer constructed by Hoffman and Hargrove from surplus commodity desktop PCs, the high performance parallel clustering algorithm was previously applied to the derivation of ecoregions from map stacks of 9 and 25 geophysical conditions or variables for the conterminous U.S. at a resolution of 1 sq km. Now applied both across space and through time, the clustering technique yields temporally-varying climate regimes predicted by transient runs of the Parallel Climate Model (PCM). Using a business-as-usual (BAU) scenario and clustering four fields of significance to the global water cycle (surface temperature, precipitation, soil moisture, and snow depth) from 1871 through 2098, the authors' analysis shows an increase in spatial area occupied by the cluster or climate regime which typifies desert regions (i.e., an increase in desertification) and a decrease in the spatial area occupied by the climate regime typifying winter-time high latitude perma-frost regions. The patterns of cluster changes have been analyzed to understand the predicted variability in the water cycle on global and continental scales. In addition, representative climate regimes were determined by taking three 10-year averages of the fields 100 years apart for northern hemisphere winter (December, January, and February) and summer (June, July, and August). The result is global maps of typical seasonal climate regimes for 100 years in the past, for the present, and for 100 years into the future. Using three-dimensional data or phase space representations of these climate regimes (i.e., the cluster centroids), the authors demonstrate the portion of this phase space occupied by the land surface at all points in space and time. Any single spot on the globe will exist in one of these climate regimes at any single point in time. By incrementing time, that same spot will trace out a trajectory or orbit between and among these climate regimes (or atmospheric states) in phase (or state) space. When a geographic region enters a state it never previously visited, a climatic change is said to have occurred. Tracing out the entire trajectory of a single spot on the globe yields a 'manifold' in state space representing the shape of its predicted climate occupancy. This sort of analysis enables a researcher to more easily grasp the multivariate behavior of the climate system.

Oglesby, Robert↗

Cross-national analysis of food security drivers: comparing results based on the Food Insecurity Experience Scale and Global Food Security Index

Abstract The second UN Sustainable Development Goal establishes food security as a priority for governments, multilateral organizations, and NGOs. These institutions track national-level food security performance with an array of metrics and weigh intervention options considering the leverage of many possible drivers. We studied the relationships between several candidate drivers and two response variables based on prominent measures of national food security: the 2019 Global Food Security Index (GFSI) and the Food Insecurity Experience Scale’s (FIES) estimate of the percentage of a nation’s population experiencing food security or mild food insecurity (FI ). We compared the contributions of explanatory variables in regressions predicting both response variables, and we further tested the stability of our results to changes in explanatory variable selection and in the countries included in regression model training and testing. At the cross-national level, the quantity and quality of a nation’s agricultural land were not predictive of either food security metric. We found mixed evidence that per-capita cereal production, per-hectare cereal yield, an aggregate governance metric, logistics performance, and extent of paid employment work were predictive of national food security. Household spending as measured by per-capita final consumption expenditure (HFCE) was consistently the strongest driver among those studied, alone explaining a median of 92% and 70% of variation (based on out-of-sample R 2 ) in GFSI and FI , respectively. The relative strength of HFCE as a predictor was observed for both response variables and was independent of the countries used for model training, the transformations applied to the explanatory variables prior to model training, and the variable selection technique used to specify multivariate regressions. The results of this cross-national analysis reinforce previous research supportive of a causal mechanism where, in the absence of exceptional local factors, an increase in income drives increase in food security. However, the strength of this effect varies depending on the countries included in regression model fitting. We demonstrate that using multiple response metrics, repeated random sampling of input data, and iterative variable selection facilitates a convergence of evidence approach to analyzing food security drivers.

42 ENGINEERING↗

Probabilistic Machine Learning Estimation of Ocean Mixed Layer Depth from Dense Satellite and Sparse In-Situ Observations

The ocean mixed layer plays an important role in the coupling between the upper ocean and atmosphere across a wide range of time scales. Estimation of the variability of the ocean mixed layer is therefore important for atmosphere-ocean prediction and analysis. The increasing coverage of in situ Argo profile data allows for an increasingly accurate analysis of the mixed layer depth (MLD) variability associated with deviations from the seasonal climatology. However, sampling rates are not sufficient to fully resolve subseasonal (<90 day) MLD variability. Yet, many multivariate observations-based analyses include implicit modeled subseasonal MLD variability. One analysis method is optimal interpolation of in situ data, but the interior analysis can be improved by leveraging surface data with regression or variational approaches. Here, we demonstrate how machine learning methods and satellite sea surface temperature, salinity, and height facilitate MLD estimation in a pilot study of two regions: the mid-latitude southern Indian and the eastern equatorial Pacific Oceans. We construct multiple machine learning architectures to produce weekly 1/2° gridded MLD anomaly fields (relative to a monthly climatology) with uncertainty estimates. We test multiple traditional and probabilistic machine learning techniques to compare both accuracy and probabilistic calibration. We validate our methodology by applying it to ocean model simulations. We find that incorporating sea surface data through a machine learning model improves the performance of spatiotemporal MLD variability estimation compared to optimal interpolation of Argo observations alone. These preliminary results are a promising first step for the application of machine learning to MLD prediction.

Machine Learning↗

Analysis models for the estimation of oceanic fields

A general model for statistically optimal estimates is presented for dealing with scalar, vector and multivariate datasets. The method deals with anisotropic fields and treats space and time dependence equivalently. Problems addressed include the analysis, or the production of synoptic time series of regularly gridded fields from irregular and gappy datasets, and the estimate of fields by compositing observations from several different instruments and sampling schemes. Technical issues are discussed, including the convergence of statistical estimates, the choice of representation of the correlations, the influential domain of an observation, and the efficiency of numerical computations.

Carter, E. F.↗

Self-reported health impacts of do-it-yourself air cleaner use in a smoke-impacted community

Smoke exposure from wildfires or residential wood burning for heat is a public health problem for many communities. Do-It-Yourself (DIY) portable air cleaners (PACs) are promoted as affordable alternatives to commercial PACs, but evidence of their effect on health outcomes is limited. Pilot test an evaluation of the effect of DIY PAC usage on self-reported symptoms, and investigate barriers and facilitators of PAC use, among members of a tribal community that routinely experiences elevated concentrations of fine particulate matter (PM 2.5 ) from smoke. We conducted studies in Fall 2021 (“wildfire study”; N = 10) and Winter 2022 (“wood stove study”; N = 17). Each study included four sequential one-to-two-week phases: 1) initial, 2) DIY PAC usage ≥8 h/day, 3) commercial PAC usage ≥8 h/day, and 4) air sensor with visual display and optional PAC use. We continuously monitored PAC usage and indoor/outdoor PM 2.5 concentrations in homes. Concluding each phase, we conducted phone surveys about participants’ symptoms, perceptions, and behaviors. We analyzed symptoms associated with PAC usage and conducted an analysis of indoor PM 2.5 concentrations as a mediating pathway using mixed effects multivariate linear regression. We categorized perceptions related to PACs into barriers and facilitators of use. No association was observed between PAC usage and symptoms, and the mediation analysis did not indicate that small observed trends were attributable to changes in indoor PM 2.5 concentrations. Small sample sizes hindered the ability to draw conclusions regarding the presence or absence of causal associations. DIY PAC usage was low; loud operating noise was a barrier to use. This research is novel in studying health effects of DIY PACs during wildfire and wood smoke exposures. Such research is needed to inform public health guidance. Recommendations for future studies on PAC use during smoke exposure include building flexibility of intervention timing into the study design.

54 ENVIRONMENTAL SCIENCES↗

The association between neighborhood obesogenic factors and prostate cancer risk and mortality: the Southern Community Cohort Study

Background: Prostate cancer is one of the leading causes of cancer-related mortality among men in the United States. We examined the role of neighborhood obesogenic attributes on prostate cancer risk and mortality in the Southern Community Cohort Study (SCCS). Methods: From the total of 34,166 SCCS male participants, 28,356 were included in the analysis. We assessed the relationship between neighborhood obesogenic factors [neighborhood socioeconomic status (nSES) and neighborhood obesogenic environment indices including the restaurant environment index, the retail food environment index, parks, recreational facilities, and businesses] and prostate cancer risk and mortality by controlling for individual-level factors using a multivariable Cox proportional hazards model. We further stratified prostate cancer risk analysis by race and body mass index (BMI). Results: Median follow-up time was 133 months [interquartile range (IQR): 103, 152], and the mean age was 51.62 (SD: ± 8.42) years. There were 1,524 (5.37%) prostate cancer diagnoses and 98 (6.43%) prostate cancer deaths during follow-up. Compared to participants residing in the wealthiest quintile, those residing in the poorest quintile had a higher risk of prostate cancer (aHR = 1.32, 95% CI 1.12–1.57, p = 0.001), particularly among non-obese men with a BMI < 30 (aHR = 1.46, 95% CI 1.07–1.98, p = 0.016). The restaurant environment index was associated with a higher prostate cancer risk in overweight (BMI ≥ 25) White men (aHR = 3.37, 95% CI 1.04–10.94, p = 0.043, quintile 1 vs. None). Obese Black individuals without any neighborhood recreational facilities had a 42% higher risk (aHR = 1.42, 95% CI 1.04–1.94, p = 0.026) compared to those with any access. Compared to residents in the wealthiest quintile and most walkable area, those residing within the poorest quintile (aHR = 3.43, 95% CI 1.54–7.64, p = 0.003) or the least walkable area (aHR = 3.45, 95% CI 1.22–9.78, p = 0.020) had a higher risk of prostate cancer death. Conclusion: Living in a lower-nSES area was associated with a higher prostate cancer risk, particularly among Black men. Restaurant and retail food environment indices were also associated with a higher prostate cancer risk, with stronger associations within overweight White individuals. Finally, residing in a low-SES neighborhood or the least walkable areas were associated with a higher risk of prostate cancer mortality.

60 APPLIED LIFE SCIENCES↗

Metabolome patterns identify active dechlorination in bioaugmentation consortium SDC-9™

Ultra-high performance liquid chromatography–high-resolution mass spectrometry (UPHLC–HRMS) is used to discover and monitor single or sets of biomarkers informing about metabolic processes of interest. The technique can detect 1000’s of molecules (i.e., metabolites) in a single instrument run and provide a measurement of the global metabolome, which could be a fingerprint of activity. Despite the power of this approach, technical challenges have hindered the effective use of metabolomics to interrogate microbial communities implicated in the removal of priority contaminants. Herein, our efforts to circumvent these challenges and apply this emerging systems biology technique to microbiomes relevant for contaminant biodegradation will be discussed. Chlorinated ethenes impact many contaminated sites, and detoxification can be achieved by organohalide-respiring bacteria, a process currently assessed by quantitative gene-centric tools (e.g., quantitative PCR). This laboratory study monitored the metabolome of the SDC-9™ bioaugmentation consortium during cis-1,2-dichloroethene (cDCE) conversion to vinyl chloride (VC) and nontoxic ethene. Untargeted metabolomics using an UHPLC-Orbitrap mass spectrometer and performed on SDC-9™ cultures at different stages of the reductive dechlorination process detected ~10,000 spectral features per sample arising from water-soluble molecules with both known and unknown structures. Multivariate statistical techniques including partial least squares-discriminate analysis (PLSDA) identified patterns of measurable spectral features (peak patterns) that correlated with dechlorination (in)activity, and ANOVA analyses identified 18 potential biomarkers for this process. Statistical clustering of samples with these 18 features identified dechlorination activity more reliably than clustering of samples based only on chlorinated ethene concentration and Dhc 16S rRNA gene abundance data, highlighting the potential value of metabolomic workflows as an innovative site assessment and bioremediation monitoring tool.

environmental monitoring↗

Regional climate change predictions from the Goddard Institute for Space Studies high resolution GCM

A new diagnostic tool is developed for examining relationships between the synoptic scale circulation and regional temperature distributions in GCMs. The 4 x 5 deg GISS GCM is shown to produce accurate simulations of the variance in the synoptic scale sea level pressure distribution over the U.S. An analysis of the observational data set from the National Meteorological Center (NMC) also shows a strong relationship between the synoptic circulation and grid point temperatures. This relationship is demonstrated by deriving transfer functions between a time-series of circulation parameters and temperatures at individual grid points. The circulation parameters are derived using rotated principal components analysis, and the temperature transfer functions are based on multivariate polynomial regression models. The application of these transfer functions to the GCM circulation indicates that there is considerable spatial bias present in the GCM temperature distributions. The transfer functions are also used to indicate the possible changes in U.S. regional temperatures that could result from differences in synoptic scale circulation between a 1XCO2 and a 2xCO2 climate, using a doubled CO2 version of the same GISS GCM.

Crane, Robert G.↗

A Step Beyond Simple Keyword Searches: Services Enabled by a Full Content Digital Journal Archive

The problems of managing and searching large archives of scientific journal articles can potentially be addressed through data mining and statistical techniques matured primarily for quantitative scientific data analysis. A journal paper could be represented by a multivariate descriptor, e.g., the occurrence counts of a number key technical terms or phrases (keywords), perhaps derived from a controlled vocabulary ( e . g . , the American Meteorological Society's Glossary of Meteorology) or bootstrapped from the journal archive itself. With this technique, conventional statistical classification tools can be leveraged to address challenges faced by both scientists and professional societies in knowledge management. For example, cluster analyses can be used to find bundles of "most-related" papers, and address the issue of journal bifurcation (when is a new journal necessary, and what topics should it encompass). Similarly, neural networks can be trained to predict the optimal journal (within a society's collection) in which a newly submitted paper should be published. Comparable techniques could enable very powerful end-user tools for journal searches, all premised on the view of a paper as a data point in a multidimensional descriptor space, e.g.: "find papers most similar to the one I am reading", "build a personalized subscription service, based on the content of the papers I am interested in, rather than preselected keywords", "find suitable reviewers, based on the content of their own published works", etc. Such services may represent the next "quantum leap" beyond the rudimentary search interfaces currently provided to end-users, as well as a compelling value-added component needed to bridge the print-to-digital-medium gap, and help stabilize professional societies' revenue stream during the print-to-digital transition.

Boccippio, Dennis J.↗

Automated Framework for Groundwater Monitoring Using DWT with LSTM and Transformers

Environmental monitoring is critical for safeguarding public health and ecological well-being. Traditional data structuring and workflow monitoring methods consume significant time and effort, hindering timely insights and effective decision-making. Our study addresses this challenge by presenting an AI framework that automates data cleaning, structuring, and modeling processes, specifically targeting applications in groundwater monitoring. By leveraging automation for data processing and model training, our framework establishes a novel and efficient paradigm for environmental monitoring, with its potential application to the vast network of over a hundred Department of Energy Environmental Management (DoE-EM) cleanup sites across the country. It analyzes data streams from a network of groundwater Internet-of-Things (IoT) sensors deployed at the Savannah River Site (SRS) for prediction modeling. This allows human experts to focus on analysis and decision-making, ultimately leading to better environmental outcomes.The framework employs multivariate time-series forecasting methods to study and model the behavior of varying chemical analytes. The continuous learning process is enabled by utilizing deep learning techniques. It allows the framework to become more nuanced in its analysis over time, adapting to the specific characteristics of the environmental site and the evolving nature of contaminant behavior. Deep learning models known for sequence modeling, LSTM, and Transformers are employed for time series forecasting. Data processing and structuring are essential components significantly impacting the final model's performance. This hypothesis was proven by presenting a comparative analysis of model performance with processed and unprocessed data. The feature engineering approach utilized was the Discrete Wavelet Transform, which works well with time series data.

Discrete Wavelet Transform (DWT)↗

Multiple burn fuel-optimal orbit transfers: Numerical trajectory computation and neighboring optimal feedback guidance

This report describes current work in the numerical computation of multiple burn, fuel-optimal orbit transfers and presents an analysis of the second variation for extremal multiple burn orbital transfers as well as a discussion of a guidance scheme which may be implemented for such transfers. The discussion of numerical computation focuses on the use of multivariate interpolation to aid the computation in the numerical optimization. The second variation analysis includes the development of the conditions for the examination of both fixed and free final time transfers. Evaluations for fixed final time are presented for extremal one, two, and three burn solutions of the first variation. The free final time problem is considered for an extremal two burn solution. In addition, corresponding changes of the second variation formulation over thrust arcs and coast arcs are included. The guidance scheme discussed is an implicit scheme which implements a neighboring optimal feedback guidance strategy to calculate both thrust direction and thrust on-off times.

Chuang, C.-H.↗

Global Nonlinear Parametric Modeling with Application to F-16 Aerodynamics

A global nonlinear parametric modeling technique is described and demonstrated. The technique uses multivariate orthogonal modeling functions generated from the data to determine nonlinear model structure, then expands each retained modeling function into an ordinary multivariate polynomial. The final model form is a finite multivariate power series expansion for the dependent variable in terms of the independent variables. Partial derivatives of the identified models can be used to assemble globally valid linear parameter varying models. The technique is demonstrated by identifying global nonlinear parametric models for nondimensional aerodynamic force and moment coefficients from a subsonic wind tunnel database for the F-16 fighter aircraft. Results show less than 10% difference between wind tunnel aerodynamic data and the nonlinear parameterized model for a simulated doublet maneuver at moderate angle of attack. Analysis indicated that the global nonlinear parametric models adequately captured the multivariate nonlinear aerodynamic functional dependence.

Morelli, Eugene A.↗

Global Nonlinear Parametric Modeling with Application to F-16 Aerodynamics

A global nonlinear parametric modeling technique is described and demonstrated. The technique uses multivariate orthogonal modeling functions generated from the data to determine nonlinear model structure, then expands each retained modeling function into an ordinary multivariate polynomial. The final model form is a finite multivariate power series expansion for the dependent variable in terms of the independent variables. Partial derivatives of the identified models can be used to assemble globally valid linear parameter varying models. The technique is demonstrated by identifying global nonlinear parametric models for nondimensional aerodynamic force and moment coefficients from a subsonic wind tunnel database for the F-16 fighter aircraft. Results show less than 10% difference between wind tunnel aerodynamic data and the nonlinear parameterized model for a simulated doublet maneuver at moderate angle of attack. Analysis indicated that the global nonlinear parametric models adequately captured the multivariate nonlinear aerodynamic functional dependence.

Morelli, Eugene A.↗

Statistical Analysis of Factors Riving Surface Ozone Variability over Continental South Africa

Statistical relationships between surface ozone (O3) concentration, precursor species and meteorological conditions in continental South Africa were examined from data obtained from measurement stations in north-eastern South Africa. Three multivariate statistical methods were applied in the investigation, i.e. multiple linear regression (MLR), principal component analysis (PCA) and –regression (PCR), and generalised additive model (GAM) analysis. The daily maximum 8-h moving average O3 concentrations were considered in these statistical models (dependent variable). MLR models indicated that meteorology and precursor species concentrations are able to explain ~50% of the variability in daily maximum O3 levels. MLR analysis revealed that atmospheric carbon monoxide (CO), temperature and relative humidity were the strongest factors affecting the daily O3 variability. In summer, daily O3 variances were mostly associated with relative humidity, while winter O3 levels were mostly linked to temperature and CO. PCA indicated that CO, temperature and relative humidity were not strongly collinear. GAM also identified CO, temperature and relative humidity as the strongest factors affecting the daily variation of O3. Partial residual plots found that temperature, radiation and nitrogen oxides most likely have a non-linear relationship with O3,while the relationship with relative humidity and CO is probably linear. An inter-comparison between O3 levels modelled with the three statistical models compared to measured O3 concentrations showed that the GAM model offered a slight improvement over the MLR model. These findings emphasise the critical role of regional-scale O3 precursors coupled with meteorological conditions in daily variances of O3 levels in continental South Africa.

multiple linear regression (MLR)↗