Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Multivariate regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Stratospheric Ozone Trends and Variability as Seen by SCIAMACHY from 2002 to 2012

Vertical profiles of the rate of linear change (trend) in the altitude range 15-50 km are determined from decadal O3 time series obtained from SCIAMACHY/ENVISAT measurements in limb-viewing geometry. The trends are calculated by using a multivariate linear regression. Seasonal variations, the quasi-biennial oscillation, signatures of the solar cycle and the El Nino-Southern Oscillation are accounted for in the regression. The time range of trend calculation is August 2002-April 2012. A focus for analysis are the zonal bands of 20 deg N - 20 deg S (tropics), 60 - 50 deg N, and 50 - 60 deg S (midlatitudes). In the tropics, positive trends of up to 5% per decade between 20 and 30 km and negative trends of up to 10% per decade between 30 and 38 km are identified. Positive O3 trends of around 5% per decade are found in the upper stratosphere in the tropics and at midlatitudes. Comparisons between SCIAMACHY and EOS MLS show reasonable agreement both in the tropics and at midlatitudes for most altitudes. In the tropics, measurements from OSIRIS/Odin and SHADOZ are also analysed. These yield rates of linear change of O3 similar to those from SCIAMACHY. However, the trends from SCIAMACHY near 34 km in the tropics are larger than MLS and OSIRIS by a factor of around two.

stratosphere↗

Sea Level Rise and City-Level Climate Action

Background: Climate change is the greatest threat to global health in the 21st century. Rising sea levels are one particularly concerning manifestation of this and many of the world’s largest cities are vulnerable to sea level rise (SLR). Thus, urban climate adaptation and mitigation policies are increasingly important to protect population health. Objectives: This study aimed to determine whether being at risk of SLR was associated with city-level climate action. It also aimed to assess the wider drivers of climate action in cities, in order to guide ongoing efforts to motivate climate action, assess public health preparedness and identify research gaps. Methods: This is an ecological cross-sectional study using secondary data from CDP, the Urban Climate Change Research Network (UCCRN), World Bank, United Nations Cities and EM-DAT (Emergency Events Database). The study population consisted of 517 cities who participated in CDP’s 2019 Cities Survey. Multivariable logistic regression was utilized to assess the relationship between risk of SLR and city-level climate action, and secondly, to assess the wider determinants of city-level climate action. Results: There was evidence of crude associations between risk of SLR and three outcome variables representing city-level climate action. However, after adjusting for confounding variables, these crude associations disappeared. World region, national income status and urban population were shown to be stronger predictors of city-level climate action. Conclusion: It is concerning for population health that there is no association demonstrated between risk of SLR and climate action. This could indicate a lack of awareness of the risks posed by SLR within urban governance. To fulfil their health protection responsibilities, it is essential that public health professionals take a leading role in advocating for climate action.

sea level rise↗

The INSTEP Monitoring Network: Merging High-and-Low Cost Measurements to Characterize California Wildfires

Despite challenges with data quality and scope, low-cost sensor networks have skyrocketed in popularity over the last 15 years, making air quality data available on refined spatial scales. More recently, studies have leveraged both high and low-quality instruments to create stronger “hybrid” models, with most studies focusing on particulate matter. Low-cost measurements typically represent ground-level emissions only, providing context for human health issues from climate change-driven events such as wildfires. Since low-cost sensors’ capabilities are localized, daily events and microclimates tend to dominate the data rather than larger regional or atmospheric trends. Likewise, their low cost explains their high uncertainty. In contrast, some regulatory-grade instruments produce column measurements as well, providing reliable information on a broader scope. To bridge this gap while expanding into gas-phase measurements, we deployed 12 air quality sensor packages in California, USA during the 2022 wildfire season. These INSTEP (Inexpensive Network Sensor Technology Exploring Pollution) monitors measure carbon monoxide (CO), carbon dioxide (CO2), ozone (O3), nitrogen dioxide (NO2), and several hydrocarbons including methane (CH4) and formaldehyde (HCHO). Half of the monitors were co-located with remote sensing spectrometers: NASA Pandora and Total Column Carbon Observing Network (TCCON). The overlap in pollutants includes NO2, O3, and HCHO between the INSTEP monitors and the Pandora column measurements. TCCON covers column CO, CO2, and CH4, rounding out our comparison. Most of the monitors were distributed throughout the San Francisco Bay area, and an additional three were located within 100 km of Los Angeles. The sites ranged in geographic and population characteristics, including desert, mountainous, coastal, and urban locations. Since varying environmental conditions such as temperature and pressure are known to challenge sensor performance, we will apply newer sensor “calibration” techniques meant to combat this. We will normalize our sensor signals by z-scoring them prior to applying a single calibration model in the form of multivariate linear regression or an artificial neural network. While this technique has been validated for the hydrocarbon and ozone sensor types (metal oxide), it has not yet been tested on electrochemical and non-dispersive infrared sensors, which are also used in the INSTEP monitors. This will serve as a test to see if this normalization technique – or another – is most effective in accounting for environmental differences among sensors. Related data analysis efforts have found success with a variety of geospatial analysis techniques, including weighted network models in which high-quality instruments are given higher weights than their low-cost counterparts. Our preliminary analysis will focus on kriging, which uses a Gaussian algorithm to assign weights, providing estimated pollution levels at locations between monitors. Smoke trajectory and evolution will also be considered using both measurement types. We also aim to baseline subtract our emission estimates from each region to determine which portion of emissions are regional and local, further characterizing burn differences in northern and southern California fires. Future directions include using INSTEP jointly with TEMPO satellite data, and mobile deployments on aircraft and uncrewed aerial vehicles (UAV).

Low-cost sensors↗

Multi-Variate LSTM Prediction of Alaska Magnetometer Chain Utilizing a Coupled Model Approach

During periods of rapidly changing geomagnetic conditions electric fields form within the Earth’s surface and induce currents known as geomagnetically induced currents(GICs), which interact with unprotected electrical systems our society relies on. In this study, we train multi-variate Long-Short Term Memory neural networks to predict magnitude of north-south component of the geomagnetic field (|BN|) at multiple ground magnetometer stations across Alaska provided by the SuperMAG database with a future goal of predicting geomagnetic field disturbances. Each neural network is driven by solar wind and interplanetary magnetic field inputs from the NASA OMNI database spanning from 2000–2015 and is fine tuned for each station to maximize the effectiveness in predicting |BN|. The neural networks are then compared against multivariate linear regression models driven with the same inputs at each station using Heidke skill scores with thresholds at the 50, 75, 85, and 99 percentiles for |BN|. The neural network models show significant increases over the linear regression models for |BN| thresholds. We also calculate the Heidke skill scores for d|BN|/dt by deriving d|BN|/dt from |BN| predictions. However, neural network models do not show clear outperformance compared to the linear regression models. To retain the sign information and thus predict BN instead of |BN|, a secondary so-called polarity model is utilized. The polarity model is run in tandem with the neural networks predicting geomagnetic field in a coupled model approach and results in a high correlation between predicted and observed values for all stations. We find this model a promising starting point for a machine learned geomagnetic field model to be expanded upon through increased output time history and fast turnaround times.

Matthew Blandin↗

Deriving Severe Hail Likelihood from Satellite Observations and Model Reanalysis Parameters using a Deep Neural Network

Geostationary satellite imagers, such as those of the Geostationary Operational Environmental Satellite (GOES) series, have been observing severe convection at 15–60-minute intervals for over 40 years. When properly assessed, such a data record can be valuable in efforts of estimating severe storm risk throughout the diurnal cycle based on automated detection of patterns consistently found atop severe storms. Furthermore, environmental conditions favorable for severe weather are well-known and are thought to be represented well by modern reanalysis products. Promoting resilience against such hazards on local and global scales is a chief goal the NASA Disasters program, which seeks to encourage use of satellite observations to mitigate risk. For instance, hail is the costliest severe weather hazard across the globe in terms of insured loss, but reporting inconsistencies for hail events globally make it difficult to develop models that can quantify the risk. Satellite observation and model reanalysis taken together have the potential to, with reasonable skill and specificity, characterize environmental conditions that are favorable for hazardous weather, and thereby enable creation of hazard climatologie. Such climatologies are particularly useful over regions without extensive radar networks or storm reporting. By mapping the multivariate combination of observed cloud features and reanalysis environmental parameters/indices to United States Next Generation Weather Radar (NEXRAD) radar-estimated Maximum Expected Size of Hail (MESH) by way of a deep neural network (DNN), estimates of likelihood for potentially severe hail can be produced. Such estimates are of greater complexity and efficiency than could be performed with previous multivariate or logistic regression analyses for observed points within convective systems. Statistical distributions of convective parameters from satellite and reanalysis are shown to highlight non-severe/severe class separation for well-known hailstorm predictors, e.g., overshooting cloud top characteristics, deep-layer wind shear, mid-level stability, helicity, and convective inhibition. These complex, multivariate predictor relationships are exploited within a DNN, which can efficiently produce a quantitative hail risk metric with better than 70% detection rate and under 30% false alarms. These hail classifications can then be aggregated across the satellite record to yield a hazard climatology for hail frequency and severity – knowledge of which is of particular interest to those who manage risk (e.g., insurers) and are seeking opportunities to identify hail-prone regions, particularly in developing nations. This NASA study uses satellite observations and model parameters in a DNN to perform climatological hailstorm analysis in support of catastrophe model development, with the hope of promoting risk resilience particularly in regions without adequate weather radar coverage.

Passive Remote Sensing↗

An Integrated Analysis of the Physiological Effects of Space Flight: Executive Summary

A large array of models were applied in a unified manner to solve problems in space flight physiology. Mathematical simulation was used as an alternative way of looking at physiological systems and maximizing the yield from previous space flight experiments. A medical data analysis system was created which consist of an automated data base, a computerized biostatistical and data analysis system, and a set of simulation models of physiological systems. Five basic models were employed: (1) a pulsatile cardiovascular model; (2) a respiratory model; (3) a thermoregulatory model; (4) a circulatory, fluid, and electrolyte balance model; and (5) an erythropoiesis regulatory model. Algorithms were provided to perform routine statistical tests, multivariate analysis, nonlinear regression analysis, and autocorrelation analysis. Special purpose programs were prepared for rank correlation, factor analysis, and the integration of the metabolic balance data.

Leonard, J. I.↗

Space Shuttle Main Engine performance analysis

For a number of years, NASA has relied primarily upon periodically updated versions of Rocketdyne's power balance model (PBM) to provide space shuttle main engine (SSME) steady-state performance prediction. A recent computational study indicated that PBM predictions do not satisfy fundamental energy conservation principles. More recently, SSME test results provided by the Technology Test Bed (TTB) program have indicated significant discrepancies between PBM flow and temperature predictions and TTB observations. Results of these investigations have diminished confidence in the predictions provided by PBM, and motivated the development of new computational tools for supporting SSME performance analysis. A multivariate least squares regression algorithm was developed and implemented during this effort in order to efficiently characterize TTB data. This procedure, called the 'gains model,' was used to approximate the variation of SSME performance parameters such as flow rate, pressure, temperature, speed, and assorted hardware characteristics in terms of six assumed independent influences. These six influences were engine power level, mixture ratio, fuel inlet pressure and temperature, and oxidizer inlet pressure and temperature. A BFGS optimization algorithm provided the base procedure for determining regression coefficients for both linear and full quadratic approximations of parameter variation. Statistical information relative to data deviation from regression derived relations was also computed. A new strategy for integrating test data with theoretical performance prediction was also investigated. The current integration procedure employed by PBM treats test data as pristine and adjusts hardware characteristics in a heuristic manner to achieve engine balance. Within PBM, this integration procedure is called 'data reduction.' By contrast, the new data integration procedure, termed 'reconciliation,' uses mathematical optimization techniques, and requires both measurement and balance uncertainty estimates. The reconciler attempts to select operational parameters that minimize the difference between theoretical prediction and observation. Selected values are further constrained to fall within measurement uncertainty limits and to satisfy fundamental physical relations (mass conservation, energy conservation, pressure drop relations, etc.) within uncertainty estimates for all SSME subsystems. The parameter selection problem described above is a traditional nonlinear programming problem. The reconciler employs a mixed penalty method to determine optimum values of SSME operating parameters associated with this problem formulation.

Santi, L. Michael↗

Experiments to Determine Whether Recursive Partitioning (CART) or an Artificial Neural Network Overcomes Theoretical Limitations of Cox Proportional Hazards Regression

New computationally intensive tools for medical survival analyses include recursive partitioning (also called CART) and artificial neural networks. A challenge that remains is to better understand the behavior of these techniques in effort to know when they will be effective tools. Theoretically they may overcome limitations of the traditional multivariable survival technique, the Cox proportional hazards regression model. Experiments were designed to test whether the new tools would, in practice, overcome these limitations. Two datasets in which theory suggests CART and the neural network should outperform the Cox model were selected. The first was a published leukemia dataset manipulated to have a strong interaction that CART should detect. The second was a published cirrhosis dataset with pronounced nonlinear effects that a neural network should fit. Repeated sampling of 50 training and testing subsets was applied to each technique. The concordance index C was calculated as a measure of predictive accuracy by each technique on the testing dataset. In the interaction dataset, CART outperformed Cox (P less than 0.05) with a C improvement of 0.1 (95% Cl, 0.08 to 0.12). In the nonlinear dataset, the neural network outperformed the Cox model (P less than 0.05), but by a very slight amount (0.015). As predicted by theory, CART and the neural network were able to overcome limitations of the Cox model. Experiments like these are important to increase our understanding of when one of these new techniques will outperform the standard Cox model. Further research is necessary to predict which technique will do best a priori and to assess the magnitude of superiority.

Kattan, Michael W.↗

Chemical studies of H chondrites. 6: Antarctic/non-Antarctic compositional differences revisited

We report data for the trace elements Au, Co, Sb, Ga, Rb, Ag, Se, Cs, Te, Zn, Cd, Bi, T1, and In (ordered by putative volatility during nebular condensation and accretion) determined by radiochemical neutron activation analysis of 14 additional H5 and H6 chondrite falls. Data for the 10 most volatile elements (Rb to In) treated by the multivariate techniques of linear discriminant analysis and logistic regression in these and 44 other falls are compared with those of 59 H4-6 chondrites from Antarctica. Various populations are tested by the multivariate techniques, using the previously developed method of randomization-simulation to assess significance levels. An earlier conclusion, based on fewer examples, that H4-6 chondrite falls are compositionally distinguishable from the Antarctic suite is verified by the additional data. This distinctiveness is highly significant because of the presence of samples from Victoria Land in the Antarctic population, which differ compositionally from falls beyond any reasonable doubt. However, it cannot be proven unequivocally that falls and Antarctic samples from Queen Maud Land are compositionally distinguishable. Trivial causes (e.g., analyst bias, weathering) cannot explain the Victoria Land (Antarctic)/non-Antarctic compositional difference for paradigmatic H4-6 chondrites. This seems to reflect a time-dependent variation of near-Earth meteoroid source regions differing in average thermal history.

Wolf, Stephen F.↗

An Assessment of the Regional Distribution of the Oxygen-Isotope Ratio in Northeastern Canada

A compilation of mean values of the oxygen-isotope ratio relative to standard mean ocean Water for 22 sites representative of conditions in north-eastern Canada is complemented with data on mean annual surface temperature, latitude, surface elevation, and mean annual shortest distance to open ocean denoted by the 10% sea-ice concentration boundary. Stepwise regression analysis is used to develop a multivariate model suitable to infer the distribution of 6 1"0 in an area of complex topography and possibly mixed source of advected water vapor. The best model is produced by a run in the backward mode at the 95% confidence level in which only temperature, latitude and distance to the open ocean remain in the model (the correlation coefficient is 0.915, the adjusted coefficient of determination is 0.809, the root mean square residual is 1.62). This model is similar to the best 6180 predictive model derived elsewhere for Greenland, suggesting a common principal source of advected moisture.

Giovinetto, Mario B.↗

Multivariate statistical analysis: Principles and applications to coorbital streams of meteorite falls

Multivariate statistical analysis techniques (linear discriminant analysis and logistic regression) can provide powerful discrimination tools which are generally unfamiliar to the planetary science community. Fall parameters were used to identify a group of 17 H chondrites (Cluster 1) that were part of a coorbital stream which intersected Earth's orbit in May, from 1855 - 1895, and can be distinguished from all other H chondrite falls. Using multivariate statistical techniques, it was demonstrated that a totally different criterion, labile trace element contents - hence thermal histories - or 13 Cluster 1 meteorites are distinguishable from those of 45 non-Cluster 1 H chondrites. Here, we focus upon the principles of multivariate statistical techniques and illustrate their application using non-meteoritic and meteoritic examples.

Wolf, S. F.↗

Statistical Evaluation of Time Series Analysis Techniques

The performance of a modified version of NASA's multivariate spectrum analysis program is discussed. A multiple regression model was used to make the revisions. Performance improvements were documented and compared to the standard fast Fourier transform by Monte Carlo techniques.

Benignus, V. A.↗

Partial Least Squares and Neural Networks for Quantitative Calibration of Laser-induced Breakdown Spectroscopy (LIBs) of Geologic Samples

The ChemCam instrument [1] on the Mars Science Laboratory (MSL) rover will be used to obtain the chemical composition of surface targets within 7 m of the rover using Laser Induced Breakdown Spectroscopy (LIBS). ChemCam analyzes atomic emission spectra (240-800 nm) from a plasma created by a pulsed Nd:KGW 1067 nm laser. The LIBS spectra can be used in a semiquantitative way to rapidly classify targets (e.g., basalt, andesite, carbonate, sulfate, etc.) and in a quantitative way to estimate their major and minor element chemical compositions. Quantitative chemical analysis from LIBS spectra is complicated by a number of factors, including chemical matrix effects [2]. Recent work has shown promising results using multivariate techniques such as partial least squares (PLS) regression and artificial neural networks (ANN) to predict elemental abundances in samples [e.g. 2-6]. To develop, refine, and evaluate analysis schemes for LIBS spectra of geologic materials, we collected spectra of a diverse set of well-characterized natural geologic samples and are comparing the predictive abilities of PLS, cascade correlation ANN (CC-ANN) and multilayer perceptron ANN (MLP-ANN) analysis procedures.

Anderson, R. B.↗

Neural network uncertainty assessment using Bayesian statistics: a remote sensing application

Neural network (NN) techniques have proved successful for many regression problems, in particular for remote sensing; however, uncertainty estimates are rarely provided. In this article, a Bayesian technique to evaluate uncertainties of the NN parameters (i.e., synaptic weights) is first presented. In contrast to more traditional approaches based on point estimation of the NN weights, we assess uncertainties on such estimates to monitor the robustness of the NN model. These theoretical developments are illustrated by applying them to the problem of retrieving surface skin temperature, microwave surface emissivities, and integrated water vapor content from a combined analysis of satellite microwave and infrared observations over land. The weight uncertainty estimates are then used to compute analytically the uncertainties in the network outputs (i.e., error bars and correlation structure of these errors). Such quantities are very important for evaluating any application of an NN model. The uncertainties on the NN Jacobians are then considered in the third part of this article. Used for regression fitting, NN models can be used effectively to represent highly nonlinear, multivariate functions. In this situation, most emphasis is put on estimating the output errors, but almost no attention has been given to errors associated with the internal structure of the regression model. The complex structure of dependency inside the NN is the essence of the model, and assessing its quality, coherency, and physical character makes all the difference between a blackbox model with small output errors and a reliable, robust, and physically coherent model. Such dependency structures are described to the first order by the NN Jacobians: they indicate the sensitivity of one output with respect to the inputs of the model for given input data. We use a Monte Carlo integration procedure to estimate the robustness of the NN Jacobians. A regularization strategy based on principal component analysis is proposed to suppress the multicollinearities in order to make these Jacobians robust and physically meaningful.

Neural Networks (Computer)↗

Ordinary chondrites - Multivariate statistical analysis of trace element contents

The contents of mobile trace elements (Co, Au, Sb, Ga, Se, Rb, Cs, Te, Bi, Ag, In, Tl, Zn, and Cd) in Antarctic and non-Antarctic populations of H4-6 and L4-6 chondrites, were compared using standard multivariate discriminant functions borrowed from linear discriminant analysis and logistic regression. A nonstandard randomization-simulation method was developed, making it possible to carry out probability assignments on a distribution-free basis. Compositional differences were found both between the Antarctic and non-Antarctic H4-6 chondrite populations and between two L4-6 chondrite populations. It is shown that, for various types of meteorites (in particular, for the H4-6 chondrites), the Antarctic/non-Antarctic compositional difference is due to preterrestrial differences in the genesis of their parent materials.

Lipschutz, Michael E.↗

Calibration or inverse regression: Which is appropriate for crop surveys using LANDSAT data?

Calibration and inverse regression estimators of crop proportions are investigated where the auxiliary variable is obtained from binary classification of multivariate LANDSAT data. The appropriate model relating classifier proportions and ground observed proportions for a given crop type is the calibration model. Under this model the inverse regression estimator is superior to the calibration estimator in estimating the crop acreage or proportion for a region of interest.

Chhikara, R. S.↗

Determinants of Time to Fatigue during Non-Motorized Treadmill Exercise

Treadmill exercise is commonly used for aerobic and anaerobic conditioning. During non-motorized treadmill exercise, the subject must provide the power necessary to drive the treadmill belt. The purpose of this study was to determine what factors affected the time to fatigue on a pair of non-motorized treadmills. Twenty subjects (10 males/10 females) attempted to complete five minutes of locomotion during separate trials at 3.22, 4.83, 6.44, 8.05, 9.66, and 11.27 km (raised dot) h(sup -1). Total exercise time (less than or equal to 5 min) was recorded. Exercise time was converted to the amount of 15 second intervals completed. Peak oxygen uptake (VO2) was measured using a graded exercise test on a standard treadmill, and anthropometric measures were collected from each subject before entering into the study. A Cox proportional hazards regression model was used to determine significant predictive factors in a multivariate analysis. Non-motorized treadmill speed and absolute peak VO2 were found to be significant predictors of exercise time, but there was no effect of anthropometric characteristics. Gender was found to be a predictor of treadmill time, but this was likely due to a higher peak VO2 in males than in females. These results were not affected by the type of treadmill tested in this study. Coaches and therapists should consider the cardiovascular fitness of an athlete or client when prescribing target speed since these factors are related to the total exercise time than can be achieved on a non-motorized treadmill.

DeWitt, John K.↗

Statistical Analysis of Factors Riving Surface Ozone Variability over Continental South Africa

Statistical relationships between surface ozone (O3) concentration, precursor species and meteorological conditions in continental South Africa were examined from data obtained from measurement stations in north-eastern South Africa. Three multivariate statistical methods were applied in the investigation, i.e. multiple linear regression (MLR), principal component analysis (PCA) and –regression (PCR), and generalised additive model (GAM) analysis. The daily maximum 8-h moving average O3 concentrations were considered in these statistical models (dependent variable). MLR models indicated that meteorology and precursor species concentrations are able to explain ~50% of the variability in daily maximum O3 levels. MLR analysis revealed that atmospheric carbon monoxide (CO), temperature and relative humidity were the strongest factors affecting the daily O3 variability. In summer, daily O3 variances were mostly associated with relative humidity, while winter O3 levels were mostly linked to temperature and CO. PCA indicated that CO, temperature and relative humidity were not strongly collinear. GAM also identified CO, temperature and relative humidity as the strongest factors affecting the daily variation of O3. Partial residual plots found that temperature, radiation and nitrogen oxides most likely have a non-linear relationship with O3,while the relationship with relative humidity and CO is probably linear. An inter-comparison between O3 levels modelled with the three statistical models compared to measured O3 concentrations showed that the GAM model offered a slight improvement over the MLR model. These findings emphasise the critical role of regional-scale O3 precursors coupled with meteorological conditions in daily variances of O3 levels in continental South Africa.

multiple linear regression (MLR)↗