Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Characterizing Spatiotemporal Uncertainty in Interpolated Meteorological Data

Interpolated meteorological data invariably contain errors. These errors have structure in time and space, particularly autocorrelation, which can cause the effects of errors to compound when model outputs are aggregated temporally or spatially. One way to account for this uncertainty is with a probabilistic model from which samples can be drawn that are coherent with respect to underlying spatial and temporal covariance structure. This work describes a probabilistic method for spatial interpolation of point-wise meteorological time series. Observational data from weather stations are generally sparse in space and dense in time (but sometimes missing). The method works by projecting time series onto orthogonal basis vectors and spatially interpolating each resulting component independently. Under suitable assumptions, and data transformations to better satisfy those assumptions, Gaussian process regression provides a complete description of the joint predictive distribution over a Gaussian random field. Spatiotemporally coherent realizations are generated as the sum of conditional (spatial) simulations of each orthogonal (temporal) component. Data-derived and generic orthogonal bases are considered. In addition to spatial interpolation, imputation of missing observational data is examined. The method is applied using near-surface air temperature over the Western United States and validated by comparing theoretical versus actual coverage of predictive distributions and analyzing the degree to which spatial and temporal covariance structure is reproduced. Computational considerations, relating to conditional simulation of random fields, are also addressed.

Conor T Doherty↗

A LANDSAT study of ephemeral and perennial rangeland vegetation and soils

The author has identified the following significant results. Several methods of computer processing were applied to LANDSAT data for mapping vegetation characteristics of perennial rangeland in Montana and ephemeral rangeland in Arizona. The choice of optimal processing technique was dependent on prescribed mapping and site condition. Single channel level slicing and ratioing of channels were used for simple enhancement. Predictive models for mapping percent vegetation cover based on data from field spectra and LANDSAT data were generated by multiple linear regression of six unique LANDSAT spectral ratios. Ratio gating logic and maximum likelihood classification were applied successfully to recognize plant communities in Montana. Maximum likelihood classification did little to improve recognition of terrain features when compared to a single channel density slice in sparsely vegetated Arizona. LANDSAT was found to be more sensitive to differences between plant communities based on percentages of vigorous vegetation than to actual physical or spectral differences among plant species.

Bentley, R. G., Jr.↗

Regression and ratio estimators to integrate AVHRR and MSS data

Regression and ratio estimators are used to integrate AVHRR-Global Area Coverage (GAC) and Landsat MSS digital data to estimate forest area in the continental United States. Forestlands are enumerated for the 48 contiguous states using five different AVHRR-GAC data sets. Results indicated that the GAC and MSS forest estimates were not highly correlated. Although the ratio of means and linear regression corrections were, on the average, closer to national U.S. Forest Service forest area estimates, these correction procedures did not consistently improve GAC estimates of forest area. GAC forest area estimates tended to be high in densely forested regions such as the northeast and low in sparsely forested areas.

Nelson, Ross↗

Global Analysis of Empirical Relationships Between Annual Climate and Seasonality of NDVI

This study describes the use of satellite data to calibrate a new climate-vegetation greenness function for global change studies. We examined statistical relationships between annual climate indexes (temperature, precipitation, and surface radiation) and seasonal attributes of the AVHRR Normalized Difference Vegetation Index (NDVI) time series for the mid-1980s in order to refine our empirical understanding of intraannual patterns and global abiotic controls on natural vegetation dynamics. Multiple linear regression results using global l(sup o) gridded data sets suggest that three climate indexes: growing degree days, annual precipitation total, and an annual moisture index together can account to 70-80 percent of the variation in the NDVI seasonal extremes (maximum and minimum values) for the calibration year 1984. Inclusion of the same climate index values from the previous year explained no significant additional portion of the global scale variation in NDVI seasonal extremes. The monthly timing of NDVI extremes was closely associated with seasonal patterns in maximum and minimum temperature and rainfall, with lag times of 1 to 2 months. We separated well-drained areas from l(sup o) grid cells mapped as greater than 25 percent inundated coverage for estimation of both the magnitude and timing of seasonal NDVI maximum values. Predicted monthly NDVI, derived from our climate-based regression equations and Fourier smoothing algorithms, shows good agreement with observed NDVI at a series of ecosystem test locations from around the globe. Regions in which NDVI seasonal extremes were not accurately predicted are mainly high latitude ecosystems and other remote locations where climate station data are sparse.

Potter, C. S.↗

Validation of the Pulmonary Function System for Use on the International Space Station

Aerobic deconditioning occurs during long duration space flight despite the use of exercise countermeasures (Convertino, 1996). As a part of International Space Station (ISS) medical operations, periodic tests designed to estimate aerobic capacity are performed to track changes in aerobic fitness and to determine the effectiveness of exercise countermeasures. These tests are performed prior to, during, and after missions of greater than 30 days in duration. Crewmembers selected for missions aboard the ISS perform a graded exercise test on a cycle ergometer approximately 270 days prior to their scheduled launch date in order to measure peak oxygen consumption (VO2PK) and peak heart rate (HRpk). Approximately 30 to 45 days prior to launch, crewmembers perform a submaximal cycle ergometer test at work rates set to elicit 25, 50 and 75% of their pre-flight VO2PK. This test, known as the Periodic Fitness Evaluation (PFE), serves as a baseline measure to which subsequent in-and post-flight exercise tests are compared. While onboard the ISS, crewmembers are normally scheduled to perform the PFE beginning with flight day (FD) 14 and every 30 days thereafter. The PFE is also conducted 5 and 30 days following flight. Using PFE data, aerobic fitness is estimated by quantifying the VO2 vs. HR relationship using linear regression and calculating the VO2 that would occur at the crewmember s previously measured HRpk. Currently, for data collected during flight, this technique assumes that the pre- vs. in-flight oxygen consumption per given cycle workload is similar. However, the validity of this assumption is based upon a sparse amount of data collected during the Skylab era (Michel, et al. 1977). The method of using heart rate and cycle ergometer work rates has been used to estimate aerobic fitness in normal gravity (Astrand and Ryhming, 1954; Lee, 1993). Due to spaceflight induced physiological alterations, such as shifts in extracellular fluid (e.g. plasma) volume, this method may not be valid during space flight. In addition, the ergometer onboard ISS is vibration-isolated and moves with the astronaut s application of force into the pedals. The effect of this movement on the VO2 of cycle exercise on ISS has not been quantified.

McCleary, Frank A.↗

An Investigation of Widespread Ozone Damage to the Soybean Crop in the Upper Midwest Determined From Ground-Based and Satellite Measurements

Elevated concentrations of ground-level ozone (O3) are frequently measured over farmland regions in many parts of the world. While numerous experimental studies show that O3 can significantly decrease crop productivity, independent verifications of yield losses at current ambient O3 concentrations in rural locations are sparse. In this study, soybean crop yield data during a 5-year period over the Midwest of the United States were combined with ground and satellite O3 measurements to provide evidence that yield losses on the order of 10% could be estimated through the use of a multiple linear regression model. Yield loss trends based on both conventional ground-based instrumentation and satellite-derived tropospheric O3 measurements were statistically significant and were consistent with results obtained from open-top chamber experiments and an open-air experimental facility (SoyFACE, Soybean Free Air Concentration Enrichment) in central Illinois. Our analysis suggests that such losses are a relatively new phenomenon due to the increase in background tropospheric O3 levels over recent decades. Extrapolation of these findings supports previous studies that estimate the global economic loss to the farming community of more than $10 billion annually.

Fishman, Jack↗

MLtool: Universal Supervised Machine Learning Tool to Model Tabulated Data

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine learning↗

MLtool Python Code

Machine Learning (ML) is a subfield of Artificial Intelligence that gives computers the ability to learn from past data without being explicitly programmed. The predictive capabilities of ML models have already been used to facilitate several scientific breakthroughs. However, the practical application of ML is often limited due to the gaps in technical knowledge of its users. The common issue faced by many scientific researchers is the inability to choose the appropriate ML pipelines that are needed to treat real-world data, which is often sparse and noisy. To solve this problem, we have developed an automated Machine Learning tool (MLtool) that includes a set of ML algorithms and approaches to aid scientific researchers. The current version of MLtool is implemented as an object-oriented Python code that is easily extensible. It includes 44 different regression algorithms used to model data. MLtool helps users select the best model for their data, based on the scoring metrics used. Besides regression algorithms, MLtool also includes a suite of pre- and post-processing techniques such as missing value imputation, categorical variable encoding, input feature normalization, uncertainty quantification, exploratory data analysis (EDA), etc. MLtool was tested on several publicly available multi-dimensional data sets and was found capable of making accurate predictions.

Machine Learning↗

Global Analysis of Empirical Relationships Between Annual Climate and Seasonality of NDVI

This paper describes the use of satellite data to calibrate a new climate-vegetation greenness relationship for global change studies. We examined statistical relationships between annual climate indexes (temperature, precipitation, and surface radiation) and seasonal attributes If the AVHRR Normalized Difference Vegetation Index (NDVI) time series for the mid-1980's in order to refine our understanding of intra-annual patterns and global abiotic controls on natural vegetation dynamics. Multiple linear regression results using global 1o gridded data sets suggest that three climate indexes: degree days (growing/chilling), annual precipitation total, and an annual moisture index together can account to 70-80 percent of the geographic variation in the NDVI seasonal extremes (maximum and minimum values) for the calibration year 1984. Inclusion of the same annual climate index values from the previous year explains no substantial additional portion of the global scale variation in NDVI seasonal extremes. The monthly timing of NDVI extremes is closely associated with seasonal patterns in maximum and minimum temperature and rainfall, with lag times of 1 to 2 months. We separated well-drained areas from lo grid cells mapped as greater than 25 percent inundated coverage for estimation of both the magnitude and timing of seasonal NDVI maximum values. Predicted monthly NDVI, derived from our climate-based regression equations and Fourier smoothing algorithms, shows good agreement with observed NDVI for several different years at a series of ecosystem test locations from around the globe. Regions in which NDVI seasonal extremes are not accurately predicted are mainly high latitude zones, mixed and disturbed vegetation types, and other remote locations where climate station data are sparse.

Potter, C. S.↗

Tropical cyclone track and genesis forecasting using satellite microwave sounder data

Although many dynamical and statistical prediction schemes are available to forecasters, tropical cyclone track errors are still large. One primary difficulty is that tropical cyclones exist over the data-sparse tropical oceans. Satellite sounders, however, routinely provide numerous data over these areas. Mean layer temperatures from the Scanning Microwave Spectrometer on board the Nimbus 6 satellite are decomposed using empirical orthogonal functions, and the expansion coefficients are related to deviations from the persistence forecast location, to speed change, to direction change and to intensity change. The significance of the regression equations is tested by a null hypothesis of zero correlation coefficient. It appears that significant information about tropical cyclone motion exists in the satellite-estimated mean layer temperatures, especially at upper levels. A physical interpretation of the statistical results is offered, and a one-storm-out independent test is used to test the stability of the equations. Finally, some further work is suggested.

Kidder, S. Q.↗

Estimation of multidimensional precipitation parameters by areal estimates of oceanic rainfall

The parameters of the multidimensional precipitation model proposed by Waymire et al. (1984) are estimated using the areal-averaged radar measurements of precipitation of the Global Atlantic Tropical Experiment (GATE) data set. The procedure followed was the fitting of the first- and second-order moments at different aggregation scales by nonlinear regression techniques. The numerical estimates of the parameters using different subsets of GATE information were reasonably stable, i.e., they were not affected by changes of the area-averaging size, temporal length of the records, and percentage of areal coverage of rainfall. This suggests that the estimation procedure is relatively robust and suitable to estimate the parameters of the multidimensional model in areas of sparse density of rain gages. The use of the space-time spectrum of rainfall to help in the determination of sampling errors due to intermittent visits of future space-borne low-altitude sensors of precipitation is also discussed.

Valdes, J. B.↗

Features of Point Clouds Synthesized from Multi-View ALOS/PRISM Data and Comparisons with LiDAR Data in Forested Areas

LiDAR waveform data from airborne LiDAR scanners (ALS) e.g. the Land Vegetation and Ice Sensor (LVIS) havebeen successfully used for estimation of forest height and biomass at local scales and have become the preferredremote sensing dataset. However, regional and global applications are limited by the cost of the airborne LiDARdata acquisition and there are no available spaceborne LiDAR systems. Some researchers have demonstrated thepotential for mapping forest height using aerial or spaceborne stereo imagery with very high spatial resolutions.For stereo imageswith global coverage but coarse resolution newanalysis methods need to be used. Unlike mostresearch based on digital surface models, this study concentrated on analyzing the features of point cloud datagenerated from stereo imagery. The synthesizing of point cloud data from multi-view stereo imagery increasedthe point density of the data. The point cloud data over forested areas were analyzed and compared to small footprintLiDAR data and large-footprint LiDAR waveform data. The results showed that the synthesized point clouddata from ALOSPRISM triplets produce vertical distributions similar to LiDAR data and detected the verticalstructure of sparse and non-closed forests at 30mresolution. For dense forest canopies, the canopy could be capturedbut the ground surface could not be seen, so surface elevations from other sourceswould be needed to calculatethe height of the canopy. A canopy height map with 30 m pixels was produced by subtracting nationalelevation dataset (NED) fromthe averaged elevation of synthesized point clouds,which exhibited spatial featuresof roads, forest edges and patches. The linear regression showed that the canopy height map had a good correlationwith RH50 of LVIS data with a slope of 1.04 and R2 of 0.74 indicating that the canopy height derived fromPRISM triplets can be used to estimate forest biomass at 30 m resolution.

LiDARD↗

Mapping tree canopy cover and canopy height with L-band SAR using LiDAR data and Random Forests

Light detection and ranging (LiDAR) data can provide direct measurements of vegetation structures but are limited by the sparse spatial coverage. Polarimetric synthetic aperture radar (SAR) can perform large-scale high-resolution mapping without weather constraints but the information about vegetation and ground subsurface are mixed in the backscatter data. In this paper, we adopted the Random Forests algorithm to train an upscaling function using tree canopy cover (TCC) and canopy height model (CHM) derived from Goddard’s LiDAR, Hyperspectral and Thermal Imager (G-LiHT) data. The regression model is then applied to the L-band Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) data acquired during the 2017 Arctic-Boreal Vulnerability Experiment (ABoVE) airborne campaign to map the TCC and CHM over the Delta Junction area in interior Alaska.

Moghaddam, Mahta↗

Classification of Dust Days by Satellite Remotely Sensed Aerosol Products

Considerable progress in satellite remote sensing (SRS) of dust particles has been seen in the last decade. From an environmental health perspective, such an event detection, after linking it to ground particulate matter (PM) concentrations, can proxy acute exposure to respirable particles of certain properties (i.e. size, composition, and toxicity). Being affected considerably by atmospheric dust, previous studies in the Eastern Mediterranean, and in Israel in particular, have focused on mechanistic and synoptic prediction, classification, and characterization of dust events. In particular, a scheme for identifying dust days (DD) in Israel based on ground PM10 (particulate matter of size smaller than 10 nm) measurements has been suggested, which has been validated by compositional analysis. This scheme requires information regarding ground PM10 levels, which is naturally limited in places with sparse ground-monitoring coverage. In such cases, SRS may be an efficient and cost-effective alternative to ground measurements. This work demonstrates a new model for identifying DD and non-DD (NDD) over Israel based on an integration of aerosol products from different satellite platforms (Moderate Resolution Imaging Spectroradiometer (MODIS) and Ozone Monitoring Instrument (OMI)). Analysis of ground-monitoring data from 2007 to 2008 in southern Israel revealed 67 DD, with more than 88 percent occurring during winter and spring. A Classification and Regression Tree (CART) model that was applied to a database containing ground monitoring (the dependent variable) and SRS aerosol product (the independent variables) records revealed an optimal set of binary variables for the identification of DD. These variables are combinations of the following primary variables: the calendar month, ground-level relative humidity (RH), the aerosol optical depth (AOD) from MODIS, and the aerosol absorbing index (AAI) from OMI. A logistic regression that uses these variables, coded as binary variables, demonstrated 93.2 percent correct classifications of DD and NDD. Evaluation of the combined CART-logistic regression scheme in an adjacent geographical region (Gush Dan) demonstrated good results. Using SRS aerosol products for DD and NDD, identification may enable us to distinguish between health, ecological, and environmental effects that result from exposure to these distinct particle populations.

satellite remote sensing↗

Probabilistic Machine Learning Estimation of Ocean Mixed Layer Depth from Dense Satellite and Sparse In-Situ Observations

The ocean mixed layer plays an important role in the coupling between the upper ocean and atmosphere across a wide range of time scales. Estimation of the variability of the ocean mixed layer is therefore important for atmosphere-ocean prediction and analysis. The increasing coverage of in situ Argo profile data allows for an increasingly accurate analysis of the mixed layer depth (MLD) variability associated with deviations from the seasonal climatology. However, sampling rates are not sufficient to fully resolve subseasonal (<90 day) MLD variability. Yet, many multivariate observations-based analyses include implicit modeled subseasonal MLD variability. One analysis method is optimal interpolation of in situ data, but the interior analysis can be improved by leveraging surface data with regression or variational approaches. Here, we demonstrate how machine learning methods and satellite sea surface temperature, salinity, and height facilitate MLD estimation in a pilot study of two regions: the mid-latitude southern Indian and the eastern equatorial Pacific Oceans. We construct multiple machine learning architectures to produce weekly 1/2° gridded MLD anomaly fields (relative to a monthly climatology) with uncertainty estimates. We test multiple traditional and probabilistic machine learning techniques to compare both accuracy and probabilistic calibration. We validate our methodology by applying it to ocean model simulations. We find that incorporating sea surface data through a machine learning model improves the performance of spatiotemporal MLD variability estimation compared to optimal interpolation of Argo observations alone. These preliminary results are a promising first step for the application of machine learning to MLD prediction.

Machine Learning↗

Comparing Ice Jam Hindcasting Models with Tree Scar Data

Hindcasting models can use historic ice jam observations and hydroclimatic data to identify conditions that form ice jams.However, historic ice jam records are often sparse or incomplete. New sources of historic ice jam data could improve hindcasting models, leading to better ice jam forecasting and flood warning systems. Because ice jams damage riparian trees, marker rings associated with historic scars include information about ice jam frequency and severity. This study examined marker rings from 56 trees along the Muskegon River to supplement the historic ice jam data on this system. The study team compared tree ring data to results from hindcasting models, which were independently validated with newspaper reports on 1,500 separate days. Logistic regression converted the marker ring data into annual ice jam probabilities. Ice jam dates from the dendrochronology data were too noisy to train or validate a hindcasting model. However, the marker ring data did confirm that ice jams on the Muskegon are nonstationary. Ice jams are significantly more likely now than they were before 1966.The marker rings also helped to distinguish between false-negatives and nondetects in the hindcasting model.

Stanford Gibson↗

Bayesian Geostatistical Modelling of PM10 and PM2.5 Surface Level Concentrations in Europe Using High-Resolution Satellite-Derived Products

Air quality monitoring across Europe is mainly based on in situ ground stations which are too sparse to accurately assess the exposure effects of air pollution for the entire continent. The demand for precise predictive modelsthat estimate gridded geophysical parameters of ambient air at high spatial resolution has rapidly grown. Here, we investigate the potential of satellite derived products to improve particulate matter (PM) estimates. Bayesiangeostatistical models addressing confounding between the spatial distribution of pollutants and remotely sensed predictors were developed to estimate yearly averages of both, fine (PM2.5) and coarse (PM10) surface PM concentrations at 1 sq.km spatial resolution over 46 European countries and were compared to geostatistical, geographically weighted and land-use regression formulations. Rigorous model selection identified the Earth observation data which contribute most to pollutants' estimation. Geostatistical models outperformed the predictive ability of the frequently employed land-use regression. The resulting estimates of PM10 and PM2.5, which represent the main air quality indicators for the urban Sustainable Development Goal, indicate that in 2016, 66.2% of the European population was breathing air above the WHO Air Quality Guidelines thresholds. Our estimates are readily available to policy makers and scientists assessing the effects of long-term exposure to pollution on human and ecosystem health.

Beloconi, Anton↗

Effects of Dose Error and Sample Size on Sonic Boom Dose-response Curves

NASA will soon be collecting noise-annoyance community survey data as the X-59 aircraft flies supersonically over several communities in the USA. Sparse measurements of the X-59 sonic thumps will be used together with physics-based simulations to estimate noise doses at survey participant locations. These dose estimates have associated error that affects the accuracy of modeled dose-response curves, which can result in misestimation of annoyance. The precision in dose-response curves is also a consideration in selecting the number of survey participants. To enable pretest studies of dose error and precision, simulated dose-response data were generated based on NASA’s Quiet Supersonic Flights 2018 test. The data included various degrees of dose error and sample size. Frequentist multilevel logistic regression models were fit to the true and perturbed dose-response data. Simple proportional relationships were identified between the model parameters and the perturbation standard deviation. The summary dose-response curves illustrate the impact on accuracy if dose error is not accounted for in the model. The precision in the dose-response curves is also shown as the number of participants and degree of participation is varied. Finally, sampling variability is illustrated by showing the dose-response curves for several replicates with random draws of participants and errors.

X-59↗