Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Gaussian process regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

30 records · Page 2

PM 2.5 Concentrations over Major Metropolitan Regions Inferred from Airborne High Spectral Resolution Lidar Measurements Using Machine Learning Regression

We use measurements of near-surface aerosol backscatter, extinction, and depolarization acquired by four NASA Langley Research Center airborne High Spectral Resolution Lidars (HSRLs) to develop a machine learning regression methodology to infer PM2.5 concentrations at the surface and aloft. These airborne HSRL measurements were acquired over major metropolitan regions in the United States and Asia during more than 170 flights since 2010. Hourly surface PM2.5 measurements from the EPA air quality system and similar networks in other countries acquired within 10 km and 15 minutes of these near-surface HSRL measurements are used to train models that compute PM2.5 concentrations from the HSRL measurements. We examine several regression methods and find that exponential Gaussian Process algorithms consistently give the best performance in terms of the lowest root-mean-square (RMS) errors and the highest correlations. Model performance varies significantly depending on various combinations of HSRL aerosol measurements (e.g., aerosol backscatter, extinction, depolarization, backscatter color ratios, lidar ratios, aerosol optical thickness) and retrievals (e.g., mixed layer height, aerosol type) used in the regressions. Models that use near-surface measurements of aerosol backscatter and aerosol intensive properties such as depolarization, backscatter color ratio, and lidar ratio typically give the best performance with RMS errors around 4 mg/m3 and correlation coefficients above 0.9. HSRL measurements were often acquired when the aircraft flew systematic “raster-scan” patterns for several hours over these cities. These flight patterns enabled measurements of the spatial, temporal, and vertical variabilities in the distributions of aerosol backscatter and aerosol intensive properties and allowed us to derive the corresponding variabilities in PM2.5 concentrations. We present examples of such variabilities over urban areas in the United States as well as Asia. We describe also how the distribution of surface PM2.5 varies with aerosol type and use these retrievals to examine model simulations of surface PM2.5 in these metropolitan regions. We also discuss how this methodology may be applied to measurements from satellite lidars such as CALIOP on CALIPSO and ATLID on EarthCARE.

lidar

Structured Covariance Gaussian Networks for Orion Crew Module Aerodynamic Uncertainty Quantification

In this paper we propose a new approach for nonlinear regression and uncertainty quantification. The method is based on a pair of neural networks which parameterize mean and dense covariance functions of a multivariate Gaussian process, trained together to maximize the log-likelihood of observing the given data. The covariance matrix is made positive definite at every input by construction. We also propose a sampling approach that produces viable surrogate function realizations from the Gaussian process. We call the proposed model a Structured Covariance Gaussian Network (SCGN). We illustrate the use of SCGNs for learning an aerodynamic response surface with built-in uncertainty for the Orion crew module. We find that SCGN provides an efficient and systematic way to learn nonlinear functional relationships and dense covariances. We compare results to a baseline Gaussian process regressor and observe that the SCGN provides comparable uncertainty descriptions with improved scalability to dataset size. The sample functions generated by SCGN are fast to evaluate online and are therefore convenient for use in trajectory simulations. These results suggest that SCGN may be a viable computational method for aerodynamic uncertainty quantification.

machine learning

Structured Covariance Gaussian Networks for Orion Crew Module Aerodynamic Uncertainty Quantification

In this paper we propose a new approach for nonlinear regression and uncertainty quantification. The method is based on a pair of neural networks which parameterize mean and dense covariance functions of a multivariate Gaussian process, trained together to maximize the log-likelihood of observing the given data. The covariance matrix is made positive definite at every input by construction. We also propose a sampling approach that produces viable surrogate function realizations from the Gaussian process. We call the proposed model a Structured Covariance Gaussian Network (SCGN). We illustrate the use of SCGNs for learning an aerodynamic response surface with built-in uncertainty for the Orion crew module. We find that SCGN provides an efficient and systematic way to learn nonlinear functional relationships and dense covariances. We compare results to a baseline Gaussian process regressor and observe that the SCGN provides comparable uncertainty descriptions with improved scalability to dataset size. The sample functions generated by SCGN are fast to evaluate online and are therefore convenient for use in trajectory simulations. These results suggest that SCGN may be a viable computational method for aerodynamic uncertainty quantification.

machine learning

A model of the human in a cognitive prediction task.

The human decision maker's behavior when predicting future states of discrete linear dynamic systems driven by zero-mean Gaussian processes is modeled. The task is on a slow enough time scale that physiological constraints are insignificant compared with cognitive limitations. The model is basically a linear regression system identifier with a limited memory and noisy observations. Experimental data are presented and compared to the model.

Rouse, W. B.

Ridge Regression Signal Processing

The introduction of the Global Positioning System (GPS) into the National Airspace System (NAS) necessitates the development of Receiver Autonomous Integrity Monitoring (RAIM) techniques. In order to guarantee a certain level of integrity, a thorough understanding of modern estimation techniques applied to navigational problems is required. The extended Kalman filter (EKF) is derived and analyzed under poor geometry conditions. It was found that the performance of the EKF is difficult to predict, since the EKF is designed for a Gaussian environment. A novel approach is implemented which incorporates ridge regression to explain the behavior of an EKF in the presence of dynamics under poor geometry conditions. The basic principles of ridge regression theory are presented, followed by the derivation of a linearized recursive ridge estimator. Computer simulations are performed to confirm the underlying theory and to provide a comparative analysis of the EKF and the recursive ridge estimator.

Kuhl, Mark R.

Estimating Dust and Water Ice Content of the Martian Atmosphere From THEMIS Data

Researchers at JPL and Arizona State University conducted a comparative study of three candidate algorithms for estimating components of the Martian atmosphere, using raw (uncalibrated) data collected by the Thermal Emission Imaging System (THEMIS). THEMIS is an instrument onboard the Mars Odyssey spacecraft that acquires image data in five visible and nine infrared (IR) wavelength bands. The algorithms under study used data collected from eight of the nine IR bands to estimate the dust and water ice content of the atmosphere. Such an algorithm could be used in onboard data processing to trigger other algorithms that search for features of scientific interest and to reduce the volume of data transmitted to Earth. The algorithms studied were based on regression models. In the study, the optical depths estimated by these algorithms were compared with optical depths estimated in ground-based processing using fully calibrated data from both THEMIS and the Thermal Emission Spectrometer (TES). TES is an instrument onboard the Mars Global Surveyor spacecraft that also observes the planet at infrared wavelengths, but at a lower spatial resolution than THEMIS does. Of the algorithms studied, the one that performed best was based on a Gaussian Support Vector Machine regression model. The test results indicated that this algorithm, operating on the raw data, had error rates that were within the uncertainty associated with the estimates obtained by the groundbased analysis of the fully calibrated data. This level of fidelity demonstrates that these algorithms are sufficiently accurate for use in an onboard setting.

Bandfield, Joshua

Linear feedback rate bounds for regressive channels

Bounds for the linear feedback capacity of m-th order Gaussian autoregressive channels are derived. The upper bound is tighter than that found by Tiernan and Schalwijk (1974) for the feedback capacity of a first-order autoregressive Gaussian channel with not necessarily linear processing. The separation between the upper and lower bounds is small, and it is conjectured that the lower bound converges to the feedback capacity of the first-order channel as the number of signals tends to infinity.

Butman, S. A.

Statistical Classification of Biosignature Information using Multiple Instrument Observations

The accurate identification of biosignatures (indications of life) from data taken from remote or in situ planetary exploration is one of the most important challenges in astrobiology, the interdisciplinary field examining habitability and the potential for extraterrestrial life. This study employs machine learning algorithms to optimize the identification of biosignatures, with an emphasis on those which are agnostic to a specific biochemical basis. We exploit the wealth of terrestrial data available from biogenic and abiogenic systems to enhance efficient feature prioritization. Our dataset, pulled from public databases and laboratory recorded measurements, includes elemental abundance, isotopic fractionation, and VNIR/Raman spectra The data curation process included standardization for detection limits and ranges. Subsequent feature extraction yielded detailed inputs for machine learning, including combinations of elemental content, isotopic ratios, and parameters of spectral peaks and troughs. Feature significance was evaluated across diverse machine learning methodologies, such as k-nearest neighbors, logistic regression, Random Forest, support vector machines, and Gaussian Naïve Bayes, along with a combined voting classifier. We utilized Receiver Operating Characteristic Area Under the Curve (ROC AUC) across 2,000 50% test-train splits as a robust metric of model performance. Results revealed a promising ROC AUC of 0.853 for the combined voting classifier. Removing elemental abundance data notably reduced model accuracy (13% decrease in AUC), highlighting its critical role in biosignature detection. Several other individual data features exhibited significance within their respective data types, offering additional granularity. This research fortifies the relevance of machine learning to astrobiology, potentially enhancing life detection missions by allowing algorithmic prioritization of high-interest samples for further investigation. Future work will refine data standardization, expand the dataset to include more terrestrial systems, and incorporate convolutional neural networks for spectral feature extraction. The potential for public data sharing is also under exploration, reinforcing our commitment to collective scientific advancement.

Statistical

NASA Tech Briefs, March 2004

Topics covered include: 1) Advanced Signal Conditioners for Data-Acquisition Systems; 2) Downlink Data Multiplexer; 3) Viewing ISS Data in Real Time via the Internet; 4) Autonomous Environment-Monitoring Networks; 5) Readout of DSN Monitor Data; 6) Parallel-Processing Equalizers for Multi-Gbps Communications; 7) AIN-Based Packaging for SiC High-Temperature Electronics; 8) Software for Optimizing Quality Assurance of Other Software; 9) The TechSat 21 Autonomous Sciencecraft Experiment; 10) Software for Analyzing Laminar-to-Turbulent Flow Transitions; 11) Elastomer Filled With Single-Wall Carbon Nanotubes; 12) Modifying Ship Air-Wake Vortices for Aircraft Operations; 13) Strain-Gauge Measurement of Weight of Fluid in a Tank; 14) Advanced Docking System With Magnetic Initial Capture; 15) Blade-Pitch Control for Quieting Tilt-Rotor Aircraft; 16) Solar Array Panels With Dust-Removal Capability; 17) Aligning Arrays of Lenses and Single-Mode Optical Fibers; 18) Automatic Control of Arc Process for Making Carbon Nanotubes; 19) Curved-Focal-Plane Arrays Using Deformed-Membrane Photodetectors; 20) Role of Meteorology in Flights of a Solar-Powered Airplane; 21) Model of Mixing Layer With Multicomponent Evaporating Drops; 22) Solution-Assisted Optical Contacting; 23) Improved Discrete Approximation of Laplacian of Gaussian; 24) Utilizing Expert Knowledge in Estimating Future STS Costs; 25) Study of Rapid-Regression Liquefying Hybrid Rocket Fuels; and 26) More About the Phase-Synchronized Enhancement Method.

Source record

Inverse sequential procedures for the monitoring of time series

When one or more new values are added to a developing time series, they change its descriptive parameters (mean, variance, trend, coherence). A 'change index (CI)' is developed as a quantitative indicator that the changed parameters remain compatible with the existing 'base' data. CI formulate are derived, in terms of normalized likelihood ratios, for small samples from Poisson, Gaussian, and Chi-Square distributions, and for regression coefficients measuring linear or exponential trends. A substantial parameter change creates a rapid or abrupt CI decrease which persists when the length of the bases is changed. Except for a special Gaussian case, the CI has no simple explicit regions for tests of hypotheses. However, its design ensures that the series sampled need not conform strictly to the distribution form assumed for the parameter estimates. The use of the CI is illustrated with both constructed and observed data samples, processed with a Fortran code 'Sequitor'.

Radok, Uwe

Using Statistical Multivariable Models to Understand the Relationship Between Interplanetary Coronal Mass Ejecta and Magnetic Flux Ropes

In-situ measurements of interplanetary coronal mass ejections (ICMEs) display a wide range of properties. A distinct subset, "magnetic clouds" (MCs), are readily identifiable by a smooth rotation in an enhanced magnetic field, together with an unusually low solar wind proton temperature. In this study, we analyze Ulysses spacecraft measurements to systematically investigate five possible explanations for why some ICMEs are observed to be MCs and others are not: i) An observational selection effect; that is, all ICMEs do in fact contain MCs, but the trajectory of the spacecraft through the ICME determines whether the MC is actually encountered; ii) interactions of an erupting flux rope (PR) with itself or between neighboring FRs, which produce complex structures in which the coherent magnetic structure has been destroyed; iii) an evolutionary process, such as relaxation to a low plasma-beta state that leads to the formation of an MC; iv) the existence of two (or more) intrinsic initiation mechanisms, some of which produce MCs and some that do not; or v) MCs are just an easily identifiable limit in an otherwise corntinuous spectrum of structures. We apply quantitative statistical models to assess these ideas. In particular, we use the Akaike information criterion (AIC) to rank the candidate models and a Gaussian mixture model (GMM) to uncover any intrinsic clustering of the data. Using a logistic regression, we find that plasma-beta, CME width, and the ratio O(sup 7) / O(sup 6) are the most significant predictor variables for the presence of an MC. Moreover, the propensity for an event to be identified as an MC decreases with heliocentric distance. These results tend to refute ideas ii) and iii). GMM clustering analysis further identifies three distinct groups of ICMEs; two of which match (at the 86% level) with events independently identified as MCs, and a third that matches with non-MCs (68 % overlap), Thus, idea v) is not supported. Choosing between ideas i) and iv) is more challenging, since they may effectively be indistinguishable from one another by a single in-situ spacecraft. We offer some suggestions on how future studies may address this.

Riley, P.

Classifying Agnostic Biosignatures using Raman, VNIR, and Elemental Data

How can we use our current wealth of terrestrial data, encompassing biogenic and abiogenic systems, to determine the distinguishing properties of life? SCOBI (Statistical Classification of Biosignature Information) uses machine learning techniques to algorithmically identify combinations of measurements that are “indicative of life”. A set of ~1000 observations, comprising elemental abundance, isotopic fractionation, VNIR reflectance, and (in progress) Raman spectra, have been assembled from existing literature and databases. The observations cover systems classified as “indicative alive” (e.g., cells, vegetation), “indicative non-alive” (e.g., fossils, teeth), “mixed indicative” (e.g., soil, pond water), or “non-indicative” (e.g., rocks, meteorites). VNIR data was preprocessed by linear interpolation from 400-2100 nm and smoothed with a Savitzky-Golay filter. To limit the amount of Earth-biochemistry-specific (non-agnostic) information included, the first five spectral features extracted were number of peaks, number of troughs, mean reflectance, mean peak width, and broadest peak width. To help further emphasize agnostic biosignatures, Earth-specific features such as chlorophylls have been manually flagged so that feature importance with and without them can be compared. Classifiers including k-nearest neighbors (KNN), Gaussian Naïve Bayes (GNB), logistic regression (LR), random forest (RF), and support vector machine (SVM) were implemented, as was a combination voting classifier. Performance metrics included false positive rates, false negative rates, and AUC with 50-50 test/train splits (Monte Carlo simulations). Key takeaways from this stage, prior to the inclusion of Raman spectra, are (1) the overall success rate of 0.933 AUC was most heavily influenced by the elemental abundance data; and (2) VNIR reflectance had the lowest classification performance with 0.52 AUC (58% of objects correctly classified). The next steps are to complete integration of Raman spectral data and to improve the approach to pre-processing and feature extraction for both types of spectral data, such as automated baseline removal, whole spectrum matching, and dimensionality reduction.

Biosignatures