Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

The role of multispectral scanners as data sources for EPA hydrologic models

An estimated cost savings of 30% to 50% was realized from using LANDSAT-derived data as input into a program which simulates hydrologic and water quality processes in natural and man-made water systems. Data from the satellite were used in conjunction with EPA's 11-channel multispectral scanner to obtain maps for characterizing the distribution of turbidity plumes in Flathead Lake and to predict the effect of increasing urbanization in Montana's Flathead River Basin on the lake's trophic state. Multispectral data are also being studied as a possible source of the parameters needed to model the buffering capability of lakes in an effort to evaluate the effect of acid rain in the Adirondacks. Water quality in Lake Champlain, Vermont is being classified using data from the LANDSAT and the EPA MSS. Both contact-sensed and MSS data are being used with multivariate statistical analysis to classify the trophic status of 145 lakes in Illinois and to identify water sampling sites in Appalachicola Bay where contaminants threaten Florida's shellfish.

Slack, R.↗

Modeling and control for closed environment plant production systems

A computer program was developed to study multiple crop production and control in controlled environment plant production systems. The program simulates crop growth and development under nominal and off-nominal environments. Time-series crop models for wheat (Triticum aestivum), soybean (Glycine max), and white potato (Solanum tuberosum) are integrated with a model-based predictive controller. The controller evaluates and compensates for effects of environmental disturbances on crop production scheduling. The crop models consist of a set of nonlinear polynomial equations, six for each crop, developed using multivariate polynomial regression (MPR). Simulated data from DSSAT crop models, previously modified for crop production in controlled environments with hydroponics under elevated atmospheric carbon dioxide concentration, were used for the MPR fitting. The model-based predictive controller adjusts light intensity, air temperature, and carbon dioxide concentration set points in response to environmental perturbations. Control signals are determined from minimization of a cost function, which is based on the weighted control effort and squared-error between the system response and desired reference signal.

NASA Discipline Life Support Systems↗

Visual Analytics of Multivariate Networks With Representation Learning and Composite Variable Construction

Multivariate networks are commonly found in real-world data-driven applications. Uncovering and understanding the relations of interest in multivariate networks is not a trivial task. This article presents a visual analytics workflow for studying multivariate networks to extract associations between different structural and semantic characteristics of the networks (e.g., what are the combinations of attributes largely relating to the density of a social network?). The workflow consists of a neural-network-based learning phase to classify the data based on the chosen input and output attributes, a dimensionality reduction and optimization phase to produce a simplified set of results for examination, and finally an interpreting phase conducted by the user through an interactive visualization interface. A key part of our design is a composite variable construction step that remodels nonlinear features obtained by neural networks into linear features that are intuitive to interpret. We demonstrate the capabilities of this workflow with multiple case studies on networks derived from social media usage and also evaluate the workflow with qualitative feedback from experts.

97 MATHEMATICS AND COMPUTING↗

Vector-Ordering Filter Procedure for Data Reduction

The vector-ordering filter (VOF) technique involves a procedure for sampling a large population of data vectors to select a subset of data vectors that fully characterize the state space of the large population. The VOF technique enables a large reduction of the volume of data that must be handled in the automated monitoring system and method discussed in the two immediately preceding articles. In so doing, the VOF technique enables the development of data-driven mathematical models of a monitored asset from sets of data that would otherwise exceed the memory capacities of conventional engineering computers. Data-driven mathematical models have been shown to offer high fidelity for purposes of control and monitoring of assets. In practice, a collection of asset-operating observations is acquired with the intention that the collection contain observations characteristic of the full dynamic range of operation of the asset. Often, such a collection contains an extremely large number of observations, many of which are redundant. The VOF technique fills the need for a means to extract, from the original collection of observational data, a reduced data matrix that excludes redundant data while maintaining the full statistical character and dynamic range of the original data. The reduced data matrix can then be used as the input data for development of a mathematical model of the monitored asset, or as training data for a neural-network substitute for an explicit mathematical model of the asset. Alternatively, the reduced data matrix can, itself, be used directly as a mathematical model of the monitored asset, as is commonly done in multivariate state-estimation techniques. The original data are collected from the asset over a range of operating states and are put in matrix form. Each column vector in the original data matrix represents the signal values acquired at a particular operational state of the asset. Thus, the number of columns of the original data matrix equals the number of observed states and the number of rows in this matrix equals the number of signals acquired at each observation. In the VOF technique, one extracts the reduced data matrix from the original data matrix through the selection of a representative subset of the column (state) vectors.

Bickford, Randall L.↗

Autonomous adaptive data acquisition for scanning hyperspectral imaging

Non-invasive and label-free spectral microscopy (spectromicroscopy) techniques can provide quantitative biochemical information complementary to genomic sequencing, transcriptomic profiling, and proteomic analyses. However, spectromicroscopy techniques generate high-dimensional data; acquisition of a single spectral image can range from tens of minutes to hours, depending on the desired spatial resolution and the image size. This substantially limits the timescales of observable transient biological processes. To address this challenge and move spectromicroscopy towards efficient real-time spatiochemical imaging, we developed a grid-less autonomous adaptive sampling method. Our method substantially decreases image acquisition time while increasing sampling density in regions of steeper physico-chemical gradients. When implemented with scanning Fourier Transform infrared spectromicroscopy experiments, this grid-less adaptive sampling approach outperformed standard uniform grid sampling in a two-component chemical model system and in a complex biological sample, Caenorhabditis elegans. We quantitatively and qualitatively assess the efficiency of data acquisition using performance metrics and multivariate infrared spectral analysis, respectively.

47 OTHER INSTRUMENTATION↗

A gradient model of vegetation and climate utilizing NOAA satellite imagery. Phase 1: Texas transect

A climatological model/variable termed the sponge (a measure of moisture availability based on daily temperature maxima and minima, and precipitation) was tested for potential biogeograhic, ecological, and agro-climatological applications. Results, depicted in tabular and graphic form, suggest that, as generalized climatic index, sponge is particularly appropriate for large-area and global vegetation monitoring. The feasibility of utilizing NOAA/AVHRR data for vegetation classification was investigated and a vegetation gradient model that utilizes sponge and AVHRR data was initiated. Along an east-west Texas gradient, vegetation, sponge, and AVHRR pixel data (channels 1 and 2) were obtained for 12 locations. The normalized difference values for the AVHRR data when plotted against vegetation characteristics (biomass, net productivity, leaf area) and sponge values along the Texas gradient suggest that a multivariate gradient model incorporating AVHRR and sponge data may indeed be useful in global vegetation stratification and monitoring.

Greegor, D.↗

Investigating Kinetic Mechanisms of Soot Formation in Plasma Pyrolysis of Methane via Active Learning (Final Technical Report)

Plasma pyrolysis of methane is an effective route for zero-carbon hydrogen production. Yet, soot generated from pyrolysis of hydrocarbons is detrimental to the climate and human health. There is ample experimental and theoretical evidence that suggests polycyclic aromatic hydrocarbons (PAHs) are the molecular precursors to soot particles. The reaction pathways of PAH formation are intricately dependent on a multitude of process parameters, whose kinetic mechanisms are not well-understood in plasma pyrolysis. This project aims to leverage advances in the kinetic modeling of soot formation in combustion, as well as in surrogate modeling and active learning, to systematically investigate the effects of process parameter on the kinetics of PAH formation in plasma pyrolysis of methane. To this end, we propose to use the PAH formation kinetics model developed by the PPPL/PU group based on the well-established ABF and HACA mechanisms, coupled with low-temperature plasma models. We will develop an active learning (AL) framework based on Bayesian optimization to systematically and data-efficiently explore the complex and multivariable parameter space of plasma pyrolysis in order to quantify the effects of plasma and feed parameters on the ABF and HACA kinetic pathways. AL is the branch of machine learning concerned with systematically querying samples from a system (experimental or computational) to train a data-driven model that maps design parameters to a performance criterion. We will use the data generated via AL to perform global sensitivity analysis, combined with uncertainty quantification, to elucidate the impact of different reaction pathways on minimizing formation of soot precursors. This study will result in an improved understanding of kinetics of PAH formation in plasma pyrolysis and can pave the way for more advanced mechanistic studies (e.g., soot nucleation mechanisms). Additionally, the findings will be useful for establishing practical strategies for increasing the pyrolysis efficiency and producing high-grade carbon for synthesis of nanomaterials.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Multivariate normality

Sets of experimentally determined or routinely observed data provide information about the past, present and, hopefully, future sets of similarly produced data. An infinite set of statistical models exists which may be used to describe the data sets. The normal distribution is one model. If it serves at all, it serves well. If a data set, or a transformation of the set, representative of a larger population can be described by the normal distribution, then valid statistical inferences can be drawn. There are several tests which may be applied to a data set to determine whether the univariate normal model adequately describes the set. The chi-square test based on Pearson's work in the late nineteenth and early twentieth centuries is often used. Like all tests, it has some weaknesses which are discussed in elementary texts. Extension of the chi-square test to the multivariate normal model is provided. Tables and graphs permit easier application of the test in the higher dimensions. Several examples, using recorded data, illustrate the procedures. Tests of maximum absolute differences, mean sum of squares of residuals, runs and changes of sign are included in these tests. Dimensions one through five with selected sample sizes 11 to 101 are used to illustrate the statistical tests developed.

Crutcher, H. L.↗

Synthesis of rotor test data for real-time simulation

A mathematical model of a hingeless tilting rotor is presented. The model was obtained by a systematic curve fit procedure applied to an extensive set of model scale wind tunnel data. The math model equations were used in a real time flight simulation model of a hingeless tilt rotor XV-15 to assess changes in flying qualities compared to those obtained using a previous rotor model. Extensive plots of the rotor derivatives are given. Discussions of attempts to apply multivariable linear regression technqiues to the data and the use of an analytical rotor representation are included.

Mcveigh, M. A.↗

M-DAS: System for multispectral data analysis

M-DAS is a ground data processing system designed for analysis of multispectral data. M-DAS operates on multispectral data from LANDSAT, S-192, M2S and other sources in CCT form. Interactive training by operator-investigators using a variable cursor on a color display was used to derive optimum processing coefficients and data on cluster separability. An advanced multivariate normal-maximum likelihood processing algorithm was used to produce output in various formats: color-coded film images, geometrically corrected map overlays, moving displays of scene sections, coverage tabulations and categorized CCTs. The analysis procedure for M-DAS involves three phases: (1) screening and training, (2) analysis of training data to compute performance predictions and processing coefficients, and (3) processing of multichannel input data into categorized results. Typical M-DAS applications involve iteration between each of these phases. A series of photographs of the M-DAS display are used to illustrate M-DAS operation.

Johnson, R. H.↗

Learning the temporal evolution of multivariate densities via normalizing flows

In this work, we propose a method to learn multivariate probability distributions using sample path data from stochastic differential equations. Specifically, we consider temporally evolving probability distributions (e.g., those produced by integrating local or nonlocal Fokker–Planck equations). Here, we analyze this evolution through machine learning assisted construction of a time-dependent mapping that takes a reference distribution (say, a Gaussian) to each and every instance of our evolving distribution. If the reference distribution is the initial condition of a Fokker–Planck equation, what we learn is the time-T map of the corresponding solution. Specifically, the learned map is a multivariate normalizing flow that deforms the support of the reference density to the support of each and every density snapshot in time. We demonstrate that this approach can approximate probability density function evolutions in time from observed sampled data for systems driven by both Brownian and Lévy noise. We present examples with two- and three-dimensional, uni- and multimodal distributions to validate the method.

97 MATHEMATICS AND COMPUTING↗

Unique signatures of synoptic features in Tiros N satellite data

The application of satellite data to the study of synoptic characteristics is analyzed. A radiative transfer model is used to generate a database of satellite information on synoptic features. Canonical discriminant analysis is employed to reveal the differences among synoptic sounding classes; and the fine structures within each sounding class are examined with rotated factor analysis. Diagrams of wind and frontal inversions are presented. It is noted that the applicability of satellite data depends on the method used to analyze it and multivariable statistical techniques may be useful for deriving additional information from satellite data.

White, G. A., III↗

Novel principal component analysis tool based on python for analysis of complex spectra of time-of-flight secondary ion mass spectrometry

Time-of-flight secondary ion mass spectrometry (ToF-SIMS) is a powerful surface analysis tool, which can simultaneously provide elemental, isotopic, and molecular information with part per million (ppm) sensitivity. However, each spectrum may be composed of hundreds of ion signals, which makes the spectra data complex. Principal component analysis (PCA) is a multivariate analysis technique that has been widely used to figure out the variances among samples in ToF-SIMS spectra data analysis and is showing great success in the explanation of complex ToF-SIMS spectra. So far, several software tools have been developed for PCA of ToF-SIMS spectra; however, none of them are freely available. Such a situation leads to some difficulties in extending applications of PCA to various research fields. More importantly, it has long been challenging for common researchers to understand PCA plots and extract chemical differences among samples. In this work, we developed a new and flexible software tool (named “advanced spectra pca toolbox”) based on python for PCA of complex ToF-SIMS spectra along with an easy-to-read manual. It can generate data analysis reports automatically to explain chemical differences among samples, allowing less experienced researchers to easily understand tricky PCA results. Moreover, it is expandable and compatible with artificial intelligence/machine learning functions. Pure goethite and different lignin adsorbed goethite samples were used as a model system to demonstrate our new software tool, proving that our software tool can be readily used in complex spectra data processing. Our new software tool is open-source, convenient, flexible, and expandable. We expect this open-source tool will benefit the ToF-SIMS community.

47 OTHER INSTRUMENTATION↗

Noninvasive assessment of mitral inertness: clinical results with numerical model validation

Inertial forces (Mdv/dt) are a significant component of transmitral flow, but cannot be measured with Doppler echo. We validated a method of estimating Mdv/dt. Ten patients had a dual sensor transmitral (TM) catheter placed during cardiac surgery. Doppler and 2D echo was performed while acquiring LA and LV pressures. Mdv/dt was determined from the Bernoulli equation using Doppler velocities and TM gradients. Results were compared with numerical modeling. TM gradients (range: 1.04-14.24 mmHg) consisted of 74.0 +/- 11.0% inertial forcers (range: 0.6-12.9 mmHg). Multivariate analysis predicted Mdv/dt = -4.171(S/D (RATIO)) + 0.063(LAvolume-max) + 5. Using this equation, a strong relationship was obtained for the clinical dataset (y=0.98x - 0.045, r=0.90) and the results of numerical modeling (y=0.96x - 0.16, r=0.84). TM gradients are mainly inertial and, as validated by modeling, can be estimated with echocardiography.

NASA Discipline Cardiopulmonary↗

A Review of Bayesian Networks for Spatial Data

We report Bayesian networks are a popular class of multivariate probabilistic models as they allow for the translation of prior beliefs about conditional dependencies between variables to be easily encoded into their model structure. Due to their widespread usage, they are often applied to spatial data for inferring properties of the systems under study and also generating predictions for how these systems may behave in the future. We review published research on methodologies for representing spatial data with Bayesian networks and also summarize the application areas for which Bayesian networks are employed in the modeling of spatial data. We find that a wide variety of perspectives are taken, including a GIS-centric focus on efficiently generating geospatial predictions, a statistical focus on rigorously constructing graphical models controlling for spatial correlation, as well as a range of problem-specific heuristics for mitigating the effects of spatial correlation and dependency arising in spatial data analysis. Special attention is also paid to potential future directions for integration of Bayesian networks with spatial processes.

97 MATHEMATICS AND COMPUTING↗

Comparison of Two Load Prediction Methods for Strain-Gage Balances

Data from a five-component semi-span balance is used to perform a systematic comparison of the load prediction accuracy of two load prediction methods. Both methods independently obtain the load prediction equations from multivariate least squares fits of balance calibration data. The first method is called the Non-Iterative Method. This approach directly uses regression models of the individual load components of a balance for the load prediction. The second method is called the Iterative Method. This alternate approach uses a load iteration equation for the load prediction that is constructed from the regression coefficients of the gage outputs of the balance. Basic characteristics of the two methods are reviewed. Afterwards, both methods are applied to calibration, check load, and wind tunnel test data of a five-component semi-span balance. Selected analysis results are compared. These comparisons confirm that the accuracy of the two methods is the same for all practical purposes.

strain-gage balance↗

Comparison of Two Load Prediction Methods for Strain-Gage Balances

Data from a high-capacity semi-span balance is used to perform a detailed comparison of the load prediction accuracy of two strain-gage balance load prediction methods. Both methods independently obtain their load prediction equations from multivariate least squares fits of balance calibration data. The first method is called Non-Iterative Method. This approach directly uses regression models of the individual load components of a balance for the load prediction. The second method is called Iterative Method. This alternate approach uses a load iteration equation for the load prediction that is constructed from the regression models of the gage outputs of the balance. Basic characteristics of the two methods are reviewed. Afterwards, both methods are applied to calibration and check load data of the chosen balance. Finally, selected analysis results are compared. These comparisons confirmed that the load prediction accuracy of the two methods is the same for all practical purposes.

wind tunnel test↗

Search for flavour-changing neutral current interactions of the top quark and the Higgs boson in events with a pair of $τ$-leptons in $pp$ collisions at $\sqrt{s}$ = 13 TeV with the ATLAS detector

A search for flavour-changing neutral current (FCNC) $tqH$ interactions involving a top quark, another up-type quark ($q = u, c$), and a Standard Model (SM) Higgs boson decaying into a $τ$-lepton pair ($H$ → $τ$ + $τ$ - ) is presented. The search is based on a dataset of pp collisions at $\sqrt{s}$ = 13 TeV that corresponds to an integrated luminosity of 139 fb -1 recorded with the ATLAS detector at the Large Hadron Collider. Two processes are considered: single top quark FCNC production in association with a Higgs boson ($pp$ → $tH$), and top quark pair production in which one of top quarks decays into $Wb$ and the other decays into $qH$ through the FCNC interactions. The search selects events with two hadronically decaying $τ$-lepton candidates ($τ$ had ) or at least one $τ$ had with an additional lepton ($e, µ$), as well as multiple jets. Event kinematics is used to separate signal from the background through a multivariate discriminant. A slight excess of data is observed with a significance of 2.3$σ$ above the expected SM background, and 95% CL upper limits on the $t → qH$ branching ratios are derived. The observed (expected) 95% CL upper limits set on the $t → cH$ and $t → uH$ branching ratios are 9.4 × 10 -4 (${4.8}^{+2.2}_{-1.4}$ × 10 -4 ) and 6.9 × 10 -4 (${3.5}^{+1.5}_{-1.0}$ × 10 -4 ), respectively. The corresponding combined observed (expected) upper limits on the dimension-6 operator Wilson coefficients in the effective $tqH$ couplings are $C$ $c$φ < 1.35 (0.97) and $C$ $u$φ < 1.16 (0.82).

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗