Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Multivariate analysis: An essential for studying complex glasses

Understanding the impact of individual compositional components on the devitrification of complex multicomponent glasses, for example, 10–50+ oxides, typically requires numerous studies to examine each component's impact. Here we apply exploratory data analysis (EDA) to a heterogeneous data set of silicate glasses to determine the cations’ individual and interacting effects on the crystallization of nepheline (nominally NaAlSiO 4 ). Our data consisted of 795 simulated high-level nuclear waste glasses composed of, on average, 50 oxide components. We determine the interactions in the heterogeneous data that cause deviations from the behavior found in simplified composition studies. Using both univariate and bivariate EDA techniques, we demonstrate the importance of including calculated structural glass parameters on nepheline's devitrification, including field strength, cation-to-anion radius ratio, and single-bond strength. Here, we also show that studies with simplified glass compositions may fall short in generating knowledge directly transferrable to complex glass compositions. The method used in this study has the potential to inform experimental design for simplified compositions (~6+ oxides) that can generate knowledge directly transferrable to complex, multivariable compositions. The observations reported here have broad implications for any study attempting to map the physical properties of a complex glass containing numerous cations.

36 MATERIALS SCIENCE↗

A Statistical Interpolation Code for Ocean Analysis and Forecasting

Abstract We present a data assimilation package for use with ocean circulation models in analysis, forecasting, and system evaluation applications. The basic functionality of the package is centered on a multivariate linear statistical estimation for a given predicted/background ocean state, observations, and error statistics. Novel features of the package include support for multiple covariance models, and the solution of the least squares normal equations either using the covariance matrix or its inverse—the information matrix. The main focus of this paper, however, is on the solution of the analysis equations using the information matrix, which offers several advantages for solving large problems efficiently. Details of the parameterization of the inverse covariance using Markov random fields are provided and its relationship to finite-difference discretizations of diffusion equations are pointed out. The package can assimilate a variety of observation types from both remote sensing and in situ platforms. The performance of the data assimilation methodology implemented in the package is demonstrated with a yearlong global ocean hindcast with a 1/4° ocean model. The code is implemented in modern Fortran, supports distributed memory, shared memory, multicore architectures, and uses climate and forecasts compliant Network Common Data Form for input/output. The package is freely available with an open source license from www.tendral.com/tsis/ .

Srinivasan, Ashwanth↗

Atmospheric condition identification in multivariate data through a metric for total variation

Identification of atmospheric conditions within a multivariable atmospheric data set is a necessary step in the validation of emerging and existing high-fidelity models used to simulate wind plant flows and operation.Atmospheric conditions relevant for wind energy research include stationary conditions, given the need for well-converged statistics for model validation, as well as conditions observed less frequently, such as extreme atmospheric events, which are used in wind turbine and wind plant design.Aggregation of observations without regard to covariance between time series discounts the dynamical nature of the atmosphere and is not sufficiently representative of atmospheric conditions.Identification and characterization of continuous time periods with atmospheric conditions that have a high value for analysis or simulation set the stage for more advanced model validation and the development of real-time control and operational strategies.The current work explores a single metric for variation in a multivariate data sample that quantifies variability within each channel as well as covariance between channels.The total variation is used to identify conditions of interest that conform to desired objective functions, such as stationary conditions, ramps or waves of wind speed, and changes in wind direction.Total variation is somewhat sensitive to the presence of outliers in the input data, and the method is best complemented by quality-control procedures to ensure reliable results.The direct detection and classification of events or conditions of interest within atmospheric data sets is vital to developing our understanding of wind plant response and to the formulation of forecasting and control models.

17 WIND ENERGY↗

The Tropical Convective Spectrum: Archetypal Vertical Structures - Part 1

A taxonomy of tropical convective and stratiform vertical structures is constructed through cluster analysis of 3 yr of Tropical Rainfall Measuring Mission (TRMM) "warm-season" (surface temperature greater than 10 C) precipitation radar (PR) vertical profiles, their surface rainfall, and associated radar-based classifiers (convective/ stratiform and brightband existence). Twenty-five archetypal profile types are identified, including nine convective types, eight stratiform types, two mixed types, and six anvil/fragment types (nonprecipitating anvils and sheared deep convective profiles). These profile types are then hierarchically clustered into 10 similar families, which can be further combined, providing an objective and physical reduction of the highly multivariate PR data space that retains vertical structure information. The taxonomy allows for description of any storm or local convective spectrum by the profile types or families. The analysis provides a quasi-independent corroboration of the TRMM 2A23 convective/ stratiform classification. The global frequency of occurrence and contribution to rainfall for the profile types are presented, demonstrating primary rainfall contribution by midlevel glaciated convection (27%) and similar depth decaying/stratiform stages (28%-31%). Profiles of these types exhibit similar 37- and 85-GHz passive microwave brightness temperatures but differ greatly in their frequency of occurrence and mean rain rates, underscoring the importance to passive microwave rain retrieval of convective/stratiform discrimination by other means, such as polarization or texture techniques, or incorporation of lightning observations. Close correspondence is found between deep convective profile frequency and annualized lightning production, and pixel-level lightning occurrence likelihood directly tracks the estimated mean ice water path within profile types.

Precipitation radar↗

Recent advancements in information extraction methodology and hardware for earth resources survey systems

The present work discusses some recent developments in preprocessing and extractive processing techniques and hardware and in user applications model development for earth resources survey systems. The Multivariate Interactive Digital Analysis System (MIDAS) is currently being developed, and is an attempt to solve the problem of real time multispectral data processing in an operational system. The main features and design philosophy of this system are described. Examples of wetlands mapping and land resource inventory are presented. A user model developed for predicting the yearly production of mallard ducks from remote sensing and ancillary data is described.

Erickson, J. D.↗

Parametric Analysis of a Hover Test Vehicle using Advanced Test Generation and Data Analysis

Large complex aerospace systems are generally validated in regions local to anticipated operating points rather than through characterization of the entire feasible operational envelope of the system. This is due to the large parameter space, and complex, highly coupled nonlinear nature of the different systems that contribute to the performance of the aerospace system. We have addressed the factors deterring such an analysis by applying a combination of technologies to the area of flight envelop assessment. We utilize n-factor (2,3) combinatorial parameter variations to limit the number of cases, but still explore important interactions in the parameter space in a systematic fashion. The data generated is automatically analyzed through a combination of unsupervised learning using a Bayesian multivariate clustering technique (AutoBayes) and supervised learning of critical parameter ranges using the machine-learning tool TAR3, a treatment learner. Covariance analysis with scatter plots and likelihood contours are used to visualize correlations between simulation parameters and simulation results, a task that requires tool support, especially for large and complex models. We present results of simulation experiments for a cold-gas-powered hover test vehicle.

Gundy-Burlet, Karen↗

Astronomical data analysis software and systems I; Proceedings of the 1st Annual Conference, Tucson, AZ, Nov. 6-8, 1991

Consideration is given to a definition of a distribution format for X-ray data, the Einstein on-line system, the NASA/IPAC extragalactic database, COBE astronomical databases, Cosmic Background Explorer astronomical databases, the ADAM software environment, the Groningen Image Processing System, search for a common data model for astronomical data analysis systems, deconvolution for real and synthetic apertures, pitfalls in image reconstruction, a direct method for spectral and image restoration, and a discription of a Poisson imagery super resolution algorithm. Also discussed are multivariate statistics on HI and IRAS images, a faint object classification using neural networks, a matched filter for improving SNR of radio maps, automated aperture photometry of CCD images, interactive graphics interpreter, the ROSAT extreme ultra-violet sky survey, a quantitative study of optimal extraction, an automated analysis of spectra, applications of synthetic photometry, an algorithm for extra-solar planet system detection and data reduction facilities for the William Herschel telescope.

Worrall, Diana M.↗

Search for ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production in the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state using proton-proton collisions at $\sqrt{s}=13\,\text {Te}\hspace{-.08em}\text {V} $

A search for ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production in the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state is presented, where H is the standard model (SM) Higgs boson. The search uses an event sample of proton-proton collisions corresponding to an integrated luminosity of 133$\,\text {fb}^{-1}$ collected at a center-of-mass energy of 13$\,\text {Te}\hspace{-.08em}\text {V}$ with the CMS detector at the CERN LHC. The analysis introduces several novel techniques for deriving and validating a multi-dimensional background model based on control samples in data. A multiclass multivariate classifier customized for the ${\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} {\text {b}{}{}} {\bar{{\text {b}{}{}}}{}{}} $ final state is developed to derive the background model and extract the signal. The data are found to be consistent, within uncertainties, with the SM predictions. The observed (expected) upper limits at 95% confidence level are found to be 3.8 (3.8) and 5.0 (2.9) times the SM prediction for the ${\text {Z}{}{}} {\text {Z}{}{}} $ and ${\text {Z}{}{}} {\text {H}{}{}} $ production cross sections, respectively.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Evaluating the potential of short-term instrument deployment to improve distributed wind resource assessment

Distributed wind projects, which are connected at the distribution level of an electricity system or in off-grid applications to serve specific or local energy needs, often rely solely on wind resource models to establish wind speed and energy generation expectations. Historically, anemometer loan programs have provided an affordable avenue for more accurate onsite wind resource assessment, and the lowering cost of lidar systems has shown similar advantages for more recent assessments. While a full 12 months of onsite wind measurement is the standard for correcting model-based long-term wind speed estimates for utility-scale wind farms, the time and capital investment involved in gathering onsite measurements must be reconciled with the energy needs and funding opportunities that drive expedient deployment of distributed wind projects. Much literature exists to quantify the performance of correcting long-term wind speed estimates with 1 or more years of observational data, but few studies explore the impacts of correcting with months-long observational periods. This study aims to answer the question of how short you can go in terms of the observational time period needed to make impactful improvements to model-based long-term wind speed estimates. Three algorithms, multivariable linear regression, adaptive regression splines, and regression trees, are evaluated for their skill at correcting long-term wind resource estimates from the European Centre for Medium-Range Weather Forecasts Reanalysis version 5 (ERA5) using months-long periods of observational data from 66 locations across the US. On average, correction with even 1 month of observations provides significant improvement over the baseline ERA5 wind speed estimates and produces median bias magnitudes and relative errors within 0.22 m s −1 and 4 percentage points of the median bias magnitudes and relative errors achieved using the standard 12 months of data for correction. However, in cases when the shortest observational periods (1 to 2 months) used for correction are not well correlated with the overlapping ERA5 reference, the resultant long-term wind speed errors are worse than those produced using ERA5 without correction. Summer months, which are characterized by weaker relative wind speeds and standard deviations for most of the evaluation sites, tend to produce the worst results for long-term correction using months-long observations. The three tested algorithms perform similarly for long-term wind speed bias; however, regression trees perform notably worse than multivariable linear regression and adaptive regression splines in terms of correlation when using 6 months or less of observational data for correction. Translating the analysis to wind energy, median relative errors in the capacity factor are on average within 10 % using 1 month of training. If the observation period used for correction is not well correlated with the reference data, however, misrepresentation of the observed capacity factor can be substantial. The risk associated with poor correlation between the observed and reference datasets decreases with increasing training period length. In the worst-correlation scenarios, the median capacity factor relative errors from using 1, 3, and 6 months are within 47 %, 26 %, and 16 %, respectively.

17 WIND ENERGY↗

MSW Variability Mapping and Conversion to Biofuel

MSW (Municipal Solid Waste) is a form of biomass which consists of categorized components of waste/trash. The general categories are paper, yard trash, construction & debris, appliances, tires, glass, metals, aluminum & steel cans, plastics, organics, inorganics, and HHW (Household Hazardous Waste). This project focuses on the factors within a region or population that contribute to variability in the composition of MSW and in turn MSW’s convertibility to biofuel. A list of contributors was determined (Social Vulnerability Index, Access to Public Transportation, Racial Distribution, GDP, Personal Income) and then JMP was used to perform a Multivariate analysis to determine correlations and a Partial Least-Squares regression to determine Variable Importance Plots for each MSW category. In addition to data analysis, the convertibility of MSW to biofuel was studied via microwave pyrolysis system in order to separate and characterize the various gaseous and bio-oil products.

09 BIOMASS FUELS↗

Switchgrass sward establishment selection is consistent across multiple environments and fertilization levels

Strong selection can occur during switchgrass sward establishment. Differences in establishment selection due to environment or management could provide information on genotype-by-environment variation and could influence strategies for breeding perennial grasses. Leaf samples were collected before sward establishment and from 3-year-old swards for two breeding groups (lowland and hybrid) at three locations. Within two locations, samples were collected from paired fertilized (112 kg N ha –1 ) and unfertilized plots. Allele frequencies from pooled DNA samples were studied through multivariate analysis of variance, genomewide trait predictions (heading date and winter survivorship), and genomically estimated breeding values (GEBVs) for individual sward survival within an independent data set. This study found only minor variations in selection due to location or management. Predicted heading dates of the hybrid population had significant changes due to fertilization and location. There were strong correlations among sward establishment survival GEBVs between growing environments (hybrid r = 0.77; gulf r = 0.97). Interestingly, this study found a small number of genotypes that were over-represented in established swards across all growing environments. This study reinforces a prior report of selection during sward establishment and indicates that only a small degree of establishment selection is location-specific within these diverse growing conditions.

59 BASIC BIOLOGICAL SCIENCES↗

Spatio-temporal multivariate cluster evolution analysis for detecting and tracking climate impacts

Recent years have seen a growing concern about climate change and its impacts. While Earth System Models (ESMs) can be invaluable tools for studying the impacts of climate change, the complex coupling processes encoded in ESMs and the large amounts of data produced by these models, together with the high internal variability of the Earth system, can obscure important source-to-impact relationships. Here, this paper presents a novel and efficient unsupervised data-driven approach for detecting statistically-significant impacts and tracing spatio-temporal source-impact pathways in the climate through a unique combination of ideas from anomaly detection, clustering and Natural Language Processing (NLP). Using as an exemplar the 1991 eruption of Mount Pinatubo in the Philippines, we demonstrate that the proposed approach is capable of detecting known post-eruption impacts/events. We additionally describe a methodology for extracting meaningful sequences of post-eruption impacts/events by using NLP to efficiently mine frequent multivariate cluster evolutions, which can be used to confirm or discover the chain of physical processes between a climate source and its impact(s).

Anomaly detection↗

An Update on Mortality in the U.S. Astronaut Corps: 1959-2009

Although it has now been over 50 years since mankind first ventured into space, the long-term health impacts of human space flight remain largely unknown. Identifying factors that affect survival and prognosis among those who participate in space flight is vitally important, as the era of commercial space flight approaches and NASA prepares for missions to Mars. The Longitudinal Study of Astronaut Health is a prospective study designed to examine trends in astronaut morbidity and mortality. The purpose of this analysis was to describe and explore predictors of overall and cause-specific mortality among individuals selected for the U.S. astronaut corps. All U.S. astronauts (n=321), regardless of flight status, were included in this analysis. Death certificate searches were conducted to ascertain vital status and cause of death through April 2009. Data were collected from medical records and lifestyle questionnaires. Multivariable Cox regression modeling was used to calculate the mortality hazard associated with embarking on space flight, adjusted for sex, race, and age at selection. Between 1959 and 2009, there were 39 (12.1%) deaths. Of these deaths, 18 (42.2%) were due to occupational accidents; 7 (17.9%) were due to other accidents; 6 (15.4%) were attributable to cancer; 6 (15.4%) resulted from cardiovascular/circulatory diseases; and 2 (5.1%) were from other causes. Participation in space flight did not significantly increase mortality hazard over time (adjusted hazard ratio=0.57; 95% confidence interval=0.26-1.26. Because our results are based on a small sample size, future research that includes payload specialists, other space flight participants, and international crew members is warranted to maximize statistical power.

Amirian, E.↗

Visualization for Insight and Data Analysis in Energy Research

This talk explores how advanced visualization technologies are transforming analytical reasoning and knowledge discovery in energy research, drawing on recent work at the National Laboratory of the Rockies' Computational Science Center. Through a series of scientific case studies, we demonstrate how immersive and high-resolution visualization environments enable scientists and engineers to identify previously unseen patterns and features - insights that often remain hidden in traditional desktop-based analysis. By embedding richer information into interactive analytics tools, these approaches support the exploration of complex, multivariate parameter spaces, where interaction itself catalyzes understanding. Beyond capability, we emphasize the critical role of visualization design grounded in perception and cognition, showing how visual encodings directly influence analytical outcomes. Spanning applications from materials science to integrated energy systems, these visualization approaches accelerate innovation and improve decision-making by enabling deeper, more reliable insight into increasingly complex energy data.

97 MATHEMATICS AND COMPUTING↗

Predicting major element mineral/melt equilibria - A statistical approach

Empirical equations have been developed for calculating the mole fractions of NaO0.5, MgO, AlO1.5, SiO2, KO0.5, CaO, TiO2, and FeO in a solid phase of initially unknown identity given only the composition of the coexisting silicate melt. The approach involves a linear multivariate regression analysis in which solid composition is expressed as a Taylor series expansion of the liquid compositions. An internally consistent precision of approximately 0.94 is obtained, that is, the nature of the liquidus phase in the input data set can be correctly predicted for approximately 94% of the entries. The composition of the liquidus phase may be calculated to better than 5 mol % absolute. An important feature of this 'generalized solid' model is its reversibility; that is, the dependent and independent variables in the linear multivariate regression may be inverted to permit prediction of the composition of a silicate liquid produced by equilibrium partial melting of a polymineralic source assemblage.

Hostetler, C. J.↗

Search for $t\overline{t}$ resonances in fully hadronic final states in $pp$ collisions at $ \sqrt{s} $ = 13 TeV with the ATLAS detector

This paper presents a search for new heavy particles decaying into a pair of top quarks using 139 fb -1 of proton-proton collision data recorded at a centre-of-mass energy of √s = 13 TeV with the ATLAS detector at the Large Hadron Collider. The search is performed using events consistent with pair production of high-transverse-momentum top quarks and their subsequent decays into the fully hadronic final states. The analysis is optimized for resonances decaying into a $t\overline{t}$ pair with mass above 1.4 TeV, exploiting a dedicated multivariate technique with jet substructure to identify hadronically decaying top quarks using large-radius jets and evaluating the background expectation from data. No significant deviation from the background prediction is observed. Limits are set on the production cross-section times branching fraction for the new Z' boson in a topcolor-assisted-technicolor model. The Z' boson masses below 3.9 and 4.7 TeV are excluded at 95% confidence level for the decay widths of 1% and 3%, respectively.

Heavy quark production↗

Interdependence in active mobility adoption: Joint modeling and motivational spillover in walking, cycling and bike-sharing

Active mobility offers an array of physical, emotional, and social well-being benefits. However, with the proliferation of the sharing economy, new nonmotorized means of transport are entering the fold, complementing some existing mobility options while competing with others. The purpose of this research study is to investigate the adoption of three active travel modes—namely walking, cycling, and bikesharing—in a joint modeling framework. Here, the analysis is based on an adaptation of the stages of change framework, which originates from the health behavior sciences. Multivariate ordered probit modeling drawing on U.S. survey data provides well-needed insights into individuals’ preparedness to adopt multiple active modes as a function of personal, neighborhood, and psychosocial factors. The research suggests three important findings. (1) The joint model structure confirms interdependence among different active mobility choices. The strongest complementarity is found for walking and cycling adoption. (2) Each mode has a distinctive adoption path with either three or four separate stages. We discuss the implications of derived stage-thresholds and plot adoption contours for selected scenarios. (3) Psychological and neighborhood variables generate more coupling among active modes than individual and household factors. Specifically, identifying strongly with active mobility aspirations, experiences with multimodal travel, possessing better navigational skills, along with supportive local community norms are the factors that appear to drive the joint adoption decisions. This study contributes to the understanding of how decisions within the same functional domain are related and help to design policies that promote active mobility by identifying positive spillovers and joint determinants.

42 ENGINEERING↗

Missing-Data Nonparametric Coherency Estimation

Chave recently proposed an estimator for multitaper spectral density where the time series contains missing values. In this article we generalize this technique to a multitaper estimator of coherence and phase and show that one can also obtain bootstrapped confidence intervals. Additionally, we give two examples. The first is a toy example in which the true coherence is known. In the second example we show that the multitaper missing-data coherence estimator computed on real data with a single gap comprising 11% of the data outperforms the Daniell-smoothed coherence estimator where there are no gaps. The case where the two time series have different missing indices is also discussed.

42 ENGINEERING↗