Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

A computer analysis of ERTS data of the Lake Gregory area of South Australia with particular emphasis on its role in terrain classification for engineering

A digital computer and multivariate statistical techniques were used to analyze 4-band multispectral data. A representation of the original data for each of the four bands allows a certain degree of terrain interpretation; however, variations in appearance of sites within and between bands, without additional criteria for deciding which representation should be preferred, create difficulties for classification. Investigation of the video data groups produced by principal components analysis and cluster analysis techniques shows that effective correlations with classifications of terrain produced by conventional methods could be carried out. The analyses also highlighted underlying relationships between the various elements. The approach used allows large areas (185 cm by 185 cm) to be classified into fundamental units within a matter of hours and can be applied to those parts of the Earth where facilities for conventional studies are poor or lacking.

Lodwick, G. D.↗

Updates to Relevance Vector Machine: Multiclass Classification, Variable Selection, and Proof-of-Concept Application to Safeguards Fresh Fuel Verification using List-Mode Neutron Collar Data

To expand the capabilities of safeguards authorities to verify the integrity of fresh fuel assemblies, Oak Ridge National Laboratory has retrofit the existing electronics of the JCC-71 uranium neutron coincidence collar, which contains 18 3 He neutron detectors and an external 241 AmLi(α, n) neutron interrogation source arranged to surround a fresh nuclear fuel assembly. The new electronics system allows analysts to record list-mode neutron multiplicity data in addition to the singles and doubles rates that are currently measured. Based on previous proof-of-concept research, analysis of these new data will identify off-normal fuel configurations in an assembly and characterize or localize the specific partial fuel defects. The purpose of this report it to document the analysis algorithm development and then to demonstrate its capability for the safeguards verification of fresh fuel assemblies using list mode neutron collar data. To analyze the complex list-mode data collected with the upgraded uranium neutron collar, multivariate classification algorithms are being developed using a novel classification method, the relevance vector machine. This approach may be applied to multiclass problems to estimate the probability that test data belongs to one of many possible classes of data. In addition, our method identifies the most useful variables/channels for making predictions, which illuminates the basis for the model’s predictions, and this interpretability is largely unique among data analytics methods. Variable selection occurs during model training and parameter tuning and does not need any external hyperparameter tuning routines. Finally, we apply the modified relevance vector machine to a simulated dataset of list-mode neutron collar data generated with the radiation transport code MCNP. The method can correctly identify off-normal fuel configurations, categorize the data according to four fuel defect scenarios, and rank the channels in the data according to prediction utility. For nuclear safeguards applications, it is concluded that this method has the potential to increase the sensitivity and reliability to detect missing fuel rods from a standard 17 x 17 Pressurized Water Reactor (PWR) fresh fuel assembly. Within this analysis, “off-normal” (i.e., missing fuel rods) were correctly classified in 17 simulated test scenarios with one quarter (25%) of the fresh fuel rods missing using a training data set of 58 simulated measurements.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Data-Driven Security Assessment of Power Grids Based on Machine Learning Approach: Preprint

Data-driven security assessment provides key indicators on power system stability using simulations on scheduling models, as opposed to dynamic simulations that are more time-consuming. This paper investigates data-driven security assessment of power grids based on machine learning. Multivariate random forest regression is used as the machine learning algorithm due to its high robustness to the input data. Three stability issues are analyzed using the proposed machine learning tool, including transient stability, frequency stability and small signal stability. The estimation values from machine learning tool are compared with those from dynamic simulations. Results show that the proposed machine learning tool can effectively predict the stability margins for the three stability metrics.

14 SOLAR ENERGY↗

Analysis of Association Between Remotely Sensed (RS) Data and Soil Transmitted Helminthes Infection Using Geographical Information Systems (GIS): Boaco, Nicaragua

Soil-transmitted helminths are intestinal nematodes that can infect all members of a population but specially school-age children living in poverty. Infection can be significantly reversed with anthelmintic drug treatments and sanitation improvement. Implementation of effective public health programs requires reliable and updated information to identify areas at higher risk and to calculate amount of drug required. Geo-referenced in situ prevalence data will be overlaid over an ecological map derived from RS data using ARC Map 9.3 (ESRI). Prevalence data and RS data matching at the same geographical location will be analyzed for correlation and those variables from RS data that better correlate with prevalence will be included in a multivariate regression model. Temperature, vegetation, and distance to bodies of water will be inferred using data from Moderate-Resolution Imaging Spectroradiometer (MODIS) and Landsat TM and ETM+. Elevation will be estimated with data from The Shuttle Radar Topography Mission (SRTM). Prevalence and intensity of infections are determined by parasitological survey (Kato Katz) of children enrolled in rural schools in Boaco, Nicaragua, in the communities of El Roblar, Cumaica Norte, Malacatoya 1, and Malacatoya 2). This study will demonstrate the importance of an integrated GIS/RS approach to define sampling clusters without the need for any ground-based survey. Such information is invaluable to identify areas of high risk and to geographically target control programs that maximize cost-effectiveness and sanitation efforts.

MorenoMadrinan, Max J.↗

Use of collateral information to improve LANDSAT classification accuracies

Methods to improve LANDSAT classification accuracies were investigated including: (1) the use of prior probabilities in maximum likelihood classification as a methodology to integrate discrete collateral data with continuously measured image density variables; (2) the use of the logit classifier as an alternative to multivariate normal classification that permits mixing both continuous and categorical variables in a single model and fits empirical distributions of observations more closely than the multivariate normal density function; and (3) the use of collateral data in a geographic information system as exercised to model a desired output information layer as a function of input layers of raster format collateral and image data base layers.

Strahler, A. H.↗

An Advanced Open-Source Platform for Air Quality Analysis, Visualization, and Prediction

Ambient air pollution is the largest environmental health risk factor, leading to several million premature deaths globally per year. The challenge of combating poor air quality is exacerbated by growing urban populations, changing emissions, and a warming climate. While there have been many advances monitoring and modeling of atmospheric composition, reflected in the dramatic increase in archived Earth Observations, there is no single measurement or method that alone can provide an accurate depiction of the entire atmosphere. The rapidly growing collections of observational and modeling data require us to be smarter about what data to include, and how such data is used. In recent years, NASA has invested significantly in advancing the concepts for Analytics Collaborative Framework (ACF) [5] and New Observing Strategies (NOS) [4] to tackle our software infrastructure need for harmonized data management and dynamic acquisition of diverse measurements for on-demand, interactive, multivariate analysis, and access [3]. It is not enough to have a big data, standalone analytics solution; it is critical that we start integrating data from remote sensing, modeling, and in-situ networks in a harmonized manner that enables timely and data-driven decision-making for air quality management. This work presents the design and development of an Air Quality Analytics Collaborative Framework (AQ ACF), as part of NASA’s Advanced Information Systems Technology (AIST) effort, to establish a data, machine-learning, and numerically driven platform for air quality analysis, visualization, and prediction.

Liu, Qian↗

Dynamic Transcriptomic and Phosphoproteomic Analysis During Cell Wall Stress in Aspergillus nidulans

The fungal cell-wall integrity signaling (CWIS) pathway regulates cellular response to environmental stress to enable wall repair and resumption of normal growth. This complex, interconnected, pathway has been only partially characterized in filamentous fungi. To better understand the dynamic cellular response to wall perturbation, a β-glucan synthase inhibitor (micafungin) was added to a growing A. nidulans shake-flask culture. From this flask, transcriptomic and phosphoproteomic data were acquired over 10 and 120 min, respectively. To differentiate statistically-significant dynamic behavior from noise, a multivariate adaptive regression splines (MARS) model was applied to both data sets. Over 1800 genes were dynamically expressed and over 700 phosphorylation sites had changing phosphorylation levels upon micafungin exposure. Twelve kinases had altered phosphorylation and phenotypic profiling of all non-essential kinase deletion mutants revealed putative connections between PrkA, Hk-8–4, and Stk19 and the CWIS pathway. Our collective data implicate actin regulation, endocytosis, and septum formation as critical cellular processes responding to activation of the CWIS pathway, and connections between CWIS and calcium, HOG, and SIN signaling pathways.

59 BASIC BIOLOGICAL SCIENCES↗

Survey of chondrule average properties in H-, L-, and LL-group chondrites - Are chondrules the same in all unequilibrated ordinary chondrites?

The petrogenetic properties of chondrules in different unequilibrated ordinary chondrites (UOCs) are compared to averaged chondrule-suite values obtained from recent analyses of several H-group, L-group, and LL-group chondrites. The purpose of the study was to develop a data base for future statistical analyses of chondrite characteristics. Mean end-member compositions of olivine (mol percent Fa) and pyroxene (mol percent Fs) were used as indices of the relative degree of 'equilibration' of each chondrule suite. It is found that the bulk chondrule geometric-mean abundances of Na, Mg, and Ni are the same from one UOC to another, and show no major systematic trends related to the H-group, L-group, of LL-group parentage of the host chondrites. The patterns of rare-earth element abundances in the chondrules are also examined, and the results are compared with statistical analyses. It is concluded that multivariate statistical analysis of pooled UOC chondrule data is justified for chondrule bulk compositions, as long as the statistical results are not misinterpreted as the primary petrogenetic features of chondrules.

Gooding, J. L.↗

Development of a Generic Creep-Fatigue Life Prediction Model

The objective of this research proposal is to further compile creep-fatigue data of steel alloys and superalloys used in military aircraft engines and/or rocket engines and to develop a statistical multivariate equation. The newly derived model will be a probabilistic fit to all the data compiled from various sources. Attempts will be made to procure the creep-fatigue data from NASA Glenn Research Center and other sources to further develop life prediction models for specific alloy groups. In a previous effort [1-3], a bank of creep-fatigue data has been compiled and tabulated under a range of known test parameters. These test parameters are called independent variables, namely; total strain range, strain rate, hold time, and temperature. The present research attempts to use these variables to develop a multivariate equation, which will be a probabilistic equation fitting a large database. The data predicted by the new model will be analyzed using the normal distribution fits, the closer the predicted lives are with the experimental lives (normal line 1 to 1 fit) the better the prediction. This will be evaluated in terms of a coefficient of correlation, R 2 as well. A multivariate equation developed earlier [3] has the following form, where S, R, T, and H have specific meaning discussed later.

Goswami, Tarun↗

Machine learning-based ethylene and carbon monoxide estimation, real-time optimization, and multivariable feedback control of an experimental electrochemical reactor

Electrochemical reduction of CO 2 gas is a novel CO 2 utilization technique that has the potential to mitigate the global climate crisis caused by anthropogenic CO 2 emissions, and enable the large-scale storage of energy generated from renewable sources in the form of carbon-based chemicals and fuels. However, due to the complexity of the electrochemical reactions, the explicit first-principles models for CO2 reduction are not available yet, and there has been a limited effort to develop process modeling, optimization and control of CO 2 electrochemical reactors. To this end, a rotating cylinder electrode (RCE) reactor has been constructed at UCLA to understand the mass transfer and reaction kinetics effects separately on the productivity. In the RCE reactor, the applied potential strongly influences the reaction energetics and the electrode rotation speed affects the hydrodynamic boundary layer and modifies the film mass transfer coefficient, which involves convective and diffusive transport. Further, the present work aims to develop a multi-input multi-output (MIMO) control scheme for the RCE reactor that integrates techniques from artificial and recurrent neural network modeling, nonlinear optimization, and process controller design. Specifically, production rates of two products from the experimental reactor, ethylene and carbon monoxide, are controlled by manipulating two inputs, applied potential and catalyst rotation speed. Process dynamics and controllability are analyzed, a feedback control strategy is designed and the controllers are tuned accordingly. The experimental electrochemical cell is employed to gather data for process modeling and implement the multivariable control system. Finally, the experimental results are presented which demonstrate excellent closed-loop performance by the control system and regulation of the outputs at three different set-points including an economically-optimal set-point.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Annual Sea Level Variability Induced by Changes in Sea Ice Extent and Accumulation on Ice Sheets: An Assessment Based on Remotely Sensed Data

Changes of mean annual net accumulation at the surface on the grounded ice sheets of East Antarctica, West Antarctica, and Greenland in response to variations in sea ice extent are estimated using grid-point values 100 km apart. The data bases are assembled principally by bilinear interpolation of remotely sensed brightness temperature (Nimbus-5 ESMR, Nimbus-7 SMMR), surface temperature (Nimbus-7 THIR), and surface elevation (ERS-1 radar altimeter). These data, complemented by field data where remotely sensed data are not available, are used in multivariate analyses in which mean annual accumulation (derived from firn emissivity) is the dependent variable; the independent variables are latitude, surface elevation, mean annual surface temperature, and mean annual distance to open ocean (as a source of energy and moisture). The last is the shortest distance measured between a grid point and the mean annual position of the 10% sea ice concentration boundary, and is used as an index of changes in sea ice extent as well as of mean concentration. Stepwise correlation analyses indicate that variations in sea ice extent of +/-50 km would lead to changes in accumulation inversely of +/-4% on East Antarctica, +/- 10% on West Antarctica, and +4% on Greenland. These results are compared with those obtained in a previous study using visually interpolated values from contoured compilations of field data; they substantiate the findings for the Antarctic ice sheets (+/-4% on East Antarctica, +/-9% in West Antarctica), and suggest a reduction by one half of the probable change of accumulation on Greenland (from +/-8%). The results also suggest a reduction of the combined contribution to sea level variability to +/- 0.19 mm/a (from +/- 0.22 mm/a).

Zwally, H. J.↗

TwinMe4AD: WGAN-based Digital Twins for Anomaly Detection

SAND2024-08373O TwinMe4AD is a Python-based software tool designed for anomaly detection using digital twins that closely mimic real, wearable healthcare datasets. The tool is invaluable for scenarios where collecting data is either expensive or impractical, serving as a privacy-preserving solution. Sensitive information is protected by training deep learning models on synthetic data derived from real datasets. One of TwinMe4AD's key features is its anomaly detection capability, which is based on fourth-order moments of parameters. This versatile approach can be applied across a range of datasets, from univariate to multivariate, making it compatible with various types of data. It also generates synthetic twins using Wasserstein Generative Adversarial Networks (WGANs), allowing users to create a small cohort of a population similar to that of a village population. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Poorey, Kunal↗

Pilot-Scale Pelleting Tests on High-Moisture Pine, Switchgrass, and Their Blends: Impact on Pellet Physical Properties, Chemical Composition, and Heating Values

In this study, we evaluated the pelleting characteristics of southern yellow pine (SYP), switchgrass (SG), and their blends for thermochemical conversion processes, such as pyrolysis and gasification. Using a pilot-scale ring-die pellet mill, we specifically assessed the impact of blend moisture, length-to-diameter (L/D) ratio in the pellet die, and ratio of pine to SG on the physico-chemical properties of the resulting pellets. We found that an increase in pine content by 25–50% marginally affected the bulk density; however, it also led to an increase in calorific value by 7% and a decrease in ash content by 72%. A moisture content of 25% (wet basis) and an L/D ratio of 5 resulted in poor pellet durability at <90% and bulk density values of <500 kg/m 3 , but increasing the L/D ratio to 9 and lowering the moisture content to 20% (w.b.) improved the pellet durability to >90% and the bulk density to >500 kg/m 3 . Blends with ≥50% pine content resulted in lower energy consumption, while a lower L/D ratio resulted in higher pelleting energy. Based on these findings, we successfully demonstrated the high-moisture pelleting of 2.5 ton of pine top residues blended with SG at 60:40 and 50:50 ratios. The quality of the pellets was monitored off-line and at-line by near infrared (NIR) spectroscopy. Multivariate models constructed by combining the NIR data and the pelleting process variables could successfully predict the pine content (R 2 = 0.99), higher heating value (R 2 = 0.98), ash (R 2 = 0.95), durability (R 2 = 0.94), and bulk density (R 2 = 0.86) of the pellets. Thus, we established how blending and densification of SYP and SG biomass could improve feedstock specifications and that NIR spectroscopy can effectively monitor the pellet properties during the high-moisture pelleting process.

09 BIOMASS FUELS↗

Gene expression of functionally-related genes coevolves across fungal species: detecting coevolution of gene expression using phylogenetic comparative methods

Researchers often measure changes in gene expression across conditions to better understand the shared functional roles and regulatory mechanisms of different genes. Analogous to this is comparing gene expression across species, which can improve our understanding of the evolutionary processes shaping the evolution of both individual genes and functional pathways. One area of interest is determining genes showing signals of coevolution, which can also indicate potential functional similarity, analogous to co-expression analysis often performed across conditions for a single species. However, as with any trait, comparing gene expression across species can be confounded by the non-independence of species due to shared ancestry, making standard hypothesis testing inappropriate. We compared RNA-Seq data across 18 fungal species using a multivariate Brownian Motion phylogenetic comparative method (PCM), which allowed us to quantify coevolution between protein pairs while directly accounting for the shared ancestry of the species. Our work indicates proteins which physically-interact show stronger signals of coevolution than randomly-generated pairs. Interactions with stronger empirical and computational evidence also showing stronger signals of coevolution. We examined the effects of number of protein interactions and gene expression levels on coevolution, finding both factors are overall poor predictors of the strength of coevolution between a protein pair. Simulations further demonstrate the potential issues of analyzing gene expression coevolution without accounting for shared ancestry in a standard hypothesis testing framework. Furthermore, our simulations indicate the use of a randomly-generated null distribution as a means of determining statistical significance for detecting coevolving genes with phylogenetically-uncorrected correlations, as has previously been done, is less accurate than PCMs, although is a significant improvement over standard hypothesis testing. These methods are further improved by using a phylogenetically-corrected correlation metric. Our work highlights potential benefits of using PCMs to detect gene expression coevolution from high-throughput omics scale data. This framework can be built upon to investigate other evolutionary hypotheses, such as changes in transcription regulatory mechanisms across species.

59 BASIC BIOLOGICAL SCIENCES↗

A variational assimilation method for satellite and conventional data: Development of basic model for diagnosis of cyclone systems

A three-dimensional diagnostic model for the assimilation of satellite and conventional meteorological data is developed with the variational method of undetermined multipliers. Gridded fields of data from different type, quality, location, and measurement source are weighted according to measurement accuracy and merged using least squares criteria so that the two nonlinear horizontal momentum equations, the hydrostatic equation, and an integrated continuity equation are satisfied. The model is used to compare multivariate variational objective analyses with and without satellite data with initial analyses and the observations through criteria that were determined by the dynamical constraints, the observations, and pattern recognition. It is also shown that the diagnoses of local tendencies of the horizontal velocity components are in good comparison with the observed patterns and tendencies calculated with unadjusted data. In addition, it is found that the day-night difference in TOVS biases are statistically different (95% confidence) at most levels. Also developed is a hybrid nonlinear sigma vertical coordinate that eliminates hydrostatic truncation error in the middle and upper troposphere and reduces truncation error in the lower troposphere. Finally, it is found that the technique used to grid the initial data causes boundary effects to intrude into the interior of the analysis a distance equal to the average separation between observations.

Achtemeier, G. L.↗

Classification by thresholding

A procedure is given which substantially reduces the processing time needed to perform maximum likelihood classification on large data sets. The given method uses a set of fixed thresholds which, if exceeded by one probability density function, makes it unnecessary to evaluate a competing density function. Proofs are given of the existence and optimality of these thresholds for the class of continuous, unimodal, and quasi-concave density functions (which includes the multivariable normal), and a method for computing the thresholds is provided for the specific case of multivariate normal densities. An example with remote sensing data consisting of some 20,000 observations of four-dimensional data from nine ground-cover classes shows that by using thresholds, one could cut the processing time almost in half.

Feiveson, A. H.↗

A Multivariate Space‐Time Dynamic Model for Characterizing the Atmospheric Impacts Following the Mt. Pinatubo Eruption

The June 1991 Mt. Pinatubo eruption resulted in a massive increase of sulfate aerosols in the atmosphere, absorbing radiation and leading to global changes in surface and stratospheric temperatures. A volcanic eruption of this magnitude serves as a natural analog for stratospheric aerosol injection, a proposed solar radiation modification method to combat a warming climate. The impacts of such an event are multifaceted and region-specific. Our goal is to characterize the multivariate and dynamic nature of the atmospheric impacts following the Mt. Pinatubo eruption. We developed a multivariate space-time dynamic linear model to understand the full extent of the spatially- and temporally-varying impacts. Specifically, spatial variation is modeled using a flexible set of basis functions for which the basis coefficients are allowed to vary in time through a vector autoregressive (VAR) structure. This novel model is cast in a Dynamic Linear Model (DLM) framework and estimated via a customized MCMC approach. We demonstrate how the model quantifies the relationships between key atmospheric parameters prior to and following the Mt. Pinatubo eruption with reanalysis data from MERRA-2 and highlight when such a model is advantageous over univariate models.

Dynamic Linear Model↗

Multivariate Bayesian Optimization of CoO Nanoparticles for CO 2 Hydrogenation Catalysis

The hydrogenation of CO 2 holds promise for transforming the production of renewable fuels and chemicals. However, the challenge lies in developing robust and selective catalysts for this process. Transition metal oxide catalysts, particularly cobalt oxide, have shown potential for CO 2 hydrogenation, with performance heavily reliant on crystal phase and morphology. Achieving precise control over these catalyst attributes through colloidal nanoparticle synthesis could pave the way for catalyst and process advancement. Yet, navigating the complexities of colloidal nanoparticle syntheses, governed by numerous input variables, poses a significant challenge in systematically controlling resultant catalyst features. We present a multivariate Bayesian optimization, coupled with a data-driven classifier, to map the synthetic design space for colloidal CoO nanoparticles and simultaneously optimize them for multiple catalytically relevant features within a target crystalline phase. The optimized experimental conditions yielded small, phase-pure rock salt CoO nanoparticles of uniform size and shape. These optimized nanoparticles were then supported on SiO 2 and assessed for thermocatalytic CO 2 hydrogenation against larger, polydisperse CoO nanoparticles on SiO 2 and a conventionally prepared catalyst. The optimized CoO/SiO 2 catalyst consistently exhibited higher activity and CH 4 selectivity (ca. 98%) across various pretreatment reduction temperatures as compared to the other catalysts. This remarkable performance was attributed to particle stability and consistent H* surface coverage, even after undergoing the highest temperature reduction, achieving a more stable catalytic species that resists sintering and carbon occlusion.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗