Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Land Model Testbed: Accelerating Development, Benchmarking and Analysis of Land Surface Models

A Land Model Testbed (LMT), designed to provide a computational framework for systematically assessing model fidelity and supporting rapid development of complex multiscale models, offers a general-purpose workflow for conducting large ensemble simulations of multiple land surface models, post-processing large volumes of model output, and evaluating model results. It leverages existing tools for launching model simulations and the International Land Model Benchmarking (ILAMB) package for assessing model fidelity through comparison with best-available observational datasets. Increased complexity and proliferation of uncertain parameters in process representations in land surface models has driven the need for frequent and intensive testing and evaluating of models to quantify uncertainties and optimize parameters such that results are consistent with observations. The LMT described here meets these needs by providing tools to run thousands of ensemble simulations simultaneously and post-process their output files, by automating execution of an enhanced version of ILAMB with site-specific benchmarks and multivariate functional relationships, and by offering ensemble diagnostics and a customizable dashboard for displaying model performance metrics and associated graphics. We envision the LMT capabilities will serve as a foundational computational resource for a proposed user facility focused on terrestrial multiscale model--data integration.

Sreepathi, Sarat↗

Multivariate Machine Learning Models of Nanoscale Porosity from Ultrafast NMR Relaxometry

Abstract Nanoporous materials are of great interest in many applications, such as catalysis, separation, and energy storage. The performance of these materials is closely related to their pore sizes, which are inefficient to determine through the conventional measurement of gas adsorption isotherms. Nuclear magnetic resonance (NMR) relaxometry has emerged as a technique highly sensitive to porosity in such materials. Nonetheless, streamlined methods to estimate pore size from NMR relaxometry remain elusive. Previous attempts have been hindered by inverting a time domain signal to relaxation rate distribution, and dealing with resulting parameters that vary in number, location, and magnitude. Here we invoke well‐established machine learning techniques to directly correlate time domain signals to BET surface areas for a set of metal‐organic frameworks (MOFs) imbibed with solvent at varied concentrations. We employ this series of MOFs to establish a correlation between NMR signal and surface area via partial least squares (PLS), following screening with principal component analysis, and apply the PLS model to predict surface area of various nanoporous materials. This approach offers a high‐throughput, non‐destructive way to assess porosity in c.a. one minute. We anticipate this work will contribute to the development of new materials with optimized pore sizes for various applications.

Fricke, Sophia N.↗

Multivariate Machine Learning Models of Nanoscale Porosity from Ultrafast NMR Relaxometry

Abstract Nanoporous materials are of great interest in many applications, such as catalysis, separation, and energy storage. The performance of these materials is closely related to their pore sizes, which are inefficient to determine through the conventional measurement of gas adsorption isotherms. Nuclear magnetic resonance (NMR) relaxometry has emerged as a technique highly sensitive to porosity in such materials. Nonetheless, streamlined methods to estimate pore size from NMR relaxometry remain elusive. Previous attempts have been hindered by inverting a time domain signal to relaxation rate distribution, and dealing with resulting parameters that vary in number, location, and magnitude. Here we invoke well‐established machine learning techniques to directly correlate time domain signals to BET surface areas for a set of metal‐organic frameworks (MOFs) imbibed with solvent at varied concentrations. We employ this series of MOFs to establish a correlation between NMR signal and surface area via partial least squares (PLS), following screening with principal component analysis, and apply the PLS model to predict surface area of various nanoporous materials. This approach offers a high‐throughput, non‐destructive way to assess porosity in c.a. one minute. We anticipate this work will contribute to the development of new materials with optimized pore sizes for various applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative data sets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student t process regression. We apply Student t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via colouring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics and classic machine learning clustering algorithms.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Large-scale analysis of structural brain asymmetries in schizophrenia via the ENIGMA consortium

Left–right asymmetry is an important organizing feature of the healthy brain that may be altered in schizophrenia, but most studies have used relatively small samples and heterogeneous approaches, resulting in equivocal findings. We carried out the largest case–control study of structural brain asymmetries in schizophrenia, with MRI data from 5,080 affected individuals and 6,015 controls across 46 datasets, using a single image analysis protocol. Asymmetry indexes were calculated for global and regional cortical thickness, surface area, and subcortical volume measures. Differences of asymmetry were calculated between affected individuals and controls per dataset, and effect sizes were meta-analyzed across datasets. Small average case–control differences were observed for thickness asymmetries of the rostral anterior cingulate and the middle temporal gyrus, both driven by thinner left-hemispheric cortices in schizophrenia. Analyses of these asymmetries with respect to the use of antipsychotic medication and other clinical variables did not show any significant associations. Assessment of age- and sex-specific effects revealed a stronger average leftward asymmetry of pallidum volume between older cases and controls. Case–control differences in a multivariate context were assessed in a subset of the data (N = 2,029), which revealed that 7% of the variance across all structural asymmetries was explained by case–control status. Subtle case–control differences of brain macrostructural asymmetry may reflect differences at the molecular, cytoarchitectonic, or circuit levels that have functional relevance for the disorder. Reduced left middle temporal cortical thickness is consistent with altered left-hemisphere language network organization in schizophrenia.

60 APPLIED LIFE SCIENCES↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative datasets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student-t process regression. We apply Student-t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via coloring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics, and classic machine learning clustering algorithms.

97 MATHEMATICS AND COMPUTING↗

Machine learning–assisted prediction of heat fluxes through thermally anisotropic building envelopes

Thermally anisotropic building envelope (TABE) is a novel active building envelope that can save energy use to maintain thermal comfort in buildings by redirecting heat and coolness from building envelopes to thermal loops. Finite element models (FEMs) can be used to compute the heat fluxes through TABEs, but the high computational cost of finite element simulations has prevented parametric studies and design optimizations. This paper proposes a domain knowledge–informed, finite element–based machine learning framework to reduce the computation cost for the energy management of buildings installed with TABE that uses a ground thermal loop. First, the training heat flux data set was generated by FEM simulations with different thermal loop schedules. Then, both shallow learning models (i.e., multivariate linear regression and eXtreme Gradient Boost, or XGBoost) and a deep learning model (i.e., deep neural network, or DNN) were trained to predict the heat fluxes. Domain knowledge was used for data preprocessing and feature selection. Finally, the suitability of the selected machine learning model was tested under different thermal loop schedules. Herein, the case study results showed that: (1) XGBoost can be as accurate as DNN (coefficient of determination equal to 0.81) with much less training time; (2) the annual energy cost savings for different thermal loop schedules obtained by the XGBoost-predicted and FEM-calculated heat fluxes are consistent, having a difference of only 4%; and (3) XGBoost can reduce the computation time for the annual energy analysis of the case study building with a given thermal loop schedule from around 12 h by using FEM to less than 1 min.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Experimental study on kinetic oxidation of graphite IG-110 by steam

Graphite is proposed for use in High-temperature Gas-cooled Reactors (HTGRs) as the fuel matrix, neutron moderator/reflector, and core structural material. One important property of nuclear grade graphite is their resistance to oxidation in high-temperature environment. Extensive investigation has been performed in the literature for graphite oxidation by air. However, available experimental data are still limited for graphite oxidation by steam under conditions comparable to a postulated steam ingress accident in HTGRs. In this study, the oxidation rate of graphite IG-110 by steam was measured at temperatures from 850 to 1100 °C with the steam partial pressure varying from 0.5 to 20.0 kPa and the hydrogen partial pressure varying from 0 to 2.0 kPa. Further analysis confirms the oxidation process in this present study is dominated by the chemical kinetics, which lends credit to the data for being used to develop numerical models. It was observed that the increase of the kinetic oxidation rate with the steam partial pressure tends to become less apparent if the steam partial pressure keeps increasing. In addition, it was found that the partitioning of hydrogen inhibits the graphite-steam reaction process even with the steam partial pressure up to 20.0 kPa. However, this inhibiting effect starts to become saturated when the hydrogen partial pressure exceeds 1.0 kPa. The oxidation rates were fitted to the conventional Langmuir-Hinshelwood (LH) and Boltzmann-enhanced Langmuir-Hinshelwood (BLH) models by a multivariable optimization algorithm. The BLH model exhibits a better accuracy than the LH model within the specified experimental conditions. The predicted oxidation rate using the BLH model shows a mean relative difference of about 24% with the maximum difference of about 55% when compared with our experimental data.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Yet Another Discriminant Analysis (YADA): A Probabilistic Model for Machine Learning Applications

This paper presents a probabilistic model for various machine learning (ML) applications. While deep learning (DL) has produced state-of-the-art results in many domains, DL models are complex and over-parameterized, which leads to high uncertainty about what the model has learned, as well as its decision process. Further, DL models are not probabilistic, making reasoning about their output challenging. In contrast, the proposed model, referred to as Yet Another Discriminate Analysis(YADA), is less complex than other methods, is based on a mathematically rigorous foundation, and can be utilized for a wide variety of ML tasks including classification, explainability, and uncertainty quantification. YADA is thus competitive in most cases with many state-of-the-art DL models. Ideally, a probabilistic model would represent the full joint probability distribution of its features, but doing so is often computationally expensive and intractable. Hence, many probabilistic models assume that the features are either normally distributed, mutually independent, or both, which can severely limit their performance. YADA is an intermediate model that (1) captures the marginal distributions of each variable and the pairwise correlations between variables and (2) explicitly maps features to the space of multivariate Gaussian variables. Numerous mathematical properties of the YADA model can be derived, thereby improving the theoretic underpinnings of ML. Validation of the model can be statistically verified on new or held-out data using native properties of YADA. However, there are some engineering and practical challenges that we enumerate to make YADA more useful.

97 MATHEMATICS AND COMPUTING↗

QProR: An Efficient Framework for Quantity-of-Interest Based Progressive Retrieval with Guaranteed Error Control

Scientific applications generate an unprecedented volume of data, overwhelming the network and file systems’ bandwidth and posing challenges for efficient and scalable data retrieval and analysis. Progressive data compression offers a promising solution by enabling on-demand retrieval at reduced size. However, existing progressive methods either fail to bound the errors in essential quantities of interest (QoIs) derived from raw data or suffer from suboptimal retrieval efficiency. In this work, we propose QProR, an efficient QoI-based progressive framework that optimizes progressive retrieval for target QoIs. Our key contributions include: (1) a systematic framework that integrates error-controlled lossy compressors with bitplane encoding while decoupling the two processes for high flexibility and adaptability; (2) a novel weighted bitplane encoding method which incorperates QoI knowledge into data refactoring to enhance retrieval efficiency; (3) an optimized retrieval strategy that accounts for the varying impacts of different variables on multivariate QoIs; (4) comprehensive evaluations using six real-world datasets from multiple scientific applications and thorough comparisons against state of the arts. Experimental results demonstrate that QProR achieves up to 80.38% reduction in the retrieval size under the same requested QoI error tolerance, when compared with the best-performing existing methods. When transferring 384 GB of scientific data to remote sites, QProR delivers up to 1.68 × speedup in the end-to-end data transfer performance.

Li, Wenbo [University of Kentucky]↗

A geo-visual analysis for exploring the socioeconomic benefits of the heating electrification using geothermal energy

In parallel to population growth and climate change, the rapid pace of urbanization worldwide has led to an enormous increase in energy demand and costs in urban areas. The subsequent energy burden has become an increasing concern for many households in the U.S. Previous studies have revealed that geothermal resources can effectively lower the electricity demand and carbon emissions in large cities. In this paper, we focus on the socioeconomic impacts of geothermal energy on urban systems by presenting an interactive visual analytics dashboard. The dashboard allows urban planners to spatially examine geothermal energy's practical benefits on energy affordability, urban livability, and resilience across the U.S. We compiled a list of socioeconomic metrics by integrating the simulation results from multiple geothermal and building models with multi-domain urban datasets (socioeconomic, demographic, and electricity utility). These metrics are created to characterize the benefits of the heating electrification of buildings using Geothermal Heat Pumps (GHP) for lowering the energy burden of middle-and low-income households nationwide. The visual dashboard employs a combination of multivariate, glyph-based, and geospatial visualization to reveal the variability and patterns in our metrics. We present a pilot study to demonstrate the GHPs' potential as a renewable and affordable solution for increasing the economic and energy grid resilience in U.S cities.

Xu, Haowen↗

Stochastic Learning Approach for Binary Optimization: Application to Bayesian Optimal Design of Experiments

Here, we present a novel stochastic approach to binary optimization suited for optimal experimental design (OED) for Bayesian inverse problems governed by mathematical models such as partial differential equations. The OED utility function, namely, the regularized optimality criterion, is cast into a stochastic objective function in the form of an expectation over a multivariate Bernoulli distribution. The probabilistic objective is then solved by using a stochastic optimization routine to find an optimal observational policy. This formulation (a) is generally applicable to binary optimization problems with soft constraints and is ideal for OED and sensor placement problems; (b) does not require differentiability of the original objective function (e.g., a utility function in OED applications) with respect to the design variable, and thus it enables direct employment of sparsity-enforcing penalty functions such as $\ell_0$, without needing to utilize a continuation procedure or apply a rounding technique; (c) exhibits much lower computational cost than traditional gradient-based relaxation approaches; and (d) can be applied to both linear and nonlinear OED problems with proper choice of the utility function. The proposed approach is analyzed from an optimization perspective with detailed convergence analysis of the optimization approach and is also analyzed from a machine learning perspective with correspondence to policy gradient reinforcement learning. The approach is demonstrated numerically by using an idealized two-dimensional Bayesian linear inverse problem and validated by extensive numerical experiments carried out for sensor placement in a parameter identification setup.

97 MATHEMATICS AND COMPUTING↗

Riders’ perceptions towards transit bus electrification: Evidence from Salt Lake City, Utah

While battery electric buses (BEBs) can lead to energy savings and reduced emissions, BEB adoption is developing slowly. Although BEBs offer quieter operations, better acceleration, and no smell of diesel or gas fumes, little focus has been placed on the user’s perspective. Here, this study investigates bus riders’ preferences toward BEBs. To achieve these objectives, a survey was designed and administered to solicit riders’ typical travel behaviors and patterns as well as preferences and opinions about BEBs’ performance in terms of emissions and noise. Statistical analysis showed that several factors influence rider perceptions towards transit bus electrification that include trip purpose, attitudes towards environmental issues and environmental impacts of BEBs, and certain non-instrumental ride factors such as ride comfort and social image. A better understanding of the importance of electrification to transit riders can help transit service providers adjust their marketing decisions and their systemwide operations to accommodate preferences towards BEBs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

High Throughput Data-Driven Design of Laser-Crystallized 2D MoS 2 Chemical Sensors: A Demonstration for NO 2 Detection

High throughput characterization and processing techniques are becoming increasingly necessary to navigate multivariable, data-driven design challenges for sensors and electronic devices. For two-dimensional materials, device performance is highly dependent upon a vast array of material properties including the number of layers, lattice strain, carrier concentration, defect density, and grain structure. In this work, laser crystallization was used to locally pattern and transform hundreds of regions of amorphous MoS 2 thin films into 2D 2H-MoS 2 . Here a high throughput Raman spectroscopy approach was subsequently used to assess the process-dependent structural and compositional variations for each illuminated region, yielding over 6000 distinct nonresonant, resonant, and polarized Raman spectra. The rapid generation of a comprehensive library of structural and compositional data elucidated important trends between structure–property processing relationships involving laser-crystallized MoS 2 , including the relationships between grain size, grain orientation, and intrinsic strain. Moreover, extensive analysis of structure/property relationships allowed for intelligent design and evaluation of major contributions to device performance in MoS 2 chemical sensors. In particular, it is found that NO 2 sensor performance is strongly dependent on the orientation of the MoS 2 grains relative to the crystal plane.

36 MATERIALS SCIENCE↗

Absolute Band Intensity of the Iodine Monochloride Fundamental Mode for Infrared Sensing and Quantitative Analysis

Iodine monochloride (ICl) is a gaseous off-product of molten salt reactors; monitoring this heteronuclear diatomic is of great interest for both environmental and safety purposes. In this paper we investigate the possibility of infrared monitoring of ICl by measuring the far-infrared absorption cross section of its fundamental band near 381 cm -1 . We have performed quantitative studies of the neat gas in a 20 cm cell at 25, 35, 50 and 70 oC at multiple pressures up to ~ 9 Torr and investigated the temperature and pressure dependence of the band’s infrared cross section. Quantitative measurements were problematic due to sample adhesion to the cell walls and windows as well as reactions/possible hydrolysis of ICl to form HCl gas. Effects were mitigated by measuring only the neat gas, using short measurement times and subtracting out the partial pressure of the HCl(g). The integrated band strength is shown to be temperature independent and was found to be equal to 9.1 x 10 -19 (cm 2 /molecule) cm-1. As expected, the temperature dependence of the band profile showed only a small effect over this limited temperature range. Furthermore, we have also investigated using the absorption data along with inverse least squares multivariate methods for the quantitative monitoring of ICl effluent concentrations under different scenarios using infrared (standoff) sensing and compare these results with traditional Beer’s Law (univariate) techniques.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Review of Bayesian Networks for Spatial Data

We report Bayesian networks are a popular class of multivariate probabilistic models as they allow for the translation of prior beliefs about conditional dependencies between variables to be easily encoded into their model structure. Due to their widespread usage, they are often applied to spatial data for inferring properties of the systems under study and also generating predictions for how these systems may behave in the future. We review published research on methodologies for representing spatial data with Bayesian networks and also summarize the application areas for which Bayesian networks are employed in the modeling of spatial data. We find that a wide variety of perspectives are taken, including a GIS-centric focus on efficiently generating geospatial predictions, a statistical focus on rigorously constructing graphical models controlling for spatial correlation, as well as a range of problem-specific heuristics for mitigating the effects of spatial correlation and dependency arising in spatial data analysis. Special attention is also paid to potential future directions for integration of Bayesian networks with spatial processes.

97 MATHEMATICS AND COMPUTING↗

The bright frontiers of microbial metabolic optogenetics

In recent years, light-responsive systems from the field of optogenetics have been applied to several areas of metabolic engineering with remarkable success. By taking advantage of light's high tunability, reversibility, and orthogonality to host endogenous processes, optogenetic systems have enabled unprecedented dynamical controls of microbial fermentations for chemical production, metabolic flux analysis, and population compositions in co-cultures. In this article, we share our opinions on the current state of this new field of metabolic optogenetics. Furthermore, we make the case that it will continue to impact metabolic engineering in increasingly new directions, with the potential to challenge existing paradigms for metabolic pathway and strain optimization as well as bioreactor operation.

59 BASIC BIOLOGICAL SCIENCES↗

Robust Predictive Control for Modular Solid-State Transformer With Reduced DC Link and Parameter Mismatch

This paper presents the analysis and implementation of a predictive control method for dc-link regulation and voltage balance in a cascaded modular reduced dc-link solid-state transformer (SST). Passive components like bulky dc links limit the power density of power converters, especially medium-voltage (MV) SST. Reduced dc-link or low-inertia converters can dramatically reduce the size, cost, and weight by tolerating larger dc-link ripples and improve the reliability with electrolytic capacitor-less dc link. However, a small dc link leads to tight coupling between the input and the output stages, which is a challenge for control design. In stacked low-inertia converters (SLIC), the low-inertia converter modules are stacked for MV applications, resulting in coupling between the modules and making the control more challenging. A new model predictive control method which can achieve deadbeat regulation on the dc link without weighting factors has been proposed to address this novel problem. This paper focuses on analyzing the condition of the low-inertia dc link up to 80% ripple, the robustness of the control under parameter mismatches, high-order terms, and important implementation issues such as model-based sampling and computation delay compensation. Significantly, the high-order terms are introduced because of the large dc-link ripple. These high-order terms are unique to the SLIC and negligible in conventional high-inertia converters. A discrete-time large-signal model is built to capture the dc-link’s nonlinear dynamics, and the eigenvalues of a small-signal Jacobian matrix are analyzed with Floquet theory to evaluate stability, using the modular soft-switching solid-state transformer (M-S4T) as an example of the SLIC. Simulation and experimental results of an MVDC M-S4T verify the analysis and the predictive control method. Finally, the general application of the predictive control to low-inertia converters is compared against a conventional PI controller using a reduced dc-link active-front-end (AFE) rectifier as an example.

14 SOLAR ENERGY↗