Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Measurement of lepton universality parameters in B + → K + ℓ + ℓ – and B 0 → K* 0 ℓ + ℓ – decays

A simultaneous analysis of the B + → K + ℓ + ℓ – and B 0 → K* 0 ℓ + ℓ – decays is performed to test muon-electron universality in two ranges of the square of the dilepton invariant mass, q 2 . The measurement uses a sample of beauty meson decays produced in proton-proton collisions collected with the LHCb detector between 2011 and 2018, corresponding to an integrated luminosity of 9 fb –1 . A sequence of multivariate selections and strict particle identification requirements produce a higher signal purity and a better statistical sensitivity per unit luminosity than previous LHCb lepton universality tests using the same decay modes. Residual backgrounds due to misidentified hadronic decays are studied using data and included in the fit model. Each of the four lepton universality measurements reported is either the first in the given q 2 interval or supersedes previous LHCb measurements. The results are compatible with the predictions of the Standard Model.

79 ASTRONOMY AND ASTROPHYSICS↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative datasets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student-t process regression. We apply Student-t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via coloring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics, and classic machine learning clustering algorithms.

97 MATHEMATICS AND COMPUTING↗

Wind Turbine Gearbox Failure Detection Through Cumulative Sum of Multivariate Time Series Data

The wind energy industry is continuously improving their operational and maintenance practice for reducing the levelized costs of energy. Anticipating failures in wind turbines enables early warnings and timely intervention, so that the costly corrective maintenance can be prevented to the largest extent possible. It also avoids production loss owing to prolonged unavailability. One critical element allowing early warning is the ability to accumulate small-magnitude symptoms resulting from the gradual degradation of wind turbine systems. Inspired by the cumulative sum control chart method, this study reports the development of a wind turbine failure detection method with such early warning capability. Specifically, the following key questions are addressed: what fault signals to accumulate, how long to accumulate, what offset to use, and how to set the alarm-triggering control limit. We apply the proposed approach to 2 years’ worth of Supervisory Control and Data Acquisition data recorded from five wind turbines. We focus our analysis on gearbox failure detection, in which the proposed approach demonstrates its ability to anticipate failure events with a good lead time.

17 WIND ENERGY↗

Machine learning–assisted prediction of heat fluxes through thermally anisotropic building envelopes

Thermally anisotropic building envelope (TABE) is a novel active building envelope that can save energy use to maintain thermal comfort in buildings by redirecting heat and coolness from building envelopes to thermal loops. Finite element models (FEMs) can be used to compute the heat fluxes through TABEs, but the high computational cost of finite element simulations has prevented parametric studies and design optimizations. This paper proposes a domain knowledge–informed, finite element–based machine learning framework to reduce the computation cost for the energy management of buildings installed with TABE that uses a ground thermal loop. First, the training heat flux data set was generated by FEM simulations with different thermal loop schedules. Then, both shallow learning models (i.e., multivariate linear regression and eXtreme Gradient Boost, or XGBoost) and a deep learning model (i.e., deep neural network, or DNN) were trained to predict the heat fluxes. Domain knowledge was used for data preprocessing and feature selection. Finally, the suitability of the selected machine learning model was tested under different thermal loop schedules. Herein, the case study results showed that: (1) XGBoost can be as accurate as DNN (coefficient of determination equal to 0.81) with much less training time; (2) the annual energy cost savings for different thermal loop schedules obtained by the XGBoost-predicted and FEM-calculated heat fluxes are consistent, having a difference of only 4%; and (3) XGBoost can reduce the computation time for the annual energy analysis of the case study building with a given thermal loop schedule from around 12 h by using FEM to less than 1 min.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Experimental study on kinetic oxidation of graphite IG-110 by steam

Graphite is proposed for use in High-temperature Gas-cooled Reactors (HTGRs) as the fuel matrix, neutron moderator/reflector, and core structural material. One important property of nuclear grade graphite is their resistance to oxidation in high-temperature environment. Extensive investigation has been performed in the literature for graphite oxidation by air. However, available experimental data are still limited for graphite oxidation by steam under conditions comparable to a postulated steam ingress accident in HTGRs. In this study, the oxidation rate of graphite IG-110 by steam was measured at temperatures from 850 to 1100 °C with the steam partial pressure varying from 0.5 to 20.0 kPa and the hydrogen partial pressure varying from 0 to 2.0 kPa. Further analysis confirms the oxidation process in this present study is dominated by the chemical kinetics, which lends credit to the data for being used to develop numerical models. It was observed that the increase of the kinetic oxidation rate with the steam partial pressure tends to become less apparent if the steam partial pressure keeps increasing. In addition, it was found that the partitioning of hydrogen inhibits the graphite-steam reaction process even with the steam partial pressure up to 20.0 kPa. However, this inhibiting effect starts to become saturated when the hydrogen partial pressure exceeds 1.0 kPa. The oxidation rates were fitted to the conventional Langmuir-Hinshelwood (LH) and Boltzmann-enhanced Langmuir-Hinshelwood (BLH) models by a multivariable optimization algorithm. The BLH model exhibits a better accuracy than the LH model within the specified experimental conditions. The predicted oxidation rate using the BLH model shows a mean relative difference of about 24% with the maximum difference of about 55% when compared with our experimental data.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

QProR: An Efficient Framework for Quantity-of-Interest Based Progressive Retrieval with Guaranteed Error Control

Scientific applications generate an unprecedented volume of data, overwhelming the network and file systems’ bandwidth and posing challenges for efficient and scalable data retrieval and analysis. Progressive data compression offers a promising solution by enabling on-demand retrieval at reduced size. However, existing progressive methods either fail to bound the errors in essential quantities of interest (QoIs) derived from raw data or suffer from suboptimal retrieval efficiency. In this work, we propose QProR, an efficient QoI-based progressive framework that optimizes progressive retrieval for target QoIs. Our key contributions include: (1) a systematic framework that integrates error-controlled lossy compressors with bitplane encoding while decoupling the two processes for high flexibility and adaptability; (2) a novel weighted bitplane encoding method which incorperates QoI knowledge into data refactoring to enhance retrieval efficiency; (3) an optimized retrieval strategy that accounts for the varying impacts of different variables on multivariate QoIs; (4) comprehensive evaluations using six real-world datasets from multiple scientific applications and thorough comparisons against state of the arts. Experimental results demonstrate that QProR achieves up to 80.38% reduction in the retrieval size under the same requested QoI error tolerance, when compared with the best-performing existing methods. When transferring 384 GB of scientific data to remote sites, QProR delivers up to 1.68 × speedup in the end-to-end data transfer performance.

Li, Wenbo [University of Kentucky]↗

Meta-analysis identifies pleiotropic loci controlling phenotypic trade-offs in sorghum

Abstract Community association populations are composed of phenotypically and genetically diverse accessions. Once these populations are genotyped, the resulting marker data can be reused by different groups investigating the genetic basis of different traits. Because the same genotypes are observed and scored for a wide range of traits in different environments, these populations represent a unique resource to investigate pleiotropy. Here, we assembled a set of 234 separate trait datasets for the Sorghum Association Panel, a group of 406 sorghum genotypes widely employed by the sorghum genetics community. Comparison of genome-wide association studies (GWAS) conducted with two independently generated marker sets for this population demonstrate that existing genetic marker sets do not saturate the genome and likely capture only 35–43% of potentially detectable loci controlling variation for traits scored in this population. While limited evidence for pleiotropy was apparent in cross-GWAS comparisons, a multivariate adaptive shrinkage approach recovered both known pleiotropic effects of existing loci and new pleiotropic effects, particularly significant impacts of known dwarfing genes on root architecture. In addition, we identified new loci with pleiotropic effects consistent with known trade-offs in sorghum development. These results demonstrate the potential for mining existing trait datasets from widely used community association populations to enable new discoveries from existing trait datasets as new, denser genetic marker datasets are generated for existing community association populations.

59 BASIC BIOLOGICAL SCIENCES↗

Observation of four-top-quark production in the multilepton final state with the ATLAS detector

This paper presents the observation of four-top-quark ($t$$\overline{t}$$t$$\overline{t}$) production in proton-proton collisions at the LHC. The analysis is performed using an integrated luminosity of 140 fb -1 at a centre-of-mass energy of 13 TeV collected using the ATLAS detector. Events containing two leptons with the same electric charge or at least three leptons (electrons or muons) are selected. Event kinematics are used to separate signal from background through a multivariate discriminant, and dedicated control regions are used to constrain the dominant backgrounds. The observed (expected) significance of the measured $t$$\overline{t}$$t$$\overline{t}$ signal with respect to the standard model (SM) background-only hypothesis is 6.1 (4.3) standard deviations. The $t$$\overline{t}$$t$$\overline{t}$ production cross section is measured to be ${22.5}^{+6.6}_{-5.5}$, consistent with the SM prediction of 12.0 ± 2.4 fb within 1.8 standard deviations. Data are also used to set limits on the three-top-quark production cross section, being an irreducible background not measured previously, and to constrain the top-Higgs Yukawa coupling and effective field theory operator coefficients that affect $t$$\overline{t}$$t$$\overline{t}$ production.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multi-Objective design of interlocking metasurfaces using conditional diffusion models

Unit cell design remains a major challenge for interlocking metasurfaces, a promising joining technology for dissimilar materials, due to the complex, competing, multivariate design space and the need for rapid adaptation to varying performance requirements. This study explores Conditional Diffusion Models as a design optimization tool for interlocking metasurfaces. Given the complex, competing, multivariate design space for interlocking metasurfaces, unit cell design remains a major challenge for this joining technology. We trained a conditional diffusion model on 25,000 finite element analysis-simulated interlocking metasurface unit cells to generate designs with tailored thermo-mechanical properties (tensile strength, shear strength, and thermal conductivity) based on specified performance criteria. The model demonstrated a success rate of approximately 72 % in producing designs that met specified property bounds. The conditional diffusion model generated both thermally resistive and conductive designs, revealing clear trends in design characteristics: taller, dendritic structures were advantageous for tensile loads, while shorter, robust designs excelled in shear applications. Our findings indicate that the model's performance is more influenced by the breadth of the design space than by the quantity of training data, highlighting the importance of expansive design domains for generating innovative solutions. This work establishes conditional diffusion models as a highly efficient and adaptable tool for rapid interlocking metasurface unit cell design, paving the way for advancements in multi-material joining technologies, as well as highlighting the justification to leverage conditional diffusion models as design tools across complex design domains.

Conditional diffusion models↗

The safety of pranlukast and montelukast during the first trimester of pregnancy: A prospective, two‐centered cohort study in Japan

Abstract For leukotriene receptor antagonists (LTRAs), especially pranlukast, safety data during pregnancy is limited. Therefore, we conducted a prospective, two‐centered cohort study using data from teratogen information services in Japan to clarify the effects of LTRA exposure during pregnancy on maternal and fetal outcomes. Pregnant women who being counseled on drug use during pregnancy at two facilities were enrolled. The primary outcome of this study was major congenital anomalies. The incidence of major congenital anomalies in women exposed to montelukast or pranlukast during the first trimester of pregnancy was compared with that of controls. Logistic regression analysis was performed to analyze the effects of maternal LTRA use during the first trimester of pregnancy on major congenital anomalies. The outcomes of 231 pregnant women exposed to LTRAs (montelukast n = 122; pranlukast n = 106; both n = 3) and 212 live births were compared with those of controls. The rate of major congenital anomalies in the LTRA group was 1.9%. Multivariable logistic regression analysis revealed that LTRA exposure was not a risk factor for major congenital anomalies (adjusted odds ratio, 0.78; 95% confidence interval, 0.23–2.05; p = 0.653). In addition, no significant difference was detected in stillbirth, spontaneous abortion, preterm birth, and low birth weight between the two groups. The present study revealed that montelukast and pranlukast were not associated with the risk of major congenital anomalies. Our findings suggest that LTRAs could be safely employed for asthma therapy during pregnancy.

Hatakeyama, Shiro↗

Search for Higgsinos in final states with low-momentum lepton-track pairs at 13 TeV

We present a search for the pair production of Higgsinos in final states with large missing transverse momentum and either two reconstructed muons or a reconstructed lepton (muon or electron) and an isolated track. The analyzed data correspond to proton-proton collisions with an integrated luminosity of 137 fb −1 , collected by the CMS experiment at $\sqrt{𝑠}$ =13 TeV in 2016, 2017, and 2018. The signal scenario assumes four nearly mass degenerate Higgsino mass eigenstates: two neutralino states $\tilde{𝜒}^0_2$ and $\tilde{𝜒}^0_1$ with a small mass difference in the range 1–10 GeV and two chargino states $\tilde{𝜒}^±_1$ with an intermediate mass. The analysis focuses on the decay of the heavier neutralino into the lighter one and a virtual 𝑍 boson, which decays into two same-flavor leptons. The leptons have small transverse momentum and/or a small opening angle between the identified muons. An isolated track is used to recover events in which only one of the two leptons is identified. Multivariate discriminants are used to enhance the sensitivity by efficiently rejecting backgrounds from SM processes or misreconstructed tracks and/or leptons. The search explores a unique phase space and probes a previously unexplored region of the signal model parameter space. Mass differences between the two neutralinos are probed down to 1.5 GeV, assuming a Higgsino mass of 100 GeV. The maximum excluded Higgsino mass is 115 GeV.

Hayrapetyan, A. [Yerevan Physics Institute]↗

Nonparametric Multiparticle Set Methods for Interpreting Environmental Samples

Collection and analysis of environmental samples is commonly used by a range of stakeholders in nuclear safeguards and security contexts. While the ubiquity of samples and their transport in the environment allow regular collection, developing and demonstrating methods for analyzing these samples is difficult. In this work, an environmental sample consists of a set of one or more individual particles. Recent advances in reactor simulation have allowed us to generate data that are more representative of real-world environmental samples, enabling statistically defensible method development and testing. The most notable of these advances is a drastic increase in the number of material depletion regions, which allows our simulations to capture the variation in isotopic composition seen at length scales consistent with environmental samples. Traditional approaches for handling multiparticle samples treat each particle in the sample individually, estimating the quantity of interest (e.g., core-average burnup) resulting from measurement and analysis of signatures (e.g., nuclide assays) from each individual particle. Individual estimates are then averaged to generate a single estimate of the quantity of interest over the entire sample. In this presentation, we introduce two novel approaches for interpreting environmental samples that comprise of multiple particles: (1) the Quantile-Quantile Comparator, which uses a multivariate generalization of quantile-quantile plots for comparing unknown statistical distributions, and (2) the Set Transformer, an attention-based neural network module designed to model interactions among elements (particles) in the input set (sample). Statistically representative sampling cannot be guaranteed as samples are passively collected and are beholden to what particles are available in the environment. These new analysis methods for set-input problems are expected to be more robust than traditional approaches to issues of sampling bias where particles are not uniformly distributed throughout regions of interest, as well as generally outperform traditional approaches by jointly considering all elements in the set. We will present results comparing the performance of traditional single particle approaches and the novel Quantile-Quantile Comparator and Set Transformer for interpretation of simulated environmental samples.

Phathanapirom, Birdy↗

National serosurvey and risk mapping reveal widespread distribution of Coxiella burnetii in Kenya

Coxiella burnetii, the causative agent of Q fever, is an emerging pathogen that has the potential to cause severe chronic infections in animals and humans worldwide. The detrimental impact on public health is projected to be higher in the low- and middle-income countries given their lower capacity to sustain effective surveillance and response measures. We implemented a national serosurvey of cattle in Kenya to map the spatial distribution of the pathogen. The study used serum samples that were collected from randomly selected cattle in different ago-ecological zones across the country. These samples were screened for the pathogen using PrioCHECK Ruminant Q Fever AB Plate ELISA kit. The laboratory findings were analyzed using INLA package to identify risk factors for C. burnetii exposure from herd- and animal-level factors, area, and bioclimatic datasets accessed from online databases. A total of 6,593 cattle were recruited for the study; of these, 7.9% (95% CI; 7.2–8.5) were seropositive. Outputs from the multivariable analysis revealed that the animal age and some of the geographical variables including wind speed, area under shrubs and “petric calcisols” type of soil were significantly associated with C. burnetii seropositivity. Being a calf, weaner or subadult was associated with lower odds of exposure compared to being an adult by 0.24 (credibility interval: 2.5% and 97.5%), 0.41 (0.30–0.55) and 0.51 (0.38–0.69), respectively. In addition, a unit increase in the wind speed increased the odds of C. burnetii seropositivity by 1.27 (1.05–1.52) while an increase on the land area under shrubs was associated with lower odds of exposure (0.67 [0.47–0.69]). The effect of petric calcisols was non-linear; an increase of the land area with this soil type was associated with an exponential increase in C. burnetii seropositivity. This study provides new data on C. burnetii seroprevalence, information of its risk factors and a prevalence map that can be used for C. burnetii risk surveillance and control. The identification of environmental risk factors for C. burnetii exposure, and the increasing awareness of the zoonotic potential of the pathogen, calls for the need to enhance the existing collaborations for the surveillance and control of C. burnetii in line with the One Health framework. The evidence generated on the potential role of environmental factors can also be used to design nature-based interventions, such as replacement of vegetation in denuded areas, to reduce potential for the aerosolization of the pathogen. Livestock vaccination in the hotspots would also reduce animal infections and hence the contamination of the environment.

60 APPLIED LIFE SCIENCES↗

High Throughput Data-Driven Design of Laser-Crystallized 2D MoS 2 Chemical Sensors: A Demonstration for NO 2 Detection

High throughput characterization and processing techniques are becoming increasingly necessary to navigate multivariable, data-driven design challenges for sensors and electronic devices. For two-dimensional materials, device performance is highly dependent upon a vast array of material properties including the number of layers, lattice strain, carrier concentration, defect density, and grain structure. In this work, laser crystallization was used to locally pattern and transform hundreds of regions of amorphous MoS 2 thin films into 2D 2H-MoS 2 . Here a high throughput Raman spectroscopy approach was subsequently used to assess the process-dependent structural and compositional variations for each illuminated region, yielding over 6000 distinct nonresonant, resonant, and polarized Raman spectra. The rapid generation of a comprehensive library of structural and compositional data elucidated important trends between structure–property processing relationships involving laser-crystallized MoS 2 , including the relationships between grain size, grain orientation, and intrinsic strain. Moreover, extensive analysis of structure/property relationships allowed for intelligent design and evaluation of major contributions to device performance in MoS 2 chemical sensors. In particular, it is found that NO 2 sensor performance is strongly dependent on the orientation of the MoS 2 grains relative to the crystal plane.

36 MATERIALS SCIENCE↗

A Review of Bayesian Networks for Spatial Data

We report Bayesian networks are a popular class of multivariate probabilistic models as they allow for the translation of prior beliefs about conditional dependencies between variables to be easily encoded into their model structure. Due to their widespread usage, they are often applied to spatial data for inferring properties of the systems under study and also generating predictions for how these systems may behave in the future. We review published research on methodologies for representing spatial data with Bayesian networks and also summarize the application areas for which Bayesian networks are employed in the modeling of spatial data. We find that a wide variety of perspectives are taken, including a GIS-centric focus on efficiently generating geospatial predictions, a statistical focus on rigorously constructing graphical models controlling for spatial correlation, as well as a range of problem-specific heuristics for mitigating the effects of spatial correlation and dependency arising in spatial data analysis. Special attention is also paid to potential future directions for integration of Bayesian networks with spatial processes.

97 MATHEMATICS AND COMPUTING↗

Yet Another Discriminant Analysis (YADA): A Probabilistic Model for Machine Learning Applications

This paper presents a probabilistic model for various machine learning (ML) applications. While deep learning (DL) has produced state-of-the-art results in many domains, DL models are complex and over-parameterized, which leads to high uncertainty about what the model has learned, as well as its decision process. Further, DL models are not probabilistic, making reasoning about their output challenging. In contrast, the proposed model, referred to as Yet Another Discriminate Analysis(YADA), is less complex than other methods, is based on a mathematically rigorous foundation, and can be utilized for a wide variety of ML tasks including classification, explainability, and uncertainty quantification. YADA is thus competitive in most cases with many state-of-the-art DL models. Ideally, a probabilistic model would represent the full joint probability distribution of its features, but doing so is often computationally expensive and intractable. Hence, many probabilistic models assume that the features are either normally distributed, mutually independent, or both, which can severely limit their performance. YADA is an intermediate model that (1) captures the marginal distributions of each variable and the pairwise correlations between variables and (2) explicitly maps features to the space of multivariate Gaussian variables. Numerous mathematical properties of the YADA model can be derived, thereby improving the theoretic underpinnings of ML. Validation of the model can be statistically verified on new or held-out data using native properties of YADA. However, there are some engineering and practical challenges that we enumerate to make YADA more useful.

97 MATHEMATICS AND COMPUTING↗

EPIsembleVis: A geo-visual analysis and comparison of the prediction ensembles of multiple COVID-19 models

In this work, we present EPIsembleVis, a web-based comparative visual analysis tool for evaluating the consistency of multiple COVID-19 prediction models. Our approach analyzes a collection of COVID-19 predictions from different epidemiological models as an ensemble and utilizes two metrics to quantify model performance. These metrics include (a) prediction uncertainty (represented as the dispersion of predictions in each ensemble) and (b) prediction error (calculated by comparing individual model predictions with the recorded data). Through an interactive visual interface, our approach provides a data-driven workflow for (a) selecting and constructing the COVID-19 model prediction ensemble based on the spatiotemporal overlap of available predictions of multiple epidemiological models, (b) quantifying the model performance using both the uncertainty of each model prediction ensemble, and the error of each ensemble member that represents individual model predictions, and (c) visualizing the spatiotemporal variability in the projection performance of individual models using a suite of novel ensemble visualization techniques, such as the data availability map, a spatiotemporal textured-tile calendar, multivariate rose chart, and time-series leaflet glyph. We demonstrate the capability of our ensemble visual interface through a case study that investigates the performance of weekly COVID-19 predictions, which are provided through the COVID-19 Forecast Hub UMass-Amherst Influenza Forecasting Center of Excellence [47] for the United States and United States Territories. The EPIsembleVis tool is implemented using open-source web technologies and adaptive system design, rendering it interoperable with Elasticsearch and Kibana for automatically ingesting COVID-19 predictions from online repositories, and it is generalizable for analyzing worldwide projections from more epidemiological models.

60 APPLIED LIFE SCIENCES↗

Long-Term Statistical Process Monitoring of an Ultrafiltration Water Treatment Process

As water treatment technology has improved, the amount of available process data has substantially increased, making real-time, data-driven fault detection a reality. One shortcoming of the fault detection literature is that methods are usually evaluated by comparing their performance on hand-picked, short-term case studies, which yields no insight into long-term performance. In this work, we first evaluate multiple statistical and machine learning approaches for detrending process data. Then, we evaluate the performance of a PCA-based fault detection approach, applied to the detrended data, to monitor influent water quality, filtrate quality, and membrane fouling of an ultrafiltration membrane system for indirect potable reuse. Based on two short case studies, the adaptive lasso detrending method is selected, and the performance of the multivariate approach is evaluated over more than a year. The method is tested for different sets of three critical tuning parameters, and we find that for long-term, autonomous monitoring to be successful, these parameters should be carefully evaluated. However, in comparison with industry standards of simpler, univariate monitoring or daily pressure decay tests, multivariate monitoring produces substantial benefits in long-term testing.

ammonia↗