Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “multivariate data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Reliable machine prognostic health management in the presence of missing data

Prognostics and health management enables the prediction of future degradation and remaining useful life (RUL) for in-service systems based on historical and contemporary data, showing promise for many practical applications. One major challenge for prognostics is the common occurrence of missing values in time-series data, often caused by disruptions in sensor communication or hardware/software failures. Another major concern is that the sufficient prior knowledge of critical component degradation with a clear failure threshold is often not readily available in practice. These issues can significantly hinder the application of advanced signal and data analysis methods and consequently degrade the health management performance. In this article, we propose a novel data-driven framework that is capable of providing accurate and reliable predictions of degradation and RUL. In this approach, one-hot health state indicators are appended to the historical time series so that the model learns end-of-life automatically. A modified gate recurrent unit based variational autoencoder is employed in generative adversarial networks to model the temporal irregularity of the incomplete time series. Furthermore, experiments on multivariate time-series datasets collected from real-world aeroengines verify that significant performance improvement can be achieved using the proposed model for robust long-term prognostics.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Quantification of the REE 3+ aqua ions and chloride species in aqueous fluids by in situ Raman spectroscopy using perturbations of the water band

Acidic NaCl-rich aqueous fluids play a crucial role in forming hydrothermal rare earth elements (REE) mineral deposits. Aqueous REE mobility is mostly controlled by the stabilities of REE 3+ and REE chloride species. Our current knowledge of REE speciation is based on solubility data, thermodynamic models and in situ spectroscopic measurements, sometimes coupled with molecular simulations. Here, in this study, we investigate Nd and Yb speciation in pH2 Cl-bearing solutions at 25 °C and 0.1 MPa with variable Cl/REE ratios using Raman Spectroscopy in solutions with 0.1 to 0.6 mol/kg NdCl 3 or YbCl 3 and 0.2 to 3.2 mol/kg NaCl. Due to the challenges in resolving the REE-Cl band, we developed a new method using the water vibrational mode and multivariate curve resolution (MCR) analysis. The Raman spectra for the vibrational band of water (2700 to 3900 cm –1 ) were collected at 25 °C and fitted by three Gaussian sub-peaks, then quantified using MCR analysis to de-convolute the water band into bulk H 2 O and the perturbations caused by of Cl – , REE 3+ , and REE chloride species. REE speciation based on the perturbations of the water band indicates that REE 3+ aqua ions dominate acidic solutions at 25 °C, but up to ~20 mol% YbCl 2+ forms at high YbCl 3 concentrations. The new method is promising for quantifying in situ speciation of the REE 3+ aqua ions and REE chloride species in aqueous fluids while providing information on the hydration of ions. This method improves our molecular level understanding of REE aqueous species stability and their role in REE mobilization during fluid-rock interaction.

58 GEOSCIENCES↗

Advanced Visualization for Scientific Data Analysis and Insight [Slides]

This talk will explore how we have used advanced visualization technologies to support analytical reasoning and knowledge discovery. Specifically, we will present several examples detailing some recent scientific successes using state-of-the-art immersive and high-resolution visualization at the National Renewable Energy Laboratory's Computational Science Center. On multiple occasions, we have observed scientists and engineers discover features in their data using advanced visualization technologies that they had not seen in prior investigations of their data on traditional desktop displays. We have embedded more information into our analytics tools, allowing engineers to explore complex multivariate spaces. We have observed how interactions seem to catalyze understanding.

97 MATHEMATICS AND COMPUTING↗

Evidence for top quark production in nucleus-nucleus collisions

We report droplets of quark-gluon plasma (QGP), an exotic state of strongly interacting quantum chromodynamics (QCD) matter, are routinely produced in heavy nuclei high-energy collisions. Although the experimental signatures marked a paradigm shift away from expectations of a weakly coupled QGP, a challenge remains as to how the locally deconfined state with a lifetime of a few fm can be resolved. The only colored particle that decays mostly within the QGP is the top quark. Here we demonstrate, for the first time, that top quark decay products are identified, irrespective of whether interacting with the medium (bottom quarks) or not (leptonically decaying W bosons). Using 1.7±0.1 nb -1 of lead-lead (A = 208) collision data recorded by the CMS experiment at a nucleon-nucleon center-of-mass energy of 5.02 TeV, we report evidence of top quark pair ($t\bar{t}$) production. Dilepton final states are selected, and the cross section ($σ_{t\bar{t}}$) is measured from a likelihood fit to a multivariate discriminator using lepton kinematic variables. The $σ_{t\bar{t}}$ measurement is additionally performed considering the jets originating from the hadronization of bottom quarks, which improve the sensitivity to the $t\bar{t}$ signal process. After background subtraction and analysis corrections, the measured $σ_{t\bar{t}}$ is 2.56 ± 0.82(tot) and 2.02 ± 0.69(tot)μb in the two cases, respectively, consistent with predictions from perturbative QCD.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Pregnancy outcome after first trimester exposure to domperidone—An observational cohort study

Abstract Aim To assess the teratogenic risk of domperidone by comparing the incidence of major malformation with domperidone to a control. Methods Pregnancy outcome data were obtained for women at two Japanese facilities that provide counseling on drug use during pregnancy between April 1988 and December 2017. The incidence of major malformation was calculated among infants born to women taking domperidone ( n = 519), nonteratogenic drugs (control, n = 1673), or metoclopramide (reference, n = 241) during the first trimester of pregnancy. Using the control group as reference, the crude odds ratio (OR) of the incidence of major malformation in the domperidone and metoclopramide groups was calculated using univariable logistic regression analysis. Adjusted OR was also calculated using multivariable logistic regression analysis adjusted for various other factors. Results The incidence of major malformation was 2.9% (14/485, 95% confidence interval [CI]: 1.6–4.8) in the domperidone group, 1.7% (27/1554, 95%CI: 1.1–2.5) in the control group, and 3.6% (8/224, 95%CI: 1.6–6.9) in the metoclopramide group. The adjusted multivariable logistic regression analysis showed no significant difference in incidence between the control and domperidone groups (adjusted OR: 1.86 [95%CI: 0.73–4.70], p = 0.191) or between the control and metoclopramide groups (adjusted OR: 2.20 [95%CI: 0.69–6.98], p = 0.183). Conclusions This observational cohort study showed that domperidone exposure during the first trimester was not associated with increased risk of major malformation in infants. These results may help alleviate the anxiety of patients who took domperidone during pregnancy.

Hishinuma, Kayoko↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

Assessment of Demographic, Genetic, and Imaging Variables Associated With Brain Resilience and Cognitive Resilience to Pathological Tau in Patients With Alzheimer Disease

Importance: Better understanding is needed of the degree to which individuals tolerate Alzheimer disease (AD)-like pathological tau with respect to brain structure (brain resilience) and cognition (cognitive resilience). Objective: To examine the demographic (age, sex, and educational level), genetic (APOE-ε4 status), and neuroimaging (white matter hyperintensities and cortical thickness) factors associated with interindividual differences in brain and cognitive resilience to tau positron emission tomography (PET) load and to changes in global cognition over time. Design, setting, an participants: In this cross-sectional, longitudinal study, tau PET was performed from June 1, 2014, to November 30, 2017, and global cognition monitored for a mean [SD] interval of 2.0 [1.8] years at 3 dementia centers in South Korea, Sweden, and the United States. The study included amyloid-β-positive participants with mild cognitive impairment or AD dementia. Data analysis was performed from October 26, 2018, to December 11, 2019. Exposures: Standard dementia screening, cognitive testing, brain magnetic resonance imaging, amyloid-β PET and cerebrospinal fluid analysis, and flortaucipir (tau) labeled with fluor-18 (18F) PET. Main outcomes and measures: Separate linear regression models were performed between whole cortex [18F]flortaucipir uptake and cortical thickness, and standardized residuals were used to obtain a measure of brain resilience. The same procedure was performed for whole cortex [18F]flortaucipir uptake vs Mini-Mental State Examination (MMSE) as a measure of cognitive resilience. Bivariate and multivariable linear regression models were conducted with age, sex, educational level, APOE-ε4 status, white matter hyperintensity volumes, and cortical thickness as independent variables and brain and cognitive resilience measures as dependent variables. Linear mixed models were performed to examine whether changes in MMSE scores over time differed as a function of a combined brain and cognitive resilience variable. Results: A total of 260 participants (145 [55.8%] female; mean [SD] age, 69.2 [9.5] years; mean [SD] MMSE score, 21.9 [5.5]) were included in the study. In multivariable models, women (standardized β = -0.15, P = .02) and young patients (standardized β = -0.20, P = .006) had greater brain resilience to pathological tau. Higher educational level (standardized β = 0.23, P < .001) and global cortical thickness (standardized β = 0.23, P < .001) were associated with greater cognitive resilience to pathological tau. Linear mixed models indicated a significant interaction of brain resilience × cognitive resilience × time on MMSE (β [SE] = -0.235 [0.111], P = .03), with steepest slopes for individuals with both low brain and cognitive resilience. Conclusions and relevance: Results of this study suggest that women and young patients with AD have relative preservation of brain structure when exposed to neocortical pathological tau. Interindividual differences in resilience to pathological tau may be important to disease progression because participants with both low brain and cognitive resilience had the most rapid cognitive decline over time.

60 APPLIED LIFE SCIENCES↗

Analysis of Inlier and Outlier Compounds with Respect to Artificial Neural Network Cetane Number Prediction Accuracy

Artificial neural networks (ANNs) are exceptional at forming non-linear correlations between multivariate input and target variables; however, they are often seen as a “black box” approach, since how ANNs form these correlations is somewhat ambiguous. Furthermore, the process underlying how ANNs learn from inlier and outlier samples within the input dataset is not fully understood. Intuitively, it is expected that training ANNs with inlier samples will increase prediction accuracy and training with outlier samples will reduce prediction accuracy; though, in practice, this is not always true. The present work identifies and analyzes inliers and outliers of existing experimental cetane number (CN) data encompassing a variety of compounds and compound groups. It also investigates how ANNs trained to predict CN perform with and without outliers included in the training data, and whether a relationship exists between inliers/outliers and ANN prediction accuracy across the whole dataset and for individual samples. Additionally, individual outlier compounds are analyzed, highlighting how they structurally differ from inlier compounds.

09 BIOMASS FUELS↗

Emergency department overcrowding and its associated factors at HARME medical emergency center in Eastern Ethiopia

Introduction: Emergency department (ED) overcrowding has become a significant concern as it can lead to compromised patient care in emergency settings. Various tools have been used to evaluate overcrowding in ED. However, there is a lack of data regarding this issue in resource-limited countries, including Ethiopia. This study aimed to validate NEDOCS, assess level of ED overcrowding and identify associated factors at HARME Medical Emergency Center, located in Hiwot Fana Comprehensive Specialized Hospital, Harar, Ethiopia. Methods: A cross-sectional study was conducted at the HARME Medical Emergency Center, Hiwot Fana Comprehensive Specialized Hospital, involving a total of 899 patients during 120 sampling intervals. The area under the receiver operating characteristic curves (AUC) was calculated to evaluate the agreement between objective and subjective assessments of ED overcrowding. A multivariable logistic regression analysis was employed to identify factors associated with ED overcrowding and statistically significant association was declared using 95% confidence level and a p-value < 0.05. Results: The interrater agreement showed a strong correlation with a Cohen's kappa (κ) of 0.80. The National Emergency Department Overcrowding Study Score demonstrated a strong association with subjective assessments from residents and case team nurses, with an AUC of 0.81 and 0.79, respectively. According to residents' perceptions, ED were considered overcrowded 65.8% of the time. Factors significantly associated with ED overcrowding included waiting time for triage (AOR: 2.24; 95% CI: 1.54–3.27), working time (AOR: 2.23; 95% CI: 1.52–3.26), length of stay (AOR: 2.40; 95% CI: 1.27–4.54), saturation level (AOR: 2.35; 95% CI: 1.31–4.20), chronic illness (AOR: 2.19; 95% CI: 1.37–3.53), and abnormal pulse rate (AOR: 1.52; 95% CI: 1.06–2.16). Conclusion: The study revealed that ED were overcrowded approximately two-thirds of the time.

60 APPLIED LIFE SCIENCES↗

teemi: An open-source literate programming approach for iterative design-build-test-learn cycles in bioengineering

Synthetic biology dictates the data-driven engineering of biocatalysis, cellular functions, and organism behavior. Integral to synthetic biology is the aspiration to efficiently find, access, interoperate, and reuse high-quality data on genotype-phenotype relationships of native and engineered biosystems under FAIR principles, and from this facilitate forward-engineering strategies. However, biology is complex at the regulatory level, and noisy at the operational level, thus necessitating systematic and diligent data handling at all levels of the design, build, and test phases in order to maximize learning in the iterative design-build-test-learn engineering cycle. To enable user-friendly simulation, organization, and guidance for the engineering of biosystems, we have developed an open-source python-based computer-aided design and analysis platform operating under a literate programming user-interface hosted on Github. The platform is called teemi and is fully compliant with FAIR principles. In this study we apply teemi for i) designing and simulating bioengineering, ii) integrating and analyzing multivariate datasets, and iii) machine-learning for predictive engineering of metabolic pathway designs for production of a key precursor to medicinal alkaloids in yeast. The teemi platform is publicly available at PyPi and GitHub.

59 BASIC BIOLOGICAL SCIENCES↗

Advanced Method Optimization for Sampling and Analysis Instrumentation

This work presents a generalized approach for analytical method optimization that branches the gap between techniques historically employed and accurate modern optimization techniques suitable for various applications. The novelty of the described strategy is the utilization of multivariate, multiobjective optimization with Karush-Kuhn-Tucker conditions to bound the optimization space to solutions within the physical limitations of instrumentation. Briefly, the basic steps outlined in this paper are to (1) determine the objective(s) that should be maximized or minimized based on the goals of the analytical application, (2) conduct a screening experiment, (3) perform ANOVA to determine the parameters which have a statistically significant effect on the objective, (4) conduct an experiment (e.g., Box-Behnken design) to collect data for fitting the objective equation, and (5) determine the physical constraints of the parameters and solve the Lagrangian to determine the optimal method parameters. A broad approach to optimization target selection allows for robust method tuning to develop improved data sets amenable for chemometrics and machine learning algorithm development. Gas chromatography-mass spectrometry was selected as a use case due to its broad use across scientific fields and time-consuming method development involving numerous parameters. In conclusion, this strategy can reduce the cost of research, improve data quality, and enable the rapid development of new analytical technique.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Validation of Stage N3 of the Eighth Edition AJCC Staging System for Nasopharyngeal Carcinoma

Objectives To validate stage nodal (N)3 of the 8th edition American Joint Committee on Cancer (AJCC) staging system for nasopharyngeal carcinoma (NPC). Methods This retrospective cohort study extracted NPC patients from the Surveillance, Epidemiology, and End Results database between 2004 and 2016. Pathologically confirmed patients with complete data of level IV, N3a, and N3b lymph node metastasis were investigated. The included patients were divided into level IV, N3a, and N3b groups. Five‐year overall survival (OS) and cancer‐specific survival (CSS) were compared among the three groups. Results A total of 693 patients were included: 285 (41.13%) patients in the level IV group, 124 (17.89%) patients in the N3a group, and 284 (40.98%) patients in the N3b group. The 5‐year OS (57.1%, 55.0%, and 55.2%) and CSS (64.4%, 63.5%, and 64.4%) were not different among the level IV, N3a, and N3b groups. Multivariate regression analysis revealed that N stage was not an independent prognostic factor for OS (hazard ratio [HR] = 1.03, 95% confidence interval [CI]: 0.91–1.17; P = .65) or CSS (HR = 1.03, 95% CI: 0.89–1.19; P = .70). Conclusion Stage N3 of the 8th edition AJCC staging system for NPC is reasonable. Level of Evidence III Laryngoscope , 131:535–540, 2021

Pan, Xin‐Bin↗

Surface-Enhanced Raman Spectroscopy Combined with Multivariate Analysis for Fingerprinting Clinically Similar Fibromyalgia and Long COVID Syndromes

Fibromyalgia (FM) is a chronic central sensitivity syndrome characterized by augmented pain processing at diffuse body sites and presents as a multimorbid clinical condition. Long COVID (LC) is a heterogenous clinical syndrome that affects 10–20% of individuals following COVID-19 infection. FM and LC share similarities with regard to the pain and other clinical symptoms experienced, thereby posing a challenge for accurate diagnosis. This research explores the feasibility of using surface-enhanced Raman spectroscopy (SERS) combined with soft independent modelling of class analogies (SIMCAs) to develop classification models differentiating LC and FM. Venous blood samples were collected using two supports, dried bloodspot cards (DBS, n = 48 FM and n = 46 LC) and volumetric absorptive micro-sampling tips (VAMS, n = 39 FM and n = 39 LC). A semi-permeable membrane (10 kDa) was used to extract low molecular fraction (LMF) from the blood samples, and Raman spectra were acquired using SERS with gold nanoparticles (AuNPs). Soft independent modelling of class analogy (SIMCA) models developed with spectral data of blood samples collected in VAMS tips showed superior performance with a validation performance of 100% accuracy, sensitivity, and specificity, achieving an excellent classification accuracy of 0.86 area under the curve (AUC). Amide groups, aromatic and acidic amino acids were responsible for the discrimination patterns among FM and LC syndromes, emphasizing the findings from our previous studies. Overall, our results demonstrate the ability of AuNP SERS to identify unique metabolites that can be potentially used as spectral biomarkers to differentiate FM and LC.

60 APPLIED LIFE SCIENCES↗

Reduction in total and major cause-specific mortality from tobacco smoking cessation: a pooled analysis of 16 population-based cohort studies in Asia

Little is known about the time course of mortality reduction following smoking cessation in Asians who have smoking behaviours distinct from their Western counterparts. We evaluated the level of reduction in all-cause, cardiovascular disease (CVD) and lung cancer mortality by years since quitting smoking, in Asia. Using Cox regression, we analysed individual participant data (n = 709 151) from 16 prospective cohorts conducted in China, Japan, Korea/Singapore, and India/Bangladesh, separately by cohorts. Cohort-specific hazard ratios (HRs) were combined using a random-effects meta-analysis. During a mean follow-up of 12.0 years, 108 287 deaths were ascertained—35 658 from CVD and 7546 from lung cancer. Among Asian men, a dose-response relationship of risk reduction in deaths from all causes, CVD and lung cancer was observed with an increase in years after smoking cessation. Compared with never smokers, however, all-cause and CVD mortality among former smokers remained elevated 10–14 years after quitting [multivariable-adjusted HR (95% confidence interval (CI) = 1.25 (1.13–1.37) and 1.20 (1.02–1.41), respectively]. Lung cancer mortality stayed almost 2-fold higher than among never smokers 15–19 years after smoking cessation [1.97 (1.41–2.73)], particularly among former heavy smokers [2.62 (1.71–4.00)]. Women who quitted for ≥5 years retained a significantly elevated mortality from all causes, CVD and lung cancer. Overall patterns of the cessation-mortality associations were similar across countries. Our findings suggest that adverse effects of tobacco smoking persist for an extended time period, even for more than two decades, which is beyond the time windows defined in current clinical guidelines for risk assessment of lung cancer and CVD.

Public, Environmental & Occupational Health↗

Utilizing the Dynamic Networks Data Processing and Analysis Experiment (DNE18) to Establish Methodologies for the Comparison of Automatic Infrasonic Signal Detectors

The Dynamic Networks Experiment 2018 (DNE18) was a collaborative effort between Los Alamos National Laboratory (LANL), Sandia National Laboratories (SNL), Lawrence Livermore National Laboratory (LLNL) and Pacific Northwest National Laboratory (PNNL) designed to evaluate methodologies for multi-modal data ingestion and processing. One component of this virtual experiment was a quantitative assessment of current capabilities for infrasound data processing, beginning with the establishment of a baseline for infrasound signal detection. To produce such baselines, SNL and LANL exploited a common dataset of infrasound data recorded across a regional network in Utah from December 2010 through February 2011. We utilize two automated signal detectors, the Adaptive F-Detector (AFD) and the Multivariate Adaptive Learning Detector (MALD) to produce automated signal detection catalogs and an analyst-produced catalog. Comparisons indicate that automatic detectors may be able to identify small amplitude, low SNR events that cannot be identified by analyst review. We document detector performance in terms of precision and recall, demonstrating that the AFD is more precise, but the MALD has higher recall. We use a synthetic dataset of signals embedded in pink noise in order to highlight shortcomings in assessing detection algorithms for low signal to noise ratio signals which are commonly of interest to the nuclear monitoring community. For comparisons utilizing the synthetic dataset, the AFD has higher recall while precision is equal for both detectors. These results indicate that both detectors perform well across a variety of background noise environments; however, both detectors fail to identify repetitive, short duration signals arriving from similar backazimuths. These failures represent specific scenarios that could be targeted for further detector development.

97 MATHEMATICS AND COMPUTING↗

Distribution Development for the RDX Regional Model at Los Alamos National Laboratory - 20385

Representing uncertainty in model inputs often means finding a balance between uncertainty and physical reality. Developing wide distributions may seem conservative in principle, but this approach may lead to unrealistic model results. Characterizing the current state of knowledge of stochastic inputs presents many challenges, especially if the data available are limited or have limited relevance to the site. Relationships among these inputs may also be important to represent but are typically complex or difficult to define. Often special adjustments must be made to account for reduced credibility in particular data. If parameters are strongly related to one another, a correlation structure may be developed for input to the model. Other techniques such as regression models may be used to incorporate relationships between the information available and the desired parameters. The process of developing distributions must consider details of the model in terms of what the distribution is meant to represent. This paper uses the example of a probabilistic fate and transport model for hexahydro-1,3,5-trinitro-1,3,5-triazine (RDX) in the regional aquifer at Los Alamos National Laboratory (LANL). For many parameters, a single draw is applied to all space and time over which the model is run, for a single iteration. This simplification is often made for many reasons, and can often be beneficial, but also adds additional complexity in the distribution development process. Defining the distributional goals as they relate to the modeling process is an important step, which should take place prior to evaluation of the data. Usually, the distributions developed are meant to characterize the average value of the parameter over the spatial and temporal domain of the model. Distribution development requires consideration of many sources of information on the parameter where available, ideally from multiple references. Examples of different sources include data from different references but also from different conditions, measurement methods, or experimental types. Depending on these conditions and the reliability or relevance of particular references, different sources of data may each contribute valuable information but have varying relevance to the site. In these cases, weighting data unequally is a useful way to incorporate this information. As an example, aqueous dispersivity data are available for a variety of rock types. Only a few values are available for the desired rock type, and this is not enough to develop a distribution. Therefore, dispersivity values from other rock types are included in distribution development but are down-weighted such that the best data have the most influence on the distribution developed. In another case, K{sub d} distributions in the model are meant to represent a known composition of multiple soil types. Data from these materials are weighted accordingly to develop a distribution for the weighted average K{sub d} across soil types. Other cases include varying reliability of different sources, and weighting data according to the confidence in these sources. In some cases, input parameters are correlated with one another. An example is advective porosity, which is positively correlated with total porosity and must be less than total porosity. Paired data with both parameters must exist to discern the relationship between the two parameters and if it is necessary to build a correlation structure into the model. In general, correlated parameters may be represented in the model by a multivariate distribution, or perhaps more desirably, capturing the correlation within the developed distributions which may be treated as independent from one another. In the example of porosity, this can be done by transforming advective porosity into a proportion of total porosity which may be drawn independently from the distribution of total porosity. This paper explores how complex data and correlations can be incorporated to meet the distributional goals of the model, using the RDX regional model as a detailed example. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative data sets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student t process regression. We apply Student t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via colouring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics and classic machine learning clustering algorithms.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗