Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Operator regression”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Improving Prediction of Surface Solar Irradiance Variability by Integrating Observed Cloud Characteristics and Machine Learning

A 5-year, 1-minute resolution observational dataset of clouds and solar radiation was produced that includes two metrics of the variability in surface solar irradiance due to cloud type and fractional sky cover. Multiple regression models were trained to fit observations of surface solar irradiance variability from those two cloud property predictors. We found that ensemble tree-based methods, Random Forest and Gradient Boosting Machine, have the least overfitting issues and showed the best performance with an R2 of 0.42. While the observational data trained in this study was only from one site, the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Southern Great Plains (SGP) site in Oklahoma, initial comparisons of the seasonality of the statistics suggest that these results are relatively weather regime independent; the generality of such a finding across sites will be tested in future work. The observational data and developed machine learning model are being used to create a numerical weather prediction model parameterization to enable day-ahead solar variability prediction in a computationally efficient way. This is a first step towards creating a new paradigm of predicting day-ahead variability with the potential to provide a new tool to improve grid operation, planning, and resilience.

Riihimaki, Laura↗

Examples of Mission-driven Data Science from Jefferson Lab and ACES

This presentation details mission-driven data science initiatives at Jefferson Lab and the Joint Institute for Advanced Computing on Environmental Studies (ACES). JLab, a U.S. Department of Energy Office of Science national laboratory, operates the Continuous Electron Beam Accelerator Facility (CEBAF), and is the lead institute for the new High Performance Data Facility (HPDF) Hub. The Joint Institute for ACES brings together interdisciplinary teams in health informatics, climate modeling, computer science, and physics to address environmental challenges, including flood modeling. The Hampton Roads region, particularly Norfolk and Virginia Beach, faces increasing flood risks, motivating the need for rapid, reliable, and risk-aware decision support. ACES’s flooding work has a focus on uncertainty quantification (UQ) and machine learning (ML) for coastal flood management. The work is motivated by the increasing vulnerability of communities such as Norfolk and Virginia Beach, Virginia, to frequent coastal flooding events, and the need for rapid, reliable decision support. The research develops computationally efficient ML surrogate models to forecast water levels and flooding risk. A central theme is the quantification and calibration of predictive uncertainty, especially for out-of-distribution (OOD) scenarios, using techniques such as Monte Carlo Dropout, Deep Ensembles, Gaussian Processes, and Deep Quantile Regression (DQR). The study demonstrates that distance-aware UQ is critical for reliable scientific AI, particularly in high-dimensional, safety-critical, and real-time applications.

McSpadden, Diana [Thomas Jefferson National Accele↗

Stability Analysis of Substituted Cobaltocenium [Bis(cyclopentadienyl)cobalt(III)] Employing Chemistry-Informed Neural Networks

Cobaltocenium derivatives are promising components of the anion exchange membranes due to their excellent thermal and alkaline stability under the operating conditions of a fuel cell. Here we present an efficient modeling approach of assessing the chemical stability of substituted cobaltocenium CoCp 2 + based on the computed electronic structure enhanced by machine learning techniques. Within the aqueous environment, the positive charge of the metal cation is balanced by the hydroxide anion through formation of the CoCp 2 + OH¯ complexes, whose dissociation is studied within the implicit solvent employing density functional theory. The data set of about 118 species based on 42 substituent groups characterized by a range of electron- donating (ED) and electron-withdrawing (EW) properties is constructed and analyzed. Given 12 carefully chosen chemistry-informed descriptors of the complexes and relevant fragments, the stability of the complexes is found to strongly correlate with the energies of the highest occupied and lowest unoccupied molecular orbitals, modulated by a switching function of the Hirshfeld charge. The latter is used as a measure of the electron withdrawing-donating character of the substituents. Based on this observation from the conventional regression analysis, two fully connected, feed-forward neural network (FNN) models with different unit structures, called the chemistry-informed (CINN) and the quadratic (QNN) neural networks, are developed. Both models predict the bond dissociation energies of the cobaltocenium complexes with mean relative errors less than 5.40% and average absolute errors less than 0.94 kcal/mol. The results show the potential of QNN to efficiently capture more complex relationships. Here, the concept of incorporating the domain (chemical) knowledge/insight into the neural network structure paves the way to applications of machine learning techniques with small data sets, ultimately leading to better predictive models compared to the conventional regression analysis.

36 MATERIALS SCIENCE↗

Probabilistic Day-Ahead Forecasting Using an Analog Ensemble Approach for Wind Farm Grid Services

Wind resource assessment and wind power forecasting are used in research and industry to anticipate future power output at scales ranging from individual wind turbines to entire wind farms. Probabilistic day-ahead wind forecasting is useful for anticipating how a wind farm could potentially participate in the day-ahead market by providing upper and lower bounds for expected power generation, thus informing grid operators of its uncertainty. Understanding this uncertainty is part of a larger project focused on building a platform that combines efforts in weather forecasting, aerodynamic and economic modeling to create maximum value of a wind plant to better provide services to the grid. This effort is also known as the Atmosphere to Electrons to Grid (A2E2G) project. One method for producing a probabilistic forecast is through the analog ensemble approach (Delle Monache et al., 2011). This method leverages historical forecasts and their corresponding observations as a training data set from which future forecasts can be made. For some future forecast, the most similar historical forecasts (analogs) are identified on a regular time basis such as once per a 3-hour window. The most similar analogs, based on a metric such as root mean square error (RMSE), are recorded and their corresponding verifying observations are used as an ensemble member for this future forecast. Prior work in this area demonstrates improvements over raw Numerical Weather Prediction (NWP) forecasts and shows skill similar to techniques such as logistic regression and machine learning (Delle Monache et al., 2013; Alessandrini et al., 2015). Here, we take the High-Resolution Rapid Refresh model (HRRR) day-ahead forecast (0-36 hours) to create a probabilistic day-ahead forecast using an analog ensemble approach. The HRRR has an hourly temporal resolution, with a spatial resolution of 3 km. The 12 UTC HRRR model run is downloaded every day for one year from August 2019 - July 2020, with the first 11 months serving as a bank of analogs from which the forecasting algorithm can create a probabilistic forecast. Once downloaded, the original HRRR forecast is temporally interpolated to 5-minutes, aligning with both the temporal resolution of the observations as well as the timescale relevant for day-ahead power forecasts. The forecast is validated at the M2 tower at the Flatirons Campus of the National Renewable Energy Laboratory (NREL) at a typical wind turbine height of 80 m. Variables such as wind speed, wind direction, and turbulence intensity are incorporated into the probabilistic forecast model and weighted according to their relative importance to the forecast. Based on metrics such as mean bias error (MBE), mean absolute error (MAE), and root mean square error, the analog ensemble forecast outperforms the raw HRRR forecast during the testing period of July 2020. Figure 1 illustrates an example day-ahead forecast compared against the verifying observations. The general variability and ramps are captured throughout the day, with potential to further improve the analog ensemble model through machine learning techniques.

numerical weather prediction↗

Occupational Experience Effects on Physiological and Perceptual Responses of Common Soldiering Tasks

Abstract Cohen BS, Redmond JE, Haven CC, Foulis SA, Canino MC, Frykman PN, Sharp MA. Occupational Experience Effects on Physiological and Perceptual Responses of Common Soldiering Tasks. J Strength Cond Res 37(4): 894–901, 2023—This study measured the impact of occupational experience (i.e., time spent deployed, in military service, and in job and task performance frequency in training, deployment, and study practice) on the physiological (heart rate [HR] and oxygen consumption [VO 2 ]) and perceptual (rate of perceived exertion [RPE]) responses to performance of critical physically demanding tasks (CPDTs). Five CPDTs (road march, build a fighting position, move under fire, evacuate a casualty, and drag a casualty to safety), common to all soldiers, were performed by 237 active duty soldiers. Linear regression models examined the association between measures of experience and physiological and perceptual performance responses to task demands. The level of significance was adjusted for multiple comparisons and set at ρ ≤ 0.0125 for this study. Significant and notable effect sizes included the impact of time spent deployed on the physiological measures of the road march (PostHR F = 24.84, p < 0.0001, β=-9.65), sandbag fill (PostHR F = 8.26, p = 0.005, β = −2.83), and sandbag carry (MeanHR F = 7.51, p = 0.007, β = −1.12; PostHR F = 7.35, p = 0.007, β = −0.87). For the road march task, there was a nearly 10 bpm decrease in postperformance HR for every year spent deployed. Road march, sandbag fill, and sandbag carry tasks PostHRs were also notably negatively associated with the experience measures of time in their MOS (job and time in military service but not for other physiological and perceptual responses, including VO 2 and RPE. Frequency of task performance in training, deployment, and study practice was not meaningfully associated with experience. The results suggest that increasing task familiarization through on-the-job occupational operational experience may result in greater proficiency and reduced physiological effort.

Sport Sciences↗

Machine-Learning Enabled Evaluation of Probability of Piping Degradation In Secondary Systems of Nuclear Power Plants

The transition to condition-based, risk-informed automated maintenance will contribute to a significant reduction of operations and maintenance costs that account for the majority of nuclear power generation costs. Furthermore, of the operations and maintenance costs in U.S. plants, approximately 80% are labor costs. To address the issue of rising operating costs and economic viability, technologies used to perform online monitoring of piping and other secondary system structural components in commercial nuclear power plants (NPPs) are under evaluation. These online monitoring systems have the potential to identify when a more detailed inspection is needed using real time measurements, rather than at a pre-determined inspection interval thus reducing the maintenance cost. This paper describes distributed high-temperature stable fiber sensors fabricated in optical fibers through a roll-to-roll laser direct writing process using femtosecond lasers. Using phase-sensitive optical time domain reflectometry, distributed acoustic and vibration sensors can be developed and deployed to critical components and systems in NPPs to perform active measurements with spatial resolution down to 0.5-meter throughout the piping systems. Complex acoustic and vibration signatures harnessed by distributed sensors are registered and analyzed by artificial intelligence algorithms for degradation detection and flaw identification. Piping elbows with machined-in flaws were instrumented with fiber sensors. High-spatial-resolution data were used to develop and validate machine learning algorithms, including both linear and nonlinear regression, and classification. Additionally, classification and sensor analysis were also performed for data analysis. The paper concludes with recommendations and future work on applications of machine learning enabled high-resolution fiber sensors for piping degradation monitoring in current or future NPPs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

CO 2 zonal injection rate allocation and plume extent evaluation through wellbore temperature analysis

Temperature analysis during a pause in injection operations, known as warmback analysis, has been used in the petroleum industry for evaluating the injection conformance and estimating the location of the flooded front in applications, such as waterflooding oil reservoirs. Here in this work, methods are introduced to extend the application of temperature warmback analysis to estimate the zonal CO 2 injection rate and zonal CO 2 plume extent during geologic CO 2 storage in a saline aquifer. First, novel analytical solutions are developed to model transient temperature in the aquifer during the injection and subsequent shut-in periods considering two-phase flow (gaseous CO 2 and aqueous brine) conditions in the aquifer. The solution involves a discretization of the aquifer into regions; the energy and mass conservation equations for the regions are solved simultaneously considering appropriate boundary conditions at the interfaces. Two solutions techniques are presented: multi-region and three-region solutions. Inverse models are developed accordingly to evaluate the injection profile and estimate the extent of the plume front in the reservoir during the injection period. The multi-region solution results in an inversion approach that requires regression analysis. However, the three-region formulation results in a simple graphical technique for inverse modeling. The analytical solutions are validated against a thermally coupled reservoir simulation tool using different synthetic cases for CO 2 injection in deep saline aquifers. The results of the developed solutions provide a good match with numerical results during forward and inverse modeling.

54 ENVIRONMENTAL SCIENCES↗

Accurate Prediction of Algal Biomass Lipid, Protein, and Carbohydrate Composition with Machine Learning Regression Modelling of Near-IR Spectra

During large scale algal biomass cultivation, it is difficult to reliably control relative composition to target levels. Rapid determination of chemical composition is feasible by using near infrared (NIR) spectral data. We sought to build and improve on reliable high-throughput screening prediction method based on partial least squares regression (PLSR) by the application of artificial neural networks (ANN) and associated optimization strategies. The algal biomass sample set was designed and created in an iterative process of culturing in physiologically diverse conditions at the GAI field site, followed by compositional analyses at NREL. The workflow allowed us to identify gaps in compositional space for informing the subsequent cultivation and sampling efforts and generated a high quality set of 210 unique samples with chemical analysis results, spectral scanning data, and cultivation metadata. We observed a significant improvement in the performance of carbohydrate content predictions using an optimized ANN model compared to PLSR, with > 16% reduction in mean absolute percent error (MAPE) when tested on the same set of reserved data. The optimized ANN models for FAME and protein prediction performed exceptionally well with 5.99% and 5.09% MAPE, respectively. Application of these methods to detection and quantification of minor biomass constituents that are relevant to certain product streams has shown positive preliminary results, opening the possibility for extensions to the outputs of this powerful data type. All models are accompanied by prediction uncertainties and unsupervised spectral outlier detection to alert an operator to unreliable spectral data. These tools can be deployed for rapid determination of algal culture status, and cultivation and biomass quality improvement.

algal biofuels↗

Aerosol emissions from water-lean solvents for post-combustion CO 2 capture

Advanced water-lean solvents (WLS) for post-combustion CO 2 capture have been gaining interest due to their ability to reduce the parasitic penalty from energy needed for solvent regeneration. Commercial implementation of these novel CO 2 capture technologies hinges on successful control of amine emissions. RTI conducted a parametric study of fundamental and operational variables influence on overall amine aerosol and vapor emissions from our water-lean solvent eCO 2 Sol™ using our 6-kW equivalent bench-scale gas absorption system. The parametric testing used a simulated flue gas with 15 % CO 2 , 2.3–4.2 % H 2 O, and 0–6 ppm sulfite (SO 3 ) to examine the impact of the presence of aerosols to the capture performance and amine emissions from the system. The SO 3 reacts with water in the flue gas to create H 2 SO 4 , which forms liquid aerosol droplets and provide nucleation sites for growth of aerosols. Scanning Mobility Particle Sizer and Aerodynamic Particle Sizer instruments monitored the aerosol particle size distribution. Parametric testing results suggested that the presence of the aerosols in the flue gas could increase the overall amine emissions by 10X compared to the baseline emissions from WLS’s vapor pressure. Principal component analysis (PCA) and projection to latent squares (PLS) developed models to predict the aerosol-based amine emissions from process data. The predictive PLS model had a correlation coefficient (Q 2 ) of 0.92 and could predict the aerosol-based emissions from the NAS process with ±15 % accuracy (average absolute deviation, AAD). The PLS regression model also identified key variables affecting aerosol-based emissions from WLS.

42 ENGINEERING↗

Accelerating phase field simulations through a hybrid adaptive Fourier neural operator with U-net backbone

Prolonged contact between a corrosive liquid and metal alloys can cause progressive dealloying. For one such process as liquid-metal dealloying (LMD), phase field models have been developed to understand the mechanisms leading to complex morphologies. However, the LMD governing equations in these models often involve coupled non-linear partial differential equations (PDE), which are challenging to solve numerically. In particular, numerical stiffness in the PDEs requires an extremely refined time step size (on the order of 10 -12 s or smaller). This computational bottleneck is especially problematic when running LMD simulation until a late time horizon is required. This motivates the development of surrogate models capable of leaping forward in time, by skipping several consecutive time steps at-once. In this paper, we propose a U-shaped adaptive Fourier neural operator (U-AFNO), a machine learning (ML) based model inspired by recent advances in neural operator learning. U-AFNO employs U-Nets for extracting and reconstructing local features within the physical fields, and passes the latent space through a vision transformer (ViT) implemented in the Fourier space (AFNO). We use U-AFNOs to learn the dynamics of mapping the field at a current time step into a later time step. We also identify global quantities of interest (QoI) describing the corrosion process (e.g., the deformation of the liquid-metal interface, lost metal, etc.) and show that our proposed U-AFNO model is able to accurately predict the field dynamics, in spite of the chaotic nature of LMD. Most notably, our model reproduces the key microstructure statistics and QoIs with a level of accuracy on par with the high-fidelity numerical solver, while achieving a significant 11, 200 × speed-up on a high-resolution grid when comparing the computational expense per time step. Finally, we also investigate the opportunity of using hybrid simulations, in which we alternate forward leaps in time using the U-AFNO with high-fidelity time stepping. We demonstrate that while advantageous for some surrogate model design choices, our proposed U-AFNO model in fully auto-regressive settings consistently outperforms hybrid schemes.

36 MATERIALS SCIENCE↗

Supervised machine learning-based multivariate regression of parallel closures for a high-collisionality deuterium-carbon plasma

Many plasmas of interest in laboratory experiments and space consist of multiple ion species. In tokamak edge plasmas, for instance, ionized impurities expelled from the vessel wall influence plasma transport. When describing multi-species plasmas using fluid equations, we need accurate closure relations to close the set of fluid equations. In this study, we introduce the development of fitting formulas for parallel closures using supervised machine learning, in conjunction with the recent closure theory, considering multi-ion collisions and arbitrary ion temperatures. We apply this approach to a high-collisionality deuterium-carbon plasma and demonstrate its effectiveness. As a result, the machine learning-based method for developing practical and accurate closures can be extended to a wider range of plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predicting U.S. federal fleet electric vehicle charging patterns using internal combustion engine vehicle fueling transaction statistics

Utilizing fueling transactions from internal combustion engine vehicles (ICEVs), the authors estimated how frequently midday public charging would be required for U.S. federal fleet battery electric vehicles (BEVs). Fueling transaction summary statistics are more widely available than trip-level telematics data, making this methodology more accessible and transferable to other researchers and fleet managers considering BEV replacements. For example, readers can easily apply a linear model using only the count of back-to-back fueling events at gas stations over 57 straight-line miles apart to predict days exceeding range. This linear regression predicted binned days exceeding 250 miles at 80% accuracy on a hold-out test set from the same fleet as the training data and 66 % accuracy on a new fleet displaying different driving behaviors. The authors additionally provide linear equations for days exceeding 200 and 300 miles as alternative range estimates to account for differences in BEV range and temperature impacts. Beyond the single-feature linear models which readers can apply, the authors tuned and trained other machine learning models on a variety of fueling transaction statistics including consecutive transaction distances, transaction distance from garage, estimated miles traveled from fuel economy and fuel quantity, and transaction periodicity. Utilizing a subset of 1678 light-duty federal fleet vehicles which contained daily vehicle miles traveled (VMT) in addition to fueling statistics, the authors determined which fueling transaction statistics were most relevant in predicting driving days exceeding 250 miles (an approximation of BEV rated driving range). In support of the U.S. federal fleet transition to zero-emission vehicles (ZEVs), the authors used these statistics and machine learning models to predict the frequency of BEV midday charging. After training models on the subset with VMT, the authors predicted days exceeding rated range for 112,902 light-duty vehicles operating in similar circumstances in the federal fleet using a Support Vector Regressor (SVR). In conclusion, they then used the projections as part of the ZEV Planning and Charging (ZPAC) tool to identify optimal candidates for BEVs for the federal fleet. An anonymized version of ZPAC is included in the supplementary materials.

25 ENERGY STORAGE↗

Code and Solution Verification Assessment of the CTF Thermal Hydraulic Subchannel Code

CTF is a thermal-hydraulics subchannel code jointly developed by Oak Ridge National Laboratory and North Carolina State University. Over the past seven years, the Consortium for Advanced Simulation of Light Water Reactors (CASL) has made a significant investment in developing CTF so it can be used to model light water reactors, including nominal operating conditions, departure from nucleate boiling analysis, and transients ranging from loss of flow to reactivity insertion accidents. In addition to implementing new modeling capabilities and developing the user input and output interface, extensive work has been performed to improve the code’s quality assurance program, resulting in a development process that conforms with NQA-1 requirements. The CASL program follows the Predictive Capability Maturity Model (PCMM) approach for assessing code quality, which emphasizes performing code verification(ensuring the code converges to the correct answer) and solution verification (ensuring the code converges for the intended application). Code and solution verification are used to identify uncertainty errors introduced by numerical approximations in the code and are important for demonstrating that the model is coded without error, which is an important aspect of the Best Estimate plus Uncertainty method. This paper presents a comprehensive overview of the code and solution verification testing that has been performed on CTF. A top-down approach is taken in which the intended CTF applications are presented, followed by the code features required for their modeling. These features are then linked to the applicable code and solution verification tests that demonstrate proper functioning. Past testing efforts are summarized, and new tests are added to help close gaps in the presented test matrix. Rather than performing “one-off” exercises, these tests are added to the automated CTF regression test suite to ensure continual code quality.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Automated System-wide Event Detection and Classification Using Machine Learning on Synchrophasor Data

As the number of phasor measurement units (PMUs) deployed in a power system increases, and their data volume streamed to the control canter intensifies, operators are facing challenges related to the analysis of such data, which need to be observed and responded to as the measurements are displayed in the Control Room. Humans are generally unable to process such large amount of data efficiently and rapidly. There is an apparent need for automated ways to analyze the data, extract actionable information about occurrence of specific events, and characterize the events quickly and cost effectively. This paper discusses the use of machine learning (ML) to facilitate such tasks by providing automated, highly computationally efficient, and cost-effective ways of extracting actionable information from synchrophasor big data in real-time. We developed Big Data Smart (BDSmart) ML-based prototype tool for the Control Room use that automatically analyses data properties from synchrophasor system measurements taken across the three grid Interconnections in the USA (Western, Eastern and ERCOT). The data collected from several hundreds of PMUs located across the Interconnections over a period of two years have been made available for our extensive study. As a result, we were able to identify a number of big data properties that influence how ML methodology is applied to select, develop, train and test the data models that can eventually be used for the tool implementation. The resulting set of candidate algorithms spans unsupervised, supervised, semi-supervised and transfer-learning approaches. Many ML techniques, such as decision trees, multinomial logistic regression, feed-forward neural networks, K-nearest neighbor, multiclass support vector machine, and single and multi-channel convolutional neural networks, are implemented, and their performance is examined. We offer the results from testing the data models. The novelty of our study is in the approaches for bad data detection and mitigation, selection of a simplified feature for event detection, and data label improvements. As a result, we came up with a list of recommendations for the utilities on how to improve the PMU recording practices to cater to the future ML applications aimed at automating the analysis of synchrophasor data.

Synchrophasors, Machine Learning, System-wide Even↗

A Bayesian nonparametric analysis for zero-inflated multivariate count data with application to microbiome study

High-throughput sequencing technology has enabled researchers to profile microbial communities from a variety of environments, but analysis of multivariate taxon count data remains challenging. Here, we develop a Bayesian nonparametric (BNP) regression model with zero inflation to analyse multivariate count data from microbiome studies. A BNP approach flexibly models microbial associations with covariates, such as environmental factors and clinical characteristics. The model produces estimates for probability distributions which relate microbial diversity and differential abundance to covariates, and facilitates community comparisons beyond those provided by simple statistical tests. We compare the model to simpler models and popular alternatives in simulation studies, showing, in addition to these additional community-level insights, it yields superior parameter estimates and model fit in various settings. The model's utility is demonstrated by applying it to a chronic wound microbiome data set and a Human Microbiome Project data set, where it is used to compare microbial communities present in different environments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Prediction of DIII-D Pedestal Structure from Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. An experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (ne) and electron temperature (Te) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (Ip), toroidal magnetic field (Bφ), neutral beam heating power (PNBI) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Prediction of DIII-D Pedestal Structure From Externally Controllable Parameters

The sharp increase of pressure at the edge of a high confinement mode (H-mode) plasma, the pedestal, strongly impacts overall plasma performance. Predicting the pedestal is a necessity to control and optimize tokamak operations. Here, an experimental data-driven machine learning (ML) approach is presented that predicts the pedestal heights and widths of electron density (n e ) and electron temperature (T e ) profiles as well as the separatrix ne from externally controllable parameters such as the plasma shape, heating method and power, and gas puff rate and integrated gas puff. The OMFIT framework was used with DIII-D data to efficiently, robustly, and automatically build a database of pedestal parameters to train machine learning models. Database creation was enabled by the search engine tool for DIII-D data, TokSearch, which parallelizes data fetching, enabling fast searches through basic signals of thousands of DIII-D shots and selection of relevant time intervals. Principal Component Analysis (PCA) separated the database into three clusters that represent classes of plasma shapes that are regularly used in DIII-D. The most important parameters for setting the pedestal structure were plasma current (I p ), toroidal magnetic field (B Φ ), neutral beam heating power (P NBI ) and shaping quantities. The Deep Jointly Informed Neural Networks (DJINN) algorithm was applied to identify suitable neural network (NN) architectures that appropriately capture the features of the pedestal database. Separate NNs were implemented for each pedestal parameter, and ensembling methods were used to improve the prediction accuracy and allowed estimation of the prediction uncertainty. The pedestal predictions of the test dataset lie within the measurement uncertainties of the pedestal parameters. The NN outperformed simple Linear Regression (LR) analysis, indicating non-linear dependencies in the pedestal structure. The presented achievements illustrate a promising path for future research, using feature extraction to infer experimental trends and thereby improve pedestal models as well as deploying NN for a fast pedestal prediction in DIII-D scenario development.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Machine Learning for Daily Forecasts of Arctic Sea Ice Motion: An Attribution Assessment of Model Predictive Skill

Physics-based simulations of Arctic sea ice are highly complex, involving transport between different phases, length scales, and time scales. Resultantly, numerical simulations of sea ice dynamics have a high computational cost and model uncertainty. We employ data-driven machine learning (ML) to make predictions of sea ice motion. The ML models are built to predict present-day sea ice velocity given present-day wind velocity and previous-day sea ice concentration and velocity. Models are trained using reanalysis winds and satellite-derived sea ice properties. We compare the predictions of three different models: persistence (PS), linear regression (LR), and a convolutional neural network (CNN). We quantify the spatiotemporal variability of the correlation between observations and the statistical model predictions. Additionally, we analyze model performance in comparison to variability in properties related to ice motion (wind velocity, ice velocity, ice concentration, distance from coast, bathymetric depth) to understand the processes related to decreases in model performance. Results indicate that a CNN makes skillful predictions of daily sea ice velocity with a correlation up to 0.81 between predicted and observed sea ice velocity, while the LR and PS implementations exhibit correlations of 0.78 and 0.69, respectively. The correlation varies spatially and seasonally: lower values occur in shallow coastal regions and during times of minimum sea ice extent. LR parameter analysis indicates that wind velocity plays the largest role in predicting sea ice velocity on 1-day time scales, particularly in the central Arctic. Regions where wind velocity has the largest LR parameter are regions where the CNN has higher predictive skill than the LR.

54 ENVIRONMENTAL SCIENCES↗