Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Dataset for Top Model Decision Tree: Selecting Segmentation Models for Reliable Quantitative Analysis in Low- and Ultralow-Dose CryoEM

Motivation Multiple deep learning model architectures can be used to segment bacterial membranes in cryoEM images. However, an AI-based tool advancement is often presented with only a single segmentation model for broad use, and this single model may show inconsistent results across datasets from different users. Here, we present the Top Model Decision Tree, a model screening framework to screen for the best model to generate bacterial inner and outer membrane masks based on user priorities. We use pre-trained segmentation models from YOLOv11, YOLO26, U-Net, Detectron2 and SAM3 fine-tuned on bacterial inner and outer membranes imaged with cryoEM. Run the Framework This notebook must be opened in Google Colab. Mount Google Drive and run with a GPU-based runtime. Open the notebook and follow steps to git clone in folders and files within this repository. There will be a repeating top_model_decision_tree.ipynb (notebook clone) that will not be used. Save your .png binary mask files and .csv table outputs within your Google Drive or download before closing the notebook. The models and all analysis/training scripts are available at [GitHub: https://github.com/Lynnicia/CryoEM_membranes_top_model_decision_tree and https://github.com/Sireesiru/Semantic-Segmentation-of-bacterial-cell-envelope-using-U-Nets.

59 BASIC BIOLOGICAL SCIENCES↗

Model averaging approaches to data subset selection

Model averaging is a useful and robust method for dealing with model uncertainty in statistical analysis. Often, it is useful to consider data subset selection at the same time, in which model selection criteria are used to compare models across different subsets of the data. Two different criteria have been proposed in the literature for how the data subsets should be weighted. We compare the two criteria closely in a unified treatment based on the Kullback-Leibler divergence and conclude that one of them is subtly flawed and will tend to yield larger uncertainties due to loss of information. Here, analytical and numerical examples are provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

FY2021 Status Report on the Computing Systems for the Yucca Mountain Project TSPA-LA Models and Testing of Selected Process Models

Sandia National Laboratories continued evaluation of the total system performance assessment (TSPA) for License Application (LA) computing systems for the previously considered Yucca Mountain Project (YMP). This was done to maintain the operational readiness of the computing infrastructure (computer hardware and software) and knowledge capability for total system performance assessment) type analysis, as directed by the National Nuclear Security Administration (NNSA), DOE 2010. The FY21 task included continued operation of the cluster; maintenance of the TSPA-LA models (with GoldSim 9.60.300); continued assessment of the status of the Infiltration Model; (a process model that feeds the TSP -LA) and preliminary assessments of the Unsaturated Zone Flow Model and the Saturated Zone Flow and Transport Model Abstraction (process models that feed the TSPA-LA). The 2014 cluster and supporting software systems are currently fully operational to support TSPA-LA type analyses.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

High Strain-Rate Strength Experiments on Molybdenum for Constitutive Model Parameter Selection

The constitutive model parameters for molybdenum obtained from the Plansee supplier have been constrained via Strength-Stabilized Rayleigh-Taylor experiments performed at the Los Alamos Neutron Science Center. Previous Hopkinson-bar measurements have produced a degenerate family of constitutive model parameters proven in the low strain rate regime. The experiments presented here provide a constraint at strain rates from 10 4 –10 6 s -1 which is used to select a preferred set of parameters for high strain-rate applications. Further analysis is warranted including determining a global fit in parameter space using all available high and low strain-rate data available.

36 MATERIALS SCIENCE↗

Microbiome-enabled genomic selection improves prediction accuracy for nitrogen-related traits in maize

Root-associated microbiomes in the rhizosphere (rhizobiomes) are increasingly known to play an important role in nutrient acquisition, stress tolerance, and disease resistance of plants. However, it remains largely unclear to what extent these rhizobiomes contribute to trait variation for different genotypes and if their inclusion in the genomic selection protocol can enhance prediction accuracy. To address these questions, we developed a microbiome-enabled genomic selection method that incorporated host SNPs and amplicon sequence variants from plant rhizobiomes in a maize diversity panel under high and low nitrogen (N) field conditions. Our cross-validation results showed that the microbiome-enabled genomic selection model significantly outperformed the conventional genomic selection model for nearly all time-series traits related to plant growth and N responses, with an average relative improvement of 3.7%. The improvement was more pronounced under low N conditions (8.4–40.2% of relative improvement), consistent with the view that some beneficial microbes can enhance N nutrient uptake, particularly in low N fields. However, our study could not definitively rule out the possibility that the observed improvement is partially due to the amplicon sequence variants being influenced by microenvironments. Using a high-dimensional mediation analysis method, our study has also identified microbial mediators that establish a link between plant genotype and phenotype. Some of the detected mediator microbes were previously reported to promote plant growth. The enhanced prediction accuracy of the microbiome-enabled genomic selection models, demonstrated in a single environment, serves as a proof-of-concept for the potential application of microbiome-enabled plant breeding for sustainable agriculture.

60 APPLIED LIFE SCIENCES↗

Methods for Computing Physically Realistic Estimates of Electric Water Heater Demand Response Resource Suitable for Bulk Power System Planning Models

Demand response is commonly called on to reduce load during system peak times or to respond to contingency events. In future power systems with higher shares of wind and solar generation (which we describe together as variable generation [VG]), demand response could have more opportunities to provide energy shifting or operating reserve services. This report evaluates the ability of residential electric water heaters, both electric resistance water heaters (ERWHs) and heat pump water heaters (HPWHs), to provide such services starting from detailed whole-building energy models that realistically represent New England single family home stock. We use a parsimonious surrogate model to represent operational flexibility in a form suitable for linear and mixed integer programming. This enables relatively fast determination of aggregate contingency reserve resource, price-taking energy shifting outcomes, and in some cases the determination of aggregate models at the megawatt (MW) scale that can be directly included in large-scale grid models. After selecting modeling methods and parameters through various computational experiments, we find interquartile ranges of contingency reserve resource in ISO-NE for about 603,400 ERWHs of 45 MW - 69 MW for Claim10 (50 minute responses provided with 10 minutes of advanced notification) and 65 MW - 102 MW for Claim30 (30 minute responses provided with 30 minutes of advanced notification), and for about 619,000 HPWHs of 48 MW - 88 MW for Claim10 and 52 MW - 90 MW for Claim30. The overall reserve resource is up to 32% of total load for ERWHs providing Claim10 service, 47% for ERWHs providing Claim30 service, 93% for HPWHs providing Claim10 service, and 97% for HPWHs providing Claim30 service. More work is required to determine if HPWHs are inherently more suitable than ERWHs for providing contingency reserve or if these results reflect idiosyncrasies of the single family home stock model used in this study. The value of this contingency resource in a Near-term VG model of ISO-NE is $\$ 0.40$ to $\$1.20$ per water heater-year, and significantly larger, $\$ 3.80$ to $\$ 5.30$ per water heater-year in a Mid-term VG model of ISONE. Aggregating surrogate models to the MW-scale for energy shifting service is more challenging than for contingency service and we only present such results for ERWHs, because we were unable to determine satisfactory ways to deal with HPWHs' time-varying and path dependent operational characteristics. Individual surrogate models suitable for evaluating the energy shifting resource from both ERWHs and HPWHs are created, however, and dispatched against day-ahead prices from the Near-Term VG and Mid-Term VG models of ISO-NE. The individual surrogate models are able to access and potentially shift all 640 GWh of HPWH load and 1,547 GWh of ERWH load we modeled in two different single family home stock models. In contrast, the most effective model of aggregate ERWH shifting resource we created only captured 34.7% of the total ERWH load. Energy shifting affected by price-taking dispatch against modeled day-ahead energy prices produces per water heater year profits of $\$19.44$ - $\$22.93$ for individual HPWHs, $\$39.11$ - $\$40.54$ for individual ERWHs, and up to $\$4.00$ - $\$4.24$ for aggregated ERWHs, with the variations mainly due to grid conditions (more or less VG). When the supply-side response to these changes is accounted for, the per water heater year production cost savings for ISO-NE are $\$7.50$ to $\$17.70$ for the most effective set of endogenously dispatched aggregate ERWHs, $\$15.60$ to $\$15.70$ for individual ERWHs dispatched against the DA prices, and $\$10.70$ to $\$11.20$ for individual HPWHs dispatched against DA prices. Those ranges primarily represent the difference between Near-Term VG and Mid-Term VG grid conditions.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Uncertainty-informed selection of CMIP6 Earth System Model subsets for use in multisectoral and impact models

Earth system models (ESMs) and general circulation models (GCMs) are heavily used to provide inputs to sectoral impact and multisector dynamic models, which include representations of energy, water, land, economics, and their interactions. Therefore, representing the full range of model uncertainty, scenario uncertainty, and interannual variability that ensembles of these models capture is critical to the exploration of the future co-evolution of the integrated human–Earth system. The pre-eminent source of these ensembles has been the Coupled Model Intercomparison Project (CMIP). With more modeling centers participating in each new CMIP phase, the size of the model archive is rapidly increasing, which can be intractable for impact modelers to effectively utilize due to computational constraints and the challenges of analyzing large datasets. In this work, we present a method to select a subset of the latest phase, CMIP6, featuring models for use as inputs to a sectoral impact or multisector dynamics models, while prioritizing preservation of the range of model uncertainty, scenario uncertainty, and interannual variability in the full CMIP6 ensemble results. This method is intended to help impact modelers select climate information from the CMIP archive efficiently for use in downstream models that require global coverage of climate information. This is particularly critical for large-ensemble experiments of multisector dynamic models that may be varying additional features beyond climate inputs in a factorial design, thus putting constraints on the number of climate simulations that can be used. We focus on temperature and precipitation outputs of CMIP6 models, as these are two of the most used variables among impact models, and many other key input variables for impacts are at least correlated with one or both of temperature and precipitation (e.g., relative humidity). Besides preserving the multi-model ensemble variance characteristics, we prioritize selecting CMIP6 models in the subset that preserve the very likely distribution of equilibrium climate sensitivity values as assessed by the latest Intergovernmental Panel on Climate Change (IPCC) report. This approach could be applied to other output variables of climate models and, possibly when combined with emulators, offers a flexible framework for designing more efficient experiments on human-relevant climate impacts. It can also provide greater insight into the properties of existing CMIP6 models.

Snyder, Abigail C.↗

Comparative Analyses of Bioequivalence Assessment Methods for In Vitro Permeation Test Data

ABSTRACT For topical, dermatological drug products, an in vitro option to determine bioequivalence (BE) between test and reference products is recommended. In particular, in vitro permeation test (IVPT) data analysis uses a reference‐scaled approach for two primary endpoints, cumulative penetration amount (AMT) and maximum flux ( J max ), which takes the within donor variability into consideration. In 2022, the Food and Drug Administration (FDA) published a draft IVPT guidance that includes statistical analysis methods for both balanced and unbalanced cases of IVPT study data. This work presents a comprehensive evaluation of various methodologies used to estimate critical parameters essential in assessing BE. Specifically, we investigate the performance of the FDA draft IVPT guidance approach alongside alternative empirical and model‐based methods utilizing mixed‐effects models. Our analyses include both simulated scenarios and real‐world studies. In simulated scenarios, empirical formulas consistently demonstrate robustness in approximating the true model, particularly in effectively addressing treatment–donor interactions. Conversely, the effectiveness of model‐based approaches heavily relies on precise model selection, which significantly influences their results. The research emphasizes the importance of accurate model selection in model‐based BE assessment methodologies. It sheds light on the advantages of empirical formulas, highlighting their reliability compared to model‐based approaches and offers valuable implications for BE assessments. Our findings underscore the significance of robust methodologies and provide essential insights to advance their understanding and application in the assessment of BE, employed in IVPT data analysis.

Leon, Sami↗

Ionization detail parameters and cluster dose: a mathematical model for selection of nanodosimetric quantities for use in treatment planning in charged particle radiotherapy

Abstract Objective . To propose a mathematical model for applying ionization detail (ID), the detailed spatial distribution of ionization along a particle track, to proton and ion beam radiotherapy treatment planning (RTP). Approach . Our model provides for selection of preferred ID parameters ( I p ) for RTP, that associate closest to biological effects. Cluster dose is proposed to bridge the large gap between nanoscopic I p and macroscopic RTP. Selection of I p is demonstrated using published cell survival measurements for protons through argon, comparing results for nineteen I p : N k , k = 2, 3, …, 10, the number of ionizations in clusters of k or more per particle, and F k , k = 1, 2, …, 10, the number of clusters of k or more per particle. We then describe application of the model to ID-based RTP and propose a path to clinical translation. Main results . The preferred I p were N 4 and F 5 for aerobic cells, N 5 and F 7 for hypoxic cells. Significant differences were found in cell survival for beams having the same LET or the preferred N k . Conversely, there was no significant difference for F 5 for aerobic cells and F 7 for hypoxic cells, regardless of ion beam atomic number or energy. Further, cells irradiated with the same cluster dose for these I p had the same cell survival. Based on these preliminary results and other compelling results in nanodosimetry, it is reasonable to assert that I p exist that are more closely associated with biological effects than current LET-based approaches and microdosimetric RBE-based models used in particle RTP. However, more biological variables such as cell line and cycle phase, as well as ion beam pulse structure and rate still need investigation. Significance . Our model provides a practical means to select preferred I p from radiobiological data, and to convert I p to the macroscopic cluster dose for particle RTP.

Engineering↗

Aboveground biomass density models for NASA’s Global Ecosystem Dynamics Investigation (GEDI) lidar mission

NASA's Global Ecosystem Dynamics Investigation (GEDI) is collecting spaceborne full waveform lidar data with a primary science goal of producing accurate estimates of forest aboveground biomass density (AGBD). This paper presents the development of the models used to create GEDI's footprint-level (~25 m) AGBD (GEDI04_A) product, including a description of the datasets used and the procedure for final model selection. The data used to fit our models are from a compilation of globally distributed spatially and temporally coincident field and airborne lidar datasets, whereby we simulated GEDI-like waveforms from airborne lidar to build a calibration database. We used this database to expand the geographic extent of past waveform lidar studies, and divided the globe into four broad strata by Plant Functional Type (PFT) and six geographic regions. GEDI's waveform-to-biomass models take the form of parametric Ordinary Least Squares (OLS) models with simulated Relative Height (RH) metrics as predictor variables. From an exhaustive set of candidate models, we selected the best input predictor variables, and data transformations for each geographic stratum in the GEDI domain to produce a set of comprehensive predictive footprint-level models. We found that model selection frequently favored combinations of RH metrics at the 98th, 90th, 50th, and 10th height above ground-level percentiles (RH98, RH90, RH50, and RH10, respectively), but that inclusion of lower RH metrics (e.g. RH10) did not markedly improve model performance. Second, forced inclusion of RH98 in all models was important and did not degrade model performance, and the best performing models were parsimonious, typically having only 1-3 predictors. Third, stratification by geographic domain (PFT, geographic region) improved model performance in comparison to global models without stratification. Fourth, for the vast majority of strata, the best performing models were fit using square root transformation of field AGBD and/or height metrics. There was considerable variability in model performance across geographic strata, and areas with sparse training data and/or high AGBD values had the poorest performance. These models are used to produce global predictions of AGBD, but will be improved in the future as more and better training data become available.

54 ENVIRONMENTAL SCIENCES↗

Small Reactors in Microgrids: Technology Modeling and Selection (Net-Zero Microgrid Program Project Report)

This report demonstrates the capabilities of the net-zero microgrid (NZM) Xendee platform for modeling an SR module with electricity, heat extraction and thermal storage in microgrids configurations. The model effectively captures the most important technical and economic considerations for SR technology specific analysis: cost and operational characteristics of SR technology and financial costs and incentives. The model can analyze multiple scenarios to establish metrics for cost-competitive and zero-carbon microgrids connected to the grid or completely isolated. The model is fully integrated within the Xendee platform for modeling and analysis of clean energy microgrids with storage and generation from renewable energy sources. The model captures the capabilities, constraints, and nuances of SR by incorporating parameters related to plant economics, design efficiency and performance, plant operation and component and fuel lifespan. The cost and operational parameters modeled in the SR module are specific to the technology selected for integration in the microgrid. Cost parameters recognize advanced nuclear technology for modular production and installation based on economies of scale from factory manufacture and related commissioning, and cost reduction through technology maturation—first-of-a-kind (FOAK) and nth-of-a-Kind (NOAK). The cost parameters include installation, operations and maintenance (O&M), fuel refueling cycle, and reactor life. Installation cost reflects economies of scale due to unit sizing at scale and colocation. O&M economies of scale for both fixed- and variable-cost fuel life-cycle costs are incurred at every refueling interval, with separate front- and back-end fuel costs, as well as waste-handling and disposition costs. This report investigates key characteristics of different SR technologies suitable for microgrid applications, including design principles, sizing, coolant properties, temperature ratings, fuel structures, and life-cycle considerations. This also includes fuel technologies applicable to these SR systems, alongside strategies for nuclear-waste and spent-fuel management and approaches to address safety, security, and proliferation challenges. Four primary groups of SR technologies are examined: water-cooled, liquid-metal-cooled, high-temperature gas-cooled, and molten-salt-cooled systems. In this report, an initial guideline for technology selection is established, aligning the characteristics of the technologies with the requirements of microgrids. The selection of technology in a microgrid is influenced by various factors, including financial capacity, location and accessibility, demand type and characteristics, reliability and resilience requirements, area constraints, and the lifespan of the microgrid. The types of electrical and non-electrical applications within the microgrid also play a significant role in technology selection. The characteristics of SRs, such as their smaller size, modularity, transportability, long refueling interval, improved safety features, ability to operate in autonomous or semi-autonomous mode, and provision of high-grade heat, are particularly appealing for microgrids. Furthermore, a list of considerations for implementing SRs in microgrids is outlined. The SR model is created to be continuously improved with the acquisition of actual data on investment and operational costs, experience with supply chains, production at scale, and field deployments. In the near term, performance data on applications in microgrids will become available from lessons learned from laboratory tests, such as those planned for the Microreactor Applications Research Validation and Evaluation Project (MARVEL), led by Idaho National Laboratory (INL). The SR model incorporates scenario data and known SR design specifications, enabling technoeconomic analysis for SR deployment in microgrids. It specifically considers the distinctive attributes of SRs as generators in technoeconomic studies. SRs can be modeled and analyzed with generation from renewable-energy sources, energy storage, and flexible loads over a range of functionality and applications. This offers a comprehensive tool for feasibility studies, scenario development, and sensitivity analysis for “what-if” consideration of any range of assumptions about SRs in microgrids and other aggregations of distributed-energy resources, including virtual power plants.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Adoption of image-driven machine learning for microstructure characterization and materials design: A Perspective

Microstructure characterization enables the development of structure-processing-property relationships critical to several research areas within the broad field of materials science, from alloy design to the assessment of corrosion resistance, and failure analysis. Conventional approaches to material characterization have relied on either qualitative inference by the human ex-pert or software applications that can extract high-level features from images, such as boundary segmentation, average grain diameter, etc. Such approaches rely heavily on subject matter expert user intervention and knowledge of what phases or more generally, what microstructural features, are of interest. The recent surge in the adoption of machine learning techniques to address problems in materials engineering has brought with it an increased interest and application of Image Driven Machine Learning (IDML) approaches. In this work, we review the applications of IDML to the field of materials characterization. A canonical hierarchy of stages is defined, which when put sequentially together completes an IDML study: problem definition, dataset building, model selection and training, model evaluation, and integration with existing instrumentation or simulation workflow. The studies reviewed in this work are analyzed from the perspective of each of these stages. Such a review permits agranular assessment of the field, for example the impact of IDML on materials characterization at the nanoscale, the size of a typical dataset required to train a semantic segmentation model on electron microscopy images, ubiquitousness of transfer learning in the domain, etc. Finally, we discuss the importance of interpretability and explainability in the field of IDML for materials characterization, and provide an overview of two emerging techniques in the field: semantic segmentation and generative adversarial networks.

Baskaran, Arun↗

Model validation and selection in metabolic flux analysis and flux balance analysis

13C-Metabolic Flux Analysis (13C-MFA) and Flux Balance Analysis (FBA) are widely used to investigate the operation of biochemical networks in both biological and biotechnological research. Both methods use metabolic reaction network models of metabolism operating at steady state so that reaction rates (fluxes) and the levels of metabolic intermediates are constrained to be invariant. They provide estimated (MFA) or predicted (FBA) values of the fluxes through the network in vivo, which cannot be measured directly. These fluxes can shed light on basic biology and have been successfully used to inform metabolic engineering strategies. Several approaches have been taken to test the reliability of estimates and predictions from constraint-based methods and to compare alternative model architectures. Despite advances in other areas of the statistical evaluation of metabolic models, such as the quantification of flux estimate uncertainty, validation and model selection methods have been underappreciated and underexplored. We review the history and state-of-the-art in constraint-based metabolic model validation and model selection. Applications and limitations of the χ 2 -test of goodness-of-fit, the most widely used quantitative validation and selection approach in 13C-MFA, are discussed, and complementary and alternative forms of validation and selection are proposed. A combined model validation and selection framework for 13C-MFA incorporating metabolite pool size information that leverages new developments in the field is presented and advocated for. Finally, we discuss how adopting robust validation and selection procedures can enhance confidence in constraint-based modeling as a whole and ultimately facilitate more widespread use of FBA in biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

Group Projected subspace pursuit for IDENTification of variable coefficient differential equations (GP-IDENT)

We propose an effective and robust algorithm for identifying partial differential equations (PDEs) with space-time varying coefficients from the noisy observation of a single solution trajectory. Identifying unknown differential equations from noisy data is a difficult task, and it is even more challenging with space and time varying coefficients in the PDE. The proposed algorithm, GP-IDENT, has three ingredients: (i) we use B-spline bases to express the unknown space and time varying coefficients, (ii) we propose Group Projected Subspace Pursuit (GPSP) to find a sequence of candidate PDEs with various levels of complexity, and (iii) we propose a new criterion for model selection using the Reduction in Residual (RR) to choose an optimal one among a pool of candidates. The new GPSP considers group projected subspaces which is more robust than existing methods in distinguishing correlated group features. We test GP-IDENT on a variety of PDEs and PDE systems, and compare it with the state-of-the-art parametric PDE identification algorithms under different settings to illustrate its outstanding performance. Furthermore, our experiments show that GP-IDENT is effective in identifying the correct terms from a large dictionary, and our model selection scheme is robust to noise.

Data-driven method↗

Use-Inspired, Process-Oriented GCM Selection: Prioritizing Models for Regional Dynamical Downscaling

Dynamical downscaling is a crucial process for providing regional climate information for broad uses, using coarser-resolution global models to drive higher-resolution regional climate simulations. The pool of global climate models (GCMs) providing the fields needed for dynamical downscaling has increased from the previous generations of the Coupled Model Intercomparison Project (CMIP). However, with limited computational resources, the need for prioritizing the GCMs for subsequent downscaling studies remains. GCM selection for dynamical downscaling should focus on evaluating processes relevant for providing boundary conditions to the regional models and be inspired by regional uses such as the response of extremes to changes in the boundary conditions. This leads to the need for metrics representing processes of relevance to diverse stakeholders and subregions of a domain. Procedures to account for metric redundancy and the statistical distinguishability of GCM rankings are required. Further, procedures for selecting realizations from ensembles of top-performing GCM simulations can be used to span the range of climate change signals in multiple ways. As a result, distinct weighting of metrics and prioritization of particular realizations may depend on user needs. We provide high-level guidelines for such region-specific evaluations and address how CMIP7 might enable dynamical downscaling of a representative sample of high-quality models across representative shared socioeconomic pathways (SSPs).

54 ENVIRONMENTAL SCIENCES↗

A Computational Information Criterion for Particle-Tracking with Sparse or Noisy Data

Traditional probabilistic methods for the simulation of advection-diffusion equations (ADEs) often overlook the entropic contribution of the discretization, e.g., the number of particles, within associated numerical methods. Many times, the gain in accuracy of a highly discretized numerical model is outweighed by its associated computational costs or the noise within the data. Herein, we address the question of how many particles are needed in a simulation to best approximate and estimate parameters in one-dimensional advective-diffusive transport. To do so, we use the well-known Akaike Information Criterion (AIC) and a recently-developed correction called the Computational Information Criterion (COMIC) to guide the model selection process. Random-walk and mass-transfer particle tracking methods are employed to solve the model equations at various levels of discretization. Numerical results demonstrate that the COMIC provides an optimal number of particles that can describe a more efficient model in terms of parameter estimation and model prediction compared to the model selected by the AIC even when the data is sparse or noisy, the sampling volume is not uniform throughout the physical domain, or the error distribution of the data is non-IID Gaussian.

97 MATHEMATICS AND COMPUTING↗

Database-wide hazard modelling of the onset of DIII-D tearing modes with field features

The rate of onset (hazard) of tearing modes is modelled probabilistically using statistical learning algorithms. Axisymmetric energy-density equilibrium fields are taken as raw high-dimensional input features which are reduced with principal component analysis. Signal processing of non-axisymmetric magnetics fluctuation array data provides the target information from which to learn. Model selection, visualization and calibration assessment procedures are detailed. Here, the analysis is deployed at large scale across the DIII-D tokamak database. Standard model selection criteria suggest that the energy-density post-processed feature is a better choice for modelling the onset rate compared to the non-processed equilibrium reconstruction solution. Two example applications of the learned rate function are demonstrated: (i) proximity-to-onset discharge monitoring and (ii) database analysis showing an (expected) observational global trend that the general hazard increases as a plasma performance metric increases. An important connection between the hazard function and its use as a conditional probability generator is reviewed in the Appendix.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗