Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Human Error”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Quantifying wildfire drivers and predictability in boreal peatlands using a two-step error-correcting machine learning framework in TeFire v1.0

Abstract. Wildfires are becoming an increasing challenge to the sustainability of boreal peatland (BP) ecosystems and can alter the stability of boreal carbon storage. However, predicting the occurrence of rare and extreme BP fires proves to be challenging, and gaining a quantitative understanding of the factors, both natural and anthropogenic, inducing BP fires remains elusive. Here, we quantified the predictability of BP fires and their primary controlling factors from 1997 to 2015 using a two-step correcting machine learning (ML) framework that combines multiple ML classifiers, regression models, and an error-correcting technique. We found that (1) the adopted oversampling algorithm effectively addressed the unbalanced data and improved the recall rate by 26.88 %–48.62 % when using multiple datasets, and the error-correcting technique tackled the overestimation of fire sizes during fire seasons; (2) nonparametric models outperformed parametric models in predicting fire occurrences, and the random forest machine learning model performed the best, with the area under the receiver operating characteristic curve ranging from 0.83 to 0.93 across multiple fire datasets; and (3) four sets of factor-control simulations consistently indicated the dominant role of temperature, air dryness, and climate extreme (i.e., frost) for boreal peatland fires, overriding the effects of precipitation, wind speed, and human activities. Our findings demonstrate the efficiency and accuracy of ML techniques in predicting rare and extreme fire events and disentangle the primary factors determining BP fires, which are critical for predicting future fire risks under climate change.

54 ENVIRONMENTAL SCIENCES↗

An Applied Strategy for Using Empirical and Hybrid Models in Online Monitoring

The monitoring of plant equipment for failure prediction is one of the key contributors to operation and maintenance (O&M) costs for a nuclear power plant (NPP) because O&M monitoring depends on labor-intensive activities that are required to meet high equipment reliability standards. These activities rely primarily on humans for information gathering, condition diagnosis, and predictive analysis. Online monitoring aims to automate these activities by relying on sensors to replace human information gathering and machine learning to replace human analysis and decision making. To facilitate automated monitoring, a systematic strategy for anomaly detection is needed to optimally use the available sensor data, empirical models, and physics-supported models. This strategy is essential to provide credible reasoning on why and when an empirical (i.e., purely data-driven) versus hybrid (i.e., physics-supported) approach should be used and to determine the ideal mix of these two approaches for a defined anomaly detection scope. The extant methods usually adopt an ad hoc trial-and-error approach that, in addition to being time-consuming and costly, is also highly subjective; it is impacted by the background and the skill set of the personnel making the decisions. Thus, such an approach cannot guarantee an optimum outcome. This represents the motivation of the current research effort, which is focused on devising a scientifically supported strategy for the optimum selection of anomaly detection methods. This report presents a detailed assessment of the main anomaly detection techniques within the empirical or hybrid method streams. Empirical methods include pattern, statistical, and causal inference. Hybrid methods include the use of physics models to train and test data methods, reduce data dimensionality, reduce data-model complexity, augment data, and reduce empirical uncertainty; hybrid methods also include the use of data to tune physics models. The listed techniques within these two streams represent the vast majority of techniques performed for anomaly detection. Using the techniques as outcomes, a strategy was developed to enable a systematic decision-making process to lead to one of these techniques. The strategy is driven by key decision points related to data relevance, simple modeling feasibility, data inference, physics-modeling value, data dimensionality, physics knowledge, method of validation, performance, data availability and suitability for training and testing, cause-effect, entropy inference, and model fitting. Each of these decision points in the strategy is explained in detail in this report with examples, along with the scientific basis behind the decisions and outcomes in common and simplified terminology. The strategy is developed for use by any NPP staff with basic engineering or science knowledge. A user-friendly graphical state flow diagram was also developed as a visual presentation of the strategy. The strategy was tested and demonstrated through two pilot projects for the application of anomaly detection at an NPP. Each pilot had two use cases: an initial case in which certain decisions were made that resulted in one or more empirical techniques and a revised use case where one or more key decisions were modified resulting in using a set of hybrid methods.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

United States multi-sector land use and land cover base maps to support human and Earth system models

Abstract Earth System Models (ESMs) require current and future projections of land use and landcover change (LULC) to simulate land-atmospheric interactions and global biogeochemical cycles. Among the most utilized land systems in ESMs are the Community Land Model (CLM) and the Land-Use Harmonization 2 (LUH2) products. Regional studies also use these products by extending coarse projections to finer resolutions via downscaling or by using multisector dynamic (MSD) models. One such MSD model is the Global Change Analysis Model (GCAM), which has its own independent land module, but often relies on CLM or LUH2 as spatial inputs for its base years. However, this requires harmonization of thematically incongruent land systems at multiple spatial resolutions, leading to uncertainty and error propagation. To resolve these issues, we develop a thematically consistent LULC system for the conterminous United States adaptable to multiple MSD frameworks to support research at a regional level. Using empirically derived spatial products, we developed a series of base maps for multiple contemporary years of observation at a 30-m resolution that support flexibility and interchangeability amongst LUH2, CLM, and GCAM classification systems.

Oliver, Jay↗

Hestia SW-IFL Onroad Fossil Fuel Carbon Dioxide (FFCO2) product: Road segment-level annual FFCO2 emissions across Arizona (2017-2022), version 1.1

The SW-IFL onroad fossil fuel carbon dioxide (FFCO2) emissions data product represents CO2 emissions from the combustion of fossil fuels by motor vehicles (e.g., passenger cars, trucks, buses, motorcycles) traveling on designated roadways. The emissions are represented geographically on each road segment within the state of Arizona spanning the 2017 to 2022 time period. This data product was developed as part of the Southwest Urban Corridor Integrated Field Laboratory (SW-IFL) project, which aims to provide new knowledge and tools that address urban environmental issues by integrating high-resolution observations, modeling, and civic engagement. The emissions data are provided in CSV (input data, ONR_FFCO2_AZ_county.csv) and GeoPackage form (output polyline objects - about 786,000 road segments, XXXX_AZ_v1.1.gpkg) designated by road class (interstates, arterials, collectors, local). The metadata file (Metadata_SW-IFL_Onroad_annualFFCO2_v1.1.docx) provides details about attributes and data formats. The GeoPackage emissions data are provided separately for local roads and nonlocal roads (interstates, arterials, collectors). The method file (Methods_SW-IFL_Onroad_annualFFCO2_v1.1.docx) describes the data processing flow and data sources. Update on 2024-04-17: Updates were made to both the input emission data file (.csv) and output segment-level emission file (.gpkg). There was an update in county-level emission input data (ONR_FFCO2_AZ_county.csv) and the entire road segments were reprocessed to reflect this update.Update on 2024-04-29: Update was made to one output segment-level emission file (Nonlocal_AZ_v1.0.gpkg). There was an error in the AADT values and the data were reprocessed to reflect this update.Update on 2024-10-22: Temporal coverage was extended to include 2022. VMT values were recalculated using new AADT data and the entire road segments were reprocessed to reflect these updates.

54 ENVIRONMENTAL SCIENCES↗

GLEAM: Galaxy Line Emission & Absorption Modeling

We present Galaxy Line Emission & Absorption Modeling (gleam), a Python tool for fitting Gaussian models to emission and absorption lines in large samples of 1D extragalactic spectra. gleam is tailored to work well in batch mode without much human interaction. With gleam, users can uniformly process a variety of spectra, including galaxies and active galactic nuclei, in a wide range of instrument setups and signal-to-noise regimes. gleam also takes advantage of multiprocessing capabilities to process spectra in parallel. With the goal of enabling reproducible workflows for its users, gleam employs a small number of input files, including a central, user-friendly configuration in which fitting constraints can be defined for groups of spectra and overrides can be specified for edge cases. For each spectrum, gleam produces a table containing measurements and error bars for the detected spectral lines and continuum and upper limits for nondetections. For visual inspection and publishing, gleam can also produce plots of the data with fitted lines overlaid. In the present paper, we describe gleam’s main features, the necessary inputs, expected outputs, and some example applications, including thorough tests on a large sample of optical/infrared multi-object spectroscopic observations and integral field spectroscopic data. gleam is developed as an open-source project hosted at https://github.com/multiwavelength/gleam and welcomes community contributions.

79 ASTRONOMY AND ASTROPHYSICS↗

Enhancing the representation of water management in global hydrological models

Abstract. This study enhances an existing global hydrological model (GHM), Xanthos, by adding a new water management module that distinguishes between the operational characteristics of irrigation, hydropower, and flood control reservoirs. We remapped reservoirs in the Global Reservoir and Dam (GRanD) database to the 0.5∘ spatial resolution in Xanthos so that a single lumped reservoir exists per grid cell, which yielded 3790 large reservoirs. We implemented unique operation rules for each reservoir type, based on their primary purposes. In particular, hydropower reservoirs have been treated as flood control reservoirs in previous GHM studies, while here, we determined the operation rules for hydropower reservoirs via optimization that maximizes long-term hydropower production. We conducted global simulations using the enhanced Xanthos and validated monthly streamflow for 91 large river basins, where high-quality observed streamflow data were available. A total of 1878 (296 hydropower, 486 irrigation, and 1096 flood control and others) out of the 3790 reservoirs are located in the 91 basins and are part of our reported results. The Kling–Gupta efficiency (KGE) value (after adding the new water management) is ≥ 0.5 and ≥ 0.0 in 39 and 81 basins, respectively. After adding the new water management module, model performance improved for 75 out of 91 basins and worsened for only 7. To measure the relative difference between explicitly representing hydropower reservoirs and representing hydropower reservoirs as flood control reservoirs (as is commonly done in other GHMs), we use the normalized root mean square error (NRMSE) and the coefficient of determination (R2). Out of the 296 hydropower reservoirs, the NRMSE is > 0.25 (i.e., considering 0.25 to represent a moderate difference) for over 44 % of the 296 reservoirs when comparing both the simulated reservoir releases and storage time series between the two simulations. We suggest that correctly representing hydropower reservoirs in GHMs could have important implications for our understanding and management of freshwater resource challenges at regional-to-global scales. This enhanced global water management modeling framework will allow the analysis of future global reservoir development and management from a coupled human–earth system perspective.

13 HYDRO ENERGY↗

Uncovering Where Compensating Errors Could Hide in ENDF/B-VIII.0

Unconstrained physics spaces between two or more nuclear data observables in a library occur when their values can be simultaneously adjusted without violating the uncertainties in either differential information or simulations of relevant integral experiments. Differential data are often too imprecise to fully bound all nuclear data observables of interest for application simulations. Integral data are simulated with combinations of nuclear data so that an error in one observable may be hidden by a counterbalancing error in another. In this manner compensating errors may lurk within nuclear data libraries and these errors have the potential to undermine the predictive power of neutron transport simulations, particularly in situations where there is no conclusive validation experiment that resembles the application of interest. The EUCLID project (Experiments Underpinned by Computational Learning for Improvements in Nuclear Data) developed a preliminary workflow to identify these unconstrained physics spaces by bringing together results from a large collection of integral experiments with their simulated counter-parts as well as differential information that have a one-to-one correspondence to nuclear data. This wealth of information is processed by machine learning tools for subsequent refinement by human experts. Here, we show how the EUCLID work-flow is executed by applying it first to 239 Pu and then to 9 Be nuclear data in ENDF/B-VIII.0.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Uncovering Where Compensating Errors Could Hide in ENDF/B-VIII.0

Unconstrained physics spaces between two or more nuclear data observables in a library occur when their values can be simultaneously adjusted without violating the uncertainties in either differential information or simulations of relevant integral experiments. Differential data are often too imprecise to fully bound all nuclear data observables of interest for application simulations. Integral data are simulated with combinations of nuclear data so that an error in one observable may be hidden by a counterbalancing error in another. In this manner compensating errors may lurk within nuclear data libraries and these errors have the potential to undermine the predictive power of neutron transport simulations, particularly in situations where there is no conclusive validation experiment that resembles the application of interest. The EUCLID project (Experiments Underpinned by Computational Learning for Improvements in Nuclear Data) developed a preliminary workflow to identify these unconstrained physics spaces by bringing together results from a large collection of integral experiments with their simulated counter-parts as well as differential information that have a one-to-one correspondence to nuclear data. This wealth of information is processed by machine learning tools for subsequent refinement by human experts. Here, we show how the EUCLID work-flow is executed by applying it first to 239 Pu and then to 9 Be nuclear data in ENDF/B-VIII.0.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Efforts to enhance reproducibility in a human performance research project

Background: Ensuring the validity of results from funded programs is a critical concern for agencies that sponsor biological research. In recent years, the open science movement has sought to promote reproducibility by encouraging sharing not only of finished manuscripts but also of data and code supporting their findings. While these innovations have lent support to third-party efforts to replicate calculations underlying key results in the scientific literature, fields of inquiry where privacy considerations or other sensitivities preclude the broad distribution of raw data or analysis may require a more targeted approach to promote the quality of research output. Methods: We describe efforts oriented toward this goal that were implemented in one human performance research program, Measuring Biological Aptitude, organized by the Defense Advanced Research Project Agency's Biological Technologies Office. Our team implemented a four-pronged independent verification and validation (IV&V) strategy including 1) a centralized data storage and exchange platform, 2) quality assurance and quality control (QA/QC) of data collection, 3) test and evaluation of performer models, and 4) an archival software and data repository. Results: Our IV&V plan was carried out with assistance from both the funding agency and participating teams of researchers. QA/QC of data acquisition aided in process improvement and the flagging of experimental errors. Holdout validation set tests provided an independent gauge of model performance. Conclusions: In circumstances that do not support a fully open approach to scientific criticism, standing up independent teams to cross-check and validate the results generated by primary investigators can be an important tool to promote reproducibility of results.

59 BASIC BIOLOGICAL SCIENCES↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

Innovating the next generation of commercial smart building software

Nearly 30% of commercial building energy use is wasted due to equipment faults and HVAC controls problems. The result is increased emissions, compromised comfort and productivity, and less reliable coordination of building power needs with a clean grid. The energy impact alone represents $17 billion in potential savings. Today’s smart building software provides a robust solution to address these operational deficiencies. Energy management and information systems (EMIS) are saving up to 9% on average, with two-year paybacks. They are being incorporated into energy management processes, commissioning services, and utility programs. As effective as they are, two barriers prevent even deeper benefits; limited personnel to fix problems once they are identified, and the expense and time to manually implement changes in control systems. In partnership with the research community, the EMIS industry is developing new capabilities to overcome these barriers. Moving beyond siloed products for either fault detection and diagnostics, or optimal control, these new capabilities empower users to not only automatically identify faults, but also to push corrective action, and control improvements to their buildings. In this paper, several areas for enhancements are documented: ‘one-time’ correction of faults such as setpoints, schedules, and economizer lockouts; short-term active testing for automated proportional integral derivative (PID) loop tuning and functional testing; and continuous supervisory control for demand flexibility and year-round efficiency. Results are presented from a pair of partner implementations out of a dozen providers integrating these enhancements into their products, including field tests from across the country, and insights into operator acceptance and integration into operations and maintenance practices.

Casillas, Armando↗

Unified Language Frontend for Physic-Informed AI/ML

Artificial intelligence and machine learning (AI/ML) are becoming important tools for scientific modeling and simulation as in several other fields such as image analysis and natural language processing. ML techniques can leverage the computing power available in modern systems and reduce the human effort needed to configure experiments, interpret and visualize results, draw conclusions from huge quantities of raw data, and build surrogates for physics based models. Domain scientists in fields like fluid dynamics, microelectronics and chemistry can automate many of their most difficult and repetitive tasks or improve the design times by use of the faster ML-surrogates. However, modern ML and traditional scientific highperformance computing (HPC) tend to use completely different software ecosystems. While ML frameworks like PyTorch and TensorFlow provide Python APIs, most HPC applications and libraries are written in C++. Direct interoperability between the two languages is possible but is tedious and error-prone. In this work, we show that a compiler-based approach can bridge the gap between ML frameworks and scientific software with less developer effort and better efficiency. We use the MLIR (multi-level intermediate representation) ecosystem to compile a pre-trained convolutional neural network (CNN) in PyTorch to freestanding C++ source code in the Kokkos programming model. Kokkos is a programming model widely used in HPC to write portable, shared-memory parallel code that can natively target a variety of CPU and GPU architectures. Our compiler-generated source code can be directly integrated into any Kokkosbased application with no dependencies on Python or cross-language interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Effects of Electrical Stimulation on hiPSC-CM Responses to Classic Ion Channel Blockers

Abstract Human induced pluripotent stem cell-derived cardiomyocytes (hiPSC-CMs) hold great potential for personalized cardiac safety prediction, particularly for that of drug-induced proarrhythmia. However, hiPSC-CMs fire spontaneously and the variable beat rates of cardiomyocytes can be a confounding factor that interferes with data interpretation. Controlling beat rates with pacing may reduce batch and assay variations, enable evaluation of rate-dependent drug effects, and facilitate the comparison of results obtained from hiPSC-CMs with those from adult human cardiomyocytes. As electrical stimulation (E-pacing) of hiPSC-CMs has not been validated with high-throughput assays, herein, we compared the responses of hiPSC-CMs exposed with classic cardiac ion channel blockers under spontaneous beating and E-pacing conditions utilizing microelectrode array technology. We found that compared with spontaneously beating hiPSC-CMs, E-pacing: (1) reduced overall assay variabilities, (2) showed limited changes of field potential duration to pacemaker channel block, (3) revealed reverse rate dependence of multiple ion channel blockers on field potential duration, and (4) eliminated the effects of sodium channel block on depolarization spike amplitude and spike slope due to a software error in acquiring depolarization spike at cardiac pacing mode. Microelectrode array optogenetic pacing and current clamp recordings at various stimulation frequencies demonstrated rate-dependent block of sodium channels in hiPSC-CMs as reported in adult cardiomyocytes. In conclusion, pacing enabled more accurate rate- and concentration-dependent drug effect evaluations. Analyzing responses of hiPSC-CMs under both spontaneously beating and rate-controlled conditions may help better assess the effects of test compounds on cardiac electrophysiology and evaluate the value of the hiPSC-CM model.

Wei, Feng↗

Reconsidering tympanal-acoustic interactions leads to an improved model of auditory acuity in a parasitoid fly

Although most binaural organisms locate sound sources using neurological structures to amplify the sounds they hear, some animals use mechanically coupled hearing organs instead. One of these animals, the parasitoid fly Ormia ochracea (O. ochracea), has astoundingly accurate sound localization abilities. It can locate objects in the azimuthal plane with a precision of 2°, equal to that of humans, despite an intertympanal distance of only 0.5 mm, which is less than 1/100th of the wavelength of the sound emitted by the crickets that it parasitizes. O. ochracea accomplishes this feat via mechanically coupled tympana that interact with incoming acoustic pressure waves to amplify differences in the signals received at the two ears. In 1995, Miles et al developed a model of hearing mechanics in O. ochracea that represents the tympana as flat, front-facing prosternal membranes, though they lie on a convex surface at an angle from the flies' frontal and transverse planes. The model works well for incoming sound angles less than ±30° but suffers from reduced accuracy (up to 60% error) at higher angles compared to response data acquired from O. ochracea specimens. Despite this limitation, it has been the basis for bio-inspired microphone designs for decades. Here, we present critical improvements to this classic hearing model based on information from three-dimensional reconstructions of O. ochracea's tympanal organ. We identified the orientation of the tympana with respect to a frontal plane and the azimuthal angle segment between the tympana as morphological features essential to the flies' auditory acuity, and hypothesized a differentiated mechanical response to incoming sound on the ipsi- and contralateral sides that depend on these features. We incorporated spatially-varying model coefficients representing this asymmetric response, making a new quasi-two-dimensional (q2D) model. The q2D model has high accuracy (average errors of under 10%) for all incoming sound angles. This improved biomechanical model may inform the design of new microscale directional microphones and other small-scale acoustic sensor systems.

36 MATERIALS SCIENCE↗

Physics guided machine learning for multi-material decomposition of tissues from dual-energy CT scans of simulated breast models with calcifications

We introduce a physics guided data-driven method for image-based multi-material decomposition for dual-energy computed tomography (CT) scans. The method is demonstrated for CT scans of virtual human phantoms containing more than two types of tissues. The method is a physics-driven supervised learning technique. We take advantage of the mass attenuation coefficient of dense materials compared to that of muscle tissues to perform a preliminary extraction of the dense material from the images using unsupervised methods. We then perform supervised deep learning on the images processed by the extracted dense material to obtain the final multi-material tissue map. The method is demonstrated on simulated breast models with calcifications as the dense material placed amongst the muscle tissues. The physics-guided machine learning method accurately decomposes the various tissues from input images, achieving a normalized root-mean-squared error of 2.75%.

Gopalakrishnan Meena, Murali↗

Development of a River Dynamical Core for E3SM to simulate compound flooding on Exascale-class heterogeneous supercomputers

Flooding events pose significant risk to human life, property, and infrastructure. Physically-consistent quantification of altered flood risks in global models requires hyper-resolution (~1 km) or fine flood simulations using two-dimensional (2D) physics schemes, both of which are unavailable in the current generation Earth System Models. Here, in this work, we have developed the River Dynamical Core (RDycore), which is an open-source, 2D shallow water equation (SWE) library for the U.S. Department of Energy's Energy Exascale Earth System Model (E3SM). RDycore uses PETSc and libCEED libraries that allows it to run efficiently on CPUs and GPUs, as well as select a time-integration algorithm at runtime without requiring any code modifications. RDycore achieves spatial error convergence rates for problems with analytical and manufactured solutions similar to those reported previously in the literature, or consistent with the implemented first-order spatial discretization scheme. RDycore's accuracy in predicting flooding for a well-studied dam break problem is comparable to existing SWE models. For a problem with 471 million grid cells, RDycore achieves a speedup of 6.6x and 7.6x on GPUs compared to CPUs when using 320 compute nodes on DOE's Perlmutter and Frontier supercomputers, respectively. The one-way coupling of the RDycore library within E3SM is demonstrated by performing multiple 5-day flooding simulations during Hurricane Harvey driven by five precipitation datasets. The E3SM--RDycore simulations at 30 m spatial resolution accurately simulate maximum water height during the hurricane when benchmarked against a previously published study and achieve a speedup of 15x (Perlmutter) and 21x (Frontier) on GPUs relative to CPUs. The work presented here is the foundational step in providing hardware and algorithmic portability framework for simulating kilometer-scale river dynamics within E3SM.

Flood Simulation↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

VirSorter2: a multi-classifier, expert-guided approach to detect diverse DNA and RNA viruses

Abstract Background Viruses are a significant player in many biosphere and human ecosystems, but most signals remain “hidden” in metagenomic/metatranscriptomic sequence datasets due to the lack of universal gene markers, database representatives, and insufficiently advanced identification tools. Results Here, we introduce VirSorter2, a DNA and RNA virus identification tool that leverages genome-informed database advances across a collection of customized automatic classifiers to improve the accuracy and range of virus sequence detection. When benchmarked against genomes from both isolated and uncultivated viruses, VirSorter2 uniquely performed consistently with high accuracy (F1-score > 0.8) across viral diversity, while all other tools under-detected viruses outside of the group most represented in reference databases (i.e., those in the order Caudovirales ). Among the tools evaluated, VirSorter2 was also uniquely able to minimize errors associated with atypical cellular sequences including eukaryotic genomes and plasmids. Finally, as the virosphere exploration unravels novel viral sequences, VirSorter2’s modular design makes it inherently able to expand to new types of viruses via the design of new classifiers to maintain maximal sensitivity and specificity. Conclusion With multi-classifier and modular design, VirSorter2 demonstrates higher overall accuracy across major viral groups and will advance our knowledge of virus evolution, diversity, and virus-microbe interaction in various ecosystems. Source code of VirSorter2 is freely available ( https://bitbucket.org/MAVERICLab/virsorter2 ), and VirSorter2 is also available both on bioconda and as an iVirus app on CyVerse ( https://de.cyverse.org/de ).

59 BASIC BIOLOGICAL SCIENCES↗