Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “expert”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Yes, No, Maybe So: Human Factors Considerations for Fostering Calibrated Trust in Foundation Models Under Uncertainty

High-stakes analytical environments require analysts to evaluate evidence and generate conclusions to inform critical decisions often under conditions of uncertainty. Probabilistic decision-making based on incomplete or inaccurate information can reduce productivity, compromise national interests, and endanger public safety. Researchers are developing expert systems built on foundation models (FMs) to support analysts’ decision-making processes by enabling human-artificial intelligence (AI) teaming, in part through the quantification and expression of uncertainty information. As FMs continue to mature, it is imperative to correspondingly consider analysts’ needs for appropriately interpreting and using uncertainty information. However, prior research indicates that it remains unclear how analysts engage with FM-generated uncertainty information and the extent to which these interactions influence trust in, and reliance on, expert systems. We plan to review the state of the science and conduct an exploratory, qualitative study to (a) understand how properly communicated uncertainty can foster calibrated trust and appropriate reliance and (b) identify approaches for effectively conveying FM-generated uncertainty information during analytical workflows. We will administer semi-structured interviews with analysts from a specific high-stakes analytical environment to collect their current experiences with job-related uncertainty and their impressions when viewing FM-generated uncertainty information. During the interview protocol, participants will be presented with several different FM outputs and invited to discuss their thoughts and beliefs about the uncertainty information displayed. Participants may provide insights into how trust and reliance may be influenced by uncertainty. The results of this study will help us to better understand how analysts currently interpret and use uncertainty information. Our findings may inform human factors recommendations for effectively conveying uncertainty information to foster calibrated trust in, and appropriate reliance on, expert systems. Interaction designers and FM developers can use this knowledge to enhance human-AI teaming and ensure the responsible deployment of FM-based expert systems in analytical workflows.

97 MATHEMATICS AND COMPUTING↗

Model-based Hierarchical Reinforcement Learning for Improved Physical Security Design: A Prototype

Prior work in FY24 developed an adversarial AI agent aid in path analysis of physical protection systems. This agent, trained using a model-based reinforcement learning algorithm, was able to successfully learn the most vulnerable path in facilities. It was able to extend the current state of practice for physical protection design by exhibiting dynamic behavior based on current environmental conditions. Whereas PathTrace largely performs a static, graph-based analysis, the AI agent was able to make decisions based on relative position in the facility, current conditions (was the adversarial agnet discovered?), and proximity to secondary targets. The agent demonstrated some novel capabilities, but had limitations that need to be resolved before it can be used for production purposes. For example, the adversarial agent generalizes poorly and takes a relatively long time to train. Nonetheless, there is still considerable promise for developing the adversarial agent further in order to explore even richer, more dynamic behaviors (e.g., adversary motivations, environmental debris, and more). This work considers a complementary idea; development of a planning agent. The planning agent is envisioned as an auto-complete-like tool that can help accelerate security system design by human experts. The agent would respect existing barriers and sensors placed by a human expert while offering cost-effective suggestions (i.e., implicitly balancing effectiveness with cost) to improve the design. The goal is for this agent to be part of an expert’s toolbox, not to totally upend the current state-of-practice, or to displace human experts. The ultimate goal would be concurrent training of both the adversarial and planning agent together, to learn entirely through self-play. This would represent an entirely new way of performing system deign. We selected a hierarchical, model-based reinforcement learning algorithm to serve as the planning agent. This is an extension of concepts used in the prior FY24 adversarial agent work. There, we had a single agent acting an environment. Here, we have two different sub-agents (policies), working together, to form a complete agent. There is a manager policy, which can select abstract goals on slower time scales, and a worker, which performs primitive actions to reach goals selected by the manager. It is worth noting that this class of algorithm is challenging to work with. From our understanding, our work is one of the first successful uses of model-based reinforcement learning (MBRL) in nuclear energy1 , and likely the first hierarchical model-based reinforcement learning application in nuclear energy. Further, this work is one of the first known attempts to apply AI to perform a design tasks in nuclear energy. Consequently, there were significant implementation challenges and the bulk of the work was focused on successful implementation and algorithm design. The results presented here are very low technology readiness level as a consequence of the lack of related literature, but still represent a significant step forward in the pursuit of applied AI for design.

42 ENGINEERING↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

Calibration of multisite raters for prospective visual reads of amyloid PET scans

Abstract INTRODUCTION In multicenter Alzheimer's disease studies, amyloid positron emission tomography (PET) visual reads are typically performed centrally by a few experts. Incorporating a broader reader network enhances scalability and generalizability. METHODS Ten neuroimaging experts from eight Alzheimer's Disease Research Centers (ADRCs) visually read 180 amyloid PET scans (30 scans and 15 duplicate scans for each of four tracers, imaged across a wide variety of scanners), using preferred reading software without anatomical imaging or quantitation. Scans were classified as elevated or non‐elevated per tracer‐specific criteria. Inter‐ and intra‐rater agreement was assessed. RESULTS Inter‐rater agreement was substantial (Fleiss’κ = 0.78), with full consensus on 69% of scans. Inter‐rater reliability was substantial to perfect across tracers (Fleiss’κ = 0.70–0.87). Intra‐rater agreement was substantial to perfect (Cohen'sκ = 0.79‐1). Scans with intermediate (10–40 Centiloid) quantitation had lower reader agreement. DISCUSSION A multicenter expert network achieved substantial agreement classifying amyloid PET scans. These scans provide a standard for reader training and reliability assurance in future studies. Highlights Calibration methods ensure reliable amyloid positron emission tomography (PET) visual reads across multiple raters. Substantial agreement is possible across readers using their preferred tools. Agreement is also substantial regardless of the amyloid PET tracer used. Scans with intermediate (10–40 Centiloid) quantitation have lower reader agreement. The calibration set will become a training tool for amyloid PET visual read studies.

Neurosciences & Neurology↗

Bayesian SegNet for Semantic Segmentation with Improved Interpretation of Microstructural Evolution During Irradiation of Materials

Understanding the relationship between the evolution of microstructures of irradiated LiAlO2pellets and tritium diffusion, retention and release could improve predictions of tritium performance. Given expert-labeled segmented images of irradiated and unirradiated pellets, we trained Deep Convolutional Neural Networks to segment images into defect, grain, and boundary classes. Qualitative microstructural information was calculated from these segmented images to facilitate the comparison of unirradiated and irradiated pellets. We tested modifications to improve the sensitivity of the model, including incorporating meta-data into the model and utilizing uncertainty quantification. The predicted segmentation was similar to the expert-labeled segmentation for most methods of microstructural qualification, including pixel proportion, defect area, and defect density. Overall, the high performance metrics for the best models for both irradiated and unirradiated images shows that utilizing neural network models is a viable alternative to expert-labeled images.

Oostrom, Marjolein T.↗

Bimodal Visualization of Industrial X-Ray and Neutron Computed Tomography Data

Advanced manufacturing creates increasingly complex objects with material compositions that are often difficult to characterize by a single modality. Our collaborating domain scientists are going beyond traditional methods by employing both X-ray and neutron computed tomography to obtain complementary representations expected to better resolve material boundaries. However, the use of two modalities creates its own challenges for visualization, requiring either complex adjustments of bimodal transfer functions or the need for multiple views. Together with experts in nondestructive evaluation, we designed a novel interactive bimodal visualization approach to create a combined view of the co-registered X-ray and neutron acquisitions of industrial objects. Using an automatic topological segmentation of the bivariate histogram of X-ray and neutron values as a starting point, the system provides a simple yet effective interface to easily create, explore, and adjust a bimodal visualization. Here, we propose a widget with simple brushing interactions that enables the user to quickly correct the segmented histogram results. Our semiautomated system enables domain experts to intuitively explore large bimodal datasets without the need for either advanced segmentation algorithms or knowledge of visualization techniques. We demonstrate our approach using synthetic examples, industrial phantom objects created to stress bimodal scanning techniques, and real-world objects, and we discuss expert feedback.

image segmentation↗

Benchmarking large language models for materials synthesis: The case of atomic layer deposition

In this work, we introduce an open-ended question benchmark, ALDbench, to evaluate the performance of large language models (LLMs) in materials synthesis, and, in particular, in the field of atomic layer deposition, a thin film growth technique used in energy applications and microelectronics. Our benchmark comprises questions with a level of difficulty ranging from the graduate level to domain expert current with the state of the art in the field. Human experts reviewed the questions along the criteria of difficulty and specificity, and the model responses along four different criteria: overall quality, specificity, relevance, and accuracy. We ran this benchmark on an instance of OpenAI’s GPT-4o. The responses from the model received a composite quality score of 3.7 on a 1–5 scale, consistent with a passing grade. However, 36% of the questions received at least one below average score. An in-depth analysis of the responses identified at least five instances of suspected hallucination. Finally, we observed statistically significant correlations between the difficulty of the question and the quality of the response, the difficulty of the question and the relevance of the response, the specificity of the question, and the accuracy of the response as graded by the human experts. Furthermore, this emphasizes the need to evaluate LLMs across multiple criteria beyond difficulty or accuracy.

Artificial intelligence↗

From Rules to Reasoning: A Survey of Large Language Model-Based Approaches to Scientific Hypothesis and Idea Generation

Scientific hypothesis generation represents a fundamental challenge in contemporary research due to exponentially expanding literature volumes and increasing disciplinary specialization. Large language models (LLMs) have emerged as transformative tools for automated scientific discovery, moving beyond traditional rule-based and literature-mining approaches. Four paradigmatic approaches define current LLM-driven hypothesis generation: direct prompting and fine-tuning methods, knowledge-enhanced frameworks integrating retrieval-augmented generation (RAG), multi-agent collaborative systems simulating research teams, and reasoning-focused approaches implementing cognitive architectures. Domain-specific applications demonstrate statistical equivalence to human expert performance in social psychology, experimental validation in biomedical research, and near-expert quality in astronomy. Evaluation methodologies encompass human expert assessment, LLM-as-judge frameworks, and comprehensive benchmarking systems. Technical challenges include hallucination management, knowledge integration limitations, and balancing novelty with feasibility. Future directions emphasize hybrid neural-symbolic architectures and sophisticated human-AI collaboration models for responsible scientific discovery acceleration.

AI-driven discovery↗

Examining Cloud Feedback Components in the Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM)

Cloud feedback remains the main source of uncertainty in climate sensitivity estimated by global climate models (GCMs), largely because subgrid cloud responses are parameterized in GCMs due to their coarse resolution. Here, this study examines cloud feedback in the global 3.25-km Simple Cloud-Resolving Energy Exascale Earth System Model (E3SM) Atmosphere Model (SCREAM 3 km) through a pair of 1-yr atmosphere-only simulations with control and +4-K sea surface temperature perturbations. SCREAM 3 km produces a positive cloud feedback that falls within but at the upper end of the range of Coupled Model Intercomparison Project phase 5 (CMIP5) and CMIP phase 6 (CMIP6) models and expert judgment. The positive cloud feedback arises from positive contributions from both high- and low-level clouds, with increases in high-cloud altitude and decreases in low-cloud amount and optical depth playing key roles. The stronger-than-CMIP-average feedback is mainly attributable to the high-cloud altitude feedback, owing to cloud tops rising nearly isothermally in SCREAM 3 km. The positive low-cloud amount feedback is weaker in SCREAM than in GCMs because estimated inversion strength (EIS) increases more dramatically with warming. A coarser 12-km resolution version of SCREAM exhibits a weaker positive cloud feedback than SCREAM 3 km, mainly because its low-cloud-radiative flux is more sensitive to EIS, leading to a stronger negative low-cloud amount feedback. With this process-level assessment of cloud feedback, this study reveals where SCREAM aligns with and diverges from conventional GCMs and expert assessment, providing insights to inform further model improvement and future expert assessment.

Cloud radiative effects↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

CORE-CM in The Greater Green River and Wind River Basins: Transforming and Advancing a National Coal Asset (Final Report)

The following document summarizes project results from “CORE-CM in the Greater Green River and Wind River Basins: Transforming and Advancing a National Coal Asset”. This project is part of the U.S. Department of Energy’s (“DOE”) National Energy Technology Laboratory’s (“NETL”) Carbon Ore, Rare Earth Elements, and Critical Minerals (CORE-CM) Initiative. This report concludes that the Greater Green River and Wind River Basins (GGRB-WRB) Area-of-Interest (AOI 9) is the ideal region for continued research and development in progressing the broader CORE-CM goals outlined by the DOE. Based upon the extensive analyses of technical, social, and community criteria, this report illustrates that the GGRB-WRB hosts numerous potential CORE-CM feedstocks (both coal- and non-coal based), diverse opportunities for utilizing existing industrial waste streams, ample infrastructure and industry to support new CORE-CM-focused technologies, and a highly motivated, well educated, and adaptable workforce to further develop the regional and national CORE-CM supply chain. Additionally, some potential solutions for technological gaps suggest that the GGRB-WRB's diverse resources can play a significant role in achieving the national goal of critical materials independence. With full community participation, meaningful involvement of regional Tribal Nations, and building upon the stakeholder engagement demonstrated here, the GGRB-WRB region presents a unique opportunity for advancing the CORE-CM Initiative. This project was designed to bring together coal-based communities and stakeholders from across the GGRB-WRB to advance new industries for CORE-CM resources. The University of Wyoming (UWyo) School of Energy Resources (SER) led a project team of experts from the Colorado Geological Survey (CGS), Colorado School of Mines (CSM), Los Alamos National Lab (LANL), and local community colleges. Input from basinal, regional, and national experts bolstered the coalition in order to advance the mission of DOE’s CORE-CM initiative and develop the domestic CORE-CM supply chain. Phase I of this project was designed to address the goal of developing and catalyzing economic growth, job creation, and technology innovation in the GGRB-WRB of Wyoming and Colorado, by increasing the supply of CORE-CM to manufacturers of non-fuel Carbon Based Products (CBP) and products reliant upon CM. The GGRB-WRB CORE-CM project worked toward providing benefit through several avenues of performance and research. • Develop a coalition team to achieve project objectives • Complete detailed assessments, including State-of-the-Art (SOTA) Data acquisition of potential CORE-CM materials across the AOI, and meaningfully contributes to DOE’s CORE-CM goals nationally. • Strategic planning for regional economic growth, job creation, and associated technology innovation around coal materials, including plans to maximize the development of potential CORE-CM resources and technology by creating regional public-private partnerships. • Define regional economic growth potential around existing strengths, energy infrastructure, business and industry, including planning for the leveraging of highly trained workforces, existing and novel coal technologies, and energy infrastructure in development of CORE-CM supply chains. • Develop a preliminary strategic plan for increasing the supply of CORE-CM materials to manufacturers of non-fuel Carbon Based Products (CBP) and products reliant upon CM, focusing on regional strengths that result in an emerging diversified CORE-CM economy. • Assemble a committed network of stakeholders and communities that learn about, accept, and grow new energy technologies within coal regions. Additionally, the project team significantly contributed to the CORE-CM Initiative’s national goals, through cross-regional scoping, collaborating with CORE-CM projects in other AOIs, and including parallel regional project experts. In addition to active inclusion and meaningful engagement and contribution to DOE-led working groups, the project team focused on engaging with regional communities including Tribal Nations, economic development groups, and regional government organizations. The project’s CORE-CM development and commercialization plan identified diverse CORECM feedstocks, potential routes towards integration with existing industries, methods for supply-chain development that leverage existing infrastructure and businesses considering the regional economy, identified entry barriers for incorporating traditional and new technologies in those supply chains, recognized opportunities for public-private partnerships to develop technology innovation centers, identified diverse workforces, and conducted stakeholder outreach and education to build a community of understanding on CORE-CM potential in the GGRB-WRB region. Detailed task descriptions can be found in each chapter.

01 COAL, LIGNITE, AND PEAT↗

WoRDMAp Outbrief: Workshop on Radiation Detection Materials [Slides]

Project Goal: Identify pathway to reliably grow and fabricate high-performance radiation detectors for use in a variety of nonproliferation applications. 1. Create 3 working groups (semiconductors, inorganic and organic scintillators) that bring together international R&D subject-matter experts (national laboratories, industry, and academia), end users, and mission stakeholders 2. Reach expert consensus views regarding current materials related technology gaps and define prioritized R&D directions required to resolve these gaps 3. Envisage the future state of next-generation radiation detection materials and quantify the benefits to nuclear security applications 4. Provide a comprehensive expert report to DNN R&D program office to serve as roadmap with recommendations for future high-impact office investment in radiation-detection materials development

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Practical Probabilistic Programming

Recent advances in probabilistic programming languages (PPLs) have provided the capability for exact inference: computing a closed-form probability distribution for a given probabilistic program. In particular, the new language Roulette uses a language oriented programming (LOP) approach, wherein analysts build new programming languages on top of a set of primitives provided by Roulette, which then translates these structures into a weighted model counting problem which can be solved by automated reasoning tools. However, because Roulette provides few convenience features, developing these new languages is challenging even for expert users. We developed a standard library of common probability functions for Roulette with the goal of improved usability. This included approximation of continuous probability density functions using discrete probability mass functions. We demonstrated this approach by modeling a cosmic ray striking a RAM controller. We found that Roulette provides a powerful interface for highly expressive probabilistic programs to be generated. In collaboration with the NNSA Advanced Simulation and Computing program, which resulted in development of a tool called Circulette, we were able to model complex circuits expressed in Verilog using probabilistic programs with an expressivity not previously possible. Our research question that motivated the development of a Roulette standard library was to determine whether non-experts could use a PPL to model relevant problems regarding radiation effects on microelectronics. This standard library improved the expressivity of Roulette by implementing common probability density functions, mathematical operators on distributions, and support for empirical distributions. While Roulette is a powerful modeling language, the untyped, LOP approach makes error messages difficult to understand and requires expert aid. We recommend further research on Roulette, especially with its error messages, to enable improved usability. At the same time, this project demonstrated that for users familiar with Roulette and the LOP approach, Roulette provides powerful new capabilities that can be integrated with other Sandia modeling capabilities.

97 MATHEMATICS AND COMPUTING↗

Validating Greater Sage-Grouse Individual-based Model (IBM) Tool (Final Report)

The project focused on validating the previously developed Greater Sage-Grouse Individual-based Model (GrSG IBM; LaGory et al. 2012, 2021). The objective was to transform this predictive, spatially and temporally explicit model into a portable resource to assist siting/resource managers in proactively assessing the cumulative impacts of wind energy development on the greater sage-grouse. Utilizing a bottom-up, individual-based approach, the GrSG IBM accounts for landscape context and species behavior, aiming to reduce uncertainty in estimating development impacts and support ecologically mindful land-based wind energy development. The validation effort covered approximately 6,540 km 2 near the Seven Mile Hill Wind Project in Wyoming. The GrSG IBM tool, built on the NetLogo platform (Tisue and Wilensky 2004), was executed over a 50-year period, with the analysis focusing on years following a 10-year initialization phase. Key results demonstrated the tool’s biological soundness across five key biological metrics: non-chick age class distribution (older than 10 weeks), adult sex ratio, life expectancy, population size, and overall population growth. For instance, the tool estimated that 58.6% of the non-chick population was reproductively immature, while the reference ranges from 51.4% to 57.8% (Patterson 1952, Rogers 1964). Experts confirmed the tool’s estimate was within a reasonable range for the species. The tool estimated average life expectancy of 1.43 years, while the reference ranges from 0.9 years to 1.1 years (Ammann 1957, Hamerstrom 1949). Experts also supported the model’s life-expectancy estimate as ecologically sound for the species in the study area. In terms of population change, the model estimated an annual shift between a 0.6% decline and a 1.0% increase over 50 years. While the reference suggests 2.9% annual decline in range-wide populations (Cortes et al. 2023), that includes many at-risk populations in South Dakota and Washington, for example. Our study area—in the northeastern part of Carbon County and western-edge of Albany County, Wyoming—is one of the remaining greater sage-grouse habitats supporting some of the most stable populations. Experts confirmed that the range of the annual population change spanning from a 0.6% decline to a 1.0% increase estimated by the tool was reasonable for our study area for this reason and confirmed that aligned with population estimates from existing studies on the greater sage-grouse and wind energy development in the study area (LeBeau et al. 2017a, Smith et al. 2024). Furthermore, the project showed that temporally explicit biological metrics generated by the GrSG IBM tool can complement the USGS’ Prioritizing Restoration of Sagebrush Ecosystems Tool (PReSET; Duchardt et al. 2021) by incorporating habitat restoration strategies into seasonal habitat suitability models to visualize population responses over time.

17 WIND ENERGY↗

A Scientist-in-the-Loop Data Analytics Framework for Intelligent Simulation Model Tuning and Validation

This project developed a scientist-in-the-loop data analytics framework for intelligent simulation model tuning and validation, targeting the Weather Research and Forecasting (WRF) model and its solar energy variant, WRF-Solar-BNL. Domain experts, such as climate scientists, depend on large-scale numerical simulations for knowledge discovery and decision-making, yet the complexity of parameter tuning and the disconnect between automated optimization and domain expertise pose significant challenges. We extended an interactive visual analytics framework that enables domain experts to observe and intervene in the computational steering process by identifying disagreements between the simulation model, surrogate model, and the expert’s domain knowledge. Using Bayesian Optimization with Gaussian Process Regression as the surrogate model, our system allows users to probe parameter relationships, analyze correlation patterns, and adjust tuning parameters in real time. We developed use cases for solar irradiance forecasting through sustained collaboration with Brookhaven National Laboratory, resolving critical model configuration challenges and achieving meaningful reductions in prediction error. The project supported one PhD student, one MS student, and eight undergraduate students across three Data Science Capstone projects, resulting in one master’s thesis.

Dasgupta, Aritra [New Jersey Institute of Technolo↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗

A Risk-Informed Approach to Trustworthiness Assessment in Digital Twins-Based Autonomous Control

In autonomous control systems, digital twins (DTs) are used to perform diagnostic and prognostic functions. The trustworthiness of these DTs is dependent on quality and coverage of the training data, model accuracy and integrity of sensor data. This work introduces a methodology to determine the trustworthiness of a DT system given faulty sensor data using a risk informed approach. Bayesian Belief Networks (BBNs) are used to propagate uncertainties and determine the probability of trustable recommendations. The decision to trust the control action provided by the DT is based on the DT output, expert opinion, and severity of problems. The performance of DTs is reliant on the data they are trained on. When they encounter out of distribution data, the trustworthiness of the recommendations decreases. To address this issue, we include an expert component that provides input on sensor degradation. For this, we utilize a generative artificial intelligence (AI) model, such as Generative Pretrained Transformer (GPT). The GPT functions as an expert with broad knowledge. The GPT is fine-tuned to understand and discriminate sensor degradation scenarios using manufactured data. This methodology is demonstrated through a case study on a Nearly Autonomous Management and Control System (NAMAC) during a steady state scenario. Various sensor degradation types with different severity levels are considered. Degraded sensor data is processed by the DT system and the fine-tuned GPT. Finally, using the BBN, we combine the GPT information and the DT output with its sources of uncertainty. This provides an output regarding the trustworthiness of the DT recommendation.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

A comparison of surrogate constitutive models for viscoplastic creep simulation of HT-9 steel

Mechanistic microstructure-informed constitutive models for the mechanical response of polycrystals are a cornerstone of computational materials science. However, as these models become increasingly more complex – often involving coupled differential equations describing the effect of specific deformation modes – their associated computational costs can become prohibitive, particularly in optimization or uncertainty quantification tasks that require numerous model evaluations. To address this challenge, surrogate constitutive models that balance accuracy and computational efficiency are highly desirable. Data-driven surrogate models, that learn the constitutive relation directly from data, have emerged as a promising solution. In this work, we develop two local surrogate models for the viscoplastic response of a steel: a piecewise response surface method and a mixture of experts model. These surrogates are designed to adapt to complex material behavior, which may vary with material parameters or operating conditions. The surrogate constitutive models are applied to creep simulations of HT-9 steel, an alloy of considerable interest to the nuclear energy sector due to its high tolerance to radiation damage, using training data generated from viscoplastic self-consistent (VPSC) simulations. In conclusion, we define a set of test metrics to numerically assess the accuracy of our surrogate models for predicting viscoplastic material behavior, and show that the mixture of experts model outperforms the piecewise response surface method in terms of accuracy.

36 MATERIALS SCIENCE↗