Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “text analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Informing Plant Asset Reliability and Availability Through AI-Driven Analysis of Operator Logs

The availability and reliability of nuclear power plant (NPP) structures, systems, and components (SSCs) are critical parameters for NPP safety. Tracking these parameters is necessary but costly and labor-intensive, requiring the collection and evaluation of SSC event data such as shutdowns, startups, and failures. To show how these events are needed for the parameters an example is given: one measure of reliability is based on the number of equipment failure events and the number of run hours (i.e., the time from a startup event to a shutdown event). Here, this work investigates using artificial intelligence (AI) to mine NPP operator log entry texts for SSC event data. Four AI approaches were explored for identifying these events, including natural language processing (NLP) methods, generative AI, generative AI combined with NLP, and topic modeling. A key challenge addressed with all four approaches is the brevity of operator log entries. Among these four a neural network–based NLP method was shown to be the most promising for this application, achieving F1 scores of 86.0% for shutdowns, 92.2% for startups, and 80.4% for failures on a subject-matter-expert-curated dataset from NPP operator logs, compared to a baseline of 66.6% for a random classifier. This shows that NLP methods can perform better than generative AI. Additionally, the NLP methods combined with generative AI were shown to perform better than generative AI alone. Generative AI was most successful at providing the background information for the NLP methods to use. This work demonstrates the potential to use AI to automate parameter collection from NPP operator log entries and other records.

97 - MATHEMATICS AND COMPUTING↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

Agent-Based, Bottom-Up Medium- and Heavy-duty Electric Vehicle Economics, Operation, Charging and Adoption (Research Performance Final Report)

This is the research performance final report for the project entitled: Agent-Based, Bottom-Up Medium- and Heavy-duty Electric Vehicle Economics, Operation, Charging and Adoption This project was able to achieve the DOE’s goals of developing new modeling tools to understand MDHD vehicle operation and adoption. The first modeling tool is a fleet-level techno-economic analysis model capable of estimating energy use and associated environmental and cost impacts for electrified and conventional vehicles of any MDHD vocation, using real-world cost and operations data, including approaches to optimizing schedules for charging and/or vehicle dispatch. The second modeling tool is a system-level, bottom-up, agent-based adoption model capable of generating geographically-resolved estimates of market projections for MDHD vehicles and charging infrastructure. These tools will be developed and published to serve dual purposes as analysis tools for researchers, and decision-support tools for decision makers within the MDHD system.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

IMAGINE BioSecurity: Mesocosm-Based Methods to Evaluate Biocontainment Strategies and Impact of Industrial Microbes Upon Native Ecosystems

Project Goals: The Integrative Modeling and Genome-scale Engineering for Biosystems Security (IMAGINE BioSecurity) SFA project seeks to establish an understanding of the behavior of engineered microbes in controlled versus environmental conditions to predictively devise new strategies for responding to biological escape. To this end, the IMAGINE Team has established a plant-soil mesocosm platform to track and quantify the fate of industrial microbes in environmental systems and assess the efficacy of biocontainment constraints upon genetically engineered microbe escape frequency and the impact of industrial microbes upon native ecological microbiomes. Abstract Text: Genetically modified industrial production microbes and their associated bioproducts have emerged as an integral component of a sustainable bioeconomy. However, the rapid development of these innovative technologies raises biosecurity concerns, namely, the risk of environmental escape. Thus, the realization of a bioeconomy hinges not only on the development and deployment of microbial production hosts, but also on the development of secure biosystems and biocontainment designs. Current laboratory-based biocontainment testing systems do not accurately reflect complexities found in natural environments, necessitating an environmentally relevant analysis pipeline that allows for the detection of rare escapees, the effect of associated bio-products, and the impact on native ecologies. To this end, we have developed an approach that utilizes soil mesocosms and integrated systems analyses to evaluate the efficacy of novel biocontainment strategies and to assess the impact of production systems upon terrestrial microbiome dynamics. We demonstrate the utility of this approach by modeling a contamination with industrial microbial chasses versus their biocontained counterparts. Here we demonstrate the broad utility of this system by highlighting findings from both strains of Saccharomyces cerevisiae that are contained with an inducible toxin anti-toxin system, and stains of Escherichia coli that are contained via genomic recoding. The resultant data demonstrate that this system has broad utility across diverse microbial chassis and biocontainment strategies, enables us to track the fate of our contaminating microbe with high sensitivity in the soil, as well as monitor broader impacts of the perturbation on the underlying soil system. The findings presented here support the use of this mesocosm-based approach to assess the environmental impact of industrial microbes and to validate biocontainment strategies.

BASIC BIOLOGICAL SCIENCES,INORGANIC, ORGANIC, PHYS↗

EMP, Attachment 1: Sampling and Analysis Plan (Rev.1)

This Sampling and Analysis Plan (SAP) is written for the Environmental Radiation Task activities related to radioactive air emissions (stack) monitoring and environmental radiological ambient air surveillance of Pacific Northwest National Laboratory (PNNL) operations at the PNNL-Richland campus and PNNL-Sequim campus. PNNL is a U.S. Department of Energy Office of Science laboratory in Richland, Washington. This plan is an attachment to PNNL’s Environmental Radiological Air Monitoring Plan (EMP) (PNNL-20919) and addresses a discrete, vital subject area that is subject to revision independent of the main text of the EMP document. This SAP provides the requirements for planning sampling events and the requirements imposed on the services provided by the analytical laboratory to the PNNL Environmental Radiation Task.

40 CFR 61 Subpart H↗

Optimizing Geospatial Assessments for Nuclear Safeguards Applications with Large Language Models

A multidisciplinary team at Argonne National Laboratory evaluated the ability of large language models (LLMs) to identify geographic locations from open-source text and assessed post-processing measures to strengthen the reliability of those extractions in support of international nuclear safeguards. The study focused on addressing challenges such as toponym ambiguity, imprecise descriptions, and misinformation, which often undermine the accuracy of LLM-derived geospatial assessments. By integrating authoritative geospatial datasets, employing rigorous validation techniques, and leveraging human-in-the-loop processes, the project aimed to enhance the precision, transparency, and reproducibility of geospatial localization workflows. The findings demonstrate that while LLMs exhibit significant potential for accelerating geospatial analysis, their outputs require systematic grounding and verification to ensure reliability in high-stakes applications. This work contributes to the broader field of geospatial intelligence and supports strategic objectives of international organizations such as the International Atomic Energy Agency (IAEA) and the U.S. Department of Energy (DOE).

97 MATHEMATICS AND COMPUTING↗

Cyote-attack Chain Estimator

Attack Chain Estimator (ACE) Application Overview The Attack Chain Estimator (ACE) Application is a sophisticated tool designed for the ingestion, classification, sequencing, and enrichment of cybersecurity threat reports. This application leverages advanced machine learning models and extensive historical data to provide comprehensive insights into cyber threats, specifically targeting Industrial Control Systems (ICS). Purpose The primary functions of the ACE Application include: Ingestion of Cybersecurity Threat Reporting: Capable of ingesting text-based threat reports in markdown or text file format. Supports ingestion of structured data from other sources in STIX/JSON format. Classification of Report’s Text-Based Events: Utilizes a DeBERTa classifier, specifically trained on cybersecurity data, to map the events to MITRE ATT&CK for ICS Tactics and Techniques. Classification is performed using multiple Jupyter notebooks and machine learning workflows hosted as FastAPI microservices: regex_data deberta_base_35_train_hft_classifier_mlflow.ipynb hft_regex_classifier_mlflow.ipynb param_train_hft_classifier_mlflow.ipynb regex_tactic_tech.ipynb Ordering of Tactics, Techniques, and Observable Events: Sequences the identified tactics, techniques, and events to form a coherent attack chain. Enrichment with Historical Attack Chain Details: Enhances the attack chain with details from historical attacks using a Markov model developed from CyOTE Precursor Analysis Report data. The Markov model is available as a FastAPI endpoint for seamless integration. Enrichment with Adversary Emulation Capabilities Data: Integrates adversary emulation capabilities data using MITRE Caldera for OT adversary abilities UUIDs. Export of Output Files: Provides options to export the enriched attack chain in JSON or CSV formats. Routing of Output to Other Applications: Facilitates routing of output to various platforms and applications, including: Threat Intelligence Platforms COREII Scout for Threat Intelligence Analysis COREII Modeling and Simulation for Adversary Emulation Technical Description The ACE Application is an advanced cybersecurity tool designed to provide detailed threat analysis and sequence generation. It is built on a robust architecture that integrates natural language processing, machine learning, and historical data modeling. Key Components: Data Ingestion Module: Handles the input of threat reports and data from various formats, ensuring flexibility in data sources. Classification Engine: Employs DeBERTa-based classifiers hosted as FastAPI microservices to analyze and classify threat report events in accordance with the MITRE ATT&CK framework for ICS. Sequence Generator: Orders the classified events into a logical attack chain, providing clear insight into the sequence of tactics and techniques used in the threat. Enrichment Engine: Integrates historical data and adversary emulation capabilities to enhance the attack chain with valuable context and additional details. The historical data enrichment is powered by a Markov model, which is available as a FastAPI endpoint. Export and Routing Module: Facilitates the export of the enriched attack chain in multiple formats and routes the output to designated applications for further analysis or emulation.

Paul, Tony [Idaho National Laboratory (INL), Idaho↗

Dataset for Cruz-O'Byrne et al (2026): "Divergent biogeochemical responses in upland coastal forest soils to repeated flooding and shifts in water chemistry"

Hydrologic disturbances from accelerated sea-level rise and the increasing frequency and intensity of storms and tidal flooding are altering biogeochemical processes in upland coastal forests, transforming these ecosystems into wetlands. However, the initial effects of flooding on belowground biogeochemistry and the mechanisms driving greenhouse gas dynamics and soil organic matter stability during the early stages of this transition remain poorly understood. This dataset presents the results of a mesocosm experiment conducted in a controlled, highly instrumented laboratory environment, in which freshwater and brackish water pulses were applied to intact soil monoliths from a temperate upland coastal forest to examine how floodwater chemistry influences soil biogeochemistry and organo-mineral interactions. All data files are plain-text CSV (comma-separated value), and no special software is required to read them. Details about the content of each file are available in the document “Dataset_readme”. The dataset consists of the following data: • rcruzobyrne_moisture: Soil volumetric water content (VWC) • rcruzobyrne_GHG: Headspace greenhouse gas (GHG) concentration and fluxes • rcruzobyrne_methane_isotopes: Headspace methane isotope signature • rcruzobyrne_porewater: Porewater chemistry • rcruzobyrne_CDOM: Porewater colored dissolved organic matter (CDOM) • rcruzobyrne_FTIR: Soil Fourier-transform infrared (FTIR) spectroscopy Details of the experimental setup, data collection, and data analysis are provided in the manuscript by Cruz-O’Byrne et al (2026) Divergent biogeochemical responses in upland coastal forest soils to repeated flooding and shifts in water chemistry. Biogeochemistry. https://doi.org/10.1007/s10533-026-01340-0

EARTH SCIENCE > ATMOSPHERE > GREENHOUSE GAS↗

Dataset for scientific paper "Simulated plant‑mediated oxygen input has strong impacts on fine‑scale porewater biogeochemistry and weak impacts on integrated methane fluxes in coastal wetlands", a modeling study based on field observation at the tidal salt marshes of the Parker River Estuary, Massachusetts, United States

This dataset is the raw and processed data for the paper "Simulated plant ‑ mediated oxygen input has strong impacts on fine ‑ scale porewater biogeochemistry and weak impacts on integrated methane fluxes in coastal wetlands". This study investigated how plant-mediated oxygen input affects subsurface biogeochemical reactions of organic carbon degradation and the resulting methane emissions of coastal wetlands by model simulation. We used the subsurface geochemical simulator PFLOTRAN for the modeling, which produced the simulated changes in porewater chemical substances and methane emissions over 10 days under different scenarios of plant-mediated oxygen input.Specifically, this dataset contains: 1) the input files for PFLOTRAN of all simulation runs conducted in this study. Those files are with an extension of ".in", containing information of the biogeochemical reaction network (stoichiometry, reaction rate, Monod constants, etc), fluid flow rate and oxygen concentration in the fluid which together simulated the plant-mediated oxygen input, the configuration of artificial reactions that simulated the methane fluxes, etc. The PFLOTRAN input files are text files, which can be opened by NotePad, but running these input files will require proper installation of PFLOTRAN (instruction: https://documentation.pflotran.org/user_guide/how_to/installation/installation.html). 2) the raw and processed model output from PFLOTRAN of all simulation runs, and 3) the python scripts used to process the raw model output, including random allocation of root cells, converting raw data into organized formats, calculating the methane fluxes based on the model output, data visualization, etc. The raw and processed model output from PFLOTRAN are in .spydata format, which can be viewed with Python. and 3) the python scripts for data processing and analysis are programming scripts, which can be opened with Python.This modeling work, in particular the model parameterization of root density and initial conditions of porewater concentrations of biogeochemical substances, was based on field measurements at the salt marsh of the Upper Parker River Estuary, Massachusetts, United States.

54 ENVIRONMENTAL SCIENCES↗

AstraAI v1

AstraAI is an open-source, structure-aware AI coding agent designed for large scientific and DOE-HPC codebases such as AMReX-based applications. Unlike general-purpose coding assistants, AstraAI combines retrieval-augmented generation (RAG) with compiler-level Abstract Syntax Tree (AST) analysis to perform precise, scope-constrained code modifications. It identifies exact function spans, enforces locality of edits, and maintains cross-file invariants, enabling deterministic and build-safe transformations in complex C++/GPU environments. AstraAI is intended for developers working on large, evolving HPC frameworks where correctness, reproducibility, and structural integrity are critical. Typical use cases include modifying physics kernels, updating GPU device lambdas, and performing multi-file refactors without breaking compilation or runtime semantics. Compared to conventional LLM-based coding agents - even those with repository access - AstraAI provides structural guarantees rather than free-form text patches. It minimizes unintended diffs, prevents scope drift, preserves formatting and build stability, and reduces structural hallucinations. By integrating compiler tooling directly into the generation loop, AstraAI transforms AI-assisted coding from probabilistic text editing into deterministic, structure-preserving program transformation suitable for mission-critical scientific software.

Natarajan, Mahesh [Lawrence Berkeley National Labo↗

Experimental study of ECH pre-ionization on J-TEXT

An experimental study on electron cyclotron heating (ECH) pre-ionization has been conducted on J-TEXT in support of the joint experiment research for ITER plasma initiation. In this experiment, ECH power was injected to the vessel before the application of loop voltage ionize the neutral gas and form the initial plasma or so-called pre-plasma. The impact of several significant factors, such as magnetic field configuration, pre-fill gas pressure, ECH toroidal injection angle and ECH power on the evolution of pre-plasma are systematically studied, aiming to identify shared features, clarify their potential relationship and optimize the discharge parameters to generate a rather high pre-plasma density. By separating ECH power from the inductive start-up, the effect of pre-plasma on tokamak start-up can be observed. A dynamic magnetic configuration facilitates the transition of pre-plasma to tokamak plasma. To assess the influence of pre-plasma density on tokamak start-up, two kinds of magnetic field configurations are examined. While the effect of different pre-plasma densities on tokamak start-up is negligible, a significant difference is observed between pure ohmic start-up and start-up with pre-ionization. The studies presented here show evolutionary trends and threshold values needed to optimize ECH pre-ionization and also a feasible way to improve pre-plasma density and a configuration to stabilize pre-plasma for transition. Eventually, these results may contribute to multi-machine research and physics analysis that can assist ITER with its optimal preparation for first plasma operation.

ECH↗

Scaling open-weight large language models for hydropower regulatory information extraction: A systematic analysis

Information extraction from regulatory and technical documents using large language models (LLMs) involves practical trade-offs between extraction quality and computational cost. We evaluate eight open-weight LLMs spanning 0.6B–70B parameters on hydropower licensing documents and report deployment-oriented evidence under a unified extraction schema and evaluation protocol. Across the model set, we observe clear scale-dependent trends in both baseline extraction quality and the effectiveness of reflective reasoning (self-checking) under our fixed-prompt, no-augmentation setting. Mid-scale models often provide a favorable balance of accuracy and efficiency, whereas the smallest models show limited or inconsistent gains from the reasoning variants tested. Larger models achieve the highest overall F1 scores but incur substantially greater compute and infrastructure requirements. We further find that reliability failure modes can distort conventional metrics in this domain: in particular, high recall can coincide with systematic extraction errors when models fabricate values for fields that are absent from the source text, underscoring the importance of conservative null handling and evidence-grounded evaluation. Overall, our study provides a reproducible resource–performance comparison for open-weight LLM-based extraction in hydropower regulatory documentation and offers practical guidance for model selection under different deployment constraints.

Evaluation protocol↗

Historic climate, cosmogenic 10Be, denudation-rate, and geospatial datasets from the Pikes Peak region, Colorado, USA

This data package contains geographic information system (GIS) layers and tabular datasets associated with the study of elevation-dependent denudation rates on Pikes Peak in the Front Range of the Rocky Mountains, Colorado, USA. The package includes GIS layers used to produce the study-area map, including sample locations, sample watershed boundaries, the Pikes Peak batholith, Pleistocene glacier extent, weather station locations, and elevation and hillshade rasters, together with comma-separated value (CSV) tables and matching CSV data dictionaries. These mapped layers provide the geographic framework for interpreting denudation patterns across the Pikes Peak region and for relating sample locations to watershed geometry, bedrock setting, glacial history, and nearby climate stations. The first group of tables reports climate and geospatial context for the study area. These files include station-based temperature and precipitation data used to characterize elevational gradients in mean annual climate and monthly climate seasonality, sample locations, denudation-rate and topographic metrics, fixed frost-cracking model parameters, frost-cracking intensity and precipitation-frequency metrics, and stream-power inversion results. Together, these data provide the basis for evaluating how denudation varies with elevation, climate, and landscape form across sampled catchments on Pikes Peak. The second group of tables reports cosmogenic nuclide and erosion-model results used in the denudation analysis. Included files contain accelerator mass spectrometry (AMS) measurements for in situ-produced cosmogenic beryllium-10 (10Be), including sample identifiers, measured 10Be:9Be ratios, analytical uncertainties, carrier mass, quartz mass, blank corrections, blank-group statistics, and calculated 10Be concentrations and uncertainties. Additional tables summarize stream-power-law inversion results for sampled catchments, including optimized model parameters, predicted erosion rates, residual metrics, channel-pixel counts, and convergence status, as well as regression equations and summary statistics used to evaluate relationships among elevation, climate, frost cracking, precipitation forcing, and denudation rate. The package contains GIS files, comma-separated value files (.csv), Microsoft Excel files (.xlsx), CSV data dictionaries, a file-level metadata table, and a readme text file.

10Be cosmogenic nuclides↗

Legacy Effects of Cropping System and Precipitation Influence the Core Camelina sativa Microbiome

Camelina ( Camelina sativa L.) is a potential biofuel crop and beneficial rotation crop in dryland cropping systems. Little is known about camelina microbiota or the legacy effect of soil origin/cropping system zones on camelina-associated microbiome assembly. To explore camelina-microbe associations, we grew camelina in the greenhouse using soil transplanted from 33 locations in the dryland wheat production area of eastern Washington. Bacterial, archaeal, and fungal communities from bulk soil, rhizosphere, and endosphere were characterized with 16S rRNA and internal transcribed spacer amplicon sequencing and were analyzed alongside site-specific climatic and edaphic data. We found that soil from the highest precipitation zone had higher alpha diversity than soil from the driest zone, but this effect was not seen in the greenhouse rhizosphere or endosphere. Plant compartment, cropping system zone, and soil origin all significantly influenced microbial composition, with soil pH and organic matter, as well as precipitation at origin, as major predictors. Analysis of abundance–occupancy distributions showed that the Actinobacteriota Aeromicrobium and Marmoricola and the fungus Pseudogymnoascus in the rhizosphere were plant-selected, while the endosphere was characterized by a number of Actinobacteriota, Rhizobium, and Clostridium. Sphingomonas amplicon sequence variants were also consistently enriched in the rhizosphere, suggesting that they are present in soils collected throughout eastern Washington and may represent good candidate biostimulants. Several lignin decomposing fungi had site-specific rhizospheric distributions, suggesting that they may be dispersal-limited or result from the legacy effect of long-term wheat cropping. Overall, this study contributes to our understanding of microbiome assembly in and on camelina roots while also highlighting the potential impact of cropping history on soil- and plant-associated microbiomes. [Formula: see text] The author(s) have dedicated the work to the public domain under the Creative Commons CC0 “No Rights Reserved” license by waiving all of his or her rights to the work worldwide under copyright law, including all related and neighboring rights, to the extent allowed by law, 2025.

Barnes, Elle M↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

PFLOTRAN modeling data and scripts associated with “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the publication “Refining the Hydrogeologic Framework of a Large River Corridor Model Using Waterborne Transient Electromagnetics” submitted to Water Resources Research (Terry et al. 2025). The data package contains the groundwater modeling dataset from PFLOTRAN software. It includes the python script for mesh generation, boundary condition setting, PFLOTRAN input deck formation and postprocessing. It couples groundwater flow and species transport for Hanford Reach river corridor and pipelines the model generation and processing. This model can be used to easily generate the model and analysis for Hanford site. It can also be adjusted to other hydrologic area with ease. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package consists of 6 folders: (1) “data” contains all necessary data as input and intermediate data for processing; (2) “mesh” contains all mesh related files to generate mesh in Hanford Reach river corridor; (3) “model_run” contains the generated script for PFLOTRAN modeling; (4) “notebooks” contains all the Python script to generate the model; (5) “output” contains all the output from the computation; (6) “postprocessing” contains the Python script to generate scientific figure for manuscript. All files are .csv (comma-separated values), .h5 (HDF5 format), .in (input files), .ipynb (Jupyter notebooks), .p (Python pickle), .png (images), .PNG (images), .py (Python scripts), .pyc (Python bytecode), .r (R scripts), .sh (shell scripts), .txt (text files), .vtu (3D mesh/visualization format), .xz (compressed archive), or .zip (compressed archive).

54 ENVIRONMENTAL SCIENCES↗

Control And Optimization Modular Modeling Application For Nuclear Deployment

The purpose of the COMMAND code is to provide a flexible, scalable tool for use in developing, integrating, and testing the technologies necessary for achieving autonomous operations of advanced nuclear reactors. The code enables users to efficiently implement custom simulations and experiments by combining key methods from different software modules. These modules are focused on: modeling and simulation tools, such as nuclear simulation tools used for high-fidelity modeling (e.g., Reactor Excursion and Leak Analysis Program [RELAP5-3D] and Monte Carlo N-Particle [MCNP]); machine learning and optimization tools (e.g., anomaly detection and data-driven modeling techniques); advanced control in its digital, high-performance, and supervisory control forms (e.g., proportional integral derivative (PID) control and model predictive control (MPC); and integration with hardware through industrial communication protocols. To ensure flexibility and scalability, COMMAND was designed to be both modular—the software “pieces” all inherit from generic building blocks and can be combined and connected to create complicated simulations—and high performing—designed for parallel processing, enabling simulations and experiments to take advantage of multi-core computers, servers, and nodes. The code is written in the Python programming language due to the language's popularity, active community, and open-source and cross-platform nature. Maintaining consistency with other simulation tools used within the nuclear energy community, users implement simulations and experiments through text input files, which define components, parameters, connections, etc., through lines of text. Given that COMMAND is written in Python, these input files are native Python scripts, and so use the standard Python structure and formatting. This also enables users to take advantage of Python's extensive package library to develop custom capabilities for their specific use cases.

Faber, Jacob [Idaho National Laboratory (INL), Ida↗

Generative AI for Power Grid Operations

Generative artificial intelligence (AI) has captured into the mainstream, demonstrating capabilities that once belonged solely to the realm of human cognition. From defeating world champions in complex games to generating human-quality text and images, Generative AI has proven its potential to revolutionize countless industries. The electric power grid is no exception. Generative AI's ability to process vast amounts of data rapidly, assist decision support and identify patterns could significantly enhance power grid operations. For example, Generative AI could improve state estimation where measurements are not available or integrate renewable energy sources more efficiently with probabilistic forecasting. The key contributions of this whitepaper are outlined below: (1) Comprehensive overview of Generative AI's applications in power grid operations: It highlights the opportunities in areas such as forecasting, state estimation, and demonstrating the potential for enhancing efficiency, reliability, and resilience. (2) Expanding Generative AI's impact through synergies with emerging technologies: The paper introduce NREL developed eGridGPT and explores how AI orchestration, multi-agent systems, and Digital Twins can collaborate to optimize grid operations, addressing the complexities of a decarbonized and electrified future. (3) In-depth analysis of challenges in implementing Generative AI: This includes considerations like data availability and quality, model validation, certification, and ethical concerns, ensuring responsible AI deployment. (4) Emphasizing human-AI collaboration: The whitepaper underscores the importance of trustworthy, transparency, and explainability in AI systems to promote seamless interaction between human operators and AI, ultimately improving decision-making. (5) Exploring future research and development: It identifies critical areas for further advancement to fully realize Generative AI's potential in power grid operations. This whitepaper serves as a valuable resource for researchers, practitioners, and policymakers looking to harness Generative AI for a more reliable, stable, and cost-effective power grid.

24 POWER TRANSMISSION AND DISTRIBUTION↗