Engineering PapersSearch

SEARCH · Engineering Papers

Results for “software requirements”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Myco-CORPSE simulations assessing mycorrhizal carbon allocation across U.S. forests and global change scenarios

Plants allocate a substantial portion of their fixed carbon belowground to mycorrhizal fungi in exchange for nutrients and other benefits. However, most current ecosystem models omit mycorrhizal processes, limiting our ability to predict plant–soil carbon dynamics under environmental change. To address this gap, we used a mycorrhiza-explicit soil biogeochemical model, Myco-CORPSE (Mycorrhizal Carbon, Organisms, Rhizosphere, and Protection in the Soil Environment), to simulate tree carbon allocation to arbuscular mycorrhizal (AM) and ectomycorrhizal (ECM) fungi in temperate forests.The dataset includes outputs from two sets of model simulations:1. Perturbation experiments: Simulations across gradients of ECM dominance (0–100%), nitrogen deposition, soil temperature, and net primary productivity (NPP) to test how these factors affect mycorrhizal C allocation and nutrient cycling.2. FIA-based simulations: Model applications to over 1,800 U.S. forest sites using site-specific data from the U.S. Forest Inventory and Analysis (FIA) program, including vegetation composition, mycorrhizal type, climate, litter traits, soil properties, and N deposition.Model outputs include simulated mycorrhizal carbon allocation and related biogeochemical variables, such as soil and microbial carbon and nitrogen stocks. Data are provided in CSV format and organized by experiment type (in separate ZIP files). Python scripts for running simulations, plotting, and spatial mapping are also included and organized similarly. No proprietary software is required. These outputs support a peer-reviewed study and were used to generate figures and tables in the associated publication.

54 ENVIRONMENTAL SCIENCES

Data for Myers-Pigg et al. (2026), "Short-term coastal forest responses to a hurricane-scale freshwater and saltwater flooding experiment"

Coastal upland forests are exposed to intensifying precipitation regimes and sea level rise, increasing tree mortality and transforming these coastal forests into wetland ecosystems. Despite these well-known risks, the differing degrees to which hydrological, biogeochemical, and biological components of upland forests respond to novel salinity exposure is relatively unknown. The Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experiment decouples two distinct disturbances associated with hydrological extremes: (1) flooding from heavy precipitation and (2) exposure to saline conditions from storm surge. This dataset includes data reported in Myers-Pigg et al. (2025), which analyzed data from the first TEMPEST flooding treatment in 2022. This includes: - Colored dissolved organic matter in porewaters - Soil temperature and oxygen - Groundwater temperature and chemistry - Dissolved organic carbon concentrations in porewaters - Soil-to-atmosphere CH4 and CO2 fluxes - Soil temperature, water content, and electrical conductivity - Root-influenced CH4 and CO2 flux - Tree sap flow velocity - The R analytical code and documentation about the computational environmental in which it was run (the "sessionInfo.txt" file) All data files are plain-text comma separated value (CSV) and no special software is required to read them.

54 ENVIRONMENTAL SCIENCES

Data from a throughfall exclusion experiment: Fine root dynamics, morphology, chemistry, and AMF colonization across four lowland Panamanian forests

Fine roots regulate forest nutrient, carbon, and water cycling, yet their variation within and among tropical forests remains under-characterized. We quantified root productivity, disappearance, and stocks to 1 m using minirhizotron imaging, and we measured morphology, elemental composition [root carbon (C), root nitrogen (N), root phosphorus (P)], and arbuscular mycorrhizal fungi (AMF) colonization to 20 cm using ingrowth cores and sequential coring. Sampling took place in four distinct lowland Panamanian forests (32 plots; 8 per forest) from 2018 through 2022 under control and throughfall-exclusion (drought) treatments in the Panama Rainforest Changes with Experimental Drying (PARCHED) experiment.The dataset is presented as an Excel workbook with six tabs. The first tab is the data dictionary. Tab S1 contains ingrowth-core production and mortality, morphology and soil moisture. Tab S2 contains sequential-coring standing stocks with associated morphology and soil moisture. Tab S3 contains minirhizotron row data records to 1 m depth, including per-frame root length and diameter, normalized length metrics, and session timing. Tab S4 contains AMF colonization. Tab S5 contains fine-root chemistry at 0–10 cm, reporting %P, %C, %N, and C:N for samples collected via ingrowth cores and sequential-coring standing stocks. CSV mirrors for each tab are provided, and a KML file supplies coordinates for all 32 plots.Key variables span live and dead fine-root biomass (and coarse fractions where applicable), specific root length (SRL) and area (SRA), diameter, root tissue density (RTD), soil moisture, AMF colonization, root %N, %C, %P, and C:N, along with minirhizotron root length and diameter. Depth, season, treatment, and plot/site identifiers are included to support cross-tab integration and analysis from 0–100 cm (minirhizotron) and 0–20 cm (cores).Units are reported in-column and missing values are coded as NA. No special software is required to open or use the files (Excel, CSV, and KML compatible).

54 ENVIRONMENTAL SCIENCES

Data for Stetten et al. (2025), "Biogeochemical controls on iron speciation and cycling across upland to shoreline gradients in freshwater and estuarine coastal soils (Lake Erie and Chesapeake Bay, United States)"

Coastal environments are dynamic interfaces that mediate carbon and nutrient exchanges between terrestrial landscapes and open waters, but it is unclear how biogeochemical reactions, in particular iron (Fe) redox transformations, affect the understanding and prediction of coastal ecosystem functions. This dataset includes measurements from two freshwater sites in the Western and Central basins of Lake Erie (Ohio, United States) and two estuarine sites in the Chesapeake Bay (Maryland, United States); the analytical results were reported by Stetten et al. (2025) in Science of the Total Environment. It was produced as part of the COMPASS-FME project, which seeks to advance a scalable, predictive understanding of the fundamental biogeochemical processes, ecological structure, and ecosystem dynamics that distinguish coastal terrestrial-aquatic interfaces from the purely terrestrial or aquatic systems to which they are coupled. The sites were sampled in November 2022 (CRC), December 2022 (MSM), February 2023 (GCW), and March 2023 (OWC); site codes follow those used by Pennington et al. (2025).The dataset consists of the following soil data:- Solid data (Fe concentration, etc.)- Porewater data (sulfate, sulfide, etc.)- Linear combination fitting results of X-ray absorption near edge structure (XANES) spectra; i.e., quantitative results of the oxidation state of Fe, indicated as a proportion of pure Fe(III) and Fe(II) model compounds- Linear combination fitting results of EXAFS (extended X-ray absorption fine structure) spectra, indicated as proportion of of Fe-model compounds (illite, smectite, etc.)Each data type has a single file in comma-separated value (CSV) format. No special software is required to read it.

54 ENVIRONMENTAL SCIENCES

Dataset for Cruz-O'Byrne et al (2026): "Divergent biogeochemical responses in upland coastal forest soils to repeated flooding and shifts in water chemistry"

Hydrologic disturbances from accelerated sea-level rise and the increasing frequency and intensity of storms and tidal flooding are altering biogeochemical processes in upland coastal forests, transforming these ecosystems into wetlands. However, the initial effects of flooding on belowground biogeochemistry and the mechanisms driving greenhouse gas dynamics and soil organic matter stability during the early stages of this transition remain poorly understood. This dataset presents the results of a mesocosm experiment conducted in a controlled, highly instrumented laboratory environment, in which freshwater and brackish water pulses were applied to intact soil monoliths from a temperate upland coastal forest to examine how floodwater chemistry influences soil biogeochemistry and organo-mineral interactions. All data files are plain-text CSV (comma-separated value), and no special software is required to read them. Details about the content of each file are available in the document “Dataset_readme”. The dataset consists of the following data: • rcruzobyrne_moisture: Soil volumetric water content (VWC) • rcruzobyrne_GHG: Headspace greenhouse gas (GHG) concentration and fluxes • rcruzobyrne_methane_isotopes: Headspace methane isotope signature • rcruzobyrne_porewater: Porewater chemistry • rcruzobyrne_CDOM: Porewater colored dissolved organic matter (CDOM) • rcruzobyrne_FTIR: Soil Fourier-transform infrared (FTIR) spectroscopy Details of the experimental setup, data collection, and data analysis are provided in the manuscript by Cruz-O’Byrne et al (2026) Divergent biogeochemical responses in upland coastal forest soils to repeated flooding and shifts in water chemistry. Biogeochemistry. https://doi.org/10.1007/s10533-026-01340-0

EARTH SCIENCE > ATMOSPHERE > GREENHOUSE GAS

Data & Code from Phoenix CPPP Phase 2 Analysis

This data and code package supports the analysis presented in “Beyond Surface Cooling: Comprehensive Field Assessment of Reflective Pavement Thermal Performance in Phoenix, Arizona” and provides fully reproducible workflows for evaluating the thermal performance of cool pavement treatments in a hot urban environment. The dataset integrates multi-modal field measurements collected across residential and nonresidential settings, including mobile air temperature traverses, stationary air temperature monitoring, residential mean radiant temperature (MRT) measurements, subsurface temperature profiles, and controlled testbed observations. The data package contains raw and processed datasets in comma-separated value (CSV) format, accompanying metadata files describing site characteristics and measurement protocols, and R scripts (.R files) used for data cleaning, time synchronization, spatial and temporal matching, quality control filtering, statistical comparison, and figure generation. All analyses were conducted using R (version ≥ 4.2.0) with commonly available packages (e.g., tidyverse, lubridate, data.table, ggplot2). No proprietary software is required to reproduce results. Field campaigns were designed to quantify the effects of high-reflectance pavement coatings on surface temperature, near-surface air temperature, subsurface heat propagation, and radiative heat exposure. Temporal alignment procedures include standardized timestamp conversion and nearest-neighbor matching of high-frequency sensor measurements to stop-based metadata within defined tolerance windows to ensure comparability across instruments. The workflows generate summary statistics, treatment–control contrasts, depth-dependent thermal gradients, and time-series visualizations used in the associated publication. By integrating mobile, stationary, radiative, and subsurface measurements within a unified and transparent processing framework, this package enables comprehensive evaluation of cool pavement performance across multiple thermal exposure pathways and supports reuse in future urban heat mitigation and climate resilience studies.

AIR TEMPERATURE

TEMPEST3 surface runoff water chemistry and organic matter composition

Coastal flooding, driven by storm surges and sea level rise, can mobilize organic matter (OM) via runoff, while introducing compositionally distinct OM (e.g., estuarine OM) into the system. To understand event-scale OM dynamics, we monitored source waters and surface runoff during an ecosystem-scale field manipulation experiment, TEMPEST (Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments), in June 2024. The TEMPEST experiment is part of the COMPASS-FME (Coastal Observations, Mechanisms, and Predictions Across Systems and Scales – Field, Measurements, and Experiments) project and designed to investigate biogeochemical and ecological impacts of freshwater and seawater flooding on coastal terrestrial-aquatic interface ecosystems by simulating freshwater and seawater storm events in two 2000m2 coastal upland forest plots (freshwater and brackish seawater plots). The temporal coverage of this dataset is during the TEMPESTⅢ event (June 11-13, 2024). This dataset contains: - Surface runoff discharge measured by flumes - Sensor data (specific conductivity, salinity, dissolved oxygen, and temperature) - Particle size distribution - Total suspended sediment concentrations (TSS), particulate and dissolved organic carbon (POC, DOC) concentrations, total nitrogen and total dissolved nitrogen (TN, TDN) concentrations - Bulk particulate and dissolved OM compositions (stable C and N isotopes of particulates and optical measurements of chromophoric dissolved OM) - High resolution mass spectrometry analysis data - Water isotope data All data files are plain-text CSV (comma-separated value), and no special software is required to read them.

COMPASS-FME

Requirements Description of the PERSENT Software

This report presents the modeling and simulation capabilities of Argonne National Laboratory’s PERSENT (PERturbation and SENsitivity for Transport) code [1] that is used in modern commercial deployment reactor technologies. The identified capabilities will be used to establish the set of PERSENT verification tasks necessary to verify PERSENT for usage on commercial projects. A similar path was followed for the REBUS [2] and DIF3D [3] software packages.

97 MATHEMATICS AND COMPUTING

Oak Ridge National Laboratory Evaluation of Stream-Trained Models in Practice

The goal of this integration is to replicate the results from the original paper Autonomous Utility Pole Identification on different camera hardware and integrate the model into a live video stream provided by the unmanned aerial system (UAS) itself while in operation. This involves retraining the original model and validating its efficacy on multiple camera modules to select the most effective device for installation. Moreover, this integration requires writing software to handle the reception of a real-time streaming protocol stream from the UAS and run each frame through the model while allowing a user to monitor the camera feed.

97 MATHEMATICS AND COMPUTING

Employing MACS/ViBRANT as a Surrogate MARVEL Reactor for Startup Reactivity Tuning and Supervisory Control Processes

Advanced nuclear reactors are a key part of the future of nuclear energy both in the United States and globally. They offer unique benefits for various energy-demanding applications, including use in remote locations, compact size, modular manufacturing, remote monitoring, low and/or variable power rating operation, and reliance on novel technologies to enhance operational safety. To achieve economic feasibility, advanced reactors must significantly reduce their workforces in comparison with the current fleet. Achieving this reduction will occur through reducing staff workloads using technology to achieve autonomous or semi-autonomous operations, demonstrated by comprehensive testing and validation activities. These operations will require both software and hardware platforms during the design and testing phases. While simulations are useful during the design phase, their performance can significantly deviate during actual deployment on hardware. This report presents the outcomes of a collaborative technical initiative between the U.S. Department of Energy (DOE) Microreactor Program (MRP) and Advanced Sensors and Instrumentation (ASI) Program. The collaboration utilized the Microreactor Automated Control System (MACS) hardware platform to bridge the gap between theoretical reactor design and actual startup and control operations. Two key use cases were investigated: facilitating the startup testing period and demonstrating supervisory control. The first use case details the key Microreactor Applications Research Validation and Evaluation (MARVEL) reactor startup physics testing activities conducted using the MACS platform. These activities included drum worth measurements, shutdown margin assessment, temperature feedback analysis, and scram time evaluation, as well as unique testing that would apply to the MARVEL reactor to demonstrate the testing methodologies in a low-risk environment. The MACS platform, serving as a surrogate representation of the MARVEL reactor, proved instrumental in performing these tests. The exercise revealed aspects that led to optimized processes, refined hardware design, and enhanced base software capabilities. By maturing methods and technologies in this manner, the initiative promises to reduce wasted time in the actual on-site reactor deployment effort, thereby saving significant time and resources. The second use case focuses on the development and implementation of supervisory control methods aimed at managing core tilt, which can result from asymmetrical operations or manufacturing imperfections in fuel rods or reactivity control devices. A key objective was to assess and compare the use of artificial intelligence (AI) for supervisory control. The effort aimed to define the role of supervisory control to enhance performance without risking control instability. This effort explored three distinct approaches: rules-based (RB) methods, optimization techniques, and reinforcement learning (RL) algorithms. Each approach was evaluated for its ease of implementation, its usability, and its effectiveness in responding to asymmetries in neutron flux. Comparative analysis of these approaches provided valuable insights into their applicability and effectiveness, offering a robust framework for advanced reactor operations. Together, these two use cases highlight the potential of hardware test beds to help streamline the design, operation, and control of advanced nuclear reactors. This collaborative effort underscores the importance of continued innovation and experimentation in achieving the next generation of safe, reliable, and economically viable nuclear energy solutions.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Field and Model Data Associated with the Manuscript “Drivers of Streamflow Intermittency in Humid Regions: 1. Evaluating Above- and Below-ground Controls of Flow Persistence in a Forested Catchment”

This package contains field data, modeling files, and scripts supporting the investigation of the drivers of streamflow intermittency in a forested catchment. It includes the field data collected from electrical resistivity tomography (ERT) surveys, ground penetrating radar (GPR), continuous self-potential (SP) monitoring, electromagnetic (EM) imaging, groundwater and stilling well. In addition, it contains the data and results of the coupled water- and electrical-flow model developed using the COMSOL Multiphysics and Advanced Terrestrial Simulator (ATS), as well as software files and Jupyter notebooks used to process the data and generate figures in the manuscript submitted for peer review. The data archive is organized in the following directories: 1) Climate Includes hourly precipitation and daily evapotranspiration time series (2024 – 2025) provided as CSV files, alongside a text file detailing dataset units. 2) Coupled_model Contains two subfolders: Synthetic and Field_Application subfolder. Synthetic subfolder contains the ATS XML input script (can be opened using any code editor) for the four synthetic hydrological cases tested (Connected and gaining, Connected and losing, Disconnected and losing, and dry stream). It also includes other experimental cases to test the influence of precipitation and concentration gradient. For each synthetic case, the flow model simulation is executed using the ATS XML scripts and the included Python script (generate_data_set.py) to convert ATS output to COMSOL-ready input. COMSOL Multiphysics template (.mph can be opened with the commercial software COMSOL and requires a license) is executed using the ATS output data to simulate the potential field. It also includes the Synthetic_model_plot.ipynb (can be opened using any code editor) to visualize the SP result and generate manuscript figures. The data subfolder contains mesh files to run both the ATS (.exo and .stl files can be viewed using Paraview; .h5 files can be opened using HDFView software and h5py Python package) and COMSOL models. Field_Application subfolder contains two subfolders: ES_MDA_inversion and Final_Model. ES_MDA_inversion contains the Python script (.py can be opened using any code editor) and SP observation data used to run the Ensemble Smoother with Multiple Data Assimilation (ES-MDA) inversion sequence to get the optimal model parameters. The Final_model subfolder contains the ATS XML input scripts, data files, output data for the two SP sites. The same workflow steps outlined for the Synthetic subfolder apply here. It also contains the Jupyter notebook (Plot_final_calib.ipynb) to visualize the results of the modeled SP, stream-groundwater exchange and moisture content. 3) Discharge Includes the electrical conductivity (EC) time series (provided as CSV files) from salt slug injections. It also includes the Jupyter notebook (Discharge_process.ipynyb) used to estimate discharge. All discharge measurements collated into rating_curve_processed.csv 4) EM Contains the CSV file of the EM data from the DUALEM-42, including spatial coordinates (x, y, z), apparent conductivity, and in-phase measurements at 2 m coil separations for horizontal coplanar (HCP) and perpendicular (PRP) geometries. 5) ERT Contains raw resistivity data (provided as CSV files), spatial location of each of the electrodes (provided as CSV files), and files used for the resistivity inversion (.resipy can be opened with the open-source ResIPy software). 6) GPR Includes GPR field datasets collected at 100 MHz and 250 MHz antenna frequencies, along with the processing/interpretation project file (GPR_process.gpz can be viewed using EKKO_Project 6, a commercial software by Sensors & Software that requires a license). 7) Slug_test Includes the slug test data at all the groundwater wells provided as CSV files, as well as the Jupyter notebook (Slug_test.ipynb) for calculating hydraulic conductivity. 8) SP Contains the SP data collected in field at the two SP sites (one in the perennial reach and the other in the intermittent reach), provided as DAT files. 9) Well_data Contains two subfolders: 1) Raw, which provides unprocessed pressure, electrical conductivity and temperature timeseries downloaded from the loggers in all the groundwater and stilling wells, and 2) Processed, which contains sorted, QA/QC timeseries data for each well. The data archive also contains data_process.ipynb, a Jupyter notebook used for field data analysis and generating figures (plotting well, SP, climate, and discharge data, as well as calculating head gradient at sites with nested groundwater wells). It also includes DTW.ipynb, a Jupyter notebook containing the code for the dynamic time warping (DTW) with sliding window to evaluate SP signal synchronicity.

ATS

National Energy Water Treatment & Speciation (NEWTS): A Water & Critical Mineral Database and Dashboard

The scarcity of water resources, the need for beneficial water reuse, and the challenges of wastewater treatment are becoming increasingly pressing in economic, social, and environmental domains. Addressing these concerns requires effective treatment strategies to manage wastewater streams and tackle environmental and economic issues. Furthermore, the recovery of critical minerals from the waste streams associated with energy production holds the promise of offsetting treatment costs and securing local sources of valuable minerals. However, relevant data on these waste streams are dispersed and challenging to locate. The process of ingesting such data into modeling software often involves multiple steps, requiring data restructuring to meet software-input requirements. The non-standardized reporting of water data makes data aggregation and reformatting a time-consuming process. Additionally, essential attributes necessary for modeling water treatment and mineral scale formation are frequently missing. Moreover, data gaps vary depending on the region of interest. Consequently, there is a pressing need for high-quality energy-water composition data that can be easily imported into water chemistry modeling software. To address this need, the National Energy Technology Laboratory has created the National Energy Water Treatment and Speciation (NEWTS) Database and Dashboard—a free online tool catering to community leaders and water researchers. NEWTS facilitates a comprehensive understanding of the composition of energy-related wastewater streams in the United States. The datasets provide detailed concentrations and speciation of major and minor aqueous compounds in energy-related wastewater streams, including power plant leachate, acid mine drainage, brackish water, and oil and gas produced water across the United States. Many of the aqueous species are critical minerals (Li, REEs) in high demand to modernize the world’s energy infrastructure. Many of the datasets also contain volumetric flow-rates needed to model the treatment and reuse scenarios in advanced aqueous chemistry software programs. The NEWTS Database and Dashboard offer public access to hitherto challenging-to-access datasets, presented in a standardized format that is tailored for easy input into aqueous chemistry modeling software. By performing the work needed to transform dispersed, disparate data sources into unified, model-ready datasets, NEWTS serves as an essential resource in advancing water treatment research and sustainable water resource management.

produced water management

Design and simulation of a SiPM-on-tile ZDC for the future EIC, and its performance with graph neural networks

We present a design for a high-granularity zero-degree calorimeter (ZDC) for the upcoming Electron-Ion Collider (EIC). The design uses SiPM-on-tile technology and features a novel staggered-layer arrangement that improves spatial resolution. To fully leverage the design’s high granularity and non-trivial geometry, we employ graph neural networks (GNNs) for energy and angle regression as well as signal classification. The GNN-boosted performance metrics meet, and in some cases, significantly surpass the requirements set in the report on science requirements and detector requirements for the EIC (Yellow Report), laying the groundwork for enhanced measurements that will facilitate a wide physics program. Our studies show that GNNs can significantly enhance the performance of high-granularity CALICE-style calorimeters by automating and optimizing the software compensation algorithms required for these systems. This improvement holds true even in the case of complicated geometries that pose challenges for image-based AI/ML methods.

Calorimeter

Multi‐Material Gradient Printing Using Meniscus‐enabled Projection Stereolithography (MAPS)

Light‐based additive manufacturing methods are widely used to print high‐resolution 3D structures for applications in tissue engineering, soft robotics, photonics, and microfluidics, among others. Despite this progress, multi‐material printing with these methods remains challenging due to constraints associated with hardware modifications, control systems, cross‐contamination, waste, and resin properties. Here, a new printing platform coined Meniscus‐enabled Projection Stereolithography (MAPS) is reported, a vat‐free method that relies on generating and maintaining a resin meniscus between a crosslinked structure and bottom window to print lateral, vertical, discrete, or gradient multi‐material 3D structures with no waste and user‐defined mixing between layers. MAPS is compatible with a wide range of resins shown and can print complex multi‐material 3D structures without requiring specialized hardware, software, or complex washing protocols. MAPS's ability to print structures with microscale variations in mechanical stiffness, opacity, surface energy, cell densities, and magnetic properties provides a generic method to make advanced materials for a broad range of applications.

bioprinting

Baseflow Identification via Explainable AI With Kolmogorov‐Arnold Networks

Abstract Hydrological models often involve constitutive laws that may not be optimal in every application. We propose to replace such laws with the Kolmogorov‐Arnold networks (KANs), a class of neural networks designed to identify symbolic expressions. We demonstrate KAN's potential on the problem of baseflow identification, a notoriously challenging task plagued by significant uncertainty. KAN‐derived functional dependencies of the baseflow components on the aridity index outperform their original counterparts; they demonstrate that water availability, rather than potential evapotranspiration, drives baseflow by constraining actual evapotranspiration under arid conditions. On a test set, they increase the Nash‐Sutcliffe efficiency (NSE) by 65%, decrease the root mean squared error by 29%, and increase the Kling‐Gupta efficiency by 34%. This superior performance is achieved while reducing the number of fitting parameters from three to two. Next, we use data from 378 catchments across the continental United States to refine the water‐balance equation at the mean‐annual scale. The KAN‐derived equations based on the refined water balance outperform both the current aridity index model, with up to a 105% increase in NSE, and the KAN‐derived equations based on the original water balance. While the performance of our model and tree‐based machine learning methods is similar, KANs offer the advantage of simplicity and transparency and require no specific software or computational tools. This case study focuses on the aridity index formulation, but the approach is flexible and transferable to other hydrological processes. Plain Language Summary Equations used in hydrologic model are often suboptimal, resulting in reduced prediction accuracy and efficiency. We implemented Kolmogorov‐Arnold networks (KAN), a machine learning algorithm for deriving symbolic formulations, to estimate groundwater recharge and showed that it outperforms an existing state‐of‐the‐art semi‐empirical formulation. In hydrology, Nash‐Sutcliffe efficiency (NSE), root mean squared error (RMSE), and Kling‐Gupta efficiency (KGE) are commonly used to evaluate model performance. Higher NSE and KGE values indicate better performance, while lower RMSE values are preferable. Our results show that NSE increased by 71%, RMSE decreased by 32%, and KGE improved by 25%. In addition, KAN identifies an optimal functional form and can be used to derive new analytical formulas using the prior knowledge. The KAN‐inspired equation outperformed the original formulation and reduced the fitting parameters. Furthermore, we refined the water‐balance equation at the mean‐annual scale and showed that, based on the new water‐balance equation, KAN can derive new formulations that are superior to the original aridity index formulations (up to 105% increase in NSE) and KAN‐derived equations based on the original water balance. These findings highlight the significant potential of KAN to advance the scientific understanding of a wide range of hydrologic processes. Key Points Kolmogorov‐Arnold networks (KANs) enhance interpretability of machine‐learned hydrological models KAN‐derived symbolic formulations outperform state‐of‐the‐art semi‐empirical aridity indices KAN‐identified functional form yields an analytical index with fewer fitting parameters and improved performance

baseflow

VISION: a modular AI assistant for natural human-instrument interaction at scientific user facilities

Scientific user facilities, such as synchrotron beamlines, are equipped with a wide array of hardware and software tools that require a codebase for human-computer-interaction. This often necessitates developers to be involved to establish connection between users/researchers and the complex instrumentation. The advent of generative AI presents an opportunity to bridge this knowledge gap, enabling seamless communication and efficient experimental workflows. Here we present a modular architecture for the Virtual Scientific Companion by assembling multiple AI-enabled cognitive blocks that each scaffolds large language models (LLMs) for a specialized task. With VISION, we performed LLM-based operation on the beamline workstation with low latency and demonstrated the first voice-controlled experiment at an x-ray scattering beamline. The modular and scalable architecture allows for easy adaptation to new instruments and capabilities. Development on natural language-based scientific experimentation is a building block for an impending future where a science exocortex—a synthetic extension to the cognition of scientists—may radically transform scientific practice and discovery.

36 MATERIALS SCIENCE

Detector Interface for Streaming, Control, and Open-source integration (DISCO) v1.0.0

This suite consists of a multi-package ecosystem featuring detector emulators, EPICS areaDetector drivers, and remote server frameworks designed for the Advanced Light Source (ALS). Engineered for high-bandwidth devices—including VFCCD, Timepix3, Timepix4, and related pixel detectors—the software simulates hardware, wraps vendor SDKs into remote-callable servers, and integrates with open-source control systems. Key Capabilities: Distributed SDK Architecture: Server packages wrap hardware-specific SDKs, allowing areaDetector drivers to execute remote framework calls. This isolates proprietary libraries from the EPICS IOC, enhancing stability and enabling distributed computing across beamline networks. Device Support: Custom drivers for VFCCD, the Timepix family, and similar sensors optimize the data path from hardware control to high-speed transport. Full-Stack Emulation: Sophisticated emulator packages allow end-to-end pipeline testing and software development without requiring physical hardware or beam time. Integrated Workflows: Supports high-bandwidth streaming for real-time analysis and robust, metadata-rich file-based workflows (e.g., HDF5/NeXus). By standardizing interfaces across heterogeneous hardware, this suite reduces technical debt. It provides the ALS with a scalable, open-source solution to manage massive data rates within a unified control environment.

Mahl, Johannes [Lawrence Berkeley National Laborat

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES