Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Toward Unified Autonomous Scattering Experiments: A Cross-Facility Case Study at ALS and PETRA III

Autonomous experiments rely on the integration of control, data acquisition, analysis, and decision-making frameworks. While such systems have been demonstrated at individual facilities, adapting them to additional instruments remains challenging due to differences in local infrastructure. We present a modular workflow that connects existing open-source tools for data access (Tiled), workflow orchestration (Prefect), analysis and visualization (pyFAI, Plotly Dash), and Gaussian-process-based adaptive sampling (gpCAM) into a unified framework for autonomous scattering experiments. The same configuration operates across two synchrotron beamlines (ALS 7.3.3 and PETRA III P03) with only minimal facility-specific adjustments, as shown in proof-of-concept demonstrations. This validates that a consistent design emphasizing modularity and shared interfaces can ease deployment across diverse experimental environments. The resulting framework provides a flexible foundation for extending autonomous control and analysis capabilities beyond a single beamline or instrument.

47 OTHER INSTRUMENTATION↗

Benchmarking image processing techniques for porosity measurement in polymer additive manufacturing: Review and experimental analysis

An image processing workflow is proposed for porosity measurement in polymer additive manufacturing. Various techniques, including global and local thresholding, region growing, and K-means clustering, were applied to microscopic images of carbon fiber reinforced acrylonitrile butadiene styrene (CF-ABS) and benchmarked for their ability to accurately measure porosity. Global methods included Otsu, minimum error, iterative, and entropy-based thresholding, while local methods included Niblack, Bernsen, Sauvola, and Bradley-Roth algorithms. Artificial uneven illumination was introduced to test local adaptive thresholds. Results showed significant differences in porosity values across methods. Otsu, region growing, and K-means clustering excelled under uniform illumination, while Sauvola and Bradley-Roth performed better with uneven illumination. Comparison with X-ray computed tomography (XCT) revealed slightly lower porosity values (2.55 %) than optimized methods (2.73–2.79 %) due to XCT's lower resolution excluding smaller pores. While XCT offers finer pore detection, it limits sample volume and underestimates porosity due to spatial variation. Validation using artificial grayscale images with 5 % porosity confirmed that Otsu, Bradley-Roth, region growing, and Sauvola algorithms produced accurate results. Although tested on a single material system, these methods can be adapted to others with optimization. In conclusion, given XCT's high computational and time costs, this study highlights suitable image processing techniques as cost-effective alternatives for porosity analysis in polymer composites.

Additive manufacturing↗

Discovering invariant spatial features in electron energy loss spectroscopy images on the mesoscopic and atomic levels

Over the last two decades, Electron Energy Loss Spectroscopy (EELS) imaging with a scanning transmission electron microscope has emerged as a technique of choice for visualizing complex chemical, electronic, plasmonic, and phononic phenomena in complex materials and structures. The availability of the EELS data necessitates the development of methods to analyze multidimensional data sets with complex spatial and energy structures. Traditionally, the analysis of these data sets has been based on analysis of individual spectra, one at a time, whereas the spatial structure and correlations between individual spatial pixels containing the relevant information of the physics of underpinning processes have generally been ignored and analyzed only via the visualization as 2D maps. Here, we develop a machine learning-based approach and workflows for the analysis of spatial structures in 3D EELS data sets using a combination of dimensionality reduction and multichannel rotationally invariant variational autoencoders. This approach is illustrated for the analysis of both the plasmonic phenomena in a system of nanowires and in the core excitations in functional oxides using low loss and core-loss EELS, respectively. The code developed in this manuscript is open sourced and freely available and provided as a Jupyter notebook for the interested reader.

36 MATERIALS SCIENCE↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗

Correlating and Simulating Socio-Demographically Driven Residential End-Use Activity Schedules

Incorporating socio-demographic and behavioral considerations into decision-support tools is crucial for identifying gaps and addressing consumer needs to ensure reliable and affordable energy solutions. In energy simulation models, the correlation between socio-demographics and time-use behavior is not well-captured. Thus, we developed a large-scale simulation workflow to generate schedules for 10 residential activities across 24 population segments defined by age, income, and employment status. Using pre-pandemic 2015-2019 American Time Use Survey (ATUS) data, we used ANOVA to confirm the correlation between demographic factors and time use. We explored three k-modes clustering methods-backward, forward, and a new hybrid approach-to delineate the occupancy patterns based on demographics. Using the probability of cluster membership for each population segment and a time inhomogeneous Markov chain to generate activity transition probabilities for each cluster, we simulated 50,000 schedules per segment and validated them against the ATUS data. The hybrid method produced the most socio-demographically differentiated clusters while demonstrating comparable performance to other approaches, with an overall root mean square error of 0.12 for both weekday and weekend schedules. Thus, the hybrid method, where each cluster is dominated by certain demographic segments and occupancy patterns, offers more modeling versatility in terms of scenario analysis. The new workflow improves the socio demographic differentiation of energy consumption by considering differences in time use. This approach enables future research on demographically segmented time of use (TOU) energy consumption, including impacts of TOU utility bills and rate analysis, long-run marginal emissions, and energy retrofits.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Boosting RDataFrame performance with transparent bulk event processing

RDataFrame is ROOT’s high-level interface for Python and C++ data analysis. Since it first became available, RDataFrame adoption has grown steadily and it is now poised to be a major component of analysis software pipelines for LHC Run 3 and beyond. Thanks to its design inspired by declarative programming principles, RDataFrame enables the development of highperformance, highly parallel analyses without requiring expert knowledge of multi-threading and I/O: user logic is expressed in terms of self-contained, small computation kernels tied together by a high-level API. This design completely decouples analysis logic from its actual execution, and opens several interesting avenues for workflow optimization. In particular, in this work we explore the benefits of moving internal data processing from an event-by-event to a bulkby-bulk loop. This refactoring dramatically reduces the framework’s runtime overheads; in collaboration with the I/O layer it improves data access patterns; it exposes information that optimizing compilers might use to auto-vectorize the invocation of user-defined computations; finally, while existing user-facing interfaces remain unaffected, it becomes possible to additionally offer interfaces that explicitly expose bulks of events, useful e.g. for the injection of GPU kernels into the analysis workflow. In order to inform similar future R&D, design challenges will be presented, as well as an investigation of the relevant timememory trade-off backed by novel performance benchmarks.

Guiraud, Enrico↗

The ATLAS Workflow Management System Evolution in the LHC Run3 and towards the High-Luminosity LHC era

The ATLAS experiment has 18+ years of experience using workload management systems to deploy and develop workflows to process and to simulate data on the distributed computing infrastructure. Simulation, processing and analysis of LHC experiment data require the coordinated work of heterogeneous computing resources. In particular, the ATLAS experiment utilizes the resources of 250 computing centers worldwide, the power of supercomputing centres, and national, academic and commercial cloud computing resources. In this contribution, we present new techniques for cost-effectively improving efficiency introduced in workflow management system software. The evolution from a mesh framework to new types of computing facilities such as cloud and HPCs is described, as well as new types of production and analysis workflows.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A New Workflow of X-ray CT Image Processing and Data Analysis of Structural Features in Rock Using Open-Source Software

X-ray computed tomography (CT) images of rock specimens often contain artifacts which must be corrected before scientific analyses are performed. Here, we present a new workflow of automated image processing to utilize poor-quality X-ray CT scan images. The workflow runs on the open-source image analysis software and efficiently separates desired features from low-contrast scanned images. The new workflow is a two-step technique using contrast enhancement and automated feature segmentation to generate noise-free binary images. The results of binary images using the proposed workflow and using a conventional thresholding technique are analyzed to show the quality of the proposed method. The paper also presents a workflow of estimating the structural geometries of features in two and three dimensions. The results of the structural feature analyses and computational time were compared between the open-source (ImageJ) and commercial image analysis software (Bruker Computed Tomography Analyzer). The commercial software was more computationally efficient, but the task-specific macros in open-source software enabled the user-desired automation in image processing and data extraction of desired structural features of comparable quality.

47 OTHER INSTRUMENTATION↗

Analysis of Bis(trifluoromethylsulfonyl)imide Interactions with Metal Cations Through a Chemical Informatics Approach

Nominally weakly coordinating anions are useful for modulating the solubility and chemical properties of metal complexes, but identification and analysis of the systematics of the interactions of anions with cationic metal complexes has not received the attention it deserves. Here, a chemical informatics approach is demonstrated for identifying and quantitatively analyzing the ways that the bis(trifluoromethylsulfonyl)imide anion (TFSI) can interact with metal-containing species. An open access computer program (PyCIFTer) was developed to facilitate large-scale structural analysis of TFSI-containing species by utilization of experimental atomic coordinate data from single-crystal X-ray diffraction (XRD) studies obtained from the Cambridge Structural Database (CSD). PyCIFTer establishes a three-dimensional vector space from the raw atomic coordinates, generating acyclic, undirected graphs that are used to rapidly analyze the structural properties (bond lengths and angles) of TFSI in individual structures in sequential/batch fashion. The structures are sorted by PyCIFTer into groups based on pre-set and chemically sensible criteria, affording a comprehensive and systematic view of TFSI structural chemistry. This approach avoids tedious one-at-a-time interrogation of structures, a prospect unreasonable in this case, and many others of contemporary chemical relevance; there were over 1500 structures in the CSD containing TFSI as of November 2024. The results demonstrate that TFSI only rarely binds to cations in the solid state, favoring the formation of species in which TFSI is found in cations’ outer coordination spheres. The prospect of applying PyCIFTer to other moieties is also discussed. PyCIFTer is also schematically compared to the commercial CSD Python application programming interface (API). Taken together, this work demonstrates the usefulness of modular workflows for sequential/batch analysis of structural data from XRD, an approach that appears poised to accelerate the translation of legacy structural results into new chemical insights and hypotheses.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Biosensor-driven strain engineering reveals key cellular processes for maximizing isoprenol production in Pseudomonas putida

Synthetic biology generates vast combinatorial designs, yet high-throughput analytical methods to screen them are poorly matched to interrogate this search space. We address this challenge by developing a biosensor-driven, growth-coupled selection strategy in Pseudomonas putida for isoprenol, a potential aviation fuel precursor. We found and characterized a noncanonical signaling pathway, revealing a functional and physical complex between a hybrid histidine kinase and an alcohol dehydrogenase, whose activity is tuned by heterodimerization. Leveraging this biosensor in a pooled CRISPRi library selection, we identified key host limitations. Iterative combinatorial strain engineering derived from these hits yielded a 36-fold titer increase to ~900 milligrams per liter. Integrated omics analysis revealed that metabolic rewiring toward amino acid catabolism was crucial for this improvement. This observation was found to be beneficial by technoeconomic analysis. Our modular workflow provides a powerful strategy for optimizing complex heterologous pathways and uncovering emergent host biology.

CRISPRi↗

A Review and Outlook on Experimental Advances and Innovations in Geological CO 2 Storage: Insights from Depleted Gas Reservoirs and Saline Aquifers

Geological storage of carbon dioxide (CO 2 ) in depleted gas reservoirs and deep saline aquifers is a key part of global decarbonization efforts. As carbon capture and storage advances toward commercial-scale deployment, the credibility and scalability of laboratory experiments are increasingly vital for guiding safe and effective field implementation. This review offers a comprehensive, cross-scale evaluation of experimental methodologies, including core flooding, high-pressure, high-temperature systems, microfluidic visualization, and emerging systems such as multilayer commingled/compartmentalized core flooding, 3D-printed micromodels, and AI-powered digital twins. These innovations are demonstrated to enhance representativeness, reproducibility, and real-time insight, thereby addressing the limitations of conventional workflows. A critical analysis of methodological gaps, such as inconsistent pressure–temperature conditions, oversimplified brine chemistry, and a lack of standardization, reveals experimental sources of scale translation errors and performance uncertainty. By comparing the unique challenges of depleted gas reservoirs (such as low water saturation and legacy well leakage) to those of saline aquifers (including pressure buildup and caprock integrity), this review identifies formation-specific priorities for experimental design. Novel contributions include a synthesis of best practices, integration strategies for model calibration, and recommendations for standardizing core handling, saturation procedures, and reporting protocols. Furthermore, this work serves as a guide for developing robust, field-relevant experimental strategies that can increase the deployment and regulatory acceptance of CO 2 storage technologies at scale.

58 GEOSCIENCES↗

A method for crystallographic mapping of an alpha-beta titanium alloy with nanometre resolution using scanning precession electron diffraction and open-source software libraries

An approach for the crystallographic mapping of two-phase alloys on the nanoscale using a combination of scanned precession electron diffraction and open-source python libraries is introduced in this paper. This method is demonstrated using the example of a two-phase α/β titanium alloy. The data were recorded using a direct electron detector to collect the patterns, and recently developed algorithms to perform automated indexing and analyse the crystallography from the results. Very high-quality mapping is achieved at a 3 nm step size. The results show the expected Burgers orientation relationships between the α laths and β matrix, as well as the expected misorientations between α laths. A minor issue was found that one area was affected by 180° ambiguities in indexing occur due to this area being aligned too close to a zone axis of the α with twofold projection symmetry (not present in 3D) in the zero-order Laue Zone, and this should be avoided in data acquisition in the future. Nevertheless, this study demonstrates a good workflow for the analysis of nanocrystalline two- or multi-phase materials, which will be of widespread use in analysing two-phase titanium and other systems and how they evolve as a function of thermomechanical treatments.

36 MATERIALS SCIENCE↗

Machine learning methods for weather forecasting

SAND2025-14466O This repository contains code for developing, training, and evaluating machine learning models for weather and climate forecasting, including forecast skill assessment, feature importance analysis, and reproducible workflows for model comparison. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Holthuijzen, Maike [Sandia National Lab. (SNL-CA),↗

S4PST: Sustainability for Programming Systems and Tools: May Workshop Report

The US Department of Energy (DOE) Exascale Computing Project (ECP) has fostered and strengthened the use of modern software engineering practices for developing applications and libraries, and this effort has resulted in the coordinated and interoperable E4S1 and xSDK2 ecosystems. Although this approach is cost-effective, it relies on robust programming systems and tools (PST) as the underlying foundation for our HPC software. At present, our primary PST stack consists of traditional high-performance computing (HPC) languages, namely Fortran, C, C++, and the popular Python language for data analysis and AI workflows. These languages support various programming frameworks and run-time abstractions that enable parallelism and concurrency across multiple node architectures and thousands of nodes through a variety of interconnect systems. However, to accommodate users’ diverse needs, certain aspects of the HPC ecosystem are delegated to vendor-specific or third-party implementations that extend beyond a particular scientific domain. This broader scope results in a multitude of specifications and variations, which leads to a complex orchestration of many-ecosystems. Unfortunately, this complexity in the ecosystem imposes additional overhead costs on consumers during the latter stages of the development cycle. In addition to the software ecosystem challenge, the upcoming conclusion of the ECP by December 2023 has raised significant concerns within the HPC programming systems community, from both the economic and social perspectives. The ECP has implemented a management structure for software development and funding decisions across all ECP participants by following a conventional hierarchical and centralized approach. However, this structure has prompted certain considerations within the community, particularly in anticipation of the Software Sustainability initiative by the DOE’s Advanced Scientific Computing Research Program (ASCR). For the success of this new initiative, it is of utmost importance to secure consistent funding and foster close engagement with researchers and core developers of existing programming-system products. This collaboration is vital to maintaining the critical capabilities of the current software during the transition phase while proactively adapting to future technology and workforce trends. The community recognizes the significance of adapting to emerging trends and is aware of the inherent fragility of the HPC software ecosystem, particularly in relation to programming systems that cater to all users. The ability to adapt and evolve is essential to staying relevant and effectively addressing these technical, economic, and social challenges. The S4PST team, which represents one of the six ASCR Software Sustainability seedling projects, is dedicated to tackling these challenges through community-based approaches that go beyond the scope of the DOE. This involves collaboration between national laboratories with academia, non-DOE institutions, hardware and system vendors, and international partners. By fostering these partnerships, we aim to create a robust and sustainable HPC software ecosystem that can effectively meet the needs of the community. This new community effort, driven by the eight DOE labs, will take on the responsibility of guiding funding decisions for programming-systems development and maintenance with transparency and consistency across all decisions. Additionally, the team will offer common technical services to the programming systems community, irrespective of their funding situations, and facilitate community-wide incubation to proactively nurture the software ecosystem. By actively engaging with stakeholders and employing a collaborative approach, we can collectively shape the future of programming systems and ensure a robust and thriving HPC software landscape. On May 11–12, 2023, the S4PST team conducted its inaugural kick-off workshop at the Innovative Computing Laboratory (ICL) in the University of Tennessee, Knoxville, hosted by Hartwig Anzt. The workshop encompassed various sessions dedicated to presentations and discussions, with the aim of comprehending the team members’ perspectives on the vision of software sustainability. Additionally, the workshop aimed to identify the technical, economic, and social requirements for sustaining the programming-systems community in the field of HPC. This report provides a summary of the S4PST effort by highlighting five major thrust areas discussed during the workshop: (i) community, (ii) technical support, (iii) training and diversity, (iv) verification, validation and correctness, and (v) emerging technologies. It also encompasses an overview of the presentations and discussions held throughout the event, our views and potential synergies with other seedling efforts, along with the outcomes and key takeaways from our initial discussions.

97 MATHEMATICS AND COMPUTING↗

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING↗

Data and Scripts associated with “Lambda-PFLOTRAN: Workflow for Incorporating Organic Matter Chemistry Informed by Ultra High Resolution Mass Spectrometry into Biogeochemical Modeling.”

This data package is associated with the publication “Lambda-PFLOTRAN: Workflow for Incorporating Organic Matter Chemistry Informed by Ultra High Resolution Mass Spectrometry into Biogeochemical Modeling” submitted to Geoscientific Model Development (Muller et al., 2024). In this manuscript, organic matter chemistry and thermodynamics are directly connected to reactive transport simulators through the newly developed Lambda-PFLOTRAN (Parallel Reactive Flow and Transport model) workflow tool that succinctly incorporates organic matter chemistry data generated from Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) into reaction networks to simulate aerobic respiration of the organic matter and the resulting biogeochemistry. Lambda-PFLOTRAN is a python-based workflow, executed through a Jupyter Notebook interface, that digests raw FTICR-MS data, develops a representative reaction network based on substrate-explicit thermodynamic modeling (also termed lambda modeling due to its key thermodynamic parameter λ used therein), and completes a biogeochemical simulation with the open source, reactive flow, and transport code PFLOTRAN. This data package contains Jupyter Notebook based workflows for two test cases for running biogeochemical simulations of organic matter oxidation identified by FTICR-MS. It contains four primary folders (workflow, data, src, and analysis), a file-level metadata file (Muller_2024_Lambda_PFLOTRAN_Manuscript_Data_Package_flmd.csv) that lists all the files contained in this data package with a short description of each, and a data dictionary (Muller_2024_Lambda_PFLOTRAN_Manuscript_Data_Package_dd.csv) file that describes the tabular column headers. The ‘workflow’ folder contains the Jupyter Notebook based workflows for running the lambda analysis, PFLOTRAN simulation, sensitivity analysis and parameter estimation. The ‘data’ folder contains the FTICR-MS data, initial conditions, and incubation data for test cases 1 and 2 in folders titled ‘WHONDRS’ and ‘Colloids’, respectively. The data folder also has a ‘Database’ folder containing a reaction network for bulk organic matter (assumed to be CH2O) and a general database for PFLOTRAN (hanford_rxn_network). The CH2O reaction network defines bulk organic matter oxidation. Biogeochemical simulations are completed for both the lambda binned organic matter and bulk organic matter reaction networks. The ‘hanford_rxn_network’ database includes information required for PFLTORAN simulations including ion size, molar mass, and charge of the aqueous species, gases, and minerals phases. The ‘src’ folder contains python source codes for performing lambda analysis, PFLOTRAN simulation, sensitivity analysis and parameter estimation. The ‘analysis’ folder contains outputs from the test cases 1 and 2 including lambda analysis, PFLOTRAN runs and the calibration results.

54 ENVIRONMENTAL SCIENCES↗

HEPTAPOD: Orchestrating High Energy Physics Workflows Towards Autonomous Agency

Many workflows in high-energy-physics (HEP) stand to benefit from recent advances in transformer-based large language models (LLMs). While early applications of LLMs focused on text generation and code completion, modern LLMs now support orchestrated agency: the coordinated execution of complex, multi-step tasks through tool use, structured context, and iterative reasoning. We introduce the HEP Toolkit for Agentic Planning, Orchestration, and Deployment (HEPTAPOD), an orchestration framework designed to bring this emerging paradigm to HEP pipelines. The framework enables LLMs to interface with domain-specific tools, construct and manage simulation workflows, and assist in common utility and data analysis tasks through schema-validated operations and run-card-driven configuration. To demonstrate these capabilities, we consider a representative Beyond the Standard Model (BSM) Monte Carlo validation pipeline that spans model generation, event simulation, and downstream analysis within a unified, reproducible workflow. HEPTAPOD provides a structured and auditable layer between human researchers, LLMs, and computational infrastructure, establishing a foundation for transparent, human-in-the-loop systems.

Menzo, Tony [Alabama U.; Fermilab] (ORCID:00000002↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗