Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parsing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Remote Instrumentation and Data Acquisition

This poster outlines the development and implementation of a remote data acquisition system for waveform analysis using a Rohde & Schwarz oscilloscope. The project involved capturing waveform data, and transferring it to a local machine for visualization and analysis. The core logic was developed in C++ with a focus on object oriented programming and the use of polymorphism so the main application can interact with any instrument without knowing its exact type, simplifying the overall logic and making it easier to add or swap out components without changing the rest of the codebase.. The system issues Standard Commands for Programmable Instruments (SCPI) via a socket connection and parses the oscilloscope s ASCII waveform data. The C++ application was containerized using Docker for ease of portability, and reproducibility. Emphasis was placed on secure networking practices, error handling, and effective data capture. The report describes the technical steps taken, challenges encountered, and future work, providing insight into the practical integration of hardware interfacing with remote computational environments.

Parikh, Jaymil [Illinois U., Urbana]↗

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING↗

GridCoPilot for Thermal Events: An LLM-Based Platform for Power Grid Reliability Analysis

Large Language Models show promise for translating natural language into database queries, but deploying such systems in safety-critical domains requires high reliability. We present an application of GridCoPilot to thermal event analysis (heatwaves and coldwaves) that affect power grid reliability. Our approach uses a LangChain SQL Agent to translate natural language queries into auditable SQL statements, with deterministic visualization routines that parse the structured query results. We introduce structural framing as a design principle, we integrate a NERC-region-level event library with county-level meteorology and decompose the combined data into three relational tables (event metadata, county-level event details, and a county-to-NERC subregion mapping), using prompt-guided joins to direct the model toward correct multi-table queries. For two core analytical patterns (identifying worst events by region and by region-year), the system achieved 100% SQL accuracy across all 16 NERC subregions and both event types (64 queries total). These results validate the approach for target use cases, though performance on diverse natural language formulations requires further investigation. We discuss design trade-offs, failure modes including JSON output truncation, and pathways for extending this approach to other hazard domains.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Expandable Log Analyzing Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, IL]↗

Expandable Log Analyzing Framework

Prior to my internship, I was informed that a previous intern had built a tool to analyse MongoDB logs and look for invalid access attempts, which served as a great reference point for my project. I was initially tasked with expanding on her prototype and filling in the gaps such as integrating it with the main monitoring tool the lab uses. Eventually, the scope grew, expanding to support other databases and a growing collection of tools. I organized the framework around an observer pattern, meaning one point in the program sending updates to the rest of the framework. Every time a log was read and parsed, it was sent to be processed by the tools, using the type of event as a means to determine which tools should get a chance to act on the log. This decouples the tools from the log reader, making future updates and additions much easier. The framework processes MongoDB logs at ~135,000 entries per second and PostgreSQL logs at ~170,500 entries per second, accurately detecting anomalies such as slow queries and connections from unknown addresses. This framework serves to fill gaps in database monitoring tools currently implemented at the lab, such as tracking failed authentication for PostgreSQL and MongoDB which had very minimal or none before this framework. National labs such as Fermilab hold sensitive data and valuable computing resources, making them attractive targets. Monitoring intrusion attempts on databases is made much easier by this comprehensive monitoring suite.

Clark, Dylan [Unlisted, IL]↗

Agile Acceleration of LLVM Flang Support for Fortran 2018 Parallel Programming

The LLVM Flang compiler ("Flang") is currently Fortran 95 compliant, and the frontend can parse Fortran 2018. However, Flang does not have a comprehensive 2018 test suite and does not fully implement the static semantics of the 2018 standard. We are investigating whether agile software development techniques, such as pair programming and test-driven development (TDD), can help Flang to rapidly progress to Fortran 2018 compliance. Because of the paramount importance of parallelism in high-performance computing, we are focusing on Fortran’s parallel features, commonly denoted “Coarray Fortran.” We are developing what we believe are the first exhaustive, open-source tests for the static semantics of Fortran 2018 parallel features, and contributing them to the LLVM project. A related effort involves writing runtime tests for parallel 2018 features and supporting those tests by developing a new parallel runtime library: the CoArray Fortran Framework of Efficient Interfaces to Network Environments (Caffeine).

Rasmussen, Katherine↗

GROWdb US River Systems - Samples

GROW Overview We developed the Genome Resolved Open Watersheds database (GROWdb), which aims to increase genomic sampling and understanding of global river microbiomes. An emphasis of GROWdb is to create a publicly available and ever-expanding microbial genome database that is focused on rivers while being interoperable with databases from other ecosystems. GROWdb is based on a network-of-networks approach to move beyond a small collection of well-studied rivers, towards a spatially distributed, global network of systematic observations. GROWdb represents the first microbial, river-focused resource parsed at various scales from genes to MAGs to community level including expression and potential based measurements that will be of interest to microbiologists, ecologists, geochemists, hydrologists, and modelers. Dataset Acknowledgement GROWdb contains data from various research campaigns, please acknowledge the following data generators, as appropriate: WHONDRS derived genomes or samples - include this statement in your acknowledgements: “This study used data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) under the River Corridor Science Focus Area (SFA) at the Pacific Northwest National Laboratory (PNNL) that was generated at the U.S. Department of Energy (DOE) Joint Genome Institute User Facility. PNNL is operated by Battelle Memorial Institute for the U.S. DOE under Contract No. DE-AC05-76RL01830. The SFA is supported by the U.S. DOE, Office of Biological and Environmental Research (BER), Environmental System Science (ESS) Program.” Total Samples loaded onto this Narrative: 178 Note: Not all GROW samples may be loaded into KBase Data Availability The data underlying GROWdb are accessible across various platforms to ensure all levels of data structure are widely available. First, all reads and MAGs are publicly hosted on National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. Second, all data related data presented here including MAG annotations, extended data tables, phylogenetic tree files, antibiotic resistance gene database files, and MAG abundance tables are available in Zenodo (link). Beyond the flat database files listed above, our aim for GROWdb was to maximize data use by making the data available in searchable and interactive platforms including the National Microbiome Data Collaborative (NMDC) data portal, the Department of Energy’s Systems Biology Knowledgebase (KBase), and a GROW specific user interface released here, GROWdb Explorer. Each platform provides different ways to interact with GROWdb: NMDC GROWdb formed a pilot project for the NMDC. Specifically, individual GROWdb datasets (metagenomes, metatranscriptomes, etc) are easily accessible and searchable through the NMDC data portal, where they are systematically connected to each other and to a rich suite of sample information and standard analysis results, following Findable, Accessible, Interoperable, and Reusable (FAIR) data practices. KBase GROWdb is publicly available within KBase, including samples (this Narrative), MAGs, and corresponding genome scale metabolic models. Access within KBase allows for immediate access and reuse of data, including comparison to private data using KBase’s 500+ analysis tools. Other linked narratives in KBase: GROW Metagenome Assembled Genomes (MAGs) GROW Metabolic Models GROWdb Explorer GROWdb data is also explorable through a graphical user interface built through the Colorado State University Geospatial Centroid (https://geocentroid.shinyapps.io/GROWdatabase/), allowing users to search and graph microbial and spatial data simultaneously. In summary, this microbial genome resource represents the first publicly available genome collection from rivers and offers data that can be leveraged across microbiome studies. GROWdb is an expanding repository to incorporate and unify global river multi-omic data for the future.

59 BASIC BIOLOGICAL SCIENCES↗

Simulation-Based Validation of An Open-Source, Scalable Framework for Building Energy Management in Small and Medium-Sized Commercial Buildings

Abstract: Small and medium-sized commercial buildings (SMCBs) represent 94% of U.S. commercial buildings but encounter substantial obstacles in adopting Building Energy Management (BEM) systems. Current approaches exhibit fundamental limitations: vendor-specific API platforms restrict interoperability through proprietary ecosystems; commercial automation software demands extensive technical expertise and licensing costs; open-source IoT solutions lack native support for building automation protocols and semantic models. This paper introduces a configuration-driven web interface framework addressing the gap between smart device advancements and accessible BEM software infrastructure for SMCBs. The framework leverages VOLTTRON middleware integrated with an automated converter that processes unified YAML configurations into heterogeneous system files, reducing required configuration artifacts from six separate files to a single unified specification. The system architecture enables vendor-agnostic operation through BACnet and Modbus protocols while supporting semantic building model integration via automated Brick Schema parsing. Configuration-driven interfaces automatically adapt to diverse HVAC types without custom development. Simulation-based validation using BOPTEST demonstrates automatic interface generation between fan coil and hydronic systems, with the automated converter successfully generating all platform-specific outputs from the single YAML input. The result demonstrates the framework's capability to streamline BEM system deployment through reduced configuration complexity. This work bridges simulation capabilities with operational deployment, demonstrating how virtual testbeds validate generalizable software frameworks for real-world building automation.

Chung, Jihoon [ORNL] (ORCID:0000000184880815)↗

Specialized Plant Growth Chamber Designs to Study Complex Rhizosphere Interactions

The rhizosphere is a dynamic ecosystem shaped by complex interactions between plant roots, soil, microbial communities and other micro- and macro-fauna. Although studied for decades, critical gaps exist in the study of plant roots, the rhizosphere microbiome and the soil system surrounding roots, partly due to the challenges associated with measuring and parsing these spatiotemporal interactions in complex heterogeneous systems such as soil. To overcome the challenges associated with in situ study of rhizosphere interactions, specialized plant growth chamber systems have been developed that mimic the natural growth environment. This review discusses the currently available lab-based systems ranging from widely known rhizotrons to other emerging devices designed to allow continuous monitoring and non-destructive sampling of the rhizosphere ecosystems in real-time throughout the developmental stages of a plant. We categorize them based on the major rhizosphere processes it addresses and identify their unique challenges as well as advantages. We find that while some design elements are shared among different systems (e.g., size exclusion membranes), most of the systems are bespoke and speaks to the intricacies and specialization involved in unraveling the details of rhizosphere processes. We also discuss what we describe as the next generation of growth chamber employing the latest technology as well as the current barriers they face. We conclude with a perspective on the current knowledge gaps in the rhizosphere which can be filled by innovative chamber designs.

59 BASIC BIOLOGICAL SCIENCES↗

Coordination and divergence in community assembly processes across co-occurring microbial groups separated by cell size

Setting the pace of life and constraining the role of members in food webs, body size can affect the structure and dynamics of communities across multiple scales of biological organization (e.g., from the individual to the ecosystem). However, its effects on shaping microbial communities, as well as underlying assembly processes, remain poorly known. Here, we analyzed microbial diversity in the largest urban lake in China and disentangled the ecological processes governing microbial eukaryotes and prokaryotes using 16S and 18S amplicon sequencing. We found that pico/nano-eukaryotes (0.22–20 μm) and micro-eukaryotes (20–200 μm) showed significant differences in terms of both community composition and assembly processes even though they were characterized by similar phylotype diversity. We also found scale dependencies whereby micro-eukaryotes were strongly governed by environmental selection at the local scale and dispersal limitation at the regional scale. Interestingly, it was the micro-eukaryotes, rather than the pico/nano-eukaryotes, that shared similar distribution and community assembly patterns with the prokaryotes. This indicated that assembly processes of eukaryotes may be coupled or decoupled from prokaryotes’ assembly processes based on eukaryote cell size. While the results support the important influence of cell size, there may be other factors leading to different levels of assembly process coupling across size classes. Additional studies are needed to quantitatively parse the influence of cell size versus other factors as drivers of coordinated and divergent community assembly processes across microbial groups. Regardless of the governing mechanisms, our results show that there are clear patterns in how assembly processes are coupled across sub-communities defined by cell size. These size-structured patterns could be used to help predict shifts in microbial food webs in response to future disturbance.

59 BASIC BIOLOGICAL SCIENCES↗

Coupled Biotic-Abiotic Processes Control Biogeochemical Cycling of Dissolved Organic Matter in the Columbia River Hyporheic Zone

A critical component of assessing the impacts of climate change on watershed ecosystems involves understanding the role that dissolved organic matter (DOM) plays in driving whole ecosystem metabolism. The hyporheic zone—a biogeochemical control point where ground water and river water mix—is characterized by high DOM turnover and microbial activity and is responsible for a large fraction of lotic respiration. Yet, the dynamic nature of this ecotone provides a challenging but important environment to parse out different DOM influences on watershed function and net carbon and nutrient fluxes. We used high-resolution Fourier-transform ion cyclotron resonance mass spectrometry to provide a detailed molecular characterization of DOM and its transformation pathways in the Columbia river watershed. Samples were collected from ground water (adjacent unconfined aquifer underlying the Hanford 300 Area), Columbia river water, and its hyporheic zone. The hyporheic zone was sampled at five locations to capture spatial heterogeneity within the hyporheic zone. Our results revealed that abiotic transformation pathways (e.g., carboxylation), potentially driven by abiotic factors such as sunlight, in both the ground water and river water are likely influencing DOM availability to the hyporheic zone, which could then be coupled with biotic processes for enhanced microbial activity. The ground water profile revealed high rates of N and S transformations via abiotic reactions. The river profile showed enhanced abiotic photodegradation of lignin-like molecules that subsequently entered the hyporheic zone as low molecular weight, more degraded compounds. While the compounds in river water were in part bio-unavailable, some were further shown to increase rates of microbial respiration. Together, river water and ground water enhance microbial activity within the hyporheic zone, regardless of river stage, as shown by elevated putative amino-acid transformations and the abundance of amino-sugar and protein-like compounds. This enhanced microbial activity is further dependent on the composition of ground water and river water inputs. Our results further suggest that abiotic controls on DOM should be incorporated into predictive modeling for understanding watershed dynamics, especially as climate variability and land use could affect light exposure and changes to ground water essential elements, both shown to impact the Columbia river hyporheic zone.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Pattern Dictionary Method for Anomaly Detection

In this paper, we propose a compression-based anomaly detection method for time series and sequence data using a pattern dictionary. The proposed method is capable of learning complex patterns in a training data sequence, using these learned patterns to detect potentially anomalous patterns in a test data sequence. The proposed pattern dictionary method uses a measure of complexity of the test sequence as an anomaly score that can be used to perform stand-alone anomaly detection. We also show that when combined with a universal source coder, the proposed pattern dictionary yields a powerful atypicality detector that is equally applicable to anomaly detection. The pattern dictionary-based atypicality detector uses an anomaly score defined as the difference between the complexity of the test sequence data encoded by the trained pattern dictionary (typical) encoder and the universal (atypical) encoder, respectively. We consider two complexity measures: the number of parsed phrases in the sequence, and the length of the encoded sequence (codelength). Specializing to a particular type of universal encoder, the Tree-Structured Lempel–Ziv (LZ78), we obtain a novel non-asymptotic upper bound, in terms of the Lambert W function, on the number of distinct phrases resulting from the LZ78 parser. This non-asymptotic bound determines the range of anomaly score. As a concrete application, we illustrate the pattern dictionary framework for constructing a baseline of health against which anomalous deviations can be detected.

97 MATHEMATICS AND COMPUTING↗

Star–Galaxy Image Separation with Computationally Efficient Gaussian Process Classification

Abstract We introduce a novel method for discerning optical telescope images of stars from those of galaxies using Gaussian processes (GPs). Although applications of GPs often struggle in high-dimensional data modalities such as optical image classification, we show that a low-dimensional embedding of images into a metric space defined by the principal components of the data suffices to produce high-quality predictions from real large-scale survey data. We develop a novel method of GP classification hyperparameter training that scales approximately linearly in the number of image observations, which allows for application of GP models to large-size Hyper Suprime-Cam Subaru Strategic Program data. In our experiments, we evaluate the performance of a principal component analysis embedded GP predictive model against other machine-learning algorithms, including a convolutional neural network and an image photometric morphology discriminator. Our analysis shows that our methods compare favorably with current methods in optical image classification while producing posterior distributions from the GP regression that can be used to quantify object classification uncertainty. We further describe how classification uncertainty can be used to efficiently parse large-scale survey imaging data to produce high-confidence object catalogs.

79 ASTRONOMY AND ASTROPHYSICS↗

Resolved Dwarf Galaxy Searches within ~5 Mpc with the Vera Rubin Observatory and Subaru Hyper Suprime-Cam*

We present a preview of the faint dwarf galaxy discoveries that will be possible with the Vera C. Rubin Observatory and Subaru Hyper Suprime-Cam in the next decade. In this work, we combine deep ground-based images from the Panoramic Imaging Survey of Centaurus and Sculptor (PISCeS) and extensive image simulations to investigate the recovery of faint, resolved dwarf galaxies in the Local Volume with a matched-filter technique. We adopt three fiducial distances - 1.5, 3.5, 5 Mpc, and quantitatively evaluate the effects on dwarf detection of varied stellar backgrounds, ellipticity, and Milky Way foreground contamination and extinction. We show that our matched-filter method is powerful for identifying both compact and extended systems, and near-future surveys will be able to probe at least ~4.5 mag below the tip of the red giant branch (TRGB) for a distance of up to 1.5 Mpc, and ~2 mag below the TRGB at 5 Mpc. This will push the discovery frontier for resolved dwarf galaxies to fainter magnitudes, lower surface brightnesses, and larger distances. Our simulations show the secure census of dwarf galaxies down to $M_{V}$$\approx$-5, -7, -8, will be soon within reach, out to 1.5 Mpc, 3.5 Mpc, and 5 Mpc, respectively, allowing us to quantify the statistical fluctuations in satellite abundances around hosts, and parse environmental effects as a function of host properties.

79 ASTRONOMY AND ASTROPHYSICS↗

Phase Doppler Interferometry for Efficient Cloud Drop Size Distribution, Number Density, and LWC Measurements

Threats to aviation safety as a result of super-cooled large drops (SLD) has been addressed by the FAA rules change (14 CFR Part 25) with the additional icing certification requirement. SLD clouds often consist of bi-modal drop size spectra leading to significant problems in simulating and characterizing these conditions in situ and in icing wind tunnels. Legacy instrumentation for measuring drop size distributions and liquid water content are challenged under these conditions. The large size range measurement problem is addressed with the development of the Phase Doppler Interferometer, Flight Probe Dual-Range (PDI FPDR). The method is described in this report along with the measurement capabilities including the dynamic measurement range and overall working size range. The PDI instrument bases drop size measurements on the light wavelength as the measurement length scale. The light wavelength is a much more robust scale, especially as compared to the light scattering intensity. Additionally, methods for accurately characterizing the sample volume in situ based on measured drop velocity and transit time are reviewed, given the importance of this parameter for merging results and measuring LWC. Droplet coincidence in the sample volume can be problematic so this condition is treated with an innovative signal parsing approach. Measurement examples acquired in the NASA IRT are provided. Measurements of LWC showed good agreement with the Artium Particle Imaging (PI) instrument but diverged from the tunnel calibration results for larger MVD values.

42 ENGINEERING↗

A novel methodology for assessing the hygroscopicity of aerosol filter samples

Abstract. Due to US regulations, concentrations of hygroscopic inorganic sulfate and nitrate have declined in recent years, leading to an increased importance of the hygroscopic nature of organic matter (OM). The hygroscopicity of OM is poorly characterized because only a fraction of the multitude of organic compounds in the atmosphere is readily measured, and there is limited information on their hygroscopic behaviors. Hygroscopicity of aerosol is traditionally measured using a humidified tandem differential mobility analyzer (HTDMA) or electrodynamic balance (EDB). EDB measures water uptake by a single particle. For ambient and chamber studies, HTDMA measurements provide water uptake and particle size information but not chemical composition. To fill this information gap, we developed a novel methodology to assess the water uptake by particles collected on Teflon filters. This method uses the same filter sample for both hygroscopicity measurements and chemical characterization, thereby providing an opportunity to link the measured hygroscopicity with ambient particle composition. To test the method, hygroscopic measurements were conducted in the laboratory for ammonium sulfate, sodium chloride, glucose, and malonic acid, which were collected on 25 mm Teflon filters using an aerosol generator and sampler. Constant-humidity solutions (CHSs), including potassium chloride, barium chloride dihydrate, and potassium sulfate, were employed in a saturated form to maintain the relative humidity (RH) at approximately 84 %, 90 %, and 97 % in small chambers. Our preliminary experiments revealed that, without the pouch, water uptake measurements were not feasible due to rapid water loss during weighing. Additionally, we observed some absorption by the aluminum pouch itself. To account for this, concurrent measurements were conducted for both the loaded and the blank filters at each RH level. Thus, the dry loaded and blank Teflon filters were placed in aluminum pouches with one side open and in RH-controlled chambers for more than 24 h. The wet loaded samples and wet blanks were then weighed using an ultramicrobalance to determine the water uptake by the respective compound and the blank Teflon filter. The net amount of water absorbed by each compound was calculated by subtracting the water uptake of the blank filter from that of the wet loaded filter. Hygroscopic parameters, including the water-to-solute (W / S) ratio, molality, mass fraction solute (mfs), and growth factors (GFs), were calculated from the measurements. The results obtained are consistent with those reported by the Extended Aerosol Inorganics Model (E-AIM) and previous studies utilizing HTDMA and EDB for these compounds, highlighting the accuracy of this new methodology. This new approach enables the hygroscopicity and chemical composition of individual filter samples to be assessed so that in complex mixtures, such as chamber and ambient samples, the total water uptake can be parsed between the inorganic and organic components of the aerosol.

54 ENVIRONMENTAL SCIENCES↗

Refactoring the elastic–viscous–plastic solver from the sea ice model CICE v6.5.1 for improved performance

This study focuses on the performance of the elastic–viscous–plastic (EVP) dynamical solver within the sea ice model, CICE v6.5.1. The study has been conducted in two steps. First, the standard EVP solver was extracted from CICE for experiments with refactored versions, which are used for performance testing. Second, one refactored version was integrated and tested in the full CICE model to demonstrate that the new algorithms do not significantly impact the physical results. The study reveals two dominant bottlenecks, namely (1) the number of Message Parsing Interface (MPI) and Open Multi-Processing (OpenMP) synchronization points required for halo exchanges during each time step combined with the irregular domain of active sea ice points and (2) the lack of single-instruction, multiple-data (SIMD) code generation. The standard EVP solver has been refactored based on two generic patterns. The first pattern exposes how general finite differences on masked multi-dimensional arrays can be expressed in order to produce significantly better code generation by changing the memory access pattern from random access to direct access. The second pattern takes an alternative approach to handle static grid properties. The measured single-core performance improvement is more than a factor of 5 compared to the standard implementation. The refactored implementation of strong scales on the Intel® Xeon® Scalable Processors series node until the available bandwidth of the node is used. For the Intel® Xeon® CPU Max series, there is sufficient bandwidth to allow the strong scaling to continue for all the cores on the node, resulting in a single-node improvement factor of 35 over the standard implementation. This study also demonstrates improved performance on GPU processors.

58 GEOSCIENCES↗

rustpix

rustpix is a high-performance, open-source Rust library with first-class Python bindings (via PyO3) for processing pixel-detector data in neutron imaging. It targets time-stamping detectors such as Timepix3 (TPX3) at ORNL's Spallation Neutron Source (VENUS beamline), where each detected neutron deposits charge across a cluster of pixels within a very high-rate event stream (96M+ hits/sec). rustpix parses TPX3 event data in parallel using memory-mapped I/O, offers four interchangeable clustering algorithms (ABS adjacency-based search, DBSCAN, graph/union-find connected components, and a parallel grid method), and extracts weighted, super-resolved centroids to produce neutron-event lists. A streaming architecture lets it process files larger than available memory. rustpix is distributed as a pip-installable Python package (with NumPy integration), Rust crates, a command-line tool, and an interactive GUI; it writes HDF5, Apache Arrow, and CSV; and it is designed to extend to TPX4 and other detector types. Released as open-source under the MIT License.

Zhang, Chen [Oak Ridge National Laboratory (ORNL),↗