Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “file transfer”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Magnetotelluric Data from Mountain Home, ID

This dataset includes magnetotelluric transfer functions in the form of EDI files for 16 stations collected by the USGS and 40 stations collected by Quantec Geoscience for Lawerence Berkeley National Lab around the Mountain Home area in Idaho. A 3D electrical resistivity model is included that images resistive and conductive bodies in the subsurface that maybe important for geothermal characterization. The model was created using ModEM using the high performance computer Yeti at the USGS.

15 GEOTHERMAL ENERGY↗

Data Management in the Continuum: Cross-facility Object-based Data Transfers

Scientific workflows are evolving from relying on a monolithic storage subsystem at a single High-Performance Computing (HPC) facility to using geographically distributed file systems, repositories, and cloud storage. As a result, storing, accessing, transferring, and managing scientific data have become highly complex and prone to performance inefficiencies. This paper delves into these challenges by exploring an optimized end-to-end interface designed to seamlessly connect various local and remote storage systems, enabling efficient data movement of objects across HPC–Cloud and HPC–HPC environments. We showcase this capability through an object-focused data management runtime system, discuss the effects of relaxed consistency semantics in distributed object scenarios, and illustrate its application in an earthquake simulation workflow. Besides reducing the amount of data by selectively transferring regions of interest, our facility-local results achieved a speedup of 45 × over an optimized HDF5 usage and 15 × over the HDF5 with caching by using the new interface in PDC-XF.

Bez, Jean Luca↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

Wave Tank Characterization Data

This data set was collected from a 16k gallon single paddle type wave tank after adjustments were made to the transfer function in the control software. The file names describe the commanded wave height and period which can then be compared with the actual measured height and period of the generated waves.

16 TIDAL AND WAVE POWER↗

Viability of S3 Object Storage for the ASC Program at Sandia

Recent efforts at Sandia such as DataSEA are creating search engines that enable analysts to query the institution’s massive archive of simulation and experiment data. The benefit of this work is that analysts will be able to retrieve all historical information about a system component that the institution has amassed over the years and make better-informed decisions in current work. As DataSEA gains momentum, it faces multiple technical challenges relating to capacity storage. From a raw capacity perspective, data producers will rapidly overwhelm the system with massive amounts of data. From an accessibility perspective, analysts will expect to be able to retrieve any portion of the bulk data, from any system on the enterprise network. Sandia’s Institutional Computing is mitigating storage problems at the enterprise level by procuring new capacity storage systems that can be accessed from anywhere on the enterprise network. These systems use the simple storage service, or S3, API for data transfers. While S3 uses objects instead of files, users can access it from their desktops or Sandia’s high-performance computing (HPC) platforms. S3 is particularly well suited for bulk storage in DataSEA, as datasets can be decomposed into object that can be referenced and retrieved individually, as needed by an analyst. In this report we describe our experiences working with S3 storage and provide information about how developers can leverage Sandia’s current systems. We present performance results from two sets of experiments. First, we measure S3 throughput when exchanging data between four different HPC platforms and two different enterprise S3 storage systems on the Sandia Restricted Network (SRN). Second, we measure the performance of S3 when communicating with a custom-built Ceph storage system that was constructed from HPC components. Overall, while S3 storage is significantly slower than traditional HPC storage, it provides significant accessibility benefits that will be valuable for archiving and exploiting historical data. There are multiple opportunities that arise from this work, including enhancing DataSEA to leverage S3 for bulk storage and adding native S3 support to Sandia’s IOSS library.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

LLNL Kimberlina 1.2 NUFT Simulations June 2018 (v2)

This dataset contains the output 6,000, 3-dimensional reactive multi-phase flow and transport aquifer simulations of brine and CO2 leakage into a protective aquiver in California’s San Joaquin Valley and input data files detailing the geologic mesh, aquifer physical properties and CO2 and brine injection rates. This data set was generated as an ongoing effort with the US DOE National Risk Assessment Partnership (NRAP) to evaluate the effectiveness of monitoring techniques to detect brine and CO2 leakage from legacy wells into underground sources of drinking water overlaying a CO2 storage reservoir. Each simulation contains a unique set of input parameters, generated stochastically. The outputs consist of these upper three geologic layers (from top): the Etchegoin, Macoma-Chanac, Santa Margarita-McLure formations. These simulations span the several distances (1, 3 and 6 km or wells W31-0.2, W31-0.5 and W31-1.0, respectively) from the CO2 injector, initiated from bottom hole pressure and saturation to calculate wellbore leakage from the storage reservoir, with low and high regional groundwater gradients and wellbore leakage into 5 leaky nodes. The dataset includes 1,000 unique simulations for each distance, which each contain a unique aquifer heterogeneity, aquifer and caprock permeability, and two model generations are included with a high permeability (prod07) and hybrid permeability (prod09). The range of permeability distributions is listed in Table 1. Each model generation consists of 3,000 simulations. Included in the dataset are the leakage rates determined from 2D wellbore models which utilize the pressure and CO2 saturation from LBL's reservoir simulations, NUFT mesh files with distributed lithology, NUFT rocktab files which describe the material properties for the geologic layers and the NUFT input files and post-processed output 'ntab' files. Each ntab file contains spatial (rows) and temporal (columns) model output tables for each model cell, the locations (x,y,z) and dimensions for each cells (dx, dy, dz). Table 1. Permeability distribution ranges for prod07 and prod09 model generations Geologic Layer: Permeability Range (log10 m^2) prod07 prod09 Etchegoin -12.92 to -10.92 -13.70 to -11.44 Macoma-Chanac -12.72 to -10.72 -13.50 to -11.24 Santa Margarita-McLure -12.70 to -10.70 -13.48 to -11.22 The input files used to generate the model include which are included in the dataset are: Time series of CO2 leakage input into the model (ex: Q_brn.W31-0.2.sim1000.layers123.tab) Time series of CO2 leakage input into the model (ex: Q_CO2.W31-0.2.sim1000.layers123.tab) Physical properties of the aquifer materials detailing the aquifer porosity, solid density, partitioning coefficients, permeabilities and van-Genuchten parameters detailed in a NUFT rocktab file: (ex: sim1000.usnt.rocktab) Numerical mesh and geologic data assigned to each model cell detailed in a NUFT genmsh format (ex: sim1000.mesh_k16.prod07.trans.genmsh) The primary output parameters are: pH (use absolute value) Change in TDS (mg/kg) Change in Pressure (Pa) Change CO2 gas saturation (fraction range 0.0-1.0) for example, the directory /p/lscratchh/mansoor1/nrap/kimberlina/prod09/mainfiles/sim1000/W31- 0.2 contains: sim1000.W31-0.2.trans.pH.red.ntab sim1000.W31-0.2.no_bg.trans.TDS.red.ntab sim1000.W31-0.2.usnt.P.deltabg.red.ntab sim1000.W31-0.2.usnt.CO2_sat.deltabg.red.ntab Each row in the NTAB files consist of model output per numerical grid cell. Each output file contains 33 columns (variables), including the information of numerical records, geologic location and sizes and the simulated parameter values over time. The first 13 variables are about numerical records and relative geologic information for a simulation grid: 1. index: simulation index 2. i: the ith grid of x-axis 3. j: the ith grid of y-axis 4. k: the ith grid of z-axis 5. element_ref: element reference 6. nuft_ind: nuft index 7. x: grid location in the x axis direction 8. y: grid location in the y axis direction 9. z: grid location in the z axis direction 10. dx: grid length in the x axis direction 11. dy: grid length in the y axis direction 12. dz: grid length in the z axis direction 13. volume: volume of the simulation grid The remainder (14, 15, 16...) variables are the simulated parameter values over time, take Pressure as an example, are: 14. 0.0y: initial pressure per cell. 15. 10.0y: simulated pressure at the end of the 10th year. 16. 20.0y: simulated pressure at the end of the 20th year. ... (repeated for every 10 years until 200 years)... The model extends 10,000 m, 5,000 m and 1,411 m in the x,y and z dimensions, respectively. The mesh consists of 164,832 cells with mesh dimensions of 101 x 51 x 32 (nx, ny, nz), with cell dimensions ranging from 100 m laterally (along x and y-axis) and model layers are as designated in the z-axis: Layer 1: atmosphere (1e-30 m thick) Layer 2: upper caprock (10 m thick) Layers 3-13: Etchegoin (536.23 m thck) Layers 14-27: Macoma-Chanac (679.04 m thick) Layers 28-32: Santa Margarita-McLure (185.94 m thick) The wellbore is placed along node i=51, j=26, and extends vertically along 5 nodes from the top to the bottom of the model. Special instructions when extracting files: Each Gzip archive (ex: prod07.sim1000-sim00099.tar.gz) contains 100 simulations. Gzip archives should be transferred into base directories (ie. In Linux: mkdir prod07; mv prod07.*.tar.gz prod07/.) before extracting, or files will be overwritten. Each sub-simulation tree should have the following file structure pattern (using the linux 'tree' command): |-- prod07 | |-- sim0001 | |-- W31-0.2 | | |-- Q_brn.W31-0.2.sim0001.layers123.tab | | |-- Q_co2.W31-0.2.sim0001.layers123.tab | | |-- sim0001.W31-0.2.no_bg.trans.TDS.red.ntab | | |-- sim0001.W31-0.2.trans.pH.red.ntab | | |-- sim0001.W31-0.2.usnt.CO2_sat.deltabg.red.ntab | | |-- sim0001.W31-0.2.usnt.P.deltabg.red.ntab | |-- W31-0.5 | | |-- Q_brn.W31-0.5.sim0001.layers123.tab | | |-- Q_co2.W31-0.5.sim0001.layers123.tab | | |-- sim0001.W31-0.5.no_bg.trans.TDS.red.ntab | | |-- sim0001.W31-0.5.trans.pH.red.ntab | | |-- sim0001.W31-0.5.usnt.CO2_sat.deltabg.red.ntab | | |-- sim0001.W31-0.5.usnt.P.deltabg.red.ntab | |-- W31-1.0 | | |-- Q_brn.W31-1.0.sim0001.layers123.tab | | |-- Q_co2.W31-1.0.sim0001.layers123.tab | | |-- sim0001.W31-1.0.no_bg.trans.TDS.red.ntab | | |-- sim0001.W31-1.0.trans.pH.red.ntab | | |-- sim0001.W31-1.0.usnt.CO2_sat.deltabg.red.ntab | | |-- sim0001.W31-1.0.usnt.P.deltabg.red.ntab | |-- sim0001.mesh_k16.prod07.trans.genmsh Disclaimer This document was prepared as an account of work sponsored by an agency of the United States government. Neither the United States government nor Lawrence Livermore National Security, LLC, nor any of their employees makes any warranty, expressed or implied, or assumes any legal liability or responsibility for the accuracy, completeness, or usefulness of any information, apparatus, product, or process disclosed, or represents that its use would not infringe privately owned rights. Reference herein to any specific commercial product, process, or service by trade name, trademark, manufacturer, or otherwise does not necessarily constitute or imply its endorsement, recommendation, or favoring by the United States government or Lawrence Livermore National Security, LLC. The views and opinions of authors expressed herein do not necessarily state or reflect those of the United States government or Lawrence Livermore National Security, LLC, and shall not be used for advertising or product endorsement purposes. Lawrence Livermore National Laboratory is operated by Lawrence Livermore National Security, LLC, for the U.S. Department of Energy, National Nuclear Security Administration under Contract DE-AC52-07NA27344. This report was reviewed and released as LLNL-MI-753464.

aquifer↗

NOvA Inclusive Nue CC Cross Section Data Release

NOvA inclusive electron neutrino cross section results presented at Neutrino2020Exposure:8.09E20 protons-on-target, neutrino-enhanced beamData release contains 3 ROOT files containing the double-differential measurement with respect to electron angle (cos theta) and electron energy, a single-differential measurement with respect to squared four-momentum transfer (Q2), and the total cross section as a function of neutrino energy. The 1D cross-section measurements include the same kinematic phase space restrictions present in the double-differential analysis.The files, xsec_{dThdE/Enu/dQ2}.root, contain the extracted cross section and associated covariance matrices for the double-differential cross section with respect to the observed electron kinematics, dThdE, neutrino energy, Enu, and squared four-momentum transfer, Q^2.In each file there will be the following histograms:— TH2D/TH1D xsec: unfolded measured cross section— TH2D xsec_cov: cross section covariance matrix— TH2D stat_cov: statistical covariance matrixIncludes text files, xsec_{dThdE/Enu/dQ2}.txt show the corresponding histograms in plain text format.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Distributed Computing for the Project 8 Experiment

The Project 8 collaboration aims to measure the absolute neutrino mass or improve on the current limit by measuring the tritium beta decay electron spectrum. We present the current distributed computing model for the Project 8 experiment. Project 8 is in its second phase of data taking with a near continuous data rate of 1Gbps. The current computing model uses DIRAC (Distributed Infrastructure with Remote Agent Control) for its workflow and data management. A detailed meta-data assignment using the DIRAC File Catalog is used to automate raw data transfers and subsequent stages of data processing. The DIRAC system is deployed on containers managed using a Kubernetes cluster to provide a scalable infrastructure. A modified DIRAC Site Director provides the ability to submit jobs using Singularity on opportunistic High-Performance Computing (HPC) sites.

Distributed Computing, Kubernetes, DIRAC, Project ↗

TRAILS Output Files

Overview This data repository contains ZIP files that store compressed versions of the output of running the WaterPaths utility planning and management tool in the DU Re-Evaluation mode (to download the tool, please see this GitHub repository). The tool was used to simulate the six-utility North Carolina Research Triangle problem. Details on the contents of each ZIP file can be seen below. Data details Temporal range: Weekly data for 2,344 weeks from 2015 to 2060 (45 years). Spatial range: Six water utilities in the North Carolina Research Triangle region (0: Chapel Hil/OWASA, 1: Durham, 2: Cary, 3: Raleigh, 4: Pittsboro, and 5: Chatham) File types: CSV and OUT Different solutions available The solution numbers correspond to the different pathway strategies (henceforth referred to as "solutions") discussed in paper's main and supporting text (abstract and link to the paper here). They are as follows: Sol92: The Durham-focused pathway strategy Sol132: The Raleigh-focused pathway strategy Sol140: The regionally-robust pathway strategy Objectives files These files can be accessed by unzipping solXX_objectives_pathways.zip that contains 1,000 Objectives_RDMXX_solsXX_to_XX.csv files. Each CSV file will consist of a row representing all the objective values for that specific solution, while every six columns represents the reliability, restriction frequency, infrastructure net present value ($ mil), peak financial cost, worst-case cost, and unit cost ($ per MG; in that order) for each of the six utilities. There will be 1,000 such files, denoting the performance of the six utilities across the 1,000 deeply uncertain states of the world (DU SOWs). Pathway files These files can be accessed by unzipping solXX_objectives_pathways.zip that contains 1,000 Pathways_sXX_RDMXX.out file. Each OUT corresponds to the set of infrastructure being triggered in a specific DU SOW, and each file will have the name file will consist of four tab-delimited columns that are described as follows: Realization: The realization in which an infrastructure options being triggered utility: The utility currently triggering infrastructure week: The week in which a specific infrastructure option is being triggered infra.: The infrastructure option being triggered If the OUT file contains only the header line, no infrastructure was triggered for that specific DU SOW. Policies files These files can be obtained by unzipping Policies.zip. Each of the 1,000 CSV files within the unzipped folder will contain weekly water use restriction policies for all 1,000 hydroclimatic realizations within a specific DU SOW. The column structure is as follows: 0rest_m: restriction multiplier for utility 0 (values between 0 and 1) 1rest_m: restriction multiplier for utility 1 (values between 0 and 1) 2rest_m: restriction multiplier for utility 2 (values between 0 and 1) 3rest_m: restriction multiplier for utility 3 (values between 0 and 1) 4rest_m: restriction multiplier for utility 4 (values between 0 and 1) 5rest_m: restriction multiplier for utility 5 (values between 0 and 1) 0transf: transfer volume for utility 0 (in MGD) 1transf: transfer volume for utility 1 (in MGD) 2transf: transfer volume for utility 2 (in MGD) 3transf: transfer volume for utility 3 (in MGD) 4transf: transfer volume for utility 4 (in MGD) 5transf: transfer volume for utility 5 (in MGD) Water Sources files These files can be obtained by unzipping WaterSources_subset.zip. Each of the 100 CSV files within the unzipped folder will contain weekly state variables at each water source for all 1,000 hydroclimatic realizations within a specific DU SOW. The column structure is as follows: Xvolume: available water volume from source X (in MGD) Xs_area: surface area of source X (in ACF) Xdemand: demand drawn from a water source from source X (in MGD) Xup_spill: upstream spillage from source X (in MGD) Xww_inflow: wastewater inflow from source X (in MGD) Xcatch_inflow: upstream catchment inflow to source X (in MGD) Xevap: evaporation multiplier for source X (values between 0 and 1) Xds_spill: downstream spillage from source X (in MGD) X_Y_alloc_cap: the allocated capacity from source X to utility Y (values between 0 and 1) X_Y_alloc_dem: the allocated demand from source X to utility Y (values between 0 and 1) Xtrmt_alloc_Y: the allocated treatment capacity from source X to utility Y (values between 0 and 1) Utilities files These files can be obtained by unzipping Utilities_subset.zip. Each of the 100 CSV files within the unzipped folder will contain weekly state variables at each utility for all 1,000 hydroclimatic realizations within a specific DU SOW. The column structure is as follows: Xst_vol: total available storage volume of utility X (in MG) Xcapacity: total storage capacity of utility X (in MG) Xnet_inf: : net inflow for all storage infrastructure for utility X (in MGD) Xst_rof: short term ROF for utility X (values between 0 and 1) Xst_stor_rof: short-term storage ROF for utility X (values between 0 and 1) Xst_trmt_rof: short-term treatment ROF for utility X (values between 0 and 1) Xlt_rof: long-term ROF for utility X (values between 0 and 1) Xlt_stor_rof: long-term storage ROF for utility X (values between 0 and 1) Xlt_trmt_rof: long-term treatment ROF for utility X (values between 0 and 1) Xrest_demand: restricted demand for utility X (in MGD) Xunrest_demand: unrestricted demand for utility X (in MGD) Xunfulf_demand: unfulfilled demand for utility X (in MGD) Xwastewater: wastewater return for utility X (in MGD) Xtreat_capacity: total treatment capacity for utility X (in MG) Xcont_fund: reserve (contingency) fund balance for utility X Xins_pout: insurance payout for utility X (% annual volumetric revenue) Xins_price: insurance price for utility X (% annual volumetric revenue) Xinfra_npv: infrastructure net present value for utility ($mil) Xst_vol: total available storage volume of utility X (in MG) Xdebt_serv: debt service for utility X (usually once per year if the infrastructure is triggered; % annual volumetric revenue) Xstor_vol: total stored volume (in MGD) Xobs_ann_dem: observed annual demand for utility X (in MGD) Xproj_dem: projected annual demand for utility X (in MGD) Xpv_debt_serv: present value of debt service payments for utility X (% annual volumetric revenue) Xgross_rev: gross revenue for utility X ($mil) Acknowledgment IM3 is a multi-institutional effort led by Pacific Northwest National Laboratory and supported by the U.S. Department of Energy's Office of Science as part of research in MultiSector Dynamics, Earth and Environmental Systems Modeling Program.

Artificial Intelligence↗

A machine-learning-driven data labeling pipeline for scientific analysis in MLExchange

This study introduces a novel labeling pipeline to accelerate the labeling process of scientific data sets by using artificial intelligence (AI)-guided tagging techniques. This pipeline includes a set of interconnected web-based graphical user interfaces (GUIs), where Data Clinic and MLCoach enable the preparation of machine learning (ML) models for data reduction and classification, respectively, while Label Maker is used for label assignment. Throughout this pipeline, data can be accessed through a direct connection to a file system or through Tiled for access through Hypertext Transfer Protocol (HTTP). Our experimental results present three use cases where this labeling pipeline has been instrumental for the study of large X-ray scattering data sets in the area of pattern recognition, the remote analysis of resonant soft X-ray scattering data and the fine-tuning process of foundation models. These use cases highlight the labeling capabilities of this pipeline, including the ability to label large data sets in a short period of time, to perform remote data analysis while minimizing data movement and to enhance the fine-tuning process of complex ML models with human involvement.

Chavez, Tanny (ORCID:0000000193172896)↗

funcX: Federated Function as a Service for Science

Here, funcX is a distributed function as a service (FaaS) platform that enables flexible, scalable, and high performance remote function execution. Unlike centralized FaaS systems, funcX decouples the cloud-hosted management functionality from the edge-hosted execution functionality. funcX's endpoint software can be deployed, by users or administrators, on arbitrary laptops, clouds, clusters, and supercomputers, in effect turning them into function serving systems. funcX's cloud-hosted service provides a single location for registering, sharing, and managing both functions and endpoints. It allows for transparent, secure, and reliable function execution across the federated ecosystem of endpoints-enabling users to route functions to endpoints based on specific needs. funcX uses containers (e.g., Docker, Singularity, and Shifter) to provide common execution environments across endpoints. funcX implements various container management strategies to execute functions with high performance and efficiency on diverse funcX endpoints. funcX also integrates with an in-memory data store and Globus for managing data that may span endpoints. We motivate the need for funcX, present our prototype design and implementation, and demonstrate, via experiments on two supercomputers, that funcX can scale to more than 130000 concurrent workers. We show that funcX's container warming-aware routing algorithm can reduce the completion time for 3,000 functions by up to 61% compared to a randomized algorithm and the in-memory data store can speed up data transfers by up to 3x compared to a shared file system.

97 MATHEMATICS AND COMPUTING↗

Processing MCNP Elemental Edit Outputs

The Monte Carlo N-Particle (MCNP) transport code version 6 (also known as MCNP6) has the capability for tracking particles on unstructured mesh (UM) geometry models embedded into constructive solid geometry (CSG) cells. A UM geometry is a collection of elements representing a solid geometry. The first step of MCNP UM modeling is using other software packages to create a finite element mesh representation of a solid 3D geometry. Computer-aided design (CAD) or computer-aided manufacturing (CAM) software is typically used to create a solid geometry model, which is later imported into mesh generation software to create a UM model. The MCNP UM feature was originally designed for models generated by the Abaqus/CAE software. The MCNP code version 6.0 and later can process UM models formatted as Abaqus input files. MCNP can process a UM model consisting of several different element types including linear tetrahedral or hexahedral elements and calculate quantities of interest such as flux and energy deposition at elements. An MCNP UM simulation provides high-fidelity elemental edit (i.e., tally) outputs, which can be further used in multiphysics calculations. The MCNP UM feature was used for multiphysics simulations where quantities of interest calculated by MCNP are used as inputs for heat transfer calculations in Abaqus. MCNP6.3 can produce two types of elemental edit output (EEOUT) file formats: ASCII and HDF5. An EEOUT file type must be requested on an EMBED card while output type (flux or energy deposition) must be requested on an EMBEE card. We wrote Python3 scripts to extract energy deposition values in an ASCII or HDF5 EEOUT file and compute a heat flux profile for an Abaqus heat transfer calculation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A case study on parallel HDF5 dataset concatenation for high energy physics data analysis

In High Energy Physics (HEP), experimentalists generate large volumes of data that, when analyzed, helps us better understand the fundamental particles and their interactions. This data is often captured in many files of small size, creating a data management challenge for scientists. In order to better facilitate data management, transfer, and analysis on large scale platforms, it is advantageous to aggregate data further into a smaller number of larger files. However, this translation process can consume significant time and resources, and if performed incorrectly the resulting aggregated files can be inefficient for highly parallel access during analysis on large scale platforms. In this paper, we present our case study on parallel I/O strategies and HDF5 features for reducing data aggregation time, making effective use of compression, and ensuring efficient access to the resulting data during analysis at scale. We focus on NOvA detector data in this case study, a large-scale HEP experiment generating many terabytes of data. Here, the lessons learned from our case study inform the handling of similar datasets, thus expanding community knowledge related to this common data management task.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Kinetic Model of HoxEFU reduction by NADH [SWR-26-087]

This repository is used to release code generated for manuscripts on the Photosynthetic Energy Transduction core program. This code simulates the reduction of HoxEFU by NADH. The electro transfer rate constants for the simulation are specified in the .csv files. The two .csv files correspond tot he two models described in Dawson et al. Cell. Rep. Phys. Sci. 2026. The code utilizes a chemical master equation, a set of differential equations, defining the time evolution of the oxidation and reduction kinetics of NAD+, NADH, a FMN flavin, and a set of iron sulfur clusters. The kinetics of HoxEFU reduction by NADH are evaluated by numerical integration of the chemical master equation using a variable-time-step Runge-Kutta algorithm.

Dahl, Peter [National Laboratory of the Rockies (N↗

Scribe Network API v.1.0.0

SAND2023-07864O Scribe Network API is an add-on tool for Scribe3D software. Scribe Network API documents tabletop exercises in trainings and plays back simulated videos of the scenarios and responses. Users can apply the software to compiled projects, or to a simple visual studio project solution package. The software runs a web application that relays information between computers using Scribe3D through a web application, or web app. The web app involves a representational state transfer (REST) application programming interface that handles sending and receiving Scribe save files and an SQL server that stores the save files. This follow-on package allows users to facilitate a networked tabletop exercise. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Noel, Todd↗