Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Using Apptainer in a Pilot-based Distributed Workload

GlideinWMS is a pilot and pressure-based workload manager for distributed scientific computing. Many experiments like CMS and Fermilab’s Neutrino experiments use it to provision elastic clusters for their analysis and simulations, split into close to a million concurrent jobs. Most user jobs require containers, and the pilots use Apptainer to set up the desired platform. For the pilots that run as regular batch jobs, Apptainer is safer, lighter, and easier to use than other containerization solutions. Many images used by the pilots are expanded SIF images distributed via the CernVM-FS: this combination is very efficient. At Fermilab, for example, we store on GitHub Dockerfiles that mimic the platform in the worker nodes of local clusters. GitHub workflows build and push the images to Docker Hub, and a service periodically pulls and converts them to the expanded SIF images in the CernVM-FS, so the scientists can find a familiar environment everywhere. Apptainer has also been used to run services inside the pilot jobs, like benchmarks that characterize the worker node being used, or a Triton Inference Server that allows sharing a GPU with all the jobs that run in parallel on a node.

Mambelli, Marco [Fermilab] (ORCID:0000000294892681↗

Secure Federated Learning Across Heterogeneous Cloud and High-Performance Computing Resources: A Case Study on Federated Fine-Tuning of LLaMA 2

Federated learning enables multiple data owners to collaboratively train robust machine learning models without transferring large or sensitive local datasets by only sharing the parameters of the locally trained models. Here, in this article, we elaborate on the design of our Advanced Privacy-Preserving Federated Learning (APPFL) framework, which streamlines end-to-end secure and reliable federated learning experiments across cloud computing facilities and high-performance computing resources by leveraging Globus Compute, a distributed function as a service platform, and Amazon Web Services. We further demonstrate the use case of APPFL in fine-tuning an LLaMA 2 7B model using several cloud resources and supercomputers.

97 MATHEMATICS AND COMPUTING↗

A Model-Free Voltage Control Approach to Mitigate Motor Stalling and FIDVR for Smart Grids

Electric power networks are large and highly nonlinear dynamical systems that present unique challenges to control design. Though there is a large number of dynamic models for power system stability and control, many models are only useful with right assumptions and wrong for other tasks. Moreover, the dynamic behavior of the grid is increasingly complex under the banner of smart grids. These lead to the difficulty of developing appropriate dynamic modeling, and thus an efficient control strategy. To avoid such modeling challenges, here we present a novel dynamic voltage control strategy based on a model-free control (MFC) approach, requiring no modeling procedure. In particular, it focuses on fault-induced delayed voltage recovery (FIDVR) events, which require complex and accurate dynamic load models to replicate such events. This work utilizes MFC as an online controller to achieve the desired voltage stability under the FIDVR event. The proposed MFC strategy allows simple implementation and low computational cost for efficient mitigation of FIDVR. For benchmarking, a reasonably accurate dynamic performance model is explored. Simulation results with the IEEE 57 bus test network demonstrate the enhanced dynamic voltage profile for load buses having induction motors with the support of reactive power resources.

24 POWER TRANSMISSION AND DISTRIBUTION↗

D2NO: Efficient handling of heterogeneous input function spaces with distributed deep neural operators

Neural operators have been applied in various scientific fields, such as solving parametric partial differential equations, dynamical systems with control, and inverse problems. However, challenges arise when dealing with input functions that exhibit heterogeneous properties, requiring multiple sensors to handle functions with minimal regularity. To address this issue, discretization-invariant neural operators have been used, allowing the sampling of diverse input functions with different sensor locations. However, existing frameworks still require an equal number of sensors for all functions. We propose a novel distributed approach to further relax the discretization requirements and solve the heterogeneous dataset challenges. Our method involves partitioning the input function space and processing individual input functions using independent and separate neural networks. A centralized neural network is used to handle shared information across all output functions. This distributed methodology reduces the number of gradient descent back-propagation steps, improving efficiency while maintaining accuracy. Here, we demonstrate that the corresponding neural network is a universal approximator of continuous nonlinear operators and present three numerical examples to validate its performance.

97 MATHEMATICS AND COMPUTING↗

Portable Software Environment for Ultrahigh-Resolution ELM Development on GPUs

This paper presents our endeavors in developing the large-scale, ultra-high-resolution E3SM Land Model (uELM), specifically designed for exascale computers furnished with accelerators such as Nvidia GPUs. The uELM is a sophisticated code that substantially relies on High-Performance Computing (HPC) environments, necessitating particular machine and software configurations. To facilitate community-based uELM developments employing GPUs, we have created a portable, standalone software environment preconfigured with uELM input datasets, simulation cases, and source code. This environment, utilizing Docker, encompasses all essential code, libraries, and system software for uELM development on GPUs. It also features a functional unit test framework and an offline model testbed for comprehensive numerical experiments. From a technical perspective, the paper discusses GPU-ready container generations, uELM code management, and input data distribution across computational platforms. Lastly, the paper demonstrates the use of environment for functional unit testing, end-to-end simulation on CPUs and GPUs, and collaborative code development.

E3SM Land Model↗

Virtual Infrastructure Twins: Software Testing Platforms for Computing-Instrument Ecosystems

Science ecosystems are being built by federating computing systems and instruments located at geographically distributed sites over wide-area networks. These computing-instrument ecosystems are expected to support complex workflows that incorporate remote, automated AI-driven science experiments. Their realization, however, requires various designs to be explored and software components to be developed, in order to support the orchestration of distributed computations and experiments. It is often too expensive, infeasible, or disruptive for the entire ecosystem to be available during the typically long software development and testing periods. We propose a Virtual Infrastructure Twin (VIT) of the ecosystem that emulates its network and computing components, and incorporates its instrument software simulators. It provides a software environment nearly identical to the ecosystem to support early development and testing, and design space exploration. We present a brief overview of previous digital infrastructure twins that culminated in the VIT concept, including (i) the virtual science network environment for developing software-defined networking solutions, and (ii) the virtual federated science instrument environment for testing the federation software stack and remote instrument control software. We briefly describe VITs for Nion microscope steering and access to GPU systems.

Rao, Nageswara↗

Virtual Infrastructure Twins: Software Testing Platforms for Computing-Instrument Ecosystems

Science ecosystems are being built by federating computing systems and instruments located at geographically distributed sites over wide-area networks. These computing-instrument ecosystems are expected to support complex workflows that incorporate remote, automated AI-driven science experiments. Their realization, however, requires various designs to be explored and software components to be developed, in order to support the orchestration of distributed computations and experiments. It is often too expensive, infeasible, or disruptive for the entire ecosystem to be available during the typically long software development and testing periods. We propose a Virtual Infrastructure Twin (VIT) of the ecosystem that emulates its network and computing components, and incorporates its instrument software simulators. It provides a software environment nearly identical to the ecosystem to support early development and testing, and design space exploration. We present a brief overview of previous digital infrastructure twins that culminated in the VIT concept, including (i) the virtual science network environment for developing software-defined networking solutions, and (ii) the virtual federated science instrument environment for testing the federation software stack and remote instrument control software. We briefly describe VITs for Nion microscope steering and access to GPU systems.

Rao, Nageswara↗

Phase Demodulation by Frequency Chirping in Coherence Microwave Photonic Interferometry

This paper presents a signal processing method to demodulate the optical interference phase of cascaded individual optical fiber intrinsic Fabry-Perot interferometric (IFPI) sensors in a coherent microwave-photonic interferometry (CMPI) distributed sensing system. The new method utilizes the chirp effect of electro-optic modulator (EOM) to create a quasi-quadrature optical interference phase shift between two adjacent pulses which correspond to two adjacent reflection points in the time domain. The phase shift can be controlled by adjusting the bias voltage that is applied to the EOM. The interference phase is calculated by elliptically fitting the phase shift. The interference phase change is proportional to the optical path difference (OPD) change of the interferometer, and the sign can be used to differentiate the increase or decrease of the OPD. The method is demonstrated for distributed strain sensing, showing good linearity, high resolution and large dynamic range.

42 ENGINEERING↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

Iteration-based Linearized Distribution-level Locational Marginal Price for Three-phase Unbalanced Distribution Systems

Distributed energy resources (DERs) are rocking the utilities’ business landscape. It calls for competitive market environments that incentivize DERs to form maximum operating efficiency. Among proposed pricing schemes, distribution-level locational marginal price (DLMP) is effective in signaling the marginal generation cost differences driven by energy losses and network constraints. It can be derived from a distribution-level optimal power flow (OPF) framework, as it essentially presents the sensitivity of optimized generation cost towards incremental loads. However, due to the high resistance-to-inductance ratio and unbalanced characteristics of distribution networks, computational affordable DLMPs are highly challenged. This article provides a linear-approximated DLMP that can be solved efficiently and generalized to account for reactive power flow, three-phase unbalanced loads and meshed network structure. The successive linear programming technique is introduced to enhance the model accuracy. Case studies on an IEEE 123-Bus system validate its accuracy against a nonlinear benchmark and capability in offering proper incentives.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Generating Massive Scale-free Networks: Novel Parallel Algorithms using the Preferential Attachment Model

Recently, there has been substantial interest in the study of various random networks as mathematical models of complex systems. As real-life complex systems grow larger, the ability to generate progressively large random networks becomes all the more important. This motivates the need for efficient parallel algorithms for generating such networks. Naïve parallelization of sequential algorithms for generating random networks is inefficient due to inherent dependencies among the edges and the possibility of creating duplicate (parallel) edges. In this article, we present message passing interface-based distributed memory parallel algorithms for generating random scale-free networks using the preferential-attachment model. Our algorithms are experimentally verified to scale very well to a large number of processing elements (PEs), providing near-linear speedups. The algorithms have been exercised with regard to scale and speed to generate scale-free networks with one trillion edges in 6 minutes using 1,000 PEs.

97 MATHEMATICS AND COMPUTING↗

Optimizing High-Throughput Inference on Graph Neural Networks at Shared Computing Facilities with the NVIDIA Triton Inference Server

Abstract With machine learning applications now spanning a variety of computational tasks, multi-user shared computing facilities are devoting a rapidly increasing proportion of their resources to such algorithms. Graph neural networks (GNNs), for example, have provided astounding improvements in extracting complex signatures from data and are now widely used in a variety of applications, such as particle jet classification in high energy physics (HEP). However, GNNs also come with an enormous computational penalty that requires the use of GPUs to maintain reasonable throughput. At shared computing facilities, such as those used by physicists at Fermi National Accelerator Laboratory (Fermilab), methodical resource allocation and high throughput at the many-user scale are key to ensuring that resources are being used as efficiently as possible. These facilities, however, primarily provide CPU-only nodes, which proves detrimental to time-to-insight and computational throughput for workflows that include machine learning inference. In this work, we describe how a shared computing facility can use the NVIDIA Triton Inference Server to optimize its resource allocation and computing structure, recovering high throughput while scaling out to multiple users by massively parallelizing their machine learning inference. To demonstrate the effectiveness of this system in a realistic multi-user environment, we use the Fermilab Elastic Analysis Facility augmented with the Triton Inference Server to provide scalable and high-throughput access to a HEP-specific GNN and report on the outcome.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

T-FSM: A Scalable Distributed Task-Based System for Frequent Subgraph Pattern Mining from a Big Graph

Finding frequent subgraph patterns in a big graph is an important problem with many applications such as classifying chemical compounds and building indexes to speed up graph queries. Since this problem is NP-hard, some recent parallel and distributed systems have been developed to accelerate the mining. However, they often have a huge memory cost, very long running time, suboptimal load balancing, poor scale-out capability, and possibly inaccurate results. In this article, we propose an efficient system called T-FSM for parallel mining of frequent subgraph patterns in a big graph. T-FSM supports a new anti-monotonic frequentness measure called Fraction-Score, which is more accurate than the widely used MNI measure. The execution engine of T-FSM supports both intra-machine parallelism and inter-machine parallelism. For intra-machine parallelism, T-FSM adopts a novel task-based execution model to ensure high multithreading concurrency, bounded memory consumption, and effective load balancing. For inter-machine parallelism, T-FSM ensures good scale-out performance with a lightweight pattern rebalancing approach that reduces workload skewness of pattern evaluations among machines. To avoid recomputing the contexts for migrated patterns, we design a novel context cache table to support concurrent and asynchronous requesting and caching of remote context data, which can timely evict and garbage collect used pattern contexts that are no longer needed to keep memory consumption bounded. Extensive experiments show that T-FSM is orders of magnitude faster than existing state-of-the-art parallel systems (more than 10×, 51×, 131×, 55× speedup over ScaleMine, DistGraph, Pangolin and Peregrine, respectively) and distributed systems (more than 42× and 88× over ScaleMine and DistGraph, respectively) for frequent subgraph pattern mining, and it scales out satisfactorily to 512 CPU cores on the Polaris supercomputer at Argonne National Laboratory.

97 MATHEMATICS AND COMPUTING↗

Clustering at Massive Scale

ClaMS provides hierarchical clustering technology for use on massive, high-dimensional datasets that require distributed memory for processing. The algorithm employed is inspired by the popular HDBSCAN algorithm but makes use of computational kernels better suited for distributed computing. ClaMS is built on scalable nearest neighbor graph construction, metric forest completion, and approximate minimum spanning tree techniques.

Stanley, ThomasA [Lawrence Livermore National Labo↗

Modular chip-integrated photonic control of artificial atoms in diamond waveguides

A central goal in creating long-distance quantum networks and distributed quantum computing is the development of interconnected and individually controlled qubit nodes. Atom-like emitters in diamond have emerged as a leading system for optically networked quantum memories, motivating the development of visible-spectrum, multi-channel photonic integrated circuit (PIC) systems for scalable atom control. However, it has remained an open challenge to realize optical programmability with a qubit layer that can achieve high optical detection probability over many optical channels. Here, we address this problem by introducing a modular architecture of piezoelectrically actuated atom-control PICs (APICs) and artificial atoms embedded in diamond nanostructures designed for high-efficiency free-space collection. The high-speed four-channel APIC is based on a splitting tree mesh with triple-phase shifter Mach–Zehnder interferometers. This design simultaneously achieves optically broadband operation at visible wavelengths, high-fidelity switching (>40dB) at low voltages, submicrosecond modulation timescales (>30MHz), and minimal channel-to-channel crosstalk for repeatable optical pulse carving. Via a reconfigurable free-space interconnect, we use the APIC to address single silicon vacancy color centers in individual diamond waveguides with inverse tapered couplers, achieving efficient single photon detection probabilities (∼15%) and second-order autocorrelation measurements g (2) (0)<0.14 for all channels. The modularity of this distributed APIC–quantum memory system simplifies the quantum control problem, potentially enabling further scaling to thousands of channels.

47 OTHER INSTRUMENTATION↗

Non-Diffusive Volume Advection with A High Order Interface Reconstruction Method

We show that non-diffusive volume advection in two-dimensions is achieved with several benchmark problems using a newly developed high-order volume of fluids (VOF) interface reconstruction method. (1) A new VOF interface reconstruction method using circular/corner facets (linear facets are a degenerate case of arcs). We create a circular interface facet in each mixed zone by matching neighbor volume with a hybrid Newton’s-bisection method and the local solution is final. In the general case, the new VOF interface reconstruction has 3rd order accuracy and can be easily made seamless. The new method addresses intrinsic issues with Young’s method such as gaps between interface facets in the case of a curved interface, and inability to define curvature nor identify corners. (2) A non-diffusive volume advection scheme. In an ALE advection step, a well-defined interface can be carried over through a Lagrange step and used to compute volume distribution into a relaxed mesh. Then, an interface reconstruction step is performed to redefine the interface in the relaxed mesh. We must point out that the interface carried over is also a solution of interface reconstruction because all the volume fractions in the relaxed mesh are naturally matched. We provide an interface tracking method compatible with our reconstruction scheme, where it is granted to use the prior info as an initial guess to capture sub-mesh resolution features. As a result, we are able to treat multiple facets inside a single mixed cell and obtain highly accurate, non-diffusive solution for advection problems with rather coarse meshes. We show our solutions for two-dimensional incompressible flows with two materials with a) the X + O diagonal translation; b) the Zalesak rotational test; and c) the single vortex spiral test.

97 MATHEMATICS AND COMPUTING↗

Assessing DER Network Cybersecurity Defences in a Power-Communication Co-Simulation Environment

Increasing penetrations of interoperable distributed energy resources (DER) in the electric power system are expanding the power system attack surface. Maloperation or malicious control of DER equipment can now cause substantial disturbances to grid operations. Fortunately, many options exist to defend and limit adversary impact on these newly-created DER communication networks, which typically traverse the public internet. However, implementing these security features will increase communication latency, thereby adversely impacting real-time DER grid support service effectiveness. In this work, a collection of software tools called SCEPTRE were used to create a co-simulation environment where SunSpec-compliant PV inverters were deployed as virtual machines and interconnected to simulated communication network equipment. Network segmentation, encryption, and moving target defence security features were deployed on the control network to evaluate their influence on cybersecurity metrics and power system performance. The results indicated that adding these security features did not impact DER-based grid control systems but improved the cybersecurity posture of the network when implemented appropriately.

97 MATHEMATICS AND COMPUTING↗

Forecasting Dynamic Line Rating with Spatial Variation Considerations

Dynamic line rating (DLR) is a technology that allows the ampacity of an electrical conductor to be calculated using real-time or forecasted weather conditions. Historically, the ampacity of a conductor has been determined using a static line rating method which assumes conservative weather assumptions. Therefore, not only can DLR give a more accurate measurement of the true ampacity of a conductor, but it can also increase its ampacity during weather conditions with greater thermal mitigations. The two primary cooling factors in the ampacity calculations are wind speed and direction. In complex terrain, wind speed and direction can have large variations over short distances. Therefore, accurately identifying the limiting span of a transmission line requires high spatial resolution of the wind along its path. One solution is to install dense weather stations along their path, though this can become costly over long distances. Therefore, researchers have investigated the use of Computation Fluid Dynamic (CFD) simulations to accurately compute the wind field along the path of a transmission line and use these results to identify the limiting section of the conductor. This work presents a case study that evaluates the coupling of CFD simulations and forecasted weather simulations using the High-Resolution Rapid Refresh (HRRR) model points over a 2-year span within a region in south eastern Idaho. The primary goal of the work is the evaluation of the number of HRRR model points used, i.e., weather stations, along the path of the line and the accuracy of the resulting ampacity. This was done using 4, 10, 17, 26, and 35 HRRR model points along two transmission line paths. The results indicate that as the number of model points are increased, the DLR ampacity of the lines decrease, yet converge as more points are added and demonstrate little change with additional HRRR points. It is expected that these results can help transmission line operators identify the number of weather stations that must be installed when coupled with CFD simulations and DLR ampacity to ensure accurate ratings and safe operations.

17 WIND ENERGY↗