Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Integrated Large-Scale Data Management Platform for Photovoltaic Power Conversion Equipment (PCE) Reliability Data

To meet the demand for accuracy and real-time capability of PV system degradation evaluation, massive volume data is needed to run high-fidelity and high-efficiency simulations and perform advanced data analysis. However, PV farm operators have a series of difficulties with PV inverter data, such as data collection from multiple channels, massive data storage, data management and massive data analysis. To address these challenges, we developed an integrated data management platform capable of data acquisition, processing, storage, query, and performing big data analysis utilizing AI algorithms. The platform can also achieve data correctness verification and provide an effective distributed data management solution to retrieve massive data and establish a connection to distributed computational frameworks.

data management platform↗

Massively scalable workflows for quantum chemistry: BigChem and ChemCloud

Electronic structure theory, i.e., quantum chemistry, is the fundamental building block for many problems in computational chemistry. Here we present a new distributed computing framework (BigChem), which allows for an efficient solution of many quantum chemistry problems in parallel. BigChem is designed to be easily composable and leverages industry-standard middleware (e.g., Celery, RabbitMQ, and Redis) for distributed approaches to large scale problems. BigChem can harness any collection of worker nodes, including ones on cloud providers (such as AWS or Azure), local clusters, or supercomputer centers (and any mixture of these). BigChem builds upon MolSSI packages, such as QCEngine to standardize the operation of numerous computational chemistry programs, demonstrated here with Psi4, xtb, geomeTRIC, and TeraChem. BigChem delivers full utilization of compute resources at scale, offers a programable canvas for designing sophisticated quantum chemistry workflows, and is fault tolerant to node failures and network disruptions. We demonstrate linear scalability of BigChem running computational chemistry workloads on up to 125 GPUs. Finally, we present ChemCloud, a web API to BigChem and successor to TeraChem Cloud. ChemCloud delivers scalable and secure access to BigChem over the Internet.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

WfBench: Automated Generation of Scientific Workflow Benchmarks

The prevalence of scientific workflows with high computational demands calls for their execution on various distributed computing platforms, including large-scale leadership-class high-performance computing (HPC) clusters. To handle the deployment, monitoring, and optimization of workflow executions, many workflow systems have been developed over the past decade. There is a need for workflow benchmarks that can be used to evaluate the performance of workflow systems on current and future software stacks and hardware platforms.We present a generator of realistic workflow benchmark specifications that can be translated into benchmark code to be executed with current workflow systems. Our approach generates workflow tasks with arbitrary performance characteristics (CPU, memory, and I/O usage) and with realistic task dependency structures based on those seen in production workflows. We present experimental results that show that our approach generates benchmarks that are representative of production workflows, and conduct a case study to demonstrate the use and usefulness of our generated benchmarks to evaluate the performance of workflow systems under different configuration scenarios.

Coleman, Taina↗

BigPanDA monitoring system evolution in the ATLAS Experiment

Monitoring services play a crucial role in the day-to-day operation of distributed computing systems. The ATLAS Experiment at LHC uses the Production and Distributed Analysis workload management system (PanDA WMS), which allows a million computational jobs to run daily at over 170 computing centers of the WLCG and opportunistic resources, utilizing 600k cores simultaneously on average. The BigPanDA monitor is an essential part of the monitoring infrastructure for the ATLAS Experiment that provides a wide range of views, from top-level summaries to a single computational job and its logs. Over the past few years of the PanDA WMS advancement in the ATLAS Experiment, several new components were developed, such as Harvester, iDDS, Data Carousel, and Global Shares. Due to its modular architecture, the BigPanDA monitor naturally grew into a platform where the relevant data from all PanDA WMS components and accompanying services are accumulated and displayed in the form of interactive charts and tables. Moreover the system has been adopted by other experiments beyond HEP. In this paper we describe the evolution of the BigPanDA monitor system, the development of new modules, and the integration process into other experiments.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Development of a Multi-Robot System for Autonomous Inspection of Nuclear Waste Tank Pits

This paper introduces the overall design plan, development timeline, and preliminary progress of the Autonomous Pit Exploration System project. This project aims to develop an advanced multi-robot system for the efficient inspection of nuclear waste-storage tank pits. The project is structured into three phases: Phase 1 involves data collection and interface definition in collaboration with Hanford Site experts and university partners, focusing on tank riser geometry and hardware solutions. Phase 2 includes the selection of sensors and robot components, detailed mechanical design, and prototyping. Phase 3 integrates all components into a cohesive system managed by a master control package which also incorporates digital twin and surrogate models, and culminates in comprehensive testing and validation at a simulated tank pit at the Idaho National Laboratory. Additionally, the system’s communication design ensures coordinated operation through shared data, power, and control signals. For transportation and deployment, an electric vehicle (EV) is chosen to support the system for a full 10 h shift with better regulatory compliance for field deployment. A telescopic arm design is selected for its simple configuration and superior reach capability and controllability. Preliminary testing utilizes an educational robot to demonstrate the feasibility of splitting computational tasks between edge and cloud computers. Successful simultaneous localization and mapping (SLAM) tasks validate our distributed computing approach. More design considerations are also discussed, including radiation hardness assurance, SLAM performance, software transferability, and digital twinning strategies.

Nuclear waste management↗

Distributed Optimal Power Management for Battery Energy Storage Systems: A Novel Accelerated Tracking ADMM Approach

Optimal power management (OPM) is critical for large-scale battery energy storage systems. Today’s methods often require formidable computational effort due to the design based on centralized numerical optimization. Thus, this paper investigates computationally distributed OPM where the agents based on the cells communicate over a network to cooperatively solve the OPM problem. We propose an accelerated tracking alternating direction method of multipliers (ADMM) algorithm to solve the distributed OPM. The proposed algorithm embeds dynamic average consensus and Nesterov’s acceleration technique in the ADMM algorithm. Not only is the proposed algorithm fully distributed without a need for fusion or aggregating nodes, but it also accelerates convergence. The paper formulates the OPM in a model predictive control framework where it seeks to regulate the charging/discharging power of each battery cell to minimize the total power losses and promote balanced use of the constituent cells while complying with the safety constraints. The paper provides ample simulation results to demonstrate the effectiveness and advantages of the proposed distributed OPM in terms of computation and convergence.

Farakhor, Amir↗

Software and computing for Run 3 of the ATLAS experiment at the LHC

The ATLAS experiment has developed extensive software and distributed computing systems for Run 3 of the LHC. These systems are described in detail, including software infrastructure and workflows, distributed data and workload management, database infrastructure, and validation. The use of these systems to prepare the data for physics analysis and assess its quality are described, along with the software tools used for data analysis itself. An outlook for the development of these projects towards Run 4 is also provided.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Scientific Data Management Beyond Traditional Computing Boundaries

Scientific data management is undergoing a fundamental transformation driven by the convergence of artificial intelligence (AI)/machine learning workflows, distributed computing and storage environments, and exponential data growth. Here, we analyze how these developments address current limitations while enabling new capabilities for cross-facility collaboration and AI-driven research.

Widener, Patrick [Oak Ridge National Laboratory (O↗

Use Case-Informed Framework for Utility Cloud Migration

This white paper presents a comprehensive methodology for assessing utilities’ cloud postures and frameworks. It aims to produce a roadmap and strategy for a cloud-enabled grid future, providing guidance for integrators, asset owners, and operators. Instead of offering a yes or no answer for cloud implementation, this framework offers strategic guidance on responsibly preparing for and deploying cloud solutions. This paper delves into cloud-service models pertinent to the electric sector, dissecting the shared responsibility model and elucidating what on-premise infrastructure as a service (IaaS), platform as a service (PaaS), and software as a service (SaaS) entail. A pivotal consideration within the context of the shared responsibility model is the allocation of responsibility for foundational security aspects—a decision that will be informed by a comprehensive risk assessment. The ensuing discussion will present a checklist of certifications necessary for a secure cloud transition, equipping utilities with the knowledge to navigate this digital transformation with confidence and with a strategic roadmap. Furthermore, the paper outlines gaps in understanding the U.S. government’s role in shaping technology development and responsible use. Its purpose is to aid decision-making by offering support for risk-informed solutions that benefit those managing assets and operating in the cloud environment. The primary objective is to enhance the resilience and future readiness of a decarbonized electric grid, with cloud solutions as one viable option. The paper synthesizes information on current and future grid architectures and applications, considering both conservative and progressive energy transitions, along with scalable and distributed computing considerations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The environmental impact, carbon emissions and sustainability of computing in the ATLAS experiment

ATLAS, a general-purpose experiment at the Large Hadron Collider (LHC), makes use of a large internationally-distributed computing infrastructure, including over 10 6 TB of managed data on disk and tape and almost one million simultaneously running CPU cores. Upgrades for the High-Luminosity LHC (HL-LHC) will increase the required computing resources by a factor of 3–4 by the beginning of the 2030s, and by an order of magnitude before the conclusion of data taking at the beginning of the 2040s. These resources are spread over around 100 computing sites worldwide. Efforts are underway within the experiment to evaluate and mitigate various aspects of the environmental impact of the sites, with the additional long-term goal of making recommendations to the sites that will significantly reduce the total expected environmental impact in the HL-LHC era. These efforts take several forms: building awareness in the experiment community, adjusting aspects of the computing policy, and modifications of data center configurations, either in ways that take advantage of particular features of ATLAS workloads or in generic ways that reduce the environmental impact of the computing resources. This paper describes the ongoing investigations and approaches that have already provided useful and actionable outcomes.

Aad, G. [CNRS/IN2P3] (ORCID:0000000266654934)↗

Numerical Investigation of Enhanced Dehumidification Processes By Using Dielectrophoresis Principles in Moist Airflows

Dispersed particle-laden flows are encountered in many building and industrial applications, such as flow in a fluidized bed, hydrocarbon transportation in pipelines, and the fouling of air-cooled heat exchangers (Kuruneru et al., 2016; Ray et al., 2019; Wang et al., 2019). Computational fluid dynamic (CFD) models have been developed in recent years to depict particle-fluid and particle-particle interactions in laminar or turbulent flows with increasing accuracy and stability. One particular particle-laden system of interest for moisture control is electrically-enhanced condensation in air and water droplet flows. Electrically-enhanced condensation consists of the use of highly charged water droplets injected in the moist air. The droplets become electric seeds that attract polar water vapor molecules to their surfaces and promote condensation. The nucleation and growth of the charged droplets deplete the vapor phase near a droplet, which is compensated for by the dielectrophoresis flow and diffusion. Dielectrophoresis flow involves surrounding vapor at a distance of about 10 to 100 nm for droplets charged by an electrospray compared to ~2 nm for a single electron charge in a droplet. As the vapor molecules collapse on the surface of the droplets, their initial electrical charge decreases with time due to the neutralization of the ions. While the physics of this phenomena is well known, engineering models for predicting the condensation rates are not available. This work computationally investigates dehumidification of moist airflow in a converging rectangular duct. The objective is to develop an engineering model that predicts water vapor condensation by employing dielectrophoresis principles. We construct a Computational Fluid Dynamics (CFD) model of the duct with electrically-enhanced condensation. The model is implemented in the open-source software OpenFOAM. We utilize the Multi-Phase Particle-In-Cell (MP-PIC) method coupled with a Population Balance Equation (PBE) approach to simulate the particle-laden system. This methodology is an Eulerian-Lagrangian approach used to simulate the droplets' behavior in the humid air. The MP-PIC approach (Andrews and O'Rourke, 1996) mitigates the computational cost by parceling several fundamental particles with similar properties (such as types, sizes, and temperature) into one computational particle. Thus, the billions of particles can be substituted by millions of computational particles without significant loss of information. The PBE was considered with the Lagrangian frame to combine the particle distribution function used in MP-PIC (Kim et al., 2020). This approach preserves mass and energy conservation between the phases in the Eulerian and Lagrangian structures. The PBE in this procedure was directly linked to the discrete parcels, making the simulation of the particle distribution computationally efficient and robust. The MP-PIC-PBE approach used in the present work was applied to the dehumidification of air. Water droplets were injected in the air stream and forced to grow according to experimentally derived correlation. The experiments were conducted on a converging duct with the same geometry and boundary conditions used to build the CFD model. This approach enabled us to approximate the effect of dielectrophoresis phenomena on the droplet and air interface. This presentation will discuss the details of the new CFD model built for the duct, the implementation of the model in OpenFOAM CFD programming language, and the experimental validation of the newly developed model. The results revealed a moderate yet measurable increase in droplet diameter due to water vapor condensation at the vapor-liquid interface of the electrically charged droplets' surface. The seed water droplet particles grew in size by capturing the water vapor in the surrounding air. The OpenFOAM model predicted reductions of humidity in the air from 5 to 10 percent.

Yel Mahi, Maliha↗

Flexible User-Defined Domain Decomposition in Kilometer-Scale E3SM Land Model Simulation

The Energy Exascale Earth System Model (E3SM) Land Model (ELM) has been extended to kilometer-scale (km-ELM) resolutions, enabling high-fidelity simulations of terrestrial processes at 1 km x 1 km grid spacing. In ELM, domain decomposition partitions the computational domain across processors, ensuring efficient parallel execution. Currently, round-robin decomposition is applied, providing a straightforward way to distribute computational workload. As ELM continues evolving at the kilometer-scale (km-scale), particularly with integrating lateral flow modeling, decomposition strategies must also account for the increased workload and data movement. This paper introduces a flexible user-defined domain decomposition framework, allowing users to customize domain partitioning based on application requirements. The impact of different decomposition strategies is evaluated across various applications concerning computation, communication, and I/O. Results demonstrate that while 1D partitioning yields superior I/O performance, k-nearest neighbors (KNN) clustering effectively reduces inter-process communication overhead. This study lays the groundwork for scalable partitioning in large-scale land surface simulations, enhancing next-generation Earth system modeling.

Wang, Dali [ORNL] (ORCID:0000000168065108)↗

Software Quality Assurance for High Performance Computing Containers

Software containers are a key channel for delivering portable and reproducible scientific software in high performance computing (HPC) environments. HPC environments are different from other types of computing environments primarily due to usage of the message passing interface (MPI) and drivers for specialized hard- ware to enable distributed computing capabilities. This distinction directly impacts how software containers are built for HPC applications and can complicate software quality assurance efforts including portability and performance. This work introduces a strategy for building containers for HPC applications that adopts layering as a mechanism for software quality assurance. The strategy is demonstrated across three different HPC systems, two of them petaflops scale with entirely different interconnect technologies and/or processor chipsets but running the same container. Performance consequences of the containerization strategy are found to be less than 5-14% while still achieving portable and reproducible containers for HPC systems.

97 MATHEMATICS AND COMPUTING↗

Cloud Services Enable Efficient AI-Guided Simulation Workflows across Heterogeneous Resources

Applications which fuse machine learning and simulation are rarely best served by a single computing resource. Highly parallel simulation codes are best deployed on super- computers, while AI tasks used to decide which simulations to perform may be best suited to specialized accelerators. Here we present a Function-as-a-Service (FaaS) system for executing complex, distributed computational campaigns that achieves performance parity with conventional workflow systems without the complexities of secure network connections between compute providers. One innovation enabling high performance is a subsystem that directly moves task data between sites, separate from the cloud-hosted FaaS system used to distribute task instructions. We also introduce a flexible scheduling system that allows us access factor of 2 trade offs between the amount of resources required to solve a problem at each compute site. We anticipate that this system will upgrade multi-site applications from demonstration projects to routine practice in computational science.

Ward, Logan↗

Explainable Neural Architecture Search (XNAS)

Code for the paper Learning Interpretable Models Through Multi-Objective Neural Architecture Search by Zachariah Carmichael, Tim Moon, and Sam Ade Jacobs. Monumental advances in deep learning have led to unprecedented achievements across a multitude of domains. While the performance of deep neural networks is indubitable, the architectural design and interpretability of such models are nontrivial. Research has been introduced to automate the design of neural network architectures through neural architecture search (NAS). Recent progress has made these methods more pragmatic by exploiting distributed computation and novel optimization algorithms. However, there is little work in optimizing architectures for interpretability. To this end, we propose a multiobjective distributed NAS framework that optimizes for both task performance and introspection. We leverage the non-dominated sorting genetic algorithm (NSGA-II) and explainable AI (XAI) techniques to reward architectures that can be better comprehended by humans. The framework is evaluated on several image classification datasets. We demonstrate that jointly optimizing for introspection ability and task error leads to more disentangled architectures that perform within tolerable error.

Carmichael, ZachariahJ↗

Extreme-scale stochastic optimization and simulation via learning-enhanced decomposition and parallelization (Final Technical Report)

Stochastic optimization and simulation models ubiquitously arise in designing and operating complex service/engineering systems. They can be extreme in scale due to high-dimensional data and decisions, and can also involve decisions made sequentially in response to newly revealed data, both causing significant computational challenge. The objective of this research is to explore a unified framework that integrates machine learning with discrete optimization and risk-averse modeling, to improve the efficiency of decomposition paradigms for stochastic optimization and simulations at extreme scale. The models we consider represent a broad class of complex decision-making problems, where 0-1 or continuous decisions are made before and/or after knowing multiple sources of uncertainties that could be correlated. We will employ machine learning methods to dynamically decide and prioritize computational procedures, including cut generation, branching, and bounding of the optimal objective. Furthermore, the research will shed new lights on the traditional decomposition algorithms for extreme-scale computing. Deliverables of the research include new modeling and computational methods for advancing the state-of-the-art research in optimization and simulation, bringing many relevant risk-averse, data-driven optimization problems in practice within the range of tractability. Examples include distributed computing server scheduling and sensor deployment for monitoring critical infrastructures. Success in this effort will enable progress in solving multiple extreme-scale problems in the complex system design and operations arising from DoE missions in energy, environment, and national security.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Characterization of ECRAM materials and devices

As the limits of Moore’s Law approaches, new computing paradigms are developing to breakthrough this bottleneck. One such computer architecture is neuromorphic computing, which models the brain. Electrochemical random-access memory (ECRAM) is a low power and energy efficient memory due to characteristics, such as in-memory compute, ion modulation of the channel conductance, and computation distribution with large scale array integration.

97 MATHEMATICS AND COMPUTING↗

Scalable training of graph convolutional neural networks for fast and accurate predictions of HOMO-LUMO gap in molecules

Abstract Graph Convolutional Neural Network (GCNN) is a popular class of deep learning (DL) models in material science to predict material properties from the graph representation of molecular structures. Training an accurate and comprehensive GCNN surrogate for molecular design requires large-scale graph datasets and is usually a time-consuming process. Recent advances in GPUs and distributed computing open a path to reduce the computational cost for GCNN training effectively. However, efficient utilization of high performance computing (HPC) resources for training requires simultaneously optimizing large-scale data management and scalable stochastic batched optimization techniques. In this work, we focus on building GCNN models on HPC systems to predict material properties of millions of molecules. We use HydraGNN, our in-house library for large-scale GCNN training, leveraging distributed data parallelism in PyTorch. We use ADIOS, a high-performance data management framework for efficient storage and reading of large molecular graph data. We perform parallel training on two open-source large-scale graph datasets to build a GCNN predictor for an important quantum property known as the HOMO-LUMO gap. We measure the scalability, accuracy, and convergence of our approach on two DOE supercomputers: the Summit supercomputer at the Oak Ridge Leadership Computing Facility (OLCF) and the Perlmutter system at the National Energy Research Scientific Computing Center (NERSC). We present our experimental results with HydraGNN showing (i) reduction of data loading time up to 4.2 times compared with a conventional method and (ii) linear scaling performance for training up to 1024 GPUs on both Summit and Perlmutter.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗