Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

UPC++ v1.0 Specification, Revision 2021.9.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Drop C: The Drop-In, Ring-of-Power Heliostat (Final Report)

This report summarizes the work performed within the Drop-C: The Drop-In, Ring-of-Power Heliostat project. The Drop-C project aimed to develop a novel heliostat with an installed cost of $\$50$/m 2 ($\$2015$) which is a drastic cost reduction compared to the state-of-the-art. The resulting 27m 2 SunRing TM heliostat’s relative small size necessitates the parallel development of a wireless solar field network and a rapid calibration system. The SunRing heliostat is an evolution of the Ring-of-Power (ROP) design from Abengoa Solar LLC which had a reported installed cost of $114/m 2 . Cost savings relative to the ROP were sought by increasing the mirror area, improving the structural efficiency, a more accurate assessment of wind loads, and an improved assembly and installation procedure. The project spanned three Budget Periods (BP) where a digital heliostat was developed in BP1 using wind tunnel testing to define load conditions.

14 SOLAR ENERGY↗

UPC++ v1.0 Specification (Rev. 2023.9.0)

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification (Revision 2022.3.0)

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2023.3.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification (Revision 2021.3.0)

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2022.9.0

UPC++ is a C++ library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

UPC++ v1.0 Specification, Revision 2020.3.0

UPC++ is a C++11 library providing classes and functions that support Partitioned Global Address Space (PGAS) programming. The key communication facilities in UPC++ are one-sided Remote Memory Access (RMA) and Remote Procedure Call (RPC). All communication operations are syntactically explicit and default to non-blocking; asynchrony is managed through the use of futures, promises and continuation callbacks, enabling the programmer to construct a graph of operations to execute asynchronously as high-latency dependencies are satisfied. A global pointer abstraction provides system-wide addressability of shared memory, including host and accelerator memories. The parallelism model is primarily process-based, but the interface is thread-safe and designed to allow efficient and expressive use in multi-threaded applications. The interface is designed for extreme scalability throughout, and deliberately avoids design features that could inhibit scalability.

97 MATHEMATICS AND COMPUTING↗

Synchronous High-frequency Distributed Readout For Edge Processing At The Fermilab Main Injector And Recycler

The Main Injector (MI) was commissioned using data acquisition systems developed for the Fermilab Main Ring in the 1980s. New VME-based instrumentation was commissioned in 2006 for beam loss monitors (BLM)[2], which provided a more systematic study of the machine and improved displays of routine operation. However, current projects are demanding more data and at a faster rate from this aging hardware. One such project, Real-time Edge AI for Distributed Systems (READS), requires the high-frequency, low-latency collection of synchronized BLM readings from around the approximately two-mile accelerator complex. Significant work has been done to develop new hardware to monitor the VME backplane and broadcast BLM measurements over Ethernet, while not disrupting the existing operations critical functions of the BLM system. This paper will detail the design, implementation, and testing of this parallel data pathway.

43 PARTICLE ACCELERATORS↗

Airport Infrastructure Planning Using Multi-Stage Stochastic Programming

The Athena project, funded by the Department of Energy, has worked to identify the critical infrastructure at Dallas Fort Worth (DFW) Airport which influences mobility between the airport and the surrounding city of Dallas. Using scalable methods that can leverage HPC resources we have developed a multi-stage stochastic infrastructure expansion model for determining parking and curb modifications to the DFW Airport over a 20-year horizon. Additionally, we have explored the impacts of congestion pricing in conjunction with infrastructure modifications. Our multi-stage stochastic model is implemented using the mpi-sppy software and solved in parallel using progressive hedging on the National Renewable Energy Laboratory's HPC system Eagle. In this talk we present results from solving this model at scale.

airport planning↗

Advanced Shuttle Strategies for Parallel QCCD Architectures

Trapped ions (TIs) are at the forefront of quantum computing implementation, offering unparalleled coherence, fidelity, and connectivity. However, the scalability of TI systems is hampered by the limited capacity of individual ion traps, necessitating intricate ion shuttling for advanced computational tasks. The quantum charge-coupled device (QCCD) framework has emerged as a promising solution, facilitating ion mobility for universal quantum computation. Current QCCD architectures predominantly feature a linear topology, which is increasingly recognized as inefficient for complex quantum operations. Anticipating the shift toward more efficacious designs, this article introduces an innovative quantum scheduling strategy optimized for parallel QCCD topologies. Our strategy proposes a probabilistic formula for ion movement, alongside ingenious methods for local layer generation and layer compression, yielding a significant reduction in ion shuttle times. Through simulations, we demonstrate that our strategy not only substantially outstrips the linear model but also exhibits better performance over other parallel strategies that employ greedy algorithms. This is achieved through our nuanced resolution of complexities, such as traffic blocks and trap capacity limitations. The consequent reduction in shuttle operations leads to lower energy consumption and an enhancement in the quantum computer's fidelity, ultimately accelerating program execution times.

43 PARTICLE ACCELERATORS↗

Parallel Programming in MCNP6

Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.

97 MATHEMATICS AND COMPUTING↗

Exploring temporal community evolution: algorithmic approaches and parallel optimization for dynamic community detection

Abstract Dynamic (temporal) graphs are a convenient mathematical abstraction for many practical complex systems including social contacts, business transactions, and computer communications. Community discovery is an extensively used graph analysis kernel with rich literature for static graphs. However, community discovery in a dynamic setting is challenging for two specific reasons. Firstly, the notion of temporal community lacks a widely accepted formalization, and only limited work exists on understanding how communities emerge over time. Secondly, the added temporal dimension along with the sheer size of modern graph data necessitates new scalable algorithms. In this paper, we investigate how communities evolve over time based on several graph metrics under a temporal formalization. We compare six different algorithmic approaches for dynamic community detection for their quality and runtime. We identify that a vertex-centric (local) optimization method works as efficiently as the classical modularity-based methods. To its advantage, such local computation allows for the efficient design of parallel algorithms without incurring a significant parallel overhead. Based on this insight, we design a shared-memory parallel algorithm DyComPar , which demonstrates between 4 and 18 fold speed-up on a multi-core machine with 20 threads, for several real-world and synthetic graphs from different domains.

97 MATHEMATICS AND COMPUTING↗

Plastic Parallel Pathways Platform - 4P Model

Global momentum is building towards a circular economy capable of keeping plastics in use and out of waste streams. Given that 79% of all plastic produced since 1950 has accumulated in landfills or the natural environment,rapid implementation of various end-of-life (EoL) management technologies will be needed to reach this target. However, it can be challenging to develop an effective plastic EoL strategy when the available options - chemical or molecular recycling, energy recovery, upcycling, downcycling, closed-loop (plastic-to-plastic) or open-loop (plastic-to-x) recycling, among others - can generate products ranging from low-grade to virgin-quality plastic and from fuels to value-added chemicals. We present a flexible material flow model capable of analyzing the effects of both plastic-to-plastic and plastic-to-x EoL management strategies on the U.S. PET economy. This Plastic Parallel Pathways Platform (4P) assesses the environmental impacts, costs, and circularity of a PET system in which waste is managed through six potential EoL pathways: landfill, incineration with energy recovery, pyrolysis to fuel oil, upcycling to glass fiber reinforced plastic (GFRP), mechanical recycling to low-grade PET, and chemical recycling (glycolysis) to bottle-grade PET. We compare the pathways across multiple metrics using multi-criteria decision analysis (MCDA) and then use a brute force algorithm to predict an optimal combination of EoL pathways to minimize greenhouse gas (GHG) emissions and costs and maximize circularity. This work highlights the need to implement a diverse portfolio of EoL strategies in parallel to enable a PET economy that meets environmental, economic, and circularity requirements simultaneously.

downcycling↗

Hybrid Solar System (Final Scientific/Technical Report)

GTI Energy (GTI) teamed with the University of California at Merced (UCM) to scaleup the hybrid solar system (HSS) technology for demonstrating its performance at the US Gypsum (USG) plant in Plaster City, California. The technology integrates two-stage concentrating solar collector with matching particle thermal transport and storage (TSS) system to deliver cost-effective, and on-demand distributed high temperature industrial process heat up to 600°C with solar thermal, in this case to a gypsum kettle, to reduce its fuel use and carbon footprint. Current solar technologies, which reach these temperatures, are not distributable (towers) or cost-effective (dish). The research team developed a conceptual system design for host site retrofit, including preliminary heat balance, process flow diagram, particle to process heat exchanger and equipment placements at the site. Subsequently, parallel efforts were carried out at UCM to design, build and test a 12 m long commercial scale prototype concentrating thermal-only collector system and at GTI to design, build and test a matching 650°C capable particle TTS system. The nominal 50 kWth collector consists of a parabolic trough and three 4 m long two-stage receivers in series. Prior to on-sun testing, a 4 m long receiver was fabricated and successfully tested at 650 °C in a laboratory setting for 100 hrs of continuous operation showing less than 15% radiation loss. A 7 m wide x 17 m long parabolic trough was then installed at UCM for on-sun testing of the 12 m long receiver, and concurrently several 4 m long receivers were built. The optics of the parabolic trough were calibrated, and on-sun test were carried out on 12 m long receivers. During tests, the intense solar radiation (53x) caused the absorber tubes in the receivers to bend, reducing the overall optical efficiency. To address the bending issue, a self-consistent algorithm that includes ray tracing, thermal and deformation models was developed to perform thermal stress analysis on absorbers for parabolic solar collectors. Results obtained with this algorithm showed a dramatic rise in deformation as absorber tube length increases. A combined efficiency parameter that includes the occluded area for the mounts was developed to obtain an optimized tube length obtained. Based on the results, a length of 2.7 m for the absorber + 0.2 m for the coupler was chosen to minimize any bending and optimize optical efficiency while maintaining ease of mounting. The associated particle TTS system was designed, built and successfully tested at GTI. It includes storage, receiving and lock hoppers and piping that simulates the transfer of captured solar energy to an actual industrial furnace. Tests over 77 charge-discharge cycles demonstrated <2% particle degradation, with no problematic particle accumulations and no flow interruptions. The piping pressure drop was about 5 psi. The team also worked with Stanley Consultants (Stanley) to prepare conceptual and preliminary engineering packages to facilitate follow-on development and commercialization efforts. These include process and instrumentation diagram’s (P&ID’s), general arrangements, electrical one-line, project definitions document, equipment data sheets, schedule, and construction cost estimate for 2 MWth system. Updated HSS technology commercialization and customer engagement plans and detailed costs and evaluated market trade-offs and manufacturing.

03 NATURAL GAS↗

A quantum processor based on coherent transport of entangled atom arrays

The ability to engineer parallel, programmable operations between desired qubits within a quantum processor is key for building scalable quantum information systems. In most state-of-the-art approaches, qubits interact locally, constrained by the connectivity associated with their fixed spatial layout. Here we demonstrate a quantum processor with dynamic, non-local connectivity, in which entangled qubits are coherently transported in a highly parallel manner across two spatial dimensions, between layers of single- and two-qubit operations. Our approach makes use of neutral atom arrays trapped and transported by optical tweezers; hyperfine states are used for robust quantum information storage, and excitation into Rydberg states is used for entanglement generation. We use this architecture to realize programmable generation of entangled graph states, such as cluster states and a seven-qubit Steane code state. Furthermore, we shuttle entangled ancilla arrays to realize a surface code state with thirteen data and six ancillary qubits and a toric code state on a torus with sixteen data and eight ancillary qubits. Finally, we use this architecture to realize a hybrid analogue–digital evolution and use it for measuring entanglement entropy in quantum simulations, experimentally observing non-monotonic entanglement dynamics associated with quantum many-body scars. Realizing a long-standing goal, these results provide a route towards scalable quantum processing and enable applications ranging from simulation to metrology.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Arrested coarsening and large density fluctuations in driven particle mixtures in two dimensions

Abstract Using molecular dynamics simulations, we study a driven, nonadditive binary mixture of spherical particles confined to move in two dimensions and immersed in an explicit solvent consisting of point particles with purely repulsive interactions. We show that, without a drive, the mixture of spherical particles phase separates and coarsens with kinetics consistent with an Ising-like conserved dynamics. Conversely, when the drive is applied, the coarsening is arrested and the system develops large density fluctuations. We show that the drive creates domains of a characteristic size which decreases with an increasing force. Furthermore, we find that these domains are anisotropic and can be oriented either parallel or perpendicular to the drive direction. Finally, we connect our findings to existing theories of strongly-driven systems, pointing out the importance of introducing the explicit solvent particles to break the Galilean invariance of the system.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

HOSS!

The Hall-D Online Skim System (HOSS) was developed to simultaneously solve two issues for the high intensity GlueX experiment. One was to parallelize the writing of raw data files to disk in order to improve bandwidth. The other was to distribute the raw data across multiple compute nodes in order to produce calibration skims of the data online. The highly configurable system employs RDMA, RAM disks, and zeroMQ driven by Python to simultaneously store and process the full high intensity GlueX data stream.

Lawrence, David↗