Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “avoiding communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

FPV Video Adaptation for UAV Collision Avoidance

First person view (FPV) technology for unmanned aerial vehicles (UAVs) provides an immersive experience for pilots and enables various personal and commercial applications such as aerial photography, drone racing, search and rescue operations, agricultural surveillance, and structural inspection. While real time video streaming from a UAV and vision-based collision avoidance strategies have been studied in literature as separate topics, in this paper we tackle collision avoidance in FPV scenarios, taking into account network delays and real time video parameters. We present a theoretical model for obstacle collisions that considers the current communication channel conditions, the real time video parameters, and the UAV's position relative to the closest obstacle. A video adaptation algorithm is then designed, using this metric, to tune the FPV video resolution, number of re-transmission attempts, and the modulation scheme to maximize the probability of avoiding collisions. This algorithm also takes into account specific latency constraints of the application. This video algorithm was evaluated in various scenarios and its ability to respond to both distances to the obstacle as well as the communication channel conditions was demonstrated. It was found that, for the considered scenarios, the performance of the proposed adaptive algorithm was, on an average, 58.63% higher than the closest non-adaptive one in terms of maximizing the probability of avoiding collision. Such collision avoidance strategies could be used to make UAV FPV applications safer and more reliable.

47 OTHER INSTRUMENTATION↗

Progress on Demonstration of a MOOSE-Based Coupled Capability for Hot Channel Factors in Fast Reactors

Hot channel factors (HCFs) are computed values that account for the impact on predicted peak fuel, cladding, and coolant temperatures due to uncertainties in the as-built reactor’s material properties and geometry as well as uncertainties due to modeling approximations. Reduction in computed HCF values via reduction or elimination of modeling approximations may translate to significant economic savings if the reactor power can be raised due to the extra temperature margin gained. While limited historical datasets exist for sodium-cooled fast reactors (SFRs), there are no available HCF data for lead-cooled fast reactors (LFRs) outside of work generated previously within NEAMS. The computation of HCFs involves insights from reactor physics, thermal fluids and heat conduction calculations to determine how the peak temperatures respond to various uncertainties in the design. Due to the significant advantages for multi-physics coupling offered by the MOOSE framework, Griffin (MOOSE-based reactor physics code), MOOSE Heat Conduction Module, and Cardinal (MOOSE-wrapped multi-physics application which includes the NekRS thermal fluids code) are being coupled together using the MOOSE MultiApp System to develop a highfidelity multi-physics modeling capability for HCF simulations. This high-fidelity coupling workflow may also be beneficial for other fast reactor applications in the future. In previous work, Griffin and NekRS were individually assessed to ensure the necessary capabilities were in place. This work describes initial efforts to couple the codes (including folding in the MOOSE Heat Conduction Module) and determining the workflow for the perturbed calculations which will leverage the Stochastic Tools Module (STM). To our knowledge, this is the first coupling of Griffin and NekRS as well as the first exploratory use of Stochastic Tools Module for Cardinal. In this report, the neutronics code Griffin, the heat conduction solver in MOOSE, and the MOOSE-wrapped application containing NekRS (Cardinal) are linked together to demonstrate the coupled capability. Griffin and Cardinal are linked dynamically by specifying shared libraries. Different coupling hierarchies are tested for selecting the most appropriate coupling strategy. A coupling scheme is selected based on the efficiency of calculation and ease of data communication. Multiple tests are performed to choose suitable mesh structure, model configurations, scheme setup and boundary conditions to avoid loss of energy due to data interpolation between different modules or weak imposition of fluxes in finite element codes. Computational experiments are performed to study the tolerance control of each type of iteration to avoid false convergence. The coupled capability is demonstrated in both single pin and 7-pin models based on LFR materials and geometry. The study finds that the use of too large a time step size in the heat conduction module can lead to temperature oscillation even though the heat conduction equation does not have a time-derivative kernel, but only the time-dependent boundary condition. A 7-pin model without duct region achieved good convergence in the coupled calculation while a 7 pin model with duct region experienced data communication issues which need to be resolved.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

GPU-enabled extreme-scale turbulence simulations: Fourier pseudo-spectral algorithms at the exascale using OpenMP offloading

Fourier pseudo-spectral methods for nonlinear partial differential equations are of wide interest in many areas of advanced computational science, including direct numerical simulation of three-dimensional (3-D) turbulence governed by the Navier-Stokes equations in fluid dynamics. This paper presents a new capability for simulating turbulence at a new record resolution up to 35 trillion grid points, on the world's first exascale computer, Frontier, comprising AMD MI250x GPUs with HPE's Slingshot interconnect and operated by the US Department of Energy's Oak Ridge Leadership Computing Facility (OLCF). Key programming strategies designed to take maximum advantage of the machine architecture involve performing almost all computations on the GPU which has the same memory capacity as the CPU, performing all-to-all communication among sets of parallel processes directly on the GPU, and targeting GPUs efficiently using OpenMP offloading for intensive number-crunching including 1-D Fast Fourier Transforms (FFT) performed using AMD ROCm library calls. With 99% of computing power on Frontier being on the GPU, leaving the CPU idle leads to a net performance gain via avoiding the overhead of data movement between host and device except when needed for some I/O purposes. Memory footprint including the size of communication buffers for MPI_ALLTOALL is managed carefully to maximize the largest problem size possible for a given node count. Detailed performance data including separate contributions from different categories of operations to the elapsed wall time per step are reported for five grid resolutions, from 2048 3 on a single node to 32768 3 on 4096 or 8192 nodes out of 9408 on the system. Both 1D and 2D domain decompositions which divide a 3D periodic domain into slabs and pencils respectively are implemented. The present code suite (labeled by the acronym GESTS, GPUs for Extreme Scale Turbulence Simulations) achieves a figure of merit (in grid points per second) exceeding goals set in the Center for Accelerated Application Readiness (CAAR) program for Frontier. The performance attained is highly favorable in both weak scaling and strong scaling, with notable departures only for 2048 3 where communication is entirely intra-node, and for 32768 3 , where a challenge due to small message sizes does arise. Communication performance is addressed further using a lightweight test code that performs all-to-all communication in a manner matching the full turbulence simulation code. Performance at large problem sizes is affected by both small message size due to high node counts as well as dragonfly network topology features on the machine, but is consistent with official expectations of sustained performance on Frontier. Overall, although not perfect, the scalability achieved at the extreme problem size of 32768 3 (and up to 8192 nodes — which corresponds to hardware rated at just under 1 exaflop/sec of theoretical peak computational performance) is arguably better than the scalability observed using prior state-of-the-art algorithms on Frontier's predecessor machine (Summit) at OLCF. New science results for the study of intermittency in turbulence enabled by this code and its extensions are to be reported separately in the near future.

3D fast Fourier transform↗

Cy-Phy ADS: Cyber Physical Anomaly Detection Framework for EV Charging Systems

Today’s large-scale Electric Vehicle (EV) infrastructures are heavily dependent on information communication technologies to maintain their operation and to support communication within sub-system components as well as the outside world. These technologies are vulnerable to various cyber and physical threats. Timely identification and mitigation of these threats are critical for improving human safety, avoiding economic losses, and preventing catastrophic system failures. By addressing this, our work presents a ResNet Autoencoder (AE) based Cyber-Physical Anomaly Detection System (Cy-Phy ADS) for detecting anomalies in EV Controller Area Network (CAN) protocol communication. It consists of four main components: Cyber-Physical Feature Extractor, ResNet AE-based Anomaly Detection Framework, Cyber-Physical Health Metric (CPHM), and Visualization Dashboard. The presented framework was trained and tested using CAN data collected from the EV charging system testbed at the Idaho National Laboratory. The presented Cy-Phy ADS compared against six widely used unsupervised anomaly detection algorithms: One Class Support Vector Machine (OCSVM), Variational Autoencoder (VAE), LSTM Autoencoder (LSTM AE), Isolation Forest (IForest), Principle Component Analysis (PCA) and Local Outlier Factor (LOF). Here the presented approach showed the highest accuracy among the compared methods. Further, the proposed approach showed comparable performance in terms of precision, F1, and False positive rate. It also showed the lowest training and inference time compared to the neural network-based baseline algorithms compared against with. Additionally, the Cy-Phy ADS has advantages such as unsupervised training, the ability to provide a holistic metric for system health characterization, and non-linear feature extraction.

99 GENERAL AND MISCELLANEOUS↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

Demonstration of a Novel Technology to Manage Electricity Demand in Grid-Independent Military Microgrids

This research was conducted by the National Renewable Energy Laboratory (NREL) in collaboration with the S&C Electric Inc. through funding provided by the ESTCP. The project demonstrates use of cybersecure Automated Demand Response (ADR) technology to effectively manage microgrid loads during grid-independent, also known as "islanded," operation. When military microgrids become isolated from the main electrical grid, they are required to balance electricity supply and demand locally. Given that local generation may be constrained, the prevailing strategy involves shedding all but the most critical loads by tripping smart circuit breakers, which then necessitate manual resetting. This approach is generally implemented at the building level, which means that the buildings with mission-critical activities are exempt from load management and remain fully powered, whereas those deemed non-critical can experience a complete loss of service. In this research we developed a method that allows building automation systems to selectively control their assets in response to load shedding request from a microgrid controller, avoiding total loss of service in contrast to the conventional control approach. A commercial OpenADR client server by GridFabric is used for communication between the microgrid controller and the building management system (BMS). The microgrid controller monitors both generation capacity and various assets within the microgrid and issues a demand reduction request when necessary. This request is communicated to the OpenADR server via Modbus. Upon receiving the request, the OpenADR server forwards it to the BMS utilizing the OpenADR protocol. The BMS is pre-configured with various levels of load reduction strategies based on the controllable assets available, allowing for a nuanced approach to demand reduction. Both lab and field tests were performed that considered load shedding needed to achieve closed transition into island mode and to accommodate changing loads and power source availability while islanded. A commercial microgrid controller was used for these tests with normal programming within the expected constraints of the system capabilities. That is, the solution did not require any specialized modification to the code base of the controller. Given the latency of the round-trip communication path between the microgrid controller and the various devices involved with the load shed processes, there are certain scenarios for which the demonstrated solution are appropriate and some which are not. The methods described in this report can be used for load shedding/restoration during transitions between islanded and grid-tied modes of operation, as well as accommodating normal variations in load and the need to remove a power source from operation for maintenance. These methods should not be used for scenarios that require load shedding within a second or two such as sudden and unanticipated significant load increases or loss of power sources through equipment faults.

24 POWER TRANSMISSION AND DISTRIBUTION↗

GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism

Graph neural networks (GNNs), an emerging class of machine learning models for graphs, have gained popularity for their superior performance in various graph analytical tasks. Mini-batch training is commonly used to train GNNs on large graphs, and data parallelism is the standard approach to scale mini-batch training across multiple GPUs. Data parallel approaches contain redundant work as subgraphs sampled by different GPUs contain significant overlap. To address this issue, we introduce a hybrid parallel mini-batch training paradigm called Split parallelism. Split parallelism avoids redundant work by splitting the sampling, loading, and training of each mini-batch across multiple GPUs. Split parallelism, however, introduces communication overheads that can be more than the savings from removing redundant work. We further present a lightweight partitioning algorithm that probabilistically minimizes these overheads. We implement spllit parllelism in GSplit and show that it outperforms state-of-the-art mini-batch training systems like DGL, Quiver, and P3.

Lim, Seung-Hwan [ORNL] (ORCID:0000000194616866)↗

InAs Terahertz Metalens Emitter for Focused Terahertz Beam Generation

Metasurfaces have opened doors to combining multiple photonic functionalities in a single compact device. In particular, the ability to generate short terahertz (THz) pulses with precise wavefront engineering in a single THz metasurface redefined the role metasurfaces can play in THz systems. Here, an InAs metalens emitter which generates and focuses a THz pulse beam is demonstrated using a 130 nm thick InAs metasurface designed as a binary‐phase Fresnel zone plate. The THz beam is focused to a spot of ≈430 μm at 1 THz with a short focal length of 5 mm and large numerical aperture of 0.5. Nanoscale InAs Mie resonators comprising the metasurface enable THz generation with an amplitude as high as 20 times compared to plasmonic THz emitters and several times compared to a 1 mm thick ZnTe crystal. This InAs metasurface emitter provides a new paradigm for designing THz imaging, spectroscopy, and communication systems, where THz beam generation and shaping are performed with a single device without compromising the generation efficiency, while eliminating losses and avoiding limitations of phase matching of conventional nonlinear optics approaches.

47 OTHER INSTRUMENTATION↗

Chloroplast Stress Signals: Control of Retrograde Signaling, Chloroplast Turn-Over, and Cell Fate Decisions

Chloroplasts (photosynthetic plastids) are semiautonomous organelles that contain their own small genomes. The proteomes of chloroplasts, however, are a mixture of plastid and nuclear-encoded proteins. Chloroplasts perform photosynthesis, which is prone to damaging the organelles, leading to the production of reactive oxygen species (ROS) that damage the cell under environmental stresses. Thus, for the cell to maintain proper chloroplast function, efficient photosynthesis, and avoid ROS damage, it relies on complex crosstalk between the chloroplast, the nucleus, other organelles within the cell, and the cytoplasm in between. This communication involves retrograde signals from chloroplasts to control nuclear gene expression, programmed cell death (PCD), and chloroplast degradation. Here we review these signals with an emphasis on the roles of the ROS singlet oxygen ( 1 O 2 ) and plastid gene expression. We cover (1) recent work on understanding how multiple 1 O 2 signaling pathways can be initiated within stressed chloroplasts, (2) how individualized post-translational regulatory systems allow chloroplasts to control their proteomes and degradation, and (3) how chloroplast signals ultimately control cell fate decisions, such as PCD, senescence, and vacuole-mediated degradation of chloroplasts (chloroplast quality control). Overall, this chapter discusses how chloroplasts can act as environmental sensors for the cell and allow plants to acclimate to stress and thrive in dynamic environments.

59 BASIC BIOLOGICAL SCIENCES↗

A Data-Driven Algorithm for Enabling Delay Tolerance in Resilient Microgrid Controls Using Dynamic Mode Decomposition

The increased implementation of smart grid technologies in the power distribution grid presents unique opportunities that enable resiliency, but also brings challenges motivating needs for novel solutions and mitigation techniques. The bi-directional power and data flow allow for the grid to operate with increased resiliency, which is the ability to avoid discontinuity of service to end-use loads during extreme events. However, in applications where control of the distribution grid or microgrid relies on communication networks, the degradation of communication systems in the form of loss or high latency can cause maloperation and result in loss of end-use loads. Here this paper presents a novel framework to enable delay tolerance of centralized microgrid control schemes to mitigate communication system latency impacts and guarantee successful control action. We demonstrate the delay tolerance on a control scheme that operates a battery energy storage system (BESS) to offset the sudden loss of generation and maintain system frequency. During periods of severely degraded communication system performance, the proposed delay-tolerant algorithm compensates for the latency by utilizing a data-driven model generated at the device level using dynamic mode decomposition (DMD) to determine the performance of the communications. The DMD technique predicts the system’s frequency using device-level terminal measurements and provides updated control signals. The HELICS cosimulation platform evaluates the cyber-physical interaction of the power system model in GridLAB-D, the centralized control agent in Python, and the discrete network model in NS-3. The framework is tested and validated on the IEEE-123 node system modified to represent a networked remote microgrid model, and the results show an improvement in the dynamic performance

24 POWER TRANSMISSION AND DISTRIBUTION↗

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗

SEAS Communication Engine: An Extensible, Flexible Wrapper for Co-Simulation Agents

When modeling and analyzing the power grid and other large scale systems, researchers often express scenarios as optimization problems and feed them into advanced software solvers. In order to allow multiple solvers to communicate with each other and share data from different domains, the National Renewable Energy Laboratory (NREL) and associated Department of Energy (DOE) labs have developed a software framework called the Hierarchical Engine for Large-scale Infrastructure Co-Simulation (HELICS). HELICS allows cosimulation via a collection of client libraries for different languages that can be called from the appropriate optimization software. However, these client libraries do not provide a higher level of abstraction beyond reading and writing data off of the shared HELICS bus. In this paper, we describe a new software library called the SEAS Communication Engine that exposes a higher-level API for running cosimulation problems. The SEAS Engine provides a class-based abstraction on top of the Python HELICS client, in order to allow users to implement their domain-specific cosimulations without needing to interact with core HELICS primitives. This will make adoption of HELICS and cosimulation in general easier, by exposing a simpler API. In the second part of the paper, we validate our library on a collection of different simulation examples, including the canonical IEEE 13 Bus Feeder. Lastly, we demonstrate using the SEAS Engine to directly call domain-specific code written in the Julia programming language. Our hope is that this will serve as a template for easily calling software in different programming languages via the SEAS Engine, thereby avoiding code duplication and complexity.

co-simulation↗

Robotic Assisted Non-Destructive Testing (NDT)

Mission Statement: Concrete wall characterization for structural integrity evaluation of H-Canyon exhaust tunnel using a remote controlled robotic arm with a NDT Concrete Instrument. Challenges: Rough and Curved Surfaces, Remote location. The UR5 is a collaborative robotic arm capable of: Payload: 11 lbs (5 kg), Reach: 33.5 in (850 mm), Footprint: 5.87 in (149 mm) diameter, Weight: 45.4 Ibs (20.6 kg). The force torque sensor reacts to a set value inputted by the user that can be utilized for sensitive products. This allows the UR5 to react to surfaces using the built-in function to search for a wall and orient itself normal to that plane. The Proceq Pundit 250 array is a nondestructive device that the user applies against a concrete wall and scans using ultrasonic transducers. The back wall and other defects can be show through the touchscreen. A fixture to hold the Proceq Pundit 250 array onto the UR5 robotic arm was designed and 3D printed. It includes openings for easy access to the buttons from the Pundit array as well as a reinforced structural design. The scan from the Pundit array can show the back-wall, rebars, and other defects to assist in determining the structural integrity. In order to scan, a program was made to search for the wall first. 1. The UR5 robotic arm will go to the desired location away from the wall. 2. It will then search for the wall by moving towards it slowly. 3. After contact, the force applied will slowly but steadily increase to a preset value inputted by the user. 4. The force torque sensor will use those values and communicate with the UR5 by adjusting the arm to become normal to the wall. 5. The arm will then apply force to push back the transducers on the Pundit array to avoid gaps. 6. It will then prompt the user to collect data and wait until finished. The UR5 teach pendant allows the user to program the robot to perform automated and repeatable tasks. It includes a free drive mode which allows the user to move the arm to a desired position by hand. The user can also apply restricted planes which include either stopped or slower motion to ensure safety. The concrete sample being tested includes rebars and different grades of surface roughness to understand the readings from the Proceq Pundit 250 array. Desired concrete samples are currently being fabricated that will simulate the terrain being tested.

3D PRINTING↗

Lightfall v0.0.1

Lightfall is a desktop application for synchrotron beamline instrument control, data acquisition, and live analysis at the Advanced Light Source (ALS). Built on Python and Qt, it provides a native graphical interface for operating beamline hardware, configuring and executing experimental scans, and visualizing results in real time. Key features include direct integration with EPICS control systems, a built-in electronic logbook, remote beamline access over secure tunnels, and an interprocess communication (IPC) architecture that coordinates with external analysis applications via ZMQ and EPICS process variables. This IPC approach allows Lightfall to orchestrate specialized analysis tools—including GPU-accelerated streaming correlators—without embedding them, avoiding the dependency conflicts common in monolithic scientific software platforms. Compared to prior approaches such as Xi-CAM's plugin-based architecture, Lightfall's design cleanly separates instrument control from domain-specific analysis, enabling feedback-driven acquisition where live analysis results can adjust scan parameters during an experiment. Its native Qt interface provides responsive performance for real-time data visualization that web-based alternatives struggle to match. Lightfall is designed for use by beamline scientists and staff operating synchrotron instruments at national user facilities.

Pandolfi, Ronald [Lawrence Berkeley National Labor↗

Integrating a Microgrid Controller with a Local OpenADR Server

When military microgrids isolate themselves from the main electrical grid, they must locally balance electricity supply and demand. Since local generation may be limited, the current strategy is to shed all but the most critical loads by tripping smart circuit breakers, which must then be reset manually (e.g., ESTCP project EW-201350). This strategy is typically applied at the building level, meaning that entire buildings housing mission critical activities must be excluded from any load management, while those considered non-critical may lose service entirely. The remotely controlled switchgear needed to manage load in this way is very expensive ($\$30,000$-$\$50,000$ per building). While effective at shedding load, this strategy disrupts installation operation and risks damaging equipment during both disconnection and re-energization. With the goals of lowering costs, protecting equipment, and enhancing the agility of DoD microgrids, this report demonstrates the use of cybersecure automated demand response (ADR) technology to manage microgrid loads during grid-independent (a.k.a. "islanded") operation. This automated approach achieves load shedding and shifting through communication signals sent to equipment controllers rather than by cutting off the flow of electricity within the microgrid itself. Because it operates only on the base network, with no connection to external entities, this strategy avoids the main cybersecurity concern raised by past applications of ADR on military bases.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data Interfaces for Automated Vehicle Services - A Municipality Perspective

As Automated Vehicle (AV) services proliferate, data sharing between AV operators and municipal agents is assuming greater importance. Information on the dynamic nature of the road system such as incidents to avoid, weather hazards (such as flooding), construction and detours, as well as active safety concerns (e.g. - riots) is important for AV operators. Such information cannot be directly sensed from a vehicle's sensor array, but instead must be communicated in a timely and trustworthy channel. Municipalities are interested in pushing this information to AV operators to support emergency response efforts, reduce traffic in construction zones, and generally improve operation of the system. Similarly, information on vehicle safety such as disengagements, as well as critical information on the use of roadway system (trips, origin and destination patterns) are important performance factors for municipalities to understand utilization and plan for appropriate infrastructure. As mobility shifts to on-demand options, the need for safe and coordinated pick-up and drop-off zones will increase (potentially reducing parking needs). For all of these reasons, communication flows between AV operators and municipalities are becoming increasingly important. This paper investigates the functions, emerging practices and protocols for sharing of such critical data, and identifies gaps in and challenges in existing practices. Additionally, case studies are used to highlight the impacts of data sharing between AV operators and municipalities.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

UPC++ as_eager Working Group Draft, Revision 2020.6.2

This draft proposes an extension for a new future-based completion variant that can be more effectively streamlined for RMA and atomic access operations that happen to be satisfied at runtime using purely node-local resources. Many such operations are most efficiently performed synchronously using load/store instructions on shared-memory mappings, where the actual access may only require a few CPU instructions. In such cases we believe it’s critical to minimize the overheads imposed by the UPC++ runtime and completion queues, in order to enable efficient operation on hierarchical node hardware using shared-memory bypass. The new upcxx::{source,operation}_cx::as_eager_future() completion variant accomplishes this goal by relaxing the current restriction that future-returning access operations must return a non-ready future whose completion is deferred until a subsequent explicit invocation of user-level progress. This relaxation allows access operations that are completed synchronously to instead return a ready future, thereby avoiding most or all of the runtime costs associated with deferment of future completion and subsequent mandatory entry into the progress engine. We additionally propose to make this new as_eager_future() completion variant the new default completion for communication operations that currently default to returning a future. This should encourage use of the streamlined variant, and may provide performance improvements to some codes without source changes. A mechanism is proposed to restore the legacy behavior on-demand for codes that might happen to rely on deferred completion for correctness. Finally, we propose a new as_eager_promise() completion variant that extends analogous improvements to promise-based completion, and corresponding changes to the default behavior of as_promise().

97 MATHEMATICS AND COMPUTING↗