Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Movement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Operational Evolution of FTS3: A DevOps Driven Approach to Elastic Operations

The File Transfer Service (FTS3) is a distributed data movement service developed at CERN and widely used to transfer data across the Worldwide LHC Computing Grid (WLCG). At Fermilab, FTS3 supports data transfers for multiple experiments, including Intensity Frontier experiments such as DUNE, enabling reliable data movement between WebDAV endpoints in Europe and the Americas.​ At CHEP 2021, we reported on the initial containerized deployment of FTS3 on OKD, the community Kubernetes distribution of Red Hat OpenShift. In this work, we present the subsequent evolution of this deployment, focusing on new operational capabilities introduced to improve scalability, robustness, and long-term maintainability.​ We describe the adoption of more secure and reproducible container build workflows, the integration of DevOps-driven operational practices, and enhancements in monitoring and automation. A key new result is the introduction of horizontal scaling and elastic resource management, allowing FTS3 components to dynamically adapt to workload variations while maintaining service reliability. We also discuss improvements in fault tolerance and operational procedures derived from production experience.​ Finally, we summarize lessons learned from operating FTS3 as a Kubernetes-native service and outline how these developments have improved the resilience and efficiency of data movement operations at Fermilab.

Munoz Flores, Victor Leopoldo [Fermilab]

Real-time tracking and analysis of gas bubble dynamics in laser powder bed fusion using in-situ X-ray characterization and machine learning

Porosity defects remain a significant challenge in the laser powder bed fusion (LPBF) process, adversely affecting the mechanical properties and reliability of additively manufactured components. Here, this study investigates the real-time formation and trajectory of gas bubbles during LPBF of Al6061 alloy using advanced in-situ X-ray characterization and machine learning. The unsupervised Gaussian mixture model and particle tracking algorithm developed are able to precisely track and quantify the properties of gas bubbles and keyhole pores. Our analysis identified five distinct types of gas bubble formation and movement patterns, emphasizing the diverse origins and behaviors of these defects. It enables precise quantification of trajectories, velocities, and morphological changes of gas bubbles, offering a granular view of the subsurface dynamics within the melt pool. Additionally, we explored keyhole-induced pore dynamics, revealing the critical role of keyhole oscillation and collapse for the formation of both large and small gas pores. It defines four different regions of gas bubble movement within the melt pool, providing a clearer understanding of how local fluid dynamics affect pore behavior. The results underscore the importance of integrating in-situ experimental observation and automated machine learning to develop a more robust predictive model for defect formation in LPBF.

In-situ X-ray imaging

Visualization of in-situ chemical flow through sand using neutron radiography

Chemical movement through soil is an important process in agriculture and ecology. Observing the spatial and temporal dynamics of these processes using conventional chemical ecology methods requires techniques that are destructive and/or lack resolution. Neutron radiography has the capability to allow chemical motion through sand/soil to be tracked with high spatial and temporal resolution, and we show that it allows for the motion of hydrophobic and hydrophilic chemicals to be distinguished. This technique can have an important impact on introducing neutron radiography to a wider community and into our understanding of chemical communication dynamics between plants and movement of applied chemicals in agricultural soils.

MORRIS, KATHRYN [Xavier University]

Queen bees offload pesticide burden to eggs when social buffering is overwhelmed

Honey bee colonies pollinate about one-third of the world’s food crops, and their rapid decline directly threatens agricultural productivity and ecosystem stability. Understanding how colony-level social defenses influence pesticide fate and the circumstances under which they fail is therefore a crucial question in pollinator biology. We used biological accelerator mass spectrometry (BioAMS), a sensitive radiotracer technique, to track the movement of a model pesticide through a small honey bee colony under laboratory conditions. We tested the hypothesis that social buffering protects honey bees from toxic accumulation and that this protection can be overcome, leading to maternal offloading of the pesticide to developing eggs. Consistent with this hypothesis, our results identified three key mechanisms governing chemical movement within a social insect colony: (1) worker bees initially decrease dietary pesticide levels by 95% through diet filtering and deposition in honeycombs, though this declines to 86% by day 10; (2) queen bees maintain markedly lower pesticide levels than workers but, over time, they accumulate the pesticide in their ovaries and transfer it into developing eggs, revealing a previously undocumented protective mechanism in reproductive individuals; and (3) the presence of a queen bee shifts colony-wide chemical distribution by concentrating worker exposure and increasing pesticide deposition in wax. Our findings show that honey bee colonies function as integrated detoxification networks, in which chemical fate depends on complex social behaviors and caste-specific physiology. When social buffering is overwhelmed, reproductive queens may survive by transferring their chemical burden to their offspring.

Biological and medical sciences

The impact of capillary heterogeneity on CO 2 flow and trapping across scales

Capillary heterogeneity has been identified over the last decade as a key control on subsurface CO 2 flow behavior during geological CO 2 sequestration. These heterogeneities can be formed in all sedimentary rocks, ranging from slight variations in the sand grain sizes to extensive sequences of interbedded sands, shales, and limestones. Capillary heterogeneity has been largely, although not entirely, overlooked in subsurface flow modeling because it is assumed to only directly influence fluid redistribution over scales of centimeters to meters. However, even small-scale fluid movements can result in dramatic impacts on the mobility and trapping of the CO 2 over kilometers. Therefore, neglecting capillary heterogeneity at multiple scales could potentially lead to errors in modeling and predicting field-scale plume migration. In this review paper, we aim to provide a consistent overview to (1) establish that capillary heterogeneity can have a major impact on CO 2 plume migration, (2) establish the respective length scales at which capillary heterogeneity matters, and (3) provide guidance for numerical modeling. This review covers pertinent literature and extracts key observations from the core to the field scales. Experimental studies have shown that millimeter-decimeter scale capillary heterogeneity can cause the so-called capillary heterogeneity trapping in addition to pore-scale residual trapping. Even at such a small scale, capillary heterogeneity can already lead to complex upscaled constitutive relationships, such as flow-rate dependent and anisotropic relative permeability, which affects field-scale CO 2 migration even when field-scale heterogeneities are present. Under gravity-dominated flow regimes, centimeter-meter scale capillary heterogeneity can entrap a significant amount of CO 2 at field scale, not just after imbibition but also during drainage. In certain cases, the presence of capillary heterogeneity can even completely stop the vertical movement of the CO 2 plume, hence greatly reducing leakage risks. At meter-kilometer scale, the influence of capillary heterogeneity is more pronounced and can hinder or redirect CO 2 migration in both lateral and vertical directions. The impact of capillary heterogeneity across multiple spatial scales poses a great challenge in modeling CO 2 migration at field scale, because it is practically impossible to build a field-scale earth model with grid blocks at millimeter scale. We recommend a hierarchical modeling approach to address this challenge. At field scale, earth models are built to capture geological features and heterogeneities in high but still practical grid resolutions. For each facies or rock type of the field-scale model, high- resolution meter-scale “conceptual” models are built with millimeter-scale grid blocks to capture representative fine-scale bedding geometries and heterogeneities in various environments of deposition, bridging the gap from subcore scale to the size of a field-scale simulation grid block. Upscaling is then used to preserve the smaller-scale flow dynamics of various rock types in field-scale simulations. Here, future work is needed to (1) refine, improve, and validate the hierarchical modeling approach; (2) build libraries of fine-scale bedding models for facies in various environments of deposition; (3) quantify multiscale capillary heterogeneity effects under subsurface uncertainties; (4) gain learning from different storage formations; and (5) establish best practices that balance accuracy and computational speed.

Capillary heterogeneity

Establishing a silica gel zone in well annulus and evaluating its performance in blocking vertical water flow

Wells are often constructed for monitoring purposes with relatively long screen lengths (e.g., >10 m). Vertical water flows can occur within the artificial or natural filterpack annulus that surrounds the screened interval, bypassing packer assemblies installed inside the wellbore. Attempts to isolate discrete vertical zones during groundwater sampling are unsuccessful when annular flow occurs and lead to remedy decisions based on biased or incorrect interpretations. Blocking vertical annular water flow and contaminant transport will help obtain more accurate concentrations of contaminants from sampling in targeted depth intervals. The application of silica gels formed from the injected colloidal silica CS suspensions is a novel approach to minimize or prevent movement of vertical movement of groundwater in the surrounding filterpack annulus. In this work, we tested the feasibility of injecting CS suspensions to target locations and developed a modified CS formulation that is injectable and prevents gravity sinking. We studied the distribution and penetration of silica gel at laboratory scale in mock well annulus with surrounding formations. We evaluated the performance of the silica gel in blocking vertical water flow in the annulus and in minimizing chemical transport through the gel zone. CS suspension formulations have been defined that are ready for injection, stay in target locations, and form gel within desired time frames. Injection of CS suspensions achieved uniform distribution in a well annulus filter pack, fully occupied the annulus pore space, and penetrated the formation surrounding the filter packer with a sufficient distance to create a hydraulic annular seal when the injection was applied at a sufficient rate. The depth of penetration into the formation was dependent on the permeability contrast between the filter pack and the surrounding formation. Silica gel that formed in the annulus blocked vertical water flow and stopped the chemical transport through the gel zone. In conclusion, this research reveals that using CS suspension injection and sequential gelation (CS-GEL) is a promising technology for blocking vertical water flow and chemical transport through the filter pack in targeted zones within the annulus of long-screened well systems.

Colloidal silica suspension

Direct observation of annealing-driven recrystallization behavior in magnesium alloy at low strain condition

Grain-twin interactions are significant in texture modification under thermodynamic driving force. In this study, annealing-driven twinning/detwinning behavior, grain growth, and corresponding texture evolution in a pre-deformed AZ31 magnesium alloy were systematically tracked and investigated via in-situ heating synchrotron X-ray diffraction and quasi in-situ electron backscattered diffraction techniques. A twinning texture is generated in the pre-deformed sample due to the activation of {$10\bar{1}2$} tensile twinning. During annealing, dislocation annihilation occurs between 100 and 280 °C, and recrystallization occurs above 280 °C, manifesting as the initial residual matrix and twins being competitively swallowed by each other, forming a bimodal texture. The recrystallization process is completed by boundary movement, which depends on the energy difference across the boundary. In addition, it is found that the grain boundaries favor movement towards the side with higher stored energy, regardless of the boundary type or the boundary energy.

In-situ observation

Design of a robot-automated flat plate/reflection geometry x-ray diffraction setup for accelerated materials discovery and structural screening

Here, we report the design, construction, and automation of a flat plate sample loading, alignment, and data acquisition system for X-ray diffraction measurements in reflection geometry implemented at the Stanford Synchrotron Radiation Lightsource. The system is built onto a single platform, enabling facile transferability, and is compartmentalized into sample storage, sample transfer, and sample position/alignment segments. The core feature of this system is a six-axis robotic arm that offers a large range of highly reproducible and programable movements. The degrees of freedom of the robot arm enable adaptability in which movements can be modified to fit various beamline environments and sample configurations. Samples are housed on 3D printed sample mounts, which are arranged onto a 6 × 2 array of sample cassettes capable of holding 7 samples. Using sample mounts designed for solid oxide electrolysis button cells (SOECs), the maximum tray capacity is 84 samples, which can be aligned and run in ~ 24 hours with long exposure scans. The sample array is additionally capable of accommodating a range of sample sizes and geometries due to the rapid 3D printed fabrication. The components of the setup will be described in detail and performance will be demonstrated with a set of representative SOEC and XRD standard samples. Opportunities for future developments and integration with the automated setup are summarized.

08 HYDROGEN

A profile monitor for proton radiography experiments at the Los Alamos Neutron Science Center

The Proton Radiography (pRad) facility at the Los Alamos Neutron Science Center utilizes pulses of protons delivered by the 800 MeV linear accelerator to produce a series of radiographic images to study the dynamic behavior of materials under extreme conditions. Radiographs taken with an empty field of view, or beam pictures, are used to normalize transmission. However, because the center of the proton beam shifts between pulses, an in situ method for measuring beam position is required to normalize images for beam movement to perform absolute radiography. The beam profile monitor described here uses an array of scintillating fibers positioned in the beam path to produce light proportional to beam intensity across the beam cross section. This light is detected using fast photodiodes and a digital oscilloscope, providing a response time of several nanoseconds—suitable for measuring the 50-ns proton pulses used in pRad. The profile monitor achieves a measured position precision of 40 μm and an intensity precision of 0.7%, allowing for beam movement corrections to be applied to images, thereby improving data accuracy and image quality.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Characterization of ELM pacing via vertical jogs on DIII-D

Edge localized mode (ELM) pacing via vertical plasma oscillations or jogging has been successfully demonstrated on DIII-D. Rapid vertical movement of the plasma toward the X-point has been shown to effectively trigger ELMs. By vertically oscillating the plasma at a rate of 10 Hz, the ELM frequency increased from ~5 Hz, the natural ELM frequency in similar DIII-D discharges, to 20 Hz. Downward jogs have been observed to trigger multiple ELMs in one cycle. ELMs triggered at higher than natural frequencies lead to smaller decreases in stored energy, from 8% to as little as below 1%. As a consequence, the peak heat flux to the divertor has been observed to be reduced by a factor of ~2. In addition, a reduction in the carbon impurity concentration has been observed. During downward jogs in the lower single null (LSN) configuration, the X-point movement is slower and smaller than the top of the plasma. As a result, a reduction in the plasma cross-section and hence volume has been observed. To understand the mechanism of ELM triggering by jogging, a toy model of the edge toroidal current has been built and tested with DIII-D experiment data. The experimental data and model suggest that when the plasma moves down toward the X-point, a net positive toroidal current is locally induced in the edge region. ELITE stability analysis suggests that this current pushes the plasma state across the peeling side of the peeling–ballooning stability boundary into the unstable region triggering ELMs.

ELM pacing

Simulation driven adaptive sampling for neutron-diffraction based strain mapping of additively manufactured parts

Neutron diffraction based strain mapping is a useful technique for measuring residual strains in additively manufactured (AM) metal parts. The measurement is traditionally done by scanning the sample in a point-wise raster pattern to extract the strain at each position. Since the overall scan can span several hours, adaptive sampling approaches using Bayesian optimization based on Gaussian process (BO-GP) regression have been introduced—demonstrating that even with a fraction of the typically made measurements the dominant strain patterns in the sample can be reconstructed. However, the parameters of the BO-GP algorithm have to be carefully chosen for best performance, and the movement time between arbitrary points can offset the time savings from a reduced number of measurement locations. In this paper, we propose algorithms to refine the BO-GP based methods by using simulations of strain patterns in AM parts based on the materials and the process used to print them. We demonstrate that the simulated strain patterns can be used to help choose better parameters for the BO-GP based framework—leading to low reconstruction error for the final strain pattern. Furthermore, we show that the strain mapping experiment can be initialized with a sampling pattern learnt from the simulation data and ordered to reduce movement time, dramatically enabling reduction in the overall time required to run the baseline BO-GP method.

Gaussian process regression

Vegetation heterogeneity reflects soil thermal state and surface soil displacement in a thawing permafrost landscape

Thawing permafrost has the potential to dramatically alter the physical and ecological structure of northern landscapes. Warming of the Arctic and subsequent degradation of permafrost have created a need to assess the stability and movement of soils on hillslopes and the potential impacts on ecosystem structure. In this work, we explore the relationships among vegetation heterogeneity, soil temperature, and soil surface displacements observed from 2019 to 2022 in a watershed in the discontinuous permafrost region on the Seward Peninsula of Alaska. Vegetation heterogeneity was measured as the standard deviation (SD) of the normalized difference vegetation index (NDVI) from 3 m PlanetScope satellite imagery around each soil temperature and active layer thickness observation. Locations of observations were clustered into three soil thermal groups, warm, intermediate, and cold, based on soil temperature and active layer thickness. Average annual horizontal surface displacements were significantly lower for soils within the warm thermal group (median = 0.033 m yr −1 ) compared to soils within the cold thermal group (median = 0.090 m yr −1 ; p < 0.001). Conversely, vegetation heterogeneity was significantly higher in the warm (median = 0.014 SD NDVI; p = 0.002) and intermediate (median = 0.015 SD NDVI; p = 0.002) groups compared with the cold thermal group (median = 0.012 SD NDVI), suggesting a warming-induced shift in vegetation community complexity. Because of the observed associations of ground surface displacement rates and vegetation heterogeneity with soil thermal state, we hypothesize that warming soil conditions induce changes in the rates and patterns of hillslope erosion due to an increase in surface movement as near-surface permafrost thaws, followed by a decrease as the permafrost table deepens and excess ice content diminishes. The transition to warm soils promotes surface ecosystem transformation, shifting the dominant vegetation at the site, given the warming climatic conditions of the region. We integrated our observations of soil temperature, vegetation heterogeneity, and soil surface displacements into a conceptual model that describes the co-evolution of hillslopes and vegetation in warming permafrost environments, which is currently unrepresented in earth system models.

54 ENVIRONMENTAL SCIENCES

Dual Channel Dual Staging: Hierarchical and Portable Staging for GPU-Based In-Situ Workflow

In-situ workflows have emerged as an attractive approach for addressing data movement challenges at very large scales. Since GPU-based architectures dominate the HPC landscapes, porting these in-situ workflows, and, specifically, the inter-application data exchange, to GPU-based systems can be challenging. Technologies such as GPUDirect RDMA (GDR), which is typically used for I/O in GPU applications as an optimization that circumvents the CPU overhead, can be leveraged to support bulk data exchanges between GPU applications. However, current GDR design often lacks performance portability across HPC clusters built with different hardware configurations. Furthermore, the local CPU may also be effectively used as an auxiliary communication mechanism to offload data exchanges. In this paper, we present a dual channel dual staging approach for efficient, scalable, and performance-portable inter-application data exchange for in-situ workflows. This approach exploits the data access pattern within in-situ workflows along with the inherent execution asynchrony to accelerate data exchanges and, at the same time, improve performance portability. Specifically, the dual channel dual staging method leverages both the local CPU and the remote data staging server to build a hierarchical joint staging area and uses this staging area to transform blocking inter-application bulk data exchanges into best-effort local data movements between GPU and CPU. The dual channel dual staging is implemented as a portability extension of the Dataspaces-GPU staging framework. We present an experimental evaluation of its performance, portability, and scalability using this implementation on three leadership GPU clusters. The evaluation results demonstrate that the dual channel dual staging method saves up to 75% in data-exchange time compared to host-based, GDR, and alternate portable designs, while maintaining scalability (up to 512 GPUs) and performance portability across the three platforms.

Zhang, Bo [University of Utah]

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue

UltraLiM: In-Memory Boolean Logic Architecture Using UltraRAM

Conventional computing architectures encounter ‘von Neumann’ and ‘memory wall’ bottlenecks which arise due to the back-and-forth data movement between the physically separate memory and processing units and the speed mismatch between them, respectively. These bottlenecks hurt both energy efficiency and the throughput of computing systems. To address these challenges, in-memory computing architectures have emerged as a promising alternative. They reduce the need for frequent data movement by executing different computing tasks inside the memory system. Here, we present UltraLiM, a logic-in-memory architecture using the UltraRAM-based memory system. UltraRAM holds the promise of developing a ‘universal memory’, overcoming the limitations of charge-based memories thanks to their non-volatile behavior with lower operating voltage. This work presents an in-memory computing architecture that integrates an UltraRAM-based memory array with a custom-designed peripheral circuitry. With this architecture, we can perform various in-memory Boolean logic operations (such as NOT, NAND, NOR, and XOR) in a single cycle. Leveraging the separate read-write paths in the UltraRAM-based memory array, we optimize read operations without encountering design conflicts. This optimization enhances the sense margin, enabling the use of simpler peripheral circuitry for in-memory logic operations.

Alam, Shamiul [University of Tennessee, Knoxville

Optimizing Traffic Signal Control to Enhance Transportation Efficiency and Maximize Pedestrian Benefits in the Road Network

Increasing urban mobility requirements demand efficient transportation system strategies for both vehicular and pedestrian movement. This study enhances the Decentralized Graph-based Multi-Agent Reinforcement Learning (DGMARL) approach, originally tailored for vehicular traffic signal timing, to incorporate pedestrian traffic dynamics. The improved algorithm considers crucial metrics such as Eco_PI, assesses vehicle fuel consumption by factoring in stops and delays, and addresses pedestrian waiting time, crucial for system efficiency while acknowledging driver waiting time impact. Utilizing Digital Twin simulation along the MLK Smart Corridor in Chattanooga, Tennessee, the algorithm's performance is compared for various pedestrian control scenarios. To evaluate the effectiveness of DGMARL, this study compared DGMARL-enabled signal management with automated pedestrian traffic detection and an actuated signal management system (real-word baseline) with pedestrian recall, which predetermingly enforces a pedestrian phase every cycle. Findings indicate substantial improvements with DGMARL, showing a 28.29% enhancement in vehicle Eco_PI, a 60.55 % reduction in pedestrian waiting time, and a 55.74% decrease in driver stop delay, on average, compared to the baseline actuated signal timing plan.

Kumarasamy, Vijayalakshmi K [The University of Ten

Using Temporal Information from Human Mobility Data to Detect Anchor Points

Spatiotemporal mobility data are available in massive quantities, but large quantities of data typically include fewer variables or data fields. Often, the only available fields are User ID, Longitude, Latitude, Timestamp (ULLT). This raises an important question: how much can we infer about human mobility patterns using only these four fields? With ULLT data, we do not know individuals' socioeconomic status information or when they are visiting their anchor points (AP) or locations (such as homes, places of employment, or schools), and it is a modern challenge to use this data to infer these characteristics. When detecting anchor locations with limited input information, verification and validation (VV) are significant challenges. This paper addresses the problem of identifying individuals' anchor locations using only temporal information from spatiotemporal datasets with limited attributes. Our approach does not explicitly use latitude and longitude during analysis. Locationbased information is only employed in the preprocessing stage to identify periods of movement (trips) and stops (dwelling). Beyond this step, all analysis is based on temporal patterns. In theory, if stops and dwell times could be detected through alternative means, our method could function entirely without location-based input. We demonstrate this methodology on the 2017 National Household Travel Survey (NHTS) data, because it includes a carefully designed and collected time use survey with representative sampling and labeled ground truth. The high-quality survey data allows us to test the accuracy of our methods because NHTS contains intended place labels and agent/user characteristics. We have also applied our validated AP identification algorithm on very large-scale GPS based trajectory data for Patterns-of-Life (PoL) assessment and other applications, but due to space limit that could not be presented here.

McBride, Liz [ORNL] (ORCID:0000000286925869)

IRIS: A Performance-Portable Framework for Cross-Platform Heterogeneous Computing

From edge to exascale, computer architectures are becoming more heterogeneous and complex. The systems typically have fat nodes, with multicore CPUs and multiple hardware accelerators such as GPUs, FPGAs, and DSPs. This complexity is causing a crisis in programming systems and performance portability. Several programming systems are working to address these challenges, but the increasing architectural diversity is forcing software stacks and applications to be specialized for each architecture. As we show, all of these approaches critically depend on their software framework for discovery, execution, scheduling, and data orchestration. To address this challenge, we believe that a more agile and proactive software framework is essential to increase performance portability and improve user productivity. To this end, we have designed and implemented IRIS: a performance-portable framework for cross-platform heterogeneous computing. IRIS can discover available resources, manage multiple diverse programming platforms (e.g., CUDA, Hexagon, HIP, Level Zero, OpenCL, OpenMP) simultaneously in the same execution, respect data dependencies, orchestrate data movement proactively, and provide for user-configurable scheduling. To simplify data movement, IRIS introduces a shared virtual device memory with relaxed consistency among different heterogeneous devices. IRIS also adds an automatic kernel workload partitioning technique using the polyhedral model so that it can resize kernels for a wide range of devices. Our evaluation on three architectures, ranging from Qualcomm Snapdragon to a Summit supercomputer node, shows that IRIS improves portability across a wide range of diverse heterogeneous architectures with negligible overhead.

97 MATHEMATICS AND COMPUTING