Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Access Patterns”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗

Spiner: Performance Portable Routines for Generic, Tabulated, Multi-Dimensional Data

We present Spiner, a new, performance-portable library for working with tabulated data. Spiner provides efficient routines for multi-dimensional interpolation and indexing on both CPUs and GPUs—as well as more exotic hardware—including interwoven interpolation and indexing access patterns, as needed for radiation transport. Importantly, Spiner defines a data format, based on HDF5, that couples the tabulated data to the information required to interpolate it, which Spiner can read and move to GPU.

97 MATHEMATICS AND COMPUTING↗

CompF2: Theoretical Calculations and Simulation Topical Group Report

This report summarizes the work of the Computational Frontier topical group on theoretical calculations and simulation for Snowmass 2021. We discuss the challenges, potential solutions, and needs facing six diverse but related topical areas that span the subject of theoretical calculations and simulation in high energy physics (HEP): cosmic calculations, particle accelerator modeling, detector simulation, event generators, perturbative calculations, and lattice QCD (quantum chromodynamics). The challenges arise from the next generations of HEP experiments, which will include more complex instruments, provide larger data volumes, and perform more precise measurements. Calculations and simulations will need to keep up with these increased requirements. The other aspect of the challenge is the evolution of computing landscape away from general-purpose computing on CPUs and toward special-purpose accelerators and coprocessors such as GPUs and FPGAs. These newer devices can provide substantial improvements for certain categories of algorithms, at the expense of more specialized programming and memory and data access patterns.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

IRI Technology Landscape – A survey of re-usable components and methodologies

This document describes technical implementation details on network access schemes connecting API-driven workflows to supercomputer centers. API-driven workflows are a central theme in connected computing, since they bring the terminal-mainframe' access pattern present since the 1970s up to the task of interfacing with modern web browser technologies. Both security (HTTPS/TLS/IPSec/VPNs/public key cryptography/digital signatures) and network protocol stacks (HTTP-REST APIs, tokens, gRPC, SRTP) have evolved to the point where implementing API-driven workflows is possible using stable, secure off-the-shelf software.

97 MATHEMATICS AND COMPUTING↗

Effectively Using Remote I/O For Work Composition in Distributed Workflows

Distributed scientific workflows are becoming more important with the interest in incorporating AI into their loops. A critical programming and performance question is how to compose workflow tasks when data is produced on one system but must be consumed on another. Since the dominant technique is composition with remote I/O, this paper explores its performance expectations. We describe BigFlowSim, a workflow I/O simulator that captures key implementation choices for remote I/O, including intensity, reuse, locality, access pattern, and data movement.With BigFlowSim, we generate a synthetic benchmark. We quantify the effects of each parameter with a performance sensitivity study. We explain trends in terms of data movement reduction and show that, under certain conditions, it is possible to establish a total order among most parameters. We apply these insights to a high energy physics workflow, Belle II Monte Carlo and simulate several I/O optimizations. Speedups range from 5% to 2×, without changing compute time.

Friese, Ryan D.↗

Quantifying Message Aggregation Optimisations for Energy Savings in PGAS Models

Upon breaking past the exascale barrier, HPC systems are facing their greatest challenge yet - a power wall that must be addressed through new methods in both hardware and software. While energy costs are becoming a major issue at all levels, of particular concern is that of the network, as the relative cost of moving data is increasing faster than ever. The partitioned global address space (PGAS) model is critical within certain HPC domains, but is known to suffer from the small message problem, where irregular many-to-many access patterns result in congesting the network with excessive numbers of small messages. To address this, the conveyor aggregation library was developed to defer individual messages and group them for subsequent bulk processing. In this paper, we investigate its impact on energy use related to the network, with a focus on the Slingshot 11 interconnect. We will demonstrate that this strategy is not only highly performant, but also crucial to reducing energy footprints to remain within target power envelopes.

Welch, Aaron [ORNL]↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

April 2020 Darshan counters from the Summit supercomputer

This dataset is the Darshan counters collected from the Summit supercomputer in a month of April 2020. 1. Description of methods used for collection/generation of data: Job submitted on Summit HPC system when completed successfully and has made I/O calls (captured by Darshan tool) writes a Darshan log file on alpine filesystem. One job can have multiple `jsrun` commands and Darshan will generate separate logs each log corresponding to an `jsrun` command, so a job can have one or more Darshan logs associated with it. 2. Methods for processing the data: To process the data, we first use `darshan-util` tool to parse the Darshan logs. Then we restructure the logs and merge data from multiple Darshan logs if they belong to the same Summit job.

97 MATHEMATICS AND COMPUTING↗

Angular Momentum in Rotating Superfluid Droplets

The angular momentum of rotating superfluid droplets originates from quantized vortices and capillary waves, the interplay between which remains to be uncovered. Here, the rotation of isolated submicrometer superfluid 4 He droplets is studied by ultrafast x-ray diffraction using a free electron laser. The diffraction patterns provide simultaneous access to the morphology of the droplets and the vortex arrays they host. In capsule-shaped droplets, vortices form a distorted triangular lattice, whereas they arrange along elliptical contours in ellipsoidal droplets. The combined action of vortices and capillary waves results in droplet shapes close to those of classical droplets rotating with the same angular velocity. The findings are corroborated by density functional theory calculations describing the velocity fields and shape deformations of a rotating superfluid cylinder.

O'Connell, Sean M.O.↗

Geospatial Analysis of Built Infrastructure and Modeled Household Driving Patterns

The level of access to opportunities for a location can be quantified in the amount of time it takes to travel from a departure point to the destinations that somebody would want or need to visit. Isochrone maps are geometric representations of the area accessible from a departure point within a set amount of time. Informed in part by the National Household Travel Survey, this report merges location data for amenities and opportunities across six frequent destination categories – employment, education, health, food, community, and transportation – with isochrone maps generated by the TravelTime API, whose departure points are census tract population-weighted centroids. Using a “Points-In-Polygon” analysis, destinations that fall within a census tract’s isochrone are tallied as accessible from the region within one of three time thresholds: 15-, 30-, and 45-minutes by the walking, cycling, public transit, and driving modes of travel. We find that access to a high number of jobs within a typical commute duration is negatively correlated with annual household vehicle miles traveled (VMT). The spatial distribution of our data suggests that the high household VMT frequently seen surrounding the edges of major cities may be related to worker commutes into the city core, and that the high household VMT frequently seen in rural tracts may be related to the longer travel distances required to access a variety of key opportunities from these areas.

99 GENERAL AND MISCELLANEOUS↗

Modeling Distributed Computing Infrastructures for HEP Applications

Predicting the performance of various infrastructure design options in complex federated infrastructures with computing sites distributed over a wide area network that support a plethora of users and workflows, such as the Worldwide LHC Computing Grid (WLCG), is not trivial. Due to the complexity and size of these infrastructures, it is not feasible to deploy experimental test-beds at large scales merely for the purpose of comparing and evaluating alternate designs. An alternative is to study the behaviours of these systems using simulation. This approach has been used successfully in the past to identify efficient and practical infrastructure designs for High Energy Physics (HEP). A prominent example is the Monarc simulation framework, which was used to study the initial structure of the WLCG. New simulation capabilities are needed to simulate large-scale heterogeneous computing systems with complex networks, data access and caching patterns. A modern tool to simulate HEP workloads that execute on distributed computing infrastructures based on the SimGrid and WRENCH simulation frameworks is outlined. Studies of its accuracy and scalability are presented using HEP as a case-study. Hypothetical adjustments to prevailing computing architectures in HEP are studied providing insights into the dynamics of a part of the WLCG and candidates for improvements.

Horzela, Maximilian↗

Satellite Data Applications for Sustainable Energy Transitions

Transitioning to a sustainable energy system poses a massive challenge to communities, nations, and the global economy in the next decade and beyond. A growing portfolio of satellite data products is available to support this transition. Satellite data complement other information sources to provide a more complete picture of the global energy system, often with continuous spatial coverage over targeted areas or even the entire Earth. We find that satellite data are already being applied to a wide range of energy issues with varying information needs, from planning and operation of renewable energy projects, to tracking changing patterns in energy access and use, to monitoring environmental impacts and verifying the effectiveness of emissions reduction efforts. While satellite data could play a larger role throughout the policy and planning lifecycle, there are technical, social, and structural barriers to their increased use. We conclude with a discussion of opportunities for satellite data applications to energy and recommendations for research to maximize the value of satellite data for sustainable energy transitions.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

TriC: Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graph analytics has emerged as an important tool in the analysis of large scale data from diverse application domains such as social networks, cyber security and bioinformatics. Counting the number of triangles in a graph is a fundamental kernel with several applications such as detecting the community structure of a graph or in identifying important vertices in a graph. The ubiquity of massive datasets is driving the need to scale graph analytics on parallel systems. However, numerous challenges exist in efficiently parallelizing graph algorithms, especially on distributed-memory systems. Irregular memory accesses and communication patterns, low computation to communication ratios, and the need for frequent synchronization are some of the leading challenges. In this paper, we present TriC, our distributed-memory implementation of triangle counting in graphs using the Message Passing Interface (MPI), as a submission to the 2020 GraphChallenge competition. Using a set of synthetic and real-world inputs from the challenge, we demonstrate a speedup of up to 90x relative to previous work on 32 processor-cores of a NERSC Cori node. We also provide details from distributed runs with up to8192 processes along with strong scaling results. The observations presented in this work provide an understanding of the system-level bottlenecks at scale that specifically impact sparse-irregular workloads and will therefore benefit other efforts to parallelize graph algorithms.

Halappanavar, Mahantesh↗

Pattern-aware prefetching using parallel log-structured file system

Techniques are provided for pattern-aware prefetching using a parallel log-structured file system. At least a portion of one or more files is accessed by detecting at least one pattern in a non-sequential access of the one or more files; and obtaining at least a portion of the one or more files based on the detected at least one pattern. The obtaining step comprises, for example, a prefetching or pre-allocation of the at least the portion of the one or more files. A prefetch cache can store the portion of the one or more obtained files. The cached portion of the one or more files can be provided from the prefetch cache to an application requesting the at least a portion of the one or more files.

Bent, John M.↗

Convergent Hydraulic Redistribution and Groundwater Access Supported Facilitative Dependency Between Trees and Grasses in a Semi‐Arid Environment

Abstract Hydraulic redistribution is the transport of water from wet to dry soil layers, upward or downward, through plant roots. Often in savanna and woodland ecosystems, deep‐rooted trees, and shallow‐rooted grasses coexist. The degree to which these different species compete for or share soil‐water derived from precipitation or groundwater, as well as how these interactions are altered by hydraulic redistribution, is unknown. We use a multilayer canopy model and field observations to examine how the presence of deep, but tree‐root accessible, groundwater impacts seasonal patterns of hydraulic redistribution, and interaction between coexisting vegetation species in a semiarid riparian woodland (US‐CMW). Based on the simulation, trees absorb moisture at the water table (∼10 m depth) and release it in the shallow soil depth (0–3 m) during the dry pre‐monsoon season. We observed the occurrence of a new convergent hydraulic redistribution pattern during the monsoon season, where moisture is transported from both the near‐surface (0–0.5 m) and the water table to intermediate soil layers (1–5 m) through tree roots. We found that hydraulic redistribution demonstrates a growth facilitation effect at this site, supporting 49% of growing season tree transpiration and 14% of the grass transpiration. Compared to a similarly structured upland savanna without accessible groundwater, the riparian site shows an increased amount of hydraulically redistributed water and more facilitative water use between coexisting grasses and trees. These results shed light on the linkage between accessible groundwater and the role of hydraulic redistribution on the interaction between deep‐rooted and shallow‐rooted vegetation.

Lee, E.↗

Tailored topotactic chemistry unlocks heterostructures of magnetic intercalation compounds

The construction of thin film heterostructures has been a widely successful archetype for fabricating materials with emergent physical properties. This strategy is of particular importance for the design of multilayer magnetic architectures in which direct interfacial spin-spin interactions between magnetic phases in dissimilar layers lead to emergent and controllable magnetic behavior. However, crystallographic incommensurability and atomic-scale interfacial disorder can severely limit the types of materials amenable to this strategy, as well as the performance of these systems. Here, we demonstrate a method for synthesizing heterostructures comprising magnetic intercalation compounds of transition metal dichalcogenides (TMDs), through directed topotactic reaction of the TMD with a metal oxide. The mechanism of the intercalation reaction enables thermally initiated intercalation of the TMD from lithographically patterned oxide films, giving access to a family of multi-component magnetic architectures through the combination of deterministic van der Waals assembly and directed intercalation chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES↗

Chromosome‐level Thlaspi arvense genome provides new tools for translational research and for a newly domesticated cash cover crop of the cooler climates

Summary Thlaspi arvense (field pennycress) is being domesticated as a winter annual oilseed crop capable of improving ecosystems and intensifying agricultural productivity without increasing land use. It is a selfing diploid with a short life cycle and is amenable to genetic manipulations, making it an accessible field‐based model species for genetics and epigenetics. The availability of a high‐quality reference genome is vital for understanding pennycress physiology and for clarifying its evolutionary history within the Brassicaceae. Here, we present a chromosome‐level genome assembly of var. MN106‐Ref with improved gene annotation and use it to investigate gene structure differences between two accessions (MN108 and Spring32‐10) that are highly amenable to genetic transformation. We describe non‐coding RNAs, pseudogenes and transposable elements, and highlight tissue‐specific expression and methylation patterns. Resequencing of forty wild accessions provided insights into genome‐wide genetic variation, and QTL regions were identified for a seedling colour phenotype. Altogether, these data will serve as a tool for pennycress improvement in general and for translational research across the Brassicaceae.

59 BASIC BIOLOGICAL SCIENCES↗