Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “memory access optimization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

VA EDH Advanced Software Pipeline Framework Report: Enhancing Automation and Scalability

The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.

97 MATHEMATICS AND COMPUTING↗

Deterministic Current–Induced Perpendicular Switching in Epitaxial Co/Pt Layers without an External Field

Current–induced spin–orbit torques (SOTs) have emerged as a powerful tool to control magnetic elements and non–uniform magnetic textures such as domain walls and skyrmions. SOT–induced switching of perpendicular magnetization generally requires an external field to break the rotational symmetry of the spin–orbit effective fields responsible for the deterministic reversal. The proposed mechanisms to eliminate this requirement often rely on complex multilayer structures that necessitate laborious optimization in the material and spin transport properties, making them less attractive for applications. Herein, current–induced, external field–free switching of an epitaxial MgO/Pt/Co trilayer with an extremely large perpendicular anisotropy in excess of 3 Tesla is reported. It is found that switching occurs due to the interplay of strong SOTs, local anisotropy fluctuations, and the Dzyaloshinkii–Moriya interaction inherent to this epitaxial system. Finally, given that these layers constitute the base stack of a magnetic tunnel junction, this switching mechanism offers the most technologically viable path toward devices such as field–free SOT–based magnetic random–access memories.

36 MATERIALS SCIENCE↗

DTLMod: A simulation framework for in situ workflow optimization

In situ processing workflows have become essential for coping with the explosion in data volume and velocity in large-scale scientific computing, providing domain scientists with early insights at runtime. Multiple frameworks implement this paradigm through a data transport layer (DTL), offering different data access modes and deployment schemes, but researchers currently lack the appropriate tools to assess design and deployment options before committing to costly real experiments. We introduce DTLMod, an open-source simulated DTL that enables performance evaluation of in situ workflow configurations at scale. Built on SimGrid, it links into any SimGrid-based simulator and is available in C++ and Python. We evaluate DTLMod along four axes: scalability (tens of thousands of simulated processes across interconnected clusters in seconds, with linear memory scaling), versatility (three implementation variants trading fidelity for speed), accuracy (simulated times faithfully reflecting real behavior), and practical utility (two use cases demonstrating evidence-based workflow design decisions).

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

A kinetic-based regularization method for data science applications

We propose a physics-based regularization technique for function learning, inspired by statistical mechanics. By drawing an analogy between optimizing the parameters of an interpolator and minimizing the energy of a system, we introduce corrections that impose constraints on the lower-order moments of the data distribution. This minimizes the discrepancy between the discrete and continuum representations of the data, in turn allowing to access more favorable energy landscapes, thus improving the accuracy of the interpolator. Our approach improves performance in both interpolation and regression tasks, even in high-dimensional spaces. Unlike traditional methods, it does not require empirical parameter tuning, making it particularly effective for handling noisy data. We also show that thanks to its local nature, the method offers computational and memory efficiency advantages over Radial Basis Function interpolators, especially for large datasets.

97 MATHEMATICS AND COMPUTING↗

Phase Stability Through Machine Learning

Understanding the phase stability of a chemical system constitutes the foundation of materials science. Knowledge of the equilibrium state of a system under arbitrary thermodynamic conditions provides valuable information about the types of phases that are likely to be synthesized and how to get there. Accessing the phase diagram in a materials system provides one with the information necessary to design materials and microstructures with optimal properties. While the materials science community has long been focused on exploiting this knowledge to navigate the materials space, recent advances in machine learning (ML) and artificial intelligence (AI) have provided the community with novel ways of interrogating the materials thermodynamics space. Furthermore, this work presents some of the most recent advances in ML/AI applied to phase stability and thermodynamics of materials. Prof. John Morral always had a passion for understanding and teaching the fundamental characteristics of phase diagrams. This review is written to honor his memory.

36 MATERIALS SCIENCE↗

Analysis of GPU Data Access Patterns on Complex Geometries for the D3Q19 Lattice Boltzmann Algorithm

GPU performance of the lattice Boltzmann method (LBM) depends heavily on memory access patterns. When implemented with GPUs on complex domains, typically, geometric data is accessed indirectly and lattice data is accessed lexicographically. Although there are a variety of other options, no study has examined the relative efficacy between them. Here, we examine a suite of memory access schemes via empirical testing and performance modeling. We find strong evidence that semi-direct is often better suited than the more common indirect addressing, providing increased computational speed and reducing memory consumption. For the layout, we find that the Collected Structure of Arrays (CSoA) and bundling layouts outperform the common Structure of Array layout; on V100 and P100 devices, CSoA consistently outperforms bundling, however the relationship is more complicated on K40 devices. When compared to state-of-the-art practices, our recommendations lead to speedups of 10–40 percent and reduce memory consumption up to 17 percent. Using performance modeling and computational experimentation, we determine the mechanisms behind the accelerations. We demonstrate that our results hold across multiple GPUs on two leadership class systems, and present the first near-optimal strong results for LBM with arterial geometries run on GPUs.

42 ENGINEERING↗

Enabling Scalable and Extensible Memory-mapped Datastores in Userspace

Exascale workloads are expected to incorporate data-intensive processing in close coordination with traditional physics simulations. These emerging scientific, data-analytics and machine learning applications need to access a wide variety of datastores in flat files and structured databases. Programmer productivity is greatly enhanced by mapping datastores into the application process's virtual memory space to provide a unified “in-memory” interface. Currently, memory mapping is provided by system software primarily designed for generality and reliability. However, scalability at high concurrency is a formidable challenge on exascale systems. Also, there is a need for extensibility to support new datastores potentially requiring HPC data transfer services. In this article, we present UMap , a scalable and extensible userspace service for memory-mapping datastores. Furthermore, through decoupled queue management, concurrency aware adaptation, and dynamic load balancing, UMap enables application performance to scale even at high concurrency. We evaluate UMap in data-intensive applications, including sorting, graph traversal, database operations, and metagenomic analytics. Our results show that UMap as a userspace service outperforms an optimized kernel-based service across a wide range of intra-node concurrency by 1.22-1.9 × . We performed two case studies to demonstrate UMap 's extensibility. First, a new datastore residing in remote memory is incorporated into UMap as an application-specific plugin. Second, we present a persistent memory allocator Metall built atop UMap for unified storage/memory.

97 MATHEMATICS AND COMPUTING↗

Computing the Properties of Matter with Leadership Computing Resources (Closeout Report for DE-SC0018121)

In order to add more capabilities to Halide, we have designed a new framework called Tiramisu and integrated this framework into Halide. Since Tiramisu enables Halide to target heterogeneous architectures, our development efforts have been refocused on Tiramisu. Most high-performance computer systems today are complex and increasingly heterogeneous; they may have CPUs, GPUs and FPGAs. Achieving best performance requires taking full advantage of all these different architectures. To address this issue, we have designed Tiramisu, an optimization framework that enables Halide (and other DSLs) to target heterogeneous architectures. Tiramisu is an optimization framework that takes as input a high level, architecture-independent representation of code and a set of scheduling and data mapping commands that guide code transformation. The input can either be generated by a domain-specific language (DSL) compiler such as Halide or directly written by a programmer. Tiramisu then applies the user-specified code and data-layout transformations and generates an architecture-specific, low-level intermediate representation (IR) that takes advantage of modern architectural features such as multicore parallelism, non-uniform memory (NUMA) hierarchies, clusters, and accelerators like GPUs and FPGAs. We integrated Tiramisu within Halide and implemented a representative set of benchmarks to evaluate this integration. Tiramisu is now open source and is available for public use (http://tiramisu-compiler.org/). A paper about Tiramisu was published, it shows that Tiramisu extends Halide with many new capabilities and that Tiramisu can generate efficient code for multicores, GPUs, FPGAs and distributed heterogeneous systems. The performance of code generated by the Tiramisu backends matches or exceeds hand optimized reference implementations. For example, the multicore backend matches the highly optimized Intel MKL library on many kernels and shows speedups reaching 4x over the original Halide. In addition to making Tiramisu more robust, we have used Tiramisu to implement a set of representative tensor operation for constructing baryon building blocks required for multi baryon contractions in LQCD. In order to implement this code, we needed to generalize Tiramisu in two ways: first we needed to support indirect array accesses, and second, we needed to add support for complex numbers to Tiramisu. The code generated by Tiramisu is 6x faster than the reference code. Our efforts towards an MPI based multi-node version of tiramisu have matured and the resulting code scales well on multiple nodes (tests up to 512 KNL nodes have been undertaken).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Improving Signal-to-Noise Ratio (SNR) for Readout Signals Using Adaptive Filters on Reconfigurable Controls Hardware

This study investigates the optimization of Signal-to-Noise Ratio (SNR) in superconducting quantum computing readout signals through adaptive filtering. Quantum computing technology has the potential to revolutionize various fields by delivering exponential speedup in solving certain computational problems. However, the technology's practical implementation is hindered by the difficulty of extracting clean, reliable signals during the readout phase, with various sources of noise presenting a significant barrier to clean signals. This noise, often present in readout profiles due to imperfect isolation, degrades the system's overall SNR, thus impeding the ability to extract the quantum state accurately. The research leverages the power of adaptive filtering to improve the SNR of quantum computing readout signals. Specifically, an adaptive filter is implemented in a PYNQ overlay on an FPGA, and eventually will be connected to a quantum computing system. The system models the noise with a Least Mean Squares (LMS) adaptive filter, and then subtracts the estimated noise from the received signal to improve the SNR. A Direct Memory Access (DMA) channel is used to handle the signal processing, delivering efficient, high-speed data transfer between the PYNQ system and the hardware. The study explores the benefits of this adaptive filtering technique, potentially providing a significant contribution to practical and fast quantum computing.

Johnson, Hans↗

Improving Signal-to-Noise Ratio (SNR) for Readout Signals Using Adaptive Filters on Reconfigurable Controls Hardware

This study investigates the optimization of Signal-to-Noise Ratio (SNR) in superconducting quantum computing readout signals through adaptive filtering. Quantum computing technology has the potential to revolutionize various fields by delivering exponential speedup in solving certain computational problems. However, the technology's practical implementation is hindered by the difficulty of extracting clean, reliable signals during the readout phase, with various sources of noise presenting a significant barrier to clean signals. This noise, often present in readout profiles due to imperfect isolation, degrades the system's overall SNR, thus impeding the ability to extract the quantum state accurately. The research leverages the power of adaptive filtering to improve the SNR of quantum computing readout signals. Specifically, an adaptive filter is implemented in a PYNQ overlay on an FPGA, and eventually will be connected to a quantum computing system. The system models the n oise with a Least Mean Squares (LMS) adaptive filter, and then subtracts the estimated noise from the received signal to improve the SNR. A Direct Memory Access (DMA) channel is used to handle the signal processing, delivering efficient, high-speed data transfer between the PYNQ system and the hardware. The study explores the benefits of this adaptive filtering technique, potentially providing a significant contribution to practical and fast quantum computing.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Characterization of throughput on the AXI DMA bus for burst data transfer over Ethernet

cThe Xilinx AXI Direct Memory Access (AXI DMA) module is an efficient solution for medium-speed data transfer in Xilinx SoC FPGAs, supporting data rates greater than 1000 Gbps even in very suboptimal operating modes. It facilitates direct transfer of AXI stream data into processor memory without constant software intervention, which reduces overhead and ensures consistent data logging. By utilizing the FPGA's available memory, large circular buffers (1-5 GiB) are used to buffer data and accommodate network limitations, enabling high-rate data bursts. In this study, we measured the performance of AXI DMA under conditions simulating its lowest practical data transfer speeds. The Arbitrary Length Data Sender was used to transmit AXI stream packets at 32-bit width and 100 MHz frequency, a narrow width and slow speed. Results show that the AXI DMA can transfer up to 3192.76 Mbps with large packet sizes but experiences reduced performance for smaller packets, as low as 2.6 Mbps for 4-byte packets. For Ethernet-limited applications, packet sizes between 8,000 and 16,000 bytes provided optimal transfer speeds of 874 to 1600 Mbps. These findings suggest that the AXI DMA is not the limiting factor in systems where packet sizes exceed 8,000 bytes.

43 PARTICLE ACCELERATORS↗

Data Analysis Approach for Large Data Volumes in a Connected Community

Recent advancements within smart neighborhoods where utilities are enabling automatic control of appliances such as heating, ventilation, and air conditioning (HVAC) and water heater (WH) systems are providing new opportunities to minimize energy costs through reduced peak load. This requires systematic collection, storage, management, and in-memory processing of large volumes of streaming data for fast performance. In this paper, we propose a multi-tier layered IoT software framework that enables effective descriptive and predictive data analysis for understanding live operation of the neighborhood, fault identification, and future opportunities for further optimization of load curves. We then demonstrate how we achieve live situational awareness of the connected neighborhood through a suite of visualization components. Finally, we discuss a few analytic dashboards that address questions such as peak load reductions obtained due to optimization, customer preference for automatic control of appliances (do they override the automatic control of HVAC?, etc.). 1 1 This manuscript has been authored by UT-Battelle, LLC under Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy. The United States Government retains and the publisher, by accepting the article for publication, acknowledges that the United States Government retains a nonexclusive, paid-up, irrevocable, world-wide license to publish or reproduce the published form of this manuscript, or allow others to do so, for United States Government purposes. The Department of Energy will provide public access to these results of federally sponsored research in accordance with the DOE Public Access Plan (http://energy.gov/downloads/doe-public-access-plan).

Chinthavali, Supriya↗

Modernization efforts for the R -Matrix code SAMMY [Abstract]

The R-Matrix code SAMMY is a widely used nuclear data evaluation code focused on the resolved range, which includes corrections for experimental effects. The code is still mostly written in Fortran 77, and uses a memory management system suitable for the time of its initial writing (1984). A modernization effort is under way to bring the code in-line with modern software development practices. A continuous-integration testing framework was added, automating the large existing set of test cases. It is run on every commit. The memory management was updated to current standard practices suitable for modern software analysis tools. The code can be obtained from https://code.ornl.gov/RNSD/SAMMY. The resonance parameters and covariance information are now stored in C++ objects shared by SAMMY and AMPX, the processing code that generates nuclear data libraries for SCALE. This allows for easier maintenance and access to the resonance parameters inside and outside of SAMMY. This feature is already used by accessing and changing parameters in memory in the Bayesian Monte Carlo Evaluation Framework for Cross Sections Nuclear Data and Integral Benchmark Experiments project, Further plans include the switch to the ENDF reading and writing routines in AMPX, as these routines are more robust, easier to maintain, and support more features. Of note here is support for the new GNDS format. Previously it wasn’t easy to share the full covariance matrix for evaluations containing more than one isotope due to limitations on the ENDF format; this is now supported in GNDS. The data are currently available in a binary SAMMY format and can be exported to GNDS to make them more widely available and sharable. The next step will be to use the same resonance processing code at 0K in AMPX and SAMMY as one of the available Reich-Moore R-Matrix formalism. The first step toward this goal is to isolate the reconstruction into a module that takes resonance parameters as its input and does not depend on SAMMY global parameters. This goal has been achieved and it should now be possible to more easily change the resonance formalism and add enhancements as the Phenomenological R-Matrix parameterization of direct, doorway, and compound nuclear reactions discussed elsewhere on this conference. This concerted modernization and enhancement effort provides multiple advantages to the nuclear data community. It will allow parameter optimization using enhanced formalisms, including experimental effects, that better match complex experimental data. Then those evaluated parameters can immediately be passed off to AMPX to be reconstructed with the exact same cross section model and be put into a data library for subsequent testing using SCALE and the Valid Benchmark suite or other suitable benchmark suites.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Ion beam etching dependence of spin-orbit torque memory devices with switching current densities reduced by Hf interlayers

We report on the fabrication of nanoscale, three-terminal in-plane spin–orbit torque switching devices with low switching current densities. Critical parameters in the fabrication process, including the ion beam etching angle and time, were optimized to avoid fabrication defects and improve device yield. Measurements of the magnetic field and current-induced switching behavior of the tunnel junctions demonstrate a sensitivity to the nanopillar aspect ratio, which dictates the nanopillars’ anisotropy and thermal stability. Additionally, we show that the current density required for switching can be reduced and the device thermal stability increased by inserting Hf interlayers into the heterostructure. Micromagnetic simulations are generally consistent with the experimentally observed switching behavior, suggesting an increase in the interfacial perpendicular anisotropy at the CoFeB/MgO interface and the reduction in the Dzyaloshinskii–Moriya interaction at the W/CoFeB interface by the Hf interlayers.

36 MATERIALS SCIENCE↗

Unified architecture for quantum lookup tables

Quantum access to arbitrary classical data encoded in unitary black-box oracles underlies interesting data-intensive quantum algorithms, such as machine learning or electronic structure simulation. The feasibility of these applications depends crucially on gate-efficient implementations of these oracles, which are commonly some reversible versions of the Boolean circuit for a classical lookup table. Here, we present a general parametrized architecture for quantum circuits implementing a lookup table that encompasses all prior work in realizing a continuum of optimal trade-offs between qubits, non-Clifford gates, and error resilience, up to logarithmic factors. Our architecture assumes only local 2D connectivity, yet recovers results, with the appropriate parameters, polylogarithmic error scaling. We also identify regimes, such as simultaneous sublinear scaling, in all parameters. These results enable tailoring implementations of the commonly used lookup table primitive to any given quantum device with constrained resources.

quantum circuits↗

Subsurface Characterization and Machine Learning Predictions at Brady Hot Springs: Preprint

Subsurface data analysis, reservoir modeling, and machine learning (ML) techniques have been applied to the Brady Hot Springs (BHS) geothermal field in Nevada, USA to further characterize the subsurface and assist with optimizing reservoir management. Hundreds of reservoir simulations have been conducted in TETRAD-G and CMG STARS to explore different injection and production fluid flow rates and allocations and to develop a training data set for ML. This process included simulating the historical injection and production since 1979 and prediction of future performance through 2040. ML networks were created and trained using TensorFlow based on multilayer perceptron (MLP), long short-term memory (LSTM), and convolutional neural network (CNN) architectures. These networks took as input selected flow rates, injection temperatures, and historical field operation data and produced estimates of future production temperatures. This approach was first successfully tested on a simplified single fracture doublet system, followed by the application to the BHS reservoir. Using an initial BHS dataset with 37 simulated scenarios, the trained and validated network predicted the production temperature for 6 production wells with the mean absolute percentage error of less than 8%. In a complementary analysis effort, the principal component analysis applied to 13 BHS geological parameters revealed that vertical fracture permeability shows the strongest correlation with fault density and fault intersection density. A new BHS reservoir model was developed considering the fault intersection density as proxy for permeability. This new reservoir model helps to explore under-exploited zones in the reservoir. A data gathering plan to obtain additional subsurface data was developed; it includes temperature surveying for three idle injection wells, at which the reservoir simulations indicate high bottom-hole temperatures. The collected data assist with calibrating the reservoir model and may lead to converting these wells to producers to access under-exploited zones in the reservoir. Data gathering activities are planned for the first quarter of 2021.

40 EE - Geothermal Technologies Office (EE-4G)↗

HARMONY: Large-Scale Architecture Search for Efficient Hybrid Language Models

As large language models scale to trillions of parameters, their computational and memory requirements present critical challenges for efficient training and deployment. While Mixture of Experts (MoE) architectures enable efficient scaling through sparse parameter activation, and state-space models like Mamba offer linear-time complexity, principled methods for combining these paradigms remain undeveloped. We introduce HARMONY (Hybrid Architecture Research for Mamba, Optimized with Neural efficiencY), a multi-objective evolutionary neural architecture search framework for discovering efficient hybrid language models that integrate Transformer attention mechanisms, Mixture-of-Experts routing, and Mamba state-space components. Through large-scale distributed search using 16,384 MI250X GPUs on the Frontier supercomputer, HARMONY explores a comprehensive design space encompassing six attention variants (MHA, MQA, GQA, MLA, SWA, and Mamba-2), variable MoE configurations with both routed and shared experts, and extensive Mamba hyperparameters. Our framework discovers heterogeneous architectures that balance training performance with computational efficiency through multi-objective optimization incorporating latency penalties and fitness-based selection. Analysis of discovered architectures reveals that optimal hybrid designs favor heterogeneous component mixing rather than homogeneous patterns, with Mamba-2 and Multi-Head Latent Attention (MLA) emerging as preferred mechanisms. Discovered architectures demonstrate superior training efficiency: our best configuration achieves a final perplexity of 1.0874 with 2.38B parameters while processing 4,320 tokens/second, outperforming significantly larger manually designed models. Full-scale evaluation shows HARMONY's top architectures achieve better loss trajectories than equivalently-sized models using state-of-the-art configurations including Mixtral, Jamba, and Samba. Additionally, we demonstrate 91% weak scaling efficiency when training discovered 36B-parameter models across 1,024 GPUs. HARMONY is released as an open framework with comprehensive tools for building and training hybrid models using expert-data-pipeline parallelism, democratizing access to automated architecture design for next-generation language models.

Herron, Emily [ORNL] (ORCID:0000000273008172)↗

EXAGRAPH: Graph and combinatorial methods for enabling exascale applications

Combinatorial algorithms in general and graph algorithms in particular play a critical enabling role in numerous scientific applications. However, the irregular memory access nature of these algorithms makes them one of the hardest algorithmic kernels to implement on parallel systems. With tens of billions of hardware threads and deep memory hierarchies, the exascale computing systems in particular pose extreme challenges in scaling graph algorithms. The codesign center on combinatorial algorithms, ExaGraph, was established to design and develop methods and techniques for efficient implementation of key combinatorial (graph) algorithms chosen from a diverse set of exascale applications. Algebraic and combinatorial methods have a complementary role in the advancement of computational science and engineering, including playing an enabling role on each other. In this paper, we survey the algorithmic and software development activities performed under the auspices of ExaGraph from both a combinatorial and an algebraic perspective. In particular, we detail our recent efforts in porting the algorithms to manycore accelerator (GPU) architectures. We also provide a brief survey of the applications that have benefited from the scalable implementations of different combinatorial algorithms to enable scientific discovery at scale. We believe that several applications will benefit from the algorithmic and software tools developed by the ExaGraph team.

97 MATHEMATICS AND COMPUTING↗