Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Distributed memory, GPU accelerated Fock construction for hybrid, Gaussian basis density functional theory

With the growing reliance of modern supercomputers on accelerator-based architecture such a graphics processing units (GPUs), the development and optimization of electronic structure methods to exploit these massively parallel resources has become a recent priority. While significant strides have been made in the development GPU accelerated, distributed memory algorithms for many modern electronic structure methods, the primary focus of GPU development for Gaussian basis atomic orbital methods has been for shared memory systems with only a handful of examples pursing massive parallelism. Here in this work, we present a set of distributed memory algorithms for the evaluation of the Coulomb and exact exchange matrices for hybrid Kohn–Sham DFT with Gaussian basis sets via direct density-fitted (DF-J-Engine) and seminumerical (sn-K) methods, respectively. The absolute performance and strong scalability of the developed methods are demonstrated on systems ranging from a few hundred to over one thousand atoms using up to 128 NVIDIA A100 GPUs on the Perlmutter supercomputer.

97 MATHEMATICS AND COMPUTING↗

A high-resolution large-eddy simulation framework for wildland fire predictions using TensorFlow

Background: Wildfires are becoming more severe, so we need improved tools to predict them over a wide range of conditions and scales. One approach towards this goal entails the use of coupled fire/atmosphere modelling tools. Although significant progress has been made in advancing their physical fidelity, existing tools have not taken full advantage of emerging programming paradigms and computing architectures to enable high-resolution wildfire simulations. Aims: The aim of this study was to present a new framework that enables landscape-scale wildfire simulations with physical representation of combustion at an affordable cost. Methods: We developed a coupled fire/atmosphere simulation framework using TensorFlow, which enables efficient and scalable computations on Tensor Processing Units. Key Results: Simulation results for a prescribed fire were compared with experimental data. Predicted fire behavior and statistical analysis for fire spread rate, scar area, and intermittency showed overall reasonable agreement. Scalability analysis was performed, showing close to linear scaling. Conclusions: While mesh refinement was shown to have less impact on global quantities, such as fire scar area and spread rate, it benefits predictions of intermittent fire behavior, buoyancy-driven dynamics, and small-scale turbulent motion. Implications: This new simulation framework is efficient in capturing both global quantities and unsteady dynamics of wildfires at high spatial resolutions.

54 ENVIRONMENTAL SCIENCES↗

Continental-Scale Controls on Hyporheic Respiration Revealed by Knowledge-Guided Machine Learning

Hyporheic zone sediments regulate organic matter turnover and in-stream respiration, yet controls on sediment respiration remain poorly constrained across heterogeneous river networks, limiting prediction of stream metabolism and carbon processing at continental scales. Here, we integrate observations from ~90 river corridors across the United States in the WHONDRS consortium with a knowledge-guided machine learning (KGML) framework that couples thermodynamic rate theory with machine learning to identify dominant controls on hyporheic respiration. Diagnostic analyses show that organic matter concentration and thermodynamic favorability define an upper bound on respiration potential, whereas biological catalytic capacity and physical accessibility jointly govern realized respiration rates through interaction effects. To represent unmeasurable accessibility constraints, we use the mechanistic model as a scaffold for KGML, allowing machine learning to target residual structure not explained by process theory. This hybrid framework improves predictive skill relative to both the mechanistic model alone and fully data-driven models while preserving interpretability. These results indicate that variability in hyporheic respiration is largely mechanistically structured and demonstrate how integrating process theory with explainable AI enhances predictive performance while enabling scalable synthesis of river corridor observations.

Zheng, Jianqiu↗

funcX: Federated Function as a Service for Science

Here, funcX is a distributed function as a service (FaaS) platform that enables flexible, scalable, and high performance remote function execution. Unlike centralized FaaS systems, funcX decouples the cloud-hosted management functionality from the edge-hosted execution functionality. funcX's endpoint software can be deployed, by users or administrators, on arbitrary laptops, clouds, clusters, and supercomputers, in effect turning them into function serving systems. funcX's cloud-hosted service provides a single location for registering, sharing, and managing both functions and endpoints. It allows for transparent, secure, and reliable function execution across the federated ecosystem of endpoints-enabling users to route functions to endpoints based on specific needs. funcX uses containers (e.g., Docker, Singularity, and Shifter) to provide common execution environments across endpoints. funcX implements various container management strategies to execute functions with high performance and efficiency on diverse funcX endpoints. funcX also integrates with an in-memory data store and Globus for managing data that may span endpoints. We motivate the need for funcX, present our prototype design and implementation, and demonstrate, via experiments on two supercomputers, that funcX can scale to more than 130000 concurrent workers. We show that funcX's container warming-aware routing algorithm can reduce the completion time for 3,000 functions by up to 61% compared to a randomized algorithm and the in-memory data store can speed up data transfers by up to 3x compared to a shared file system.

97 MATHEMATICS AND COMPUTING↗

Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees

We present differentiable predictive control (DPC), a method for offline learning of constrained neural control policies for nonlinear dynamical systems with performance guarantees. We show that the sensitivities of the parametric optimal control problem can be used to obtain direct policy gradients. Specifically, we employ automatic differentiation (AD) to efficiently compute the sensitivities of the model predictive control (MPC) objective function and constraints penalties. To guarantee safety upon deployment, we derive probabilistic guarantees on closed-loop stability and constraint satisfaction based on indicator functions and Hoeffding’s inequality. We empirically demonstrate that the proposed method can learn neural control policies for various parametric optimal control tasks. In particular, we show that the proposed DPC method can stabilize systems with unstable dynamics, track time-varying references, and satisfy nonlinear state and input constraints. Our DPC method has practical time savings compared to alternative approaches for fast and memory-efficient controller design. Specifically, DPC does not depend on a supervisory controller as opposed to approximate MPC based on imitation learning. We demonstrate that, without losing performance, DPC is scalable with greatly reduced demands on memory and computation compared to implicit and explicit MPC while being more sample efficient than model-free reinforcement learning (RL) algorithms.

97 MATHEMATICS AND COMPUTING↗

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING↗

PLANC: Parallel Low-rank Approximation with Nonnegativity Constraints

In this work, we consider the problem of low-rank approximation of massive dense nonnegative tensor data, for example, to discover latent patterns in video and imaging applications. As the size of data sets grows, single workstations are hitting bottlenecks in both computation time and available memory. We propose a distributed-memory parallel computing solution to handle massive data sets, loading the input data across the memories of multiple nodes, and performing efficient and scalable parallel algorithms to compute the low-rank approximation. We present a software package called Parallel Low-rank Approximation with Nonnegativity Constraints, which implements our solution and allows for extension in terms of data (dense or sparse, matrices or tensors of any order), algorithm (e.g., from multiplicative updating techniques to alternating direction method of multipliers), and architecture (we exploit GPUs to accelerate the computation in this work). We describe our parallel distributions and algorithms, which are careful to avoid unnecessary communication and computation, show how to extend the software to include new algorithms and/or constraints, and report efficiency and scalability results for both synthetic and real-world data sets.

97 MATHEMATICS AND COMPUTING↗

A GPU-Accelerated Population Generation, Sorting, and Mutation Kernel for an Optimization-Based Causal Inference Model

We develop a GPU-accelerated machine learning generative adversarial network model that can be used with observational data for the purpose of constructing causal inferences. The theoretical basis of our machine learning model is novel and is conceptualized to be operable and scalable for high performance computing platforms. Our GPU-accelerated code enables large-scale parallelization of the computation within a common and accessible computing environment. This will expand the reach of our model and empower research in new substantive domains while maintaining the underlying theoretical properties.

Cho, Wendy K. Tam↗

From Edge to HPC: Investigating Cross-Facility Data Streaming Architectures

In this paper, we investigate three cross-facility data streaming architectures, Direct Streaming (DTS), Proxied Streaming (PRS), and Managed Service Streaming (MSS). We examine their architectural variations in data flow paths and deployment feasibility, and detail their implementation using the Data Streaming to HPC (DS2HPC) architectural framework and the SciStream memory-to-memory streaming toolkit on the production-grade Advanced Computing Ecosystem (ACE) infrastructure at Oak Ridge Leadership Computing Facility (OLCF). We present a workflow-specific evaluation of these architectures using three synthetic workloads derived from the streaming characteristics of scientific workflows. Through simulated experiments, we measure streaming throughput, round-trip time, and overhead under work sharing, work sharing with feedback, and broadcast and gather messaging patterns commonly found in AI-HPC communication motifs. Our study shows that DTS offers a minimal-hop path, resulting in higher throughput and lower latency, whereas MSS provides greater deployment feasibility and scalability across multiple users but incurs significant overhead. PRS lies in between, offering a scalable architecture whose performance matches DTS in most cases.

George, Anjus [ORNL] (ORCID:0000000179737061)↗

Development and Validation of a Two-Phase Thermal-Hydraulic CFD Code NEK-2P

A project is underway to develop, verify and validate an advanced two-phase flow modeling capability for the highly-scalable, high-performance Computational Fluid Dynamics (CFD) code NEK5000. The goal of this work is to verify and validate the two-phase version of the NEK5000 code, named NEK-2P, to simulate the two-phase flow and heat transfer phenomena that occur in a Boiling Water Reactor (BWR) fuel bundle under various operating conditions. The NEK-2P two-phase flow models follow the approach used for the Extended Boiling Framework (EBF) previously developed at Argonne but include more fundamental physical models of boiling phenomena and advanced numerical algorithms for improved computational accuracy, robustness, and computational speed. The development of the NEK-2P two-phase solver and the implementation of the Extended Boiling Framework two-phase models were initially supported by Argonne National Laboratory (Argonne) through a Laboratory Directed Research and Development (LDRD) project during FY2014-2016. The development and validation of the two-phase models through analyses of selected two-phase boiling flow experiments was supported by the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program in FY2017-2020. This report focuses on verification and validation of the water-steam boiling model NEK-2P Two-Phase, CFD code. The NEK-2P was validated with Nuclear Power Engineering Corporation (NUPEC) Pressurized Water Reactor (PWR) Sub-channel and Bundle Test (PSBT) void distribution benchmark. Three different simulations were performed and analyzed for various operating conditions such as wall-heat flux and sub-cooled inlet temperatures. Reasonably good agreement with measured data was obtained in predicting the measured void distributions. Simulations were performed for Virginia Tech. (VT) 3x3 rod bundle geometry with and without spacers. The preliminary results were presented for Simplified Spacer Grid (SSG). In addition, the implementation of interface reconstruction model was tested with one of the Becker benchmark Critical Heat Flux (CHF) experiments.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

The Portals 4.3 Network Programming Interface

This report presents a specification for the Portals 4 network programming interface. Portals 4 is intended to allow scalable, high-performance network communication between nodes of a parallel computing system. Portals 4 is well suited to massively parallel processing and embedded systems. Portals 4 represents an adaption of the data movement layer developed for massively parallel processing platforms, such as the 4500-node Intel TeraFLOPS machine. Sandia's Cplant cluster project motivated the development of Version 3.0, which was later extended to Version 3.3 as part of the Cray Red Storm machine and XT line. Version 4 is targeted to the next generation of machines employing advanced network interface architectures that support enhanced offload capabilities.

97 MATHEMATICS AND COMPUTING↗

CRADA Number NFE-19-07845 with American Nanotechnologies, Inc. (CRADA Final Report)

Cooperative Research and Development Agreement (CRADA) NFE-19-07845 between Oak Ridge National Laboratory (ORNL) and American Nanotechnologies, Inc. (ANI) focused on studying the use of dielectrophoresis as a mechanism to purify nanoparticles, particularly semiconducting carbon nanotubes (CNTs). The successful development of low-cost purification processes is a significant bottleneck in the adoption of semiconducting CNTs in commercial semiconducting devices. ANI is a startup company developing scalable systems for performing nanoparticle dielectrophoresis, which has previously been used only in microfluidic devices. The work done under this CRADA is exploring the use of ANI’s patent resonant dielectrophoresis (rDEP) technology and its ability to purify semiconducting CNTs at large scale. While continued work is on-going, r-DEP has proven to be a viable way to control nanoparticles in bulk dispersions. Optimization of the process is now underway to reach minimum viable product and begin material sales. Additionally, ANI and ORNL continue to collaborate on leveraging this technology to reach down the value chain and create new commercial devices.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Advanced Alkaline Membrane H 2 /Air Fuel Cell System with Novel Technique for Air CO 2 Removal

Over the course of this project, significant progress was achieved in developing the hydroxide exchange membrane fuel cell (HEMFC) and the electrochemically-driven CO₂ separator (EDCS), with a focus on improving performance, durability, and scalability. Key milestones were met, and the technology demonstrated potential for a wide range of applications, including fuel cell vehicles, direct air capture (DAC), and life support systems.

08 HYDROGEN↗

Control-Agnostic Beam Instrumentation with Redis at the Core

Redis isn’t a database — it’s our protocol. Fermilab’s RedisAdapter provides a high-performance, control-system-agnostic bridge between digitized beam data and downstream consumers such as ACNET and EPICS. It forms the foundation of three new software components deployed across MicroTCA-based digitizers: GMMDM, a runtime for memory-mapped data movement from Zynq-based platforms; GRAFE, a front end for Redis-to-ACNET presentation; and GREFE, an EPICS IOC front end. Together, these tools enable modular, standardized instrumentation pipelines. Precision timing is handled via White Rabbit PPS distribution, allowing nanosecond-scale synchronization across crates. This architecture, originally prototyped in Booster BPM systems, is now deployed on modern hardware and designed to meet the performance, modularity, and scalability requirements of the PIP-II era.

Steinkamp, Derek [Fermilab] (ORCID:000900027228626↗

Implementation of Manifold-Based Combustion Models in a Highly Scalable Low Mach Number Reacting Flow Solver: Preprint

Manifold-based representations of the thermochemistry are often employed in conjunction with large eddy eimulation (LES) to lower the cost of combustion simulations. This work describes steps taken to implement this modeling approach in PeleLM, a scalable and performance-portable low Mach number flow solver. Most significantly, this includes adapting the projection method used by PeleLM to satisfy the mass conservation constraint for use with manifold-based models. The implementation is designed to be general across manifold-based models, including both those that employ traditional tabulation and those that employ neural networks. An initial demonstration for simple test cases is presented and will be used for performance assessment.

high-performance computing↗

LivChat.....So Far

This presentation provides an overview of LivChat, a managed ChatGPT service developed by Lawrence Livermore National Laboratory (LLNL) under the auspices of the U.S. Department of Energy. The initiative was driven by a high demand for generative AI services, the need for enhanced security, and the goal of increasing productivity across various use cases. The development journey involved exploring open-source models and iterating with OpenAI/Azure solutions. The presentation delves into the architecture of LivChat, highlighting its user interface (UI) and backend components. The backend is designed to be separate, RESTful, integrative, and scalable, ensuring robust performance and adaptability. Despite the advanced technology, the presentation emphasizes that LivChat is not a magical solution but a sophisticated tool that requires realistic expectations. Looking ahead, the presentation outlines future directions for LivChat, including training, retrieval-augmented generation (RAG), and innovative ingestion methods. These advancements aim to further enhance the capabilities and applications of LivChat, ensuring it remains at the forefront of generative AI services.

Computer science↗

Evaluate data lake design for the accelerator control system

Increasing precision in automation for modern particle accelerators not only creates a requirement to gather data from all devices but also demands scalable and high-performance data infrastructure with the capability of handling vast incoming device data. A well architected data lake is suitable for such a system which integrates real-time data acquisition, transient data caching, and long-term storage. This paper evaluates data lake architecture for an Accelerator Control System (ACS), focusing on two critical components of a data lake, data cache and long-term storage.

Jaikar, Amol [Fermilab]↗

Facile and scalable dry surface doping technique to enhance the electrochemical performance of LiNi 0.64 Mn 0.2 Co 0.16 O 2 cathode materials

Lithium nickel manganese cobalt oxide (NMC) is one of the dominant cathode materials in lithium-ion batteries. In this study a simple, efficient and scalable surface doping technique is successfully demonstrated, which can be readily used in mass production of cathode materials. For the first time neodymium oxide (Nd 2 O 3 ) has been employed as the surface doping agent. The Nd-doped NMC shows greatly improved cycling and rate performance, and the enhanced cycling stability has been demonstrated in full pouch cells, with a 17.5% increase in capacity retention after 300 cycles. Fewer cracks have been observed in the doped NMC after cycling, and in situ X-ray diffraction reveals the suppressed lattice collapse by Nd doping. Greatly suppressed surface phase change has been confirmed by HR-TEM and EELS. The result suggests great promise in using this dry doping technique to enhance the electrochemical performance of NMC cathodes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗