Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “accelerate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Pseudodiagonalization Method for Accelerating Nonlinear Subspace Diagonalization in Density Functional Theory

In density functional theory, each self-consistent field (SCF) nonlinear step updates the discretized Kohn-Sham orbitals by solving a linear eigenvalue problem. The concept of pseudodiagonalization is to solve this linear eigenvalue problem approximately, and specifically utilizing a method involving a small number of Jacobi rotations that takes advantage of the good initial guess to the solution given by the approximation to the orbitals from the previous SCF iteration. The approximate solution to the linear eigenvalue problem can be very rapid, particularly for those steps near SCF convergence. Here, we adapt pseudodiagonalization to finite-temperature and metallic systems, where partially-occupied orbitals must be individually resolved with some accuracy. We apply pseudodiagonalization to the subspace eigenvalue problem that arises in Chebyshev-filtered subspace iteration. In tests on metallic and other systems for a range of temperatures, we show that pseudodiagonalization achieves similar rates of SCF convergence to exact diagonalization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accessing and Photo-Accelerating Low-Overpotential Pathways for CO 2 Reduction: A Bis-Carbene Ruthenium Terpyridine Catalyst

A ruthenium catalyst bearing a bidentate bis(carbene) ligand is prepared and studied as a catalyst for CO 2 electroreduction. The catalyst [Ru(tpy)(bis-mim)(MeCN)][PF 6 ] 2 (tpy) is 2,2′,:6′,2″-terpyridine; bis-mim is (methylenebis(N-methylimidazol-2-ylidene)) mediates reduction of CO 2 into CO with a turnover frequency of 630 s –1 and Faradaic efficiency (FE) of 30% at an overpotential of 730 mV. The strongly donating bis(carbene) ligand also enables access to a pathway operating at a lower overpotential of ca. 310 mV. While low-overpotential catalysis is slow in the dark (TOF = 0.01 s –1 ), visible light illumination increases the rate 10-fold (TOF = 0.11 s –1 ). Here, a full mechanistic picture is developed using kinetic analysis from cyclic voltammetry, spectroelectrochemistry, and computational methods, with the bis-mim ligand facilitating rapid CO 2 activation at low overpotentials. Comparisons with other ruthenium catalysts yield insight into the ability to tune the rate of chemical steps (e.g., ligand dissociation and CO 2 nucleophilic attack) and the overpotential by tailoring the primary coordination sphere while retaining the “redox-active” tpy ligand.

CO2 reduction↗

FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters

In this paper, we propose a framework called FPDeep, which uses a hybrid of model and layer paral- lelism to configure distributed reconfigurable clusters to train DNNs. This approach has numerous benefits. First, the design does not suffer from batch size growth. Second, novel workload and weight partitioning leads to balanced loads of both among nodes. And third, the entire system is fine-grained pipeline. This leads to high parallelism and utilization and also minimizes the time features need to be cached while waiting for back-propagation.

Wang, Tianqi↗

Scalable Incremental Checkpointing using GPU-Accelerated De-Duplication

Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs' high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.

Tan, Nigel↗