Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scalable performance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A higher-order finite-element implementation of the nonlinear Fokker–Planck collision operator for charged particle collisions in a low density plasma

Collisions between particles in a low density plasma are described by the Fokker–Planck collision operator. In applications, this nonlinear integro-differential operator is often approximated by linearised or ad-hoc model operators due to computational cost and complexity. In this work, we present an implementation of the nonlinear Fokker–Planck collision operator written in terms of Rosenbluth potentials in the Rosenbluth–MacDonald–Judd (RMJ) form. The Rosenbluth potentials may be obtained either by direct integration or by solving partial differential equations (PDEs) similar to Poisson's equation: we optimise for performance and scalability by using sparse matrices to solve the relevant PDEs. We represent the distribution function using a tensor-product continuous-Galerkin finite-element representation and we derive and describe the implementation of the weak form of the collision operator. We present tests demonstrating a successful implementation using an explicit time integrator and we comment on the speed and accuracy of the operator. Finally, we speculate on the potential for applications in the current and next generation of kinetic plasma models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Structural Phase Transitions between Layered Indium Selenide for Integrated Photonic Memory

The primary mechanism of optical memoristive devices relies on phase transitions between amorphous and crystalline states. The slow or energy-hungry amorphous–crystalline transitions in optical phase-change materials are detrimental to the scalability and performance of devices. Leveraging an integrated photonic platform, nonvolatile and reversible switching between two layered structures of indium selenide (In 2 Se 3 ) triggered by a single nanosecond pulse is demonstrated. The high-resolution pair distribution function reveals the detailed atomistic transition pathways between the layered structures. With interlayer “shear glide” and isosymmetric phase transition, switching between the α- and β-structural states contains low re-configurational entropy, allowing reversible switching between layered structures. Broadband refractive index contrast, optical transparency, and volumetric effect in the crystalline–crystalline phase transition are experimentally characterized in molecular-beam-epitaxy-grown thin films and compared to ab initio calculations. Finally, the nonlinear resonator transmission spectra measure of incremental linear loss rate of 3.3 GHz, introduced by a 1.5 µm-long In 2 Se 3 -covered layer, resulted from the combinations of material absorption and scattering.

36 MATERIALS SCIENCE↗

Recent Advances in Scalable, High‐Mass Loaded Electrodes for Grid‐Scale Energy Storage

Abstract The increasing electrification of daily life as well as the intermittent characteristic of renewable energy sources require viable solutions for grid‐scale energy storage. Critical considerations for grid storage applications are electrode mass loading and electrode thickness as these features govern battery pack energy density, an important factor in determining manufacturing costs. For this reason, there is increased interest in finding new ways of creating electrodes with high mass loading. In this review, various high‐mass loading fabrication approaches are considered for positive electrode materials used in batteries. The benchmark used for high mass loading is above 20 mg cm −2 , which is higher than the practical limit of conventional tape‐cast electrodes. Several different electrode approaches are described including templating, laser patterning, direct ink writing, and electrodeposition. A variety of materials are covered with the most prominent being LiFe(PO 4 ) (LFP), LiCoO 2 (LCO), and MnO 2 . In research to date, scalable electrochemical performance has been achieved with mass loadings over 100 mg cm −2 . Areal capacities as high as 14.7 mAh cm −2 at 1.82 mA cm −2 have been achieved in non‐aqueous electrolytes and 9.8 mAh cm −2 at 10 mA cm −2 in aqueous electrolytes. These results establish that the mass loading of electrodes can be scaled up without compromising their electrochemical properties.

White, Makena [Department of Materials Science and↗

A Self‐Consistent Model for Sorption and Transport in Polyimide‐Derived Carbon Molecular Sieve Gas Separation Membranes

Abstract Demand for energy‐efficient gas separations exists across many industrial processes, and membranes can aid in meeting this demand. Carbon molecular sieve (CMS) membranes show exceptional separation performance and scalable processing attributes attractive for important, similar‐sized gas pairs. Herein, we outline a mathematical and physical framework to understand these attributes. This framework shares features with dual‐mode transport theory for glassy polymers; however, physical connections to CMS model parameters differ from glassy polymer cases. We present evidence in CMS membranes for a large volume fraction of microporous domains characterized by Langmuir sorption in local equilibrium with a minority continuous phase described by Henry's law sorption. Using this framework, expressions are provided to relate measurable parameters for sorption and transport in CMS materials. We also outline a mechanism for formation of these environments and suggest future model refinements.

Sanyal, Oishi↗

Multifrontal Non-negative Matrix Factorization

Non-negative matrix factorization (Nmf) is an important tool in high-performance large scale data analytics with applications ranging from community detection, recommender system, feature detection and linear and non-linear unmixing. While traditional Nmf works well when the data set is relatively dense, however, it may not extract sufficient structure when the data is extremely sparse. Specifically, traditional Nmf fails to exploit the structured sparsity of the large and sparse data sets resulting in dense factors. We propose a new algorithm for performing Nmf on sparse data that we call multifrontal Nmf (Mf-Nmf) since it borrows several ideas from the multifrontal method for unconstrained factorization (e.g. LU and QR). We also present an efficient shared memory parallel implementation of Mf-Nmf and discuss its performance and scalability. We conduct several experiments on synthetic and realworld datasets and demonstrate the usefulness of the algorithm by comparing it against standard baselines. We obtain a speedup of 1.2x to 19.5x on 24 cores with an average speed up of 10.3x across all the real world datasets.

Sao, Piyush↗

Balanced k -means clustering on an adiabatic quantum computer

Adiabatic quantum computers are a promising platform for efficiently solving challenging optimization problems. Therefore, many are interested in using these computers to train computationally expensive machine learning models. We present a quantum approach to solving the balanced k-means clustering training problem on the D-Wave 2000Q adiabatic quantum computer. In order to do this, we formulate the training problem as a quadratic unconstrained binary optimization (QUBO) problem. Unlike existing classical algorithms, our QUBO formulation targets the global solution to the balanced k-means model. We test our approach on a number of small problems and observe that despite the theoretical benefits of the QUBO formulation, the clustering solution obtained by a modern quantum computer is usually inferior to the solution obtained by the best classical clustering algorithms. Nevertheless, the solutions provided by the quantum computer do exhibit some promising characteristics. We also perform a scalability study to estimate the run time of our approach on large problems using future quantum hardware. Finally, as a final proof of concept, we used the quantum approach to cluster random subsets of the Iris benchmark data set.

97 MATHEMATICS AND COMPUTING↗

Scalable balanced training of conditional generative adversarial neural networks on image data

Here, we propose a distributed approach to train deep convolutional generative adversarial neural network (DC-CGANs) models. Our method reduces the imbalance between generator and discriminator by partitioning the training data according to data labels, and enhances scalability by performing a parallel training where multiple generators are concurrently trained, each one of them focusing on a single data label. Performance is assessed in terms of inception score, Fréchet inception distance, and image quality on MNIST, CIFAR10, CIFAR100, and ImageNet1k datasets, showing a significant improvement in comparison to state-of-the-art techniques to training DC-CGANs. Weak scaling is attained on all the four datasets using up to 1000 processes and 2000 NVIDIA V100 GPUs on the OLCF supercomputer Summit.

97 MATHEMATICS AND COMPUTING↗

Local convergence analysis of an inexact trust-region method for nonsmooth optimization

In Baraldi, we introduced an inexact trust-region algorithm for minimizing the sum of a smooth nonconvex function and a nonsmooth convex function in Hilbert space—a class of problems that is ubiquitous in data science, learning, optimal control, and inverse problems. Furthermore, this algorithm has demonstrated excellent performance and scalability with problem size. In this paper, we enrich the convergence analysis for this algorithm, proving strong convergence of the iterates with guaranteed rates. In particular, we demonstrate that the trust-region algorithm recovers superlinear, even quadratic, convergence rates when using a second-order Taylor approximation of the smooth objective function term.

97 MATHEMATICS AND COMPUTING↗

dCache: The Storage System of Choice for Data-Intensive Applications

The ever-increasing volumes of data produced by modern scientific facilities like EuXFEL and LHC put significant stress on data management infrastructure operated by laboratories and research centers. The challenges to be addressed span the entire data life cycle, from ingest and efficient data analysis to long-term preservation, typically involving large tape libraries. dCache, a storage system developed in collaboration between the Deutsches Elektronen-Synchrotron (DESY), Fermi National Accelerator Laboratory, and Nordic e-Infrastructure Collaboration (NeIC), is designed to manage a large number of disk servers and to facilitate transparent data migration to and from archival storage. Its multifaceted approach offers a unified method to support a variety of scientific use cases with the same storage infrastructure, including high-throughput data ingest, data sharing over wide area networks, efficient access from HPC clusters, and long-term data preservation on tertiary storage. Initially developed for high energy physics (HEP) experiments, dCache is now used by various scientific communities, including astrophysics, biomedical research, and life sciences, each having specific requirements. This paper presents architecture, deployment strategies, performance and scalability enhancements, and recent advancements in dCache addressing the needs of scientific communities. Finally, we touch on the development and release process, ensuring the software’s high quality.

DCache↗

Adiabatic quantum support vector machines

Adiabatic quantum computers can solve difficult optimization problems (e.g., the quadratic unconstrained binary optimization problem), and they seem well suited to train machine learning models. In this paper, we describe an adiabatic quantum approach for training support vector machines. We show that the time complexity of our quantum approach is an order of magnitude better than the classical approach. Next, we compare the test accuracy of our quantum approach against a classical approach that uses the Scikit-learn library in Python across five benchmark datasets (Iris, Wisconsin Breast Cancer (WBC), Wine, Digits, and Lambeq). We show that our quantum approach obtains accuracies on par with the classical approach. Finally, we perform a scalability study in which we compute the total training times of the quantum approach and the classical approach with an increasing number of features and an increasing number of data points in the training dataset. In conclusion, our scalability results show that the quantum approach obtains a 3.5–4.5x speedup over the classical approach on datasets with many (millions of) features.

Computational Complexity↗

Blockchain Smart Contract Reference Framework and Program Logic Architecture for Transactive Energy Systems

This paper proposes a reference framework for a transactive energy market based on blockchain. The framework was designed based on the engineering requirements of a distribution-scale market; including participant needs, expected market transactions, and the cybersecurity constructs required to support a fair, secure and efficient market operation. It leverages the existing blockchain primitives to provide clear value propositions to the transactive market, including identity management (access control), data security (integrity), resiliency (decentralization, scalability and performance). The validity of the proposed framework is demonstrated using a real-time 5-min double-auction market. The results highlight its benefits while providing strong validation of applicability to blockchain within transactive energy systems.

Gourisetti, Sri Nikhil Gupta↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Distributed approximate minimal Steiner trees with millions of seed vertices on billion-edge graphs

In this report, we present a parallel 2-approximation Steiner minimal tree algorithm and its MPI-based distributed implementation. In place of expensive distance computations between all pairs of seed vertices, the solution we employ exploits a cheaper Voronoi cell computation. Our design leverages asynchronous processing and message prioritization to accelerate convergence of distance computations, and harnesses vertex and edge centric processing to offer fast time-to-solution. We demonstrate scalability and performance using real-world graphs with up to 128 billion edges and 512 compute nodes, and show the ability to find Steiner trees with up to one million seed vertices. Using 12 data instances, we present comparison with the state-of-the-art exact solver, SCIP-Jack, and two sequential 2-approximate algorithms. We empirically show that, on average, the total distance of the Steiner tree identified by our solution is 1.1290 times greater than the Steiner minimal tree – well within the theoretical approximation bound of 2.

97 MATHEMATICS AND COMPUTING↗

Recent advancements in high performance polymer electrolyte fuel cell electrode fabrication – Novel materials and manufacturing processes

The global effort to introduce polymer electrolyte fuel cells for clean and renewable energy to the market is increasing the demand for high performance, robust and affordable membrane electrode assemblies (MEAs). There is not yet a standard method for large scale production of MEAs, or the methods employed are generally unsatisfactory in terms of quality and performance. A large number of published data of newly developed catalyst and electrolyte materials, claim to improve the state of the art, but are often not fully comparable due to different experimental studies and experimental designs. This article summarizes the trends in material developments and emerging MEA-manufacturing techniques. The materials and techniques are systematically compared in terms of cell performance and scalability. Current and future scientific challenges are identified and analysed based on published findings over the past five years. Finally, the results of the cited papers have been quantitatively compared to each other and to the internal benchmarks used in each cited work to provide a complete picture of the state of the art in PEFC MEA manufacturing.

25 ENERGY STORAGE↗

Molten Lithium-Brass/Zinc Chloride System as High-Performance and Low-Cost Battery

Batteries with high safety, low cost, and reasonable energy density are essential for grid-scale energy storage and still remain elusive. In this paper, we report a solid electrolyte-based liquid lithium-brass/zinc chloride (SELL-brass/ZnCl 2 ) battery using garnet-type lithium-ion solid electrolyte, lithium anode, and brass/ZnCl 2 cathode. The chemistry of the cell reaction and the ability of being assembled in discharged state ensures a high safety. The use of low-cost ZnCl 2 cathode can realize a low cell material cost of $16 kWh –1 . The adoption of lithium anode guarantees a high theoretical energy density of 750 Wh kg –1 and 2,250 Wh L –1 . Moreover, by using brass powder as a Zn source in the cathode, the Zn particle growth issue is successfully solved, and a good cycling stability of the battery can be obtained. As full cell performance and scalability are also verified, our SELL-brass/ZnCl 2 battery shows a high potential for practical use in grid energy storage.

25 ENERGY STORAGE↗

AstroPix: A pixelated HVCMOS sensor for space-based gamma-ray measurement

A next-generation medium-energy gamma-ray telescope targeting the MeV range would address open questions in astrophysics regarding how extreme conditions accelerate cosmic-ray particles, produce relativistic jet outflows, and more. One concept, AMEGO-X, relies upon the mission-enabling CMOS Monolithic Active Pixel Sensor silicon chip AstroPix. AstroPix is designed for space-based use, featuring low noise, low power consumption, and high scalability. Desired performance of the device include an energy resolution of 5 keV (or 10% FWHM) at 122 keV and a dynamic range per-pixel of 25–700 keV, enabled by the addition of a high-voltage bias to each pixel which supports a depletion depth of 500 μ m. This work reports on the status of the AstroPix development process with emphasis on the current version under test, version three (v3), and highlights of version two (v2). Version 3 achieves energy resolution of 10.4 ± 3.2% at 59.5 keV and 94 ± 6 μ m depletion in a low-resistivity test silicon substrate.

Astrophysics instrumentation↗